Fatigue detection method based on driver face recognition

Through the combination of multi-task cascaded convolutional network and lightweight neural network, the driver's global and local facial features are extracted and feature fusion is performed, which solves the problem of insufficient fatigue detection accuracy in the existing technology and achieves high-precision fatigue warning.

CN120472434AActive Publication Date: 2025-08-12GUANGXI UNIVERSITY OF TECHNOLOGY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510581673.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-12
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing fatigue detection technology is difficult to achieve accurate fatigue detection under interference from external environment, resulting in poor detection accuracy and inability to effectively warn the driver of dangerous states.

Method used

A multi-task cascade convolutional network is used to recognize driver facials, extract global and local facial features, and combine lightweight neural network analysis to integrate feature fusion through the global feature extraction module and the local flow fusion module to achieve high-precision fatigue detection.

Benefits of technology

It improves the accuracy and robustness of fatigue detection, can effectively identify the driver's fatigue status in complex environments, promptly warn, and reduce the risk of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472434A_ABST
    Figure CN120472434A_ABST
Patent Text Reader

Abstract

The invention discloses a fatigue detection method based on driver face recognition, and the method comprises the steps: inputting a driving image into a preset face extraction neural network model for recognition, and obtaining a global face image and a local eye region image; inputting the global face image into a lightweight neural network to obtain a global feature map, and inputting the local eye region image into the lightweight neural network to obtain a double local feature map; inputting the global feature map into a global feature extraction model for feature extraction to obtain global feature information, and performing fusion analysis on a first eye feature map and a second eye feature map in the double local feature maps to obtain local feature information; and performing multi-fusion processing on the global feature information and the local feature information, and classifying the eye state of the driver according to a processing result to determine the fatigue degree of the driver and give an alarm, thereby ensuring traffic safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fatigue driving detection, and in particular to a fatigue detection method based on driver facial recognition. Background Art

[0002] Driver fatigue in commercial vehicles is a key factor in major traffic accidents. This is particularly true in long-distance freight and passenger transport scenarios, where accidents are often caused by driver fatigue. To address this safety hazard, intelligent fatigue detection technology has become a core component of commercial vehicle active safety systems.

[0003] Existing fatigue detection technologies include vehicle behavior analysis (such as analyzing steering wheel rotation, braking frequency, and vehicle speed). This method is easily interfered with by external environmental factors and is difficult to achieve accurate fatigue detection, resulting in poor detection accuracy and an inability to effectively warn drivers of dangerous fatigue states and avoid traffic accidents.

[0004] It can be seen that how to effectively detect fatigue driving and maintain traffic safety has become a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention

[0005] The present invention provides a fatigue detection method based on driver facial recognition, which solves the problem of how to extract and enhance features of the driver's face from global and local perspectives to improve the accuracy of the model output results and achieve timely warning of traffic accidents.

[0006] In order to solve the above technical problems, an embodiment of the present invention provides a method for detecting driver fatigue based on facial recognition, comprising:

[0007] Inputting the acquired driving image of the driver inside the target vehicle into a preset facial extraction neural network model for recognition to obtain a global facial image and a local eye area image of the driver;

[0008] Inputting the global facial image into a pre-trained first lightweight neural network to output a global feature map, and inputting the local eye area image into a pre-trained second lightweight neural network to output a dual local feature map;

[0009] Inputting the global feature map into a global feature extraction model for feature extraction to obtain global feature information corresponding to the global facial image, and performing fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye region image;

[0010] performing multiple fusion processing on the global feature information and the local feature information, and classifying the driver's eye state according to the processing results;

[0011] A fatigue detection result corresponding to the driver's fatigue level is obtained according to the classification result.

[0012] Furthermore, the process of acquiring the driver's global facial image and local eye area image includes:

[0013] Inputting the driving image into the facial extraction neural network model, and outputting the coordinates of the boundary vertices and key points of the driver's entire face; the facial extraction neural network model is obtained by pre-training a multi-task cascade convolutional network;

[0014] The driving image is cropped according to the boundary vertex coordinates and the key point coordinates to obtain the global facial image and the local eye region image.

[0015] Furthermore, the step of inputting the global feature map into a global feature extraction model to perform feature extraction to obtain global feature information corresponding to the global facial image includes:

[0016] In the global feature extraction model, performing a dimensionality reduction process of global average pooling on the global feature map to obtain a first feature tensor;

[0017] Inputting the first feature tensor into the fully connected layer configured in the global feature extraction model for channel adjustment to obtain a second feature tensor;

[0018] Performing nonlinear optimization on the second feature tensor using an activation function to obtain a third feature tensor;

[0019] The third feature tensor is integrated with the global feature map to determine the global feature information.

[0020] Furthermore, the performing fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye area image includes:

[0021] After inputting the dual local feature maps into a selected local fusion model, the first eye feature map and the second eye feature map are fused element by element to generate a local fused feature map;

[0022] The local fusion feature map is sequentially subjected to convolution and nonlinear activation processing, and an attention weighted modulation mechanism is introduced during the processing to generate a local enhanced feature map;

[0023] The local enhanced feature map is integrated with the pre-selected second eye feature map and then feature splitting is performed to determine the local feature information.

[0024] Furthermore, the performing multiple fusion processing on the global feature information and the local feature information, and classifying the driver's eye state according to the processing results, includes:

[0025] The global feature information and the local feature information are sequentially subjected to weight analysis, channel enhancement, and weighted fusion processing to obtain target facial feature data;

[0026] The target facial feature data is input into a preset classifier model to classify the closed and open states of the driver's eyes.

[0027] Furthermore, the step of sequentially performing weight analysis, channel enhancement, and weighted fusion processing on the global feature information and the local feature information to obtain target facial features includes:

[0028] Performing vector splicing on the global feature information and the local feature information, and assigning weights to the spliced global feature information and the local feature information respectively;

[0029] Performing first channel enhancement and spatial alignment processing on the weighted global feature information in sequence to obtain global optimization features, and performing second channel enhancement on the weighted local feature information to determine local optimization features;

[0030] The global optimization feature and the local optimization feature are weightedly fused to obtain the target facial feature.

[0031] Furthermore, obtaining a fatigue detection result corresponding to the driver's fatigue level according to the classification result includes:

[0032] Determining the driver's eye state distribution probability based on the classification result;

[0033] Calculating a target number of frames of eye closure for the driver within a preset time period based on the eye state distribution probability;

[0034] Determining a quantitative result corresponding to the fatigue degree by calculating the percentage of the target frame number to the total frame number in a preset time period;

[0035] When the quantification result exceeds a preset fatigue threshold, a warning is issued, and the fatigue detection result is generated according to the warning content and the quantification result.

[0036] Another embodiment of the present invention provides a driver fatigue detection system based on facial recognition, comprising:

[0037] A facial feature detection module is used to input the acquired driving image of the driver inside the target vehicle into a preset facial extraction neural network model for recognition, thereby obtaining a global facial image and a local eye area image of the driver;

[0038] a feature extraction module, configured to input the global facial image into a pre-trained first lightweight neural network to output a global feature map, and input the local eye region image into a pre-trained second lightweight neural network to output a dual local feature map;

[0039] a feature optimization module, configured to input the global feature map into a global feature extraction model for feature extraction to obtain global feature information corresponding to the global facial image, and to perform a fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye region image;

[0040] a classification module, configured to perform multiple fusion processing on the global feature information and the local feature information, and classify the driver's eye state according to the processing results;

[0041] The fatigue detection module is used to obtain a fatigue detection result corresponding to the driver's fatigue level according to the classification result.

[0042] Another embodiment of the present invention provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the fatigue detection method based on driver facial recognition as described above is implemented.

[0043] Yet another embodiment of the present invention provides a computer-readable storage medium storing a computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the fatigue detection method based on driver facial recognition as described above is implemented.

[0044] Compared with the prior art, the embodiments of the present invention have the following advantages:

[0045] The embodiment of the present invention uses a task-cascaded convolutional network for face detection, and simultaneously extracts global facial features and local eye features. Therefore, even if part of the facial information is lost, effective judgment can still be made through local features, thereby improving the robustness of the model; local facial features such as the left and right eye areas and global facial features are analyzed separately by lightweight neural networks, greatly improving detection efficiency; the flow fusion module and the global feature extraction module are combined to fuse and analyze local features and global features respectively, which can retain facial information of different scales to achieve high-precision fatigue driving detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flowchart of a method for detecting driver fatigue based on facial recognition in one embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the overall process of driver facial recognition in one embodiment of the present invention;

[0048] Figure 3 is a diagram of a global feature analysis architecture in one embodiment of the present invention;

[0049] Figure 4 is a diagram of a local feature fusion architecture in one embodiment of the present invention;

[0050] Figure 5 Schematic diagram of the structure of a fatigue detection system based on driver facial recognition in one embodiment of the present invention;

[0051] Figure 6 A structural block diagram of a preferred embodiment of a computer device provided by the present invention;

[0052] Reference numerals: M1, facial feature detection module; M2, feature extraction module; M3, feature optimization module; M4, classification module; M5, fatigue detection module. DETAILED DESCRIPTION

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0054] In the description of this application, the terms "first," "second," "third," etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first," "second," "third," etc. may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0055] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the two components. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are for illustrative purposes only, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0056] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this application have the same meanings as those commonly understood by those skilled in the art. The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. Those skilled in the art will understand the specific meanings of the above terms in this application in specific circumstances.

[0057] In order to solve the problem that the existing methods are not accurate in detecting the driver's driving status, an embodiment of the present invention provides a fatigue detection method based on driver facial recognition. For details, please refer to Figure 1 , Figure 1 The figure shows a flowchart of a method for detecting driver fatigue based on facial recognition in one embodiment of the present invention, which includes the following steps:

[0058] S1. Inputting the acquired driving image of the driver inside the target vehicle into a preset facial extraction neural network model for recognition, and obtaining the driver's global facial image and local eye area image.

[0059] In an embodiment of the present invention, a single-frame image in a video stream captured by an RGB camera inside the target vehicle can be selected as the input of the face extraction neural network model. It can be understood that the multi-task cascaded convolutional network (MTCNN) has significantly improved the robustness, accuracy and real-time performance of driver fatigue detection in actual image recognition applications through mechanisms such as cascade structure, multi-task learning, and key point positioning. It performs particularly well under occlusion and complex lighting conditions, providing an efficient solution for fatigue detection tasks. Therefore, preferably, this embodiment selects a multi-task cascaded convolutional network to pre-train the face extraction neural network model.

[0060] After training, the driving image is input into the facial extraction neural network model, and the boundary vertex coordinates and key point coordinates of the driver's entire face are output. It can be understood that in this embodiment, the MTCNN model will output a rectangular box of the facial boundary (including the four vertex coordinates of the boundary box) and 5 key point coordinates. The 5 coordinate values include the coordinates of the center point of the left eye, the center point of the right eye, the nose, the left corner of the mouth, and the right corner of the mouth. For details, please refer to Figure 2 As shown on the far left, Figure 2 The figure shows the overall process of driver face recognition in one embodiment of the present invention. In this embodiment, the driving image is cropped according to the boundary vertex coordinates and key point coordinates, thereby obtaining a global face image and a local eye area image. Figure 2 It can be seen from the figure that the local eye area image includes the driver's left eye and right eye.

[0061] For example, in some embodiments of the present invention, the vertex distribution of the facial rectangle output by MTCNN is set as follows: upper left corner P1 = (x1 + y1), upper right corner P2 = (x2 + y2), lower right corner P3 = (x3 + y3), lower left corner P4 = (x4 + y4). The width of the box is expressed as: W face , height is expressed as: H face For the eye area, in some embodiments of the present invention, E L =(x L ,y L ) to represent the center point of the left eye, using E R =(x R ,y R ) represents the center of the right eye. Each region follows a set of proportional relationships based on the eyelid size. Based on the coordinates of the four vertices of the facial rectangle, a global facial image reflecting the driver's overall facial features is cropped, with a size of 64*64*3. Correspondingly, the left and right eye regions are cropped based on the coordinates of the left and right eye centers, respectively, with a size of 32*32*3.

[0062] Specifically, the overall face is represented as:

[0063] Face=(x1,y1,W face ,H face )

[0064] The left eye area can be represented as:

[0065]

[0066] The right eye area can be represented as:

[0067]

[0068] Wherein, α and β are the ratios of the width and height of the eye to the size of the face. In this example, α is set to 0.25 and β is set to 0.17.

[0069] S2. Input the global facial image into a pre-trained first lightweight neural network, and output a global feature map; input the local eye area image into a pre-trained second lightweight neural network, and output a dual local feature map.

[0070] Combine Figure 2 It can be seen that this embodiment adopts a dual-stream architecture based on a lightweight neural network (MobileNetV3) for the feature extraction stage of global facial images and local eye area images. The lightweight network is used as the backbone network and is suitable for real-time application scenarios such as ADAS (Advanced Driver Assistance System). The trained lightweight model is directly deployed on the vehicle side to achieve offline recognition, which can better protect the user's privacy. MobileNetV3-Large and MobileNetV3-Small in this embodiment are two different configurations of the MobileNetV3 model, which can be optimized for different application scenarios.

[0071] The first lightweight neural network applied to global facial images is MobileNetV3-Large. Its large receptive field effectively captures semantic features widely distributed across facial regions, such as the spatial relationships between areas like the eyes and mouth, and the overall facial contour. This provides semantic priors for subsequent extraction of fine-grained local features. In this example, the global facial image is fed into a pre-trained MobileNetV3-Large network as the global stream for feature analysis, ultimately outputting a 4x4x704 global feature map.

[0072] The second lightweight neural network applied to the local eye area image is MobileNetV3-Small, which is more capable of capturing fine-grained appearance and texture details from the local eye image, including the contour changes of the eyelid edge, the light and dark distribution of the pupil and iris areas, the degree of opening of the palpebral fissure, etc. In this embodiment, the left eye area and the right eye area of the local eye image are input into the pre-trained MobileNetV3-Small as local stream 1 and local stream 2 respectively for feature extraction, and the corresponding output is a dual local feature map reflecting the features of the left eye area and the right eye area of the driver. Exemplarily, local feature maps of size 4*4*704 can be output respectively. It should be understood that this embodiment adopts a multi-stream parallel processing strategy, which can significantly improve the detection performance.

[0073] S3. Input the global feature map into a global feature extraction model to perform feature extraction to obtain global feature information corresponding to the global facial image, and perform fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye area image.

[0074] This step is a further feature analysis process. In the face of the three feature maps output by the lightweight network, in this embodiment, they are respectively input into different neural network modules for feature processing.

[0075] For the global feature map corresponding to the global facial image, please refer to Figure 3 , Figure 3 The figure shows a global feature analysis architecture diagram in one embodiment of the present invention. It can be seen that, in this embodiment, preferably, the global feature map is input into the global feature extraction model, namely the global feature extraction module (GFEM). Integrating this module into the neural network architecture can significantly suppress redundant noise and enhance the weight of key features. In this embodiment, the global feature extraction module can be divided into two different stages: the "squeeze" stage, which aggregates global spatial information through global average pooling to reduce sensitivity to local noise; and the "excitation" stage, which adjusts the importance of each channel through a fully connected layer and uses a Sigmoid activation function for optimization adjustment.

[0076] Specifically, during the "squeeze" phase, the global feature map undergoes dimensionality reduction using global average pooling to produce the first feature tensor. As can be appreciated, global average pooling compresses the spatial dimensions of each channel, accurately capturing the channel's global statistical characteristics. This first feature tensor is then fed into the two fully connected layers configured in the global feature extraction model for channel conditioning, yielding the second feature tensor.

[0077] In the "excitation" stage, the activation function is used to perform nonlinear optimization on the second eigenvalue tensor to obtain the third eigenvalue tensor.

[0078] For example, for the input global feature map X∈R H×W×C , H, W, C are the height, width and number of channels of the feature map respectively. After global average pooling, X is reduced to a 1×1×C tensor, i.e. the first feature tensor X pool .

[0079] Next, the pooled feature map X pool Through two consecutive fully connected layers, the output tensor is generated That is, the second characteristic tensor Z, expressed as:

[0080] Z=((X pool ·W1+b1)·W2+b2)

[0081] Where W1 is the weight of the first fully connected layer, b1 is the bias term of the first fully connected layer; W2 is the weight of the second fully connected layer, b2 is the bias term of the second fully connected layer.

[0082] Then, the Sigmoid activation function σ is applied to Z, which normalizes the weight of each channel to between 0 and 1, and obtains the tensor S∈R 1×1×C :

[0083] S=σ(Z)

[0084] The tensor S is further resized to a shape of 1×1×C and then broadcasted to increase the dimension, obtaining the third feature tensor S broadcast It should be understood that the dimensionality increase process replicates and expands S along the spatial dimensions (height H and width W) to generate a weight matrix with the same dimensions as the input feature map X, thereby achieving channel calibration.

[0085] After the above two stages, the third feature tensor is integrated with the global feature map to determine the global feature information. For example, the original feature map X is combined with S broadcast Perform element-by-element penalty to generate the output feature map, that is, the global feature information Output∈R H×W×C :

[0086] Output=X·S broadcast

[0087] For the dual local feature maps corresponding to the local eye area image, please refer to Figure 4 , Figure 4 The figure shows a local feature fusion architecture diagram in one embodiment of the present invention. It can be seen that in this embodiment, the first eye feature map and the second eye feature map in the dual local feature maps are preferably input into the selected local fusion model, namely the local fusion module (SFM). This module can effectively merge the outputs from the eye feature stream. In this embodiment, the local fusion module integrates skip connections, which can alleviate the gradient vanishing problem of deep networks. It performs element-by-element fusion on the first eye feature map and the second eye feature map to generate a local fusion feature map.

[0088] Specifically, this embodiment performs convolution and nonlinear activation processing on the local fusion feature map in sequence, introduces the attention weighted modulation mechanism in the processing process, and generates a local enhanced feature map. In this embodiment, the input dual-stream feature map is subjected to channel grouping convolution, such as Figure 4 As shown in the figure, the eyes are divided into two groups and 1×1 convolution is performed on each group. Batch normalization and GuLU function activation are then performed.

[0089] Exemplarily, 1×1 convolution is performed on the first eye feature map as input stream 1 and the second eye feature map as input stream 2, respectively, to generate two independent feature streams X1 and X1, both of size 4×4×30.

[0090] Then, X1 and X1 are summed element by element to generate the feature map X sum , expressed as:

[0091] X sum =X1+X2

[0092] The fused feature map X sum After activation by the GeLU function, the feature map X is generated GELU , the GeLU function can enhance the nonlinear relationship between features by calculating the normal distribution and capture more complex feature relationships in the feature map.

[0093] Then X GeLU Apply 1×1 convolution again to generate feature map X conv , and calculate the feature map X through the Sigmoid function conv Attention weights can guide the network to focus on relevant areas and suppress irrelevant information while maintaining computational efficiency. conv The local enhanced feature map X generated after Sigmoid processing sigmoid It is expressed as follows:

[0094] X sigmoid =σ(X conv ),X sigmoid ∈R 4×4×2C

[0095] Finally, the local enhanced feature map is integrated with the pre-selected second eye feature map and then feature splitting is performed to determine the local feature information. That is, the feature map obtained by performing a 1×1 convolution on the original feature map of input stream 2 is combined with the local enhanced feature map X sigmoid By multiplying the elements one by one, a 4×4×2C feature map is obtained. The feature map is then split with the middle channel as the breakpoint to generate two 4×4×C local feature maps, which include the local feature information of the left eye feature map and the right eye feature map.

[0096] S4. Perform multiple fusion processing on the global feature information and the local feature information, and classify the driver's eye state according to the processing results.

[0097] In this embodiment, the designed multi-fusion processing is a three-stream fusion architecture, that is, two local information streams and one global feature information stream are fused. The three-stream fusion architecture sequentially performs weight analysis, channel enhancement and weighted fusion processing on the global feature information and the local feature information. Specifically:

[0098] First, the global feature information and the local feature information are vector-joined. In this embodiment, after global average pooling is performed on each input stream, vector concatenation is performed. The joint feature X formed after concatenation is expressed as:

[0099] X=[x1,x2,x3]

[0100] Where x1, x2, and x3 are the global feature information flow, local information flow 1, and local information flow 2 after pooling, respectively.

[0101] The concatenated global feature information and local feature information are weighted through the linear layer, and the probability distribution of the classification is output through the Softmax layer.

[0102] Secondly, channel enhancement and spatial alignment are performed on the weighted global and local feature information. Specifically, the weighted global feature information is sequentially subjected to a first channel enhancement and spatial alignment process to promote semantic alignment between different streams (such as the association between global facial features and local iris features) to obtain globally optimized features. Simultaneously, a second channel enhancement process is performed on the weighted local feature information to determine the locally optimized features. Preferably, lightweight convolution can be used for channel enhancement.

[0103] Finally, a weighted fusion of the global and local optimized features is performed to obtain the target facial features, which are then fed into a pre-set classifier model to classify the driver's eye status as closed or open. In some embodiments of the present invention, the classifier model is a binary classification probability distribution algorithm, such as logistic regression, likelihood estimation, or a Bernoulli distribution model. The final output is a binary classification result: the probability distribution of eyes open / closed in the current single-frame image.

[0104] In some embodiments of the present invention, the classification result is output through a linear layer during the classification process.

[0105] S5. Obtain a fatigue detection result corresponding to the driver's fatigue level according to the classification result.

[0106] The driver's eye state (open or closed) is determined based on the distribution probability of the driver's eye state output by the softmax layer. For each frame of the collected video stream, the eye state distribution probability will be calculated, and the target number of frames with the driver's eyes closed will be counted within the preset time period. Then, by calculating the percentage of the target number of frames in the total number of frames in the preset time period, the quantitative result corresponding to the degree of fatigue is determined. For example, 300 seconds are taken to calculate the percentage of the number of frames with closed eyes in the total number of frames in 300 seconds, that is, the PERCLOS value. PERCLOS is a bio-indicator algorithm for evaluating human visual attention and fatigue, which can accurately evaluate the driver's fatigue level. A higher percentage of eye closure means that the driver's eyes have been closed for a long time, indicating that the driver may be tired or sleepy. The calculation formula is as follows:

[0107]

[0108] Among them, n close Indicates the number of frames in which the eyes are closed, N total Indicates the total number of frames captured during the observation period. The relationship between the PERCLOS indicator and fatigue status can be seen in the following table:

[0109] PERCLOS value Fatigue state 0-0.05 Normal wakefulness, eyes closed normally, no signs of fatigue 0.05-0.2 Mild fatigue, possibly starting to feel slightly sleepy 0.2-0.4 Moderate fatigue, obvious drowsiness, and decreased concentration 0.4-0.7 Severe fatigue, obvious tiredness, easy distraction, slow reaction speed 0.7 and above Severe fatigue, extreme sleepiness, and the possibility of falling asleep at any time, posing a great safety risk

[0110] PERCLOS value and fatigue status relationship table

[0111] When the calculated PERCLOS value, a quantitative measure of fatigue, exceeds a preset fatigue threshold (typically 0.7 or 70%), the driver is considered fatigued and a warning is issued. Based on the warning and the quantitative result, fatigue detection results are generated, such as a detection log or a report covering the entire detection process.

[0112] For example, if the driver is judged to be moderately fatigued, a voice announcement reminder is triggered, such as reminding the driver to concentrate. If the driver is judged to be severely fatigued, a seat vibration + emergency braking request is triggered.

[0113] In summary, the embodiments of the present invention provide a driver fatigue detection method based on multi-stream facial feature analysis. By combining a multi-task cascaded convolutional network and a lightweight neural network, global and local features are extracted from the collected driving images of the driver to improve the anti-occlusion capability. By combining the global feature extraction module and the local stream fusion module for feature fusion, facial features of different scales can be more fully captured, ensuring the complementarity of local and global information, and improving recognition accuracy by fusing local features and global features.

[0114] An embodiment of the present invention provides a fatigue detection system based on driver facial recognition. For details, see Figure 5 , Figure 5 FIG. 1 is a schematic diagram of a system for detecting fatigue based on driver facial recognition in one embodiment of the present invention, comprising:

[0115] The facial feature detection module M1 is used to input the acquired driving image of the driver inside the target vehicle into a preset facial extraction neural network model for recognition, and obtain the driver's global facial image and local eye area image;

[0116] A feature extraction module M2 is configured to input the global facial image into a pre-trained first lightweight neural network to output a global feature map, and input the local eye area image into a pre-trained second lightweight neural network to output a dual local feature map;

[0117] A feature optimization module M3 is configured to input the global feature map into a global feature extraction model for feature extraction to obtain global feature information corresponding to the global facial image, and to perform a fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye region image;

[0118] a classification module M4, configured to perform multiple fusion processing on the global feature information and the local feature information, and classify the driver's eye state according to the processing results;

[0119] The fatigue detection module M5 is used to obtain a fatigue detection result corresponding to the driver's fatigue level according to the classification result.

[0120] like Figure 6 As shown, an embodiment of the present invention further provides a computer device, Figure 6 This is a structural block diagram of a preferred embodiment of a computer device provided by the present invention, wherein the computer device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the above-mentioned method when executing the computer program.

[0121] Preferably, the computer program can be divided into one or more modules / units (e.g., computer program 1, computer program 2, ...), which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units can be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device.

[0122] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can be any conventional processor. The processor is the control center of the terminal device, and uses various interfaces and lines to connect the various parts of the terminal device.

[0123] The memory mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system, an application program required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card, etc., or the memory can also be other volatile solid-state storage devices.

[0124] It should be noted that the above terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 6 The structural block diagram is only an example of a terminal device and does not constitute a limitation of the terminal device. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above-mentioned methods. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0125] Accordingly, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to perform the steps in the method of the above embodiment, for example Figure 1 Steps S1 to S5 described in .

[0126] The technical features and technical effects of the fatigue detection system based on driver facial recognition proposed in an embodiment of the present invention are the same as the technical features and technical effects of the fatigue detection method based on driver facial recognition proposed in an embodiment of the present invention, and will not be repeated here.

[0127] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A method for detecting driver fatigue based on facial recognition, characterized in that: include: Inputting the acquired driving image of the driver inside the target vehicle into a preset facial extraction neural network model for recognition to obtain a global facial image and a local eye area image of the driver; Inputting the global facial image into a pre-trained first lightweight neural network to output a global feature map, and inputting the local eye area image into a pre-trained second lightweight neural network to output a dual local feature map; Inputting the global feature map into a global feature extraction model for feature extraction to obtain global feature information corresponding to the global facial image, and performing fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye region image; performing multiple fusion processing on the global feature information and the local feature information, and classifying the driver's eye state according to the processing results; A fatigue detection result corresponding to the driver's fatigue level is obtained according to the classification result.

2. The driver fatigue detection method based on facial recognition according to claim 1, characterized in that: The process of acquiring the driver's global facial image and local eye area image includes: Inputting the driving image into the facial extraction neural network model, and outputting the coordinates of the boundary vertices and key points of the driver's entire face; the facial extraction neural network model is obtained by pre-training a multi-task cascade convolutional network; The driving image is cropped according to the boundary vertex coordinates and the key point coordinates to obtain the global facial image and the local eye region image.

3. The driver fatigue detection method based on facial recognition according to claim 1, characterized in that: Inputting the global feature map into a global feature extraction model to perform feature extraction to obtain global feature information corresponding to the global facial image includes: In the global feature extraction model, performing a dimensionality reduction process of global average pooling on the global feature map to obtain a first feature tensor; Inputting the first feature tensor into the fully connected layer configured in the global feature extraction model for channel adjustment to obtain a second feature tensor; Performing nonlinear optimization on the second feature tensor using an activation function to obtain a third feature tensor; The third feature tensor is integrated with the global feature map to determine the global feature information.

4. The driver fatigue detection method based on facial recognition according to claim 1, characterized in that: The fusing and analyzing the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye region image includes: After inputting the dual local feature maps into a selected local fusion model, the first eye feature map and the second eye feature map are fused element by element to generate a local fused feature map; The local fusion feature map is sequentially subjected to convolution and nonlinear activation processing, and an attention weighted modulation mechanism is introduced during the processing to generate a local enhanced feature map; The local enhanced feature map is integrated with the pre-selected second eye feature map and then feature splitting is performed to determine the local feature information.

5. The driver fatigue detection method based on facial recognition according to claim 1, characterized in that: The performing multiple fusion processing on the global feature information and the local feature information, and classifying the driver's eye state according to the processing results, includes: The global feature information and the local feature information are sequentially subjected to weight analysis, channel enhancement, and weighted fusion processing to obtain target facial feature data; The target facial feature data is input into a preset classifier model to classify the closed and open states of the driver's eyes.

6. The driver fatigue detection method based on facial recognition according to claim 5, characterized in that: The step of sequentially performing weight analysis, channel enhancement, and weighted fusion processing on the global feature information and the local feature information to obtain target facial features includes: Performing vector splicing on the global feature information and the local feature information, and assigning weights to the spliced global feature information and the local feature information respectively; Performing first channel enhancement and spatial alignment processing on the weighted global feature information in sequence to obtain global optimization features, and performing second channel enhancement on the weighted local feature information to determine local optimization features; The global optimization feature and the local optimization feature are weightedly fused to obtain the target facial feature.

7. The driver fatigue detection method based on facial recognition according to claim 1, characterized in that: Obtaining a fatigue detection result corresponding to the driver's fatigue level according to the classification result includes: Determining the driver's eye state distribution probability based on the classification result; Calculating a target number of frames of eye closure for the driver within a preset time period based on the eye state distribution probability; Determining a quantitative result corresponding to the fatigue degree by calculating the percentage of the target frame number to the total frame number in a preset time period; When the quantification result exceeds a preset fatigue threshold, a warning is issued, and the fatigue detection result is generated according to the warning content and the quantification result.

8. A driver fatigue detection system based on facial recognition, characterized in that: include: A facial feature detection module is used to input the acquired driving image of the driver inside the target vehicle into a preset facial extraction neural network model for recognition, thereby obtaining a global facial image and a local eye area image of the driver; a feature extraction module, configured to input the global facial image into a pre-trained first lightweight neural network to output a global feature map, and input the local eye region image into a pre-trained second lightweight neural network to output a dual local feature map; a feature optimization module, configured to input the global feature map into a global feature extraction model for feature extraction to obtain global feature information corresponding to the global facial image, and to perform a fusion analysis on the first eye feature map and the second eye feature map in the dual local feature maps to obtain local feature information corresponding to the local eye region image; a classification module, configured to perform multiple fusion processing on the global feature information and the local feature information, and classify the driver's eye state according to the processing results; The fatigue detection module is used to obtain a fatigue detection result corresponding to the driver's fatigue level according to the classification result.

9. A computer device, characterized in that: The system comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for detecting fatigue based on driver facial recognition according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the fatigue detection method based on driver facial recognition as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A driver fatigue detection system and a fatigue detection method thereof

    CN109740477A

  • Double-flow fusion network iris living body detection method based on light field image sequence

    CN111914646A

  • Facial expression recognition method and device, facial expression model training method and device, equipment and storage medium

    CN116343287A

  • Driver fatigue detection method fusing local and global face features

    CN116630942A

  • Driving fatigue detection method and system based on lightweight neural network image enhancement

    CN117789181A