Fatigue driving detection method and fatigue driving detection system based on multi-feature fusion
The multi-feature fusion method using CNNs and autoencoders addresses redundancy and noise in fatigue detection systems, enhancing accuracy and adaptability by dynamically adjusting weights and incorporating attention mechanisms.
Patent Information
- Application Number
- CN202510462863.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
The existing fatigue driving detection systems have problems such as feature redundancy and noise interference, information loss and suboptimal decision-making, especially in complex scenarios, the reliability of single-modal detection has significantly decreased.
The multi-feature fusion method is adopted to extract facial features and autoencoder through the CNN model, combine dynamic feature fusion and cross-modal attention mechanism, dynamically adjust modal weights, realize the timing correlation between facial movements and steering wheel operations, reduce vehicle data dimensions and remove redundancy, and use a two-way LSTM network to classify fatigue states.
It improves the robustness and accuracy of fatigue driving detection, can dynamically adapt to environmental changes in complex scenarios, reduce noise interference and redundant characteristics in multimodal data, and enhances the interpretability and discrimination of the system.
Smart Images

Figure CN120318802A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of road traffic, and in particular to a fatigue driving detection method and a fatigue driving detection system that integrate multiple features. Background Art
[0002] Traditional anti-fatigue measures (such as mandatory rest regulations and driver self-management) have limited implementation effects due to the lack of real-time monitoring and objective evaluation criteria. With the development of artificial intelligence, biosensing, and computer vision technologies, fatigue driving detection systems (such as multi-modal sensor technologies based on facial expression recognition, heart rate monitoring, and steering wheel operation analysis) already have the ability to give real-time warnings.
[0003] Current fatigue driving detection technologies mainly use single-sided features based on the face, vehicle data, physiological signals, etc. as the judgment basis for fatigue detection and warning. However, there are still technical bottlenecks - the limitations of single-modal detection: existing solutions are mostly based on a single data source, such as facial images or steering wheel operations, and their reliability significantly decreases in complex scenarios. In recent years, however, the dual-modal fusion method has gradually become a research hotspot, which solves the defects of traditional methods that ignore data dynamics and context dependence. Typical solutions include combinations of "vision + vehicle data" and "vision + physiology".
[0004] Regarding the method of "face + vehicle data", it can be roughly divided into the following two types.
[0005] One is to use CNN and time-domain analysis to extract facial data features and vehicle data features respectively, and then directly splice them into a joint feature vector. Finally, the joint feature vector is input into a fully connected network (such as MLP) for binary classification (fatigue / normal). The limitations of this method are: feature redundancy and noise interference. If the facial data is blocked (such as by sunglasses) or the vehicle data is interfered by complex road conditions (such as continuous curves), the noise features will contaminate the global features through splicing, resulting in model confusion. Ignoring cross-modal correlation: simple splicing cannot model the temporal causal relationship between facial actions (such as frequent blinking) and small steering wheel corrections, losing key fatigue dynamic patterns.
[0006] The other is independent model training, which uses respectively: a facial model - training a classifier based on CNN to output the fatigue probability; a vehicle model - training a classifier based on a time-series model (such as HMM) to output the fatigue probability. Finally, the fatigue probability results of the two are weighted and averaged, and then judged according to logical rules. The limitations of this method are: information loss and suboptimal decision-making. During independent training, the facial and vehicle models do not share intermediate features and cannot learn cross-modal complementarity (such as the joint pattern of closing eyes and steering wheel stagnation). Static weights cannot adapt to dynamic scenarios: fixed weights fail when the day-night switches or road conditions change. Summary of the Invention
[0007] In order to solve the technical defects of feature redundancy and noise interference, information loss and sub-optimal decision-making existing in the existing fatigue driving detection system with multi-feature fusion, the present invention provides a fatigue driving detection method with multi-feature fusion and its fatigue driving detection system.
[0008] The present invention is implemented by the following technical solutions: A fatigue driving detection method with multi-feature fusion, which includes the following steps:
[0009] Extract facial features from the face image of the driver;
[0010] Extract vehicle features from the vehicle driving parameters of the vehicle driven by the driver;
[0011] Fuse the facial features and the vehicle features to determine whether the driver is fatigued while driving;
[0012] Among them, the facial features are extracted through a specific CNN model: 1) Basic convolutional feature extraction; 2) Local convolutional feature extraction; 3) Global convolutional feature extraction; 4) Pooling operation; 5) Aggregate global features; 6) Dimensionality reduction mapping;
[0013] The vehicle features are also extracted through a specific autoencoder: Adopt a symmetric deep neural network structure, and compress high-dimensional time-series data to a low-dimensional latent space through non-linear mapping.
[0014] As a further improvement of the above solution, the autoencoder includes:
[0015] 1) Encoder: Input → 256 → 128 → 64;
[0016] Input layer: Input, 256 nodes;
[0017] Hidden layer 1: Fully connected layer, 128 nodes, ReLU activation;
[0018] Hidden layer 2: Fully connected layer, 64 nodes, ReLU activation;
[0019] Latent layer: Fully connected layer, 64 nodes, linear activation, and add L1 sparse regularization;
[0020] 2) Decoder: Output 64 → 128 → 256;
[0021] Hidden layer 3: Fully connected layer, 64 nodes, ReLU activation;
[0022] Hidden layer 4: Fully connected layer, 128 nodes, ReLU activation;
[0023] Output layer: Fully connected layer, 256 nodes, Sigmoid activation to match the input normalization range.
[0024] Preferably, the encoder input vector is non-linearly dimensionally reduced through a three-layer fully-connected network:
[0025] h1 = ReLU(w1x + b1)
[0026] h2 = ReLU(w2h1 + b2)
[0027] h3 = w3h2 + b3
[0028] h1 is the input layer, h2 is the intermediate layer, and h3 is the output layer; w1, w2, and w3 are the weight matrices of each layer, used to map the input or the output of the previous layer to the current layer; b1, b2, and b3 are the bias terms of each layer, used to adjust the offset of the output; x is the input data; ReLU: activation function, used to introduce non-linear characteristics;
[0029] The decoder reconstructs the input through a reverse structure:
[0030]
[0031] W6, W5, W4: weight matrices; b4, b5, and b6 are bias terms;
[0032] The training objective function combines the reconstruction error and the sparse constraint:
[0033]
[0034] Γ: the value of the loss function, representing the total loss of the model;
[0035] N: the number of samples, representing the total number of input data;
[0036] x i : the input vector of the i-th sample;
[0037] The reconstructed output vector of the i-th sample, generated by the decoder;
[0038] Reconstruction loss, representing the Euclidean distance between the input vector and the reconstructed output vector;
[0039] λ: the weight parameter of the regularization term, used to balance the influence of the reconstruction loss and the regularization term;
[0040] ‖z‖1: regularization term, representing the L1 norm of the latent representation, used to prevent overfitting;
[0041] Finally, the extracted 64-dimensional latent vector z can effectively characterize the vehicle motion characteristics.
[0042] As a further improvement of the above scheme, the CNN model includes:
[0043] 1) Basic Convolutional Feature Extraction Module: Perform a convolution operation on the face image through a 3×3 convolutional kernel to extract basic features and obtain a feature map of 112×112×64 pixels;
[0044] 2) Local Convolutional Feature Extraction Module: Perform a convolution operation on the feature map to extract local features and obtain a secondary feature map of 56×56×256 pixels;
[0045] 3) Global Convolutional Feature Extraction Module: Perform a convolution operation on the secondary feature map to construct global features and obtain a tertiary feature map of 28×28×512 pixels;
[0046] 4) Pooling Operation Module: Gradually compress the size of the tertiary feature map through 3 times of 2×2 max pooling operations, gradually compressing from 112×112 pixels to 14×14 pixels;
[0047] 5) Aggregate Global Feature Module: Perform global average pooling on the pooled tertiary feature map to obtain 256 scalar values representing the overall features;
[0048] 6) Dimensionality Reduction Mapping Module: Linearly transform the 256-dimensional feature vector to 64 dimensions to obtain the feature vector of the facial features.
[0049] As a further improvement of the above solution, the fusion method of the facial features and the vehicle features includes the following steps:
[0050] 1) Spatiotemporal Alignment;
[0051] 2) Feature Concatenation and Normalization;
[0052] 3) Cross-Modal Attention Fusion.
[0053] Further, for the spatiotemporal alignment: Use the PPS signal of GPS to synchronize the timestamps of the facial feature sequence and the vehicle feature sequence; among them, the facial feature sequence is composed of multiple corresponding facial features that change over time, and the vehicle feature sequence is respectively formed by multiple corresponding facial features that change over time.
[0054] Further, for the feature concatenation and normalization: Concatenate along the feature dimension.
[0055] Further, for the cross-modal attention fusion: Use the vehicle features as the query, the facial features as the key and value, retrieve relevant information through similarity, calculate the attention of the vehicle features to each facial feature component, and generate dynamic weights.
[0056] As a further improvement of the above solution, the face image is preprocessed and then input into the CNN model to extract the facial features: First, the RetinaFace model is used to detect the face bounding box in the face image, and the coordinates of 5 key points, namely the two eye pupils, the tip of the nose, and the corners of the mouth, are located. The RetinaFace model simultaneously outputs the detection box position, the key point coordinates, and the confidence level through multi-task learning; Second, only the detection results with a confidence level higher than 0.95 are retained; Subsequently, based on the detected 5 key point coordinates and the preset standard face coordinates, through affine transformation matrix calculation, the non-standard face is projected and corrected to the front view angle to generate a standardized face image of 112×112 pixels; Finally, pixel normalization processing is performed on the aligned face image, and the RGB channel values are linearly scaled from [0, 255] to the range of [-1, 1] to form the normalized input data for direct processing by the CNN model.
[0057] Further, the key coordinate points are:
[0058]
[0059] Standard face coordinates: The two eyes are horizontally symmetric, the tip of the nose is centered, and the corners of the mouth are aligned.
[0060] As a further improvement of the above solution, the vehicle driving parameters are preprocessed and then the vehicle features are extracted through a specific autoencoder: The data with different dimensions of acceleration and vehicle speed are normalized by Z-score; The time series data is segmented into windows of a fixed length - each 5 seconds is a sample; The data is smoothed through a filtering algorithm to remove noise.
[0061] As a further improvement of the above solution, the fatigue driving detection method further includes: Inputting the fatigue feature vector after fusing the facial features and the vehicle features into a bidirectional LSTM network, outputting the driver's fatigue state through a fully connected layer, and giving an alarm according to the fatigue state classification.
[0062] The present invention also provides a fatigue driving detection system with multi-feature fusion, which uses any of the above multi-feature fusion fatigue driving detection methods to determine whether the driver is fatigued. The fatigue driving detection system includes: a feature extraction module for feature extraction, a multi-modal fusion module for multi-modal fusion, and a classification and alarm module for classification and alarm.
[0063] The present invention introduces the following strategies:
[0064] 1) Dynamic feature fusion: Use a gating mechanism (Gating Network) to dynamically adjust the modal weights according to the scene (such as light intensity, road curvature).
[0065] 2) Cross-modal attention: Design a spatio-temporal attention module (CNN model) to capture the temporal correlation between facial actions and steering wheel operations;
[0066] 3) Autoencoder neural network structure to extract vehicle data and perform dimensionality reduction and redundancy removal: Vehicle data has high dimensions and redundancy (such as high correlation between adjacent time points). The autoencoder can extract key dynamic patterns, reducing the dimension to 10% - 20% of the original data; the autoencoder automatically captures non-linear relationships to avoid human bias.
[0067] 4) The present invention, through the collaborative combination of the improved CNN model and the improved autoencoder, completely solves the technical defects of feature redundancy and noise interference, information loss and sub-optimal decision-making existing in the existing fatigue driving detection system with multi-feature fusion. Because of the unique capabilities of the CNN model and the autoencoder in processing different data types and feature extraction, their collaborative work systematically solves the problems of noise, redundancy and information fragmentation in multi-modal data through feature complementarity, dynamic weight allocation and joint optimization mechanisms. Its advantages lie in balancing model performance and computational efficiency, and at the same time having strong scene adaptation ability, providing a feasible technical solution for fatigue driving detection. Specifically as follows:
[0068] a. Complementary feature extraction and noise reduction collaboration
[0069] Local feature extraction ability of the CNN model: Through multi-layer convolution and pooling operations, the CNN model can efficiently capture fine-grained features in facial images (such as eye closure state, drooping corners of the mouth, etc.). Its structural characteristics (local receptive field, weight sharing) make it have a certain robustness to image noise.
[0070] Data compression and redundancy removal of the autoencoder: For high-dimensional temporal vehicle data, the autoencoder realizes non-linear dimensionality reduction through an encoding-decoding mechanism, effectively filtering out redundant information (such as strongly correlated data at adjacent time points), while retaining key dynamic patterns.
[0071] Combined advantage: The combination of the two reduces noise interference and redundant features in multi-modal data from the source, enhancing the robustness and accuracy of the system.
[0072] b. Feature fusion with dynamic scene adaptability
[0073] Gating mechanism and attention model: Through the cross-modal attention mechanism, vehicle data is used as a "query" to dynamically retrieve associated facial features (such as the temporal correlation between eye closing actions and steering wheel stagnation), generating scene-adaptive fusion weights. For example, when strong light interference causes facial features to fail, the system automatically increases the decision weight of vehicle data; in complex road conditions, it relies on the stability of facial features.
[0074] Sparse Regularization and Information Retention: The autoencoder introduces L1 sparse constraints, forcing the latent layer to activate only the key vehicle features related to fatigue and avoiding redundancy with the facial features extracted by the CNN model.
[0075] Combined Advantages: In complex driving scenarios (such as day-night switching, complex road conditions), the collaboration between the CNN model and the autoencoder can dynamically adapt to environmental changes and avoid the failure of a single modality. The collaboration between the two significantly improves the interpretability and discriminability of multi-modal features.
[0076] c. Enhanced Robustness in Complex Scenarios
[0077] Adversarial Data Incompleteness: When single-modal data is missing or of low quality (such as facial occlusion, vehicle sensor failure), the vehicle features reconstructed by the autoencoder and the facial features extracted by the CNN model can be complementary through the attention mechanism, avoiding the problem of noise diffusion in traditional splicing fusion.
[0078] Temporal Consistency Assurance: Through the hard synchronization of GPS-PPS signals and the dynamic time warping algorithm, the spatio-temporal alignment accuracy of multi-modal data is ensured, avoiding feature mis-matching caused by temporal misalignment.
[0079] Combined Advantages: The combination of the CNN model and the autoencoder provides high-quality input features for the bidirectional LSTM, enabling it to more accurately model the fatigue state of the driver. Brief Description of the Drawings
[0080] Figure 1 It is a flowchart of a fatigue driving detection method for multi-feature fusion.
[0081] Figure 2 To implement Figure 1 The structural diagram of the fatigue driving detection model for the method in
[0082] Figure 3 For Figure 2 The structural diagram of the CNN model adopted by the CNN network module in
[0083] Figure 4 For Figure 3 The schematic diagram of the model training loss of the CNN model in
[0084] Figure 5 It is the recognition schematic diagram of multi-modal mode 1 without using the fatigue driving detection method of the present invention.
[0085] Figure 6 For Figure 1 The recognition schematic diagram of multi-modal mode 2 using the method in Detailed Implementation Manner
[0086] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0087] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.
[0088] Please refer to Figure 1 , the fatigue driving detection method with multi-feature fusion of the present invention mainly includes the following steps: a) Step 1, data collection; b) Step 2, data preprocessing; c) Step 3, feature extraction; d) Step 4, multimodal fusion; e) Step 5, temporal modeling; f) Step 6, classification and warning. Based on the thinking concept of the fatigue driving detection method with multi-feature fusion, combined with the trained fatigue driving detection model, the corresponding fatigue driving detection system with multi-feature fusion can be built. The fatigue driving detection system includes a data collection module for data collection (i.e., including Figure 2 the vehicle data collection module and the facial data collection module in Figure 2 ), a data preprocessing module for data preprocessing (i.e., including Figure 2 the vehicle data preprocessing module and the facial data preprocessing module in Figure 2 ), a feature extraction module for feature extraction (i.e., including Figure 2 the autoencoder neural network module and the CNN network module in
[0089] The technical key points of the present invention are that the following strategies are introduced in this solution: 1) Dynamic feature fusion: Use a gating mechanism (Gating Network) to dynamically adjust the modal weights according to the scenario (such as light intensity, road curvature); 2) Cross-modal attention: Design a spatio-temporal attention module to capture the temporal correlation between facial movements and steering wheel operations; 3) Use the autoencoder neural network structure to extract vehicle data and perform dimensionality reduction and redundancy removal: The vehicle data has a high dimension and redundancy (such as high correlation between adjacent time point data). The autoencoder can extract key dynamic patterns, and the dimension is reduced to 10% - 20% of the original data; the autoencoder automatically captures non-linear relationships to avoid human bias. The following will be described in detail one by one.
[0090] a) Step 1: Data acquisition.
[0091] Collect three - dimensional acceleration, angular velocity, and speed at a frequency of 10Hz through the GPS / IMU module and transmit them via the CAN bus; collect facial image data using an in - vehicle infrared camera (640*480) and crop it to the 112*112 ROI area.
[0092] In this embodiment, dataset acquisition is to collect facial videos at 30fps using an in - vehicle infrared camera (640*480 resolution), crop the 112*112 RIO area (annotate the coordinates of both eyes, the tip of the nose, and the corners of the mouth); collect three - dimensional acceleration, angular velocity, and vehicle speed information using the GPS / IMU module. It is obtained from 6 drivers of different ages and genders, with a total number of frames equal to 5119k and a total time exceeding 36 hours. Different vehicle speeds and light conditions are used, covering different passenger orders, driving routes, and traffic conditions. The dataset includes drivers with different facial features (with / without beards, long / short hair, etc.).
[0093] b) Step 2: Data pre - processing: 1) Vehicle data pre - processing; 2) Image data pre - processing.
[0094] 1) Vehicle data pre - processing.
[0095] Normalize data with different dimensions such as acceleration and vehicle speed through Z - score; segment time - series data into fixed - length windows - each 5 seconds as a sample; smooth the data through a filtering algorithm to remove noise.
[0096] 2) Image data pre - processing.
[0097] Image data pre - processing is mainly divided into four steps.
[0098] First, detect the face bounding box in the image through the improved RetinaFace model and locate the coordinates of 5 key points including the pupils of both eyes, the tip of the nose, and the corners of the mouth. The model simultaneously outputs the position of the detection box, the coordinates of the key points, and the confidence level through multi - task learning.
[0099] Second, only retain high - quality detection results with a confidence level higher than 0.95 to filter out blurred or occluded faces.
[0100] Subsequently, based on the detected 5 key points and the preset standard face coordinates (both eyes horizontally symmetric, the tip of the nose centered, and the corners of the mouth aligned), calculate through the affine transformation matrix to project non - standard pose faces such as tilted and side - face ones to the front - facing view and generate a standardized face image of 112×112 pixels. The coordinates of the target key points are:
[0101]
[0102] Finally, pixel normalization is performed on the aligned image, linearly scaling the RGB channel values from [0, 255] to the range of [-1, 1], eliminating the influence of lighting differences, and forming normalized input data that can be directly processed by the neural network.
[0103] In this embodiment, data preprocessing is comprehensively labeled based on PERCLOS (eyelid closure time > 80% is fatigue), the number of consecutive yawns (> 3 times / minute), and abnormal vehicle behaviors (such as lane departure > 2 times / minute); normalization processing is performed on vehicle data, and data such as acceleration and vehicle speed are normalized by Z-score. The time series data is segmented into samples every 5 seconds, and the data is smoothed through a filtering algorithm to remove noise. For facial image data, an improved version of the RetinaFace model is used to detect the face bounding box and key point coordinates, and the detection results with a confidence level lower than 0.90 are filtered out. The face with a non-standard pose is corrected to a frontal view through affine transformation, and pixel normalization processing is performed.
[0104] c) Step three, feature extraction: 1) Vehicle data feature extraction; 2) Facial feature extraction.
[0105] 1) Vehicle data feature extraction.
[0106] Construct an autoencoder: adopt a symmetric deep neural network structure, and compress high-dimensional time series data into a low-dimensional latent space through non-linear mapping;
[0107] Input the vector by the encoder, and perform non-linear dimensionality reduction through a three-layer fully connected network:
[0108] h1 = ReLU(w1x + b1)
[0109] h2 = ReLU(w2h1 + b2)
[0110] z = w3h2 + b3
[0111] h1 is the input layer, h2 is the intermediate layer, and h3 is the output layer; w1, w2, and w3 are the weight matrices of each layer, used to map the input or the output of the previous layer to the current layer; b1, b2, and b3 are the bias terms of each layer, used to adjust the offset of the output; x is the input data; ReLU: activation function, used to introduce non-linear characteristics.
[0112] The decoder reconstructs the input through the reverse structure:
[0113]
[0114] W6, W5, W4: weight matrices, used to map the input or the output of the previous layer to the current layer;
[0115] b1, b2, b3: Bias terms used to adjust the offset of the output;
[0116] ReLU: Activation function used to introduce non - linear characteristics.
[0117] The training objective function combines the reconstruction error and the sparsity constraint:
[0118]
[0119] Γ: The value of the loss function, representing the total loss of the model;
[0120] N: The number of samples, representing the total number of input data;
[0121] x i : The input vector of the i - th sample;
[0122] The reconstructed output vector of the i - th sample, generated by the decoder;
[0123] Reconstruction loss, representing the Euclidean distance between the input vector and the reconstructed output vector;
[0124] λ: The weight parameter of the regularization term, used to balance the influence of the reconstruction loss and the regularization term;
[0125] ‖‖z‖‖1: Regularization term, representing the L1 norm of the latent representation, used to prevent overfitting.
[0126] The finally extracted 64 - dimensional latent vector z can effectively characterize the vehicle motion features.
[0127] 2) Facial feature extraction.
[0128] Input the pre - processed face image into the constructed CNN model. The model automatically learns and extracts the key features of the face through convolutional layers and pooling layers. The CNN model constructed in the present invention includes: a basic convolutional feature extraction module, a local convolutional feature extraction module, a global convolutional feature extraction module, a pooling operation module, an aggregated global feature module, and a dimensionality reduction mapping module.
[0129] Basic Convolutional Feature Extraction Module: Perform a convolution operation on the face image using a 3×3 convolutional kernel to extract basic features and obtain a feature map of 112×112×64 pixels. Local Convolutional Feature Extraction Module: Perform a convolution operation on the feature map to extract local features and obtain a secondary feature map of 56×56×256 pixels. Global Convolutional Feature Extraction Module: Perform a convolution operation on the secondary feature map to construct global features and obtain a tertiary feature map of 28×28×512 pixels. Pooling Operation Module: Gradually compress the size of the tertiary feature map through 3 times of 2×2 max pooling operations, gradually compressing from 112×112 pixels to 14×14 pixels. Aggregate Global Feature Module: Perform global average pooling on the pooled tertiary feature map to obtain 256 scalar values representing the overall features. Dimensionality Reduction Mapping Module: Linearly transform the 256-dimensional feature vector to 64 dimensions to obtain the feature vector of the facial features.
[0130] Therefore, the convolutional feature extraction is divided into 3 layers: Convolution Conv1: The convolutional kernel size is 3×3, and the number of convolutional kernels is 64. The obtained convolutional result is: 112×112×64 → Extract low-level features such as edges / textures; Convolution Conv2: The convolutional kernel size is 3×3, and the number of convolutional kernels is 256. The obtained convolutional result is: 56×56×256 → Combine low-level features into local features such as eyes and mouth corners; Convolution Conv3: The convolutional kernel size is 3×3, and the number of convolutional kernels is 512. The obtained convolutional result is: 28×28×512 → Construct global features such as the shape of the nose bridge and facial contour. Pooling layer: Through 3 times of 2×2 max pooling → The size of the feature map is gradually compressed from 112×112 to 14×14; Global feature aggregation: Global average pooling → Take the average of the 14×14 values of each feature map → Obtain 256 scalar values representing the overall features. Dimensionality reduction mapping: The fully connected layer linearly transforms 256 dimensions → to 64 dimensions → Output the final feature vector.
[0131] Step 3 is a quite crucial technical point in this case. Feature extraction: Use an autoencoder to perform feature extraction on vehicle data to obtain a 64-dimensional latent vector z; Input the preprocessed face image into the CNN model, and extract facial features through the convolutional layer and pooling layer, and finally output a 64-dimensional feature vector. Please refer to Figure 2 and Figure 3 , which shows the neural network model structure adopted by the present invention (including a convolutional neural network (CNN) for feature extraction and a bidirectional long short-term memory network (Bi-LSTM) for time series modeling). The CNN model automatically learns key features in the face image, such as local features like eyes and mouth corners, and global features like the shape of the nose bridge and facial contour, through multiple convolutional and pooling operations. The Bi-LSTM network is used to capture the temporal dependencies before and after in the fused feature sequence, so as to more accurately judge the fatigue state of the driver.
[0132] Extract the facial features through a specific CNN model: 1) Basic convolutional feature extraction; 2) Local convolutional feature extraction; 3) Global convolutional feature extraction; 4) Pooling operation; 5) Aggregate global features; 6) Dimensionality reduction mapping. Facial feature analysis: Based on computer vision, real-time capture micro-expression features such as the frequency of eye closure, the number of yawns, and the deviation of head posture, and use an improved RetinaFace model for image preprocessing. Then input the preprocessed image: Input the face image that has undergone pixel normalization into the CNN model, and then perform convolutional feature extraction. 1) Conv1: Perform a convolution operation on the image through a 3×3 convolutional kernel to extract low-level features such as edges and textures, and obtain a feature map of 112×112×64. 2) Conv2: Perform a convolution operation on the feature map output by Conv1 to further combine low-level features and extract local features such as eyes and mouth corners, and obtain a feature map of 56×56×256. 3) Conv3: Perform a convolution operation on the feature map output by Conv2 to construct global features such as the shape of the nose bridge and facial contour, and obtain a feature map of 28×28×512. 4) Pooling operation: Through 3 times of 2×2 max-pooling operations, gradually compress the size of the feature map, from 112×112 to 14×14 step by step, while retaining the most important feature information. 5) Global feature aggregation: Perform global average pooling on the pooled feature map, take the average of the 14×14 values of each feature map, and obtain 256 scalar values representing the overall features. 6) Dimensionality reduction mapping: Linearly transform the 256-dimensional feature vector to 64 dimensions through a fully connected layer to obtain the final feature vector.
[0133] Vehicle data extraction: Quantify the abnormal fluctuations of driving operations through behavior data such as braking / acceleration frequency and lane departure times. Integrate vehicle CAN bus data (such as vehicle speed), combine GPS positioning to analyze operation anomalies, and normalize the extracted data. Input the processed data into an autoencoder neural network. Autoencoder structure design: Adopt a symmetric three-layer fully connected network.
[0134] Encoder (input → 256 → 128 → 64)
[0135] Input layer: Input(shape=(256))
[0136] Hidden layer 1: Fully connected layer, 128 nodes, ReLU activation.
[0137] Hidden layer 2: Fully connected layer, 64 nodes, ReLU activation.
[0138] Latent layer: Fully connected layer, 64 nodes, linear activation, and add L1 sparse regularization.
[0139] Decoder (64 → 128 → 256)
[0140] Hidden layer 3: Fully connected layer, 64 nodes, ReLU activation.
[0141] Hidden layer 4: Fully connected layer, 128 nodes, ReLU activation.
[0142] Output layer: Fully connected layer, 256 nodes, Sigmoid activation (matching the input normalization range).
[0143] d) Step 4, Multimodal fusion: 1) Spatiotemporal alignment; 2) Feature concatenation and normalization; 3) Cross-modal attention fusion. Multimodal fusion uses the PPS signal of GPS to synchronize the timestamps of vehicle data and facial image data, and cubic spline interpolation is used to fill in the missing packets with a delay exceeding 5 ms. The vehicle feature sequence and the facial feature sequence are aligned through the dynamic time warping algorithm, and feature concatenation and normalization are performed. The cross-modal attention fusion mechanism is adopted, with vehicle features as queries, and facial features as keys and values, to calculate the attention of vehicle features to each facial feature component, generate dynamic weights, and mix vehicle data features with attention-enhanced features.
[0144] 1) Spatiotemporal alignment.
[0145] Timestamp synchronization: Hardware-level alignment, using the PPS signal of GPS, align the system clocks of the IMU sensor and the camera to the precision level; for packets with a delay exceeding 5 ms, cubic spline interpolation is used to fill in the missing frames;
[0146]
[0147] z vehical,t : Represents a certain state of the vehicle at time t;
[0148] a k : Represents the polynomial coefficients of cubic spline interpolation, and the value range of k is from 0 to 3; these coefficients are determined by fitting the known data points;
[0149] t: Represents the current time;
[0150] t base : Represents the reference time, usually a reference time point within the interpolation interval;
[0151] k: The exponent of the polynomial term, with a value range from 0 to 3, representing a cubic polynomial.
[0152] Dynamic time warping: Perform cost matrix calculation, for the vehicle feature sequence and the facial feature sequence, construct the cost matrix:
[0153]
[0154] C i,j: represents an element in the cost matrix, indicating the cost between the i-th point of the vehicle feature sequence and the j-th point of the facial feature sequence
[0155] point;
[0156] z vehical,i : represents the i-th feature vector of the vehicle feature sequence;
[0157] z face,j : represents the j-th feature vector of the facial feature sequence;
[0158] represents the square of the Euclidean distance between the vehicle feature and the facial feature, used to measure
[0159] the similarity of two feature points;
[0160] |i - j|: is the absolute difference between the indices of the vehicle feature sequence and the facial feature sequence, used to constrain the continuity and monotonicity of the path.
[0161] Then, use the dynamic programming algorithm to search for the minimum cost path from to, with the constraints: the path is continuous (only allowing right, up, and diagonal moves), and the path is monotonic (the i and j indices are strictly increasing).
[0162] Feature resampling interpolates and aligns the feature sequences of the two modalities according to the path to obtain the synchronized feature sequences:
[0163]
[0164] represents the synchronized vehicle feature sequence after feature resampling;
[0165] represents the synchronized facial feature sequence after feature resampling;
[0166] R T*d : represents the dimension of the feature sequence, where T is the time step and d is the dimension of the feature.
[0167] 2) Feature concatenation and normalization.
[0168] Concatenate along the feature dimension:
[0169]
[0170] z contact : the concatenated feature sequence, with a dimension of Tmin × 64;
[0171] T min : the minimum value of the time step, representing the time length of the feature sequence;
[0172] Add position encoding:
[0173] z concat + = SinusoidalPosEnc(T min , 64)
[0174] The feature sequence after adding position encoding;
[0175] SinusoidalPosEnc: Sinusoidal position encoding function.
[0176] Layer normalization: Perform layer normalization on the feature vectors at each time step:
[0177]
[0178] z normal,t : The normalized feature vector;
[0179] z concat,t : The feature vector of the concatenated feature sequence at time step t;
[0180] μ t : The mean of the feature vector Zconcat,t;
[0181] σ t : The standard deviation of the feature vector z concat,t;
[0182] ε: A very small constant used to prevent division by zero errors.
[0183] 3) Cross-modal attention fusion.
[0184] Query-Key-Value matrix generation, using vehicle features as the query (Query), facial features as the key (Key) and value (Value), and retrieving relevant information through similarity.
[0185] Calculate the attention of vehicle features to each facial feature component through scaled dot product:
[0186]
[0187] Q: Query, representing the vehicle feature sequence;
[0188] K: Key, representing the facial feature sequence;
[0189] V: Value, representing the facial feature sequence;
[0190] dk: Dimension of the key, which is 64 here;
[0191] Softmax: Activation function used to normalize attention scores into a probability distribution.
[0192] The dot product calculates the cosine similarity between the vehicle and the facial features, measuring the correlation between the vehicle features and the facial features; softmax normalizes it into a probability distribution, and weighted summation is used to obtain the context vector.
[0193] Concatenate the vehicle features and the attention output, and generate dynamic weights through a fully connected layer:
[0194] g = σ(W g [z v_norm ; Attention(Q, K, V)] + b g )
[0195] g: Gating weight, used to control the fusion ratio of the vehicle features and the attention-enhanced features;
[0196] σ: Activation function, usually the sigmoid function, used to limit the output within the range of (0, 1);
[0197] Wg: Weight matrix, used to map the concatenated features to the gating weight;
[0198] bg: Bias term;
[0199] z v_norm : Normalized vehicle features;
[0200] Attention(Q, K, V): Output of the attention mechanism.
[0201] Mix the vehicle data features and the attention-enhanced features according to the gating weight:
[0202] Z fused = g · z v_norm + (1 - g) · Attention(Q, K, V)
[0203] z fused : Fused features;
[0204] g: Gating weight;
[0205] z v_norm : Normalized vehicle features;
[0206] Attention(Q, K, V): Output of the attention mechanism.
[0207] Motion state feedback: When the confidence of the vehicle features is high (such as sudden braking, sharp turning), g → 1, dominating the decision-making; when the confidence of the facial features is high (such as closing eyes, yawning), g → 0, emphasizing visual information.
[0208] In this embodiment, multi-source data fusion: spatio-temporally align two types of data, and use the PPS signal of GPS to synchronize the timestamps of vehicle data and facial image data. Align the vehicle feature sequence and the facial feature sequence through the dynamic time warping algorithm, and perform feature splicing and normalization. Use the vehicle features as queries, and the facial features as keys and values, retrieve relevant information through similarity, calculate the attention of vehicle features to each facial feature component, generate dynamic weights, and mix the vehicle data features with the attention-enhanced features.
[0209] e) Step Five: Temporal sequence modeling: 1) Bi-LSTM temporal sequence modeling; 2) Temporal feature aggregation; 3) Fatigue state classification.
[0210] 1) Bi-LSTM temporal sequence modeling.
[0211] Divide the above synchronously aligned fusion feature sequence into a training set, a validation set, and a test set according to the ratio of 7:2:1;
[0212] Input the synchronously aligned fusion feature sequence with (30 time steps * 64 dimensions) into a bidirectional LSTM network to capture the temporal dependencies before and after:
[0213]
[0214] The hidden state of the forward LSTM at time step t;
[0215] z fused,t : The input of the fusion feature sequence at time step t;
[0216] The hidden state of the forward LSTM at time step t-1;
[0217] The hidden state of the backward LSTM at time step t;
[0218] The hidden state of the backward LSTM at time step t+1;
[0219] ht: The final hidden state, concatenated by the hidden states of the forward and backward LSTMs, with a dimension of 128.
[0220] 2) Temporal feature aggregation.
[0221] Mean pooling: Aggregate all hidden states within the time window;
[0222] 3) Fatigue state classification.
[0223] Fully connected classification layer:
[0224] p = sigmoid(W ch pool +b c )
[0225] P: Probability of fatigue state, with a value range of (0, 1)
[0226] Wc: Weight matrix of the fully connected layer;
[0227] bc: Bias term of the fully connected layer;
[0228] sigmoid: Activation function, used to limit the output within the range of (0, 1), representing the probability of fatigue state.
[0229] (1) In this embodiment, based on the LSTM time series model: the fused fatigue feature vector is input into a bidirectional LSTM network, and the driver's fatigue state is output through a fully connected layer. Then, based on the output result, the fatigue state is judged, and whether the driver is in a fatigue driving state and the level of fatigue degree are output.
[0230] f) Step six, classification and alarm.
[0231] When the system determines that it is in a fatigue driving state, a reminder is issued.
[0232] Based on the thinking concept of the above multi-feature fusion-based fatigue driving detection method, a corresponding fatigue driving detection model is established, then the model is trained, and finally the trained model is applied in practice.
[0233] The specific steps of model training include: a) obtaining a data set; b) establishing a loss function required for training the model; c) determining the network hyperparameters required for training the network; d) training, iteratively updating the weight coefficients, and obtaining the combination of weight coefficients of the neural network model with the optimal accuracy and generalization performance; e) combining the optimal weight coefficient combination with the model to construct the final fatigue driving detection model, which can accurately judge whether the driver is in a fatigue state in the actual task of driver fatigue detection.
[0234] Among them, as Figure 4 shown, the loss function design: due to the multi-modal nature of the fused data, a multi-task joint loss function is used to decouple and complement features: the reconstruction loss and sparse constraint are jointly optimized, forcing the autoencoder to extract low-dimensional sparse features of vehicle data and avoid redundancy with facial features. The classification loss guides the multi-modal fusion module to focus on cross-modal feature interactions strongly related to fatigue. The task weight coefficients are aligned according to task priorities and dimensions, and automatic weight adjustment based on task uncertainty.
[0235] L total = λ1L recon + λ2L sparse + λ3L class
[0236] L total : The total loss function, representing the comprehensive loss of the model;
[0237] λ1, λ2, λ3: Weight coefficients, used to balance the impacts of reconstruction loss, sparse constraint, and classification loss.
[0238] Reconstruction loss: The mean squared error of the autoencoder, constraining the low-dimensional representation ability of vehicle data
[0239]
[0240] L reconl : Reconstruction loss, measuring the difference between the input data and the reconstructed data.
[0241] N: The number of samples;
[0242] x i : The input vector of the i-th sample;
[0243] The reconstructed output vector of the i-th sample;
[0244] The square of the Euclidean distance between the input vector and the reconstructed output vector.
[0245] Sparse constraint: The KL divergence penalizes the sparsity of the latent vector z, preventing overfitting
[0246]
[0247] L sparse : Sparse constraint loss, used to penalize the sparsity of the latent vector;
[0248] ρ: The expected sparsity, representing the proportion of non-zero elements in the latent vector;
[0249] The actual sparsity of the j-th element of the latent vector.
[0250] Classification loss: The weighted cross-entropy for fatigue state classification, addressing class imbalance:
[0251]
[0252] L class : Classification loss, used to measure the accuracy of model classification;
[0253] Wc: Class weights, used to address the class imbalance problem;
[0254] yc: The true label of the c-th class;
[0255] pc: The predicted probability of the c-th category;
[0256] N total : The total number of samples;
[0257] N C : The number of samples of the c-th category.
[0258] Among them, for model configuration and training: The Adam optimizer is selected as the iterative calculation scheme for model parameters, and the initial learning rate is set to 3×10 -4 , and the cosine annealing scheduler and mixed-precision training are used to balance the convergence speed and stability.
[0259] The training is divided into two stages: a) Pre-training stage: Independently optimize the vehicle autoencoder (100 epochs, minimizing reconstruction and sparsity losses) and the facial CNN (50 epochs, cross-entropy loss), and freeze the backbone network parameters; b) Joint training stage: Unfreeze all parameters, input multi-modal data in synchronous batches, dynamically integrate features through cross-modal attention gating, input into a two-layer LSTM to model temporal dependencies, and train for a total of 200 epochs.
[0260] Based on the F1 score of the validation set, trace back the optimal weights; Monitor the total loss, modal alignment error, and classification confidence throughout the training process to ensure the collaborative optimization of multi-modal features.
[0261] The specific steps for the actual application of the fatigue driving detection system include: f) Preprocess the image or video data in the actual detection task to make it meet the input requirements of the model; g) Use the preprocessed data as input, perform fatigue driving detection through the constructed model, and obtain the fatigue detection result; h) According to the detection result, judge whether the driver is in a fatigue state and take corresponding measures, such as issuing an alarm, etc., to remind the driver to pay attention to safety.
[0262] Among them, after obtaining the fatigue detection result, analyze the fatigue probability to judge the fatigue state. Data collection: The in-vehicle infrared camera captures facial images at 30 fps, and detects the eye closure frequency (PERCLOS), the number of yawns, and the head pose; Synchronously obtain vehicle data such as vehicle speed, steering angle, and acceleration through GPS; Single-modal threshold trigger: When the proportion of eyelid closure time > 80%, the number of yawns > 3 times within 10 minutes, or the number of lane departures > 2 times, it is marked as abnormal, as shown in Table 1.
[0263] Table 1 Feature fatigue thresholds
[0264]
[0265] Among them, dynamic weight fusion: if the vehicle brakes suddenly (acceleration < -0.5g) or turns violently, the weight of vehicle features is increased to 90%; if the facial features fail (such as strong light interference), the decision-making depends on vehicle data.
[0266] Sequential modeling classification: The fused features are input into a bidirectional LSTM to analyze the continuous 5-second behavior pattern, and the fatigue probability value P is output. fatigue 。
[0267] b) Evaluation metrics: accuracy, recall rate, F1 score. As shown in Table 2, the performance of single-modal and multi-modal fusion is compared:
[0268] Multi-modal (vehicle + face) method 1 (as Figure 5 shown): The facial data features and vehicle data features are extracted using CNN and time-domain analysis respectively, and then directly concatenated into a joint feature vector. Finally, the joint feature vector is input into a fully connected network (such as MLP) for binary classification (fatigue / normal).
[0269] Multi-modal (vehicle + face) method 2 (as Figure 6 shown): Independent models are used for training respectively: facial model - a classifier is trained based on CNN to output the fatigue probability; vehicle model - a classifier is trained based on a sequential model (such as HMM) to output the fatigue probability. Finally, the fatigue probability results of the two are weighted and averaged, and then judged according to logical rules.
[0270] Multi-modal (this scheme): An autoencoder neural network model and an improved lightweight CNN are used to extract vehicle data and facial data, then attention mechanism fusion and dynamic gating weight monitoring are performed, and finally input into LSTM for training and classification, and finally judged according to the results.
[0271] Table 2 Comparison of the performance of traditional fusion and this scheme's fusion
[0272] Model variant Accuracy F1 score Multimodal (vehicle + face) method 1 89.9% 0.85 Multimodal (vehicle + face) method 2 77.8% 0.76 Multimodal (this solution) 93.3% 0.90
[0273] Among them, based on the obtained fatigue probability value, grading alarms are carried out. The system adopts a three-level progressive alarm mechanism according to the fatigue probability calculated in real time and the continuous determination times, as shown in Table 3.
[0274] Table 3 Alarm mechanism
[0275]
[0276] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0277] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.
Claims
1. A fatigue driving detection method based on multi - feature fusion, comprising the following steps: Extract facial features from the driver's face image; Extract vehicle features from the vehicle driving parameters of the vehicle driven by the driver; Fuse the facial features and vehicle features to determine whether the driver is fatigued; It is characterized in that the facial features are extracted by designing a CNN model: 1) Basic convolution feature extraction; 2) Local convolution feature extraction; 3) Global convolution feature extraction; 4) Pooling operation; 5) Aggregating global features; 6) Dimensionality reduction mapping; The vehicle features are also extracted by designing an auto - encoder: adopting a symmetric deep neural network structure, and compressing high - dimensional time - series data into a low - dimensional latent space through non - linear mapping.
2. The fatigue driving detection method with multi-feature fusion according to claim 1, characterized in that, The auto - encoder includes: 1) Encoder: input → 256 → 128 → 64; Input layer: input, 256 nodes; Hidden layer 1: fully - connected layer, 128 nodes, ReLU activation; Hidden layer 2: fully - connected layer, 64 nodes, ReLU activation; Latent layer: fully - connected layer, 64 nodes, linear activation, and add L1 sparse regularization; 2) Decoder: output 64 → 128 → 256; Hidden layer 3: fully - connected layer, 64 nodes, ReLU activation; Hidden layer 4: fully - connected layer, 128 nodes, ReLU activation; Output layer: fully - connected layer, 256 nodes, Sigmoid activation to match the input normalization range.
3. The fatigue driving detection method based on multi - feature fusion according to claim 2, characterized in that The input vector is input by the encoder, and non - linear dimensionality reduction is performed through a three - layer fully - connected network: h1 = ReLU(w1x + b1) h2 = ReLU(w2h1 + b2) h3 = w3h2 + b3 h1 is the input layer, h2 is the intermediate layer, and h3 is the output layer; w1, w2, w3 are the weight matrices of each layer, used to map the input or the output of the previous layer to the current layer; b1, b2, b3 are the bias terms of each layer, used to adjust the offset of the output; x is the input data; ReLU: activation function, used to introduce non - linear characteristics; The decoder reconstructs the input through an inverse structure: W6, W5, W4: weight matrices; b4, b5, b6 are bias terms; The training objective function combines the reconstruction error and the sparse constraint: Γ: the value of the loss function, representing the total loss of the model; N: the number of samples, representing the total number of input data; xi: the input vector of the i - th sample; The reconstructed output vector of the i-th sample, generated by the decoder; Reconstruction loss, representing the Euclidean distance between the input vector and the reconstructed output vector; λ: the weight parameter of the regularization term, used to balance the influence of the reconstruction loss and the regularization term; ‖‖z‖‖1: regularization term, representing the L1 norm of the latent representation, used to prevent overfitting; Finally, the extracted 64 - dimensional latent vector z can effectively represent the vehicle motion characteristics.
4. The fatigue driving detection method with multi-feature fusion according to claim 1, wherein The CNN model includes: 1) Basic convolution feature extraction module: perform convolution operation on the face image through a 3×3 convolution kernel, extract basic features, and obtain a feature map of 112×112×64 pixels; 2) Local convolution feature extraction module: perform convolution operation on the feature map, extract local features, and obtain a secondary feature map of 56×56×256 pixels; 3) Global Convolution Feature Extraction Module: Perform convolution operations on the secondary feature map to construct global features, obtaining a tertiary feature map of 28×28×512 pixels; 4) Pooling Operation Module: Gradually compress the size of the tertiary feature map through 3 times of 2×2 max pooling operations, gradually compressing from 112×112 pixels to 14×14 pixels; 5) Aggregate Global Feature Module: Perform global average pooling on the pooled tertiary feature map to obtain 256 scalar values representing the overall features; 6) Dimensionality Reduction Mapping Module: Linearly transform the 256-dimensional feature vector to 64 dimensions to obtain the feature vector of the facial features.
5. The fatigue driving detection method with multi-feature fusion according to claim 1, characterized in that The fusion method of the facial features and the vehicle features includes the following steps: 1) Spatiotemporal alignment; 2) Feature splicing and normalization; 3) Cross-modal attention fusion.
6. The fatigue driving detection method with multi-feature fusion according to claim 5, wherein The spatiotemporal alignment: Use the PPS signal of GPS to synchronize the timestamps of the facial feature sequence and the vehicle feature sequence; among them, the facial feature sequence is composed of multiple corresponding facial features that change over time, and the vehicle feature sequence is respectively formed by multiple corresponding facial features that change over time; And / or, the feature splicing and normalization: Splice along the feature dimension. And / or, the cross-modal attention fusion: Use the vehicle features as queries, the facial features as keys and values, retrieve relevant information through similarity, calculate the attention of the vehicle features to each facial feature component, and generate dynamic weights.
7. The fatigue driving detection method with multi-feature fusion according to claim 1, characterized in that, The face image is preprocessed and then input into the CNN model to extract the facial features: First, detect the face bounding box in the face image through the RetinaFace model, and locate the coordinates of 5 key points including the pupils of both eyes, the tip of the nose, and the corners of the mouth. The RetinaFace model simultaneously outputs the detection box position, key point coordinates, and confidence through multi-task learning; Secondly, only retain the detection results with a confidence higher than 0.95; Subsequently, based on the detected 5 key point coordinates and the preset standard face coordinates, calculate through the affine transformation matrix, and project the non-standard face to the frontal view to generate a standardized face image of 112×112 pixels; Finally, perform pixel normalization processing on the aligned face image, linearly scale the RGB channel values from [0, 255] to the range of [-1, 1], forming the normalized input data for direct processing by the CNN model; And / or, the vehicle driving parameters are preprocessed and then the vehicle features are extracted through a specific autoencoder: Normalize the data with different dimensions such as acceleration and vehicle speed through Z-score; Cut the time series data into windows of fixed length - each 5 seconds as a sample; Smooth the data through a filtering algorithm to remove noise.
8. The fatigue driving detection method with multi-feature fusion according to claim 7, wherein The coordinates of the key coordinate points are: Standard face coordinates: The eyes are horizontally symmetric, the tip of the nose is centered, and the corners of the mouth are aligned.
9. The fatigue driving detection method with multi-feature fusion according to claim 1, wherein The fatigue driving detection method further includes: Input the fatigue feature vector after fusing the facial features and the vehicle features into a bidirectional LSTM network, output the driver's fatigue state through a fully connected layer, and issue an alarm according to the fatigue state classification.
10. A fatigue driving detection system with multi-feature fusion, characterized in that, It uses the fatigue driving detection method with multi-feature fusion described in any one of claims 1 to 9 to determine whether the driver is fatigued. The fatigue driving detection system includes: a feature extraction module for feature extraction, a multi-modal fusion module for multi-modal fusion, and a classification and warning module for classification and warning.
Citation Information
Cited By
Driver fatigue detection method and system based on reconstruction enhanced evidence network
CN120713525A
Brain fatigue dynamic early warning method and device, computer equipment and storage medium
CN120899254A
A brain fatigue dynamic early warning method and device, computer equipment and storage medium
CN120899254B
Compression method and system for multi-source heterogeneous time series data, and motor fault prediction method and system
CN121036770A
Abnormal driving recognition method based on multiple modes of pedal data and facial expressions
CN121188645A