Fatigue driving detection method based on NIR and RGB image feature fusion
By constructing a multimodal detection model that fuses NIR and RGB image features, the problems of insufficient light sensitivity and feature interaction in existing fatigue driving detection are solved. It achieves high-precision and robust detection under different lighting conditions, adapts to complex driving environments, and has good generalization ability and real-time performance.
Patent Information
- Application Number
- CN202510697837.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-10-21
AI Technical Summary
Existing fatigue driving detection methods are easily affected by lighting conditions. Single-modal image detection suffers from problems such as light sensitivity or lack of color information. Multimodal feature interaction is insufficient, resulting in insufficient detection accuracy and robustness. Furthermore, the model training process suffers from poor adaptability to changes in lighting and the risk of overfitting.
A multimodal detection model based on the fusion of NIR and RGB image features is constructed. Features are extracted through the MobileNetV2 backbone network and adaptively fused by a cross-modal attention module. The AdamW optimizer and OneCycleLR scheduler are used to optimize training. Data augmentation strategies are introduced to achieve adaptive feature fusion and improve model stability.
It improves the accuracy and robustness of fatigue driving detection, maintains stable detection performance under different lighting conditions, has good generalization ability and real-time performance, and the lightweight design of the model meets the real-time requirements of vehicle-mounted equipment.
Smart Images

Figure CN120823583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a fatigue driving detection method based on NIR and RGB image feature fusion. Background Art
[0002] With the continued growth of vehicle ownership, driver fatigue has become a major factor in traffic accidents. Accurately detecting driver fatigue is crucial for improving driving safety. Existing fatigue detection methods primarily rely on single-modality image information, such as RGB or near-infrared (NIR) images.
[0003] Traditional fatigue detection methods based on RGB images are susceptible to lighting conditions. At night or in strong light, image quality degrades significantly, leading to inaccurate facial feature extraction and, in turn, impacting fatigue detection accuracy. While NIR image-based detection methods are insensitive to lighting variations and can capture clear facial images in low-light conditions, NIR images lack color information, making it difficult to fully describe the driver's facial state. Consequently, their effectiveness is limited when used alone.
[0004] Currently, some studies have attempted to combine multimodal data for fatigue detection, but most employ simple feature concatenation or early fusion strategies, failing to fully exploit the complementary information between RGB and NIR images across different spectra. Existing methods also underutilize multiscale features, making it difficult to capture subtle facial changes in the driver, resulting in poor robustness in complex scenarios. Furthermore, existing technologies lack adaptive mechanisms for handling intermodal feature interactions, making it impossible to dynamically adjust fusion strategies based on varying inputs, further limiting improvements in detection performance.
[0005] In terms of model training and optimization, existing technologies also have shortcomings: traditional fatigue detection methods rely on single-modality images (such as RGB or NIR), which are sensitive to light (RGB) or lack color information (NIR). The differences in the category distribution of the dataset are not fully considered during the training process of some models, resulting in poor detection results of the model on certain categories. Existing multimodal fusion methods are also mostly simple feature splicing, which does not fully explore complementary information and lacks dynamic adjustment capabilities, resulting in limited detection accuracy and robustness. Moreover, in the selection of optimization algorithms, some methods cannot strike a good balance between convergence speed and training accuracy, and are prone to falling into local optimal solutions, affecting the overall performance of the model. Summary of the Invention
[0006] To address the limitations of single-modality image detection in existing technologies, as well as the inadequacy of multimodal feature interaction, low model efficiency, and poor generalization, the present invention provides a fatigue driving detection method based on NIR and RGB image feature fusion. By constructing a multimodal fusion model, it effectively integrates the advantageous features of the two images, thereby improving the accuracy and robustness of fatigue detection. The method mainly includes:
[0007] S1: Acquire driver facial images under different lighting conditions, perform data preprocessing and dataset construction;
[0008] S2: Build a fatigue driving detection network model, which includes MobileNetV2, a cross-modal attention module, a convolutional layer, a global average pooling layer, a fully connected layer, and an output layer;
[0009] Use the pre-trained MobileNetV2 as the backbone network to extract features of RGB and NIR images respectively;
[0010] A cross-modal attention module is used to process RGB and NIR features and adaptively fuse image features;
[0011] The fused features are input into the fusion convolution layer for feature integration to obtain a two-dimensional feature map;
[0012] Use global average pooling to convert the two-dimensional feature map into a one-dimensional vector;
[0013] Input the one-dimensional vector into the fully connected layer for fatigue state classification;
[0014] The output layer uses the softmax function to generate the probability distribution of various fatigue states;
[0015] S3: Use the dataset to train the fatigue driving detection network model. During training, class weights are calculated and the AdamW optimizer is used to decouple weight decay. The OneCycleLR learning rate scheduler is used to dynamically adjust the learning rate based on the training rounds, and an early stopping mechanism is used to avoid overfitting.
[0016] S4: Input the actual facial image into the trained fatigue driving detection network model to obtain the probability distribution of the fatigue state, and then judge the driver's fatigue state.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above method.
[0019] A computer program product includes a computer program or instructions, which implement the steps of the above method when the program or instructions are executed by a processor.
[0020] The beneficial effects brought about by the technical solution provided by the present invention are:
[0021] (1) By constructing a multimodal fusion model and combining the color detail information of RGB images with the illumination robustness advantages of NIR images, the present invention uses the proposed cross-modal attention module to dynamically allocate the fusion weights of the two modal features of RGB and NIR, making full use of the complementary information of the two modalities under different illumination conditions to enhance the richness of feature representation. This breaks through the limitations of traditional fixed fusion strategies and achieves adaptive feature fusion. Unlike traditional simple feature splicing or early fusion strategies, the present invention can automatically learn the importance of the two modal features based on the content of the input image, achieving more efficient fusion.
[0022] (2) During the training process, the model’s training accuracy and convergence speed are effectively balanced through category weight balancing, AdamW optimizer, OneCycleLR dynamic learning rate scheduling and early stopping mechanism, optimizing the model’s convergence speed and generalization ability, avoiding overfitting, and enabling the model to perform well in different data sets and actual scenarios. The introduction of OneCycleLR scheduling and AdamW optimizer significantly improves the model’s convergence efficiency and stability. The data enhancement strategy and multimodal fusion mechanism enable the model to adapt to complex actual driving environments, reducing the impact of factors such as illumination changes and posture changes on the detection results. In tests simulating different illumination intensities, angles, and different postures of drivers, the model of the present invention can still maintain stable detection performance, while the detection accuracy of the traditional single-modality model has dropped significantly. The pre-trained backbone network and reasonable model structure design enable the model to quickly adapt to the facial features of different drivers without the need for extensive training for specific individuals, and has strong generalization ability. When testing new driver samples that have not participated in training, the accuracy of the model of the present invention can still reach a high level, demonstrating its good generalization performance.
[0023] (3) The lightweight MobileNetV2 backbone network is adopted to reduce the number of model parameters and computational complexity, making the model run faster while ensuring detection accuracy, improving the model's operating efficiency and meeting the real-time requirements of practical applications. At the same time, combined with the SELayer attention mechanism, it focuses on key facial areas related to fatigue detection (such as eyes and mouth), ensuring the accuracy of feature extraction on a lightweight basis. While reducing model complexity, it improves the pertinence of feature extraction and achieves high-precision feature extraction in a lightweight model.
[0024] It can be seen that the detection method of the present invention has high detection accuracy, strong robustness, good generalization ability and high model efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0026] Figure 1 This is a flow chart of a fatigue driving detection method based on NIR and RGB image feature fusion in an embodiment of the present invention;
[0027] Figure 2 is a schematic diagram of an adaptive fusion module in an embodiment of the present invention;
[0028] Figure 3 is a schematic diagram of a fatigue driving detection network model in an embodiment of the present invention;
[0029] Figure 4 Schematic diagram of the training results of a dual-stream network based on NIR and RGB fusion in an embodiment of the present invention;
[0030] Figure 5 Schematic diagram of the training results of a single-stream NIR network according to an embodiment of the present invention;
[0031] Figure 6 Schematic diagram of the training results of an RGB single-stream network in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.
[0033] Example 1
[0034] Please refer to Figure 1 , Figure 1 This is a flow chart of a fatigue driving detection method based on NIR and RGB image feature fusion in an embodiment of the present invention, which specifically includes:
[0035] S1. Data preprocessing and dataset construction
[0036] Driver facial images under different lighting conditions are acquired, with each sample consisting of a pair of RGB and NIR images and a corresponding fatigue state label (e.g., awake, slightly asleep, yawning). The images are preprocessed, including resizing to 224×224 pixels to unify the image size for subsequent model processing. Data augmentation operations such as random horizontal flipping, color jittering, random rotation, and random cropping are also used to increase data diversity, improve model generalization, and reduce overfitting. The preprocessed images are converted into tensors and normalized. RGB images are normalized using a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225], while NIR images are normalized using a mean of [0.5] and a standard deviation of [0.5] to uniformly process the data and accelerate model convergence.
[0037] S2. Build a fatigue driving detection network model
[0038] (1) Feature extraction
[0039] A pre-trained MobileNetV2 backbone network was used to extract features from RGB and NIR images, respectively. Compared to traditional convolutional networks, MobileNetV2 is lightweight, reducing model parameters and computational complexity while maintaining feature extraction capabilities, thereby improving model efficiency. The trained single-stream model was 12.23MB in size, while the two-stream model was 28.63MB in size. For RGB images, the backbone network outputs a 320-dimensional feature map; for NIR images, the same 320-dimensional feature map was obtained using the MobileNetV2 backbone network.
[0040] During feature extraction, attention mechanisms such as the SELayer (Squeeze-and-Excitation Layer) are introduced to weight the RGB and NIR feature maps separately. The SELayer adaptively adjusts the weights along the channel dimension to highlight key facial features, such as the eyes and mouth, which are closely associated with fatigue. This suppresses interference from irrelevant background information and improves the targeted nature of feature extraction.
[0041] (2) Adaptive fusion of image features
[0042] After obtaining the features of RGB and NIR images in step S2, the cross-modal attention module is used to process the features of RGB and NIR. This is an attention mechanism specially designed for processing multimodal data, which is used to capture the complementary information between RGB and NIR images of two different modalities. This module is based on the query-key-value paradigm of the attention mechanism, and takes the features of one modality as the query (Query), and the features of the other modality as the key (Key) and value (Value). By calculating the attention weight matrix, the adaptive fusion of the two modal features is achieved. This method can fully explore the complementary information between the two modalities and enhance the richness of feature representation. Adaptive fusion module such as Figure 2 As shown in Figure 2, the cross-modal attention fusion module contains the following key components:
[0043] 1) Feature Dimensionality Reduction Layer: 1×1 convolution is used to reduce feature dimensions and computational complexity. RGB and NIR feature maps output from the MobileNetV2 backbone network are reduced in dimension using a 1×1 convolution kernel with an 8:1 ratio, reducing the dimension from 320 to 40. This preserves key feature information while significantly reducing the complexity of subsequent attention calculations.
[0044] 2) Bidirectional attention calculation unit: Calculate the attention weights in two directions: RGB guides NIR and NIR guides RGB. Let the RGB feature map be F RGB ∈R C×H×W , the NIR characteristic graph is F NIR ∈R C×H×W , where C is the number of channels, H and W are the height and width of the feature map respectively. First, the feature map is reshaped into a sequence form: F RGB →R C×(H×W) , F NIR →R C×(H×W) .
[0045] Calculate the attention weight matrix:
[0046] When using RGB features as queries and NIR features as keys and values:
[0047]
[0048] When using NIR features as queries and RGB features as keys and values:
[0049]
[0050] Where W Q and W k is the learnable transformation matrix, d k is a scaling factor used to stabilize the gradient.
[0051] 3) Feature fusion layer: Generate enhanced feature representation through weighted fusion. Calculate weighted features:
[0052]
[0053] Where W V is the transformation matrix of the values, Indicates A RGB→NIR The transpose of Indicates A NIR→RGB The transpose of .
[0054] 4) Residual connection mechanism: The final cross-modal fusion features achieve progressive feature fusion through learnable parameters.
[0055]
[0056] Among them, α is a learnable parameter used to control the ratio of the fusion of the two modal features. and Represents the fused RGB features and NIR features.
[0057] The fused features are input into the fusion convolutional layer, where 1×1 convolutions are used for feature integration, reducing the number of channels and computational complexity. Next, global average pooling is used to convert the two-dimensional feature map into a one-dimensional vector, and fatigue state classification is performed through a fully connected layer. The fully connected layer contains multiple hidden layers and uses the ReLU activation function to increase the model's nonlinear expressiveness. A dropout layer is also used to prevent overfitting. The output layer uses a softmax function to generate the probability distribution of the three fatigue states (awake, slightly asleep, and yawning).
[0058] The core difference between the cross-modal attention mechanism of the present invention and the ordinary attention mechanism is that the query, key, and value of the ordinary attention mechanism all come from the same input source, and mainly calculate the correlation between different positions within a single feature map to achieve self-attention of the features; while the present invention innovatively derives the query from the source modality (such as RGB), and the key and value from the target modality (such as NIR), to achieve true cross-modal feature interaction.
[0059] Specifically, the approach utilizes: 1) a bidirectional computational strategy, guiding NIR with RGB and vice versa, creating mutually reinforcing positive feedback; 2) an 8:1 dimensionality reduction ratio is used to optimize computational complexity, reducing the computational effort by approximately 85% compared to full-dimensional attention; and 3) learnable parameters are introduced for progressive feature fusion, with the initial value set to 0 to ensure training stability. This design allows the color detail of RGB and the illumination robustness of NIR to be mutually transferred and enhanced, achieving deep semantic interaction of multimodal features rather than traditional simple feature concatenation, significantly improving the accuracy and robustness of fatigue driving detection under various lighting conditions.
[0060] S3, classification and model training
[0061] During model training, the cross-entropy loss function is used as the optimization objective. Taking into account the possible differences in the class distribution of the dataset, class weights are calculated. The weight of each class is determined by the number of samples in the class and the total number of samples. This allows the model to pay more attention to classes with fewer samples during training, thereby balancing the class distribution differences in the dataset.
[0062] The AdamW optimizer is used. Compared to the traditional Adam optimizer, AdamW adds a decoupled implementation of weight decay to better prevent model overfitting. The initial learning rate is set to 1e-4, the weight decay is set to 1e-3, and the OneCycleLR learning rate scheduler is used to dynamically adjust the learning rate based on the number of training rounds. Training is performed with 100 iterations and a batch size of 16. In the early stages of training, the learning rate is gradually increased to its maximum value to accelerate model convergence; in the later stages of training, the learning rate is gradually decreased to ensure more stable convergence to the optimal solution.
[0063] During training, we use an early stopping mechanism. When the validation set accuracy does not improve within 10 consecutive epochs, we stop training to avoid overfitting. We also record various metrics during training, such as training loss, training accuracy, validation loss, and validation accuracy.
[0064] The dataset used in this training contains 38,000 images, 19,000 RGB-NIR image pairs, and 288 RGB-NIR video clips, of which 80% are used as training data and 20% as validation data. When constructing the dataset, it is necessary to take into account the diversity of different drivers, different lighting conditions, and different fatigue states to ensure the comprehensiveness of the data. At the same time, the data is grouped according to the driver ID to facilitate subsequent leave-one-out cross-validation and more accurate evaluation of model performance. Figure 3 shown.
[0065] Figure 4 It is based on the training results of the NIR and RGB fusion dual-stream network. Figure 5 It is the training result based on the NIR single-stream network. Figure 6 It is the training result based on the RGB single-stream network, Figure 4 、 Figure 5 、 Figure 6 Analysis of the training and test results shows that the dual-stream network based on NIR and RGB fusion is superior to the single-stream model trained on a single image in terms of model stability and detection accuracy.
[0066] After 100 rounds of training, the model with the highest accuracy was used for testing and verification. The detection accuracy of the three models is shown in Table 1:
[0067] Table 1 Detection accuracy of three models
[0068] Model Accuracy RGB single-stream model 91.67% NIR single-stream model 92.52% Two-stream model 96.30%
[0069] The size of the trained model is shown in Table 2:
[0070] Table 2: Model size after training
[0071] Model size RGB single-stream model 12.23MB NIR single-stream model 12.23MB Two-stream model 28.63MB
[0072] S4: Input the actual facial image into the trained fatigue driving detection network model to obtain the probability distribution of the fatigue state, and then judge the driver's fatigue state.
[0073] The present invention fully utilizes the advantages of color detail information and lighting robustness by fusing the features of RGB and NIR images, and can maintain a high detection accuracy under different lighting conditions. Experiments conducted according to the technical solution of the present invention show that the accuracy of the dual-stream fusion model on the test set is significantly higher than that of the single modality model (RGB single stream 91.67%, NIR single stream 92.52%), which effectively improves the reliability of fatigue driving detection. In actual tests, for fatigue state detection under complex lighting conditions, the accuracy of the model of the present invention reached 96.30%, which is 4.63% higher than the model based on a single RGB image and 3.78% higher than the model based on a single NIR image. In scenarios such as lighting changes and posture interference, the detection stability is improved by about 5%. It shows stronger robustness and real-time performance in scenarios such as complex lighting and posture changes, providing an efficient and reliable solution for fatigue driving detection, with actual deployment value, and the lightweight design of the model (the dual-stream model is only 28.63MB) meets the real-time requirements of vehicle-mounted equipment.
[0074] Example 2
[0075] A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0076] Example 3
[0077] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above method.
[0078] Example 4
[0079] A computer program product includes a computer program or instructions, which implement the steps of the above method when the program or instructions are executed by a processor.
[0080] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A fatigue driving detection method based on NIR and RGB image feature fusion, characterized in that: include: S1: Acquire driver facial images under different lighting conditions, perform data preprocessing and dataset construction; S2: Build a fatigue driving detection network model, which includes MobileNetV2, a cross-modal attention module, a convolutional layer, a global average pooling layer, a fully connected layer, and an output layer; Use the pre-trained MobileNetV2 as the backbone network to extract features of RGB and NIR images respectively; A cross-modal attention module is used to process RGB and NIR features and adaptively fuse image features; The fused features are input into the fusion convolution layer for feature integration to obtain a two-dimensional feature map; Use global average pooling to convert the two-dimensional feature map into a one-dimensional vector; Input the one-dimensional vector into the fully connected layer for fatigue state classification; The output layer uses the softmax function to generate the probability distribution of various fatigue states; S3: Use the dataset to train the fatigue driving detection network model. During training, class weights are calculated and the AdamW optimizer is used to decouple weight decay. The OneCycleLR learning rate scheduler is used to dynamically adjust the learning rate based on the training rounds, and an early stopping mechanism is used to avoid overfitting. S4: Input the actual facial image into the trained fatigue driving detection network model to obtain the probability distribution of the fatigue state, and then judge the driver's fatigue state.
2. The fatigue driving detection method based on NIR and RGB image feature fusion as claimed in claim 1, characterized in that: In S1, each acquired image includes a pair of RGB and NIR images, and a corresponding fatigue state label, where the fatigue state label includes awake, slightly asleep, and yawning.
3. The fatigue driving detection method based on NIR and RGB image feature fusion as claimed in claim 1, characterized in that: In S1, the image is preprocessed, including image size adjustment and data augmentation operations. By adjusting the image size, an image of uniform size is obtained; the data augmentation operations include random horizontal flipping, color jittering, random rotation and random cropping, which increase the diversity of the data.
4. The fatigue driving detection method based on NIR and RGB image feature fusion as claimed in claim 3, characterized in that: The preprocessed images are converted into tensors and normalized to obtain the dataset.
5. The fatigue driving detection method based on NIR and RGB image feature fusion as claimed in claim 1, characterized in that: In S2, RGB features are used as queries and NIR features as keys and values, or vice versa, and attention weights are calculated to achieve adaptive fusion of the two modal features. Specifically: When RGB features are used as queries and NIR features as keys and values, the attention weights are: When NIR features are used as queries and RGB features as keys and values, the attention weights are: Among them, F RGB is the RGB feature, F NIR is the NIR feature, W Q and W k is the learnable transformation matrix, d k is a scaling factor used to stabilize the gradient; Generate enhanced feature representation through weighted fusion and calculate weighted features: Among them, W V is the transformation matrix of the values; Indicates A RGB→NIR The transpose of Indicates A NIR→RGB The transpose of The fusion formula is: in, and Represents the fused RGB features and NIR features, and α is a learnable parameter used to control the ratio of the fusion of the two modal features.
6. The fatigue driving detection method based on NIR and RGB image feature fusion as claimed in claim 1, characterized in that: In S2, the fully connected layer includes multiple hidden layers, and the ReLU activation function is used to increase the nonlinear expression ability of the model, while the Dropout layer is used to prevent overfitting.
7. The fatigue driving detection method based on NIR and RGB image feature fusion as claimed in claim 1, characterized in that: In S3, the cross entropy loss function is used as the optimization target, and the weight of each category is determined by the number of samples in the category and the total number of samples, so that the model pays more attention to categories with fewer samples during training and balances the category distribution differences of the data set.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the fatigue driving detection method based on NIR and RGB image feature fusion as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that A computer program is stored, and when the program is executed by a processor, the steps of the fatigue driving detection method based on NIR and RGB image feature fusion according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the steps of the fatigue driving detection method based on NIR and RGB image feature fusion as described in any one of claims 1 to 7.
Citation Information
Cited By
Resistance spot welding quality prediction and optimization method based on multi-modal data and cross-modal attention network fusion
CN121561815A