Driver fatigue detection method based on DLS model
By lightweight transformation of Swin-Transformer and building dual-stream SwinTransformer and space-time fusion model DASTFM, the problem of poor real-time fatigue driving detection in the existing technology is solved, and more efficient fatigue monitoring identification and judgment accuracy is achieved.
Patent Information
- Application Number
- CN202510406676.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-27
AI Technical Summary
The existing fatigue driving detection methods based on neural networks are poor in real time and cannot be applied to high-demand application scenarios such as fast vehicle speeds and complex road conditions.
Using the driver fatigue detection method based on the DLS model, the dual-stream SwinTransformer and the space-time fusion model DASTFM are built to reduce the computational complexity and improve the recognition accuracy.
It improves the real-time and judgment accuracy of fatigue monitoring recognition, can be applied to more scenarios, reduces the computational complexity, and ensures detection accuracy.
Smart Images

Figure CN120220124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fatigue driving detection, and specifically provides a driver fatigue detection method based on the DLS model. Background Art
[0002] Fatigue driving, characterized by symptoms such as yawning or long periods of closed eyes, poses a serious threat to road safety. Developing methods to detect driver fatigue to prevent accidents is crucial. Researchers have classified fatigue detection methods into three main categories: physiological signals, vehicle behavior, and computer vision.
[0003] Physiological and vehicle behavior methods are costly and invasive, and sensor-based devices may impede driving. Additionally, hardware failures and environmental factors can affect the accuracy of experiments and prediction results. With the widespread application of neural networks, some technicians have used the CNN convolutional neural network in computer vision algorithms to detect facial features. However, although computer vision algorithms have improved accuracy, traditional neural networks have difficulties in comprehensive feature extraction and capturing long-term data dependencies due to relying on stacked convolutions and pooling. Although performance has been improved, complexity has increased due to redundant self-attention calculations for smaller classification tasks, resulting in longer judgment times for fatigue driving, making it unsuitable for application scenarios with high real-time requirements for fatigue driving recognition, such as high vehicle speeds and complex road conditions. Summary of the Invention
[0004] To solve the problem of poor real-time performance of existing neural network-based fatigue driving detection methods, the present invention provides a driver fatigue detection method based on the DLS model, which can improve detection accuracy, reduce computational complexity, improve the real-time performance of fatigue monitoring and recognition, and be applicable to more scenarios.
[0005] The technical solution of the present invention is as follows: A driver fatigue detection method based on the DLS model, characterized in that it includes the following steps: S1: Lightweight transformation of Swin-Transformer; Swin-Transformer includes: multiple stages, and each stage includes multiple identical SwinTransformer blocks; Reduce the value of M in the stage on the original basis, and denote the transformed block as: Light SwinBlock; Replace the original Swin Transformer block with the Light Swin Block to obtain the transformed Swin-Transformer; S2: Construct a dual-stream SwinTransformer based on the modified Swin-Transformer; The dual-stream SwinTransformer includes: two modified Swin-Transformers with the same structure in parallel; one Swin-Transformer is used to extract spatial features, and the other is used to extract optical flow features; S3: Construct a spatio-temporal fusion model DASTFM; The spatio-temporal fusion model DASTFM is used to perform multi-modal fusion on the spatial features and optical flow features extracted at the same stage in the two parallel Swin-Transformers. The fusion process is as follows:
[0006] In the formula, s i is the result of multiplying the spatial feature extracted in the i-th stage by the dynamic change factor, t i is the result of multiplying the optical flow feature extracted in the i-th stage by the dynamic change factor; s ie is the enhanced version of the spatial feature, and t ie is the enhanced version of the optical flow feature; (st) 1 ie represents the enhanced representation of the integrated feature with time features as the main features; (st) 2 ie represents the enhanced representation of the integrated feature with spatial features as the main features; α and σ are the dynamic change factors corresponding to the spatial feature and the optical flow feature respectively; (s i ) t represents the enhanced spatial flow feature, and (t i ) s represents the enhanced optical flow feature; F s represents the finally output spatial feature, and F t represents the finally output optical flow feature; F i is the fusion feature output by the spatio-temporal fusion model DASTFM; S4: Construct a DLS model and construct a fatigue driving detection model based on the DLS model; The DLS model includes: a dual-stream SwinTransformer, a spatio-temporal fusion model DASTFM, a hierarchical decoding module, and a classification and judgment module; In the dual-stream Swin Transformer, the spatial features and optical flow features extracted by the Swin-Transformer at the same stage in two parallel Swin-Transformers are fed into the same spatio-temporal fusion model DASTFM for multi-modal fusion; The fused features output by the spatio-temporal fusion model DASTFM are denoted as: F i ; Each F i is respectively fed into the hierarchical decoding module. After the four input layers are hierarchically decoded by the hierarchical decoding module and then feature fusion is performed, the final output feature F final ; The final output feature F final is fed into the classification and judgment module to obtain the classification and judgment result; S5: The fatigue driving detection model is trained through a pre-constructed training set to obtain the trained fatigue driving detection model; S6: Based on the camera, the facial video segment of the driver is captured in real time, denoted as: the video to be recognized; The video to be recognized is decomposed into video frames, denoted as: the video frames to be recognized; The optical flow data is generated from the video frames to be recognized based on the FarneBack algorithm, and the optical flow data is maintained as an event sequence, denoted as: the optical flow data to be recognized; S7: The video frames to be recognized and the optical flow data to be recognized are fed into the trained fatigue driving detection model, and are respectively used as the inputs of two parallel Swin-Transformers. After being calculated and recognized by the fatigue driving detection model, the fatigue driving recognition result is output.
[0007] It is further characterized in that: In the dual-stream Swin Transformer, the number of Swin Transformer Blocks set in four consecutive stages in each Swin-Transformer is respectively: 2, 2, 6, 2; Each Swin block in each stage divides the input feature map based on W-MSA into M×M small windows; the values of M in the four stages are respectively set to: 3, 6, 12, 24; The values of α and σ corresponding to the four states are respectively: α1 to α4 are respectively: 0.8, 0.6, 0.4, 0.2; σ1 to σ4 are respectively: 0.2, 0.4, 0.6, 0.8; The hierarchical decoding module includes: an ASPP model and multiple GSFM models; The output of the last stage outputs the fused F i is directly fed into the said ASPP model, and the remaining F i are respectively fed into a GSFM model; The said ASPP model uses the F corresponding to the last stage i as the input and performs an upsampling operation. During the upsampling process, when the size of the feature map is the same as that of the feature map of a certain stage, the output feature set F4 i is obtained and fed into the GSFM model corresponding to this stage; The GSFM model corresponding to each stage uses F i and F4 i as the input, performs feature fusion based on channels, and outputs the fused feature G i+1 ; The outputs of all GSFM models are concatenated to obtain the final output feature F final ; The calculation process in the said GSFM model is: ; Among them, CS is the Channel Shuffle operation; The representation of the feature set F4 i is: ; The representation of the feature set F i is: ; In the formula, n represents the number of channels; The said classification and judgment module includes: a Linear layer, a Tanh layer, a LayerNorm layer, and a SoftMax layer connected in sequence.
[0008] A driver fatigue detection method based on the DLS model provided by this application lightweightly transforms the traditional Swin-Transformer, reduces the number of small window partitions in the W-MSA, reduces the burden of the multi-scale attention (MSA) mechanism, simplifies the computational requirements in the forward propagation process, and effectively reduces the computational complexity. At the same time, a spatio-temporal fusion model DASTFM is constructed to integrate multi-modal images and features extracted at different scales to obtain richer semantic features. In DASTFM, for the features of images with different resolutions, a calculation formula for enhancing the representation of integrated features at different stages is redesigned, and dynamic change factors α and σ are introduced to solve the problem that feature information becomes increasingly blurred as the image resolution decreases. Through dynamic transformation, DASTFM enables the enhanced representation of integrated features to maximize the spatio-temporal features of the image, thereby improving the recognition accuracy of the model. When constructing the DLS (Dual-Stream Swin-Transformer) model of this method, video frames are extracted from real-time videos, and spatial features and optical flow features are obtained based on the video frames as model inputs. The Farneback optical flow method is combined to obtain the temporal dimension features of the driver by calculating the movement of pixels in the video sequence. Through the lightweight transformation of Swin-Transformer, the computational complexity of the model is reduced. At the same time, based on DASTFM, multi-modal images and features extracted at different scales are integrated to obtain richer semantic features, improving the detection accuracy while reducing the computational complexity, ensuring the real-time performance and judgment accuracy of this method for fatigue driving recognition. Description of the Drawings
[0009] Figure 1 is the system model structure diagram of this application; Figure 2 is the DLS model structure diagram; Figure 3 is the DASTFM model diagram; Figure 4 is the GSFM structure diagram; Figure 5 is the schematic diagram of the channel transformation strategy; Figure 6 is the FLOPs comparison of different models; Figure 7 is the experimental result of the YawDD dataset; Figure 8 is the experimental result of the Tsinghua University - DDD dataset in Taiwan Province, China. Detailed Implementation Manner
[0010] This application includes a driver fatigue detection method based on the DLS model, which includes the following steps.
[0011] Computer vision-based detection methods typically involve placing a camera in front of the driver to capture video clips. After preprocessing the video clips, video frames are obtained. By analyzing each video frame and the transitions between video frames, spatial and temporal features of the driver can be extracted. Then, an intelligent system deeply analyzes these features to identify behavioral patterns that may indicate fatigue during driving. This method can evaluate the driver's fatigue level without interrupting the driver's normal driving activities. The specific process is as Figure 1 shown.
[0012] Specifically, the system analyzes this visual information to collect detailed spatial and temporal features, including the driver's head position, eye movements, and other physiological cues that may indicate fatigue. The intelligent system processes these features to indicate the driver's state. Using computer vision methods can improve driving safety without distracting the driver or causing discomfort.
[0013] Moreover, the extraction effect of driver features will directly affect the final classification and judgment results. Compared with traditional convolutional neural networks, Swin-Transformer realizes the effective transfer and fusion of feature maps at different scales based on the window splitting and shifting strategy. In order to more efficiently process a large number of video frames in real-time tasks, this application constructs a fatigue driving detection model based on Swin-Transformer.
[0014] S1: Lightweight transformation of Swin-Transformer.
[0015] The traditional Swin-Transformer includes multiple stages connected in sequence, and each stage includes multiple identical Swin Transformer blocks.
[0016] Among them, the first stage includes: a Linear Embedding layer and multiple identical SwinTransformer blocks, and the remaining stages include a Patch Partion layer and multiple Swin Transformer blocks.
[0017] Specifically, the input image is first fed into the Patch Partition layer for image segmentation, and then into the Linear Embedding layer in the first stage for feature mapping. After mapping, the data is fed into the Swin Transformer block in the first stage. In each subsequent stage, the input feature map first undergoes downsampling through the Patch Partion layer to generate a hierarchical representation, and then is processed based on the Swin Transformer block in each stage.
[0018] Among them, each Swin Transformer block in each stage simultaneously includes: W-MSA (Window-based Multi-head Self-Attention) and SW-MSA (Shifted Windows Multi-Head Self-Attention). The input feature map is processed by W-MSA and then by SW-MSA.
[0019] W-MSA divides the input feature map to obtain M×M non-overlapping small windows, and independently applies the multi-head self-attention mechanism within each small window to calculate the self-attention weights. This window design reduces the complexity of the model, and can effectively capture information even for the long-range dependencies between small windows after dividing large-sized images.
[0020] However, the recognition metrics for driver fatigue driving are relatively clear. Moreover, when applied to real-time recognition tasks, large-scale image classification tasks based on in-vehicle camera acquisitions usually require a large number of classification parameters, and too many parameters will increase the burden on the multi-scale attention (MSA) mechanism. Therefore, this application optimizes the structure of the swing-transformer to reduce complexity, and lightweight processes the number of detection heads of W-MSA to reduce the computational amount.
[0021] Specifically, each Swin Transformer block in each stage divides the input feature map based on W-MSA into M×M non-overlapping small windows, and independently applies the multi-head self-attention mechanism within each small window. In this application, the value of M in the stage is made smaller on the original basis, simplifying the computational requirements in the forward propagation process. Then the modified block is denoted as: Light Swin Block; replacing the original Swin Transformer block with the Light Swin Block to obtain the modified Swin-Transformer.
[0022] The specific value of M is adaptively adjusted according to the calculation accuracy. In this embodiment, the number of detection heads in the original four stages {4, 8, 16, 32} is streamlined, and finally the values of M in the four stages are set to: {3, 6, 12, 24}.
[0023] At the same time, the structure of the swing-transformer is optimized. Specifically, in the two-stream SwinTransformer, the number of Swin Transformer Blocks set in each consecutive four stages of each Swin-Transformer is set to: 2, 2, 6, 2; that is, in the third stage, the number of stacked blocks of the swin-transformer is reduced from 18 to 6. The depth of the model is fine-tuned. In the transfer learning stage, the model complexity is reduced while not affecting the overall performance of the model.
[0024] S2: Construct a two-stream SwinTransformer based on the modified Swin-Transformer.
[0025] The two-stream SwinTransformer includes: two modified Swin-Transformers with the same structure in parallel; one Swin-Transformer is used to extract spatial features, and the other is used to extract optical flow features.
[0026] The RGB image can capture the spatial details of the driver, while the Farneback optical flow can provide in-depth understanding of the driver's actions and changes over time. By constructing a two-stream SwinTransformer in this application to integrate these two types of features, the network's understanding of global and detailed features can be enhanced.
[0027] S3: Construct a spatio-temporal fusion model DASTFM (Dynamic Allocation Spatial–Temporal Fusion Model).
[0028] The spatio-temporal fusion model DASTFM is used to perform multimodal fusion on the spatial features and optical flow features extracted at the same stage in the two parallel Swin-Transformers. This multimodal fusion effectively enriches the information available for analysis. As Figure 3 shown, the fusion process is as follows:
[0029] In the formula: s i is the spatial feature s i extracted in the i-th stage and the dynamic change factor α iAssign the result of multiplication to s i After that, input it into DASTFM, t i is the optical flow feature extracted in the i-th stage, t i multiplied by the dynamic change factor σ i Assign the result of multiplication to t i After that, input it into DASTFM. s ie is the enhanced version of the spatial feature, t ie is the enhanced version of the optical flow feature. ECA (Efficient channel attention) adaptively determines the kernel size of one-dimensional convolution and effectively captures the interdependence between channels. This method does not require complex dimensionality reduction and expansion techniques, thus maintaining the efficiency and lightweight features of the network.
[0030] (1 - α)s i +(1 - σ)t i and (α * s i + σ * t i ) represent the initial combination of temporal features and spatial features in the same stage, (st) 1 ie represents the enhanced representation of the integrated feature with temporal features as the main features; (st) 2 ie represents the enhanced representation of the integrated feature with spatial features as the main features. (st) 1 ie and (st) 2 ie also determine the fusion weights of s i and t i expressed based on α and σ for the two features, enabling the network to perform soft selection or weighted averaging between s i and t i . GAP (Global Average Pooling) replaces the spatial information directly with the mean value, resulting in fewer network parameters, preventing overfitting and reducing the computational cost. Moreover, GAP more naturally strengthens the connection between categories and feature maps, which is beneficial to improving the detection accuracy. Among them, α and σ are the dynamic change factors corresponding to the spatial feature and the optical flow feature respectively; the values of α and σ corresponding to the four states are: α1 to α4 are: 0.8, 0.6, 0.4, 0.2; σ1 to σ4 are: 0.2, 0.4, 0.6, 0.8. In the four stage phases, the sum of α i and σ i corresponding to each stage is 1. By adjusting α i and σ iThe specific value of
[0031] (s i ) t represents the enhanced spatial flow feature, and (t i ) s represents the enhanced optical flow feature; during the calculation of (s i ) t , the previously enhanced (st) 1 ie and s ie are convolved element by element, and the spatial flow feature is multiplied by their corresponding weights to further enhance the feature and its weight. The calculation process of (t i ) s is the same as that of (s i ) t , and the same processing is performed on the optical flow feature. GMP (Global Max Pooling) extracts the maximum value for each channel of the feature map, resulting in a one-dimensional vector with the same number of channels as the feature map. GMP can highlight the most significant information in the feature map, thereby reducing the overfitting phenomenon to a certain extent. At the same time, due to the global maximum pooling operation being independent for each channel, the robustness is better.
[0032] F s represents the final output spatial feature, and F t represents the final output optical flow feature; F i is the fusion feature output by the spatio-temporal fusion model DASTFM. During the feature fusion process, first, ECA is used to capture the mutual dependencies between channels, then GAP is used to reduce the parameters to prevent overfitting, and finally, GMP is used to highlight the most significant features of each channel. This method designs the operation sequence and operation content for each channel to ensure that the computational amount can be effectively reduced and the detection accuracy can be improved.
[0033] In this method in DASTM, for the features of images with different resolutions, the calculation formula for the enhanced representation of the integrated features in different stages is redesigned, and the dynamic change factors α and σ are introduced to solve the problem that the feature information becomes increasingly blurred as the image resolution decreases. Through dynamic transformation, it can be ensured that the enhanced representation of the integrated features can maximize the spatio-temporal features of the image, thereby improving the recognition accuracy of the model.
[0034] S4: Construct a DLS (Dual-Stream Swin-Transformer) model, and build a fatigue driving detection model based on the DLS model.
[0035] The structure of the DLS model is as Figure 1As shown in the figure, the model includes: a two-stream Swin Transformer, a spatio-temporal fusion model DASTFM, a hierarchical decoding module, and a classification and judgment module. Figure 1 In the [reference], the PatchEmbedding operation and the Linear Embedding operation in the original model are combined and represented as a Patch Merging layer.
[0036] In the two-stream Swin Transformer, the spatial features s i and the optical flow features t i extracted by the stages at the same stage in the two parallel Swin-Transformers are respectively i multiplied by the corresponding dynamic change factors α i and σ i and the original spatial features s i and the optical flow features t Figure 1 are replaced. In [reference], the spatial feature output by the first stage is s1, the optical flow feature output is t1, there are a total of 4 stages, the output spatial features are s1~s4, and the temporal features are t1~t4.
[0037] s i is multiplied by α i , and the value of s i *α i is assigned to s i , t i is multiplied by σ i , and the value of t i *σ i is assigned to t i , and then s i and t i are sent into the corresponding DASTFM. The fused features output by the spatio-temporal fusion model DASTFM are denoted as: F i ; the fused features output by the DASTFM corresponding to the four stages are F1~F4.
[0038] Each F i is respectively sent into the hierarchical decoding module. After the hierarchical decoding module decodes the four inputs hierarchically, final fusion is performed to obtain the final output feature F final ; the final output feature F final is sent into the classification and judgment module to obtain the final classification and judgment result.
[0039] This application starts by parsing the video frame by frame, directing individual frame images to an enhanced swing-transformer to extract spatial features. Meanwhile, the original video data is sent to an improved FarneBack module to generate optical flow data, which preserves the event sequence and is then processed by the same swing-transformer to extract temporal features. The fusion of these spatio-temporal features, refined by a dynamically allocated spatio-temporal fusion model (DASTFM) and a global saliency fusion model (GSFM), is used for driver state classification.
[0040] Specifically, the main role of the Patch Merging layer is to perform downsampling, reduce the resolution of the feature map, and adjust the number of channels, thus forming a hierarchical design. Patch Embedding is the process of dividing an image into multiple small patches and mapping each patch to a high-dimensional vector (embedding). The main role of PatchEmbedding is to convert the original two-dimensional image data into a one-dimensional vector sequence, so that the model can better process the image data. By dividing the image into small patches and mapping each patch to a high-dimensional vector, these vectors can then be input into the Transformer model for processing. This processing method enables the model to better capture the local features in the image.
[0041] Among them, the dynamic change factors α1-α4 are set to 0.8, 0.6, 0.4, and 0.2 respectively, and the dynamic change factors σ1-σ4 are 0.2, 0.4, 0.6, and 0.8 respectively. The reason for setting α and σ in this application is that in the Swin-Transformer model, as the image is downsampled, the feature information of ordinary pictures and the feature information of Farneback optical flow are in opposite states, that is, the feature information of ordinary pictures becomes more and more blurred as the downsampling progresses; while the feature information of Farneback optical flow becomes more and more specific as the downsampling progresses. Therefore, this application designs a dynamically transformed model based on the dynamic change factors to solve this problem.
[0042] Large-scale image classification tasks usually require a large number of classification parameters. The driver fatigue index is relatively clear, and too many parameters will increase the burden on the multi-scale attention (MSA) mechanism. This application optimizes the structure of the swing-transformer to reduce the complexity. Specifically, the number of stacked blocks in the swing-transformer in the third stage is reduced from 18 to 6, and the number of detection heads has been streamlined from the initial {4, 8, 16, 32} to a more efficient configuration {3, 6, 12, 24}, simplifying the computational requirements in the forward propagation process.
[0043] The hierarchical decoding module in this application includes: an ASPP model and multiple GSFM models; the fused F4 output by the last stage is directly fed into the ASPP model, and the remaining F1~F3 are respectively fed into a GSFM model; the ASPP model takes F4 as the input and performs an upsampling operation. During the upsampling process, when the size of the feature map is the same as that of the feature map of a certain stage, the feature set F4 is output i , and is fed into the GSFM model corresponding to this stage; the GSFM model corresponding to each stage takes F i and F4 i as the input, performs feature fusion based on the channels, and outputs the fused feature G' i+1 ; the outputs of all GSFM models are concatenated to obtain the final output feature F final .
[0044] As Figure 4 shown, the calculation process in the GSFM model is: ; where CS is the Channel Shuffle operation of channel transformation; The channels of the features obtained by initial fusion are different in each stage and contain different semantic information. In order to integrate them into the target features, this application uses the GSFM module to process the features of each stage. The features output by DASTFM are used as the input of GSFM. In GSFM, they are concatenated along the channel dimension to form F c , which integrates the feature information of two layers. After that, the features are successively passed through the CS (Channel Shuffle) channel transformation, global multi-scale pooling (GMP), and then multiplied element by element with the original features (F i or F4 i ) to obtain G' P and G'q, and then G' P and G'q are added element by element to generate the final fused feature G' i+1 of this stage, and this feature G' i+1 is then input into the next stage.
[0045] The channel transformation strategy adopted in this application is as Figure 5 shown. The representation of the feature set F4 i is: ; The representation of the feature set F i is: ; In the formula, n represents the number of channels.
[0046] The output is rearranged as: , in this application, the channel transformation strategy is designed based on the principle of permutation and combination to ensure a more thorough mixing and switching function. The channel transformation operation mixes the characteristics at the channel level, which helps the network better understand and utilize this information, thereby improving the overall accuracy.
[0047] The classification and judgment module in this application includes: a Linear layer, a Tanh layer, a LayerNorm layer, and a SoftMax layer connected in sequence. The final output feature F final Undergoes linear processing through the Linear layer, introduces non-linear features through the Tanh layer, and then undergoes normalization activation through the LayerNorm layer. Finally, the final classification result is generated through the SoftMax layer.
[0048] S5: Train the fatigue driving detection model through a pre-constructed training set to obtain a trained fatigue driving detection model.
[0049] The construction and selection of the specific training set and validation set, as well as the training process, are all implemented based on existing technologies.
[0050] S6: Based on the camera, capture the facial video segments of the driver in real time, denoted as: the video to be recognized; Decompose the video to be recognized into video frames, denoted as: the video frames to be recognized; Generate optical flow data from the video frames to be recognized based on the FarneBack algorithm, and keep the optical flow data in an event sequence, denoted as: the optical flow data to be recognized.
[0051] S7: Send the video frames to be recognized and the optical flow data to be recognized into the trained fatigue driving detection model, and use them as the inputs of two parallel Swin-Transformers respectively to obtain the recognition result output by the fatigue driving detection model.
[0052] This application designs a dynamically allocated spatio-temporal fusion model (DASTFM). DASTFM dynamically fuses spatio-temporal features, and GSFM fuses features from different layers of DASTFM to improve the detection efficiency. The simulation results show that the accuracy of DLS (Dual-Stream Swin-Transformer) in this application has increased by 0.33%, and the computational complexity has decreased by 49.3%. The running time of each test epoch of DLS has decreased by 33.1%.
[0053] To verify the performance of the model in this application, the following experimental content is constructed.
[0054] In the experiments of this application, a high-performance server was used, equipped with 64GB of memory, a 16-core AMD Ryzen 9 5950X processor, and two graphics cards, GeForce RTX 4060Ti and GeForce RTX 3060. The software environment of this application includes CUDA 11.6 and Python 3.10, which helps this application perform computationally intensive tasks.
[0055] In the training and prediction phases, two comprehensive video datasets were used: Tsinghua University in Taiwan, China - DDD and YaWDD. These datasets were carefully preprocessed to extract frame images and optical flow images, which are crucial for the analysis of this application. To ensure a thorough evaluation of the model of this application, the datasets were carefully divided into a training set, a test set, and a validation set in a ratio of 7:2:1.
[0056] The model parameters of DLS are shown in Table 1, including parameters such as input image size, training epochs, learning rate, learning optimizer, etc. In Table 1, Parameters is the parameter name and Value is the parameter value.
[0057] Table 1: Experimental Parameters
[0058] The DADCNN model was proposed in the paper "Efficient driver anomaly detection via conditional temporal proposal and classification network" by Lang Su, Chen Sun, Dongpu Cao, and Amir Khajepour. IEEE Transactions on Computational Social Systems, 10(2): 736–745, 2023.
[0059] The swwin - res model was proposed in "Fpirst: Fatigue driving recognition method based on feature parameter images and a residual swin transformer" by Weichu Xiao, Hongli Liu, Ziji Ma, Weihong Chen, and Jie Hou. Sensors, 24(2), 2024.
[0060] Jing Bai, Wentao Yu, Zhu Xiao, Vincent Havyarimana, Amelia C. Regan, Hongbo Jiang, and Licheng Jiao proposed 2s-STGCN in the paper "Two-stream spatial–temporal graph convolutional networks for driver drowsiness detection. IEEE Transactions on Cybernetics, 52(12):13821–13833, 2022."
[0061] In this experiment, the above three models were used together with the DLS model in this application for experiments to evaluate the performance of the models. The accuracy (Acc), recall, and F1 score of different models are shown in Table 2. Since the original code for accurately replicating the experimental environment could not be obtained, a comparison model was created according to the structure and parameters in the original literature. The research focus of this application is to achieve the lightweight of the model while ensuring accuracy. The model in this application achieved an accuracy of 96.78%.
[0062] Table 2: Experimental results
[0063] Based on the results in Table 2, it can be seen that the DLS model proposed in this application has good performance. Compared with other models, the accuracy of the DLS model has increased by 6.34%, 6.27%, and 0.33% respectively.
[0064] Figure 6 The floating-point operations per second (FLOPs) of various models are shown. After optimization, the FLOPs of the DLS model in this application reached 18.19 GFLOPs, achieving lightweight model calculations while ensuring the accuracy of the model.
[0065] The results of driver drowsiness detection are as Figure 7 and Figure 8 shown, which shows the visualization and analysis results of the feature maps of the DLS and other comparison models. The proposed DLS model can effectively extract key features. In contrast, the feature extraction of other models is relatively scattered and cannot clearly reflect the state of the driver. Therefore, the DLS model in this application has good feature recognition and extraction capabilities.
[0066] This application evaluates the impact of the DASTFM and GSFM modules on the performance of the DLS model through ablation experiments. Table 3 shows three cases: (1) Using both the DASTFM module and the GSFM module in the DLS model (DLS in the table), (2) Not using the DASTFM module in the DLS model (Without STFM in the table), (3) Not using the GSFM module in the DLS model (Without GSFM in the table); The mean accuracy (mAcc) and mean average precision (mAp) of the DLS model. The results show that these two modules are indispensable for improving accuracy. To avoid display problems, DASTFM in this application is represented as STFM in Table 3.
[0067] Table 3: Results of ablation experiments Based on the data in Table 3, it can be seen that the DASTFM module is the most effective in improving the accuracy of the DLS model, while the impact of the GSFM module on improving the accuracy of the DLS model is relatively limited. This is mainly because the DASTFM module can effectively integrate features at different stages, while the role of the GSFM module is to further optimize the details.
[0069] The dual-stream lightweight SwinTransformer (DLS) model for detecting driver fatigue proposed in this method combines the Farneback optical flow method to obtain the temporal features of the driver by calculating the motion of pixels in the video sequence. Based on DASTFM, the features extracted from multi-modal images and different scales are integrated to obtain richer semantic features. This method optimizes the structure of the swing-transformer to reduce complexity, and the optimization of the structure does not affect the overall performance of the model. The DLS model reduces the computational complexity during the forward propagation process by reducing the total number of parameters. The simulation results show that compared with other models, the DLS model framework maintains comparable accuracy while reducing approximately 49.3% of the floating-point operations.
Claims
1. A driver fatigue detection method based on DLS model, characterized in that: It includes the following steps: S1: Lightweight transformation of Swin-Transformer; Swin-Transformer includes: multiple stages, each stage includes multiple identical SwinTransformer blocks; The value of M in stage is reduced from the original value, and the modified block is recorded as: Light Swin Block; Light Swin Block replaces the original Swin Transformer block to obtain the modified Swin-Transformer; S2: Building a two-stream SwinTransformer based on the modified Swin-Transformer; The dual-stream SwinTransformer includes: two modified Swin-Transformers with the same structure connected in parallel; one Swin-Transformer is used to extract spatial features, and the other is used to extract optical flow features; S3: Construct the spatiotemporal fusion model DASTFM; The spatiotemporal fusion model DASTFM is used to perform multimodal fusion of spatial features and optical flow features extracted at the same stage in two parallel Swin-Transformers. The fusion process is as follows: ; ; ; ; ; ; In the formula, s i is the result of multiplying the spatial features extracted in the i-th stage by the dynamic change factor. t i It is the result of multiplying the optical flow features extracted in the i-th stage by the dynamic change factor; s ie is the spatial feature enhanced version, t ie It is an enhanced version of the optical flow feature; (st) 1 ie It represents the enhanced representation of the integrated features with time features as the main features; (st) 2 ie It represents the enhanced representation of integrated features with spatial features as the main features; α and σ are the dynamic change factors corresponding to the spatial features and optical flow features respectively; (s i ) t represents the enhanced spatial flow feature, (t i ) s Represents enhanced optical flow features; F s Represents the spatial features of the final output, F t Optical flow features representing the final output; F i It is the fusion feature output by the spatiotemporal fusion model DASTFM; S4: construct a DLS model, and build a fatigue driving detection model based on the DLS model; The DLS model includes: a dual-stream SwinTransformer, a spatiotemporal fusion model DASTFM, a hierarchical decoding module and a classification judgment module; In the dual-stream SwinTransformer, the spatial features and optical flow features extracted by the same stage in the two parallel Swin-Transformers are sent to the same spatiotemporal fusion model DASTFM for multimodal fusion; The fusion feature output by the spatiotemporal fusion model DASTFM is recorded as: i ; Each F i The four input layers are respectively sent to the layered decoding module, and the layered decoding module decodes the four input layers and then performs feature fusion to obtain the final output feature F final ; The final output feature F final Send it to the classification judgment module to obtain the classification judgment result; S5: training the fatigue driving detection model using a pre-constructed training set to obtain the trained fatigue driving detection model; S6: The facial video clip of the driver is captured in real time by the camera, which is recorded as the video to be recognized; Decomposing the video to be identified into video frames, which are recorded as: video frames to be identified; The video frame to be identified generates optical flow data based on the FarneBack algorithm, and the optical flow data is kept as an event sequence, which is recorded as: optical flow data to be identified; S7: Send the video frame to be identified and the optical flow data to be identified into the trained fatigue driving detection model as inputs of two parallel Swin-Transformers respectively. After calculation and identification by the fatigue driving detection model, the fatigue driving identification result is output.
2. The driver fatigue detection method based on the DLS model according to claim 1, characterized in that: In the dual-stream SwinTransformer, the numbers of SwinTransformer Blocks set in four consecutive stages in each Swin-Transformer are 2, 2, 6, and 2 respectively.
3. The driver fatigue detection method based on the DLS model according to claim 1 is characterized in that: The Swin block in each stage divides the input feature map into M×M small windows based on W-MSA; the values of M in the four stages are set to 3, 6, 12, and 24, respectively.
4. The driver fatigue detection method based on the DLS model according to claim 1 is characterized in that: The values of α and σ corresponding to the four states are: α1~α4 are: 0.8, 0.6, 0.4, 0.2 respectively; σ1~σ4 are 0.2, 0.4, 0.6, and 0.8 respectively.
5. The driver fatigue detection method based on the DLS model according to claim 1 is characterized in that: The layered decoding module includes: an ASPP model and multiple GSFM models; The last stage outputs the fused F i directly fed into the ASPP model, and the remaining F i Each is fed into a GSFM model; The ASPP model is based on the F corresponding to the last stage i As input, upsampling operation is performed. During the upsampling process, when the feature map size is the same as the feature map size of a certain stage, the output feature set F4 i , and sent to the GSFM model corresponding to this stage; The GSFM model corresponding to each stage is represented by F i and F4 i As input, feature fusion is performed based on channels, and fusion feature G' is output i+1 ; The outputs of all GSFM models are concatenated to obtain the final output feature F final .
6. The driver fatigue detection method based on the DLS model according to claim 5 is characterized by: The calculation process in the GSFM model is: ; Among them, CS is the Channel Shuffle operation; Feature Set F4 i The expression is: ; Feature Set F i The expression is: ; In the formula, n represents the number of channels.
7. The driver fatigue detection method based on the DLS model according to claim 1 is characterized by: The classification judgment module includes: a Linear layer, a Tanh layer, a LayerNorm layer and a SoftMax layer connected in sequence.