Multi-scale Spatiotemporal Feature Fusion Fatigue Driving Detection Method for Large Model Prompt Words

Through the multi-scale spatiotemporal feature fusion method of large model prompt words and mixed expert models, the problem of insufficient detection accuracy of fatigue driving in the existing technology is solved, and more efficient and accurate fatigue driving detection is achieved.

CN119832530BActive Publication Date: 2025-06-20HUNAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510300526.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing fatigue driving detection methods are difficult to fully capture the subtle changes in the driver's facial expressions, and lack an efficient spatio-temporal motion feature modeling mechanism, resulting in low detection accuracy.

Method used

The detection method of multi-scale spatiotemporal and spatial characteristics fusion of large-scale prompt words is used to generate prompt words corresponding to fatigue-related expression motion units through pre-training the big model, combined with the mixed expert model design and the cascade structure of space-time Transformer, multi-scale spatiotemporal and spatial characteristics are extracted, and comprehensive analysis is performed through a lightweight Transformer network.

Benefits of technology

It achieves higher fatigue driving detection accuracy, can fully capture the subtle changes in the driver's facial expressions, and efficiently model the spatial and temporal characteristics, providing an efficient and accurate fatigue driving detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832530B_ABST
    Figure CN119832530B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of fatigue driving detection, and specifically relates to a fatigue driving detection method for multi-scale spatio-temporal feature fusion of large model prompt words. This method uses a pre-trained large model to generate corresponding prompt words and extract semantic embedding vectors of the prompt words. Then, it performs multi-scale sampling on the video frames input into the mixture of experts model, cascades the prompt word features with the channels of each scale video sequence, fuses the multi-scale features to generate a micro-expression spatio-temporal motion feature map, and inputs it into a lightweight Transformer network. Combining with the dynamically updated comprehensive learning vector, it outputs the fatigue probability through a classifier, judges whether it is fatigue driving based on the threshold warning, and finally uses the cross-entropy loss function combined with an optimizer to train the mixture of experts model. This method can comprehensively capture the subtle changes in the driver's facial expressions, and through the guidance of large model prompt words, efficiently model the spatio-temporal motion features, providing an efficient and accurate solution for vehicle driver fatigue driving detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fatigue driving detection, and particularly relates to a fatigue driving detection method for multi-scale spatio-temporal feature fusion of large model prompt words. Background Art

[0002] Vehicles are one of the most important means of transportation in the world today, and fatigue driving has a great impact on the safe driving of vehicles and social public traffic order. According to statistics, fatigue driving will cause a significant decline in the driver's reaction speed and seriously damage the judgment. The economic losses and casualties caused by fatigue driving traffic accidents in China are relatively high every year.

[0003] Using facial expression recognition to automatically detect fatigue driving is one of the important artificial intelligence fatigue driving detection methods at present. However, existing methods mainly extract time features or single-scale spatial features through traditional methods, such as existing patents with publication numbers CN108053615B, CN118887810B, CN117789181B, and existing patent applications with publication number CN118097636A. These solutions are difficult to comprehensively capture the subtle changes in the driver's facial expressions and lack an efficient spatio-temporal motion feature modeling mechanism, resulting in low detection accuracy. Summary of the Invention

[0004] The present invention provides a fatigue driving detection method for multi-scale spatio-temporal feature fusion of large model prompt words. This method realizes higher detection accuracy through large model prompt word-guided multi-scale comprehensive spatio-temporal feature learning and analysis. The method includes the following steps:

[0005] S1. Define expression movement units related to fatigue according to the Facial Action Coding System, use a pre-trained large model to generate prompt words corresponding to the expression movement units, extract semantic embedding vectors of the prompt words, and generate a prompt word feature matrix P ;

[0006] S2. Design and process the mixture of experts model:

[0007] S2.1. Perform multi-scale sampling on the video frame sequence input to the mixture of experts model S to generate L different-scale video sequences S l ; Concatenate the prompt word feature matrix P with each scale video sequence S l at the channel dimension to realize the fusion of semantic embedding vectors and motion features, and generate a fused multi-scale feature sequence PS l ;

[0008] S2.2. Extract multi-scale spatial features PS l and multi-scale spatio-temporal features F l from each fused feature sequence; G l

[0009] S2.3. Upsample the spatio-temporal features of different scales G l to a unified spatial size, cascade them, and integrate them through a convolutional layer to generate a micro-expression spatio-temporal motion feature map M ;

[0010] S2.4. Use a lightweight Transformer network to comprehensively analyze the spatio-temporal motion features of the micro-expression spatio-temporal motion feature map M and update the initialized comprehensive learning vector Y to Y’ ;

[0011] S3. Based on Y’ output the fatigue state probability through a fully connected layer classifier, and combine a preset threshold to determine whether the driver is in a fatigue state;

[0012] S4. Use the cross-entropy loss function combined with an optimizer to train the mixture of experts model.

[0013] In a specific embodiment, in step S1, the expression movement unit includes at least one of eyelid closure, yawning, and head tilt, and the pre-trained large model is CLIP or GPT .

[0014] In a specific embodiment, in step S1, the shape of the prompt word feature matrix P is N , C P , where N is the number of prompt words, and C P is the feature channel number of each prompt word.

[0015] In a specific embodiment, in step S2.1, multi-scale sampling of the input video frame sequence S includes sampling at least three scales of 1, 1 / 4, 1 / 8, 1 / 16, 1 / 32 of the original image size; generating L different-scale video sequences S l in which L ≥3, and each scale video sequence S lThe shape is T , H l , W l , C S , T is the number of frames, H l 、 W l are the spatial dimensions, C S is the original number of channels of the input video frames.

[0016] In a specific embodiment, in step S2.2, spatial Transformer is used to extract multi-scale spatial features F l , and the shape of the multi-scale spatial features F l is T , H l , W l , C F . Temporal Transformer is used to extract multi-scale spatio-temporal features G l , and the shape of the multi-scale spatio-temporal features G l is T , H l , W l , C G .

[0017] In a specific embodiment, the specific operation of step S2.3 is: first, bilinear interpolation upsampling is performed on the spatio-temporal features G l at each scale to the maximum scale spatial dimension, then the upsampled features are cascaded in the time dimension to generate cascaded features, and finally the number of channels of the cascaded features is compressed through a 3×3 convolutional layer to obtain the micro-expression spatio-temporal motion feature map M .

[0018] In a specific embodiment, in step S2.4, the initial shape of the comprehensive learning vector Y is C M , and the output Transformer of the lightweight Y’ network retains the same channel dimension as Y .

[0019] In a specific embodiment, in step S2.5, the preset threshold is 0.5 - 0.8, and when the fatigue probability output by the classifier exceeds the threshold, a warning signal is triggered.

[0020] In a specific embodiment, the specific operation of step S4 is as follows:

[0021] a) Perform data augmentation on the input video frame sequence S to generate training samples;

[0022] b) Input the training samples into the mixture of experts model to generate a comprehensive learning vector Y’ ;

[0023] c) Calculate the cross - entropy loss between the output of the classifier and the true label;

[0024] d) Based on the calculated cross - entropy loss, use AdamW the optimizer to backpropagate and update the network parameters of the mixture of experts model, and combine the weight decay strategy to prevent overfitting;

[0025] e) Repeat steps a) - d) until the mixture of experts model converges.

[0026] In a specific embodiment, in step a), the data augmentation includes at least one of random cropping, horizontal flipping, and color jitter; in step d), the cosine annealing learning rate scheduling method is used to dynamically adjust the learning rate; in step d), use Dropout and the weight decay technique to further prevent overfitting.

[0027] The present invention has at least the following beneficial effects:

[0028] 1. Through the expression movement unit prompt word learning mechanism, combined with the semantic understanding ability of the pre - trained large model, the micro - expression movement features related to fatigue are accurately extracted, solving the problem of insufficient sensitivity of traditional methods to subtle expression changes.

[0029] 2. The mixture of experts model is designed for multi - scale sampling and is combined with the spatial - temporal Transformer cascade structure to synchronously extract local details and global context information, enhancing the adaptability to expression actions of different resolutions and overcoming the limitations of single - scale feature modeling.

[0030] 3. By cascading the semantic embedding vector of the prompt word with the multi - scale video sequence sampled by the mixture of experts model S l realize the fusion of the semantic embedding vector and the motion features, and enhance the spatio - temporal features. Combined with the dynamic update of the comprehensive learning vector by the lightweight Transformer network, the representation ability of fatigue - related features is improved, solving the problem of spatio - temporal motion feature fragmentation in traditional methods.

[0031] 4. The mixture of experts model adopts a phased feature fusion strategy that integrates upsampling and convolution, and uses a lightweight Transformer network design to retain key information while reducing computational redundancy.

[0032] In summary, the method of the present invention guides the mixture of experts model based on time, multi-scale space, and micro-expression action units for fatigue driving detection through pre-training of large model expression action unit prompt word learning, which can comprehensively capture the subtle changes in the driver's facial expressions and efficiently model spatio-temporal features, providing an efficient and accurate solution for vehicle driver fatigue driving detection. Brief Description of the Drawings

[0033] Figure 1 is the flowchart of the method of the present invention. Detailed Embodiment

[0034] Please refer to Figure 1 , a large model prompt word multi-scale spatio-temporal feature fusion fatigue driving detection method provided by the present invention specifically includes the following steps:

[0035] S1. Define the expression action units related to fatigue according to the Facial Action Coding System FACS . Among them, the expression action units include at least one of eyelid closure, yawning, and head tilt. Use the pre-trained large model to generate prompt words corresponding to the expression action units, such as eye closure, yawning, head drooping, etc., and extract the semantic embedding vectors of the prompt words to generate a prompt word feature matrix P so as to convert the above-mentioned prompt words into a numerical vector representation form suitable for machine learning models. Among them, the pre-trained large model can adopt CLIP , or can also adopt GPT . The shape of the prompt word feature matrix P is N , C P , where N is the number of prompt words, and C P is the number of feature channels for each prompt word.

[0036] S2. Mixture of experts model processing:

[0037] S2.1. Perform multi-scale sampling on the video frame sequence S input into the mixture of experts model to capture the detailed information of the video frame sequence S at different resolutions, and generate L different-scale video sequences S l . Among them, the video frame sequenceS is a pre - processed video frame. The pre - processing includes existing basic pre - processing operations such as normalizing the video frame size, normalizing, and color space conversion. The multi - scale sampling includes at least three scales of 1, 1 / 4, 1 / 8, 1 / 16, 1 / 32 of the original image size. L ≥3, each scale video sequence S l has a shape of T , H l , W l , C S , T is the number of frames, H l , W l are the height and width, and C S is the original number of channels of the input video frame. Concatenate the prompt word feature matrix P with each scale video sequence S l at the channel dimension level, so that at each scale, fuse the previously obtained semantic embedding vector with the motion features of the video frame sequence to generate a fused multi - scale feature sequence PS l . This allows the mixture - of - experts model to consider specific signs of fatigue motion when analyzing each frame.

[0038] S2.2. For each fused feature sequence PS l , use spatial Transformer to extract multi - scale spatial features F l . The multi - scale spatial features F l has a shape of T , H l , W l , C F to capture local patterns in the image and their relationships, and use temporal Transformer to extract multi - scale spatio - temporal features G l . The multi - scale spatio - temporal features G l has a shape of T , H l , W l , C G, so as to facilitate the extraction of dynamic change information in the time dimension, which is crucial for identifying subtle facial expression changes over time.

[0039] S2.3. First, for the spatio-temporal features at each scale G l perform bilinear interpolation upsampling to a unified spatial size, for example, upsampling to the spatial size of the largest scale, and then concatenate the upsampled features in the time dimension to generate concatenated features. Compress the number of channels of the concatenated features through a 3×3 convolutional layer to integrate information from different scales, and obtain the micro-expression spatio-temporal motion feature map M .

[0040] S2.4. Input the micro-expression spatio-temporal motion feature map M into the lightweight Transformer network, and introduce a randomly initialized comprehensive learning vector Y in the input. The initial shape of the comprehensive learning vector Y is C M . The lightweight Transformer network comprehensively analyzes the spatio-temporal features of the micro-expression spatio-temporal motion feature map M and updates the comprehensive learning vector Y to Y’ , and Y’ retains the same channel dimension as Y . This step deeply learns the fused features through a lightweight Transformer network, updates the comprehensive learning vector Y , and strengthens the understanding of the features of time, space, and facial expression motion units by the mixture of experts model.

[0041] Y’ S3. Based on Y’ output the fatigue state probability through a fully connected layer classifier, and combine a preset threshold to determine whether the driver is in a fatigue state. For example, set the preset threshold to 0.5 - 0.8. When the fatigue probability output by the classifier exceeds this threshold, trigger a warning signal to determine that the driver is in a fatigue state.

[0042] S4. Use the cross-entropy loss function combined with an optimizer to train the mixture of experts model. The specific operations are as follows:

[0043] a) Perform data augmentation on the input video frame sequence S to generate training samples. Among them, data augmentation includes at least one of random cropping, horizontal flipping, and color jittering to improve the generalization ability of the model.

[0044] b) Input the training samples into the mixture of experts model to generate a comprehensive learning vector Y’ .

[0045] c) Calculate the cross-entropy loss between the output of the classifier and the true labels.

[0046] d) Based on the calculated cross-entropy loss, update the network parameters of the mixture-of-experts model through backpropagation using an AdamW optimizer, and combine the weight decay strategy to prevent overfitting. In this step, the cosine annealing learning rate scheduling method is used to dynamically adjust the learning rate of the mixture-of-experts model to improve training stability, and Dropout and the weight decay technique are used to further prevent overfitting.

[0047] e) Repeat steps a)-d) until the mixture-of-experts model converges.

[0048] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions and substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. A large-model prompt word multi-scale spatiotemporal feature fusion fatigue driving detection method, characterized in that: The steps include: S1. Define the expression movement units related to fatigue according to the facial action coding system, use the pre-trained large model to generate the prompt words corresponding to the expression movement units, extract the semantic embedding vector of the prompt words, and generate the prompt word feature matrix P ; S2. Hybrid expert model design and processing: S2.

1. Video frame sequence for input hybrid expert model S Perform multi-scale sampling to generate L Video sequences of different scales S l ; The prompt word feature matrix P With each scale video sequence S l Perform channel dimension cascading to achieve the fusion of semantic embedding vector and motion feature, and generate a fused multi-scale feature sequence PS l ; S2.

2. For each fused feature sequence PS l , extract multi-scale spatial features F l and multi-scale spatiotemporal features G l ; S2.

3. Spatiotemporal features of different scales G l Upsample to a uniform spatial size, cascade and integrate through convolutional layers to generate micro-expression spatiotemporal motion feature maps M ; S2.4, Use lightweight Transformer Network-based spatiotemporal motion feature map of micro-expressions M The spatiotemporal motion characteristics are comprehensively analyzed and the initialized comprehensive learning vector Y Updated to Y’ ; S3, based on Y’ The fatigue state probability is output by the fully connected layer classifier, and combined with the preset threshold, it is determined whether the driver is in a fatigue state; S4. Use the cross entropy loss function combined with the optimizer to train the hybrid expert model.

2. The fatigue driving detection method based on large model prompt word multi-scale spatiotemporal feature fusion according to claim 1 is characterized in that: In step S1, the expression movement unit includes at least one of eyelid closure, yawning, and head tilt, and the pre-trained large model is CLIP or GPT .

3. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 1 is characterized in that: In step S1, the prompt word feature matrix P The shape is [ N , C P ],in N is the number of prompt words, C P The number of feature channels for each cue word.

4. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 2 is characterized in that: In step S2.1, the input video frame sequence S The multi-scale sampling includes sampling at least three scales of 1, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image size; generate L Video sequences of different scales S l middle, L ≥3, video sequence at each scale S l The shape is [ T , H l , W l , C S ], T is the number of frames, H l , W l is the space size, C S is the original number of channels of the input video frame.

5. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 4 is characterized in that: In step S2.2, use the space Transformer Extracting multi-scale spatial features F l , multi-scale spatial features F l The shape is [ T , H l , W l , C F ], usage time Transformer Extracting multi-scale spatiotemporal features G l , multi-scale spatiotemporal features G l The shape is [ T , H l , W l , C G ].

6. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 5 is characterized in that: The specific operation of step S2.3 is: first, the spatiotemporal features of each scale G l Perform bilinear interpolation upsampling to the maximum scale space size, then cascade the upsampled features in the time dimension to generate cascade features, and finally compress the number of cascade feature channels through a 3×3 convolutional layer to obtain the micro-expression spatiotemporal motion feature map M .

7. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 6 is characterized in that: In step S2.4, the comprehensive learning vector Y The initial shape is [ C M ], lightweight Transformer Output of the network Y’ Retention and Y Same channel dimensions.

8. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 7 is characterized in that: In step S2.5, the preset threshold is 0.5-0.8, and a warning signal is triggered when the fatigue probability output by the classifier exceeds the threshold.

9. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 7 is characterized in that: The specific operations of step S4 are: a) Input video frame sequence S Perform data enhancement and generate training samples; b) Input the training samples into the mixed expert model to generate a comprehensive learning vector Y’ ; c) Calculate the cross entropy loss between the classifier output and the true label; d) Based on the calculated cross entropy loss, AdamW The optimizer back-propagates to update the network parameters of the hybrid expert model and combines the weight decay strategy to prevent overfitting; e) Repeat steps a)-d) until the hybrid expert model converges.

10. The method for detecting fatigue driving by integrating multi-scale spatiotemporal features of large model prompt words according to claim 9 is characterized in that: In step a), the data enhancement includes at least one of random cropping, horizontal flipping, and color jittering; in step d), the learning rate is dynamically adjusted using a cosine annealing learning rate scheduling method ... Dropout And weight decay techniques to further prevent overfitting.

Citation Information

Patent Citations

  • Driver fatigue detection method based on micro-expressions

    CN108053615B

  • Driving fatigue detection method and system based on lightweight neural network image enhancement

    CN117789181B

  • Image vision method and system for automatically detecting driver fatigue based on artificial intelligence

    CN118097636A

  • A continuous drive-through automatic inspection method, system and medium

    CN118887810B

  • Driver fatigue detection method based on CT-Net model

    CN118366134A