A fatigue driving detection method based on large model classification adaptor

By introducing a classification adapter at each stage of the pre-trained large model, optimizing features and performing classification and scoring phase by phase, the problems of low identification accuracy and high training cost in the prior art are solved, and fatigue driving detection with higher accuracy and lower cost are achieved.

CN119832531BActive Publication Date: 2025-05-23HUNAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510312667.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-23
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing fatigue driving detection method based on facial expression recognition has low recognition accuracy, and the cost of training large models from scratch is extremely high, making it difficult to popularize deployment.

Method used

The existing pre-trained large model is adopted and a classification adapter is designed at each stage. Through phase-by-stage feature optimization and classification scoring, the model's adaptability and accuracy in fatigue driving detection tasks are improved.

Benefits of technology

It significantly improves the accuracy of fatigue driving detection, avoids the high training costs brought by re-training large models, and achieves more efficient detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832531B_ABST
    Figure CN119832531B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of fatigue driving detection, and specifically relates to a fatigue driving detection method based on a large model classification adaptor, comprising the following steps: firstly collecting a driver's facial video, converting the video into a static image I . Then the image I Input the pre-trained large model, extract the multi-stage visual feature map, and optimize the visual feature map stage by stage through the classification adaptor. Then use the optimized feature map of the final stage Fl <supgt; ”
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fatigue driving detection, and in particular relates to a fatigue driving detection method based on a large model classification adaptor. Background Art

[0002] In today's society, vehicles are one of the most important means of transportation in my country. Fatigue driving has a great impact on the safe driving of vehicles and the safety of the public environment. Fatigue driving is one of the important causes of traffic accidents. When a driver is driving fatigued, it will seriously affect his reaction speed, judgment and attention, resulting in the driver's inability to respond promptly and accurately when faced with emergencies, thereby increasing the risk of traffic accidents. These accidents may not only cause casualties and property losses, but may also cause more serious social problems. Through fatigue driving detection, the driver's fatigue state can be detected in time, reminding him to rest, thereby effectively avoiding potential dangers and improving driving safety.

[0003] At present, the industry usually uses facial expression recognition technology based on artificial intelligence to determine whether the vehicle driver is driving fatigued. There are two main methods for detecting fatigue driving based on facial expression recognition. The first method relies on traditional artificial intelligence and pattern recognition models for expression recognition, such as the announcement number CN 108053615 B The existing patent collects facial images in normal and fatigue states by designing fatigue stimulation experiments, identifies facial textures under corresponding micro-expressions, and establishes a corresponding micro-expression library of facial texture features; through the minimum distance discriminant function, the vector distance between the current facial texture features and the facial texture features of the micro-expressions in the micro-expression library is calculated to judge the driver's micro-expressions. This type of method has low recognition accuracy due to its limited ability to extract facial expression features and feature-based classification. Another example is the announcement number CN 105769120 B The existing patent and publication number is CN 104881955 B The existing patent of also has this defect. The second is to train a large model from scratch, such as the one in the announcement number CN 118277835 B The existing patented solution has extremely high training costs and is difficult to deploy on a large scale. Summary of the invention

[0004] The present invention provides a fatigue driving detection method based on a large model classification adaptor. On the one hand, the existing pre-trained large model is adopted to greatly save the model training cost. On the other hand, by constructing a classification adaptor and coordinating with each stage of the pre-trained large model, the trained large model can more accurately recognize the driver's expression, thereby more accurately judging whether the driver is fatigued driving.

[0005] The method comprises the following steps:

[0006] S1, collect the driver's facial video through the in-car optical camera, based on rank - pooling The method extracts facial dynamic motion information and converts the video into a static image. I ;

[0007] S2, the image I Input to the local vehicle GPU In the pre-trained large model, a multi-stage visual feature map is extracted, and the visual feature map is optimized stage by stage through a classification adaptor; wherein the pre-trained large model includes L stages, and the visual feature map extracted in each stage is recorded as feature map Fl , l The value range is 1~ L integer;

[0008] S3, classification adaptor to feature map Fl The steps for feature optimization include:

[0009] S3.1. Feature Map Fl Down-sampling is performed and the adaptive feature map is obtained by fine-tuning the convolution layer. Fl ’ ;

[0010] S3.2. Adaptive feature map Fl ’ After two linear layers and sigmoid Activation function processing generation 2 K Category scoring, K is the number of candidate expression categories;

[0011] S3.3, then the adaptive feature map Fl ’ After upsampling and feature map Fl And classification and scoring concatenation, integrated through another linear layer to obtain the optimized feature map Fl ” , and input to the next stage;

[0012] S4. The optimized feature map obtained in the final stage of the pre-trained large model Fl ”It is cascaded with the classification and scoring of the classification adaptor at all stages, and outputs the driver's expression category and fatigue driving judgment results through a multi-layer perceptron;

[0013] S5. Select part of the training data to fine-tune the parameters in step S3 and step S4 to optimize the parameters of the pre-trained large model.

[0014] In a specific embodiment, in step S1, based on rank - pooling The method extracts facial dynamic motion information and converts the video into a static image. I , specifically including: dynamically aggregating features of continuous video frames to generate a single-frame image containing time dimension information.

[0015] In a specific embodiment, the pre-trained large model is a lightweight Transformer Architecture, including DeepSeek - 7B or its derivative models.

[0016] In a specific embodiment, in step S3.1, downsampling uses average pooling or maximum pooling operation, and the convolution layer uses a 1×1 convolution kernel to adjust the channel dimension.

[0017] In a specific embodiment, in step S3.2, a binary probability distribution is generated for each candidate expression category, including the probability of belonging to the category and the probability of not belonging to the category.

[0018] In a specific embodiment, in step S3.3, the upsampling operation uses bilinear interpolation or transposed convolution, and the cascaded feature dimension is mapped through a linear layer to match the input feature dimension of the next stage.

[0019] In a specific embodiment, in step S5, the parameters in step S3 and step S4 are fine-tuned using low-rank adaptation technology.

[0020] Beneficial effects of the present invention: The method of the present invention is based on the existing pre-trained large model, and a classification adaptor is designed at each stage of the large model to process the feature map stage by stage through the classification adaptor. Fl , generate adaptive features Fl’ and 2 K The feature map can help the large model better adapt to the fatigue driving detection task. K The classification scoring can effectively reduce the category retrieval space and improve the accuracy of fatigue driving detection. This method combines multi-stage classification adaptors in cascade, which can achieve progressive, coarse-to-fine classification prediction, thereby generating more accurate fatigue driving detection results, thereby improving the accuracy of fatigue driving detection. Compared with the existing technology, this method improves the detection accuracy while avoiding the problem of extremely high training costs caused by training a large model from scratch. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0022] See also Figure 1 The present invention provides a fatigue driving detection method based on a large model classification adaptor, which includes the following steps.

[0023] S1, collect the driver's facial video through the in-car optical camera, based on rank - pooling The method extracts facial dynamic motion information, including: aggregating dynamic features of continuous video frames to generate a single frame image containing time dimension information, thereby converting the video into a static image I Since the optical flow method needs to calculate the motion vector frame by frame, D The amount of convolution calculation is very large, and the embodiment of the present invention adopts rank - pooling This method extracts timing features in a lightweight manner, which can ensure both processing efficiency and the ability to retain dynamic features, and is more suitable for in-vehicle edge devices.

[0024] S2, the image I Input to the local vehicle GPU In the pre-trained large model, multi-stage visual feature maps are extracted and the features of the visual feature maps are optimized stage by stage through the classification adaptor. I After the first stage of processing of the pre-trained large model, the extracted visual feature map is recorded as the feature map F 1. Feature map F 1 After feature optimization by the classification adaptor, the optimized feature map is obtained F 1 ” , optimize the feature map F 1 ” As the input of the pre-trained large model in the second stage. Assume that the pre-trained large model has L stages, each stage generates a feature map Fl , the classification adaptor generates feature maps for each stage Fl Perform feature optimization to obtain the optimized feature map Fl ” , l The value range is 1~ L integer, and the obtained optimized feature map Fl ” As the input of the next stage. In the embodiment of the present invention, the pre-trained large model adopts lightweight Transformer Architectural DeepSeek - 7B Of course, in other embodiments of the present invention, the pre-trained large model can also be DeepSeek - 7BThe derivative model of .

[0025] S3, classification adaptor to feature map Fl The steps of feature optimization include the following steps.

[0026] S3.1. Feature Map Fl Downsampling is performed, and average pooling or maximum pooling is used to reduce the amount of calculation. Adaptive feature maps are obtained by fine-tuning the convolution layer. Fl ’ , The convolution layer uses a 1×1 convolution kernel to adjust the channel dimension.

[0027] S3.2. Adaptive feature map Fl ’ After two linear layers and sigmoid The activation function is processed to predict the driver's expression category. To ensure the accuracy of classification, a binary classification method is used to generate a binary probability distribution for each candidate expression category, including the probability of belonging to the category and the probability of not belonging to the category. K candidate expression categories, and finally generate 2 K Rating of categories.

[0028] S3.3, then the adaptive feature map Fl ’ Upsampling to feature map Fl Same resolution and same as feature map Fl And classification and scoring cascade, and then integrate these features through another linear layer to obtain the optimized feature map Fl ” , and optimize the feature map Fl ” Input to DeepSeek - 7B The next stage of the large model. In the embodiment of the present invention, the upsampling operation adopts bilinear interpolation, and the cascaded feature dimensions are mapped to match the input feature dimensions of the next stage through a linear layer. Of course, in other embodiments of the present invention, the upsampling operation can also adopt transposed convolution.

[0029] By introducing a classification adaptor after each stage, supervision signals related to expression classification (such as fatigue features such as eye closure and drooping mouth corners) are gradually injected to make DeepSeek - 7B The large model continuously focuses on the key areas of the task during feature extraction, achieving dynamic feature optimization. Fl ’ With feature map Fl The classification and scoring cascade not only retains the general representation ability of the pre-trained large model, but also incorporates task-specific features to prevent the deep features from losing detailed information.

[0030] S4. The optimized feature map obtained in the final stage of the pre-trained large model Fl ” Compared with the classification adaptor in all stages L ×2 K The present invention provides fine-grained supervision signals through classification and scoring at each stage, guides the model to focus on fatigue-related expressions (such as frequent blinking), and realizes multi-level supervision. The fatigue driving judgment result is a binary classification output result, which can be achieved by analyzing the correlation between the expression category and the preset fatigue characteristics (such as the high correlation between the "sleepy" expression and fatigue driving). The present invention cascades multi-level features and classification supervision signals, and uses MLP It makes multi-task decisions and realizes end-to-end driver status analysis, which significantly improves the coordination and accuracy of fatigue detection and expression recognition, while being able to adapt to the real-time requirements of on-board equipment.

[0031] S5. Select part of the training data and use low-rank adaptation technology to fine-tune the parameters newly added to the pre-trained large model, that is, the parameters in step S3 and step S4, to optimize the parameters of the pre-trained large model. The present invention only trains the parameters newly added to the pre-trained large model, and the number of parameters is relatively small, while the parameters of the pre-trained large model itself are fixed and not trained, thereby greatly saving training costs.

[0032] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions and substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A fatigue driving detection method based on a large model classification adaptor, characterized in that: The steps include: S1, collect the driver's facial video through the in-car optical camera, based on Rank-pooling The method extracts facial dynamic motion information and converts the video into a static image. I ; S2, the image I Input to the local vehicle GPU In the pre-trained large model, a multi-stage visual feature map is extracted, and the visual feature map is optimized stage by stage through a classification adaptor; wherein the pre-trained large model includes L stages, and the visual feature map extracted in each stage is recorded as feature map Fl , l The value range is 1~ L integer; S3, classification adaptor to feature map Fl The steps for feature optimization include: S3.

1. Feature Map Fl Down-sampling is performed and the adaptive feature map is obtained by fine-tuning the convolution layer. Fl ’ ; S3.

2. Adaptive feature map Fl ’ After two linear layers and sigmoid Activation function processing generation 2 K Category scoring, K is the number of candidate expression categories; S3.3, then the adaptive feature map Fl ’ After upsampling and feature map Fl And classification and scoring concatenation, integrated through another linear layer to obtain the optimized feature map Fl ” , and input to the next stage; S4. The optimized feature map obtained in the final stage of the pre-trained large model Fl ” It is cascaded with the classification and scoring of the classification adaptor at all stages, and outputs the driver's expression category and fatigue driving judgment results through a multi-layer perceptron; S5. Select part of the training data to fine-tune the parameters in step S3 and step S4 to optimize the parameters of the pre-trained large model.

2. The fatigue driving detection method based on large model classification adaptor according to claim 1 is characterized in that: In step S1, based on Rank-pooling The method extracts facial dynamic motion information and converts the video into a static image. I , specifically including: dynamically aggregating features of continuous video frames to generate a single-frame image containing time dimension information.

3. The fatigue driving detection method based on large model classification adaptor according to claim 1 is characterized in that: The pre-trained large model is lightweight Transformer Architecture, including DeepSeek-7B or its derivative models.

4. The fatigue driving detection method based on large model classification adaptor according to claim 1 is characterized in that: In step S3.1, downsampling uses average pooling or maximum pooling operation, and the convolution layer uses a 1×1 convolution kernel to adjust the channel dimension.

5. The fatigue driving detection method based on large model classification adaptor according to claim 1 is characterized in that: In step S3.2, a binary probability distribution is generated for each candidate expression category, including the probability of belonging to the category and the probability of not belonging to the category.

6. The fatigue driving detection method based on large model classification adaptor according to claim 1 is characterized in that: In step S3.3, the upsampling operation uses bilinear interpolation or transposed convolution, and the cascaded feature dimensions are mapped through a linear layer to match the input feature dimensions of the next stage.

7. The fatigue driving detection method based on large model classification adaptor according to claim 1 is characterized in that: In step S5, the parameters in step S3 and step S4 are fine-tuned using low-rank adaptation technology.

Citation Information

Patent Citations

  • A driver fatigue driving detection method and system

    CN104881955B

  • Fatigue driving detection methods and devices

    CN105769120B

  • Driver fatigue detection method based on micro-expressions

    CN108053615B

  • A driving fatigue identification method and system based on large model

    CN118277835B

  • Fatigue driving early-warning method based on expression recognition

    CN108664947A