Driver distraction behavior detection method and system based on improved YOLOv8 model
By improving the YOLOv8 model, replacing the network structure and adding new modules, the existing driver distraction behavior detection methods in terms of accuracy and real-time balance, efficient and accurate detection are achieved, and road traffic safety is improved.
Patent Information
- Application Number
- CN202510276856.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing driver distraction behavior detection methods are poor in terms of accuracy and real-time balance, and over-reliance on complex models leads to high computing resources.
Using the improved YOLOv8 model, replace the CSPLayer_2Conv module in the network as the LarK module, add CAFM module, and replace the original DecoupledHead with RT-DETR Decoder in the detection head to improve feature extraction and detection performance.
It achieves higher detection accuracy and real-time performance, reduces the consumption of computing resources, overcomes the shortcomings of relying on complex models in the prior art, and improves road traffic safety and ability to prevent distracted driving accidents.
Smart Images

Figure CN120107937A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information processing technology, and in particular relates to a driver distraction behavior detection method and system based on an improved YOLOv8 model. Background Art
[0002] At present, driver distraction behavior detection based on computer vision has become a key research field in intelligent transportation systems. However, in reality, driver distraction behavior detection methods are not well balanced between accuracy and real-time performance. Vehicle status detection methods are susceptible to interference, physiological signal methods have complex equipment, and depth cameras in visual detection methods have limitations. Although ordinary cameras are good, the existing model feature extraction and fusion are insufficient, resulting in low detection accuracy in complex scenes. Therefore, the development of accurate, efficient, and well-balanced detection strategies is of far-reaching significance and has certain practical value.
[0003] Existing distraction detection methods mainly rely on local feature extraction, but this method is insufficient in capturing global feature information, which may lead to unstable performance of the model in complex scenarios. In addition, how to effectively integrate the local feature extraction capability of convolutional neural networks with the global feature capture advantages of Transformer models and ensure the overall real-time and accuracy of the system is still an urgent problem to be solved.
[0004] In view of the above situation, the present invention proposes a driver distraction behavior detection method and system based on the YOLOv8 model, which can effectively improve the existing technology and overcome its shortcomings. Summary of the invention
[0005] The purpose of the present invention is to provide a driver distraction behavior detection method and system based on an improved YOLOv8 model to solve the problems of poor accuracy and real-time performance, over-reliance on complex models and high consumption of computing resources in the prior art.
[0006] To achieve the above object, the present invention provides a method for detecting driver distraction behavior based on an improved YOLOv8 model, comprising the following steps:
[0007] Step 1: Based on the State Farm Distracted Driver Detection dataset, normal samples are removed to construct a distracted driving dataset, and the Mosaic data enhancement technology is used to preprocess the distracted driving dataset, and the training set and validation set are divided after manual annotation;
[0008] Step 2: Build an improved YOLOv8 model framework; replace the CSPLayer_2Conv module in Stage Layer 1, Stage Layer 2, Stage Layer 3, and Stage Layer 4 in the original YOLOv8 network Backbone with the LarK module, add the CAFM module to the Neck part, and replace the original detection head DecoupledHead with RT-DETR Decoder in the Head part;
[0009] Step 3: Use the training set to train the improved YOLOv8 model, and use some images in the Driver Distraction Dataset as the test set to test the model performance;
[0010] Step 4: Use accuracy, precision, recall, F1 value, parameter amount and GFLOPs as evaluation indicators to evaluate the performance of the detection results of step 3.
[0011] Preferably, the LarK module is composed of a depthwise separable convolution (DW), a SE Block, and a feedforward network (FFN) module with a GRN unit. Feature extraction is performed through the LarK module. The specific process is as follows: BN is selected and incorporated into the convolution layer; in the DW conv part, based on pixel correlation, an architecture combining a convolution layer parallel to a large kernel and a dilated convolution module is adopted to capture high-quality feature maps, which are converted into an equivalent kernel Dilated Re-param Block, and default hyperparameter values are given, wherein the default hyperparameter values are an equivalent kernel size K=13, a parallel convolution layer size k=(5,7,3,3,3) and a dilation rate r=(1,2,3,4,5) of the parallel convolution layer.
[0012] Preferably, the CAFM module is used to extract local features and global features to capture more extensive data information. The specific process is as follows:
[0013] The local branch of CAFM is used to extract local features and achieve comprehensive denoising; the calculation expression is as follows:
[0014] F conv =W 3×3×3 (CS(W 1×1 (Y)));
[0015] Among them, F conv is the output of the local branch, W 1×1 represents a convolution kernel of size 1×1, W 3×3×3 represents a convolution kernel of size 3×3×3, CS represents the channel shuffle operation, and Y is the input feature map;
[0016] The self-attention mechanism is used for the global branch of CAFM to capture more extensive data information; query (Q), key (K) and value (V) are generated through 1×1 convolution and 3×3 deep convolution to obtain three tensors of shape ^H×^W×^C, and then Q is reshaped into Reshape K into Finally, the attention map is calculated by K and Q The calculation expression of the global branch is as follows:
[0017]
[0018] Here, α is a learnable scaling parameter that controls the and The matrix product size, F att Represents the output of the global branch;
[0019] Get the output result F of the CAFM module out , the expression is as follows:
[0020] F out =F att +F conv .
[0021] Preferably, RT-DETR Decoder adopts a multi-level iterative optimization mechanism to improve the expressiveness of target features and finally output the category information and bounding box position of the target. The specific contents are as follows: the core of RT-DETR lies in the TransformerDecoder module, which iteratively optimizes the target through a multi-layer structure, generates target categories and bounding boxes, and improves the detection performance with the help of an auxiliary prediction head; each layer of the Decoder consists of a self-attention mechanism, a cross-attention mechanism, and a feedforward network module, wherein the self-attention mechanism allows target queries to exchange information with each other and simulates the global relationship between them; the cross-attention mechanism enables the target query to interact with the output characteristics of the encoder, extract relevant information from multi-scale image features and integrate global context; the feedforward network module further processes the query embedding through nonlinear transformations to enhance its expressiveness.
[0022] Preferably, the calculation expressions for accuracy, precision, recall and F1 value in step 4 are as follows:
[0023]
[0024] Where Accuracy represents accuracy, Precision represents precision, Recall represents recall, FP represents the number of negative predictions as positives, TP represents the number of positive predictions as positives, TN represents the number of negative predictions as negatives, and FN represents the number of positive predictions as negatives.
[0025] The present invention also provides a system for detecting driver distraction behavior based on an improved YOLOv8 model, comprising a data set construction module, a feature extraction module, a model optimization module, a behavior recognition module and a performance improvement module;
[0026] Dataset building module: used to create and preprocess a driving behavior dataset containing various distracting behaviors (such as texting, calling, etc.), divide the dataset into training set and validation set to meet the needs of model training and evaluation; and introduce some images in DriverDistractionDataset as test sets to test model accuracy and generalization ability;
[0027] Feature extraction module: used to collect and extract the driver's status information and behavior characteristics in real time through the improved YOLOv8 network. Information collection is achieved using on-board equipment and cameras.
[0028] Model optimization module: used to introduce large kernel convolution LarK, convolutional attention hybrid CAFM mechanism and RT-DETR Decoder module into the YOLOv8 network. By optimizing the network structure, the model can adaptively adjust the feature extraction strategy to achieve efficient recognition of driver distraction behavior.
[0029] Behavior recognition module: used to adjust and output the recognition results of driver distraction behavior in real time according to the model optimization results, ensuring the accuracy and real-time performance of the system in identifying driver distraction behavior;
[0030] Performance improvement module: used to globally optimize the preliminary recognition results based on large kernel convolution LarK, convolutional attention hybrid CAFM mechanism and RT-DETRDecoder module. The optimization process involves fitness evaluation of the recognition strategy to generate the optimal driver distraction behavior detection strategy.
[0031] Preferably, the behavior recognition module extracts image features based on LarK through an improved YOLOv8 model, calculates output results based on local branches and global branches, and uses RT-DETR Dec oder to accurately identify driver distraction behaviors.
[0032] Therefore, the present invention adopts the above-mentioned driver distraction behavior detection method and system based on the improved YOLOv8 model, and performs distraction behavior detection by performing data fusion, feature extraction and preprocessing on the collected driving behavior data according to a preset processing flow; this method overcomes the shortcomings of existing detection methods in accuracy and real-time performance; at the same time, compared with traditional detection methods, it effectively improves the detection accuracy, establishes a distraction behavior recognition model based on the improved YOLOv8 model, and effectively overcomes the problems of reliance on complex models and high consumption of computing resources in the prior art, which is crucial to improving road traffic safety and preventing distracted driving accidents, and also provides more safety protection for drivers.
[0033] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is an overall flow chart of a driver distraction behavior detection method based on an improved YOLOv8 model of the present invention;
[0035] Figure 2 A structural block diagram of a system used in a driver distraction behavior detection method based on an improved YOLOv8 model of the present invention;
[0036] Figure 3 Schematic diagram of the structure of the improved YOLOv8 according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] See also Figure 1 , a driver distraction behavior detection method based on an improved YOLOv8 model, comprising the following steps:
[0039] Step 1: Based on the State Farm Distracted Driver Detection dataset, normal samples were removed to construct a distracted driving dataset, and the distracted driving dataset was preprocessed using Mosaic data enhancement technology. After manual annotation, the dataset was divided into a training set and a validation set. More specifically, the original dataset contains 10 categories, a total of 17,462 training images and 4,961 test images. These images are from 26 subjects of different body shapes and skin colors. Normal driving samples were removed, and 10,760 representative images were selected as the training dataset. The dataset was further divided into a training set (7,531 images) and a validation set (3,229 images) in a ratio of 7:3.
[0040] Step 2: Build an improved YOLOv8 model framework; replace the CSPLayer_2Conv module in Stage Layer 1, Stage Layer 2, Stage Layer 3, and Stage Layer 4 in the original YOLOv8 network Backbone with the LarK module, use the LarK module and the dilated convolution mechanism to improve feature extraction, promote inter-channel communication and spatial aggregation through SE Block, and increase feature depth at the same time; add the CAFM module to the Neck part, extract local features and capture more extensive data information through local branches and global branches respectively; and replace the original detection head DecoupledHead with RT-DETR Decoder in the Head part, iteratively optimize the target through a multi-layer structure, generate target categories and bounding boxes, and use the auxiliary prediction head to improve detection performance;
[0041] Among them, the LarK module consists of a depthwise separable convolution (DW), a SE Block, and a feedforward network (FFN) module with a GRN unit. Feature extraction is performed through the LarK module. The specific process is as follows: BN is selected and incorporated into the convolution layer; in the DW conv part, based on pixel correlation, an architecture combining a convolution layer parallel to a large kernel and a dilated convolution module is adopted to capture high-quality feature maps, which are converted into an equivalent kernel Dilated Re-param Block, and default hyperparameter values are given. The default hyperparameter values are the equivalent kernel size K = 13, the parallel convolution layer size k = (5, 7, 3, 3, 3) and the dilation rate of the parallel convolution layer r = (1, 2, 3, 4, 5).
[0042] The CAFM module extracts local and global features to capture more extensive data information. The specific process is as follows:
[0043] The local branch of CAFM is used to extract local features and achieve comprehensive denoising; the calculation expression is as follows:
[0044] Fconv =W 3×3×3 (CS(W 1×1 (Y)));
[0045] Among them, F conv is the output of the local branch, W 1×1 represents a convolution kernel of size 1×1, W 3×3×3 represents a convolution kernel of size 3×3×3, CS represents the channel shuffle operation, and Y is the input feature map;
[0046] The self-attention mechanism is used for the global branch of CAFM to capture more extensive data information; query (Q), key (K) and value (V) are generated through 1×1 convolution and 3×3 deep convolution, and three shapes are obtained. ^ H× ^ W× ^ C, and then reshape Q into Reshape K into Finally, the attention map is calculated by K and Q The calculation expression of the global branch is as follows:
[0047]
[0048] Here, α is a learnable scaling parameter that controls the and The matrix product size, F att Represents the output of the global branch;
[0049] Get the output result F of the CAFM module out , the expression is as follows:
[0050] F out =F att +F conv .
[0051] RT-DETR Decoder adopts a multi-level iterative optimization mechanism to continuously improve the expressiveness of target features, and finally outputs the category information and bounding box position of the target. At the same time, by introducing an auxiliary prediction head, the detection effect and accuracy are further enhanced; the specific contents are as follows: the core of RT-DETR lies in the Transformer Decoder module, which iteratively optimizes the target through a multi-layer structure, generates target categories and bounding boxes, and uses auxiliary prediction heads to improve detection performance; each layer of the Decoder consists of a self-attention mechanism, a cross-attention mechanism, and a feedforward network module. Among them, the self-attention mechanism allows target queries to exchange information with each other and simulate the global relationship between them; the cross-attention mechanism enables the target query to interact with the output characteristics of the encoder, extract relevant information from multi-scale image features and integrate global context; the feedforward network module further processes the query embedding through nonlinear transformations to enhance its expressiveness.
[0052] Step 3: Use the training set to train the improved YOLOv8 model, and use some images in the Driver Distraction Dataset as the test set to test the model performance; More specifically, the algorithm of the present invention is based on the Windows 10 operating system, Intel (R) core (TM) i9-10900K CPU, Nvidia GeForce RTX 309024G GPU, based on the deep learning framework of Pytorch1.10, and the supporting environment is CUDA11.1. During the network training process, the input image is scaled to a uniform size (640 pixels × 640 pixels) and normalized, SGD is selected as the network optimizer, the learning rate is set to 0.01, the momentum factor is set to 0.937, the weight decay factor is set to 0.0005, the Batch Size is set to 32, and a total of 50 rounds of training; Finally, the accuracy and recall rate of the driver distraction behavior detection method of the present invention on the test set are as high as 95.85% and 94.12%, respectively.
[0053] Step 4: Use accuracy, precision, recall, F1 value, parameter amount and GFLOPs as evaluation indicators to evaluate the performance of the detection results of step 3; the calculation expressions of accuracy, precision, recall and F1 value are as follows:
[0054]
[0055] Where Accuracy represents accuracy, Precision represents precision, Recall represents recall, FP represents the number of negative predictions as positives, TP represents the number of positive predictions as positives, TN represents the number of negative predictions as negatives, and FN represents the number of positive predictions as negatives.
[0056] See also Figure 2 , a system for detecting driver distraction behavior based on an improved YOLOv8 model, comprising a data set construction module, a feature extraction module, a model optimization module, a behavior recognition module and a performance improvement module;
[0057] Dataset building module: used to create and preprocess a driving behavior dataset containing various distracting behaviors (such as texting, calling, etc.), divide the dataset into training set and validation set to meet the needs of model training and evaluation; and introduce some images in DriverDistractionDataset as test sets to test model accuracy and generalization ability;
[0058] Feature extraction module: used to collect and extract the driver's status information and behavior characteristics in real time through the improved YOLOv8 network. Information collection is achieved using on-board equipment and cameras.
[0059] Model optimization module: used to introduce large kernel convolution LarK, convolutional attention hybrid CAFM mechanism and RT-DETR Decoder module into the YOLOv8 network. By optimizing the network structure, the model can adaptively adjust the feature extraction strategy to achieve efficient recognition of driver distraction behavior.
[0060] Behavior recognition module: used to adjust and output the recognition results of driver distraction behavior in real time according to the model optimization results, to ensure the accuracy and real-time performance of the system in identifying driver distraction behavior; specifically, the behavior recognition module uses the improved YOLOv8 model to extract image features based on LarK, calculates the output results based on local branches and global branches, and uses RT-DETR Decoder to accurately identify driver distraction behavior.
[0061] Performance improvement module: used to globally optimize the preliminary recognition results based on large kernel convolution LarK, convolutional attention hybrid CAFM mechanism and RT-DETRDecoder module. The optimization process involves fitness evaluation of the recognition strategy to generate the optimal driver distraction behavior detection strategy.
[0062] Therefore, the present invention adopts the above-mentioned driver distraction behavior detection method and system based on the improved YOLOv8 model, which effectively resolves the dilemma of existing detection means in balancing accuracy and real-time performance, breaks through the limitations of traditional convolutional neural networks in global feature extraction and the bottlenecks of Transformer models in resource occupation and reasoning time. The established distraction behavior recognition model based on the improved YOLOv8 model can keenly capture the subtle features of driver distraction behavior in complex driving scenarios, greatly improve detection accuracy, and ensure computational efficiency at the same time, laying a solid foundation for the driving safety warning mechanism and effectively reducing the risk of traffic accidents.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A driver distraction behavior detection method based on an improved YOLOv8 model, characterized in that: The following steps are involved: Step 1: Based on the State Farm Distracted Driver Detection dataset, normal samples are removed to construct a distracted driving dataset, and the distracted driving dataset is preprocessed using Mosaic data enhancement technology. After manual annotation, the dataset is divided into a training set and a validation set. Step 2: Build an improved YOLOv8 model framework; replace the CSPLayer_2Conv module in Stage Layer1, StageLayer2, Stage Layer3, and Stage Layer4 in the original YOLOv8 network Backbone with the LarK module, add the CAFM module to the Neck part, and replace the original detection head Decoupled Head with RT-DETR Decoder in the Head part; Step 3: Use the training set to train the improved YOLOv8 model, and use Driver Distraction Dataset as the test set to test the model performance; Step 4: Use accuracy, precision, recall, F1 value, parameter amount and GFLOPs as evaluation indicators to evaluate the performance of the detection results of step 3.
2. The method for detecting driver distraction based on the improved YOLOv8 model according to claim 1, characterized in that: The LarK module consists of a depthwise separable convolution, a SE Block, and a feedforward network module with a GRN unit. Feature extraction is performed through the LarK module. The specific process is as follows: BN is selected and incorporated into the convolution layer; in the DW conv part, based on pixel correlation, an architecture combining a convolution layer in parallel with a large kernel and a dilated convolution module is adopted to capture high-quality feature maps, which are converted into an equivalent kernel Dilated Re-param Block, and default hyperparameter values are given. The default hyperparameter values are the equivalent kernel size K = 13, the parallel convolution layer size k = (5, 7, 3, 3, 3) and the dilation rate of the parallel convolution layer r = (1, 2, 3, 4, 5).
3. The method for detecting driver distraction based on the improved YOLOv8 model according to claim 2, characterized in that: The CAFM module is used to extract local features and global features. The specific process is as follows: The local features are extracted using the local branch of CAFM; the calculation expression is as follows: F conv =W 3×3×3 (CS(W 1×1 (Y))); Among them, F conv is the output of the local branch, W 1×1 represents a convolution kernel of size 1×1, W 3×3×3 represents a convolution kernel of size 3×3×3, CS represents the channel shuffle operation, and Y is the input feature map; A self-attention mechanism is used for the global branch of CAFM; Q, K, and V are generated through 1×1 convolution and 3×3 depth convolution to obtain three tensors of shape ^H×^W×^C, and then Q is reshaped into Reshape K into Finally, the attention map is calculated by K and Q The calculation expression of the global branch is as follows: Here, α is a learnable scaling parameter that controls the and The matrix product size, F att Represents the output of the global branch; Get the output result F of the CAFM module out , the expression is as follows: F out =F att +F conv 。 4. The method for detecting driver distraction based on the improved YOLOv8 model according to claim 3, characterized in that: RT-DETR Decoder adopts a multi-level iterative optimization mechanism to improve the expressiveness of target features and ultimately outputs the target’s category information and bounding box position. The specific contents are as follows: The core of RT-DETR lies in the Transformer Decoder module, which iteratively optimizes the target through a multi-layer structure to generate target categories and bounding boxes; each layer of the Decoder consists of a self-attention mechanism, a cross-attention mechanism, and a feedforward network module. The self-attention mechanism allows target queries to exchange information with each other and simulates the global relationship between them; the cross-attention mechanism enables the target query to interact with the output characteristics of the encoder, extracting relevant information from multi-scale image features and integrating global context; the feedforward network module further processes query embedding through nonlinear transformations.
5. The method for detecting driver distraction based on the improved YOLOv8 model according to claim 4, characterized in that: The calculation expressions for accuracy, precision, recall and F1 value in step 4 are as follows: Where Accuracy represents accuracy, Precision represents precision, Recall represents recall, FP represents the number of negative predictions as positives, TP represents the number of positive predictions as positives, TN represents the number of negative predictions as negatives, and FN represents the number of positive predictions as negatives.
6. A system for detecting driver distraction behavior based on an improved YOLOv8 model as claimed in any one of claims 1 to 5, characterized in that: It includes data set construction module, feature extraction module, model optimization module, behavior recognition module and performance improvement module; Dataset building module: used to create and preprocess a driving behavior dataset containing multiple distracting behaviors, divide the dataset into training set and validation set, and introduce Driver Distraction Dataset as a test set to test the model accuracy and generalization ability; Feature extraction module: used to collect and extract the driver's status information and behavior characteristics in real time through the improved YOLOv8 network. Information collection is achieved using on-board equipment and cameras. Model optimization module: used to introduce large kernel convolution LarK, convolutional attention hybrid CAFM mechanism and RT-DETR Decoder module into the YOLOv8 network. By optimizing the network structure, the model can adaptively adjust the feature extraction strategy to achieve efficient recognition of driver distraction behavior. Behavior recognition module: used to adjust and output the recognition results of driver distraction behavior in real time based on the model optimization results; Performance improvement module: used to globally optimize the preliminary recognition results based on large kernel convolution LarK, convolutional attention hybrid CAFM mechanism and RT-DETRDecoder module.
7. The system used in the driver distraction behavior detection method based on the improved YOLOv8 model according to claim 8, characterized in that: The behavior recognition module uses the improved YOLOv8 model to extract image features based on LarK, calculates the output results of local branches and global branches, and uses the RT-DETR Decoder to accurately identify driver distraction behaviors.
Citation Information
Patent Citations
Convolutional neural network and attention mechanism sign language recognition method based on YOLOv8
CN119049121A
Factory vehicle detection method based on improved YOLO v8
CN119360301A
Power transmission line multi-defect detection method and system based on improved YOLOv8s
CN119559479A