Driver dangerous driving behavior visual identification method and device, terminal and medium
By improving the C2f module of the YOLOv8n model as a C2f-DC module and using the SAIOU loss function, the problems of large amount of calculation and low detection accuracy in driver dangerous driving behavior recognition are solved, and more efficient behavior recognition and accident prevention are achieved.
Patent Information
- Application Number
- CN202510668142.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The existing driver hazard driving behavior recognition technology has problems such as large calculation volume, low detection accuracy and imbalance in samples, resulting in poor accident prevention results.
Using the improved YOLOv8n model, by reconstructing the C2f module into the C2f-DC module and introducing the DC module, combined with the SAIOU loss function, the detection accuracy and lightweight level are improved, and more efficient recognition of driver dangerous driving behaviors is achieved.
It improves the recognition accuracy and detection accuracy of drivers' dangerous driving behavior, reduces the amount of calculation, and enhances accident prevention capabilities.
Smart Images

Figure CN120340003A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly relates to a visual recognition method, device, terminal and medium for drivers' dangerous driving behaviors. Background Art
[0002] The driving state of a driver directly affects the driving quality. When a driver engages in dangerous behaviors such as making a call, typing, having a distracted conversation, eating or drinking, it is very easy to cause distracted driving, affecting the reaction and response in an emergency, thus leading to tragedies and being an important cause of traffic accidents. With the continuous development of deep learning and visual recognition technologies, using deep learning technology to identify drivers' dangerous driving behaviors, if timely reminders or measures such as autonomous driving can be taken in advance, the accident rate can be reduced, protecting the driving safety of passengers and drivers, which has important research significance. Summary of the Invention
[0003] This application provides a visual recognition method, device, terminal and medium for drivers' dangerous driving behaviors, and its advantage is that it has a better recognition accuracy rate for drivers' dangerous driving behaviors.
[0004] The technical solution of this application is as follows: On the one hand, this application discloses a visual recognition method for drivers' dangerous driving behaviors, including the following steps: S1: Collect image data of the driver during driving; S2: Use a dangerous driving behavior recognition model to detect and recognize the image data of the driver during driving, and output whether the driver has dangerous driving behaviors; The dangerous driving behavior recognition model is obtained through the following steps: S21: Collect image data of the driver during dangerous driving, and construct a driver dangerous driving state data set based on this; S22: Construct a dangerous driving behavior recognition model, which is an improvement of the YOLOv8n model. Use the DC module to reconstruct the C2f module in the backbone network of the YOLOv8n model into a C2f-DC module; S23: Use the driver dangerous driving state data set to train the dangerous driving behavior recognition model.
[0005] Further, in step S21, it further includes the step of annotating the collected image data of the driver during dangerous driving, and dividing the driver dangerous driving state data set into a training set and a validation set.
[0006] Further, in step S21, the annotation types of the image data of the driver during dangerous driving include nine categories: making a call with the right hand, typing with the right hand, making a call with the left hand, typing with the left hand, tuning the radio, drinking, having a conversation with a passenger, leaning to the right side while driving, and touching the head.
[0007] Furthermore, in the DC module: First, the Split operation is used on the input signal to split it into three parts: Q, V, and K. The Q signal channel undergoes feature input processing through the Attn Mat module with an attention mechanism to obtain signal A, which is then multiplied by the K signal to get Uatt: ; A 3×3 depthwise separable convolution DWConv is introduced to decompose Uatt into a depthwise convolution and a pointwise convolution. After passing through the ReLU activation function and batch normalization operations, the resulting U2 signal is obtained. The channel number is adjusted through a 1×1 convolution to combine U2 and Uatt to get U3: ; The V signal first undergoes an adaptive average pooling operation and is divided into V1 and V2. The V1 signal expands the input channels through a 1×1 convolution, introduces the depthwise separable convolution DWConv, decomposes it into a depthwise convolution and a pointwise convolution, passes through the GAM global attention mechanism, and finally adjusts the channel number through a 1×1 convolution to obtain the signal V3; The final output connects V2, V3, and U3 through a residual connection to achieve the signal output. The output signal Y is expressed as: .
[0008] Furthermore, the reconstruction of the C2f module is as follows: Replace the Bottleneck bottleneck layer structure in the backbone network C2f module with the DC module to obtain the C2f-DC module.
[0009] Furthermore, the improvement of the YOLOv8n model also includes: Using the loss function SAIOU to replace the loss function CIOU in the YOLOv8n model. The form of the loss function SAIOU is as follows:
[0010] In the formula, IOU represents the intersection over union, B p and B g represent the regions of the predicted box and the ground truth box respectively, represents the angular loss, represents the distance loss, represents the shape loss, w g and h g are the width and height of the ground truth box, w p and h p are the width and height of the predicted box, v represents a function for measuring the similarity of aspect ratios, ρ x and ρ y represent the normalized squared distances in the x and y axis directions respectively, ω w and ω hrespectively represent the width and height shape penalty coefficients of the ground truth box and the predicted box, and θ represents the hyperparameter for controlling the shape loss.
[0011] Further, it further includes step S3: when detecting that the driver has dangerous driving behavior, sending a reminder to the driver and / or activating the vehicle's automatic driving.
[0012] On the other hand, the present application provides a visual recognition device for a driver's dangerous driving behavior, including: An image acquisition unit for acquiring image data when the driver is driving; A detection unit for using a dangerous driving behavior recognition model to detect and recognize the image data when the driver is driving, and outputting whether the driver has dangerous driving behavior; the dangerous driving behavior recognition model is obtained through the following steps: S21: Acquire image data when the driver is driving dangerously, and construct a driver's dangerous driving state data set therefrom; S22: Construct a dangerous driving behavior recognition model, which is an improvement of the YOLOv8n model. The DC module is used to reconstruct the C2f module in the backbone network of the YOLOv8n model into a C2f-DC module; S23: Use the driver's dangerous driving state data set to train the dangerous driving behavior recognition model; And a reminder unit for sending a reminder to the driver when detecting that the driver has dangerous driving behavior.
[0013] On the other hand, the present application provides a visual recognition terminal for a driver's dangerous driving behavior, including a processor and a memory. When the computer program stored in the memory is called and executed by the processor, the visual recognition method for a driver's dangerous driving behavior as described above is implemented.
[0014] On the other hand, the present application provides a computer-readable medium, which stores a computer program. When the computer program is called and executed by a computer, the visual recognition method for a driver's dangerous driving behavior as described above is implemented.
[0015] In summary, the beneficial effects of the present application are: 1. In view of the problem that the use of residual connections in the C2f module of the YOLOv8n algorithm for multi-branch structure fusion features easily leads to an increase in computational complexity, a new C2f-DC module is designed. The DC module introduces a separable convolution module, which improves the lightweight level of the module. At the same time, it gives full play to the advantages of the Attn Mat and GAM attention mechanism modules to improve the detection accuracy of the module. In this way, the C2f-DC module can not only have the lightweight advantage but also ensure the detection accuracy, avoiding the information redundancy caused by the traditional C2f residual connection; 2. In view of the problems of low detection accuracy and imbalance between easy and difficult samples in the CIOU loss function of the YOLOv8n algorithm, this application proposes a new loss function SAIOU to improve the prediction accuracy between the detection head bounding box and the ground truth box. Compared with the CIOU loss function, the SAIOU loss function overcomes the defect of using the arctangent function to solve the angle deviation in the CIOU function, and has better positioning accuracy and convergence speed. Description of the Drawings
[0016] Figure 1 is a schematic structural diagram of the dangerous driving behavior recognition model of this application; Figure 2 is a schematic structural diagram of the DC module network of this application; Figure 3 is a schematic diagram of a partial data set of this application; Figure 4 is a schematic diagram of the detection effect of the YOLOv8n algorithm before improvement; Figure 5 is a schematic diagram of the detection effect of the improved Driver-YOLOv8 algorithm; Figure 6 is the P-R curve graph of the YOLOv8n algorithm before improvement; Figure 7 is the P-R curve graph of the improved Driver-YOLOv8 algorithm. Detailed Description of the Preferred Embodiments
[0017] The following will describe in detail the specific embodiments of this application with reference to the accompanying drawings.
[0018] A specific embodiment of this application provides a visual recognition method for drivers' dangerous driving behaviors, including the following steps: S1: Collect image data of the driver during driving; S2: Use the dangerous driving behavior recognition model to detect and recognize the image data of the driver during driving, and output whether the driver has dangerous driving behaviors; The structure of the dangerous driving behavior recognition model is as Figure 1 shown, and the dangerous driving behavior recognition model is obtained through the following steps: S21: Collect image data of the driver during dangerous driving, and construct a data set of the driver's dangerous driving state based on this; In step S21, it further includes the step of annotating the collected image data of the driver during dangerous driving, and dividing the data set of the driver's dangerous driving state into a training set and a validation set.
[0019] In step S21, the annotation types of the driver's dangerous driving image data include nine categories: right hand making a call (Right_p), right hand typing (Right_w), left hand making a call (Left_p), left hand typing (Left_w), tuning the radio (Play_video), drinking (Drinking), talking to passengers (Talking), leaning to the right while driving (Taking), and touching the head (Head).
[0020] In this example, the StateFarm dataset launched by the Kaggle organization is adopted, and the driver's dangerous driving images are selected to construct a driver's dangerous driving state dataset. The driver's dangerous driving state dataset contains 10,760 photos, covering the above 9 common types of dangerous driving behaviors.
[0021] The dataset is labeled using Labelimg and divided into a training set and a validation set, with 8,758 and 2,002 photos respectively. Part of the dataset is as Figure 3 shown.
[0022] S22: Construct a dangerous driving behavior recognition model, which is an improvement of the YOLOv8n model. As Figure 2 , the C2f module in the backbone network of the YOLOv8n model is reconstructed into a C2f-DC module using the DC module; In the DC module, for the input signal, the Split operation is first used to split the channels into three parts: Q, V, and K. The Q signal channel undergoes feature input processing through the Attn Mat module with an attention mechanism to obtain signal A, and then it is multiplied by the K signal to obtain Uatt: ; A 3×3 depthwise separable convolution DWConv is introduced to decompose Uatt into a depthwise convolution and a pointwise convolution. After passing through the ReLU activation function and batch normalization operation, the result U2 signal is obtained. The channel number is adjusted through a 1×1 convolution to combine U2 and Uatt to obtain U3: ; The V signal first undergoes an adaptive average pooling operation and is divided into V1 and V2. The V1 signal expands the input channels through a 1×1 convolution, introduces the DWConv depthwise separable convolution, decomposes it into a depthwise convolution and a pointwise convolution, passes through the GAM global attention mechanism, and finally adjusts the channel number through a 1×1 convolution to obtain the signal V3; The final output realizes the output of the signal by connecting V2, V3, and U3 through a residual. The output signal Y is expressed as: .
[0023] The reconstruction of the C2f module is as follows: Replace the Bottleneck bottleneck layer structure in the backbone network C2f module with the DC module to obtain the C2f-DC module.
[0024] The improvement of the YOLOv8n model also includes: Using the loss function SAIOU to replace the loss function CIOU in the YOLOv8n model to improve the prediction accuracy between the detection head bounding box and the ground truth box. The form of the loss function SAIOU is as follows:
[0025] In the formula, IOU represents the intersection over union, B p and B g represent the regions of the predicted box and the ground truth box respectively, represents the angular loss, represents the distance loss, represents the shape loss, w g and h g are the width and height of the ground truth box, w p and h p are the width and height of the predicted box, v represents the function for measuring the aspect ratio similarity, ρ x and ρ y represent the normalized squared distances in the x and y axis directions respectively, ω w and ω h represent the shape penalty coefficients of the width and height of the ground truth box and the predicted box respectively, and θ represents the hyperparameter for controlling the shape loss.
[0026] The improved algorithm model is as Figure 1 shown, and we name it the Driver-YOLOv8 algorithm.
[0027] S23: Use the driver's dangerous driving state dataset to train the dangerous driving behavior recognition model.
[0028] S3: When it is detected that the driver has dangerous driving behavior, send a reminder to the driver and / or activate the vehicle's automatic driving. Sending a reminder to the driver is a voice reminder or a video text reminder.
[0029] To verify the performance of the improved algorithm in driver dangerous driving behavior detection, we conducted ablation experiments and comparative experiments. The experiments were carried out on the Ubuntu20.04 operating system, with the CPU configured as Inter(R)i9-13900K and the GPU configured as NVIDIA RTX A5000. The experimental environment selected the Pytorch 2.2 deep learning framework, the editing language was selected as python3.10.8, the CUDA version was 11.8.0, the Adam optimizer was selected, the number of iterations was set to 100, the batchsize was set to 8, the momentum factor was set to 0.937, and the weight decay factor was set to 0.0005.
[0030] Since there are 9 definitions of drivers' dangerous driving behaviors, different from traditional object detection tasks, drivers' abnormal behaviors usually involve small-scale hand or head movements, such as "typing with the right hand" or "adjusting the radio". The detection of these behaviors is more difficult and requires more precise model optimization strategies. Therefore, to evaluate the performance of the improved algorithm, precision (P), recall (R), and mean average precision (mAP) are selected as evaluation indicators. Among them, TP is the number of positive samples correctly identified; FP is the number of misdetected negative samples; FN is the number of missed positive samples; n represents the total number of dataset categories; i represents the number of detections.
[0031] The mean average precision mAP reflects the overall detection precision ability of the algorithm model and is an important evaluation indicator, mainly divided into mAP0.5 and mAP0.5:0.95.
[0032] To verify the effect of the Driver-YOLOv8 algorithm in drivers' dangerous driving behaviors, ablation experiments were carried out, and separate and combined experimental analyses were performed on two improved modules, C2f-DC and SAIOU. The experimental results are shown in Table 1. Experiment 1 represents the experimental results using the C2f-DC and SAIOU modules, that is, the calculation results of the YOLOv8n algorithm. Experiment 2 represents the experimental results after integrating the C2f-DC module into YOLOv8n. Experiment 3 represents the experimental results after integrating the SAIOU module into YOLOv8n. Experiment 4 represents the results after integrating both the C2f-DC and SAIOU modules into YOLOv8n, that is, the results of the Driver-YOLOv8 algorithm.
[0033] Table 1 Results of ablation experiments Through Table 1, the accuracy rate, recall rate, mAP0.5, and mAP0.5:0.95 of the YOLOv8n basic algorithm in the recognition of drivers' dangerous driving behaviors reached 98.9%, 92.5%, 95.6%, and 80.0% respectively. When the C2f-DC and SAIOU modules were introduced respectively, the mAP0.5 index was improved, with an increase of 2.6 and 2.4 percentage points respectively. Especially when the two were used together, the mAP0.5 value reached the best of 98.7%, an increase of 3.1 percentage points compared to the basic algorithm. Finally, compared with the algorithm before improvement, although the accuracy rate of the Driver-YOLOv8 algorithm decreased by 0.6 percentage points, the mAP0.5, mAP0.5:0.95, and recall rate reached 98.7%, 80.7%, and 98.3% respectively, an increase of 3.1, 0.7, and 5.8 percentage points respectively compared to the algorithm before improvement, proving that the improved algorithm has superior detection accuracy advantages.
[0034] Randomly select drivers' dangerous driving behaviors from the dataset and send them into the algorithms before and after improvement for verification. The detection effect diagrams of 9 types of dangerous driving behaviors are as Figure 4 、 5 shown. It can be analyzed from the detection effect comparison diagram that the detection accuracy of Driver-YOLOv8 for dangerous driving behaviors has been improved compared to the YOLOv8n algorithm before improvement.
[0035] The P-R curves before and after improvement are shown in Figures 6 and 7. The mean average precision of the improved algorithm is 98.7%, and the mean average precision of the YOLOv8n algorithm is 95.6%. The mAP of the improved algorithm has increased by 3.1 percentage points compared to the algorithm before improvement. In addition, in the comparison of the mean average precision of nine types of dangerous driving behaviors, only the mean average precision of typing with the right hand is slightly lower than that of the algorithm before improvement, and the mean average precision of the recognition of the remaining dangerous driving behaviors has been improved compared to before improvement, proving that the improved algorithm has better detection performance.
[0036] Another embodiment of the present application provides a visual recognition device for drivers' dangerous driving behaviors, including: An image acquisition unit for acquiring image data when the driver is driving; A detection unit for detecting and recognizing the image data when the driver is driving by using a dangerous driving behavior recognition model, and outputting whether the driver has dangerous driving behaviors; the dangerous driving behavior recognition model is obtained through the following steps: S21: Acquire image data when the driver is driving dangerously, and construct a dataset of the driver's dangerous driving state based on this; S22: Construct a dangerous driving behavior recognition model, which is an improvement of the YOLOv8n model. The DC module is used to reconstruct the C2f module in the backbone network of the YOLOv8n model into a C2f-DC module; S23: Train the dangerous driving behavior recognition model using the driver's dangerous driving state dataset; And a reminder unit, which is used to send a reminder to the driver when a dangerous driving behavior of the driver is detected.
[0037] Another embodiment of the present application provides a visual recognition terminal for a driver's dangerous driving behavior, including a processor and a memory. When the computer program stored in the memory is called and executed by the processor, the above-mentioned visual recognition method for a driver's dangerous driving behavior is implemented.
[0038] Another embodiment of the present application provides a computer-readable medium. When the computer program stored in the computer-readable medium is called and executed by a computer, the above-mentioned visual recognition method for a driver's dangerous driving behavior is implemented.
[0039] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A visual recognition method for drivers' dangerous driving behaviors, characterized in that, It includes the following steps: S1: Collect image data of the driver while driving; S2: Use the dangerous driving behavior recognition model to detect and recognize the image data of the driver while driving, and output whether the driver has dangerous driving behavior; The dangerous driving behavior recognition model is obtained through the following steps: S21: Collect image data of the driver during dangerous driving and construct a driver dangerous driving state data set based on this; S22: Construct a dangerous driving behavior recognition model, which is an improvement of the YOLOv8n model. The DC module is used to reconstruct the C2f module in the backbone network of the YOLOv8n model into a C2f-DC module; S23: Use the driver dangerous driving state data set to train the dangerous driving behavior recognition model.
2. The visual recognition method for dangerous driving behavior of a driver according to claim 1, wherein, In step S21, it also includes the steps of annotating the collected driver dangerous driving image data and dividing the driver dangerous driving state data set into a training set and a validation set.
3. The visual recognition method for the driver's dangerous driving behavior according to claim 2, characterized in that, In step S21, the annotation types of the driver dangerous driving image data include nine categories: making a call with the right hand, typing with the right hand, making a call with the left hand, typing with the left hand, tuning the radio, drinking, talking to passengers, leaning to the right while driving, and touching the head.
4. The visual recognition method for the dangerous driving behavior of a driver according to claim 1, characterized in that, In the DC module: For the input signal, first use the Split operation to split the channels into three parts: Q, V, and K. The Q signal channel undergoes feature input processing through the Attn Mat module with an attention mechanism to obtain signal A, and then multiplies it with the K signal to obtain Uatt: ; Introduce a 3×3 depthwise separable convolution DWConv to decompose Uatt into a depthwise convolution and a pointwise convolution. After passing through the ReLU activation function and batch normalization operation, the result U2 signal is obtained. Adjust the number of channels through a 1×1 convolution to combine U2 and Uatt to obtain U3: ; The V signal first undergoes an adaptive average pooling operation and is divided into V1 and V2. The V1 signal expands the input channels through a 1×1 convolution, introduces the DWConv depthwise separable convolution, decomposes it into a depthwise convolution and a pointwise convolution, passes through the GAM global attention mechanism, and finally adjusts the number of channels through a 1×1 convolution to obtain the signal V3; The final output realizes the output of the signal by connecting V2, V3, and U3 through a residual. The output signal Y is expressed as: 。 5. The visual recognition method for dangerous driving behavior of a driver according to claim 4, wherein The reconstruction of the C2f module is: Replace the Bottleneck bottleneck layer structure in the backbone network C2f module with the DC module to obtain the C2f-DC module.
6. The visual recognition method for dangerous driving behavior of a driver according to claim 1, characterized in that, The improvement of the YOLOv8n model also includes: Using the loss function SAIOU to replace the loss function CIOU in the YOLOv8n model. The form of the loss function SAIOU is as follows: ; where IOU represents the intersection over union, B p and B g represent the regions of the predicted bounding box and the ground truth bounding box respectively, represents the angular loss, represents the distance loss, represents the shape loss, w g and h g are the width and height of the ground truth bounding box, w p and h p are the width and height of the predicted bounding box, v represents the function for measuring the similarity of aspect ratio, ρ x and ρ y represent the normalized squared distances in the x and y axis directions respectively, ω w and ω h represent the shape penalty coefficients of the width and height of the ground truth bounding box and the predicted bounding box respectively, and θ represents the hyperparameter for controlling the shape loss.
7. The visual recognition method for dangerous driving behavior of a driver according to claim 1, characterized in that It also includes step S3: When it is detected that the driver has dangerous driving behavior, give a reminder to the driver and / or activate the vehicle's automatic driving.
8. A visual recognition device for drivers' dangerous driving behaviors, characterized in that, It includes: An image acquisition unit for collecting image data of the driver while driving; A detection unit for using the dangerous driving behavior recognition model to detect and recognize the image data of the driver while driving, and output whether the driver has dangerous driving behavior; The dangerous driving behavior recognition model is obtained through the following steps: S21: Collect the image data of the driver during dangerous driving and construct a dataset of the driver's dangerous driving states based on this; S22: Construct a dangerous driving behavior recognition model, where the dangerous driving behavior recognition model is an improvement of the YOLOv8n model. The C2f module in the backbone network of the YOLOv8n model is reconstructed into a C2f-DC module by using a DC module; S23: Use the dataset of the driver's dangerous driving states to train the dangerous driving behavior recognition model; And a reminder unit for sending a reminder to the driver when it is detected that the driver has dangerous driving behavior.
9. A visual recognition terminal for drivers' dangerous driving behaviors, characterized in that, It includes a processor and a memory. The memory stores a computer program. When the computer program is called and executed by the processor, it implements the visual recognition method for the driver's dangerous driving behavior as described in any one of claims 1-7.
10. A computer-readable medium, characterized in that, The computer-readable medium stores a computer program. When the computer program is called and executed by a computer, it implements the visual recognition method for the driver's dangerous driving behavior as described in any one of claims 1-7.
Citation Information
Patent Citations
Driver face detection method based on lightweight improved YOLOv8 model and medium
CN117935333A
FPC (Flexible Printed Circuit) substrate reinforcing sheet re-pasting defect detection method based on 2D (Two-Dimensional) image
CN118533839A
Dangerous driving visual detection method based on YOLO-SGC
CN118552939A
Man-machine perception and decision-making method based on generalized driving behavior map
CN118839588A
Lightweight driver distraction behavior detection method
CN119904845A
Cited By
Landslide detection method and system based on glaf and sa-iou optimization
CN122530833A