Driver behavior identification method based on yo-v8

By improving the driver behavior recognition method of yolo-v8, using data augmentation and feature extraction technology, combined with CIoU loss function, the problems of low driver distraction behavior detection accuracy and complex model in the prior art are solved, and accurate identification and real-time monitoring of driver behavior are achieved.

CN120107936AInactive Publication Date: 2025-06-06GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510268952.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of low detection accuracy and complex model in the detection of driver distraction behavior, and it is difficult to accurately identify whether the driver's driving behavior is safe.

Method used

Using a driver behavior recognition method based on improved yolo-v8, the driver distracted images are downloaded and preprocessed on the coco data set, image augmentation and feature extraction are performed, deformable neural network and adaptive tracing calculations are used, and the overlapping prediction box and data subsidence problems are solved by combining CIoU loss function.

Benefits of technology

It realizes accurate and rapid identification of driver facial features, improves the accuracy and speed of driver behavior detection, and reduces model complexity and detection error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107936A_ABST
    Figure CN120107936A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of deep learning. The invention relates to a driver behavior recognition method and discloses a driver behavior recognition method based on yo-v8. According to the method, a yolk target detection framework is used. In order to improve the target detection precision, a deformable convolutional neural network is selected. The method is optimized on the basis of an original algorithm. In the aspect of a backbone network, the loss function is greatly improved. The Yolov8 model is pre-programmed, and the same result can be obtained in different environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is an innovative work in the field of traffic management. A novel target detection method based on computer vision is introduced. The method can accurately and quickly identify whether the driver's driving behavior is safe. Bad driving by the driver will cause the car to deviate from the road, resulting in a car accident that may affect the life and property safety of the driver and others. Therefore, the driver's driving behavior is monitored in real time to reduce the driver's bad driving behavior and improve traffic safety. Background Art

[0002] With the rapid development of the economy, the number of cars has increased year by year, and the problem of car driving safety has become increasingly prominent, and the driver's bad driving behavior is the key to the problem of car driving safety. According to the survey results released by the National Highway Traffic Administration of the United States, 80% of traffic accidents are caused by bad driving of drivers. Bad driving of drivers mainly includes: drivers making phone calls, drinking water, chatting, and inattention while driving on the road. In view of the problems of low detection accuracy and complex models in the existing driver distraction behavior detection methods, a driver distraction driving behavior detection based on improved yolo-v8 is proposed. This method can accurately and quickly identify the driver's facial features to ensure the driver's safe driving. Summary of the invention

[0003] In order to solve the problems existing in the background, the present invention proposes a driver behavior recognition method based on yolo-v8 to achieve accurate and rapid recognition of driver facial features.

[0004] A driver behavior recognition method based on yolo-v8 is characterized by comprising the following steps:

[0005] S1: First download driver distraction images on the coco dataset.

[0006] S2: Preprocess the dataset, including image resizing and data classification.

[0007] S3: This data set is divided into nine categories: driver drinks water, driver looks away, driver is on the phone, driver drives normally, driver closes eyes, driver opens eyes, driver wears seat belt, sends text message, and yawns.

[0008] S4: Perform image augmentation on the dataset to enhance the detection of small objects.

[0009] S5: The preprocessed image is passed to the training model yolo-v8n, which has been trained many times and has the ability to accurately recognize the driver's face.

[0010] S6: In the yolo-v8n model, a deformable neural network is used to extract features from images.

[0011] S7: Perform adaptive frame calculation and automatically adjust the frame value for each training.

[0012] S8: In order to solve the overlapping prediction boxes, the model uses a loss function to solve the problem of data redundancy.

[0013] S9: The final data output test results include the detection accuracy and complexity of the backbone network, as well as the detection accuracy and complexity of the loss function.

[0014] Furthermore: the method of S2 preprocesses the data set including adjusting the image size, data augmentation, and data classification.

[0015] Furthermore: the method of S1 performs data augmentation on the data set. This method can achieve more accurate detection of small targets and augment the image by random cropping and random arrangement.

[0016] Further: According to the method of claim 1, the loss function of s5 is: the loss function is to measure the accuracy of the ML model in predicting the expected result and it is also an important basis for calculating the matching degree between the predicted box and the real box. Therefore, CIoU can be expressed as

[0017]

[0018] Where d is the distance between the center of the predicted box and the true box, c is the diagonal distance of the minimum bounding rectangle, and v is a correction factor used to further adjust the loss function, taking into account the shape and direction of the target box.

[0019] Further: the output detection result of S6 includes the detection accuracy and complexity of the backbone network, and the detection accuracy and complexity of the loss function. Accuracy: indicates the reliability of the detection result. Recall: indicates the ability of the model to find the predicted target from the data set. Mean average precision mAP: indicates the mean AP of all categories. The larger the value, the higher the accuracy of the model training.

[0020] The present invention adopts yolo-v8n with fast detection speed and high detection accuracy as the detection model. The driver's driving behavior is monitored in real time using computer vision methods. The computer vision method has the advantage of non-contact with the driver, thereby indirectly avoiding the driver's distraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 The flowchart of the driver behavior detection method of the present invention is

[0022] Figure 2Training flow chart of the method of the present invention

[0023] Figure 3 Comparison training results for the method of the present invention DETAILED DESCRIPTION

[0024] In order to improve the accuracy and effectiveness of the method of the present invention, the present invention is further described in detail below in combination with image data and specific implementation methods.

[0025] Example 1

[0026] like Figure 1 The present invention uses the yolo-v8n algorithm to identify the characteristics of the driver's driving behavior.

[0027] First, download the driver distraction images from the coco dataset in advance. Use labelimg to adjust the image size and mark the distraction behaviors.

[0028] The processed images are passed into the yolo-v8n model, which has been extensively trained and can accurately identify driver distraction.

[0029] In the yolo-v8n model, a deformable convolutional neural network is used to extract features from images, which is used to reduce the effective area of ​​the detection target, reduce the network complexity, and improve the accuracy and robustness of the network.

[0030] In order to make the predicted box and the real box more adaptable, this model uses the loss function CIoU to make the target box regression more stable, so that the detection result is more accurate.

[0031] Example 2

[0032] Figure 2 The training flow chart of the method of the present invention is

[0033] Build the model and build the yolo-v8n model in advance

[0034] Deploy the driver distraction dataset downloaded from the coco dataset to the yolo-v8n model

[0035] Feature extraction from images using deformable convolutional neural networks

[0036] Finally, we get the Driver_Distraction.pt model of driver distraction.

[0037] Implementation Case 3

[0038] Figure 3 The comparison result of the method of the present invention is shown in FIG.

[0039] Comparison of the results after modifying the backbone network in yolo-v8 to Conv convolutional neural network and before modification

[0040] Compare the results of replacing the loss function CIoU in yolo-v8 with EIOU and before the replacement

[0041] Finally, we get the confusion matrix and two evaluation parameters: precision and average precision (mAP).

[0042] Finally, by calculating these three indicators, the final improvement results are obtained.

[0043]

[0044] Precision P is how many samples are correctly predicted as positive samples. FP represents the number of correct predictions, and FN represents the number of incorrect predictions.

[0045]

[0046] The mean average precision (mAP) represents the mean AP of all categories. The larger the value, the higher the accuracy of model training. i It represents the average precision of the i-th category, where i is the number of categories.

Claims

1. A driver behavior recognition method based on yolo-v8, comprising the following steps: S1: First download driver distraction images on the coco dataset. S2: Preprocess the dataset, including image resizing and data classification. S3: This data set is divided into nine categories: driver drinks water, driver looks away, driver is on the phone, driver drives normally, driver closes eyes, driver opens eyes, driver wears seat belt, sends text message, and yawns. S4: Perform image augmentation on the dataset to enhance the detection of small objects. S5: The preprocessed image is passed to the training model yolo-v8n, which has been trained many times and has the ability to accurately recognize the driver's face. S6: In the yolo-v8n model, a deformable neural network is used to extract features from images. S7: Perform adaptive frame calculation and automatically adjust the frame value for each training. S8: In order to solve the overlapping prediction boxes, the model uses a loss function to solve the problem of data redundancy. S9: The final data output test results include the detection accuracy and complexity of the backbone network, as well as the detection accuracy and complexity of the loss function.

2. The method according to claim 1, wherein s2 is characterized in that Dataset collection and labeling.

3. The method according to claim 1, wherein s3 is characterized in that Download 5561 images of driver distraction from the coco dataset.

4. The method according to claim 1, wherein s4 is characterized by the system detecting and augmenting images.

5. The method according to claim 1, characterized in that The loss function of s8 is as follows: The loss function is used to measure the accuracy of the ML model in predicting the expected result. It is also an important basis for calculating the matching degree between the predicted box and the true box. Therefore, CIoU can be expressed as: Where: d is the distance between the center of the predicted box and the true box, c is the diagonal distance of the minimum bounding rectangle, and v is a correction factor used to further adjust the loss function, taking into account the shape and direction of the target box. EIOU can be expressed as: Where: b represents the predicted box, represents the true box, α is a parameter used to balance the ratio, wc and hc are the widths of the minimum bounding rectangles of the predicted bounding box and the true bounding box, and ρ is the Euclidean distance between the two points.

6. The method according to claim 1, characterized in that The S9 outputs the detection results, which include the detection accuracy and complexity of the backbone network, and the detection accuracy and complexity of the loss function.