A fatigue driving recognition method based on improved YOLOv5 neural network
By improving the YOLOv5 neural network model, the DCN convolution module, TA attention module and Wingloss function were introduced, which solved the problem of face detection accuracy in low-light environments and achieved more efficient fatigue driving recognition.
Patent Information
- Application Number
- CN202410646223.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-05-23
AI Technical Summary
The prior art is difficult to accurately identify the driver's facial features in low-light environments, affecting the accuracy of fatigue detection.
Using the improved YOLOv5 neural network model, the DCN convolution module, TA attention module and Wingloss function are introduced to enhance the face detection capability of complex backgrounds and low-light environments.
It improves the detection accuracy of key local areas of the face, enhances the robustness of the system and the ability to capture details, and significantly improves the accuracy of fatigue driving recognition.
Smart Images

Figure CN118644841B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a fatigue driving recognition method based on an improved YOLOv5 neural network. Background Art
[0002] With the rapid development of my country's economy and society and the rise of tourism, the traffic volume on mountain roads has increased significantly. At present, traffic accidents are still one of the main life-threatening risks. Lack of road safety awareness, drunk driving and fatigue driving are all key hidden dangers to traffic safety. The harm of fatigue driving cannot be underestimated. It can significantly slow down the driver's reaction time and judgment, thereby increasing the possibility of accidents. Therefore, how to effectively monitor and prevent fatigue driving has become an important topic in the field of traffic safety research.
[0003] At present, fatigue driving detection methods are mainly divided into contact and non-contact methods. The contact method determines the driver's fatigue state by measuring his physiological indicators, such as heart rate, brain waves, etc. This type of method requires the person being tested to wear various information detection sensors, which not only affects the normal driving of the person being tested, but also has very limited practical applications in special environments such as mountain roads. Non-contact fatigue monitoring technology is less invasive and does not affect driving behavior. This method usually uses an on-board camera to continuously monitor the driver's facial expressions, eye movements, etc. However, the low-light environment common on mountain roads poses a huge challenge to the application of face detection technology. Traditional face detection algorithms often cannot accurately identify the driver's facial features under low-light conditions, thus affecting the accuracy of fatigue detection. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a fatigue driving identification method based on an improved YOLOv5 neural network.
[0005] The objective of the present invention is achieved through the following technical solutions:
[0006] A fatigue driving recognition method based on an improved YOLOv5 neural network, comprising:
[0007] Obtain an image to be recognized;
[0008] Using the trained fatigue driving recognition model, the driver’s fatigue state in the image to be recognized is identified.
[0009] in,
[0010] The fatigue driving recognition model includes face and key point recognition modules and fatigue state recognition modules.
[0011] The face and key point recognition module uses an improved YOLOv5 neural network to detect faces and facial key points;
[0012] The fatigue state recognition module recognizes the driver's fatigue state based on the facial key points determined by the face and key point recognition module.
[0013] Improvements to the YOLOv5 neural network include:
[0014] Use DCN convolutional modules to replace the CBS convolutional modules in the YOLOv5 backbone neural network and the neck network;
[0015] A TA attention module is introduced into the backbone network and the neck of the YOLOv5 neural network. Specifically, in the improved YOLOv5 neural network, the backbone network includes a first DCN convolution module, a second DCN convolution module, a first C3 module, a third DCN convolution module, a second C3 module, a TA attention module, a fourth DCN convolution module, a second C3 module, a fifth DCN convolution module, a third C3 module and an SPPF module in the direction of the signal flow, that is, the TA attention module is introduced between the second C3 module and the fourth DCN convolution module in the backbone network; the neck network includes a first DCN convolution module in the neck, a first upsampling module, a connection module connecting the third C3 module of the backbone network and the first upsampling module of the neck network, a first C3 module in the neck, a second DCN convolution module in the neck, a TA attention module and a second upsampling module in the neck in the direction of the signal flow, that is, the TA attention module is introduced between the second DCN convolution module in the neck and the second upsampling module;
[0016] The Wingloss function is used to replace the loss function of the bounding box regression part in the YOLOv5 neural network.
[0017] Furthermore, the training of the fatigue driving recognition model includes:
[0018] Acquire training sample images, each of which corresponds to a labeling information, wherein the labeling information includes target frame information, face key point information, and fatigue status information;
[0019] The improved YOLOv5 neural network is trained using training sample images and corresponding annotation information.
[0020] Furthermore, the DCN convolution module calculates the offset Δp of the sampling point of the convolution kernel on the input feature map through a convolution layer. n , and bilinear interpolation is used to determine the output feature map after deformable convolution.
[0021] Furthermore, the TA attention module realizes the interaction between dimensions through a three-branch structure, generates interaction results of three dimensions, namely, a first output tensor, a second output tensor and a third output tensor, and then aggregates the three output tensors together through an averaging method to obtain the output tensor of the TA attention module, wherein the TA attention module includes a Z-Pool layer to combine the average pooling feature and the maximum pooling feature, which is expressed by the formula:
[0022] Z-Pool(χ)=[MaxPool 0d (χ),AvgPool 0d (χ)],
[0023] Among them, MaxPool 0d (·) represents the maximum pooling operation, AvgPool 0d (·) represents the average pooling operation, χ represents the input tensor, C corresponds to the channel dimension of the input tensor, i.e., the C dimension, H corresponds to the length dimension of the input tensor, i.e., the H dimension, W corresponds to the width dimension of the input tensor, i.e., the W dimension, and Z-Pool(·) represents the operation of the Z-Pool layer.
[0024] Furthermore, for the interaction between the H dimension and the C dimension, the following operations are performed:
[0025] For the input tensor Rotate 90 degrees counterclockwise along the H axis to obtain the rotation tensor
[0026] Rotate the tensor After processing by the Z-Pool layer, we get the tensor
[0027] The tensor After convolution and batch normalization, the intermediate processing result is formed
[0028] The intermediate processing results The first attention weight is generated after being processed by the sigmoid function.
[0029] Apply the first attention weight to the rotation tensor And rotate 90 degrees clockwise along the H axis to get the first output tensor.
[0030] Furthermore, for the interaction between the C dimension and the W dimension, the following operations are performed:
[0031] For the input tensor Rotate 90 degrees counterclockwise along the W axis to obtain the rotation tensor
[0032] Rotate the tensor After processing by the Z-Pool layer, we get the tensor
[0033] The tensor After convolution and batch normalization, the intermediate processing result is formed
[0034] The intermediate processing results The second attention weight is generated after being processed by the sigmoid function;
[0035] Apply the second attention weight to the rotation tensor Then rotate 90 degrees clockwise along the W axis to get the second output tensor.
[0036] Furthermore, for the interaction between the H dimension and the W dimension, the following operations are performed:
[0037] For the input tensor After processing by the Z-Pool layer, we get the tensor
[0038] The tensor After convolution and batch normalization, the intermediate processing result is formed
[0039] The intermediate processing results The third attention weight is generated after being processed by the sigmoid function;
[0040] Apply the third attention weight to the input tensor χ, resulting in a third output tensor.
[0041] Furthermore, the facial key points include eye key points and mouth key points, and the method also includes determining an eye fatigue state and a mouth fatigue state, and the fatigue driving state is determined by combining the eye fatigue state and the mouth fatigue state.
[0042] Furthermore, the determination of the eye fatigue state includes judging the eye fatigue state using the PERCLOS criterion.
[0043] Furthermore, the determination of the mouth fatigue state includes judging the mouth fatigue state by the number of yawns per unit time.
[0044] The beneficial effects of the present invention are:
[0045] The present invention adopts an improved YOLOv5 neural network model, which uses the DCN convolution module to enhance the adaptability to facial expressions and posture changes under complex backgrounds, and improves the detection accuracy of key local areas (such as eyes and mouth); at the same time, in this model, the introduction of the TA attention mechanism enables the model to focus more on important information in the image, effectively suppresses irrelevant background noise, and enhances the overall robustness of the system; in addition, the high sensitivity of the Wingloss function to small errors helps the model to capture subtle changes in the face more finely.
[0046] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0048] Figure 1 It is an improved YOLOv5 network structure;
[0049] Figure 2 It is a different sampling method of variable convolution;
[0050] Figure 3 It is a schematic diagram of bilinear interpolation;
[0051] Figure 4 It is a schematic diagram of variable convolution;
[0052] Figure 5 It is a schematic flow chart of the implementation of variable convolution;
[0053] Figure 6 This is a schematic diagram of the TA attention mechanism module;
[0054] Figure 7 This is the network structure diagram of TA attention mechanism;
[0055] Figure 8 It is a schematic diagram of the aspect ratio of the eye;
[0056] Fig. 9 It is a diagram of the aspect ratio of the mouth;
[0057] Fig.10 It is based on the eye opening and closing curve of PERCLOS;
[0058] Fig.11 This is the detection effect diagram. DETAILED DESCRIPTION
[0059] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.
[0060] Since the existing YOLOv5 basic model faces certain challenges in facial feature detection and driver fatigue status judgment in complex mountainous road environments, this application proposes a fatigue driving recognition method based on an improved YOLOv5 neural network, which includes:
[0061] Obtain an image to be recognized;
[0062] The trained fatigue driving recognition model is used to identify the driver's fatigue state in the image to be identified, where:
[0063] The fatigue driving recognition model includes face and key point recognition modules and fatigue state recognition modules.
[0064] The face and key point recognition module uses an improved YOLOv5 neural network to detect faces and facial key points;
[0065] The fatigue state recognition module recognizes the driver's fatigue state based on the facial key points determined by the face and key point recognition module.
[0066] In the traditional YOLOv5 basic network, the input image is extracted through the CBS convolution module (i.e., convolution Conv + batch normalization Batch normalization + activation function Silu). However, since the basic model is not flexible enough for the feature extraction of highly deformed targets and targets of different shapes in dynamic environments and low-light environments. Therefore, the present invention uses the DCN convolution module to replace the CBS convolution module in the backbone network and the neck network of the YOLOv5 model.
[0067] Figure 1 This is the improved YOLOv5 neural network structure diagram, refer to Figure 1 The present invention improves the YOLOv5 neural network by using the DCN convolution module (i.e. Figure 1The DNC-Conv module in the YOLOv5 neural network replaces the CBS convolution module in the YOLOv5 neural network. Among them, DCN convolution is Deformable Convolutional Networks convolution, also known as variable convolution; the TA attention module is introduced in the backbone network and neck of the YOLOv5 neural network. Specifically, in the improved YOLOv5 neural network, the backbone network includes the first DCN convolution module, the second DCN convolution module, the first C3 module, the third DCN convolution module, the second C3 module, the TA attention module, the fourth DCN convolution module, the second C3 module, the fifth DCN convolution module, the third C3 module and the SPPF module in the direction of signal flow, that is, the TA attention module is introduced between the second C3 module and the fourth DCN convolution module; in the neck network, the neck includes the first DCN convolution module, the first upsampling module (that is, Figure 1 Upsample in the neck network), the connection module connecting the third C3 module of the backbone network and the first upsampling module of the neck network (i.e. Figure 1 Contact in the neck), the first C3 module in the neck, the second DCN convolution module in the neck, the TA attention module, and the second upsampling module, that is, the TA attention module in the neck is located between the second DCN convolution module in the neck and the second upsampling module.
[0068] It should be noted that both the C3 module and the SPPF module in the YOLOv5 neural network have CBS convolution modules. The CBS convolution modules in the C3 module and the SPPF module are not replaced in the improved YOLOv5 neural network model, and the structure of the C3 module and the SPPF module in the existing YOLOv5 is the same.
[0069] When the feature map of the previous layer arrives, the DCN convolution module will cover the pixel area corresponding to the feature map according to the size of the convolution kernel; use the learned offset to adjust the sampling position of the convolution kernel on the feature map; control the stride of the convolution kernel movement and the filling method of the feature map edge through the step size and padding parameters; call the weight and bias parameters to complete the feature map related calculations; finally, the processed feature map is output by the output channel.
[0070] Figure 2 are different sampling methods of variable convolution, where Figure 2 (a) is the standard 3×3 convolution kernel sampling method; Figure 2 (b) is the change of deformable convolution sampling points after offset adjustment; Figure 2 (c) and Figure 2 (d) shows two specific examples of convolution with variability (i.e., the offsets show a certain regularity).
[0071] Deformable convolution relies on the offsets learned by the network, which make the sampling points of the convolution kernel on the input feature map adaptively adjusted according to the needs of the target area. In this method, each sampling point is assigned an offset calculated by the feature map through additional convolutional layers. These offsets are usually expressed in decimal form, so that the convolution operation can focus on the key areas in the image more accurately.
[0072] The variable convolution operation is shown in formula (1):
[0073]
[0074] Among them, p0 represents any point on the input feature map, p n Represents the offset of each point in the convolution kernel relative to the center point, ω(p n ) indicates that the offset in the convolution kernel is p n The weight corresponding to the ground position, x(·) represents the pixel value on the input feature map, R0 represents the set of points in the convolution kernel, and y(p0) represents the pixel value at position p0 on the output feature map.
[0075] After the offset is introduced, since the offset is usually a decimal, the offset pixel points can be obtained by bilinear interpolation to replace the actual pixel points on the input feature map.
[0076] Bilinear interpolation means setting weighted weights by calculating the distances of the four nearest actual pixels around the interpolation point on the horizontal and vertical coordinates, and then setting the pixel value of the interpolation point to the weighted sum of these four points to finally obtain the pixel value of the interpolation point. Then the calculation of f at the interpolation point P = (x, y) is as follows. First, perform linear interpolation in the x direction:
[0077]
[0078] Then do linear interpolation in the y direction and calculate as follows:
[0079]
[0080] The value of point P = (x, y) is obtained by combining the results of two linear interpolations.
[0081]
[0082] Figure 3 This is a schematic diagram of bilinear interpolation. Figure 3 In the above formula, x and y represent the length and width, f() is the pixel value of the corresponding point, and Q 11 =(x1,y1),Q 12=(x1,y2),Q 21 =(x2,y1),Q 22 =(x2,y2), the weighted sum of these four points is used to calculate the pixel value of point P, and the weight of each point is determined by the distance from each point to point P.
[0083] Figure 4 The schematic diagram of deformable convolution is shown. In this structure, the offsets are generated by a separate convolution layer, which is responsible for calculating the offsets and does not participate in the final convolution operation. Figure 4 The N in represents the size of the convolution kernel. For example, for a 3×3 convolution kernel, the value of N is 9. Figure 4 The process along the arrow direction in the upper part of the figure represents the process of this dedicated convolution layer learning the offset, where the number of channels of the offset field is 2N, which means that each convolution kernel position learns the offset in two directions (i.e., x and y directions).
[0084] like Figure 4 As shown in the figure, on the input feature map, the sampling area of the standard convolution is usually a square area (indicated by the red box) that matches the size of the convolution kernel, while the sampling area of the deformable convolution is realized by the offset points shown in the small blue rectangular box, which are adjusted according to the learned offset. This shows the main difference between deformable convolution and traditional convolution, that is, deformable convolution can dynamically adjust the sampling position according to the actual needs of the feature.
[0085] according to Figure 4 , the size of the convolution sampling area corresponding to a point on an output feature map to the input feature map is K×K. According to the deformable convolution operation, each convolution sampling point in this K×K area must learn an offset, and the offset is expressed in coordinates, so one output needs to learn 2×K×K parameters. Assuming that the size of an output is H×W, a total of 2×K×K×H×W parameters need to be learned.
[0086] Figure 5This is the implementation flow chart of variable convolution, where B is Batch size, which means the batch size, that is, the number of samples input into the model at one time; inC is Number of input channels, which means the number of channels of the input feature map; H is Height, which means the height of the feature map; W is Width, which means the width of the input feature map; kernel=3, which means the convolution size is 3×3; stride=1, which means the stride is 1; Input(B,inC,H,W) means the dimension and shape of the input feature map.
[0087] First, input Input and calculate the offset (i.e. Δp in Formula 1). n ), and then find p_n (i.e. p in formula 1 n ) and p_0 (i.e. p0 in formula 1), and then p=p_n+p_0+offset is obtained, and then p is rotated and deformed, and then bilinear interpolation is performed, and the result after bilinear interpolation is rotated and deformed again, and finally the output is obtained through convolution.
[0088] Figure 5 In the example, the offset field (N=K×K) has a dimension of B×2×K×K×H×W.
[0089] Assume that the dimension of the input feature map is B×C×H×W, and the feature maps in a batch (a total of C) share an offset field, that is, the offset used by each feature map in a batch is the same.
[0090] Since the deformable convolution does not change the size of the input feature map, the size of the output feature map is also H×W.
[0091] The improved YOLOv5 model uses the DCNC (Deformable Convolutional Networks Conv, also known as DCN convolution) module to realize image feature extraction and enhance the ability to capture targets with irregular shapes and complex backgrounds.
[0092] In some embodiments of the present invention, the improvement of YOLOv5 also includes introducing a TA attention (Triplet Attention) module (e.g., triple attention) in the backbone network (i.e., Backbone) and the neck (i.e., neck) of the YOLOv5 neural network. Figure 1Specifically, a TA attention module is introduced between the second C3 module and the fourth DCN-Conv convolution module in the backbone network, and a TA attention mechanism is introduced between the second DCN-Conv convolution module and the second Upsample module in the neck.
[0093] The TA attention module realizes the interaction between dimensions through a three-branch structure, generates three-dimensional interaction results, namely the first output tensor, the second output tensor and the third output tensor, and then aggregates the three output tensors together through the averaging method to obtain the output tensor of the TA attention module. The TA attention module activates the interdependence between dimensions by performing a rotation operation on the input tensor, and then uses a residual transformation to further strengthen this dependency, which can more accurately focus on important features in the image and suppress irrelevant background noise.
[0094] Figure 6 This is a schematic diagram of the TA attention mechanism module. Figure 6 As shown in Figure 1, the TA attention module contains a Z-Pool layer, which is specially designed to reduce the dimension of the tensor while retaining its rich representation. It further reduces the computational burden by combining the average pooled features and the maximum pooled features. This structural design not only ensures the secure interaction of information, but also optimizes the computational efficiency. The Z-Pool layer operation is expressed by formula (5):
[0095] Z-Pool(χ)=[MaxPool 0d (χ),AvgPool 0d (χ)], (5)
[0096] Among them, MaxPool 0d (·) represents the maximum pooling operation, AvgPool 0d (·) represents the average pooling operation, χ represents the input tensor, C corresponds to the channel dimension of the input tensor, i.e., the C dimension, H corresponds to the length dimension of the input tensor, i.e., the H dimension, W corresponds to the width dimension of the input tensor, i.e., the W dimension, and Z-Pool(·) represents the operation of the Z-Pool layer.
[0097] The TA attention mechanism is a complex network structure that achieves effective interaction between different dimensions by performing specific transformations on the input tensor. The TA mechanism consists of three main branches, each responsible for capturing the associations between different dimensions.
[0098] Figure 7 This is a schematic diagram of the TA attention mechanism network structure. Figure 7 , for the interaction between the H dimension and the C dimension ( Figure 7 ), do the following:
[0099] For the input tensor Rotate 90 degrees counterclockwise along the H axis to obtain the rotation tensor
[0100] Rotate the tensor After processing by the Z-Pool layer, we get the tensor
[0101] The tensor After convolution and batch normalization, the intermediate processing result is formed
[0102] The intermediate processing results The first attention weight is generated after being processed by the sigmoid function.
[0103] Apply the first attention weight to the rotation tensor And rotate 90 degrees clockwise along the H axis to get the first output tensor, whose size is C×H×W, and its shape is the same as the input tensor.
[0104] For the interaction between C and W dimensions ( Figure 7 In the middle branch), do the following:
[0105] For the input tensor Rotate 90 degrees counterclockwise along the W axis to obtain the rotation tensor
[0106] Rotate the tensor After processing by the Z-Pool layer, we get the tensor
[0107] The tensor After convolution and batch normalization, the intermediate processing result is formed
[0108] The intermediate processing results The second attention weight is generated after being processed by the sigmoid function;
[0109] Apply the second attention weight to the rotation tensor Then rotate 90 degrees clockwise along the W axis to obtain the second output tensor, whose size is also C×H×W.
[0110] For the interaction between H and W dimensions ( Figure 7 ), do the following:
[0111] For the input tensor After processing by the Z-Pool layer, we get the tensor
[0112] The tensor After convolution and batch normalization, the intermediate processing result is formed
[0113] The intermediate processing results The third attention weight is generated after being processed by the sigmoid function;
[0114] Applying the third attention weight to the input tensor χ results in a third output tensor, also of size C×H×W.
[0115] After obtaining the above three output tensors, the three output tensors are aggregated together by the averaging method to obtain the output tensor y0 of the TA attention module, which is expressed as:
[0116]
[0117] In formula (6), σ(·) represents the sigmoid function, ψ1(·), ψ2(·) and ψ3(·) all represent convolution and batch normalization processing (i.e., batch normalization); the upper horizontal line in the first two terms connected by the plus sign “+” in the brackets in formula (6) represents a rotation operation (i.e., rearranging the matrix).
[0118] In some embodiments of the present invention, the improvement of YOLOv5 also includes replacing the loss function of the bounding box regression part in the YOLOv5 neural network with the Wingloss function, so as to be used in application scenarios with high sensitivity to small errors in face key point detection. The core idea of Wingloss is to provide a larger gradient when the error is small, which helps the network to adjust the predicted position of the key point more accurately. The loss function consists of two parts: a linear part and a nonlinear part. Through this structure, the stability and convergence speed in the learning process can be effectively balanced. The calculation formula of the Wingloss function is shown in formula (7):
[0119]
[0120] In formula (7), y represents the true value of the feature point, Represents the predicted value of the feature point, ω and ε are parameters that control the shape of the curve, and C0 is a constant calculated from ω and ε, which is used to provide a smooth transition between the two parts of the loss function.
[0121] In some embodiments, the key points of the face include eye key points and mouth key points. The present invention first uses the eye key points and mouth key points to judge the eye fatigue state and mouth fatigue state respectively, and then determines the fatigue driving state by combining the eye fatigue state and mouth fatigue state.
[0122] After the face is recognized, the model can also mark the key feature points of the eyes on the face, and then calculate the eye aspect ratio (EAR) based on the Euclidean distance, which is used as a criterion for judging the state of eye fatigue. If the calculated EAR value is greater than the set state threshold, it means that the eyes are open at this time. Figure 8 This is a diagram of the eye aspect ratio, for reference Figure 8 , the EAR calculation formula is as follows:
[0123]
[0124] After completing face detection, the model will mark the key feature points of the mouth, and then use the Euclidean distance to calculate the mouth aspect ratio (MAR), which is used as a criterion for evaluating the mouth state. If the calculated MAR value is greater than the set state threshold, it means that the eyes are open at this time. The MAR calculation formula is as follows:
[0125]
[0126] In some embodiments of the present invention, the determination of the eye fatigue state includes using the PERCLOS criterion to judge the eye fatigue state. PERCLOS, i.e., Percentage Of Eyelid Closure Over The Pupil Over Time, is defined as the proportion of pupils covered by eyelid closure within a certain period of time, and is one of the important indicators for evaluating fatigue or sleepiness. Fig.10 It is based on the PERCLOS eye proportional curve. This indicator is measured by quantifying the duration of the degree of eye closure in a specific period of time. It is calculated as follows:
[0127]
[0128] In formula (10), N closeFrame Indicates the number of frames in which the eyes are closed within a given time, N totalFrameIndicates the total number of frames in the same time period. The PERCLOS method uses the EM, P70, and P80 criteria to determine whether the eyes are open or closed during that time. Specifically, when the degree of eye closure reaches or exceeds 50%, 70%, or 80%, the person's eyes are considered to be closed according to the corresponding criteria. Among these criteria, P80 is considered to have the strongest correlation with fatigue status, so the P80 standard is used as the basis for assessing whether the subject is in a state of fatigue.
[0129] In some embodiments of the present invention, the determination of the mouth fatigue state includes judging the mouth fatigue state by the number of yawns per unit time. If the PYawn value exceeds a set threshold, it is determined that the person being tested is in a fatigue state. The fatigue parameter formula of the mouth state is as follows:
[0130]
[0131] In formula (11), N Yawn represents the number of times the subject yawns, and T represents the total time.
[0132] The data set is strictly selected based on the PERCLOS parameter. When the degree of eye closure is greater than or equal to 80%, the eye state is marked as closed, otherwise it is marked as open. In the detection and judgment process, there may be cases where eye details are blocked or cannot be detected. Therefore, the number of yawns per unit time is added as an auxiliary basis for fatigue judgment, and the two fatigue judgment methods are combined as the fatigue driving judgment standard.
[0133] The present invention also provides a method for training a fatigue driving recognition model, the method comprising:
[0134] Acquire training sample images, each of which corresponds to a labeling information, wherein the labeling information includes target frame information, face key point information, and fatigue status information;
[0135] The improved YOLOv5 neural network is trained using training sample images and corresponding annotation information.
[0136] In order to better illustrate the advantages of the present invention, specific experimental results are provided below.
[0137] First, the accuracy and average accuracy are defined and calculated as follows:
[0138]
[0139]
[0140] In formula (12), N TP,lIndicates the number of correct recognitions in the lth state; N FP,l represents the number of misidentifications in the lth state; P l represents the accuracy under l states; in formula (13), P mA It represents the average accuracy.
[0141] The dataset used for driver face detection comes from the public WIDER FACE dataset. The WIDER FACE dataset selected 32,203 images for face annotation, and a total of 393,703 faces were labeled, with an average of 12 faces per image. Each face is annotated with detailed category information, including expression, illumination, blur, occlusion, and pose. In view of the particularity of the background mountainous roads in this experiment, a dataset related to the faces of drivers on mountainous roads was built using the annotation tool LabelImg. These two datasets can effectively simulate the face detection scene during driving on mountainous roads, and can more realistically evaluate the accuracy of face detection of the improved model.
[0142] Table 1 Comparison of experimental results
[0143]
[0144] From the experimental results in Table 1, we can see that the basic model YOLOv5, after training with the training data set, has an average precision (PmA) of 73.8%. The average detection accuracy of the YOLOv5 model is improved by 4.4% by replacing the YOLOv5 feature extraction convolution module with a deformable convolution module, adding the TA attention mechanism, and using the Wingloss function. Through the analysis of the experimental results, after improving and optimizing the modules at all levels of the original model and adopting a more scientific and efficient training strategy, the performance of the model in detecting difficult targets can be effectively improved, and the average accuracy of target detection can be improved. In the face of different scenarios, the accuracy and speed of the model can be weighed, and the original model can be improved accordingly, so as to find a more suitable and efficient detection model.
[0145] The dataset used for fatigue detection comes from the public YawDD video dataset. The YawDD dataset covers both wearing glasses and not wearing glasses. The dataset simulates four common driving scenarios: silence, speaking, fatigue yawning, and singing. Through the study and investigation of human fatigue status, this experiment uses eye features and yawning fatigue status as the main basis for fatigue judgment. In the image detection fatigue experiment, 2,000 images are extracted through key frames, and the extracted images are divided into four categories: open eyes, closed eyes, open mouth, and closed mouth using the Labelmg annotation tool. The image resolution is uniformly set to 640×640, and training and verification are carried out in a ratio of 6:4. In the video detection fatigue experiment, 30 awake video samples and 40 fatigue video samples are obtained through editing processing to determine the fatigue status of the driver.
[0146] Table 2 Comparison of detection results of different models
[0147]
[0148] As shown in Table 2, the fatigue detection and recognition accuracy of the YOLOv5 basic model in the image data extracted from the YawDD video dataset is 82.5%, while the DCN-YOLOv5+TA+Wingloss model has an accuracy of 89.9%. From the image data experiment, it can be seen that the DCN-YOLOv5+TA+Wingloss model has better detection performance.
[0149] Table 3 DCN-YOLOv5+TA+Wingloss algorithm detection results
[0150]
[0151] The edited YawDD dataset is used to measure the performance of the model. The dataset not only marks the mental state of the test subjects in different videos, but also includes common scenes on mountain roads, which is closer to the real detection environment. This dataset can be used to more effectively test and improve the performance of the model. Fig.11 It is a detection effect diagram, wherein the detection results are shown in Table 3. 25 out of 30 samples in the awake state are judged correctly, with an accuracy rate of 83.3%. 36 out of 40 samples in the fatigue state are judged correctly, with an accuracy rate of 90.0%. The average accuracy rate of the improved fatigue judgment model of the present invention is 86.65%.
[0152] The edited YawDD dataset is used to measure the performance of the model. The dataset not only marks the mental state of the test subjects in different videos, but also includes common scenes on mountain roads, which is closer to the real detection environment. This dataset can be used to more effectively test the performance of the improved model. The comparison of the detection results is shown in Table 3. Among the 30 samples in the awake state, 25 samples are correctly judged, with an accuracy of 83.3%. Among the 40 samples in the fatigue state, 36 samples are correctly judged, with an accuracy of 90.0%. The average accuracy of the fatigue judgment model improved in this paper is 86.65%.
[0153] The present invention aims to explore the problem of determining the driver's fatigue state in a mountainous road environment. Due to factors such as insufficient light and tree branch shadows that are unique to mountainous areas, the driver's face will be affected, the noise of the video or image will be increased, and thus face detection and fatigue state recognition will be challenged. In order to cope with these difficulties, a fatigue detection model combining a DCN convolution module, a TA attention mechanism and a Wingloss function is proposed. The DCN convolution module enhances the adaptability to changes in facial expressions and postures in complex backgrounds and key areas (such as eyes, mouth, etc.); the TA attention mechanism helps the model focus on key information in the image, thereby ignoring irrelevant background noise and improving robustness; the Wingloss function makes the model more sensitive to small errors and improves the ability to capture details. In the embodiments of the present invention, the WIDER FACE dataset and the YawDD dataset are used to train and verify the model face detection and fatigue judgment, respectively, and the PERCLOS value of the eye state and the yawning state are used as the basis for fatigue judgment. The experimental results show that the model performs well in the recognition of fatigue driving state.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A fatigue driving identification method based on an improved YOLOv5 neural network, characterized in that: include: Obtain an image to be recognized; Using the trained fatigue driving recognition model, the driver’s fatigue state in the image to be recognized is identified. in, The fatigue driving recognition model includes face and key point recognition modules and fatigue state recognition modules. The face and key point recognition module uses an improved YOLOv5 neural network to detect faces and facial key points; The fatigue state recognition module recognizes the driver's fatigue state based on the facial key points determined by the face and key point recognition module. Improvements to the YOLOv5 neural network include: Use the DCN convolution module to replace the CBS convolution module in the backbone network and the neck network in the YOLOv5 neural network; A TA attention module is introduced into the backbone network and the neck of the YOLOv5 neural network. Specifically, in the improved YOLOv5 neural network, the backbone network includes a first DCN convolution module, a second DCN convolution module, a first C3 module, a third DCN convolution module, a second C3 module, a TA attention module, a fourth DCN convolution module, a second C3 module, a fifth DCN convolution module, a third C3 module and an SPPF module in the direction of the signal flow, that is, the TA attention module is introduced between the second C3 module and the fourth DCN convolution module in the backbone network; the neck network includes a first DCN convolution module in the neck, a first upsampling module, a connection module connecting the third C3 module of the backbone network and the first upsampling module of the neck network, a first C3 module in the neck, a second DCN convolution module in the neck, a TA attention module and a second upsampling module in the neck in the direction of the signal flow, that is, the TA attention module is introduced between the second DCN convolution module in the neck and the second upsampling module; The Wingloss function is used to replace the loss function of the bounding box regression part in the YOLOv5 neural network.
2. The fatigue driving identification method based on the improved YOLOv5 neural network according to claim 1 is characterized in that: The training of the fatigue driving recognition model includes: Acquire training sample images, each of which corresponds to a labeling information, wherein the labeling information includes target frame information, face key point information, and fatigue status information; The improved YOLOv5 neural network is trained using training sample images and corresponding annotation information.
3. The fatigue driving identification method based on the improved YOLOv5 neural network according to claim 1 is characterized in that: The DCN convolution module calculates the offset Δp of the sampling point of the convolution kernel on the input feature map through a convolution layer. n , and bilinear interpolation is used to determine the output feature map after deformable convolution.
4. The fatigue driving identification method based on the improved YOLOv5 neural network according to claim 1 is characterized in that: The TA attention module realizes the interaction between dimensions through a three-branch structure, generates the interaction results of three dimensions, namely the first output tensor, the second output tensor and the third output tensor, and then aggregates the three output tensors together through the averaging method to obtain the triple attention output tensor of the TA attention module, wherein the TA attention module includes a Z-Pool layer to combine the average pooling feature and the maximum pooling feature, which is expressed by the formula: Z-Pool(x)=[MaxPool 0d (x),AvgPool 0d (x)], Among them, MaxPool 0d (·) represents the maximum pooling operation, AvgPool 0d (·) represents the average pooling operation, χ represents the input tensor, C corresponds to the channel dimension of the input tensor, i.e., the C dimension, H corresponds to the length dimension of the input tensor, i.e., the H dimension, W corresponds to the width dimension of the input tensor, i.e., the W dimension, and Z-Pool(·) represents the operation of the Z-Pool layer.
5. A fatigue driving identification method based on an improved YOLOv5 neural network according to claim 4, characterized in that: For the interaction between the H dimension and the C dimension, perform the following operations: For the input tensor Rotate 90 degrees counterclockwise along the H axis to obtain the rotation tensor Rotate the tensor After processing by the Z-Pool layer, we get the tensor The tensor After convolution and batch normalization, the intermediate processing result is formed The intermediate processing results The first attention weight is generated after being processed by the sigmoid function. Apply the first attention weight to the rotation tensor And rotate 90 degrees clockwise along the H axis to get the first output tensor.
6. A fatigue driving identification method based on an improved YOLOv5 neural network according to claim 4, characterized in that: For the interaction between the C and W dimensions, perform the following operations: For the input tensor Rotate 90 degrees counterclockwise along the W axis to obtain the rotation tensor Rotate the tensor After processing by the Z-Pool layer, we get the tensor The tensor After convolution and batch normalization, the intermediate processing result is formed The intermediate processing results The second attention weight is generated after being processed by the sigmoid function; Apply the second attention weight to the rotation tensor Then rotate 90 degrees clockwise along the W axis to get the second output tensor.
7. The fatigue driving identification method based on the improved YOLOv5 neural network according to claim 4 is characterized in that: For the interaction between the H dimension and the W dimension, perform the following operations: For the input tensor After processing by the Z-Pool layer, we get the tensor The tensor After convolution and batch normalization, the intermediate processing result is formed The intermediate processing results The third attention weight is generated after being processed by the sigmoid function; Apply the third attention weight to the input tensor χ, resulting in a third output tensor.
8. The fatigue driving identification method based on the improved YOLOv5 neural network according to claim 1 is characterized in that: The facial key points include eye key points and mouth key points. The method also includes determining an eye fatigue state and a mouth fatigue state, and the fatigue driving state is determined by combining the eye fatigue state and the mouth fatigue state.
9. The method for identifying fatigue driving based on an improved YOLOv5 neural network according to claim 8, characterized in that: The determination of the eye fatigue state includes using the PERCLOS criterion to judge the eye fatigue state.
10. The method for identifying fatigue driving based on an improved YOLOv5 neural network according to claim 8, characterized in that: The determination of the mouth fatigue state includes judging the mouth fatigue state by the number of yawns per unit time.