Method for detecting dangerous driving behavior based on line-of-sight direction time relationship learning

By combining head and eye orientation with time relationship learning, a time-localization network for gaze direction is constructed, which solves the problem of inaccurate driver gaze state localization and enables effective detection and warning of dangerous driving behaviors.

CN115661800BActive Publication Date: 2026-01-09HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211366926.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2026-01-09
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

Existing technologies cannot effectively combine driver's line of sight and time information during driving, resulting in inaccurate positioning of line of sight direction and difficulty in providing early warning of dangerous driving behaviors.

Method used

A joint network of head orientation and eye orientation was designed. By learning the temporal relationship between gaze direction, a temporal positioning network for gaze direction was constructed to continuously detect the driver's gaze state and issue a warning when the duration of a dangerous gaze state exceeds the safe duration.

Benefits of technology

It can handle situations where the head and eyes are not aligned, and robustly handle changes in gaze direction over time, thus achieving effective detection of dangerous driving behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661800B_ABST
    Figure CN115661800B_ABST
Patent Text Reader

Abstract

The application discloses a dangerous driving behavior detection method based on a line-of-sight direction time relationship learning. A convolutional neural network is designed to estimate the head orientation and the binocular orientation of a driver. In view of the possible inconsistency between the head orientation and the binocular orientation, a head orientation and binocular orientation joint network is designed to estimate the line-of-sight direction of the driver. In view of the problem that the line-of-sight direction changes with time during driving and it is difficult to accurately determine the dangerous line-of-sight direction state, a Gaussian time weight-based line-of-sight direction time relationship learning is designed, a line-of-sight direction time positioning network is constructed, and reliable time positioning of the dangerous line-of-sight direction is realized. When the duration of the dangerous line-of-sight direction exceeds a threshold value, a safety warning is given to the driver. The application can handle the inconsistency between the head orientation and the binocular orientation, can robustly handle different line-of-sight direction time change processes, and can effectively realize dangerous driving behavior detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of driver line-of-sight direction detection and time positioning, and particularly relates to a dangerous driving behavior detection method based on line-of-sight direction time relationship learning. BACKGROUND

[0002] In order to ensure good traffic order and the safety of people's lives and property, it is necessary to monitor the dangerous driving behavior of a driver who is driving. With the rapid development of deep learning and computer vision, methods for detecting dangerous driving behavior based on video information are gradually attracting attention in the industry.

[0003] Chinese Patent Application Publication No. CN114005093A, "Driving Behavior Warning Method and Device Based on Video Analysis, Equipment and Medium", proposes a driving behavior warning method based on video analysis. The position information, speed information and driving trajectory information of the target vehicle and other vehicles in the image are labeled according to the object features, and a plurality of dangerous driving features are pre-acquired to identify images in which the target vehicle exhibits dangerous driving behavior. When the number of images in which the target vehicle exhibits dangerous driving behavior in a preset unit of time is greater than a preset threshold, the driver of the target vehicle is warned. However, this method only uses external information of the vehicle for detection and does not combine the driving state of the driver, so it cannot achieve early warning. Chinese Patent Application Publication No. CN113942450A, "Vehicle-mounted Intelligent Driving Warning System and Vehicle", proposes a vehicle-mounted intelligent driving warning system and vehicle, in which a line-of-sight detection module is used to obtain the line-of-sight state of the driver to control the warning execution module. However, the prediction of the line-of-sight state of the driver does not combine time information, and when the line-of-sight state of the driver changes greatly, the information obtained by this method will be lost, thereby causing misjudgment.

[0004] Kellnhofer et al. in "Gaze360: Physically Unconstrained Gaze Estimation in the Wild" proposed a time-based line-of-sight estimation model and an error estimation loss function, which extracted a relatively reliable line-of-sight direction. Eunji Chong et al. in "Detecting Attended Visual Targets in Video" solved the problem of detecting attention targets in video, identifying where each person in each frame of the video is looking, and correctly handling cases where the gaze target is outside the frame.

[0005] However, as Figure 8The above method does not consider the inconsistency of head orientation and eye orientation in the line-of-sight direction that may occur, and the problem that it is difficult to accurately locate the dangerous line-of-sight direction state in the driving process when the line-of-sight direction changes greatly over time. Therefore, we propose a safe driving behavior detection method based on line-of-sight direction time relationship learning. The method designs a head orientation and eye orientation joint network to estimate the driver's line-of-sight direction, designs a Gaussian time weight based line-of-sight direction time relationship learning, constructs a line-of-sight direction time positioning network, and realizes reliable time positioning of dangerous line-of-sight direction. SUMMARY

[0006] The purpose of the present application is to overcome the shortcomings of the prior art and provide a dangerous driving behavior detection method based on line-of-sight direction time relationship learning.

[0007] The present application is implemented by the following technical solutions:

[0008] A dangerous driving behavior detection method based on line-of-sight direction time relationship learning, continuously detects the line-of-sight state of the driver and the time positioning of the state during the driving process of the driver, and when the line-of-sight state is in a dangerous line-of-sight state and the duration is greater than a safe duration, the driver is reminded, and specifically includes the following steps:

[0009] Step 1, input the safe driving data set, train the head orientation estimation network, and obtain the head orientation estimation network parameter model;

[0010] Step 2, input the safe driving data set, train the eye line-of-sight direction estimation network, and obtain the eye line-of-sight direction estimation network parameter model;

[0011] Step 3, input the safe driving data set, train the head and eye joint line-of-sight direction estimation network, and obtain the head and eye joint line-of-sight direction estimation network parameter model;

[0012] Step 4, input the safe driving data set, train the line-of-sight state time positioning network, and obtain the line-of-sight state time positioning network parameter model;

[0013] Step 5, during the driving process of the driver, estimate the line-of-sight state time positioning of the driver, and detect the dangerous driving behavior of the driver.

[0014] Step 1, input the safe driving data set, train the head orientation estimation network, and obtain the head orientation estimation network parameter model, specifically including the following steps:

[0015] Step 1-1: input the head detection data set, and train the head detection network model based on Yolov5;

[0016] Step 1-2: Input the safe driving training set, use the head detection network model trained in step 1-1 to detect the head region of the input image, and crop the head region image after detection;

[0017] Step 1-3: Normalize the head region image obtained in step 1-2 to make the size uniform, and obtain the image center point O head , and the head center point is represented by the image center point O head , and a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis;

[0018] Step 1-4: The normalized head region image is input into the head orientation estimation network, which is composed of a ResNet-34 network and three fully connected layers, to obtain a head orientation vector , respectively representing the horizontal coordinate and the vertical coordinate;

[0019] Step 1-5: Calculate the loss function between the head orientation vector obtained in step 1-4 and the true vector , respectively representing the horizontal coordinate and the vertical coordinate, and the loss function formula is as follows:

[0020]

[0021] Step 1-6: Use the loss function of step 1-5 to train the head orientation estimation network in step 1-4 to obtain the head orientation estimation network parameter model.

[0022] Step 2: Input the safe driving data set to train the dual eye gaze direction estimation network to obtain the dual eye gaze direction estimation network parameter model, which includes the following steps:

[0023] Step 2-1: Input the human eye detection data set to train the left eye detection network model based on Yolov5;

[0024] Step 2-2: Input the safe driving training set, use the method of step 1-2 to obtain the head region image, and then use the left eye detection network model trained in step 2-1 to detect the left eye region in the head region image, and crop the left eye region image after detection;

[0025] Step 2-3: Normalize the left eye region image obtained in step 2-2 to make the size uniform, and obtain the image center point O left_eye , and the left eye center point is represented by the image center point O left_eye , and a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis;​

[0026] Step 2-4: The normalized left eye region image in step 2-3 passes through the left branch of the binocular gaze direction estimation network, which is composed of a ResNet-18 network, to generate the gaze direction vector of the left eye respectively represent the horizontal coordinate and the vertical coordinate;

[0027] Step 2-5: Input the human eye detection dataset to train the right eye detection network model based on Yolov5;

[0028] Step 2-6: Input the safe driving training set, use the method in step 1-2 to obtain the head region portrait, and then use the right eye detection network model trained in step 2-5 to detect the right eye region in the head region image, and obtain the right eye region image after cropping;

[0029] Step 2-7: The right eye region image obtained in step 2-6 is normalized respectively to make the size uniform, and the image center point O is obtained right_eye , the right eye center point is represented by O right_eye , and a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis;

[0030] Step 2-8: The normalized right eye region image in step 2-7 passes through the right branch of the binocular gaze direction estimation network, which is composed of a ResNet-18 network, to generate the gaze direction vector of the right eye respectively represent the horizontal coordinate and the vertical coordinate;

[0031] Step 2-9: The left eye gaze direction vector obtained in step 2-4 and the right eye gaze direction vector obtained in step 2-8 pass through a multilayer perceptron φ eye with one hidden layer to generate a binocular gaze direction vector respectively represent the horizontal coordinate and the vertical coordinate:

[0032] α bin_eye = φ eye (α left_eye ,α right_eye )

[0033] Step 2-10: Calculate the loss function between the binocular gaze direction vector in step 2-9 and the real vector , and the loss function formula is as follows:

[0034] ​

[0035] Step 2-11: Train the binocular gaze direction estimation network composed of step 2-4, step 2-8 and step 2-9 using the loss function in step 2-10 to obtain the binocular gaze direction estimation network parameter model.

[0036] Step 3: Input the safe driving data set to train the head and binocular joint gaze direction estimation network to obtain the head and binocular joint gaze direction estimation network parameter model, which specifically includes the following steps:

[0037] Step 3-1: Input the safe driving training set, use the head detection network model trained in step 1-1 to detect the head region of the input image, and cut the head region image after cutting;

[0038] Step 3-2: Use the head orientation estimation network model trained in step 1 to extract the head orientation vector

[0039] Step 3-3: Use the binocular gaze direction estimation network model trained in step 2 to extract the binocular gaze direction vector

[0040] Step 3-4: Obtain the head orientation vector head and the binocular gaze direction vector bin_eye obtained in step 3-3 through a multilayer perceptron φ(·) containing a hidden layer, and the output result is represented as the normalized head and binocular joint gaze direction vector represent the horizontal and vertical coordinates respectively:

[0041] α union =φ(α head ,α bin_eye )

[0042] Step 3-5: Calculate the loss function between the head and binocular joint gaze direction vector obtained in step 3-4 and the real vector , and the loss function formula is as follows:

[0043]

[0044] Step 3-6: Train the head and binocular joint gaze direction estimation network using the loss function in step 3-5 to obtain the head and binocular joint gaze direction estimation network parameter model.

[0045] The gaze state time positioning network is trained by inputting the safe driving data set, and a gaze state time positioning network parameter model is obtained, and the specific steps include the following steps:

[0046] Step 4-1: Input the safe driving data set, and continuously sample the original video containing the driver's head to obtain a video frame sequence;

[0047] Step 4-2: Using the head and binocular joint gaze direction estimation network model trained in step 3, the head and binocular joint gaze direction of the driver in each video frame is estimated by the method of steps 3-1 to 3-4, and the joint gaze direction vector of the driver in all video frames is obtained

[0048] Step 4-3: The joint gaze direction of the driver in all video frames obtained in step 4-2 is converted into a gaze angle feature, and the conversion formula is as follows:

[0049] Step 4-4: The gaze angle features of the driver in all video frames obtained in step 4-3 are combined to form a gaze angle feature sequence T is the time length, t is a time, t∈{1,2,...,T}, θ t represents the gaze angle feature of the driver in the video frame at time t;

[0051] Step 4-5: The gaze angle feature sequence Θ obtained in step 4-4 is passed through two one-dimensional convolution layers, one maximum pooling layer and one one-dimensional time sequence convolution layer to obtain a new gaze angle feature sequence θ′ t is the gaze angle feature at time t;

[0052] Step 4-6: For the gaze angle feature θ′ t at time t, a Gaussian kernel G t is used to represent the time scale of θ′ t , t∈{1,2,...,T};

[0053] Step 4-6-1: The gaze angle feature sequence Θ′ obtained in step 4-5 is passed through a one-dimensional convolution layer to obtain the standard deviation sequence t of the Gaussian kernel G t of all gaze angle features, and each standard deviation is limited to (0, 1) by a sigmoid operation, σ t represents the standard deviation of the Gaussian kernel G t ;

[0054] ​Step 4-6-2: T is the length of time, define Z as a normalization constant, i∈{1,2,...,T}, t∈{1,2,...,T}, μ t represents the Gaussian kernel G t , p i is the parameter of the Gaussian kernel G t , then the standard deviation sequence learned in step 4-6-1 is used The weight of the Gaussian kernel of the line-of-sight angle feature θ' at time t is represented as: t

[0055]

[0056] Step 4-6-3: The center position of the line-of-sight angle feature θ' at time position t is represented as: t

[0057]

[0058] Step 4-6-4: define r d as the time scale ratio, using the standard deviation sequence learned in step 4-6-1 The width of the line-of-sight angle feature θ' at time position t is represented as: t

[0059]

[0060] Step 4-7: For all Gaussian kernels obtained in step 4-6, use the Gaussian kernel fusion algorithm to fuse two Gaussian kernels that are adjacent and have a large overlap, obtain the Gaussian kernel set after fusion and the time position set of the Gaussian kernel after fusion;

[0061] Step 4-7-1: define the length of the time intersection between the Gaussian kernel at time t1 and the Gaussian kernel at time t2 as The length of the time union is The overlap degree is

[0062] Step 4-7-2: define the original Gaussian kernel set G Define the Gaussian kernel set after the fusion process as G end , define the time position set of the Gaussian kernel after fusion as T';

[0063] Step 4-7-3: input the original Gaussian kernel set G start , initialize G end as an empty set, define q∈{1,2,...,T}, z∈{1,2,...,T}, q,z both represent time positions; ​​​

[0064] Step 4-7-4: Make q point to G start the first Gaussian kernel in, z points to G start the second Gaussian kernel in, i.e. initialize q = 1, z = 2;

[0065] Step 4-7-5: Calculate and the overlap IoU between two Gaussian kernels, σ q denotes the standard deviation of the Gaussian kernel , σ z denotes the standard deviation of the Gaussian kernel , μ q denotes the mathematical expectation of the Gaussian kernel , μ z denotes the mathematical expectation of the Gaussian kernel ;

[0066] Step 4-7-5-1: Calculate and the length H of the time intersection of two Gaussian kernels q,z , the calculation formula is as follows, center q denotes the center position of the line-of-sight angle feature at time position q, center z denotes the center position of the line-of-sight angle feature at time position z, width q denotes the time width of the line-of-sight angle feature at time position q, width z denotes the time width of the line-of-sight angle feature at time position z,:

[0067] H q,z = length ((center q -width q , center q + width q )∩(center z -width z , center z + width z ))

[0068] Step 4-7-5-2: Calculate and the length L of the time union of two Gaussian kernels q,z , the calculation formula is as follows:

[0069] L q,z = length ((center q -width q , center q + width q)∪(center z -width z ,center z +width z ))

[0070] Step 4-7-5-3: Calculate the overlap IoU between two Gaussian kernels and The formula is as follows: q,z

[0071] IoU q,z =H q,z / L q,z

[0072] Step 4-7-6: According to the IoU obtained in step 4-7-5-3 q,z , compare the size of IoU q,z with 0.7;

[0073] Step 4-7-6-1: If IoU q,z ≥ 0.7, according to the fusion formula as follows:

[0074]

[0075]

[0076] fusion and Save the fusion result to time (q+z) / 2 to the set T';

[0077] Step 4-7-6-2: If IoU q,z < 0.7, add Gaussian kernel to the set G end , add time q to the set T', q = z,

[0078] Step 4-7-7: Point z to the next Gaussian kernel in G start , that is, z = z + 1;

[0079] Step 4-7-8: Compare the size of q and T;

[0080] Step 4-7-8-1: When q ≤ T, the traversal is not finished, then repeat steps 4-7-5 to 4-7-8;

[0081] Step 4-7-8-2: When q > T, the traversal is finished, then execute step 4-7-9;

[0082] ​Step 4-7-9: After step 4-7-8, the set of Gaussian kernels G after fusion is obtained end and the set of time positions T' of the Gaussian kernels after fusion.

[0083] Step 4-8: Using the set of fusion Gaussian kernels G obtained in step 4-7, the weighted sum of each feature in the feature sequence Θ end is calculated according to the weight in the fusion Gaussian curve, to obtain the line-of-sight angle fusion feature sequence Θ'' = {θ''i , θ''i t is the line-of-sight angle fusion feature at time t, t ∈ {1, 2,..., T}, i ∈ {1, 2,..., T}, t' ∈ T', W t′ [i] is the weight of the fusion Gaussian curve at time t', and the line-of-sight angle fusion feature calculation formula is as follows:

[0084]

[0085] Step 4-9: According to the fusion feature sequence Θ'' obtained in step 4-8, the line-of-sight state classification result sequence Y = {y classify , y t ∈ {1, 2,..., T} is obtained by using the threshold classification method through the classification function φ t (·), and the classification function is as follows:

[0086]

[0087] where β1 is the lower boundary of the safe line-of-sight angle, and β2 is the upper boundary of the safe line-of-sight angle.

[0088] Step 4-10: According to the method in step 4-6-3, the center position center t of the line-of-sight angle feature θ' t t at time t is obtained, to form the center position value sequence

[0089] Step 4-11: According to the method in step 4-6-4, the width width t of the line-of-sight angle feature θ' t head at time t is obtained, to form the width value sequence

[0090] Step 4-12: The classification result sequence Y obtained in step 4-9, the center position value sequence obtained in step 4-10, and the width value sequence obtained in step 4-11 are traversed to obtain the starting position and the ending position j is the segment number of the gaze state, j∈A, A is the set of segment gaze state segment numbers;

[0091] Step 4-13: according to the starting position of each segment gaze state obtained in step 4-12 and the ending position and the real starting position and the width Calculate the positioning loss, and the loss function is as follows:

[0092]

[0093] Step 4-14: using the loss function in step 4-13, training the gaze state time positioning network to obtain the gaze state time positioning network model parameters.

[0094] Step 5 describes the driving process of the driver, estimates the gaze state time positioning of the driver, and detects the dangerous driving behavior of the driver, which specifically includes the following steps:

[0095] Step 5-1: during the driving process of the driver, the camera continuously captures the video containing the head of the driver;

[0096] Step 5-2: frame the captured video continuously;

[0097] Step 5-3: using the method of steps 4-2 to 4-3, the gaze angle features of the driver in all video frames are obtained, and a gaze angle feature sequence is formed;

[0098] Step 5-4: the gaze angle feature sequence obtained in step 5-3 is input into the gaze state time positioning network model for detection, and the starting position and ending position of each segment gaze state are obtained;

[0099] Step 5-5: according to the starting position and ending position of each segment gaze state obtained in step 5-4, the duration of each segment gaze state is obtained;

[0100] Step 5-6: detecting the duration of each segment gaze state obtained in step 5-5, when the gaze state is in a dangerous gaze state and the duration is greater than the safe duration, it is determined as a dangerous driving behavior, and the system reminds the driver.

[0101] The advantages of the present application are: the present application can handle the case that the head and the eyes are not consistent, and can also robustly handle different gaze direction time change processes, and can effectively realize dangerous driving behavior detection. BRIEF DESCRIPTION OF DRAWINGS

[0102] Figure 1Flowchart of a dangerous driving behavior detection method based on line-of-sight temporal relationship learning;

[0103] Figure 2 Diagram showing the division of line of sight ( Figure 2 (a) is a diagram of the line of sight based on angle; Figure 2 (b) is a schematic diagram of two line-of-sight states;

[0104] Figure 3 This is a schematic diagram of head orientation vector extraction;

[0105] Figure 4 A schematic diagram of binocular gaze direction vector extraction;

[0106] Figure 5 A schematic diagram illustrating the extraction of the combined gaze direction of the head and both eyes;

[0107] Figure 6 Flowchart for Gaussian kernel learning;

[0108] Figure 7 A schematic diagram of time positioning for line-of-sight status;

[0109] Figure 8 A comparison chart showing learning using time relationships versus learning without time relationships. Detailed Implementation

[0110] like Figure 1 As shown, a dangerous driving behavior detection method based on gaze direction temporal relationship learning is presented. During the driver's driving process, a camera continuously captures video including the driver's head, continuously sampling frames, one frame every four frames, for a total of 32 frames. Based on the video frame sequence, the driver's gaze state temporal localization is estimated (gaze state segmentation is as follows). Figure 2 As shown, the detection of dangerous driving behavior by drivers includes the following steps:

[0111] Step 1: Input the safe driving dataset, train the head orientation estimation network, and obtain the head orientation estimation network parameter model;

[0112] Step 1-1: Input the head detection dataset and train a head detection network model based on Yolov5;

[0113] Step 1-2: Input the safe driving training set, use the head detection network model trained in Step 1-1 to detect the head region of the input image, and obtain the head region image after cropping;

[0114] Steps 1-3: Normalize the head region image obtained in Step 1-2 to ensure uniform size and obtain the image center point O. head Using the image center point O headThe head center point is represented, a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis;

[0115] Steps 1-4: The normalized head region image is input into the head orientation estimation network, the head orientation estimation network is composed of a ResNet-34 network and three fully connected layers, and a head orientation vector is obtained The horizontal and vertical coordinates are represented by x and y respectively;

[0116] Step 1-5: Calculate the head orientation vector obtained in step 1-4 and the real vector The loss function between them is The horizontal and vertical coordinates are represented by x and y respectively, and the loss function formula is as follows:

[0117]

[0118] Step 1-6: Use the loss function of step 1-5 to train the head orientation estimation network in step 1-4 to obtain the head orientation estimation network parameter model.

[0119] Step 2, input the safe driving data set, and train the binocular line of sight direction estimation network to obtain the binocular line of sight direction estimation network parameter model;

[0120] Step 2-1: Input the human eye detection data set, and train the left eye detection network model based on Yolov5;

[0121] Step 2-2: Input the safe driving training set, use the method of step 1-2 to obtain the head region portrait, and then use the left eye detection network model trained in step 2-1 to detect the left eye region in the head region image, and obtain the left eye region image after cropping;

[0122] Step 2-3: Normalize the left eye region image obtained in step 2-2 to make the size uniform, and obtain the image center point O left_eye The left eye center point is represented by O left_eye A rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis;

[0123] Step 2-4: Make the normalized left eye region image in step 2-3 pass through the left branch of the binocular line of sight direction estimation network, and the left branch of the binocular line of sight direction estimation network is composed of a ResNet-18 network to generate the left eye line of sight direction vector The horizontal and vertical coordinates are represented by x and y respectively;

[0124] Step 2-5: Input the human eye detection dataset, and train the right eye detection network model based on Yolov5;

[0125] Step 2-6: Input the safe driving training set, obtain the head region portrait using the method of step 1-2, and then use the right eye detection network model trained in step 2-5 to detect the right eye region in the head region image, and obtain the right eye region image after cropping;

[0126] Step 2-7: Normalize the right eye region image obtained in step 2-6 respectively to make the size uniform, and obtain the image center point O right_eye , and the right eye center point is represented by O right_eye . Establish a rectangular coordinate system with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis;

[0127] Step 2-8: Make the normalized right eye region image in step 2-7 pass through the right branch of the binocular gaze direction estimation network, and the right branch of the binocular gaze direction estimation network is composed of a ResNet-18 network to generate the gaze direction vector of the right eye , respectively representing the horizontal coordinate and the vertical coordinate;

[0128] Step 2-9: Pass the left eye gaze direction vector obtained in step 2-4 and the right eye gaze direction vector in step 2-8 through a multilayer perceptron φ eye with one hidden layer to generate a binocular gaze direction vector , respectively representing the horizontal coordinate and the vertical coordinate:

[0129] α bin_eye =φ eye (α left_eye ,α right_eye )

[0130] Step 2-10: Calculate the loss function between the binocular gaze direction vector in step 2-9 and the real vector , and the loss function formula is as follows:

[0131]

[0132] Step 2-11: Use the loss function in step 2-10 to train the binocular gaze direction estimation network composed of step 2-4, step 2-8 and step 2-9 to obtain the binocular gaze direction estimation network parameter model.

[0133] Step 3, input the safe driving dataset, and train the head and binocular joint gaze direction estimation network to obtain the head and binocular joint gaze direction estimation network parameter model;​

[0134] Step 3-1: input the safe driving training set, use the head detection network model trained in step 1-1 to detect the head region of the input image, and crop the head region image;

[0135] Step 3-2: use the head orientation estimation network model trained in step 1 to extract the head orientation vector from the head region image obtained in step 3-1 by using the method of step 1-4

[0136] Step 3-3: use the dual eye gaze direction estimation network model trained in step 2 to extract the dual eye gaze direction vector in the head region image by using the method of step 2-1 to step 2-9 from the head region image obtained in step 3-1

[0137] Step 3-4: obtain the head orientation vector α head obtained in step 3-3 bin_eye Through a multilayer perceptron φ(·) containing a hidden layer, the output result is represented as a normalized head and dual eye joint gaze direction vector respectively represent the horizontal coordinate and the vertical coordinate:

[0138] α union = φ(α head , α bin_eye )

[0139] Step 3-5: calculate the loss function between the head and dual eye joint gaze direction vector obtained in step 3-4 and the real vector The loss function formula is as follows:

[0140]

[0141] Step 3-6: use the loss function in step 3-5 to train the head and dual eye joint gaze direction estimation network to obtain the head and dual eye joint gaze direction estimation network parameter model.

[0142] Step 4, input the safe driving data set, and train the gaze state time positioning network to obtain the gaze state time positioning network parameter model;

[0143] Step 4-1: input the safe driving data set, continuously sample the original video containing the driver's head to obtain a video frame sequence;

[0144] Step 4-2: Using the head and eyes joint gaze direction estimation network model trained in step 3, the head and eyes joint gaze direction of the driver in each video frame is estimated by the method of step 3-1 to step 3-4, and the joint gaze direction vector of the driver in all video frames is obtained

[0145] Step 4-3: The joint gaze direction of the driver in all video frames obtained in step 4-2 is converted into a gaze angle feature, and the conversion formula is as follows:

[0146]

[0147] Step 4-4: The gaze angle features of the driver in all video frames obtained in step 4-3 are combined to form a gaze angle feature sequence T is the length of time, t is a time, t ∈ {1, 2,..., T}, θ t represents the gaze angle feature of the driver in the video frame at time t;

[0148] Step 4-5: The gaze angle feature sequence Θ obtained in step 4-4 is passed through two one-dimensional convolution layers, one maximum pooling layer and one one-dimensional time sequence convolution layer to obtain a new gaze angle feature sequence θ' t is the gaze angle feature at time t;

[0149] Step 4-6: For the gaze angle feature θ' at time t t , a Gaussian kernel G t is used to represent the time scale of θ' t , t ∈ {1, 2,..., T};

[0150] Step 4-6-1: The gaze angle feature sequence Θ' obtained in step 4-5 is passed through a one-dimensional convolution layer to obtain the standard deviation sequence t of the Gaussian kernel G t of all gaze angle features, and each standard deviation is limited to (0, 1) by sigmoid operation, σ t represents the standard deviation of the Gaussian kernel G t ; Step 4-6-2: T is the length of time, define Z as the normalization constant, i ∈ {1, 2,..., T}, t ∈ {1, 2,..., T}, μ t represents the mathematical expectation in the Gaussian kernel G i , p t is the Gaussian kernel G t ​the standard deviation sequence learned in step 4-6-1 the line-of-sight angle feature θ' at time t is represented as t the weight of the Gaussian kernel is represented as

[0152]

[0153] the center position of the line-of-sight angle feature θ' at time t is represented as t

[0154]

[0155] step 4-6-4: define r d as the time scale ratio, and use the standard deviation sequence learned in step 4-6-1 the width of the line-of-sight angle feature θ' at time t is represented as t

[0156]

[0157] step 4-7: for all Gaussian kernels obtained in step 4-6, use a Gaussian kernel fusion algorithm to fuse two adjacent Gaussian kernels with a large overlap degree, obtain the Gaussian kernel set after fusion and the time position set of the Gaussian kernel after fusion;

[0158] step 4-7-1: define the length of the time intersection between the Gaussian kernel at time t1 and the Gaussian kernel at time t2 as the length of the time union is the overlap degree is

[0159] step 4-7-2: define the original Gaussian kernel set G define the Gaussian kernel set after fusion as G end , and define the time position set of the Gaussian kernel after fusion as T';

[0160] step 4-7-3: input the original Gaussian kernel set G start , initialize G end as an empty set, define q ∈ {1, 2,..., T} and z ∈ {1, 2,..., T}, both q and z represent time positions;

[0161] step 4-7-4: make q point to the first Gaussian kernel in G start and z point to the second Gaussian kernel in G start , that is, initialize q = 1 and z = 2;

[0162] step 4-7-5: calculate and​​ the overlap IoU between two Gaussians, σ q denotes a Gaussian kernel the standard deviation of the Gaussian, σ z denotes a Gaussian kernel the standard deviation of the Gaussian, μ q denotes a Gaussian kernel the mathematical expectation in the Gaussian, μ z denotes a Gaussian kernel the mathematical expectation in the Gaussian;

[0163] Step 4-7-5-1: Calculate and the length H of the time intersection of two Gaussians q,z , which is calculated as follows, q denotes the center position of the line-of-sight angle feature at time position q, z denotes the center position of the line-of-sight angle feature at time position z, q denotes the time width of the line-of-sight angle feature at time position q, z denotes the time width of the line-of-sight angle feature at time position z,

[0164] H q,z = length((center q -width q , center q + width q )∩(center z -width z , center z + width z ))

[0165] Step 4-7-5-2: Calculate and the length L of the time union of two Gaussians q,z , which is calculated as follows:

[0166] L q,z = length((center q -width q , center q + width q )∪(center z -width z , center z + width z ))

[0167] Step 4-7-5-3: Calculate and Overlap between two Gauss kernels, IoU q,z , the formula is as follows:

[0168] IoU q,z = H q,z / L q,z

[0169] Step 4-7-6: According to the IoU q,z obtained in step 4-7-5-3, compare the size of IoU q,z with 0.7;

[0170] Step 4-7-6-1: If IoU q,z ≥ 0.7, according to the fusion formula as follows:

[0171]

[0172]

[0173] fusion and save the fusion result to time (q+z) / 2 into the set T';

[0174] Step 4-7-6-2: If IoU q,z < 0.7, add the Gauss kernel into the set G end , add the time q into the set T', q = z,

[0175] Step 4-7-7: Point z to the next Gauss kernel in G start , that is, z = z + 1;

[0176] Step 4-7-8: Compare the size of q and T;

[0177] Step 4-7-8-1: When q ≤ T, the traversal is not finished, then repeat steps 4-7-5 to 4-7-8;

[0178] Step 4-7-8-2: When q > T, the traversal is finished, then execute step 4-7-9;

[0179] Step 4-7-9: After executing step 4-7-8, obtain the set G end of Gauss kernels after the fusion process is finished and the set T' of time positions of the Gauss kernels after the fusion is finished;

[0180] Step 4-8: Use each Gauss kernel in the set G end of the fused Gauss kernels obtained in step 4-7 to calculate the feature sequence according to the weight in the fused Gauss curve The weighted sum of each feature in the sequence yields the gaze angle fusion feature sequence Θ″={θ″. t}, θ″ t Let W be the gaze angle fusion feature at time t, where t∈{1,2,...,T}, i∈{1,2,...,T}, t′∈T′. t′ [i] represents the weight of the fused Gaussian curve at time t′. The formula for calculating the gaze angle fusion feature is as follows:

[0181]

[0182] Steps 4-9: Based on the fused feature sequence Θ″ obtained in Step 4-8, use a threshold classification method, and pass it through the classification function φ. classify (·) Obtain the gaze state classification result sequence of each fusion feature. For t∈{1,2,...,T}, the classification function is as follows:

[0183]

[0184] Where β1 is the lower boundary of the safe line of sight angle, and β2 is the upper boundary of the safe line of sight angle;

[0185] Step 4-10: Based on the method in Step 4-6-3, obtain the line-of-sight angle feature θ′ at time position t. t center t , forming a sequence of central position values

[0186] Step 4-11: Based on the method in Step 4-6-4, obtain the line-of-sight angle feature θ′ at time position t. t width t , forming a width value sequence

[0187] Step 4-12: Traverse the classification result sequence Y obtained in Step 4-9 and the center position value sequence obtained in Step 4-10. and the width value sequence obtained in step 4-11 Obtain the starting position of each line of sight state. and end position j is the segment number of the line of sight state, j∈A, where A is the set of segment numbers of each line of sight state;

[0188] Step 4-13: Based on the starting position of each line of sight obtained in Step 4-12 and end position With the actual starting position and width Compute the position loss, the loss function is as follows:

[0189]

[0190] Step 4-14: using the loss function in step 4-13, train the line-of-sight state time positioning network to obtain the line-of-sight state time positioning network model parameters.

[0191] Step 5, in the driving process of the driver, estimate the line-of-sight state time positioning of the driver, and detect the dangerous driving behavior of the driver;

[0192] Step 5-1: during the driving process of the driver, the camera continuously captures video containing the driver's head;

[0193] Step 5-2: continuously frame the captured video;

[0194] Step 5-3: using the method of steps 4-2 to 4-3, obtain the line-of-sight angle features of the driver in all video frames in step 5-2, and form a line-of-sight angle feature sequence;

[0195] Step 5-4: input the line-of-sight angle feature sequence obtained in step 5-3 into the line-of-sight state time positioning network model for detection to obtain the starting position and ending position of each line-of-sight state;

[0196] Step 5-5: according to the starting position and ending position of each line-of-sight state obtained in step 5-4, obtain the duration of each line-of-sight state;

[0197] Step 5-6: detect the duration of each line-of-sight state obtained in step 5-5, when the line-of-sight state is in a dangerous line-of-sight state and the duration is greater than the safe duration, it is determined as dangerous driving behavior, and the system reminds the driver.

Claims

1. A dangerous driving behavior detection method based on line-of-sight direction time relationship learning, characterized in that: During the driving process of the driver, the camera continuously captures a video containing the head of the driver, and frames are continuously extracted from the video; According to the sequence of video frames, the gaze state time positioning of the driver is estimated, and the dangerous driving behavior of the driver is detected, specifically including the following steps: Step 1, input the safe driving data set, perform head orientation estimation network training, and obtain a head orientation estimation network parameter model; Step 2, input the safe driving data set, perform binocular gaze direction estimation network training, and obtain a binocular gaze direction estimation network parameter model; Step 3, input the safe driving data set, perform head and binocular joint gaze direction estimation network training, and obtain a head and binocular joint gaze direction estimation network parameter model; Step 4, input the safe driving data set, perform gaze state time positioning network training, and obtain a gaze state time positioning network parameter model; Step 5, during the driving process of the driver, the gaze state time positioning of the driver is estimated, and the dangerous driving behavior of the driver is detected; Step 1, input the safe driving data set, perform head orientation estimation network training, and obtain a head orientation estimation network parameter model, specifically including the following steps: Step 1-1, input the head detection data set, and train a head detection network model based on Yolov5; Step 1-2, input the safe driving training set, use the head detection network model trained in step 1-1 to detect the head region of the input image, and obtain the head region image after cropping; Step 1-3: normalizing the head region image obtained in step 1-2 to unify the size, and obtaining an image center point , the image center point is represented as a head center point, and a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis; Step 1-4: The normalized head region image is input into the head orientation estimation network, which is composed of a ResNet-34 network and three fully connected layers, to obtain a head orientation vector , , respectively represent the horizontal and vertical coordinates; Steps 1-5: Calculate the head orientation vector obtained in Steps 1-4 With real vectors The loss function between , The x and y axes represent the coordinates respectively, and the loss function formula is as follows: , Step 1-6, use the loss function of step 1-5 to train the head orientation estimation network in step 1-4, and obtain a head orientation estimation network parameter model; Step 3, input the safe driving data set, perform head and binocular joint gaze direction estimation network training, and obtain a head and binocular joint gaze direction estimation network parameter model, specifically including the following steps: Step 3-1, input the safe driving training set, use the head detection network model trained in step 1-1 to detect the head region of the input image, and obtain the head region image after cropping; Step 3-2: For the head region image obtained in step 3-1, use the head orientation estimation network model trained in step 1 to extract the head orientation vector using the method of step 1-4 ; Step 3-3: For the head region image obtained in Step 3-1, using the binocular gaze direction estimation network model trained in Step 2, extract the binocular gaze direction vector in the head region image using the method of Step 2 ; Step 3-4: Head orientation vector is obtained from the head part in step 3-2 Eye gaze direction vector is obtained from the eye part in step 3-3 By a multi-layer perceptron with one hidden layer The output result is represented as normalized head and eye joint gaze direction vector , , X and Y represent the horizontal and vertical coordinates respectively: , Step 3-5: Calculate the head with the binocular joint gaze direction vector from the head obtained in step 3-4 between the real vector and the predicted vector , the loss function is as follows: , Step 3-6, use the loss function of step 3-5 to train the head and binocular joint gaze direction estimation network, and obtain a head and binocular joint gaze direction estimation network parameter model. 2.The dangerous driving behavior detection method based on line-of-sight direction-time relationship learning according to claim 1, wherein, Step 2, input the safe driving data set, perform binocular gaze direction estimation network training, and obtain a binocular gaze direction estimation network parameter model, specifically including the following steps: Step 2-1, input the eye detection data set, and train a left eye detection network model based on Yolov5; Step 2-2, input the safe driving training set, use the method of step 1-2 to obtain the head region image, then use the left eye detection network model trained in step 2-1 to detect the left eye region in the head region image, and obtain the left eye region image after cropping; Step 2-3: normalizing the left eye region image obtained in step 2-2 respectively to make the size uniform, and obtaining the image center point , the image center point is represented as the left eye center point, and a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis. Step 2-4: the normalized left eye region image in step 2-3 is input into the left branch of the binocular gaze direction estimation network, which is composed of a ResNet-18 network, to generate the gaze direction vector of the left eye , , respectively represent the horizontal coordinate and the vertical coordinate; Step 2-5, input the eye detection data set, and train a right eye detection network model based on Yolov5; Step 2-6: input the safe driving training set, use the method of step 1-2 to obtain the head portrait in the head region, and then use the right eye detection network model trained in step 2-5 to detect the right eye region in the head region image, and obtain the right eye region image after cropping; Step 2-7: Normalizing the right eye region images obtained in Step 2-6 respectively to unify their sizes, and obtaining image center points , and the image center points represent the right eye center points, and a rectangular coordinate system is established with the head center point as the coordinate origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis. Step 2-8: the normalized right eye region image in step 2-7 passes through the right branch of the binocular gaze direction estimation network, the right branch of the binocular gaze direction estimation network is composed of a ResNet-18 network, to generate the gaze direction vector of the right eye , , respectively represent the horizontal coordinate and the vertical coordinate; Step 2-9: The left eye gaze direction vector obtained in step 2-4 and the right eye gaze direction vector of step 2-8 are passed through a multi-layer perceptron with one hidden layer , resulting in a binocular gaze direction vector , , denote the horizontal and vertical coordinates, respectively: , Step 2-10: Calculate the binocular gaze direction vector in step 2-9 between the real vector and the loss function formula is as follows: , Step 2-11: using the loss function in step 2-10, training the binocular gaze direction estimation network composed of step 2-4, step 2-8 and step 2-9, obtaining the binocular gaze direction estimation network parameter model. 3.The dangerous driving behavior detection method based on line-of-sight direction-time relationship learning according to claim 1, characterized in that: The input safe driving data set in step 4 is used to train the gaze state time positioning network, and the gaze state time positioning network parameter model is obtained, which includes the following steps: Step 4-1: input the safe driving data set, and continuously sample the original video containing the driver's head to obtain a video frame sequence; Step 4-2: Using the head and eyes joint gaze direction estimation network model trained in step 3, the head and eyes joint gaze direction of the driver in each video frame is estimated by using the method of step 3-1 to step 3-4, and the joint gaze direction vector of the driver in all video frames is obtained ; Step 4-3: Convert the joint gaze direction of the driver in all video frames obtained in Step 4-2 into gaze angle features, using the following conversion formula: Step 4-3: Convert the joint gaze direction of the driver in all video frames obtained in Step 4-2 into gaze angle features, using the following conversion formula: Step 4-3: Convert the joint gaze direction of the driver in all video frames obtained in Step 4- , Step 4-4: the line-of-sight angle feature of the driver in each video frame obtained in step 4-3 is composed into a line-of-sight angle feature sequence , is a time length, is a certain moment, , represents the line-of-sight angle feature of the driver in the video frame at the moment. Step 4-5: Obtain a new line-of-sight angle feature sequence from the line-of-sight angle feature sequence obtained in step 4-4 through two one-dimensional convolution layers, one maximum pooling layer, and one one-dimensional time sequence convolution layer , , is the line-of-sight angle feature at time t. Step 4-6: Line of sight angle feature for time t , using a Gaussian kernel to represent the time scale, ; Step 4-6-1: Obtain the line-of-sight angle feature sequence from step 4-5 , get the standard deviation sequence of all line-of-sight angle features through a one-dimensional convolution layer , and limit each standard deviation to , , where represents the standard deviation of the Gaussian kernel​ Step 4-6-2: is the time length, and is a normalization constant, , , , , denotes the Gaussian kernel , and is the parameter of the Gaussian kernel , then the weight of the Gaussian kernel of the line-of-sight angle feature at time t is expressed as: ​ , Step 4 - 6-3: The line of sight angle feature at time position t The center position of the line of sight angle feature is represented as: , Step 4-6-4: Definition For the time scale ratio, the standard deviation sequence learned in Step 4-6-1 is used The width of the line-of-sight angle feature at time position t is expressed as: ​ , Step 4-7: for all Gaussian kernels obtained in step 4-6, use Gaussian kernel fusion algorithm to fuse two Gaussian kernels with adjacent and large overlap, obtain the Gaussian kernel set after fusion and the time position set of the Gaussian kernel after fusion; Step 4-7-1: Definition The length of the time intersection between the Gaussian kernels of time instants and is The length of the time union is The degree of overlap is ; Step 4-7-2: Define the original set of Gaussian kernels , define the set of Gaussian kernels after the fusion process as , define the set of time positions of the Gaussian kernels after the fusion as ; Step 4-7-3: input original set of Gaussian kernels , initialize to empty set, define , , all denote temporal positions; Step 4 - 7 - 4: Make point to the first Gaussian kernel in , point to the second Gaussian kernel in , i.e. initialize , ; Step 4 - 7-5: Compute and the degree of overlap between , denotes the standard deviation of the Gaussian kernel , denotes the standard deviation of the Gaussian kernel , denotes the mathematical expectation in the Gaussian kernel , denotes the mathematical expectation in the Gaussian kernel ; Step 4-7-5-1: Calculation and Length of the time intersection of two Gaussian kernels The calculation formula is as follows: Indicates the time position as The center position of the line-of-sight angle characteristics, Indicates the time position as The center position of the line-of-sight angle characteristics, Indicates the time position as The time span of the viewing angle characteristics Indicates the time position as The time width of the line-of-sight angle characteristics: Step 4-7-5-2: Compute and Length of the time union of two Gaussian kernels The formula is as follows: , Step 4-7-5-3: Calculation and Overlap between two Gaussian kernels , the calculation formula is as follows: , Step 4-7-6: According to the compound obtained in Step 4-7-5-3 , the size of the comparison with 0.7; Step 4 - 7 - 6 - 1 : If , according to the fusion formula as follows: , , fusion with save the fusion result to , time join set ; Step 4 - 7 - 6 - 2: If , the Gaussian kernel is added to the set , the time is added to the set , , ; Step 4-7-7: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] point to The next Gaussian kernel in the process, namely ; Step 4 - 7-8: Comparison with the size of Step 4-7-8-1 : When the traversal is not yet finished, steps 4-7-5 to 4-7-8 are repeated. Step 4-7-8-2: When the traversal is finished, then step 4-7-9 is performed; Step 4-7-9: After step 4-7-8 is executed, a set of Gaussian kernels after the fusion process is completed is obtained and a set of time positions of the Gaussian kernels after the fusion is completed ; Step 4-8: Obtain the fusion Gaussian kernel set using the fusion Gaussian kernels in step 4-7 Step 4-9: Calculate the weighted sum of each feature in the feature sequence Step 4-10: Obtain the line-of-sight angle fusion feature sequence , Step 4-11: Obtain the line-of-sight angle fusion feature at time t , , , Step 4-12: Obtain the weight of the fusion Gaussian curve at time t The line-of-sight angle fusion feature calculation formula is as follows: , Step 4-9: according to the fusion feature sequence obtained in step 4-8 , using the threshold classification method, through a classification function to obtain the line-of-sight state classification result sequence of each fusion feature , The classification function is as follows: , wherein is a lower boundary of the safety view angle, is an upper boundary of the safety view angle; Step 4-10: Obtain the line-of-sight angle feature at time position t according to the method in Step 4-6-3 of the center position , the sequence of center position values ; Step 4-11: Based on the method in Step 4-6-4, obtain the line-of-sight angle feature at time position t. width , forming a width value sequence ; Step 4-12: Obtain the start position and the end position of each segment of the line-of-sight state from the sequence of classification results obtained in Step 4-9, the sequence of center position values obtained in Step 4-10, and the sequence of width values obtained in Step 4-11. is a set of segment numbers of the segments of the line-of-sight state.​​​​​​​ Step 4-13: Based on the starting position of each line of sight obtained in Step 4-12 and end position With the actual starting position and width Calculate the localization loss using the following loss function: , Step 4-14: using the loss function in step 4-13, training the gaze state time positioning network, obtaining the gaze state time positioning network model parameters.

4. The dangerous driving behavior detection method based on line-of-sight direction-time relationship learning according to claim 3, characterized in that: The gaze state time positioning of the driver is estimated in the driving process of the driver, and the dangerous driving behavior of the driver is detected, which includes the following steps: Step 5-1: during the driving process of the driver, the camera continuously shoots the video containing the driver's head; Step 5-2: continuously sampling the video; Step 5-3: using the method of step 4-2 to step 4-3, obtaining the gaze angle features of the driver in all video frames, and forming a gaze angle feature sequence; Step 5-4: input the gaze angle feature sequence obtained in step 5-3 into the gaze state time positioning network model for detection, and obtain the starting position and ending position of each gaze state; Step 5-5: according to the starting position and ending position of each gaze state obtained in step 5-4, the duration of each gaze state is obtained; Step 5-6: detecting the duration of each gaze state obtained in step 5-5, when the gaze state is in the dangerous gaze state and the duration is greater than the safe duration, it is determined as dangerous driving behavior, and the system reminds the driver.

Citation Information

Patent Citations

  • Vehicle-mounted intelligent driving early warning system and vehicle

    CN113942450A

  • Driving behavior warning method and device based on video analysis, equipment and medium

    CN114005093A