A driving behavior analysis method, device, electronic device and storage medium

By analyzing the driver's face image, integrating the position information of the key face points and the driver's posture information, the problem of low accuracy in driving behavior analysis in the prior art is solved, and more efficient identification and prevention of irregular driving behaviors is achieved.

CN114255504BActive Publication Date: 2025-06-03NANJING LINGXING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111611041.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-06-03
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The prior art has low accuracy in driving behavior analysis, making it difficult to effectively identify and prevent irregular driving behaviors.

Method used

By analyzing the driver's face image, the position information of the key points of the face is obtained, the first face posture information and the second face posture information are integrated, and then the driver is analyzed whether the driver has irregular driving behavior.

Benefits of technology

It improves the accuracy of driving behavior analysis, can more effectively identify and prevent irregular driving behaviors, and reduces the risk of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255504B_ABST
    Figure CN114255504B_ABST
Patent Text Reader

Abstract

The present application discloses a driving behavior analysis method, device, electronic device and storage medium, belonging to the technical field of vehicle operation. The method includes: analyzing the acquired face image of the driver to obtain face information including at least the position information of face key points; determining the first face pose information of the driver based on the position information of the face key points; determining the second face pose information of the driver based on the eye region data in the face image; fusing the first face pose information and the second face pose information; and analyzing whether the driver has an irregular driving behavior based on the fused target face pose information. In this way, based on the face key points and eye region data of the driver, relatively accurate target face pose information can be obtained. Based on the relatively accurate target face pose information, it is beneficial to accurately analyze whether the driver has an irregular driving behavior. Therefore, the accuracy of driving behavior analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of vehicle operation, and in particular, to a driving behavior analysis method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of Internet technology, many vehicle operation enterprises have transferred their offline operation models to the online, resulting in the emergence of online car-hailing services. Online car-hailing services are available on demand, bringing great convenience to passengers' travel.

[0003] During the process of driving a vehicle, if a driver exhibits non-standard driving behaviors such as fatigue driving or turning back to chat with passengers, it is very easy to cause traffic accidents such as rear-end collisions, bringing harm to the driver and passengers, and further increasing the management difficulty of the vehicle operation platform. Therefore, it is necessary to analyze the driving behavior of the driver to reduce the occurrence frequency of non-standard driving behaviors. And how to accurately analyze the driving behavior of the driver is an urgent problem to be solved. Summary of the Invention

[0004] Embodiments of this application provide a driving behavior analysis method, apparatus, electronic device, and storage medium to solve the problem of relatively low accuracy in analyzing driving behavior in related technologies.

[0005] In a first aspect, embodiments of this application provide a driving behavior analysis method, including:

[0006] Analyze the acquired face image of the driver to obtain face information, where the face information at least includes the position information of face key points;

[0007] Based on the position information of the face key points, determine the first face pose information of the driver;

[0008] Based on the eye region data in the face image, determine the second face pose information of the driver;

[0009] Fuse the first face pose information and the second face pose information to obtain target face pose information;

[0010] Based on the target face pose information, analyze whether the driver exhibits non-standard driving behaviors. In some embodiments, the face information further includes an indication information on whether a mask is worn, and

[0011] When the indication information indicates that the driver is not wearing a mask, based on the position information of the face key points, determine the first face pose information of the driver.

[0012] In some embodiments, the face information further includes position correction information of face key points, and further includes:

[0013] When the indication information indicates that the driver wears a mask, based on the position correction information, perform correction processing on the position information of the first key point among the face key points, where the first key point is the face key point blocked by the mask;

[0014] Based on the position information of the face key points, determine the first face pose information of the driver, including:

[0015] Based on the corrected position information of the first key point and the position information of the second key point, determine the first face pose information of the driver, where the second key point refers to the face key points other than the first key point.

[0016] In some embodiments, fuse the first face pose information and the second face pose information according to the following formula to obtain the target face pose information:

[0017]

[0018] pitch = pitch1 + α * pitch2

[0019] roll = roll1;

[0020] where yaw1 is the yaw angle in the first face pose information, pitch1 is the pitch angle in the first face pose information, roll1 is the roll angle in the first face pose information; yaw2 is the yaw angle in the second face pose information, pitch2 is the pitch angle in the second face pose information, yaw is the yaw angle in the target face pose information, pitch is the pitch angle in the target face pose information, roll is the roll angle in the target face pose information, when pitch1 is less than the set value, α = 1; when pitch1 is not less than the set value, α = 0.

[0021] In some embodiments, analyze the face image of the driver obtained to obtain face information, including:

[0022] Analyze the face image through a key point analysis model to obtain the face information;

[0023] where the key point analysis model is trained according to the following steps:

[0024] Obtain face image samples, where the face image samples include first face image samples without wearing masks and second face image samples obtained by adding masks to some of the first face image samples;

[0025] Determine the position correction annotation information corresponding to the face image sample, where the position correction annotation information corresponding to the first face image sample is a preset value, and the position correction annotation information corresponding to the second face image sample is determined based on the position transformation relationship between the face key points annotated in the second face image sample and the face key points annotated in the first face image sample;

[0026] Use the face image sample as the input, and the position correction annotation information corresponding to the face image sample, the position annotation information of the face key points, and the annotation information of whether a mask is worn as the output to train a preset model to obtain the key point analysis model.

[0027] In a second aspect, an embodiment of the present application provides a driving behavior analysis device, including:

[0028] An image analysis module, configured to analyze the face image of the driver obtained to obtain face information, where the face information includes at least the position information of the face key points;

[0029] A first determination module, configured to determine the first face pose information of the driver based on the position information of the face key points;

[0030] A second determination module, configured to determine the second face pose information of the driver based on the eye region data in the face image;

[0031] An information fusion module, configured to fuse the first face pose information and the second face pose information to obtain target face pose information;

[0032] A behavior analysis module, configured to analyze whether the driver has an irregular driving behavior based on the target face pose information.

[0033] In some embodiments, the face information further includes an indication information of whether a mask is worn.

[0034] The first determination module is specifically configured to, when the indication information indicates that the driver is not wearing a mask, determine the first face pose information of the driver based on the position information of the face key points.

[0035] In some embodiments, the face information further includes position correction information of the face key points, and further includes:

[0036] A correction module, configured to, when the indication information indicates that the driver is wearing a mask, perform a correction process on the position information of the first key point in the face key points based on the position correction information, where the first key point is the face key point blocked by the mask;

[0037] The second determination module is specifically configured to determine the first face pose information of the driver based on the position information of the first key point after correction and the position information of the second key point, where the second key point refers to the face key points other than the first key point.

[0038] In some embodiments, the information fusion module is specifically configured to fuse the first face pose information and the second face pose information according to the following formula to obtain the target face pose information:

[0039]

[0040] pitch = pitch1 + α * pitch2

[0041] roll = roll1;

[0042] where yaw1 is the yaw angle in the first face pose information, pitch1 is the pitch angle in the first face pose information, roll1 is the roll angle in the first face pose information; yaw2 is the yaw angle in the second face pose information, pitch2 is the pitch angle in the second face pose information, yaw is the yaw angle in the target face pose information, pitch is the pitch angle in the target face pose information, roll is the roll angle in the target face pose information, and when pitch1 is less than the set value, α = 1; when pitch1 is not less than the set value, α = 0.

[0043] In some embodiments, the image analysis module is specifically configured to:

[0044] Analyze the face image through a key point analysis model to obtain the face information;

[0045] where the key point analysis model is trained according to the following steps:

[0046] Obtain face image samples, where the face image samples include first face image samples without wearing masks and second face image samples obtained by adding masks to some of the first face image samples;

[0047] Determine the position correction annotation information corresponding to the face image samples, where the position correction annotation information corresponding to the first face image samples is a preset value, and the position correction annotation information corresponding to the second face image samples is determined based on the position transformation relationship between the face key points annotated in the second face image samples and the face key points annotated in the first face image samples;

[0048] Using the face image sample as the input, and the position correction annotation information corresponding to the face image sample, the position annotation information of the face key points, and the annotation information on whether a mask is worn as the output, a preset model is trained to obtain the key point analysis model.

[0049] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, where:

[0050] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned driving behavior analysis method.

[0051] In a fourth aspect, an embodiment of the present application provides a storage medium, and when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the above-mentioned driving behavior analysis method.

[0052] In the embodiment of the present application, the obtained face image of the driver is analyzed to obtain face information, and the face information at least includes the position information of the face key points. Based on the position information of the face key points, the first face pose information of the driver is determined. Based on the eye region data in the face image, the second face pose information of the driver is determined. The first face pose information and the second face pose information are fused to obtain the target face pose information. Furthermore, based on the target face pose information, it is analyzed whether the driver exhibits irregular driving behavior. In this way, based on the face key points and eye region data of the driver, the face pose information of the driver is determined respectively, and then the determined face pose information is fused to obtain relatively accurate target face pose information. Based on the relatively accurate target face pose information, it is beneficial to accurately analyze whether the driver exhibits irregular driving behavior. Therefore, the accuracy of driving behavior analysis can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0054] Figure 1 FIG. is an application scenario diagram of a driving behavior analysis method provided by an embodiment of the present application;

[0055] Figure 2 FIG. is a flowchart of a driving behavior analysis method provided by an embodiment of the present application;

[0056] Figure 3 FIG. is a flowchart of another driving behavior analysis method provided by an embodiment of the present application;

[0057] Figure 4 It is a flowchart of another driving behavior analysis method provided by an embodiment of the present application;

[0058] Figure 5 It is a flowchart of a method for training a key point analysis model provided by an embodiment of the present application;

[0059] Figure 6 It is a schematic diagram of the process of distraction detection provided by an embodiment of the present application;

[0060] Figure 7 It is another schematic diagram of the process of distraction detection provided by an embodiment of the present application;

[0061] Figure 8 It is a schematic structural diagram of a driving behavior analysis device provided by an embodiment of the present application;

[0062] Figure 9 It is a schematic hardware structure diagram of an electronic device for implementing the driving behavior analysis method provided by an embodiment of the present application. Specific embodiments

[0063] To solve the problem of relatively low accuracy in analyzing driving behavior in related technologies, an embodiment of the present application provides a driving behavior analysis method, device, electronic device, and storage medium.

[0064] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0065] The driving behavior analysis method provided by the embodiment of the present application is applicable to all scenarios that require driving behavior analysis. Subsequently, the solution of the embodiment of the present application will be introduced taking online car-hailing as an example.

[0066] Figure 1 It is an application scenario diagram of a driving behavior analysis method provided by an embodiment of the present application, including a terminal 11, a server 12, and a terminal 13, where:

[0067] The terminal 11, such as a mobile phone, an iPad, etc., is installed with a client for booking a vehicle. Passengers can control the terminal to send a vehicle booking request to the server through the client. The vehicle booking request may include the departure place and destination specified by the passenger.

[0068] After receiving the vehicle reservation request sent by the terminal 11, the server 12 can select a vehicle from the vehicles that can carry passengers according to the departure place and destination in the vehicle reservation request, such as selecting a vehicle that is closest to the departure place. Thereafter, the server 12 sends the characteristic information of the selected vehicle, such as the license plate number, vehicle color, etc., to the terminal 11, and can send a dispatch instruction to the terminal 13 corresponding to the selected vehicle, the dispatch instruction including the departure place, destination, and passenger contact information in the vehicle reservation request.

[0069] The terminal 13, such as a mobile phone, iPad, etc., is installed with a client for accepting orders. After receiving the dispatch instruction sent by the server through the client, the driver can drive the vehicle to the corresponding departure place to pick up the passenger and send the passenger to the destination.

[0070] After introducing the application scenario of the embodiment of the present application, the driving behavior analysis method of the embodiment of the present application is introduced below with reference to a specific flowchart.

[0071] Figure 2 A flowchart of a driving behavior analysis method provided in an embodiment of the present application, the execution subject of the method may be Figure 1 The server in the embodiment includes the following steps:

[0072] 201: Analyze the acquired facial image of the driver to obtain facial information, where the facial information at least includes position information of key facial points.

[0073] Generally, online ride-hailing vehicles are equipped with a driver monitoring system (DMS), which is mainly used to record images of the driving position in the vehicle. The images are from image acquisition devices installed on the left and / or right of the steering wheel. The field of view covers the driver's face. Therefore, the driver's facial image can be obtained through the DMS.

[0074] In a specific implementation, after obtaining the driver's facial image, a face frame may be first determined in the face image, and then facial key points may be detected in the face frame to obtain position information of facial key points in the face image.

[0075] 202: Determine first facial posture information of the driver based on the position information of facial key points.

[0076] The first face posture information includes a yaw angle yaw1, a pitch angle pitch1, and a roll angle roll1.

[0077] For example, the position information of facial key points is input into the first facial posture analysis model for posture analysis to obtain the first facial posture information of the driver, wherein the first facial posture analysis model is obtained by learning the relationship between the position information of facial key points in the facial image sample and the facial posture information of the driver.

[0078] 203: Determine the second face pose information of the driver based on the eye region data in the face image.

[0079] Among them, the second face pose information includes yaw angle yaw2, pitch angle pitch2, and roll angle roll2.

[0080] For example, input the eye region data in the face image into the second face pose analysis model for pose analysis, so as to obtain the second face pose information of the driver. Among them, the second face pose analysis model is learned from the relationship between the eye region data in the face image samples and the face pose information of the driver.

[0081] Considering that the eye region in the face image is relatively small, in order to focus on the eye features, the smallest rectangular frames where the two eyes are located in the face image can be determined first, and then the image regions corresponding to the two rectangular frames are spliced together as the eye region data. In this way, both the redundant image regions between the two eyes can be removed, improving the pose analysis speed of the second face pose analysis model, and the eye contact information can be highlighted, improving the pose analysis accuracy of the second face pose analysis model.

[0082] 204: Fuse the first face pose information and the second face pose information to obtain the target face pose information.

[0083] For example, the first face pose information and the second face pose information can be fused according to the following formula to obtain the target face pose information:

[0084]

[0085] pitch = pitch1 + α * pitch2

[0086] roll = roll1;

[0087] Among them, yaw1 is the yaw angle in the first face pose information, pitch1 is the pitch angle in the first face pose information, roll1 is the roll angle in the first face pose information; yaw2 is the yaw angle in the second face pose information, pitch2 is the pitch angle in the second face pose information, yaw is the yaw angle in the target face pose information, pitch is the pitch angle in the target face pose information, roll is the roll angle in the target face pose information. When pitch1 is less than the set value, α = 1; when pitch1 is not less than the set value, α = 0.

[0088] Generally, yaw1, pitch1, and roll1 are all between -90° and 90°, so the set value can be 5°.

[0089] 205: Analyze whether the driver exhibits irregular driving behavior based on the target face pose information.

[0090] Among them, irregular behaviors include distracted behaviors such as looking out of the window for a long time, looking down at the mobile phone navigation, and turning back to chat with passengers.

[0091] In specific implementation, when the target face pose information of consecutive multiple face images meets the preset conditions, it can be determined that the driver exhibits irregular driving behavior. Among them, the preset conditions are such that for any one of the pose angles of yaw angle, pitch angle, and roll angle, the proportion of face images in which the pose angle exceeds the given angle range, such as (-40°, 40°), in consecutive multiple face images exceeds the preset proportion, such as 90%.

[0092] In addition, it should be noted that there is no strict sequence relationship between the above 202 and 203.

[0093] Figure 3 The flowchart of another driving behavior analysis method provided by the embodiment of the present application includes the following steps:

[0094] 301: Analyze the face image of the driver obtained to obtain face information, where the face information includes at least the position information of face key points and the indication information of whether a mask is worn.

[0095] In specific implementation, the face frame can be determined in the face image first, and then, face key point detection is performed in the face frame to obtain the position information of face key points in the face image, and mask recognition can be performed on the face frame. Based on the recognition result, the indication information of whether a mask is worn is determined.

[0096] 302: When the indication information indicates that the driver is not wearing a mask, determine the first face pose information of the driver based on the position information of the face key points.

[0097] Considering that when determining the first face pose information of the driver based on the position information of the face key points, the accuracy of the first face pose information depends more on the key points on the chin and cheeks of the face, and when wearing a mask, due to the mask occlusion, the positions of the key points at these positions are not very accurate. Therefore, when it is determined that the driver is not wearing a mask, the first face pose information of the driver can be determined based on the position information of the face key points to ensure the accuracy of the obtained first face pose information, thereby ensuring the accuracy of the subsequent analysis of driving behavior.

[0098] 303: Determine the second face pose information of the driver based on the eye region data in the face image.

[0099] For the implementation of this step, reference can be made to the implementation of 203, which will not be elaborated here.

[0100] 304: Fuse the first face pose information and the second face pose information to obtain the target face pose information.

[0101] For the implementation of this step, refer to the implementation of 204, which will not be elaborated here.

[0102] 305: Based on the target face pose information, analyze whether the driver exhibits irregular driving behavior.

[0103] For the implementation of this step, refer to the implementation of 205, which will not be elaborated here.

[0104] Figure 4 The flowchart of another driving behavior analysis method provided by the embodiments of the present application includes the following steps:

[0105] 401: Analyze the acquired face image of the driver to obtain face information, where the face information at least includes the position information of face key points, the indication information of whether a mask is worn, and the position correction information of face key points.

[0106] In some embodiments, the position information of face key points, the indication information of whether a mask is worn, and the position correction information of face key points can be determined respectively. For example, first determine the face frame in the face image, then perform face key point detection in the face frame to obtain the position information of face key points in the face image, and perform mask recognition on the face frame, and determine the indication information of whether a mask is worn based on the recognition result. In addition, the face image can also be input into a position correction model to obtain the position correction information of face key points, where the position correction model is learned from the position conversion relationship between face key points in the face image samples of the driver wearing a mask and the face image samples without wearing a mask.

[0107] In some embodiments, the face image can be analyzed by a key point analysis model to obtain face information, that is, simultaneously obtain the position information of face key points, the indication information of whether a mask is worn, and the position correction information of face key points.

[0108] Specifically, during implementation, it can be based on Figure 5 The shown process trains the key point analysis model, and this process includes the following steps:

[0109] 501a: Obtain face image samples, where the face image samples include the first face image samples without wearing a mask and the second face image samples obtained by adding a mask to some of the first face image samples.

[0110] Considering that the time cost and labor cost of collecting face images of the same driver without wearing a mask and wearing a mask are relatively high, face masks can be randomly added to the collected face images of the driver without wearing a mask to obtain face images of the driver wearing a mask.

[0111] 502a: Determine the position correction annotation information corresponding to the face image sample. Among them, the position correction annotation information corresponding to the first face image sample is a preset value, and the position correction annotation information corresponding to the second face image sample is determined based on the position transformation relationship between the face key points marked in the second face image sample and the face key points marked in the first face image sample.

[0112] Assume that the position correction annotation information corresponding to each face image sample is in matrix form. For example, Then, for the first face image sample, since the position information of the face key points detected in the first face image sample is itself accurate, M can be a diagonal matrix, and the values of the elements on the diagonal of the diagonal matrix can be numbers close to 1. For example, the above preset value is 1; for the second face image sample, a 1 、a 2 、a 3 、a 4 、t x 、t y can be determined based on the position transformation relationship between the face key points marked in the second face image sample and the face key points marked in the first face image sample. Moreover, when determining a 1 、a 2 、a 3 、a 4 、t x 、t y generally considered are the key points covered by the mask, such as the key points on the chin and cheeks of the face.

[0113] 503a: Use the face image sample as the input, and the position correction annotation information corresponding to the face image sample, the position annotation information of the face key points, and the annotation information of whether a mask is worn as the output to train the preset model to obtain a key point analysis model.

[0114] Specifically, during implementation, the preset model can be trained with the face image sample as the input, and the position correction annotation information corresponding to the face image sample, the position annotation information of the face key points, and the annotation information of whether a mask is worn as the output. The training is terminated until it is determined that the face analysis accuracy of the preset model reaches the preset accuracy, and the current preset model is used as the key point analysis model.

[0115] 402: Judge whether the driver wears a mask based on the indication information. If not, go to 403; if so, go to 404.

[0116] 403: Determine the first face pose information of the driver based on the position information of the face key points.

[0117] 404: Based on the position correction information, perform correction processing on the position information of the first key point among the face key points, where the first key point is the face key point blocked by the mask.

[0118] For example, for each first key point, the position information of the first key point can be corrected according to the following formula:

[0119]

[0120] where (x * , y * ) are the coordinates of the first key point after correction, (x, y) are the coordinates of the first key point before correction, is the position correction information of the face key point.

[0121] In this way, the accuracy of the position of the face key point blocked by the mask can be corrected, thereby improving the accuracy of subsequent analysis of driving behavior.

[0122] 405: Based on the position information of the first key point after correction and the position information of the second key point, determine the first face pose information of the driver, where the second key point refers to the face key points other than the first key point.

[0123] 406: Determine the second face pose information of the driver based on the eye region data in the face image.

[0124] 407: Fuse the first face pose information and the second face pose information to obtain the target face pose information.

[0125] 408: Analyze whether the driver shows irregular driving behavior based on the target face pose information.

[0126] The above process will be introduced below taking distracted driving as an example of irregular driving behavior.

[0127] Figure 6A schematic diagram of the process for distracted driving detection provided by an embodiment of this application. Multiple images collected by a DMS camera are obtained, and face region detection is performed on each image to obtain the face bounding box of the driver. Then, face key point detection is performed on the face bounding box to obtain the position information of the face key points, and mask recognition is performed within the face bounding box. If a mask is recognized, the position information of the key points covered by the mask is corrected. Then, based on the corrected position information of the face key points, the position information of the face key points not covered by the mask, and the image data of the two eyes within the face bounding box, face pose analysis is performed on the driver to obtain the distraction probability of the image. Then, based on the distraction probabilities of multiple images, it is determined whether the driver exhibits distracted behavior.

[0128] To achieve the above objective, a face position detection model, a key point analysis model, and two face pose analysis models can be pre-trained. Among them, the face position detection model is used to determine the position of the face bounding box of the driver in the image; the key point analysis model is used to detect the position information of the face key points in the face image, the information on whether a mask is worn on the face, and the position correction information of the key points; one face pose analysis model is used to determine the face pose information of the driver based on the position information of the face key points; and the other face pose analysis model is used to determine the face pose information of the driver based on the image data corresponding to the two eyes. Here, any face pose information includes yaw angle, pitch angle, and roll angle.

[0129] The following separately introduces these several models.

[0130] 1. Face position detection model.

[0131] The loss function of the face position detection model is defined as follows:

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140] Among them, \(t\) i represents the predicted face bounding box of the \(i\)-th image sample, Denote the annotated face bounding box of the \(i\)-th image sample, denote \(IoU\) as a function of ; denote \(\beta\) 2 as a function of ; are the coordinates of the upper left vertex and the lower right vertex of the predicted face bounding box in the \(i\)-th image sample respectively, are the coordinates of the upper left vertex and the lower right vertex of the annotated face bounding box in the \(i\)-th image sample respectively.

[0141] 2. Key point analysis model.

[0142] When performing face key point detection, the face images usually do not have masks. Due to the epidemic, drivers need to wear masks. At this time, the key points covered by the masks are likely to have position offsets. Therefore, the following measures are taken for optimization:

[0143] 1) Optimize the training data. Add masks to some of the collected face image samples to increase the robustness of the key point analysis model.

[0144] 2) Optimize the loss function. The loss function \(L\) related to face key point detection pts is defined as follows.

[0145]

[0146] where \(n\) is the total number of face image samples, \(m\) is the total number of face image samples with masks, \(C\) is the number of key points in a face image sample, is the coordinate of the \(c\)-th key point in the \(i\)-th face image sample, and \(\alpha\) is the ratio of the loss value \(Loss\) of generating key point detection data with masks in the previous step to the loss \(Loss\) of key points without masks during the training process.

[0147] 3) In practical applications, the accuracy of the face pose analysis model corresponding to the face key points (i.e., the above-mentioned first face pose analysis model) depends highly on the key points of the face chin and cheeks. When the driver wears a mask, these key points will be blocked by the mask, so the positions of the detected key points are inaccurate. To improve the analysis accuracy of the first face pose analysis model for face poses and also to improve the analysis accuracy of subsequent driving behaviors, the positions of the key points blocked by the mask can be corrected.

[0148] In specific implementation, a face image of the same person without wearing a mask can be obtained as the first face image sample, and a face image of the same person wearing a mask can be obtained as the second face image sample. The positions of the face key points in the first face image sample and the second face image sample are respectively marked. For the target key points covered by the mask, such as the face key points on the chin and cheeks, the transformation matrix from the second face image sample to the first face image sample can be determined based on the positions of the target key points in the second face image sample and the first face image sample. The transformation matrix can represent the position difference of the target key points when covered by the mask and when not covered by the mask. For example, the transformation matrix M is defined as follows:

[0149]

[0150]

[0151] where (x *' , y *' ) are the coordinates of the target key point when not covered by the mask, and (x', y') are the coordinates of the target key point when covered by the mask.

[0152] That is, the key point analysis model can learn the conversion relationship between the positions of the target key points when covered by the mask and when not covered by the mask. Therefore, subsequently, when the driver wears a mask, based on the position correction information of the key points output by the key point analysis model, the position of the detected key points covered by the mask (i.e., the target key points) can be corrected to improve the accuracy of key point detection for faces wearing masks.

[0153] 3. Face pose analysis model.

[0154] Whether training the face pose analysis model based on key points (i.e., the above-mentioned first face pose analysis model) or training the face pose analysis model based on eye region data (i.e., the above-mentioned second face pose analysis model), the loss function of face pose regression can be defined as follows:

[0155]

[0156] where λ y , λ p , λ r are the loss weights for multi-task training, corresponding to the preset weights of yaw, pitch, and roll respectively. And λ p is a relatively large ratio because there is less pitch data in the training data. l y is the yaw output by the model, is the labeled yaw, l p is the pitch output by the model, is the marked pitch, l r is the roll output by the model, is the marked roll.

[0157] In practical applications, when the driver looks left or right for a long time, the lines of sight of the left and right eyes may be different. The state of the driver's line of sight when frequently checking the mobile phone and being distracted and the normal driving state are also different when the driver's face is in a normal posture. Therefore, information about the eye dimension is added on the basis of key points to improve the accuracy of distracted driving detection.

[0158] Figure 7 FIG. is a schematic diagram of the process of another distracted driving detection provided by an embodiment of the present application. The input has two dimensions. One is the coordinate information of the face key points, such as the coordinates of 96 key points, and the other is the left and right eye image data. The sizes of the left and right eye images are both 64x64, and they are spliced into a 128x64 image. Based on the positions of 96 key points and the eye images, a convolutional neural network (CNN) can be used to perform numerical regression to predict distracted driving. When predicting, a set of key points and a set of eye images both obtain a set of face pose information: yaw, pitch, and roll.

[0159] It is assumed that 96 key points are input into a 2-layer hidden layer CNN to regress yaw1, pitch1, and roll1. The eye images are subjected to feature extraction through convolutional layers, pooling layers, etc., and then yaw2, pitch2, and roll2 are regressed through a linear activation function.

[0160] Further, the face pose information on the two dimensions can be fused according to the following formula:

[0161]

[0162] pitch = pitch1 + α * pitch2

[0163] roll = roll1;

[0164] Among them, yaw1, pitch1, and roll1 are all between -90 and 90. When pitch1 < -5, α = 1, and in other cases α = 0. The correction of roll2 is ignored.

[0165] In specific implementation, among N consecutive face images, if the proportion of the number of images in which any one of yaw1, pitch1, and roll1 exceeds a given angle range, such as (-40, 40), exceeds 90%, it is determined that the driver has a distracted driving behavior, and a voice alarm can be given to remind the driver to maintain safe driving, thereby reducing traffic safety hazards and also reducing the management difficulty of the vehicle operator. N is an integer.

[0166] Based on the same inventive concept, an embodiment of the present application further provides a driving behavior analysis device. The principle of the driving behavior analysis device for solving problems is similar to that of the above-mentioned driving behavior analysis method. Therefore, for the implementation of the driving behavior analysis device, reference can be made to the implementation of the driving behavior analysis method, and repeated parts will not be elaborated. Figure 8 FIG. 4 is a schematic structural diagram of a driving behavior analysis device provided by an embodiment of the present application, including an image analysis module 801, a first determination module 802, a second determination module 803, an information fusion module 804, and a behavior analysis module 805.

[0167] The image analysis module 801 is configured to analyze the acquired face image of the driver to obtain face information, and the face information at least includes position information of face key points.

[0168] The first determination module 802 is configured to determine first face pose information of the driver based on the position information of the face key points.

[0169] The second determination module 803 is configured to determine second face pose information of the driver based on the eye region data in the face image.

[0170] The information fusion module 804 is configured to fuse the first face pose information and the second face pose information to obtain target face pose information.

[0171] The behavior analysis module 805 is configured to analyze whether the driver has an irregular driving behavior based on the target face pose information.

[0172] In some embodiments, the face information further includes indication information on whether a mask is worn.

[0173] The first determination module 802 is specifically configured to, when the indication information indicates that the driver is not wearing a mask, determine the first face pose information of the driver based on the position information of the face key points.

[0174] In some embodiments, the face information further includes position correction information of face key points, and further includes:

[0175] A correction module 806 is configured to, when the indication information indicates that the driver is wearing a mask, perform correction processing on the position information of a first key point among the face key points based on the position correction information, where the first key point is a face key point blocked by the mask.

[0176] The second determination module 803 is specifically configured to determine the first face pose information of the driver based on the corrected position information of the first key point and the position information of the second key point, where the second key point refers to the face key points other than the first key point.

[0177] In some embodiments, the information fusion module 804 is specifically configured to fuse the first face pose information and the second face pose information according to the following formula to obtain the target face pose information:

[0178]

[0179] pitch = pitch1 + α * pitch2

[0180] roll = roll1;

[0181] where yaw1 is the yaw angle in the first face pose information, pitch1 is the pitch angle in the first face pose information, roll1 is the roll angle in the first face pose information; yaw2 is the yaw angle in the second face pose information, pitch2 is the pitch angle in the second face pose information, yaw is the yaw angle in the target face pose information, pitch is the pitch angle in the target face pose information, roll is the roll angle in the target face pose information, and when pitch1 is less than the set value, α = 1; when pitch1 is not less than the set value, α = 0.

[0182] In some embodiments, the image analysis module 801 is specifically configured to:

[0183] Analyze the face image through a key point analysis model to obtain the face information;

[0184] where the key point analysis model is trained according to the following steps:

[0185] Obtain face image samples, where the face image samples include first face image samples without wearing masks and second face image samples obtained by adding masks to some of the first face image samples;

[0186] Determine the position correction annotation information corresponding to the face image samples, where the position correction annotation information corresponding to the first face image samples is a preset value, and the position correction annotation information corresponding to the second face image samples is determined based on the position transformation relationship between the face key points marked in the second face image samples and the face key points marked in the first face image samples;

[0187] Taking the face image sample as the input, and the position correction annotation information corresponding to the face image sample, the position annotation information of the face key points, and the annotation information on whether a mask is worn as the output, a preset model is trained to obtain the key point analysis model.

[0188] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module may be integrated in one processor, or may exist physically alone, or two or more modules may be integrated in one module. The coupling between each module can be realized through some interfaces, and these interfaces are usually electrical communication interfaces, but it does not exclude the possibility of being mechanical interfaces or other forms of interfaces. Therefore, the modules described as separate components may or may not be physically separated, and may be located in one place, or may be distributed to different positions of the same or different devices. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0189] Figure 9 FIG. 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device includes physical devices such as a transceiver 901 and a processor 902. Among them, the processor 902 may be a central processing unit (CPU), a microprocessor, an application specific integrated circuit, a programmable logic circuit, a large scale integrated circuit, or a digital processing unit, etc. The transceiver 901 is used for the electronic device to perform data transceiver with other devices.

[0190] The electronic device may further include a memory 903 for storing software instructions executed by the processor 902. Of course, it may also store some other data required by the electronic device, such as the identification information of the electronic device, the encryption information of the electronic device, user data, etc. The memory 903 may be a volatile memory, such as a random access memory (RAM); the memory 903 may also be a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the memory 903 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 903 may be a combination of the above memories.

[0191] In the embodiments of the present application, the specific connection medium between the above-mentioned processor 902, memory 903, and transceiver 901 is not limited. In the embodiments of the present application, Figure 9 only the case where the memory 903, the processor 902, and the transceiver 901 are connected through the bus 904 is taken as an example for illustration. The bus is represented by a thick line in Figure 9 The connection manners between other components are only for schematic illustration and are not restrictive. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0192] The processor 902 can be dedicated hardware or a processor running software. When the processor 902 can run software, the processor 902 reads the software instructions stored in the memory 903 and, under the drive of the software instructions, executes the driving behavior analysis method involved in the foregoing embodiments.

[0193] The embodiments of the present application also provide a storage medium. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the driving behavior analysis method involved in the foregoing embodiments.

[0194] In some possible implementation manners, each aspect of the driving behavior analysis method provided in the present application can also be implemented in the form of a program product. The program product includes program code. When the program product runs on the electronic device, the program code is used to cause the electronic device to execute the driving behavior analysis method involved in the foregoing embodiments.

[0195] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0196] In the embodiments of the present application, the program product for driving behavior analysis may be a CD-ROM and include program code, and can run on a computing device. However, the program product of the present application is not limited to this. In this document, a readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0197] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0198] The program code contained on a readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, radio frequency (RF), etc., or any suitable combination of the above.

[0199] The program code for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network such as a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0200] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0201] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0202] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0203] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0204] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0205] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0206] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0207] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A driving behavior analysis method, characterized in that, it includes: Input the obtained face image of the driver into a pre-trained key point analysis model to obtain face information, where the face information includes the position information of face key points, the indication information of whether a mask is worn, and the position correction information of face key points; When the indication information indicates that the driver wears a mask, based on the position correction information, correct the position information of the first key point among the face key points. The first key point is the face key point blocked by the mask. Based on the corrected position information of the first key point and the position information of the second key point, determine the first face pose information of the driver. The second key point refers to the face key points other than the first key point; Based on the eye region data in the face image, determine the second face pose information of the driver; Fuse the first face pose information and the second face pose information to obtain target face pose information, where the target face pose information includes yaw angle, pitch angle and roll angle; When the target face pose information corresponding to multiple consecutive face images meets a preset condition, it is determined that the driver has an irregular driving behavior. The preset condition includes that the proportion of face images with pose angles exceeding a given angle range among the multiple face images exceeds a preset proportion, and the pose angle is any one of the yaw angle, pitch angle and roll angle.

2. The method according to claim 1, characterized in that, it further includes: When the indication information indicates that the driver does not wear a mask, based on the position information of the face key points, determine the first face pose information of the driver.

3. The method according to claim 1 or 2, characterized in that, Fuse the first face pose information and the second face pose information according to the following formula to obtain target face pose information: pitch = pitch1 + α * pitch2 roll = roll1; where, yaw1 is the yaw angle in the first face pose information, pitch1 is the pitch angle in the first face pose information, roll1 is the roll angle in the first face pose information; yaw2 is the yaw angle in the second face pose information, pitch2 is the pitch angle in the second face pose information, yaw is the yaw angle in the target face pose information, pitch is the pitch angle in the target face pose information, roll is the roll angle in the target face pose information. When pitch1 is less than the set value, α = 1; when pitch1 is not less than the set value, α = 0.

4. The method according to claim 1, characterized in that, The key point analysis model is trained according to the following steps: Obtain face image samples, where the face image samples include first face image samples without wearing masks and second face image samples obtained by adding masks to some of the first face image samples; Determine the position correction annotation information corresponding to the face image sample, wherein the position correction annotation information corresponding to the first face image sample is a preset value, and the position correction annotation information corresponding to the second face image sample is determined based on the position transformation relationship between the face key points annotated in the second face image sample and the face key points annotated in the first face image sample; Use the face image sample as the input, and the position correction annotation information corresponding to the face image sample, the position annotation information of the face key points, and the annotation information indicating whether a mask is worn as the output to train a preset model to obtain the key point analysis model.

5. A driving behavior analysis device, Characterized in that, It includes: An image analysis module, configured to input the acquired face image of the driver into a pre-trained key point analysis model to obtain face information, where the face information includes the position information of the face key points, the indication information indicating whether a mask is worn, and the position correction information of the face key points; A first determination module, configured to, when the indication information indicates that the driver wears a mask, based on the position correction information, perform correction processing on the position information of the first key point among the face key points, where the first key point is the face key point blocked by the mask, and based on the corrected position information of the first key point and the position information of the second key point, determine the first face pose information of the driver, and the second key point refers to the face key points other than the first key point; A second determination module, configured to determine the second face pose information of the driver based on the eye region data in the face image; An information fusion module, configured to fuse the first face pose information and the second face pose information to obtain target face pose information, where the target face pose information includes yaw angle, pitch angle, and roll angle; A behavior analysis module, configured to determine that the driver has an irregular driving behavior when the target face pose information corresponding to multiple consecutive face images meets a preset condition, where the preset condition includes that the proportion of face images with pose angles exceeding a given angle range among the multiple face images exceeds a preset proportion, and the pose angle is any one of the yaw angle, pitch angle, and roll angle.

6. The device according to claim 5, Characterized in that, The first determination module is further configured to: When the indication information indicates that the driver does not wear a mask, determine the first face pose information of the driver based on the position information of the face key points.

7. An electronic device, Characterized in that, It includes: At least one processor, and a memory communicatively connected to the at least one processor, wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.

8. A storage medium, Characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Image recognition method and device, computer equipment and storage medium

    CN111507138A

  • Fatigue state identification method and system based on deep learning, and storage medium

    CN112163470A

  • Mask-wearing face recognition method based on eye attention mechanism

    CN112818901A