Attention detection method and device, electronic equipment and storage medium

By extracting feature and detecting the driver's face images, identifying fatigue behavior and degree, and outputing attention warning information, the problem of insufficient adaptability, accuracy and sensitivity of attention detection in the prior art is solved, and efficient driver attention monitoring is achieved.

CN120088762APending Publication Date: 2025-06-03BEIJING UCAS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411939379.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve high adaptability, high accuracy and high sensitivity attention detection, especially when different camera installation positions and driver race, gender, and age differences.

Method used

By obtaining the driver's face image, feature extraction and object detection are performed, facial feature points, eye feature points and mouth feature points, target fatigue behavior and fatigue degree, and output attention warning information based on object detection results.

Benefits of technology

It improves the accuracy and sensitivity of attention detection, can be applied to a variety of scenarios, effectively reducing safety hazards during driving and reducing traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088762A_ABST
    Figure CN120088762A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an attention detection method and device, electronic equipment and a storage medium, and can solve the problem that attention detection with high adaptability, high accuracy and high sensitivity cannot be realized by some attention monitoring methods which are commonly used at present. The method comprises the following steps: acquiring a to-be-detected image, wherein the to-be-detected image comprises a face image of a driver; feature extraction and object detection are carried out on the to-be-detected image, face feature points and an object detection result included in the to-be-detected image are obtained, and the face feature points at least comprise face feature points, eye feature points and mouth feature points; determining a target fatigue behavior of the driver according to the face feature points; determining the fatigue degree of the driver according to the target fatigue behavior; according to the fatigue degree and the object detection result, attention early warning information is output, and the attention early warning information is used for indicating the attention distraction condition of the driver and the corresponding early warning action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of attention detection, and in particular, to an attention detection method, apparatus, electronic device, and storage medium. Background Art

[0002] According to research by the National Highway Traffic Safety Administration, traffic accidents caused by driver distraction account for more than 70%. Real-time monitoring of the driving conditions of drivers, accurate identification of driving behaviors related to driver distraction, and timely and effective early warning can effectively reduce potential safety hazards during driving and are of great significance for reducing the occurrence of traffic accidents. Due to the diversity of camera installation positions in different practical scenarios, as well as differences in race, gender, and age among drivers, and the fact that the head postures of drivers are constantly changing, some commonly used attention monitoring methods may not be able to achieve highly adaptable, highly accurate, and highly sensitive attention detection. Summary of the Invention

[0003] Based on this, it is necessary to provide an attention detection method, apparatus, electronic device, and storage medium for the above technical problems.

[0004] In a first aspect, the embodiments of the present application provide an attention detection method, and the attention detection method includes: obtaining an image to be detected, where the image to be detected includes a face image of a driver;

[0005] Performing feature extraction and object detection on the image to be detected to obtain face feature points and an object detection result included in the image to be detected, where the face feature points at least include: facial feature points, eye feature points, and mouth feature points;

[0006] Determining a target fatigue behavior of the driver according to the face feature points;

[0007] Determining the fatigue level of the driver according to the target fatigue behavior;

[0008] Outputting an attention warning message according to the fatigue level and the object detection result, where the attention warning message is used to indicate the attention distraction situation of the driver and the corresponding warning action.

[0009] As an optional implementation manner, in the first aspect of the embodiments of the present application, the determining the target fatigue behavior of the driver according to the face feature points includes:

[0010] Determining the head deflection fatigue behavior of the driver according to the facial feature points;

[0011] Determine the eye closing fatigue behavior of the driver based on the eye feature points;

[0012] Determine the mouth opening and closing fatigue behavior of the driver based on the mouth feature points.

[0013] As an optional implementation manner, in the first aspect of the embodiments of the present application, the determining the head deflection fatigue behavior of the driver according to the facial feature points includes:

[0014] Determine target horizontal feature points and target vertical feature points from the facial feature points;

[0015] Calculate the real-time three-dimensional posture of the head according to the target horizontal feature points and the target vertical feature points;

[0016] Compare the real-time rotation angle with the pre-stored standard three-dimensional posture to obtain the head deflection fatigue behavior.

[0017] As an optional implementation manner, in the first aspect of the embodiments of the present application, the determining the eye closing fatigue behavior of the driver according to the eye feature points includes:

[0018] Determine a plurality of eye closing feature points and a plurality of eye reference feature points among the eye feature points;

[0019] Calculate the eye closing area according to the plurality of eye closing feature points;

[0020] Calculate the eye reference area according to the plurality of eye reference feature points;

[0021] Determine the eye closing fatigue behavior according to the eye closing area and the eye reference area.

[0022] As an optional implementation manner, in the first aspect of the embodiments of the present application, the determining the mouth opening and closing fatigue behavior of the driver according to the mouth feature points includes:

[0023] Determine a plurality of mouth opening and closing feature points and a plurality of mouth reference feature points among the mouth feature points;

[0024] Calculate the mouth opening and closing area according to the plurality of mouth opening and closing feature points;

[0025] Calculate the mouth reference area according to the plurality of mouth reference feature points;

[0026] Determine the mouth opening and closing fatigue behavior according to the mouth opening and closing area and the mouth reference area.

[0027] As an alternative implementation, in the first aspect of the embodiments of the present application, determining the fatigue level of the driver according to the target fatigue behavior includes:

[0028] Determining a corresponding fatigue value according to the target fatigue behavior;

[0029] Continuously monitoring the target fatigue behavior and statistically analyzing the fatigue values within a preset time period to obtain target fatigue data;

[0030] Determining the fatigue level of the driver according to the target fatigue data.

[0031] As an alternative implementation, in the first aspect of the embodiments of the present application, the method further includes:

[0032] Real-time detecting the driving speed of the vehicle;

[0033] When it is detected that the driving speed is greater than the function activation speed threshold, starting to acquire the image to be detected;

[0034] When it is detected that the driving speed is less than the function stop speed threshold, stopping to acquire the image to be detected.

[0035] In a second aspect, an attention detection device provided by the embodiments of the present application includes: an acquisition module for acquiring an image to be detected, where the image to be detected includes a face image of a driver;

[0036] A processing module for extracting features from the image to be detected to obtain face feature points included in the image to be detected, where the face feature points at least include: facial feature points, eye feature points, and mouth feature points;

[0037] The processing module is further configured to determine the target fatigue behavior of the driver according to the face feature points;

[0038] The processing module is further configured to determine the fatigue level of the driver according to the target fatigue behavior;

[0039] An output module for outputting a fatigue warning message corresponding to the fatigue level according to the fatigue level.

[0040] In a third aspect, an electronic device provided by the embodiments of the present application includes:

[0041] A memory storing executable program code;

[0042] A processor coupled to the memory;

[0043] The processor calls the executable program code stored in the memory and executes the attention detection method in the first aspect of the embodiments of the present application.

[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium that stores a computer program, and the computer program causes a computer to execute the attention detection method in the first aspect of the embodiments of the present application. The computer-readable storage medium includes ROM / RAM, a magnetic disk, an optical disc, or the like.

[0045] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute some or all of the steps of any one of the methods in the first aspect.

[0046] In a sixth aspect, an embodiment of the present application provides an application publishing platform. The application publishing platform is used to publish a computer program product. When the computer program product runs on a computer, it causes the computer to execute some or all of the steps of any one of the methods in the first aspect.

[0047] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0048] The embodiments of the present application provide an attention detection method, device, electronic device, and storage medium. An image to be detected is obtained, and the image to be detected includes a face image of a driver. Feature extraction and object detection are performed on the image to be detected to obtain face feature points and an object detection result included in the image to be detected. The face feature points at least include: facial feature points, eye feature points, and mouth feature points. According to the face feature points, the target fatigue behavior of the driver is determined. According to the target fatigue behavior, the fatigue level of the driver is determined. According to the fatigue level and the object detection result, an attention warning message is output, and the attention warning message is used to indicate the attention dispersion situation of the driver and the corresponding warning actions. In this solution, whether the driver is fatigued is judged by various feature points recognized in the facial image. In addition, through the preset object detection in the image, it can be judged whether the driver is engaged in dangerous driving, thereby realizing the attention detection of the driver, which can effectively improve the accuracy and sensitivity of attention detection and can be applied to a large number of scenarios. Description of the Drawings

[0049] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0050] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0051] Figure 1 is a flowchart illustration of an attention detection method provided by an embodiment of the present application Figure 1 ;

[0052] Figure 2 is a flowchart illustration of an attention detection method provided by an embodiment of the present application Figure 2 ;

[0053] Figure 3 is a schematic structural diagram of an attention detection device provided by an embodiment of the present application;

[0054] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0055] In order to more clearly understand the above objects, features, and advantages of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0056] The terms "first" and "second" in the specification and claims of the present application are used to distinguish different objects, rather than to describe a specific order of the objects.

[0057] The terms "including" and "having" in the embodiments of the present application and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0058] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0059] According to research by the National Highway Traffic Safety Administration, traffic accidents caused by driver distraction account for more than 70%. It can be seen that driver distraction has become the main factor leading to traffic accidents. Therefore, real-time monitoring of the driving status of drivers, accurate identification of driving behaviors related to driver distraction, and timely and effective early warning can effectively reduce potential safety hazards during driving and are of great significance for reducing the occurrence of traffic accidents.

[0060] According to the type of perception, driver attention monitoring methods can be mainly divided into three categories: methods based on vehicle driving status, methods based on driver physiological characteristics, and methods based on computer vision. Among them, the method based on computer vision has become the current mainstream research direction due to its advantages such as ease of use, generality, and low cost in practical applications.

[0061] Currently, it is challenging to monitor the driver's attention state in a real driving environment based on vision in real time. Existing technologies use cameras to collect real-time images of drivers, extract the facial features of drivers and identify relevant objects through neural networks, identify the eye closure, yawning, and head posture of drivers based on their facial features, and identify dangerous driving actions such as smoking and making phone calls based on the information of the identified objects. However, in different practical scenarios, the installation positions of cameras are diverse, and there are differences in race, gender, and age among drivers, and the head postures of drivers also keep changing, which results in low accuracy of traditional methods and may not be able to adapt to different application scenarios.

[0062] To solve some or all of the above technical problems, embodiments of the present application provide an attention detection method, apparatus, electronic device, and storage medium. An image to be detected is obtained, and the image to be detected includes a face image of a driver. Feature extraction and object detection are performed on the image to be detected to obtain face feature points and object detection results included in the image to be detected. The face feature points at least include: facial feature points, eye feature points, and mouth feature points. According to the face feature points, the target fatigue behavior of the driver is determined. According to the target fatigue behavior, the fatigue level of the driver is determined. According to the fatigue level and the object detection result, an attention warning message is output, and the attention warning message is used to indicate the attention dispersion situation of the driver and the corresponding warning actions. In this solution, whether the driver is fatigued is determined through various feature points recognized in the facial image. In addition, through the preset object detection in the image, whether the driver is engaged in dangerous driving can be determined, thereby realizing the attention detection of the driver, which can effectively improve the accuracy and sensitivity of attention detection and can be applied to a large number of scenarios.

[0063] As Figure 1 shown, Figure 1 FIG. is a flowchart of an attention detection method provided by an embodiment of the present application. The method may include the following steps:

[0064] 101. Obtain an image to be detected.

[0065] In an embodiment of the present application, the image to be detected may include a face image of a driver.

[0066] In some embodiments, the image to be detected may be obtained by real-time acquisition by a starlight camera, an infrared camera (meeting the usage requirements of night scenes), etc. provided inside the vehicle.

[0067] In some embodiments, the image to be detected may be obtained after a series of preprocessing. The preprocessing may include: size adjustment, aspect ratio stretching, normalization, color channel conversion, and other processing methods.

[0068] 102. Perform feature extraction and object detection on the image to be detected to obtain face feature points and object detection results included in the image to be detected.

[0069] In an embodiment of the present application, since the driver's attention dispersion usually includes: fatigue driving and attention transfer, whether the driver is fatigued can be determined through face feature points, and whether the attention is transferred can be determined through the object detection result.

[0070] In some embodiments, feature extraction is performed on the image to be detected, which may specifically include: inputting the image to be detected into a face feature point detection neural network model, such as the MediaPipe face detection model, to obtain high-precision positioning information of face feature points. The face feature point detection neural network model can be selected according to the specific requirements of the system, such as neural network models like MediaPipe face detection model, OpenFace, OpenVINO, OpenCV DNN, Dlib HOG, Dlib CNN, etc. According to the functional requirements of driver attention detection, relevant feature point sets such as head three-dimensional pose, eyes, mouth, reference position, etc. are screened and integrated. That is to say, the face feature points at least include: facial feature points, eye feature points, and mouth feature points.

[0071] In some embodiments, object detection is performed on the image to be detected, which may specifically include: inputting the image to be detected into an object detection neural network model to obtain object detection results, and the object detection results may indicate information such as object category, position, size, confidence, etc.

[0072] In some embodiments, the object detection results may specifically include information such as object category, object recognition box and its confidence, etc. By parsing the object recognition box information, the object center point position and object size information are extracted, so as to screen out the objects to be recognized (such as mobile phones, cigarettes, etc.) according to the position and size.

[0073] In some embodiments, in the actual scenario, there may be some misidentifications in object recognition. For example: an object that is not a mobile phone is identified as a mobile phone, or a mobile phone is recognized in the image, but the mobile phone is far from the driver and the driver has not operated it, etc. Therefore, in order to improve the accuracy of object recognition, misidentified objects can be filtered, which may specifically include: judging whether the center point position of the recognized object is located within the corresponding reference area of the face; verifying whether the width and height values of the object recognition box are respectively less than the preset width and height thresholds of the corresponding object.

[0074] In some embodiments, judging whether the center point position of the recognized object is located within the corresponding reference area of the face may specifically include:

[0075] Select a set of vertical feature points in the face reference area, which are P1 and P2 from top to bottom respectively, and the center point P3 of the known object to be measured. Connect the three points respectively to obtain the straight line L1 determined by P1 and P2, the straight line L2 determined by P3 and P1, the straight line L3 determined by P3 and P2, the straight line L4 passing through P1 and perpendicular to L1, and the straight line L5 passing through P2 and perpendicular to L1. Calculate the absolute value of the included angle θ between the straight line L1 and the straight line L2 L1L2 , the absolute value of the included angle θ between the straight line L1 and the straight line L3 L1L3 , judge θL1L2 and θ L1L3 are both less than 90 degrees. When θ L1L2 and θ L1L3 are both less than 90 degrees, it indicates that P3 is between the two parallel lines L4 and L5.

[0076] Similarly, select a set of horizontal feature points in the face reference region, which are P4 and P5 from left to right respectively, and the center point P3 of the known object to be measured. Connect the three points respectively to obtain the straight line L6 determined by P4 and P5, the straight line L7 determined by P3 and P4, the straight line L8 determined by P3 and P5, the straight line L9 passing through P3 and perpendicular to L6, and the straight line L10 passing through P4 and perpendicular to L6. Calculate the absolute value of the included angle θ L6L7 between the straight line L6 and the straight line L7, and the absolute value of the included angle θ L6L8 between the straight line L6 and the straight line L8, and determine whether θ L6L7 and θ L6L8 are both less than 90 degrees. When θ L6L7 and θ L6L8 are both less than 90 degrees, it indicates that P3 is between the two parallel lines L9 and L10.

[0077] When θ L1L2 , θ L1L3 , θ L6L7 , θ L6L8 are all less than 90 degrees, it indicates that the center point P3 of the object is within the square region formed by the two parallel lines L4 and L5 and the two parallel lines L9 and L10 (i.e., the corresponding face reference region).

[0078] In some embodiments, when the shape of the required reference region is other shapes, the horizontal (or vertical) judgment in the face reference region can be modified. Specifically, it may include: select a set of horizontal feature points in the face reference region, which are P4 and P5 from left to right respectively, define the horizontal distance between the feature points P4 and P5 as the reference distance D1. Then, calculate the distance D2 from the center point P3 of the object to be measured to the straight line L1. The determination condition is whether D2 is less than the product of D1 and a preset coefficient threshold T, and the coefficient threshold T is dynamically adjusted according to the ratio of the distances from P3 to P1 and P2 in the direction of the straight line L1. Through this method, various reference region shapes including but not limited to circular shapes can be flexibly defined.

[0079] It should be noted that in the traditional range judgment method based on fixed reference points and distance thresholds, when the head posture changes, it will be affected by perspective projection distortion, resulting in an obvious deviation of the actual judgment area relative to the expected face reference area. In this method, since the reference area is constructed based on the horizontal and vertical feature points of the face, the reference area can be dynamically adjusted with the change of the face posture, and the above deviation will not occur in this scenario, which can provide an accurate representation and judgment of the face reference area, and has high adaptability and accuracy in actual application scenarios, meeting the actual application requirements. When the object confidence level screened by the above method is higher than the preset confidence level threshold of the corresponding object and the duration exceeds the predetermined time threshold, it can be determined that there is a dangerous driving behavior (such as making a call or smoking) according to the category of the object (such as a mobile phone or a cigarette).

[0080] 103. Determine the target fatigue behavior of the driver according to the face feature points.

[0081] In the embodiments of the present application, the target fatigue behavior may include: head deflection, eye closure, mouth opening and closing, etc. Therefore, the target fatigue behavior of the driver can be determined according to the facial feature points, eye feature points, and mouth feature points in the face feature points.

[0082] 104. Determine the fatigue level of the driver according to the target fatigue behavior.

[0083] In the embodiments of the present application, different fatigue behaviors may correspond to different fatigue levels, and in the fatigue behavior, different levels can also be subdivided. Therefore, the fatigue level of the driver can be determined for the target fatigue behavior.

[0084] 105. Output an attention warning message according to the fatigue level and the object detection result.

[0085] In the embodiments of the present application, the attention warning message is used to indicate the attention dispersion situation of the driver and the corresponding warning actions. That is to say, the attention dispersion situation can indicate whether the driver has an attention dispersion situation, such as: looking at the mobile phone, closing eyes, lowering the head for a long time, etc.; the manifestation forms of the warning actions are various, and may include but are not limited to various means such as auditory alarms, visual prompts, and tactile feedback, so as to timely remind the driver to correct improper behaviors and thus ensure driving safety.

[0086] In some embodiments, for the same type of fatigue and dangerous driving behaviors, the warning actions can be executed at a certain cooling time interval to prevent frequent execution of warning actions during the continuous actions of the driver with different time lengths (for example, continuously executing warning actions during a call).

[0087] An embodiment of the present application provides an attention detection method, which includes obtaining an image to be detected, where the image to be detected includes a face image of a driver; performing feature extraction and object detection on the image to be detected to obtain face feature points and object detection results included in the image to be detected, and the face feature points at least include: facial feature points, eye feature points, and mouth feature points; determining the target fatigue behavior of the driver according to the face feature points; determining the fatigue level of the driver according to the target fatigue behavior; and outputting an attention warning message according to the fatigue level and the object detection result, where the attention warning message is used to indicate the attention dispersion situation of the driver and the corresponding warning actions. In this solution, it is determined whether the driver is fatigued by various feature points recognized in the facial image. In addition, through the preset object detection in the image, it can be determined whether the driver is driving dangerously, so as to realize the attention detection of the driver, which can effectively improve the accuracy and sensitivity of the attention detection and can be applied to a large number of scenarios.

[0088] As Figure 2 shown, Figure 2 FIG. is a flowchart of an attention detection method provided by an embodiment of the present application, and the method may further include the following steps:

[0089] 201. Real-time detect the driving speed of the vehicle.

[0090] In the embodiment of the present application, for scenarios where there is no need to strictly monitor the driver's attention during parking or slow driving, the start and stop of the driver attention monitoring function can be controlled based on a function activation vehicle speed threshold and a function stop vehicle speed threshold, so the driving speed of the vehicle can be detected in real time.

[0091] 202. When it is detected that the driving speed is greater than the function activation vehicle speed threshold, start to obtain the image to be detected.

[0092] In the embodiment of the present application, when the vehicle speed is higher than the set function activation vehicle speed threshold, the driver attention monitoring function will be activated; when the vehicle speed is lower than the set function stop vehicle speed threshold, these functions will be deactivated. Compared with the method that only relies on a single function activation vehicle speed threshold, this method can effectively avoid the frequent start and stop of the monitoring function when the vehicle speed fluctuates near the activation vehicle speed threshold.

[0093] In some embodiments, the above driver attention monitoring function includes but is not limited to eye closing behaviors at various levels, long-term head-down, left and right head-turning, left and right head-tilting behaviors, as well as smoking and making phone calls. However, this method is not applicable to the monitoring of yawning behavior. Since yawning is an involuntary behavior caused by fatigue, even in a parked or slow-driving state, yawning can still reflect the fatigue state of the driver. Therefore, the monitoring of yawning behavior is not affected by changes in vehicle speed.

[0094] 203. Extract features and detect objects from the image to be detected, and obtain the facial feature points and object detection results included in the image to be detected.

[0095] In the embodiments of the present application, for the description of step 203, please refer to the detailed description of step 102 in the above embodiments, and the embodiments of the present application will not be elaborated herein.

[0096] 204. Determine the head deflection fatigue behavior of the driver according to the facial feature points.

[0097] In the embodiments of the present application, using the facial feature points, through the spatial angle algorithm calculation, the three-dimensional head pose is obtained. By comparing the real-time three-dimensional head pose with the statistically normal driving three-dimensional head pose, behaviors such as long-term low head-up, left and right head turns, and left and right head tilts are identified.

[0098] In some embodiments, from the facial feature points, target horizontal feature points and target vertical feature points are determined; according to the target horizontal feature points and target vertical feature points, the real-time three-dimensional head pose is calculated; according to the real-time rotation angle and the pre-stored standard three-dimensional pose for comparison, the head deflection fatigue behavior is obtained.

[0099] It should be noted that using the three-dimensional coordinate information of the set of face feature points related to the three-dimensional head pose obtained through, for example, the MediaPipe face detection neural network model, through the spatial angle algorithm calculation, the three-dimensional head pose is obtained, including the rotation angles of the head relative to the image on the X, Y, and Z axes. When selecting multiple groups of horizontal and vertical feature points, for the calculated multiple groups of head X, Y, and Z axis rotation angles, data processing methods such as calculating weighted / unweighted averages according to the confidence levels of each group of horizontal and vertical feature points can be used for processing to obtain more stable and accurate head X, Y, and Z axis rotation angles.

[0100] Select one (or more) groups of horizontal and vertical feature points on the face, and select facial feature points that are not easily occluded under normal circumstances. For example, the horizontal feature points are the left outer canthus P1 and the right outer canthus P2, and the vertical feature points are the center of the eyebrows P3 and the tip of the nose P4. All feature points are represented by three-dimensional coordinates (x, y, z).

[0101] In some embodiments, the head X-axis rotation angle algorithm is to calculate the Euclidean distance D of the projections of points P3 and P4 on the plane formed by the Y-axis and the Z-axis x , and the Z-axis difference Δz between points P3 and P4 x , use D x and Δz x to perform an arcsine calculation to obtain the radian value θ of the X-axis rotation r,x , and finally convert θ r,x to the X-axis rotation angle θ d,x, that is, the rotation angle of the head along the X-axis.

[0102] Among them, the Euclidean distance D x is calculated by the formula: The difference in the Z-axis Δz x is calculated by the formula: Δz x = z (P3) - z (p4) ; The radian value θ of the rotation along the X-axis r,x is calculated by the formula: The rotation angle θ of the head along the X-axis d,x is calculated by the formula:

[0103] In some embodiments, the algorithm for the rotation angle of the head along the Y-axis is to calculate the Euclidean distance D of the projections of two points P1 and P2 on the plane formed by the X-axis and the Z-axis y , and the difference in the Y-axis Δz between the two points P1 and P2 y , use D y and Δz y to perform an arcsine calculation to obtain the radian value θ of the rotation along the Y-axis r,y , and finally convert θ r,y to the rotation angle θ of the head along the Y-axis d,y , that is, the rotation angle of the head along the Y-axis.

[0104] Among them, the Euclidean distance D y is calculated by the formula: The difference in the Y-axis Δz y is calculated by the formula: Δz y = z (P1) - z (p2) ; The radian value θ of the rotation along the Y-axis r,y is calculated by the formula: The rotation angle θ of the head along the Y-axis d,y is calculated by the formula:

[0105] In some embodiments, the algorithm for the rotation angle of the head along the Z-axis is to calculate the Euclidean distance D of the projections of two points P3 and P4 on the plane formed by the X-axis and the Y-axis z , and the difference in the X-axis Δx between the two points P1 and P2 z , use D z and Δx z to perform an arcsine calculation to obtain the radian value θ of the rotation along the Z-axis r,z , and finally convert θ r,z to the rotation angle θ of the head along the Z-axis d,z , that is, the rotation angle of the head along the Z-axis.

[0106] Among them, the Euclidean distance D z is calculated by the formula: The difference in the X-axis Δxz The calculation formula for it is: Δx z = x (P3) - x (p4) ; The radian value θ of the Z - axis rotation r,z The calculation formula for it is: The Z - axis rotation angle θ d,z The calculation formula for it is:

[0107] In some embodiments, to identify behaviors such as long - time low head - up, left - and - right head - turning, and left - and - right head tilting, it is necessary to statistically calculate the real - time three - dimensional head pose of the driver. Due to different installation positions of the camera and different body shapes and postures of the driver, the normal driving three - dimensional head poses of different drivers are unpredictable. Therefore, it is also necessary to statistically calculate the normal driving three - dimensional head pose of the driver. To statistically calculate the real - time three - dimensional head pose and the normal driving three - dimensional head pose of the driver, it is necessary to store and record the rotation angles of the head around the X, Y, and Z axes in chronological order. For example, the rotation angles of the X, Y, and Z axes are respectively stored in a data structure such as a data queue in a format similar to (time T, rotation angle θ d ), and optionally, data filtering is performed before storage. For example, the maximum and minimum values of the data are filtered. For example, data with an absolute value of the rotation angle θ d greater than 80 degrees is filtered.

[0108] When calculating the real - time three - dimensional head pose (the rotation angles of the head around the X, Y, and Z axes) of the driver using the rotation angles of the head around the X, Y, and Z axes stored in chronological order, methods such as unweighted or weighted moving average based on a time window or a data - quantity window can be adopted. For example: calculating the arithmetic mean of the data points within a time window of the last 1 s or a window of the last 20 data samples. Optionally, different weights can be assigned to the data at different time points when calculating the mean value to calculate the weighted mean.

[0109] When calculating the normal driving three - dimensional head pose of the driver using the rotation angles of the head around the X, Y, and Z axes stored in chronological order, because in the actual driving scenario, the driver has a few moments to view the left and right rearview mirrors and other scenarios. At this time, there are significant differences between the three - dimensional head pose of the driver and the three - dimensional head pose when driving straight ahead normally. It is necessary to filter out the rotation angles of the head around the X, Y, and Z axes of the driver in such scenarios. Therefore, methods such as unweighted or weighted percentage filtering average based on a time window or a data - quantity window can be adopted. For example: filtering out data points with data values less than the 25% threshold and data values greater than the 75% threshold within a time window of the last 2 minutes or a window of the last 2400 data samples, and then calculating the arithmetic mean of the data points. Optionally, different weights can be assigned to the data at different time points according to actual needs when calculating the mean value to calculate the weighted mean.

[0110] In some embodiments, the calculated real-time three-dimensional head pose of the driver is compared with the normal driving three-dimensional head pose of the driver. When the rotational angle changes of the real-time three-dimensional head pose of the driver with respect to the normal driving three-dimensional head pose of the driver on the X, Y, and Z axes exceed a certain angle threshold and the duration exceeds a predetermined time threshold, it is determined as behaviors such as low head-up, left and right head-turning, and left and right head-tilting.

[0111] Exemplarily, when the rotational angle change of the real-time three-dimensional head pose of the driver on the X axis with respect to the rotational angle of the normal driving three-dimensional head pose of the driver on the X axis exceeds -15 degrees and remains for more than 2 seconds, it is determined as a dangerous driving behavior of long-term head-down; when the rotational angle change on the X axis exceeds 15 degrees and remains for more than 2 seconds, it is determined as a dangerous driving behavior of long-term head-up. When the rotational angle change on the Y axis exceeds -25 degrees and remains for more than 3 seconds, it is determined as a dangerous driving behavior of long-term right head-turning; when the rotational angle change on the Y axis exceeds 25 degrees and remains for more than 3 seconds, it is determined as a dangerous driving behavior of long-term left head-turning. When the rotational angle change on the Z axis exceeds -20 degrees and remains for more than 3 seconds, it is determined as a dangerous driving behavior of long-term right head-tilting; when the rotational angle change on the Z axis exceeds 20 degrees and remains for more than 3 seconds, it is determined as a dangerous driving behavior of long-term left head-tilting.

[0112] In some embodiments, if it is necessary to determine the driver's gaze area, the threshold ranges of the rotational angles of the head on the X, Y, and Z axes in each gaze area can be calibrated according to the specific installation position of the camera to distinguish the driver's gaze area.

[0113] 205. Determine the eye closure fatigue behavior of the driver according to the eye feature points.

[0114] In the embodiments of the present application, the relationship between the eye closure feature value and the reference area feature value is calculated using the set of eye-related feature points as the eye closure degree. By comparing the real-time eye closure degree of the driver with the adaptive thresholds of the eye closure degrees at all levels, different levels of eye closure behaviors are identified.

[0115] In some embodiments, among the eye feature points, multiple eye closure feature points and multiple eye reference feature points are determined; the eye closure area is calculated according to the multiple eye closure feature points; the eye reference area is calculated according to the multiple eye reference feature points; and the eye closure fatigue behavior is determined according to the eye closure area and the eye reference area.

[0116] It should be noted that the relationship between the eye closure feature value and the reference area feature value is calculated using the set of eye-related feature points as the eye closure degree, and the eye closure degree is stored and recorded in chronological order.

[0117] Currently, the commonly used method for evaluating the degree of eye closure mainly relies on calculating the eye closure eigenvalue through eye-related feature points, such as the Eye Aspect Ratio (EAR). This method quantifies the opening and closing state of the eyes by calculating the EAR, so as to distinguish whether the eyes are open or closed. However, when the driver's head rotates significantly relative to the camera, the feature points in the face image captured by the camera will appear perspective projection deformation in the rotation direction, resulting in the shortening or elongation of the distance between feature points, and further causing deviation in the EAR calculated based on eye-related feature points.

[0118] Therefore, the use of reference region feature points can be introduced. The reference region feature points do not change their positions during the eye closure process, thus providing a stable benchmark, and will undergo the same perspective projection deformation as the eye-related feature points when the driver's head rotates. Therefore, by calculating the relationship (such as the ratio relationship) between the eye closure eigenvalue and the reference region eigenvalue as the eye closure degree, the deviation of the eye closure eigenvalue introduced by the perspective projection deformation can be effectively eliminated. Specifically, the eye closure eigenvalue and the reference region eigenvalue can be quantified using area features or aspect ratio features. This method can improve the accuracy and robustness of the eye closure state judgment under the condition of head pose change.

[0119] In some embodiments, a group of eye-related feature points are selected, such as five feature points P1, P2, P3, P4, P5 from left to right on the upper eyelid, and five feature points P6, P7, P8, P9, P10 from left to right on the lower eyelid. The above ten feature points are used to enclose the eye region and calculate its area. A group of reference region feature points are selected, such as five feature points P11, P12, P13, P14, P15 from left to right above the eyebrows, and five feature points P16, P17, P18, P19, P20 from left to right on the lower eyelid, which are used to enclose the reference region and calculate its area.

[0120] Among them, the method for calculating the area enclosed by the feature points may include: dividing the region enclosed by the ten eye-related feature points into four sub-regions, each sub-region consists of four feature points. Sub-region 1 is composed of feature points P1, P2, P6, P7, sub-region 2 is composed of feature points P2, P3, P7, P8, sub-region 3 is composed of feature points P3, P4, P8, P9, and sub-region 4 is composed of feature points P4, P5, P9, P10. Calculate the quadrilateral area of each sub-region, and add the areas of these four sub-regions to obtain the eye opening and closing area. To simplify the calculation, the quadrilateral area of each sub-region can be approximated as a trapezoidal area for calculation, so as to improve the calculation efficiency and real-time performance. The calculation method of the reference region area is similar. By calculating the ratio of the eye opening and closing area to the reference region area as the eye closure degree.

[0121] In some embodiments, it is also necessary to consider the accuracy differences of the left and right eye feature points of the driver's head at different rotation angles. Specifically, when the camera is located on the left side of the driver's face, the left eye is closer to the camera and has a smaller deflection angle compared to the right eye. At this time, the detection accuracy of the left eye feature points is relatively high. Vice versa. Therefore, it is necessary to dynamically and adaptively allocate weights to the eye closure degrees of the left and right eyes according to factors such as the rotation angle of the driver's head. By performing weighted averaging on the eye closure degrees of the left and right eyes, a comprehensive eye closure degree is calculated.

[0122] In some embodiments, according to the head Y-axis rotation angle θ d,y dynamically calculate the weights of the left and right eyes, and at the same time set a Y-axis rotation angle threshold θ t for the validity of feature points. When the head Y-axis rotation angle θ d,y is 0 degree, the weights of both the left and right eyes are 0.5; when the absolute value of the head Y-axis rotation angle θ d,y is greater than the Y-axis rotation angle threshold θ t , if the head rotates to the left, the weight of the right eye is 1 and the weight of the left eye is 0. Conversely, the weight of the left eye is 1 and the weight of the right eye is 0. When the head Y-axis rotation angle θ d,y is in the range from 0 to θ t , the weights of the left and right eyes are calculated according to the ratio of θ d,y to θ t . Finally, the comprehensive eye closure degree ECR is calculated by performing weighted averaging on the eye closure degrees ECR l and ECR r of the left and right eyes.

[0123] Among them, the weight adjustment amount W ad is calculated by the formula: W ad = θ d,y / θ t / 2; the weight of the right eye W r is calculated by the formula: W r = max(0, min(1, 0.5 + W ad )); the weight of the left eye W l is calculated by the formula: W l = 1 - W r ; the formula for calculating the eye closure degree ECR is: ECR = ECR l × W l + ECR r × W r .

[0124] In some embodiments, by comparing the obtained high-precision real-time eye closure degree of the driver with the adaptive thresholds of eye closure degrees at each level, different degrees of eye closure behaviors are identified.

[0125] In an actual fatigue driving scenario, as the degree of drowsiness deepens, the eye state of the driver usually undergoes a series of changes, gradually transitioning from a normal state to a long-term semi-closed state, and finally even developing into a long-term closed-eye state. During this process, the driver's attention gradually decreases, and the driving risk also increases accordingly. If only the closed-eye behavior of the driver is recognized, it will lead to a significant neglect of the driver's fatigue state and miss the valuable opportunity for early warning of the fatigue state. This method cannot effectively monitor and warn most fatigue driving situations, and thus it is difficult to meet the requirements of practical applications. Even for the same driver, there are differences in their eye states under different driving scenarios (different driving times, different driving states, different driving light conditions, etc.). Therefore, based on the highly accurate eye closure degree obtained above, the adaptive thresholds of different levels of eye closure degree can be calculated in real time, and the eye closure degrees of different levels of different drivers can be recognized according to these adaptive thresholds. For example: based on the statistical calculation method of eye closure degree data, to obtain the normal eye closure degree and the minimum eye closure degree of the driver, and further divide the adaptive thresholds of different levels of eye closure degree.

[0126] In some embodiments, the normal eye closure degree of the driver is obtained through the statistical calculation of the eye closure degree data. The calculation of the normal eye closure degree of the driver can adopt non-weighted or weighted moving average methods based on a time window or a data quantity window, and record the maximum moving average value as the normal eye closure degree of the driver. The purpose of recording the maximum moving average value as the normal eye closure degree of the driver is to prevent the continuous decrease of the driver's eye closure degree as the driver's drowsiness increases, resulting in the decrease of the normal eye closure degree benchmark, so that the eye closure behavior of the driver cannot be accurately recognized accordingly. The minimum eye closure degree of the driver is obtained through the statistical calculation of the eye closure degree in the normal blinking state of the driver in the eye closure degree data. The calculation of the minimum eye closure degree of the driver can adopt methods such as averaging based on a minimum heap. Within the range of the normal eye closure degree and the minimum eye closure degree, different levels of adaptive thresholds of eye closure degree are divided according to a certain ratio, and the number of levels increases with the increase of the severity of eye closure. At the same time, use the adaptive thresholds of different levels of eye closure degree of the driver obtained this time to update and optimize the stored thresholds of different levels of eye closure degree of the corresponding driver, and use this as a temporary threshold when the adaptive thresholds of different levels of eye closure degree have not been calculated.

[0127] Calculate the arithmetic mean of the eye closure degree data within a time window of the last 2 minutes or within a data sample window of the last 2400 data samples. Further, when calculating the mean value, different weights can be assigned to the data at different time points to calculate the weighted mean value, and the maximum value of this mean value is recorded as the normal eye closure degree ECR of the driver. norm 。

[0128] In addition, store the eye closure degree data in a min-heap with a capacity of 100. When the data statistical time in the min-heap exceeds 3 minutes, calculate the arithmetic mean of the data in the heap, and use this as the minimum eye closure degree ECR of the driver. min . The minimum eye closure degree ECR of the driver min and the numerical range covered by the normal eye closure degree ECR of the driver norm are divided according to the ratio of R 1 / R 2 / R 3 / ... / R n to determine the eye closure degrees at each level.

[0129] Among them, the calculation formula for the adaptive threshold ECR of the eye closure degree at the m-th level is: tm

[0130] In some embodiments, for the recognition of eye closure behaviors at different levels, the real-time eye closure degree of the driver calculated based on different time windows or data quantity windows can be compared with the adaptive thresholds of the eye closure degrees at each level, so as to recognize eye closure behaviors at different levels. Specifically, for high-level eye closure behaviors, a shorter time window is allocated; for low-level eye closure behaviors, a longer time window is allocated, so as to ensure a quick response to high-level eye closure behaviors and reduce the misjudgment rate of low-level eye closure behaviors.

[0131] Exemplarily, the eye closure behaviors are divided into four levels, and the level number increases with the increase of the eye closure severity. For the eye closure behavior at the fourth level (the highest level), the system allocates a 1-second (or other time) time window to calculate the average value of the driver's eye closure degree in the most recent 1 second, and compares this average value with the adaptive threshold at the fourth level to identify whether there is an eye closure behavior at the fourth level. For the eye closure behavior at the third level, the time window is set to 1.5 seconds (or other time); for the second level, the time window is set to 2 seconds (or other time); for the first level, the time window is set to 2.5 seconds (or other time). The recognition process for each level follows the same mechanism as that of the fourth level, and by comparing the average value of the eye closure degree within their respective time windows with the adaptive threshold at the corresponding level, it is determined whether there is an eye closure behavior at the corresponding level.

[0132] 206. Determine the opening and closing fatigue behavior of the driver's mouth according to the mouth feature points.

[0133] ​In the embodiments of the present application, a set of mouth-related feature points is used to calculate the relationship between the mouth opening and closing feature value and the reference area feature value as the mouth opening and closing degree. By comparing the real-time mouth opening and closing degree of the driver with the mouth opening and closing degree threshold, yawning behavior is identified.

[0134] In some embodiments, among the mouth feature points, a plurality of mouth opening and closing feature points and a plurality of mouth reference feature points are determined; according to the plurality of mouth opening and closing feature points, the mouth opening and closing area is calculated; according to the plurality of mouth reference feature points, the mouth reference area is calculated; and according to the mouth opening and closing area and the mouth reference area, mouth opening and closing fatigue behavior is determined.

[0135] The selected reference area feature points do not change their relative positions during the mouth opening and closing process, thus providing a stable reference, and will undergo the same perspective projection deformation as the mouth-related feature points when the driver's head rotates. Therefore, by calculating the relationship (such as a proportional relationship) between the mouth opening and closing feature value and the reference area feature value as the mouth opening and closing degree, the deviation of the mouth opening and closing feature value introduced by the perspective projection deformation can be effectively eliminated. Specifically, the mouth opening and closing feature value and the reference area feature value can be quantified using area features or aspect ratio features.

[0136] When calculating the real-time mouth opening and closing degree of the driver using the mouth opening and closing degrees stored in chronological order, methods such as non-weighted or weighted moving average based on a time window or a data quantity window can be adopted. When the real-time mouth opening and closing degree of the driver is higher than the mouth opening and closing degree threshold and the duration exceeds a predetermined time threshold, it is determined as yawning behavior.

[0137] In some embodiments, a set of mouth-related feature points is selected, such as five feature points P1, P2, P3, P4, P5 from left to right on the upper lip and five feature points P6, P7, P8, P9, P10 from left to right on the lower lip. The above ten feature points are used to enclose the mouth area and calculate its area.

[0138] A set of reference area feature points is selected, such as five feature points P11, P12, P13, P14, P15 from left to right on the cheekbone and five feature points P16, P17, P18, P19, P20 from left to right on the chin, for enclosing the reference area and calculating its area.

[0139] The area calculation method is the same as the above-mentioned eye area calculation method. By calculating the ratio of the mouth opening and closing area to the reference area, it is used as the mouth opening and closing degree. Calculate the arithmetic mean of the mouth opening and closing degree within the time window of the last 1 s or the window of the last 20 data samples. Further, when calculating the mean value, different weights can be assigned to the data at different time points to calculate the weighted mean value as the real-time mouth opening and closing degree of the driver. When the real-time mouth opening and closing degree of the driver exceeds the mouth opening and closing degree threshold (such as 0.6 or other values) and remains for more than 2 s (or other values), it is determined as a yawning behavior.

[0140] 207. Determine the corresponding fatigue value according to the target fatigue behavior.

[0141] 208. Continuously monitor the target fatigue behavior and count the fatigue values within a preset duration to obtain the target fatigue data.

[0142] 209. Determine the fatigue level of the driver according to the target fatigue data.

[0143] In the embodiment of the present application, the driver fatigue value is updated according to the identified fatigue driving behavior, the driver fatigue level is divided according to the driver fatigue value, and different levels of warning actions are performed when detecting fatigue driving behavior and dangerous driving behavior according to the driver fatigue level.

[0144] In some embodiments, for different detected fatigue driving behaviors, different numerical accumulations are performed on the driver's fatigue value according to different severities. For the same type of fatigue driving behavior, the system will update the fatigue value according to the preset cooling time interval. In addition, when it is detected that the driver maintains a normal driving state for a period of time, the fatigue value will be appropriately reduced. Different driver fatigue levels are divided according to the driver fatigue value.

[0145] 210. Output an attention warning message according to the fatigue level and the object detection result.

[0146] In the embodiment of the present application, for the description of step 210, please refer to the detailed description of step 105 in the above embodiment, and the embodiment of the present application will not be repeated.

[0147] 211. When it is detected that the driving speed is less than the functional vehicle stop speed threshold, stop acquiring the image to be detected.

[0148] The embodiment of the present application provides an attention detection method, which determines whether a driver is fatigued by various feature points recognized in the facial image. In addition, through the preset object detection in the image, it can be determined whether the driver is engaged in dangerous driving, thereby realizing the attention detection of the driver. This can effectively improve the accuracy and sensitivity of attention detection, and compare the real-time state with the standard state, making this method applicable to a large number of scenarios.

[0149] In some embodiments, during the driving process of the driver, the driving state parameters can be continuously recorded, including but not limited to information such as cumulative driving time, driver fatigue value, driver fatigue level, eye closure degree, mouth closure degree, real-time head three-dimensional pose, and current driving state. When a fatigued driving behavior or a dangerous driving behavior is detected, image capture can be performed according to a preset cooling time interval, and the captured images can be stored for subsequent analysis. These images record the state of the driver at the moment when the fatigued or dangerous driving behavior occurs.

[0150] In some embodiments, a display interface can also be set in the vehicle. The change curve of the driver's fatigue value during the driving process is displayed on this interface in the form of a time axis, and the state information of the driver at that time is displayed according to the selected time point. At the time point when a fatigued or dangerous driving behavior is detected, it can be significantly marked on the time axis, and an option is provided to display the captured image at the corresponding moment.

[0151] In some embodiments, the above embodiments can be implemented by a vision-based driver attention monitoring system. The vision-based driver attention monitoring system can include:

[0152] Driver face feature point acquisition unit: used to collect the real-time image of the driver, input the real-time image into the face feature point detection neural network model after preprocessing such as size adjustment, aspect ratio maintenance, normalization, and color channel conversion, obtain the face feature points, and screen and integrate the set of feature points required for each function of the driver attention monitoring system for use by subsequent modules.

[0153] Driver head pose recognition unit: using the set of feature points related to the three-dimensional head pose, calculating the three-dimensional head pose through the spatial angle algorithm, and identifying behaviors such as long-term low head-up, left and right head turns, and left and right head tilts according to the comparison between the real-time three-dimensional head pose and the statistically normal driving three-dimensional head pose.

[0154] Driver eye closure recognition unit: using the set of feature points related to the eyes, calculating the relationship between the eye closure feature value and the reference area feature value as the eye closure degree, and identifying different levels of eye closure behaviors by comparing the real-time eye closure degree of the driver with the adaptive thresholds of each level of eye closure degree.

[0155] Driver yawning recognition unit: Using the set of mouth-related feature points, calculate the relationship between the mouth opening and closing feature value and the reference area feature value as the mouth opening degree. By comparing the real-time mouth opening degree of the driver with the mouth opening degree threshold, recognize the yawning behavior.

[0156] Driver dangerous driving behavior recognition unit: Input the preprocessed real-time image of the driver into the object detection neural network model to obtain the object detection result. By screening and discriminating the information such as object category, position, size, and confidence in the object detection result, recognize dangerous driving behaviors such as making a phone call and smoking.

[0157] Driver distracted attention warning unit: Update the driver's fatigue value according to the recognized fatigue driving behavior, divide the driver's fatigue level according to the driver's fatigue value, and perform different degrees of warning actions when detecting fatigue driving behavior and dangerous driving behavior according to the driver's fatigue level.

[0158] Driver status recording and display unit: Record the driver's driving status, and statistically analyze and display the driver's status.

[0159] As Figure 3 shown, an embodiment of the present application provides an attention detection device, which may include:

[0160] An acquisition module 301, configured to acquire an image to be detected, where the image to be detected includes a face image of a driver;

[0161] A processing module 302, configured to extract features from the image to be detected to obtain face feature points included in the image to be detected, where the face feature points at least include: facial feature points, eye feature points, and mouth feature points;

[0162] The processing module 302 is further configured to determine the target fatigue behavior of the driver according to the face feature points;

[0163] The processing module 302 is further configured to determine the fatigue level of the driver according to the target fatigue behavior;

[0164] An output module 303, configured to output fatigue warning information corresponding to the fatigue level according to the fatigue level.

[0165] In some embodiments, the processing module 302 is specifically configured to determine the head deflection fatigue behavior of the driver according to the facial feature points;

[0166] The processing module 302 is specifically configured to determine the eye closure fatigue behavior of the driver according to the eye feature points;

[0167] The processing module 302 is specifically configured to determine the fatigue behavior of the driver's mouth opening and closing according to the mouth feature points.

[0168] In some embodiments, the processing module 302 is specifically configured to determine the target lateral feature points and the target longitudinal feature points from the facial feature points;

[0169] The processing module 302 is specifically configured to calculate the real-time three-dimensional pose of the head according to the target lateral feature points and the target longitudinal feature points;

[0170] The processing module 302 is specifically configured to compare the real-time rotation angle with the pre-stored standard three-dimensional pose to obtain the fatigue behavior of the head deflection.

[0171] In some embodiments, the processing module 302 is specifically configured to determine a plurality of eye closing feature points and a plurality of eye reference feature points among the eye feature points;

[0172] The processing module 302 is specifically configured to calculate the eye closing area according to the plurality of eye closing feature points;

[0173] The processing module 302 is specifically configured to calculate the eye reference area according to the plurality of eye reference feature points;

[0174] The processing module 302 is specifically configured to determine the eye closing fatigue behavior according to the eye closing area and the eye reference area.

[0175] In some embodiments, the processing module 302 is specifically configured to determine a plurality of mouth opening and closing feature points and a plurality of mouth reference feature points among the mouth feature points;

[0176] The processing module 302 is specifically configured to calculate the mouth opening and closing area according to the plurality of mouth opening and closing feature points;

[0177] The processing module 302 is specifically configured to calculate the mouth reference area according to the plurality of mouth reference feature points;

[0178] The processing module 302 is specifically configured to determine the mouth opening and closing fatigue behavior according to the mouth opening and closing area and the mouth reference area.

[0179] In some embodiments, the processing module 302 is specifically configured to determine the corresponding fatigue value according to the target fatigue behavior;

[0180] The processing module 302 is specifically configured to continuously monitor the target fatigue behavior and statistically analyze the fatigue values within a preset duration to obtain the target fatigue data;

[0181] The processing module 302 is specifically configured to determine the fatigue degree of the driver according to the target fatigue data.

[0182] In some embodiments, the acquisition module 301 is further configured to detect the driving speed of the vehicle in real time;

[0183] The acquisition module 301 is further configured to start acquiring the image to be detected when it is detected that the driving speed is greater than the function activation vehicle speed threshold;

[0184] The acquisition module 301 is further configured to stop acquiring the image to be detected when it is detected that the driving speed is less than the function stop vehicle speed threshold.

[0185] In the embodiments of the present application, each module can implement the attention detection method provided in the above method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0186] As Figure 4 shown, the embodiments of the present application further provide an electronic device, which may include:

[0187] A memory 401 storing executable program code;

[0188] A processor 402 coupled to the memory 401;

[0189] Wherein, the processor 402 calls the executable program code stored in the memory 401 and executes the attention detection method executed by the electronic device in the above method embodiments.

[0190] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the attention detection method in the above method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0191] The embodiments of the present application further provide a computer program product, which stores a computer program. When the computer program is executed by a processor, it implements each process of the attention detection method in the above method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0192] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0193] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0194] In the present application, the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0195] In the present application, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0196] In this application, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the computer-readable medium includes permanent and non-permanent, removable and non-removable storage media. The storage medium can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (Parallel Random Access Memory, PRAM), static random access memory (Static Random Access Memory, SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), programmable read-only memory (Programmable Read-only Memory, PROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), other types of random access memory (Random Access Memory, RAM), read-only memory (Read-Only Memory, ROM), one-time programmable read-only memory (One-time Programmable Read-Only Memory, OTPROM), electrically-erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), flash memory or other memory technologies, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0197] It should be noted that, in this text, relative terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the element.

[0198] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application. The above-mentioned multiple embodiments are not necessarily multiple independent embodiments. Dividing them into multiple embodiments is only used to highlight the different technical features in different embodiments. Those skilled in the art should be aware that the above-mentioned multiple embodiments can also be combined arbitrarily.

[0199] In various embodiments of the present application, it should be understood that the magnitude of the serial numbers of the above processes does not necessarily mean the inevitable sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0200] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0201] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0202] When the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute some or all of the steps of the above methods in the various embodiments of this application.

[0203] The above are only the specific implementation manners of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting attention, characterized in that: The method comprises: Acquire an image to be detected, wherein the image to be detected includes a face image of a driver; Performing feature extraction and object detection on the image to be detected to obtain facial feature points and object detection results included in the image to be detected, wherein the facial feature points at least include: facial feature points, eye feature points, and mouth feature points; Determining target fatigue behavior of the driver according to the facial feature points; determining the driver's fatigue level according to the target fatigue behavior; Attention warning information is output according to the fatigue level and the object detection result, where the attention warning information is used to indicate the driver's attention distraction and corresponding warning actions.

2. The method according to claim 1, characterized in that: The step of determining the target fatigue behavior of the driver according to the facial feature points includes: determining the driver's head deflection fatigue behavior according to the facial feature points; determining the driver's eye closure fatigue behavior according to the eye feature points; The driver's mouth opening and closing fatigue behavior is determined according to the mouth feature points.

3. The method according to claim 2, characterized in that The step of determining the driver's head deflection fatigue behavior according to the facial feature points includes: Determine a target transverse feature point and a target longitudinal feature point from the facial feature points; Calculating the real-time three-dimensional posture of the head according to the target transverse feature points and the target longitudinal feature points; The head deflection fatigue behavior is obtained by comparing the real-time rotation angle with a pre-stored standard three-dimensional posture.

4. The method according to claim 2, characterized in that: The step of determining the driver's eye closure fatigue behavior according to the eye feature points includes: Among the eye feature points, determining a plurality of eye closure feature points and a plurality of eye reference feature points; Calculating the eye closure area according to the plurality of eye closure feature points; Calculating an eye reference area according to the plurality of eye reference feature points; The eye closure fatigue behavior is determined based on the eye closure area and the eye reference area.

5. The method according to claim 2, characterized in that: The determining the driver's mouth opening and closing fatigue behavior according to the mouth feature points includes: Among the mouth feature points, determining a plurality of mouth opening and closing feature points and a plurality of mouth reference feature points; Calculating the mouth opening and closing area according to the plurality of mouth opening and closing feature points; Calculating a mouth reference area according to the plurality of mouth reference feature points; The mouth opening and closing fatigue behavior is determined according to the mouth opening and closing area and the mouth reference area.

6. The method according to any one of claims 1 to 5, characterized in that: The step of determining the driver's fatigue level according to the target fatigue behavior includes: Determining a corresponding fatigue value according to the target fatigue behavior; Continuously monitoring the target fatigue behavior and collecting statistics on fatigue values ​​within a preset time period to obtain target fatigue data; The driver's fatigue level is determined according to the target fatigue data.

7. The method according to claim 1, characterized in that The method further comprises: Real-time detection of vehicle speed; When it is detected that the driving speed is greater than the function activation speed threshold, starting to acquire the image to be detected; When it is detected that the driving speed is less than the function deactivation speed threshold, the acquisition of the image to be detected is stopped.

8. An attention detection device, characterized in that: The attention detection device comprises: An acquisition module, used for acquiring an image to be detected, wherein the image to be detected includes a face image of the driver; A processing module, used for extracting features from the image to be detected to obtain facial feature points included in the image to be detected, wherein the facial feature points include at least facial feature points, eye feature points, and mouth feature points; The processing module is further used to determine the target fatigue behavior of the driver according to the facial feature points; The processing module is further used to determine the driver's fatigue level according to the target fatigue behavior; The output module is used to output fatigue warning information corresponding to the fatigue level according to the fatigue level.

9. An electronic device, characterized in that: The electronic device comprises: A memory storing executable program code; and a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the attention detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the attention detection method according to any one of claims 1 to 7 is implemented.