A method and device for detecting fatigue driving
By performing distortion correction on the face images collected by the camera and using lightweight convolutional neural network for fatigue driving detection, the problem of low detection accuracy in the existing methods is solved, and higher detection accuracy and timely early warning are achieved.
Patent Information
- Application Number
- CN202410612672.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-05-16
AI Technical Summary
The existing fatigue driving detection methods are easily affected by light, weather and camera placement, resulting in poor image effects and low detection accuracy.
By acquiring the face images collected by the camera, performing distortion correction, and using the corrected images for fatigue driving detection, a lightweight convolutional neural network is designed for head posture estimation.
It improves the accuracy of fatigue driving detection, can more accurately capture the driver's facial expressions and head movements, and timely sends out warning messages to ensure driving safety.
Smart Images

Figure CN118552940B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of vehicle driving. Specifically, the present invention relates to a fatigue driving detection method and device. Background Art
[0002] Traffic accidents are a common type of accident, and the casualties and property losses caused by these accidents cannot be ignored. The occurrence of accidents often stems from the uneven quality of drivers and the frequent occurrence of various violations. The most common behaviors include speeding, running red lights, fatigue driving, etc. Fatigue driving, as one of the main causes of traffic accidents, poses a threat to the lives and safety of drivers, passengers, and people on and around the road.
[0003] In order to reduce the potential safety hazards caused by fatigue driving, at present, there are already some fatigue driving monitoring systems based on cameras and sensors to detect the behaviors of drivers. However, the existing detection methods are easily affected by light, weather, and the placement position of the camera, resulting in poor image quality of the captured images, and thus leading to low detection accuracy for fatigue driving. Summary of the Invention
[0004] The embodiments of the present invention provide a fatigue driving detection method and device to at least solve the problem of low detection accuracy for fatigue driving caused by poor image quality of the captured images.
[0005] According to an embodiment of the present invention, a fatigue driving detection method is provided, including:
[0006] Obtaining a face image collected by a camera;
[0007] When the face image is distorted, correcting the face image;
[0008] Performing fatigue driving detection based on the corrected face image;
[0009] When the corrected face image meets the conditions of fatigue driving, outputting a warning message.
[0010] According to an embodiment of the present invention, a fatigue driving detection device is provided, including:
[0011] An obtaining module, configured to obtain a face image collected by a camera;
[0012] A correction module, configured to correct the face image when the face image is distorted;
[0013] A detection module, configured to perform fatigue driving detection based on the corrected face image;
[0014] An output module, configured to output a warning message when the corrected face image meets the conditions of fatigue driving.
[0015] According to another embodiment of the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above-described fatigue driving detection methods when running.
[0016] According to another embodiment of the present invention, there is also provided an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-described fatigue driving detection methods.
[0017] Through the present invention, in the case where the captured image is distorted, the face image is corrected for distortion, and the corrected face image is used for fatigue driving detection, which can improve the accuracy of fatigue driving detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a hardware structure diagram of a mobile terminal of a fatigue driving detection method according to an embodiment of the present invention;
[0019] Figure 2 is a flowchart of a fatigue driving detection method according to an embodiment of the present invention;
[0020] Figure 3 is a schematic diagram of coordinate transformation for image correction according to an embodiment of the present invention;
[0021] Figure 4 is a lightweight architecture model diagram according to an embodiment of the present invention;
[0022] Figure 5 is a structure diagram of a building block according to an embodiment of the present invention;
[0023] Figure 6 is a schematic diagram of a rotation angle according to an embodiment of the present invention;
[0024] Figure 7 is an eye extraction feature map according to an embodiment of the present invention;
[0025] Figure 8 is a mouth extraction feature map according to an embodiment of the present invention;
[0026] Figure 9 is a structure diagram of a fatigue driving detection system according to an embodiment of the present invention;
[0027] Figure 10 is a call relationship structure diagram of a fatigue driving detection system according to an embodiment of the present invention;
[0028] Figure 11 It is a schematic diagram of the fatigue driving detection interface according to an embodiment of the present invention;
[0029] Figure 12 It is a structural diagram of a fatigue driving detection device according to an embodiment of the present invention;
[0030] Figure 13 It is a structural diagram of a fatigue driving detection device according to an embodiment of the present invention. Specific embodiments
[0031] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0032] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0033] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 It is a hardware structure block diagram of a mobile terminal of a fatigue driving detection method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more (only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only schematic, and it does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0034] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to a fatigue driving detection method in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0036] This embodiment provides a fatigue driving detection method, wherein, Figure 2 is a flowchart of a fatigue driving detection method according to an embodiment of the present invention, as Figure 2 shown, the process includes the following steps:
[0037] Step S201, obtain a face image collected by a camera.
[0038] In order to monitor the facial expression and head posture of the driver, obtain a face image collected by the camera in real time. Since the position of the camera moves during the vehicle driving process, and due to the influence of factors such as environmental changes, the face image may be distorted.
[0039] Step S202, correct the face image when the face image is distorted.
[0040] When the face image is distorted, the error of the recognition result obtained by performing facial expression recognition based on the obtained face image is relatively large. Therefore, first correct the face of the face image so that the face of the corrected face image is not distorted.
[0041] Optionally, the correcting the face image when the face image is distorted includes:
[0042] Obtain the center of the face image;
[0043] Perform face correction on the face image based on the center of the face image to obtain the corrected face image.
[0044] In this embodiment, the distorted face image is corrected to obtain a corrected face image.
[0045] As Figure 3 shown, Figure 3 is the face image of the driver captured by the camera. Point O is the center of the camera optical axis. C0 represents the real camera system. x0 and z0 respectively represent the x-axis and z-axis of the camera system C0. The y-axis (not shown) is perpendicular to the x-axis and z-axis; C R represents a virtual camera system with the head center aligned with the virtual optical axis. x R and z R respectively represent the x-axis and z-axis of the camera system C R . The y-axis (not shown) is perpendicular to the x-axis and z-axis. When the face image is distorted, it is difficult to accurately obtain the real rotation angle and facial condition of the face image. Therefore, the facial image of the face image can be extracted, and the center of the facial boundary is used as the center of the face image. The face image is corrected with the center of the face image, and the face image is corrected to the image in the normal shooting state. The center of the corrected face image is aligned with the optical axis of the virtual optical system.
[0046] The corrected image obtained by the above method can more truly reflect the state of the driver, can more clearly obtain the nodding characteristics of the driver, such as the head rotation angle, and can improve the accuracy of fatigue driving detection.
[0047] Optionally, the fatigue driving detection based on the corrected face image includes:
[0048] Obtain the first rotation angle of the corrected face image in the first coordinate system, where the center of the corrected face image is located on the extension line of the coordinate axis of the first coordinate system;
[0049] Determine the second rotation angle of the face image in the second coordinate system based on the first rotation angle. The second coordinate system is the coordinate system corresponding to the optical axis of the camera;
[0050] Perform fatigue driving detection on the face image based on the second rotation angle.
[0051] In this embodiment, after the face image is corrected, the center of the corrected face image is aligned with the optical axis of the virtual optical system. When multiple frames of images are acquired, the first rotation angle of the face image in the first coordinate system can be obtained according to the differences between multiple frames of face images, and the acquired images can also be compared with the face image of the driver in the normal state to determine the first rotation angle. Specifically, the corrected face image can be input into a convolutional neural network to obtain the rotation angle of the corrected face image in the virtual camera system, that is, the first rotation angle.
[0052] To obtain a more realistic rotation angle of the face image, the first rotation angle in the first coordinate system corresponding to the virtual camera system is converted to the coordinate system of the real camera system, that is, the second rotation angle corresponding to the second coordinate system. The second rotation angle can fit the actual rotation angle, improving the accuracy of nod detection. Among them, the first coordinate system can be Figure 3 the real camera system C0 shown in Figure 3 and the second coordinate system can be R the virtual camera system C shown in R . After the face image is corrected, the center of the face image is located on the extension line of the coordinate axis z
[0053] of the virtual camera system. For example, the facial region of the corrected face image is extracted and then used as the input of the convolutional neural network to estimate the head pose of the face image in the virtual camera system C R . The output of the convolutional neural network is transformed back to the camera coordinate system C0 as the final head pose estimate. To evaluate the head pose estimation algorithm, a benchmark data set containing real head poses is required. These real head poses are measured or estimated relative to the real camera system C0.
[0054] To train the convolutional neural network to estimate the head pose in the virtual camera system C R , the real head poses provided by the data set must be appropriately transformed. To evaluate the accuracy of the head pose estimation algorithm, the head pose estimates obtained from the convolutional neural network are converted back to the camera system C0 for accuracy evaluation. These head pose estimates are usually represented in the form of Euler angles and can be converted into the form of a rotation matrix as shown in the following formula:
[0055]
[0056] where α, β, and γ are the pitch angle, yaw angle, and roll angle respectively. The benchmark true head pose provided by the data set in the camera system C0 can be converted into the same matrix form and denoted as M0. Similarly, M R represents the rotation matrix of the head pose in the virtual camera system C R . CR The Euler angles in it can be calculated by the following formula:
[0057]
[0058] where α R , β R and γ R respectively represent the pitch, yaw and roll in C R . M Rmn represents the element in the m-th row and n-th column of the matrix M R . Through the above coordinate transformation, the rotation angles in the virtual camera system are converted to the rotation angles in the real camera system, so as to obtain the rotation angles of the face image in each direction in the real camera system.
[0059] To achieve the above effect, a lightweight convolutional neural network is constructed to accurately estimate the head pose based on the corrected image. To achieve this goal, group convolution and depthwise separable convolution technologies are used, and its structure is as Figure 4 shown. This network model is relatively small and is designed specifically for this application.
[0060] Depthwise separable convolution is a method of decomposing convolution, which consists of two parts: depthwise convolution (DW) and pointwise convolution (PW). By repeatedly applying group convolution and depthwise separable convolution technologies in the building blocks, a lightweight network can be effectively constructed, which is widely used in many CNN-based recognition algorithms. MobileNetV2 adopts a strategy of first introducing an expansion layer (pointwise convolution), and then using depthwise separable convolution to significantly improve the performance. In the building block of MobileNetV2, the expansion layer is used to map the input features to a high dimension, then filtered by depthwise convolution, and finally the features are mapped to a low dimension by pointwise convolution.
[0061] Based on MobileNetV1 and MobileNetV2, the present invention adopts a lightweight network structure for head pose estimation. This network consists of a series of building blocks, where the building blocks are constructed using pointwise group convolution (PWG) and depthwise separable convolution technologies, as Figure 5 shown.
[0062] Pointwise group convolution is a special type of group convolution that can be regarded as an implementation of group convolution using a kernel with a spatial size of 1×1. Compared with standard pointwise convolution, pointwise group convolution uses fewer parameters, which helps to reduce the scale of the network. We use pointwise group convolution to map the input features to a high dimension. The high-dimensional features are then further processed by depthwise separable convolution, where the expansion factor is set to 2 and the number of groups is set to 4 to achieve a balance between performance and network scale. In the building blocks of the network, a batch normalization operation and an activation layer are included after each convolutional layer to improve the performance of the model.
[0063] In the proposed network, the following two types of activation functions are used: ReLU and h-swish. Previous work on MobileNetV3 has shown that the introduction of h-swish can improve network performance at the cost of more computational effort compared with ReLU. These two activation functions are used in the network of this system to achieve a better balance between model accuracy and computational cost.
[0064] In this embodiment, fatigue driving detection is performed on the corrected face image, and the detection process is shown in Figure 6 as follows. The face image in the camera system C0 is obtained, and the face image is detected. In the case of distortion, the head center is obtained and corrected with the head center as the center. The center of the corrected face image is aligned with the virtual optical axis of the virtual camera. The corrected face image is input into the estimation network, that is, the convolutional neural network. The rotation angle of the face image in the virtual camera system is obtained through the neural network, and coordinate transformation is performed to obtain the rotation angle of the face image in the real camera system. Fatigue driving detection is performed based on this rotation angle.
[0065] By using image correction to reduce the influence of perspective distortion and designing a lightweight convolutional neural network for head pose estimation. When the head center of the input image of the neural network is not aligned with the camera coordinate system, mathematical correction is performed to align the head center with the optical axis of the virtual camera coordinate system.
[0066] Step S203, perform fatigue driving detection based on the corrected face image.
[0067] In this step, based on the corrected face image, the facial expression and head movement of the driver are detected to determine whether the driver is in a fatigue driving state. For example, it is detected whether the driver blinks frequently, nods frequently, leaves the driver's seat, etc. Performing fatigue driving detection based on the corrected face image can improve the detection accuracy.
[0068] Optionally, the fatigue driving detection of the face image based on the second rotation angle includes:
[0069] Determine the nodding operation of the face image based on the second rotation angle, where the nodding operation includes at least one of a vertical nodding operation, a horizontal nodding operation, and an inclined nodding operation;
[0070] When the nodding amplitude and nodding frequency of the nodding operation determined based on the second rotation angle are within a first preset range, determine that the corrected face image meets the conditions of fatigue driving.
[0071] Convert the first rotation angle to the corresponding second rotation angle in the real camera system, and determine whether the face image has a nodding operation based on the second rotation angle. For example, determine the nodding amplitude based on the rotation angle. When the nodding amplitude is within a preset amplitude range and the nodding frequency is within a preset frequency range, the above factors can also be combined to determine the nodding operation.
[0072] The main parameter for judging drowsiness is the Euler angle. The schematic diagram of the rotation angle is as Figure 7 shown. It can be obtained from the figure that when the driver is dozing off, obviously the head will make movements similar to nodding and tilting. According to the head postures shown by ordinary people when dozing off, there will obviously be few movements in Yaw, and the main focus is on the behaviors of Pitch and Roll. Therefore, set the parameter ratio threshold. For example, within a time period of 10s, when the nodding angle in Pitch and Roll is greater than 20° and the duration ratio exceeds 0.3, it is determined that the driver is in a dozing state and a warning is issued.
[0073] Among them, the vertical nodding can be the nodding in the picth direction as shown in Figure 7 the figure, the horizontal nodding can be the nodding in the roll direction as shown in the figure, and the inclined nodding can be the nodding between the picth and roll directions. In addition, it can also include irregular nodding, such as inclined nodding, with irregular nodding frequency and amplitude, but nodding that conforms to the characteristics of drowsy nodding.
[0074] The nodding characteristics of the driver can be obtained more clearly through the corrected face image, improving the accuracy of fatigue driving detection.
[0075] Optionally, the fatigue driving detection based on the corrected face image includes:
[0076] Detect the characteristics of the blinking operation of the corrected face image, where the characteristics of the blinking operation include at least one of the blinking frequency and speed, the blinking duration, and the irregular blinking operation;
[0077] When the characteristics of the blinking operation are within a second preset range, determine that the corrected face image meets the conditions of fatigue driving.
[0078] Since the corrected face image can capture the features of the face image more accurately, such as eye features and mouth features.
[0079] To detect fatigue driving features, this embodiment detects the blinking operation of the face image, including any one or more of the blinking frequency, blinking speed, blinking duration, irregular blinking operation, etc.
[0080] (a) Blinking frequency and speed analysis: The blinking frequency, for example, the number of blinks per minute, and the blinking speed reflect the time required for the blinking process. The normal blinking speed is very fast, while the blinking speed in the drowsy state is slower. Normal blinking and blinking caused by fatigue driving are different. For example, frequent and rapid blinking may indicate fatigue, while normal blinking is more stable and regular.
[0081] (b) Blinking duration: A longer period of eye closure may indicate that the driver's eyes are fatigued or that they are trying to resist falling asleep. This can be used to more accurately identify potential fatigue driving situations.
[0082] (c) Irregular blinking behavior: Such as slow blinking, intermittent blinking, or asymmetric blinking. These abnormal blinking behaviors may indicate the driver's fatigue or distraction, as they are usually different from blinking in the normal alert state.
[0083] The blinking function module usually provides data visualization, allowing the driver and the monitoring personnel to view the real-time and historical records of the blinking data. This helps the driver understand their alert state and can take actions when needed.
[0084] By calculating the eye aspect ratio (EAR), it can be determined whether there is fatigue driving. When the human eye is open, the EAR fluctuates around a certain value. When the human eye is closed, the EAR drops rapidly and theoretically approaches zero. In this system, a threshold is set for the EAR. If the value is lower than the set threshold, it is determined that the eyes are in the closed state. To detect the number of blinks, the consecutive frame numbers of the same blink need to be set. The blinking speed is relatively fast, and generally, the blinking action is completed in 1 - 3 frames. As Figure 8 shown, the calculation formula of the eye aspect ratio (EAR) is shown in Equation (3):
[0085]
[0086] By detecting the above blinking features, the fatigue driving state of the driver can be detected, and corresponding measures can be taken.
[0087] Optionally, the fatigue driving detection based on the corrected face image includes:
[0088] Detect the characteristics of the yawning operation of the corrected face image, where the characteristics of the yawning operation include at least one of the mouth opening degree, mouth opening time, mouth opening frequency, and mouth opening speed;
[0089] When the characteristics of the yawning operation are within a third preset range, determine that the corrected face image meets the conditions of fatigue driving.
[0090] Yawning is a common sign of fatigue, usually indicating a decrease in the driver's alertness and the need to rest or take other measures to maintain safe driving. Yawn detection includes the number of yawns, as well as the speed and frequency of yawns. The speed of opening the mouth can reflect the time required for a yawn, and the frequency of yawns reflects the frequency. Frequent and rapid yawns may indicate a sharp decline in the driver's alertness, which is a potential danger sign.
[0091] The present invention uses a double-threshold method for yawn detection, that is, detecting the inner contour: combining the mouth opening degree and the mouth opening time. In this system, Yawn is the number of frames that meet the yawn condition, N is the total number of frames within 1 minute, and the threshold is set to 10%. When Freq>10%, it is considered that a deep yawn or at least two consecutive shallow yawns have occurred, and a fatigue reminder is given at this time.
[0092] Use the Dlib model for face recognition, detect 68 feature points of the face, obtain the index of the facial landmarks of the mouth, and perform grayscale processing on the video stream through OpenCV to detect the position information of the human mouth.
[0093] From Figure 9 it can be seen that yawning can be calculated using the coordinate points at the mouth. 51, 59, 53, and 57 are the vertical coordinates, represented by y; 49 and 55 are the horizontal coordinates, represented by x, to calculate the mouth opening degree. The Euclidean distance formula for calculating yawns is shown in Equation (4). By calculating the distance to determine whether the mouth is open and the mouth opening time, it can be determined whether a person is yawning. At the same time, this threshold should be reasonable and needs to be distinguished from normal speaking or humming.
[0094]
[0095] Determining whether a yawn occurs based on the corrected face image can improve the accuracy of detection.
[0096] Step S204: When the corrected face image meets the conditions of fatigue driving, output a warning message.
[0097] When it is determined through the face image that the driver meets the conditions of fatigue driving, a warning prompt can be output, or other measures can be taken to prevent the vehicle from being in danger.
[0098] In the embodiment of the present invention, the distortion correction of the face image is performed, and the corrected face image is used for fatigue driving detection, which can improve the accuracy of fatigue driving detection.
[0099] To facilitate a clearer understanding of this embodiment, specific embodiments are given as examples below. The fatigue driving detection system includes five major modules, and the system schematic diagram is as follows Figure 10 As shown, it reflects the hierarchical structure among the various modules of the system. The system captures images from the camera, monitors the fatigue state, and then determines whether fatigue driving has occurred through the monitoring module. If fatigue driving occurs, the result is output through the alarm module.
[0100] Figure 11 It reflects the call relationship among the various modules. From the sensor layer to the data acquisition and preprocessing layer: The sensor layer captures data from the camera and then transfers these raw data to the data acquisition and preprocessing layer. At this layer, the data is preprocessed, including denoising, calibration, and time synchronization, to ensure the quality and consistency of the data. In the case where the captured image is distorted, the face image is corrected for distortion and then passed to the detection module.
[0101] From the data acquisition and preprocessing layer to the monitoring module: The processed data is transferred from the data acquisition and preprocessing layer to the monitoring module. These modules are responsible for monitoring the driver's facial expressions and head postures. Different functional monitoring modules, such as the blink function module, yawn function module, and drowsy nod function module, first extract the corresponding features and further analyze the feature data to monitor for signs of fatigue. Among them, the blink function module is used to detect the driver's blinking operation, the yawn function module is used to detect the driver's yawning operation, and the drowsy nod function module is used to detect the driver's nodding operation. The implementation of the above functional modules is based on the corrected face image.
[0102] From the monitoring function module to the data integration and analysis layer: This part transfers the analysis results of the monitoring module to the data integration and analysis layer. This layer integrates the results of different modules to comprehensively evaluate the driver's state.
[0103] From the data integration and analysis layer to the alarm trigger layer: The data integration and analysis layer determines whether there are potential signs of fatigue. If so, it will trigger the corresponding alarm. These alarms can include sound alarms, vibration alarms, or visual cues.
[0104] From the alarm trigger layer to the user interface layer: The alarm trigger layer conveys the alarm to the driver through the user interface layer.
[0105] The alarm function module is used to issue alarms to remind the driver to take measures to maintain vigilance and safety. The alarm function module is responsible for triggering alarms based on the analysis results of other modules (such as the blinking function module, yawning function module, drowsy nodding function module). If the system detects potential signs of fatigue, the alarm function module will initiate the alarm process.
[0106] The alarm function module usually supports multiple different types of alarms, including sound alarms, vibration alarms, or a combination. Different types of alarms can better adapt to the driver's perception methods. Usually, the alarm will last for a period of time to ensure that the driver is effectively reminded. Persistence can also automatically extend the alarm time when the signs of fatigue persist. In addition to sound and vibration, the alarm function module can also use other methods, such as displaying text or icons on the driver's dashboard.
[0107] To facilitate the detection of the driver's fatigue driving, the human-machine interface of the driving detection system can be designed, such as Figure 12 shown.
[0108] Among them, the operation console can select the video source. The video source module is the entry of the system and is responsible for real-time monitoring of the driver. This module usually includes a camera or other video capture devices to obtain the driver's facial expressions, head postures, and behaviors. Through the video source, the system obtains real-time visual information for subsequent analysis. The fatigue monitoring part is the core part of the system and can be selectively monitored. By analyzing data such as the driver's yawning, blinking, and nodding, it is judged whether there is a fatigue driving state. Usually, if this state persists within a certain period of time (such as 3 seconds), the system will determine it as a fatigue driving state. The task of out-of-position monitoring is to monitor whether the driver has left the driver's seat. Out-of-position monitoring helps to warn the driver in a timely manner to avoid potential dangers caused by distraction.
[0109] The status output bar is the result presentation part of the system, and it is responsible for real-time output of the system's determination results on the driver's status. This module usually presents the information to the driver in a visual way to remind them whether their status is normal. This can be achieved through alarm sounds, vibrations, graphical interfaces, or other means to ensure that the driver can quickly take measures to maintain vigilance.
[0110] In the embodiments of the present invention, by introducing an image correction algorithm to correct face images, the system can effectively reduce the negative impact of perspective distortion on face images, thereby improving the accuracy of head pose estimation. Moreover, by adopting a lightweight convolutional neural network based on depthwise separable convolution, the system has a small model size and fast processing speed while maintaining high accuracy. The network model is small, reducing memory and storage requirements, providing efficient application possibilities for platforms with limited resources. The lightweight convolutional neural network has relatively few parameters and computational complexity, so it can perform excellently in terms of real-time performance, achieve fast response, and be able to issue fatigue warnings in a timely manner.
[0111] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0112] In this embodiment, a fatigue driving detection device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0113] Figure 13 is a structural diagram of a fatigue driving detection device according to an embodiment of the present invention, as Figure 13 shown, the device includes:
[0114] An acquisition module 1301, configured to acquire face images collected by a camera;
[0115] A correction module 1302, configured to correct the face image when the face image is distorted;
[0116] A detection module 1303, configured to perform fatigue driving detection based on the corrected face image;
[0117] An output module 1304, configured to output a warning message when the corrected face image meets the conditions of fatigue driving.
[0118] Optionally, the correction module includes:
[0119] A first acquisition sub-module, configured to acquire the center of the face image;
[0120] A correction sub-module, configured to perform face correction on the face image based on the center of the face image to obtain the corrected face image.
[0121] Optionally, the detection module includes:
[0122] A second acquisition sub-module, configured to acquire a first rotation angle of the corrected face image in a first coordinate system, wherein the center of the corrected face image is located on the extension line of the coordinate axis of the first coordinate system;
[0123] A first determination sub-module, configured to determine a second rotation angle of the face image in a second coordinate system based on the first rotation angle, where the second coordinate system is a coordinate system corresponding to the optical axis of the camera;
[0124] A first detection sub-module, configured to perform fatigue driving detection on the face image based on the second rotation angle.
[0125] Optionally, the first detection sub-module includes:
[0126] A first determination unit, configured to determine a nodding operation of the face image based on the second rotation angle, where the nodding operation includes at least one of a vertical nodding operation, a horizontal nodding operation, and an inclined nodding operation;
[0127] A second determination unit, configured to determine that the corrected face image meets the condition of fatigue driving when the nodding amplitude and nodding frequency of the nodding operation determined based on the second rotation angle are within a first preset range.
[0128] Optionally, the detection module includes:
[0129] A second detection sub-module, configured to detect features of a blinking operation of the corrected face image, where the features of the blinking operation include at least one of a blinking frequency and speed, a blinking duration, and an irregular blinking operation;
[0130] A second determination sub-module, configured to determine that the corrected face image meets the condition of fatigue driving when the features of the blinking operation are within a second preset range.
[0131] Optionally, the fatigue driving detection based on the corrected face image includes:
[0132] A third detection sub-module, configured to detect features of a yawning operation of the corrected face image, where the features of the yawning operation include at least one of mouth opening degree, mouth opening time, mouth opening frequency, and mouth opening speed;
[0133] A third determination sub-module, configured to determine that the corrected face image meets the condition of fatigue driving when the features of the yawning operation are within a third preset range.
[0134] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: all the above modules are located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0135] In the embodiment of the present invention, by introducing an image correction algorithm to correct the face image, the system can effectively reduce the negative impact of perspective distortion on the face image, thereby improving the accuracy of head pose estimation. And by using a lightweight convolutional neural network based on depthwise separable convolution, the system has a small model size and fast processing speed while maintaining high accuracy. The network model is small, reducing memory and storage requirements, providing efficient application possibilities for platforms with limited resources. The lightweight convolutional neural network has relatively few parameters and computational complexity, so it can perform well in terms of real-time performance, achieve fast response, and be able to issue fatigue warnings in a timely manner.
[0136] The embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0137] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.
[0138] The embodiment of the present invention also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0139] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0140] For the specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and details thereof will not be repeated herein.
[0141] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A fatigue driving detection method, characterized in that, Including: Obtain a face image collected by a camera; When the face image is distorted, correct the face image; Perform fatigue driving detection based on the corrected face image; When the corrected face image meets the conditions of fatigue driving, output a warning message; Among them, the fatigue driving detection based on the corrected face image includes: Obtain a first rotation angle of the corrected face image in a first coordinate system, where the first coordinate system includes a real camera coordinate system; Determine a second rotation angle of the face image in a second coordinate system based on the first rotation angle. The second coordinate system is the coordinate system corresponding to the optical axis of the camera. Among them, the center of the corrected face image is located on the extension line of the coordinate axis of the second coordinate system; Among them, the facial area of the corrected face image is extracted and then used as the input of a convolutional neural network to estimate the head pose of the face image in the second coordinate system; The output of the convolutional neural network is transformed back to the first coordinate system as the final head pose estimate; The convolutional neural network consists of grouped convolution and depthwise separable convolution; Perform fatigue driving detection on the face image based on the second rotation angle.
2. The method according to claim 1, wherein The correcting the face image when the face image is distorted includes: Obtain the center of the face image; Perform facial correction on the face image based on the center of the face image to obtain the corrected face image.
3. The method according to claim 1, wherein The fatigue driving detection based on the second rotation angle of the face image includes: Determine the nodding operation of the face image based on the second rotation angle. The nodding operation includes at least one of a vertical nodding operation, a horizontal nodding operation, and an inclined nodding operation; When the nodding amplitude and nodding frequency of the nodding operation determined based on the second rotation angle are within a first preset range, determine that the corrected face image meets the conditions of fatigue driving.
4. The method according to claim 1, wherein The fatigue driving detection based on the corrected face image includes: Detect the characteristics of the blinking operation of the corrected face image. The characteristics of the blinking operation include at least one of blinking frequency and speed, blinking duration, and irregular blinking operation; When the characteristics of the blinking operation are within a second preset range, determine that the corrected face image meets the conditions of fatigue driving.
5. The method according to claim 1, wherein The fatigue driving detection based on the corrected face image includes: Detect the characteristics of the yawning operation of the corrected face image. The characteristics of the yawning operation include at least one of mouth opening degree, mouth opening time, mouth opening frequency, and mouth opening speed; When the characteristics of the yawning operation are within a third preset range, determine that the corrected face image meets the conditions of fatigue driving.
6. A fatigue driving detection device, characterized in that, Including: An acquisition module, configured to obtain a face image collected by a camera; A correction module, configured to correct the face image when the face image is distorted; A detection module, configured to perform fatigue driving detection based on the corrected face image; An output module, configured to output a warning message when the corrected face image meets the conditions of fatigue driving; Wherein, the detection module includes: A second acquisition sub-module, configured to acquire a first rotation angle of the corrected face image in a first coordinate system, where the first coordinate system includes a real camera coordinate system; A first determination sub-module, configured to determine a second rotation angle of the face image in a second coordinate system based on the first rotation angle, where the second coordinate system is a coordinate system corresponding to the optical axis of the camera, and the center of the corrected face image is located on the extension line of the coordinate axis of the second coordinate system; wherein, the facial region of the corrected face image is extracted and then used as the input of a convolutional neural network to estimate the head pose of the face image in the second coordinate system; the output of the convolutional neural network is transformed back to the first coordinate system as the final head pose estimate; the convolutional neural network consists of grouped convolution and depthwise separable convolution; A first detection sub-module, configured to perform fatigue driving detection on the face image based on the second rotation angle.
7. The device according to claim 6, characterized in that The correction module includes: A first acquisition sub-module, configured to acquire the center of the face image; A correction sub-module, configured to perform facial correction on the face image based on the center of the face image to obtain the corrected face image.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when running.
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for detecting fatigue driving
CN106203293A
Fatigue driving detection method and system based on facial feature fusion
CN110532887A
Face image pose estimation and correction method and system, medium and electronic equipment
CN113011401A