Fatigue driving detection method and system based on camera and heart rate monitoring

CN120145299APending Publication Date: 2025-06-13CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510211978.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing fatigue driving detection algorithm has the problem of poor real-time performance, and the reliance on fixed thresholds leads to a high misjudgment rate, so individual differences cannot be effectively considered.

Method used

A multi-feature fusion method based on camera and heart rate monitoring is adopted to weighted fusion through the driver's facial features, lane yaw rate and heart rate data detected by smart wearable devices, and the fatigue threshold is dynamically calculated to improve the real-time and accuracy of the detection.

Benefits of technology

It effectively reduces detection time, improves real-timeness, significantly improves the accuracy of driving fatigue detection, reduces the misjudgment rate, and can dynamically adjust the threshold according to individual characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145299A_ABST
    Figure CN120145299A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of fatigue driving detection, and discloses a fatigue driving detection method and system based on a camera and heart rate monitoring, and the method comprises the steps: collecting video data; performing feature extraction on the video frame in the cab to obtain driver facial features, obtaining individual difference indexes based on the driver facial features, calculating a fatigue state judgment value according to the individual difference indexes, and judging a driver facial fatigue state; performing feature extraction on the road video frame to obtain a lane line, calculating an included angle between the lane line and the lower edge of the image according to the slope of the lane line, and further calculating a yaw rate; and performing weighted fusion on the face fatigue state of the driver, the yaw rate and the detection result of the intelligent wearable device to obtain a comprehensive fatigue state judgment value, and judging the driving state of the driver according to the comprehensive fatigue state judgment value. According to the invention, the accuracy of driving fatigue detection is effectively improved while the detection time of fatigue driving is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fatigue driving detection, and particularly to a fatigue driving detection method and system based on a camera and heart rate monitoring. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] The rapid development of the transportation industry has brought great convenience to human life and promoted the sound and rapid sustainable development of the national economy. However, at the same time, the number of serious road traffic safety accidents has also increased synchronously, which is recognized as the greatest threat and killer to humans outside of war. Fatigue driving is one of the main reasons for frequent traffic accidents. During daily driving, fatigue driving often occurs. In many cases, drivers are not aware of their poor mental state and continue to drive the vehicle, resulting in car accidents. If a warning can be issued in time when the driver is fatigued to remind the driver to take appropriate measures, such accidents can be effectively avoided. Therefore, the effective detection of fatigue driving is of great significance for improving driving safety.

[0004] As a driving behavior with extremely high risk, fatigue driving requires the detection system to have good real-time performance. The detection system needs to continuously and quickly detect the fatigue characteristics of the driver in consecutive image frames, and has extremely high requirements for the stability of the algorithm. Existing fatigue driving detection algorithms have problems with poor real-time performance.

[0005] Currently, most fatigue driving detection algorithms extract fatigue characteristics based on the driver's image as the basis for fatigue determination. However, the facial features and driving habits of different drivers are different. Judging the fatigue state only based on the driver's features has certain limitations. And existing fatigue driving detection algorithms use fixed thresholds to distinguish the driver's facial features, ignoring the individual differences of drivers, which will obviously also lead to a high false positive rate. Summary of the Invention

[0006] To solve the above problems, the present invention proposes a fatigue driving detection method and system based on a camera and heart rate monitoring. It uses a method of multi-feature fusion to describe the fatigue state of the driver, calculates the fatigue threshold according to individual differences, and uses an improved tracking algorithm and detection algorithm for fusion, reducing the detection time, improving the real-time performance, and effectively improving the accuracy of driving fatigue detection.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In the first aspect, the present invention provides a fatigue driving detection method based on a camera and heart rate monitoring, including the following steps:

[0009] Collect the video data inside the cab and the road video data;

[0010] Extract features from the video frames inside the cab to obtain the driver's facial features, obtain the individual difference index based on the driver's facial features, calculate the fatigue state judgment value according to the individual difference index, and judge the driver's facial fatigue state according to the fatigue state judgment value;

[0011] Extract features from the road video frames to obtain the left and right lane lines, calculate the angles between the left and right lane lines and the lower edge of the image according to the slopes of the left and right lane lines, and calculate the yaw rate based on the angles;

[0012] Perform weighted fusion on the driver's facial fatigue state, yaw rate, and the detection results of the intelligent wearable device to obtain the comprehensive fatigue state judgment value, and judge the driver's driving state according to the comprehensive fatigue state judgment value.

[0013] As an alternative implementation, the individual difference index includes the maximum value of the eye aspect ratio and the minimum value of the mouth aspect ratio.

[0014] As an alternative implementation, the calculation formula for the fatigue state judgment value is:

[0015]

[0016] where EAR_Max is the maximum value of the eye aspect ratio, MAR_Min is the minimum value of the mouth aspect ratio, EAR_Now is the currently detected eye aspect ratio, and MAR_Now is the currently detected mouth aspect ratio.

[0017] As an alternative implementation, judging the driver's facial fatigue state according to the fatigue state judgment value is specifically:

[0018] When the duration or frequency of a < a Th in the video is greater than the set threshold, it is judged that the driver is in a fatigued state;

[0019] When the frequency of b > b Th in the video is greater than the set threshold, it is judged that the driver is in a fatigued state;

[0020] where a Th is the judgment threshold for the driver's eye fatigue state, and b Th is the judgment threshold for the driver's mouth fatigue state.

[0021] As an alternative implementation, the calculation formula for the yaw rate is:

[0022]

[0023] where θ l and θr They are respectively the inner included angles between the left and right lane lines and the lower edge of the image.

[0024] As an alternative implementation, the detection results of the driver's facial fatigue state, yaw rate, and intelligent wearable device are weighted and fused. Specifically:

[0025] Normalize the detection results of the driver's facial fatigue state, yaw rate, and intelligent wearable device, and use the weighted average method to assign corresponding weights to each fatigue feature after normalization to obtain a comprehensive fatigue state judgment value.

[0026] In a second aspect, the present invention provides a fatigue driving detection system based on a camera and heart rate monitoring, including:

[0027] A data acquisition module, configured to: acquire the in-cab video data and road video data;

[0028] A facial fatigue state judgment module, configured to: extract features from the in-cab video frames to obtain the driver's facial features, obtain individual difference indicators based on the driver's facial features, calculate a fatigue state judgment value according to the individual difference indicators, and judge the driver's facial fatigue state according to the fatigue state judgment value;

[0029] A yaw rate calculation module, configured to: extract features from the road video frames to obtain the left and right lane lines, calculate the included angles between the left and right lane lines and the lower edge of the image according to the slopes of the left and right lane lines, and calculate the yaw rate based on the included angles;

[0030] A comprehensive judgment module, configured to: perform weighted fusion on the detection results of the driver's facial fatigue state, yaw rate, and intelligent wearable device to obtain a comprehensive fatigue state judgment value, and judge the driver's driving state according to the comprehensive fatigue state judgment value.

[0031] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in the first aspect is completed.

[0033] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, the method described in the first aspect is implemented.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] The present invention proposes a fatigue driving detection method and system based on camera and heart rate monitoring, comprehensively considers the fatigue characteristics of different information sources, adopts a multi-feature fusion method to describe the driver's fatigue state, extracts fatigue feature indicators from the eye, mouth and lane departure features, adopts a weighted average method to fuse different fatigue features to form a more reliable fatigue judgment condition, and combines the intelligent wearable device to detect the heart rate to comprehensively judge the current driving state, sets a three-level fatigue judgment benchmark, and establishes a fatigue detection algorithm that can accurately judge the driver's fatigue level. If the state is not good, an alarm prompt is issued.

[0036] The present invention proposes a fatigue driving detection method and system based on a camera and heart rate monitoring. Aiming at the problem of poor real-time performance of existing fatigue driving detection methods, the present invention innovatively proposes an improved face tracking prediction algorithm to improve the tracking effect of the algorithm. The improved tracking algorithm is integrated with the detection algorithm to reduce the detection time and improve the real-time performance.

[0037] The present invention proposes a fatigue driving detection method and system based on camera and heart rate monitoring, which aims at the problem that the thresholds cannot be unified due to individual differences in the size of the driver's eyes and mouth. According to the actual eye size and mouth opening and closing size of the driver, the present invention selects the maximum value of the eye ratio and the minimum value of the mouth opening and closing ratio of the previous several video frames and continuously replaces them, selects a specific fatigue threshold suitable for each driver to judge the degree of fatigue, and has high accuracy. Effectively improve the accuracy of driving fatigue detection.

[0038] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0040] Figure 1 A flowchart of a method for detecting fatigue driving based on a camera and heart rate monitoring provided in Example 1 of the present invention;

[0041] Figure 2 A flowchart of a tracking prediction algorithm provided in Example 1 of the present invention;

[0042] Figure 3 A schematic diagram of the structure of the lane detection model provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0044] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0045] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0046] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0047] Embodiment 1

[0048] As Figure 1 shown, this embodiment provides a fatigue driving detection method based on a camera and heart rate monitoring, including the following steps:

[0049] S1. Collect the video data inside the cab and the road video data;

[0050] S2. Extract features from the video frames inside the cab to obtain the driver's facial features, obtain individual difference indicators based on the driver's facial features, calculate a fatigue state judgment value according to the individual difference indicators, and judge the driver's facial fatigue state according to the fatigue state judgment value;

[0051] S3. Extract features from the road video frames to obtain the left and right lane lines, calculate the angles between the left and right lane lines and the lower edge of the image according to the slopes of the left and right lane lines, and calculate the yaw rate based on the angles;

[0052] S4. Perform weighted fusion on the driver's facial fatigue state, yaw rate, and the detection results of the intelligent wearable device to obtain a comprehensive fatigue state judgment value, and judge the driver's driving state according to the comprehensive fatigue state judgment value.

[0053] In the existing methods, most of the driver fatigue detection algorithms are based on a single eye or mouth feature as the judgment basis. The single feature lacks rich fatigue information and is easily interfered by the driver's personal habits or the external environment, resulting in the failure of the deviation detection algorithm. Therefore, the driver fatigue detection should comprehensively consider various influencing factors.

[0054] The present invention combines the facial features of the driver, the navigation features of the vehicle, and the physical state indicators of the driver, effectively improving the accuracy of driving fatigue detection.

[0055] (1) Facial features of the driver

[0056] In the field of driver state detection, detection based on facial features is one of the most commonly used means. It mainly monitors the movement of the driver's eyes and mouth and the position change of the head, and then determines the degree of fatigue. However, most of the existing detection algorithms are based on the values of the mouth aspect ratio (MAR) and the eye aspect ratio (EAR), and use the threshold of the feature points of the driver's eyes and mouth as the judgment feature to determine whether the driver is fatigued. Most of these algorithms ignore the individual characteristics of the driver and use fixed thresholds to judge the states of the eyes and mouth, which will obviously lead to a high misjudgment rate. In fact, the eye sizes of different drivers are different, the opening amplitudes of the mouths are also different, and the corresponding parameter thresholds are also different.

[0057] The present invention first extracts the facial features of the driver from the collected video data, calculates the maximum value of the eye aspect ratio (EAR_Max) and the minimum value of the mouth aspect ratio (MAR_Min) according to the actual eye size and mouth opening size of the driver, and calculates the ratios of the EAR and MAR values of each subsequent frame of image to EAR_Max and MAR_Min and assigns new values a and b. Among them, a and b are used as fatigue state judgment values, and their calculation formulas are:

[0058]

[0059] Among them, EAR_Max is the maximum value of the eye aspect ratio, MAR_Min is the minimum value of the mouth aspect ratio, EAR_Now is the currently detected eye aspect ratio, and MAR_Now is the currently detected mouth aspect ratio. During the driving process of the driver, a and b are dynamically adjusted. The facial fatigue state of the driver is judged according to the fatigue state judgment value, specifically:

[0060] When the duration or frequency of a < a Th in the video is greater than the set threshold, it is judged that the driver is in a fatigued state;

[0061] When the frequency of b > b Th in the video is greater than the set threshold, it is judged that the driver is in a fatigued state;

[0062] Among them, a Th is the judgment threshold for the eye fatigue state of the driver, and b Th is the judgment threshold for the mouth fatigue state of the driver. a Th and b ThSet according to the actual situation.

[0063] The present invention selects a specific fatigue threshold applicable to each driver to judge the degree of fatigue, with high precision.

[0064] In addition, as a driving behavior with extremely high risk, fatigue driving requires the detection system to have good real-time performance. The detection system needs to quickly detect the fatigue characteristics of the driver continuously in consecutive image frames, and also has extremely high requirements for the algorithm stability. The CamShift algorithm is an iterative optimization technology for object tracking. The traditional CamShift algorithm needs to non-autonomously select the area containing the tracking target in the initial frame through manual annotation. The colors similar to the target have a great impact on the algorithm stability during the tracking process. When the tracking target has strong motility, it is easy to cause the tracking window to deviate or even completely lose the tracking target.

[0065] In view of the above problems, the present invention combines an improved CamShift algorithm with Kalman filtering and proposes a target tracking and prediction algorithm as follows:

[0066] Taking the improved CamShift algorithm for realizing moving target tracking as the basic algorithm. At the same time, using the optimal estimation of the state value obtained by Kalman filtering to reasonably adjust the search box to obtain the best tracking result.

[0067] As Figure 2 shown, first, use the face detection algorithm to extract the face of the driver in the input video frame as the tracking target, and initialize the Kalman filter and the CamShift target tracking algorithm with this. Then judge whether the occlusion or color interference is greater than the threshold. If it is greater than the threshold, add the predicted value of the Kalman filter as the observed value of the next frame to update the position and size of the search window, and then continue to track with the CamShift algorithm. If it is less than the threshold, it is considered that there is no color interference and serious occlusion, and directly output the calculation result of the CamShift target tracking algorithm. Repeat the above steps until the tracking ends.

[0068] The present invention combines the target detection algorithm and the tracking algorithm, while ensuring better detection effect, greatly reducing the detection time and improving the real-time performance. When the target detection algorithm detects the face image in the initial frame, the tracking and prediction algorithm starts.

[0069] (2) Navigation characteristics of the vehicle

[0070] When a driver is fatigued, due to the disorder of physiological and psychological functions, phenomena such as slow action or decreased judgment may occur, and even the ability to control the vehicle may be lost, resulting in the vehicle deviating from its original track and driving, which is extremely likely to cause traffic accidents. Therefore, the lane departure state should be an important basis for judging driver fatigue. By using the lane line detection algorithm to obtain the lane line parameters in real time and combining the lane departure strategy to obtain the position information of the vehicle in the current lane, the accuracy and real-time performance of the lane line detection algorithm are crucial for lane departure detection. Due to environmental factors such as road occlusion and lane blur, as well as the inherent sparse characteristics of lane lines in road images, the performance of traditional lane line detection algorithms is poor, and it is difficult for ordinary convolutional neural networks to extract accurate lane line features from road images. To address this problem, a high-precision real-time lane line detection model is designed.

[0071] As Figure 3 shown, it includes an encoder, a feature enhancement module, a decoder, and a lane line prediction branch. First, the modified lightweight network RepVgg-A0 is used to encode the road image. Then, a multi-size asymmetric shuffle convolution module is proposed according to the characteristics of lane lines to enhance the model's ability to extract lane line features. Finally, an adaptive upsampling module is proposed as the decoder to upsample the feature map to the original resolution for pixel-level classification detection, and a lane line prediction branch is added to output the confidence of the existence of lane lines. The structural design of each sub-module will be described in detail below.

[0072] 1. Encoder Structure

[0073] After the road image is input, the lightweight network RepVgg-A0 is used as the encoder of the model to initially extract lane line features.

[0074] The original RepVgg-A0 network downsamples the input image to 1 / 32 of the original image through three convolutional layers with a stride of 2, reducing the image resolution while increasing the receptive field. However, the overly small resolution will cause a large amount of spatial information in the encoded image to be lost, and it is difficult for the subsequent decoding process to repair this information, affecting the lane detection accuracy. Therefore, the structure of the RepVgg-A0 network is adjusted, and the strides of the convolutions in the last two layers of the network are both set to 1, reducing the downsampling ratio from 32 times to 8 times. After the above operations, the size of the encoded feature map increases relatively, retaining more original lane information, but another problem arises. The receptive field of the network becomes smaller, making it difficult to learn global features. Therefore, dilated convolutions are introduced in the last two layers of the network. The regular 3×3 convolutions in the last two layers of the RepVgg-A0 network are replaced with 3×3 dilated convolutions with dilation rates of 2 and 4 respectively, expanding the receptive field of the network without introducing additional computational complexity. After the initially extracted features from the input 3-channel image pass through the modified RepVgg-A0 network, the number of channels becomes 1280, and the resolution is reduced to 1 / 8 of the original. To reduce the computational complexity of subsequent operations and at the same time fuse the extracted features, a 1×1 convolution is added to the last layer of the encoder to compress the number of channels to 128.

[0075] 2. Feature Enhancement Module

[0076] Currently, image segmentation algorithms based on the encoder-decoder structure usually first use a pre-trained classification neural network for feature encoding, and then directly upsample the encoded feature map to extract semantic information. However, conventional classification neural networks have translational invariance, which is a feature friendly to the target classification task. Simply put, no matter what position transformation the target image undergoes, the same response output will be obtained. This means that the position information in the image is more likely to be ignored, while the segmentation task pays more attention to the spatial features in the image. Pixels at the same position may have different semantic information in different instances. When the position of the target to be detected changes, the instance mask output by the network should also change accordingly. Therefore, directly decoding the features extracted by the classification network will affect the performance of the lane detection model. To address the above problems, a lightweight feature enhancement module is proposed to aggregate the lane line features extracted by the encoder, significantly improving the accuracy of lane detection while only introducing a small number of parameters.

[0077] Lane line detection is different from the detection of conventional objects. It has the inherent characteristics of being sparse and slender. A lane line usually spans the entire image, which requires the network to have a sufficiently large receptive field. For instance segmentation networks, an effective way to increase the receptive field is to use convolutional kernels of larger sizes. Inspired by the ShuffleNet V2 architecture, the present invention designs a multi-size shuffle convolution module that includes convolutional kernels of three sizes: 3×3, 5×5, and 7×7. Among them, the 3×3 convolution is used to extract the detailed features of the lane lines, and the 5×5 and 7×7 convolutions have a larger receptive field and can capture lane line features at a larger scale. After the feature map is input, the channels are first evenly divided into two branches. The secondary branch performs an equal mapping, and the primary branch sequentially performs convolutions of three sizes: 3×3, 5×5, and 7×7, and adds non-linear factors using the FReLu activation function after each convolution. Finally, the channels of the primary branch and the secondary branch are concatenated and then a full-channel shuffle operation is performed to promote the fusion of feature information between channels.

[0078] Compared with directly using large convolutional kernels, using the shuffle convolution module reduces the computational cost to a certain extent. However, the 5×5 and 7×7 convolutions still account for a relatively large amount of computation. Therefore, the present invention introduces asymmetric convolutions to further reduce the computational cost. Asymmetric convolutions replace the conventional k×k convolution with k×1 and 1×k convolutions, which can significantly reduce the computational cost. For a conventional k×k convolution, the number of parameters and the computational amount for one convolution are respectively:

[0079] Params=k 2 C i C o

[0080] FLOPs=k 2 C i C o H o W o

[0081] Among them, Hi and Wi are respectively the height and width of the input feature map, and Ci and Co are respectively the number of channels of the input and output feature maps. The number of parameters and the computational amount of the asymmetric convolution equivalent to the k×k convolution are respectively:

[0082] Params a =2kC i C o

[0083] FLOPs a =kC i C o H o (2W o +k-1)

[0084] As can be seen from the above formula, the larger the size of the convolutional kernel, the more obvious the reduction in the number of parameters and the amount of computation when it is converted into an asymmetric convolution. In addition, research shows that the application of asymmetric convolution in the middle layer of the network has better effects. Therefore, the 5×5 and 7×7 convolutions in the multi-size shuffle convolution module are replaced with asymmetric convolutions. For the multi-size asymmetric shuffle convolution module of the present invention, for an input feature map with a fixed size of 46×80×128, the number of parameters and the amount of computation of the module are reduced by 60.24% and 61.47% respectively.

[0085] Six multi-size asymmetric shuffle convolution modules are stacked to form the feature enhancement module of the lane line detection model of the present invention. Among them, the latter 5 modules use dilated convolutions to further expand the receptive field, and the dilation rates are set to 2, 4, 6, 8, and 10 respectively. The feature enhancement module further extracts the lane line information existing in the feature map output by the encoder and inputs the result into the lane line prediction branch and the decoder structure.

[0086] 5. Lane Line Prediction Branch

[0087] The lane line detection model designed by the present invention can detect up to 6 lane lines at the same time. The lane line prediction branch is used to judge the presence of each lane line. First, the number of channels is reduced to 7 through a 1×1 convolution, and after being activated by Softmax, it is downsampled to 23×40×7 using average pooling with a stride of 2. Subsequently, 2 fully connected layers are continuously used and activated by ReLU and Sigmoid respectively, and a one-dimensional feature vector with a length of 6 is output, which respectively represents the probabilities of the existence of 6 preselected lane lines. In actual use, a confidence threshold is set. When the confidence is greater than the threshold, it means that the lane line exists, otherwise it does not exist. In this embodiment, the threshold is set to 0.5.

[0088] 4. Decoder

[0089] The role of the decoder is to upsample the low-resolution feature map containing rich feature information to the size of the input image, so as to classify each pixel in the feature map. The commonly used upsampling methods mainly include bilinear interpolation and transposed convolution.

[0090] Adding the bilinear interpolation and transposed convolution upsampling results directly together has achieved certain effects, but the applicability of the two upsampling methods to specific image regions is not considered. On this basis, the present invention proposes an adaptive upsampling module, allowing the network to select the weights of the two upsampling methods at each position by itself, which can extract image features more efficiently.

[0091] After the adaptive upsampling module takes as input a feature map of size H×W×C, it first performs upsampling using bilinear interpolation and transposed convolution respectively to initially obtain two upsampled feature maps E and F of size 2H×2W×C / 2. Then, E and F are concatenated in the channel dimension to obtain a feature map G of size 2H×2W×C. Next, a 3×3 convolution is performed on G to extract the spatial attention description S(2H×2W×2), and the Softmax function is used on S to extract two attention weights of size 2H×2W×1 each. Finally, the attention weights are weighted and summed with E and F respectively to obtain the final upsampling result (2H×2W×C / 2).

[0092] Stacking the adaptive upsampling module three times constitutes the decoder of the lane line detection model of the present invention. After being decoded by the decoder, the input feature map is upsampled to the original image size, and the number of channels is reduced to 7. The first channel is used to predict the lane line background, and the remaining channels directly predict the pixel coordinates of the lane line instances, which has a faster detection speed compared to the algorithm that first performs semantic segmentation and then fits the lane lines.

[0093] Based on the lane line detection model of the present invention, the pixel coordinates of the lane line instances can be directly obtained. For lane departure detection, parameters such as the slope provided by the lane line detection need to be used to determine the lane departure situation. Considering that the lane lines in the near field close to the vehicle show linear characteristics, the lower half area of the lane image is regarded as the near field, and the least squares method is used to fit the pixel points of the lane lines in the near field to obtain the linear regression equation of the lane lines in the lower half of the image:

[0094] v = -ku + b

[0095] where (u, v) are the coordinates of the lane line in the pixel coordinate system, k is the slope of the lane line, and b is the ordinate of the intersection point of the lane line and the v-axis.

[0096] In the pixel coordinate system, if the vehicle is driving in the middle of the road, the slope of the left lane line is negative, and the right one is positive, and the absolute value of the slope increases as the distance from the vehicle decreases. Lane departure detection mainly focuses on the left and right two lane lines adjacent to the vehicle. Therefore, the two lane lines with the largest positive slope and the smallest negative slope are respectively used as the research objects for lane departure detection. The slopes of the fitting lines of the left and right two lane lines are respectively used for lane departure analysis. According to the slope, the angles between the left and right lane lines and the lower edge of the image can be obtained as:

[0097] θ l = arctan((-k l )

[0098] θ r = arctan(k l )

[0099] where θl and θ r are respectively the inner included angles between the left and right lane lines and the lower edge of the image, and K l and K r are respectively the slopes of the left and right lane lines.

[0100] The calculation formula of the yaw rate is:

[0101]

[0102] Among them, θ l and θ r are respectively the inner included angles between the left and right lane lines and the lower edge of the image.

[0103] Among them, δ represents the yaw rate. When the vehicle is driving in the center of the road, the yaw rate is close to 1. When yawing to the left, the yaw rate increases, and when yawing to the right, the yaw rate decreases. Therefore, the change of the yaw rate within a certain period of time can sensitively reflect whether the vehicle yaws.

[0104] (III) Driver's physical state indicators

[0105] Some studies have found that there is a certain correlation between heart rate and fatigue driving. When the human body is in a fatigued state, it will cause the heart rate to rise. The present invention uses a smart wearable device to detect the driver's heart rate. As an auxiliary feature, combined with the above-mentioned driver's facial fatigue characteristics and driving yaw rate, a comprehensive judgment of fatigue driving is made.

[0106] When the driver is fatigued, various fatigue characteristics will appear, such as longer eye-closure time, increased number of yawns, deeper vehicle deviation, and rising heart rate, etc. Some of these characteristics have a high correlation with the driver's fatigue state, such as long-term eye closure, and some can only be used as an auxiliary basis for judging the fatigue state, such as lane deviation and rising heart rate, etc. Although the eye characteristics largely reflect the driver's fatigue state, in actual applications, it may be affected by environmental factors such as the driver wearing sunglasses or vehicle jitter, resulting in the failure of eye state detection, and there are also differences in the performance of individuals when fatigued. Therefore, relying solely on a single index to judge the driver's fatigue state has certain limitations. In order to improve the reliability of the fatigue detection algorithm, the present invention comprehensively considers the fatigue characteristics of different information sources and uses a multi-feature fusion method to describe the driver's fatigue state.

[0107] Since the dimension units of different fatigue judgment indexes are different, in order to simplify the subsequent calculation process of fatigue feature fusion, it is necessary to normalize the driver's facial fatigue state, yaw rate, and driver's heart rate data, and use the weighted average method to assign corresponding weights to each normalized fatigue feature to obtain a comprehensive fatigue state judgment value.

[0108] The calculation formula of the comprehensive fatigue state judgment value is:

[0109]

[0110] Among them, i represents each fatigue determination index after normalization, and W i is the corresponding weight.

[0111] The fatigue value Fatigue is classified into levels to obtain a consistent description of the driver's fatigue state. Table 1 shows the fatigue level classification rules after multi-feature fusion. The higher the Fatigue value, the higher the driver's fatigue level. According to the magnitude of Fatigue, the driver's fatigue state is divided into three levels: awake, mildly fatigued, and severely fatigued.

[0112] Table 1 Fatigue level classification table after multi-feature fusion

[0113]

[0114] When the driver's fatigue state is severely fatigued, an alarm is issued.

[0115] Embodiment 2

[0116] This embodiment provides a fatigue driving detection system based on a camera and heart rate monitoring, including:

[0117] A data acquisition module, configured to: acquire the in-cab video data and the road video data;

[0118] A facial fatigue state judgment module, configured to: extract features from the in-cab video frames to obtain the driver's facial features, obtain the individual difference index based on the driver's facial features, calculate the fatigue state judgment value according to the individual difference index, and judge the driver's facial fatigue state according to the fatigue state judgment value;

[0119] A yaw rate calculation module, configured to: extract features from the road video frames to obtain the left and right lane lines, calculate the angles between the left and right lane lines and the lower edge of the image according to the slopes of the left and right lane lines, and calculate the yaw rate based on the angles;

[0120] A comprehensive judgment module, configured to: perform weighted fusion on the driver's facial fatigue state, the yaw rate, and the detection results of the intelligent wearable device to obtain a comprehensive fatigue state judgment value, and judge the driver's driving state according to the comprehensive fatigue state judgment value.

[0121] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0122] In more embodiments, the following is also provided:

[0123] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0124] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0125] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0126] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0127] The method in Embodiment 1 can be directly embodied as being executed and completed by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules may be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0128] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.

[0129] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed locally or within a distributed device. In a distributed device, program modules can be located in local and remote storage media.

[0130] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0131] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0132] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0133] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for detecting fatigue driving based on camera and heart rate monitoring, characterized in that: The following steps are involved: Collecting video data inside the cab and on the road; Extract features from the video frames inside the cab to obtain facial features of the driver, obtain individual difference indicators based on the facial features of the driver, calculate fatigue state judgment values ​​based on the individual difference indicators, and judge the facial fatigue state of the driver based on the fatigue state judgment values; Extract features from road video frames to obtain left and right lane lines, calculate the angle between the left and right lane lines and the bottom edge of the image based on the slopes of the left and right lane lines, and calculate the yaw rate based on the angle; The driver's facial fatigue state, yaw rate and detection results of the smart wearable device are weightedly fused to obtain a comprehensive fatigue state judgment value, and the driver's driving state is judged based on the comprehensive fatigue state judgment value.

2. A method for detecting fatigue driving based on camera and heart rate monitoring as claimed in claim 1, characterized in that: The individual difference indicators include a maximum value of the eye length-to-width ratio and a minimum value of the mouth length-to-width ratio.

3. A method for detecting fatigue driving based on camera and heart rate monitoring as claimed in claim 1, characterized in that: The calculation formula of fatigue state judgment value is: Among them, EAR_Max is the maximum value of the eye aspect ratio, MAR_Min is the minimum value of the mouth aspect ratio, EAR_Now is the real-time detected eye aspect ratio, and MAR_Now is the real-time detected mouth aspect ratio.

4. A method for detecting fatigue driving based on camera and heart rate monitoring as claimed in claim 3, characterized in that: The driver's facial fatigue state is judged according to the fatigue state judgment value, specifically: When the video Th When the duration or frequency of fatigue is greater than the set threshold, the driver is judged to be in fatigue state;​ When the video Th When the frequency is greater than the set threshold, the driver is judged to be in a fatigue state; Among them, a Th is the driver's eye fatigue state judgment threshold, b Th It is the threshold for judging the driver's mouth fatigue state.

5. The method for detecting fatigue driving based on camera and heart rate monitoring as claimed in claim 1, characterized in that: The calculation formula of yaw rate is: Among them, θ l and θ r are the inner angles between the left and right lane lines and the lower edge of the image respectively.

6. The method for detecting fatigue driving based on camera and heart rate monitoring as claimed in claim 1, characterized in that: The driver's facial fatigue status, yaw rate and the detection results of smart wearable devices are weighted fused as follows: The driver's facial fatigue state, yaw rate and detection results of smart wearable devices are normalized, and the weighted average method is used to assign corresponding weights to each normalized fatigue feature to obtain the comprehensive fatigue state judgment value.

7. A fatigue driving detection system based on camera and heart rate monitoring, characterized in that: include: The data acquisition module is configured to: acquire the video data inside the cab and the video data on the road; The facial fatigue state judgment module is configured to: extract features from the video frames inside the cab to obtain facial features of the driver, obtain individual difference indicators based on the facial features of the driver, calculate a fatigue state judgment value based on the individual difference indicators, and judge the facial fatigue state of the driver based on the fatigue state judgment value; The yaw rate calculation module is configured to: extract features from the road video frame to obtain left and right lane lines, calculate the angle between the left and right lane lines and the lower edge of the image according to the slopes of the left and right lane lines, and calculate the yaw rate based on the angle; The comprehensive judgment module is configured to: perform weighted fusion on the driver's facial fatigue state, yaw rate and detection results of the smart wearable device to obtain a comprehensive fatigue state judgment value, and judge the driver's driving state according to the comprehensive fatigue state judgment value.

8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Fatigue driving detection method in infrared monitoring scene

    CN121482756A

  • A method for detecting fatigue driving in an infrared monitoring scene

    CN121482756B