Driver fatigue state detection system based on computer vision processing

By combining multi-dimensional information on facial blinking frequency, upper body movement frequency and grip strength, using high-definition cameras and grip strength sensors for comprehensive judgment, the existing system's inaccurate recognition of light and occlusion underneath is solved, and more accurate fatigue state detection is achieved.

CN120227031AActive Publication Date: 2025-07-01SHANDONG JIANZHU UNIV

Patent Information

Application Number
CN202510306010.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Under the influence of factors such as light and line of sight occlusion, the existing driver fatigue monitoring system is inaccurate in recognition of facial features, resulting in inaccurate fatigue status judgment.

Method used

Combining multi-dimensional information such as facial blink frequency, upper body movement frequency and grip strength, data is collected through high-definition cameras and grip strength sensors, and a deep neural network and decision tree model are used to make comprehensive judgments.

Benefits of technology

Provide richer and more comprehensive data support, improves the accuracy of fatigue state judgment, adapts to complex driving environments and reduces misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120227031A_ABST
    Figure CN120227031A_ABST
Patent Text Reader

Abstract

The invention discloses a driver fatigue state detection system based on computer vision processing. The system comprises a facial feature extraction unit which selects a high-definition camera to capture a driver facial image; the upper body feature monitoring unit collects upper body image data through a high-definition camera to accurately extract the action frequency of the upper body; the hand-held steering wheel characteristic monitoring unit is used for sensing pressure applied by hands of a driver by arranging a grip strength sensor on a steering wheel through a protective sleeve; the data processing unit receives data information collected by the facial feature extraction unit, the upper body feature monitoring unit and the handheld steering wheel feature monitoring unit for feature fusion; the fatigue state judgment unit judges the fatigue state of the driver according to a preset fatigue judgment rule; the face blinking frequency is combined with multi-dimensional information such as the upper body action frequency and the grip strength, the state of a driver can be reflected from different angles, richer and more comprehensive data support is provided for fatigue judgment, and the fatigue state can be judged more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of driver fatigue state detection, and particularly to a driver fatigue state detection system based on computer vision processing. Background Art

[0002] After a truck driver has been driving continuously for a long time, the body and brain are in a fatigued state and the reaction speed will decrease significantly. According to research, in a fatigued state, the driver's attention concentration time will be greatly shortened from the normal 30 - 45 minutes to 10 - 15 minutes or even shorter. When facing suddenly emerging obstacles or other emergencies ahead during driving, a fatigued driver cannot make correct operations such as braking and avoidance in time, greatly increasing the risk of accidents.

[0003] Due to the time requirements for transporting goods or traffic control, many large trucks drive long distances at night. Since large trucks are not as intelligent as small cars, for many large trucks (such as traditional trucks), to avoid fatigue of the large truck driver during long - term night driving, in many cases, it relies on the driver himself or the reminder of the co - driver to ensure driving safety. There are also some new large trucks equipped with driver fatigue monitoring systems. These systems use high - precision cameras and advanced artificial intelligence algorithms to monitor the driver's facial features, eye movements, and emotional states in real - time, effectively preventing fatigue driving and distracted driving. However, the single - camera monitoring may be affected by factors such as light and line - of - sight occlusion, resulting in inaccurate facial feature recognition, thus affecting the judgment of the driver's state. Summary of the Invention

[0004] By providing a driver fatigue state detection system based on computer vision processing in the embodiments of the present application, using multi - dimensional information such as the combination of facial blink frequency, upper - body movement frequency, and grip strength, it can reflect the driver's state from different angles, provide richer and more comprehensive data support for fatigue judgment, and can more accurately judge the fatigue state.

[0005] The embodiments of the present application provide a driver fatigue state detection system based on computer vision processing, including a facial feature extraction unit, which selects a high - definition camera to capture the driver's facial image. The high - definition camera is fixed at a suitable position in the cab of the large truck using a vacuum suction cup and a metal - shaped flexible rod (which does not interfere with the driver's normal driving of the vehicle and can also capture the driver's face and upper body in a large range);

[0006] An upper - body feature monitoring unit, which collects upper - body image data through the high - definition camera and uses a method combining background subtraction technology and optical flow method to accurately extract the movement frequency of the upper body;

[0007] The hand - held steering wheel feature monitoring unit uses a protective cover to set grip sensors on the steering wheel. It can accurately sense the pressure exerted by the driver's hand by collecting data at a frequency of 70Hz.

[0008] The data processing unit adopts a feature fusion architecture based on a deep neural network, constructs a deep neural network with multiple hidden layers, receives the data information collected from the facial feature extraction unit, the upper body feature monitoring unit, and the hand - held steering wheel feature monitoring unit, and takes the pre - processed facial, upper body, and hand - held steering wheel features as different input branches respectively. The features of different branches are fused in a certain hidden layer using a fusion method based on the attention mechanism.

[0009] The fatigue state judgment unit judges the fatigue state of the driver based on the analysis results of the data processing unit according to the preset fatigue judgment rules.

[0010] The data processing unit receives the image information taken by the facial feature extraction unit, uses a convolutional neural network model to optimize the detection of facial feature points and outputs the coordinates of the key facial feature points, accurately depicts the facial contour and the shape of key parts. First, it determines the eye closure time according to the eye features. Usually, specific feature points on the upper and lower eyelids are selected for the eyes, and the distance between the specific feature points on the upper and lower eyelids is calculated to judge whether the eyes are closed.

[0011] Blink events are identified by monitoring the alternating changes in the closed and open states of the eyes. However, in the actual process, the blink frequency may be interfered by various factors, so dynamic blink frequency analysis is introduced to dynamically analyze and correct the blink frequency.

[0012] By analyzing the eye states in multiple consecutive frames of images, normal blinks and abnormal blinks are identified. For normal blinks, a standard blink cycle range T min to T max (unit: s) is set. When a blink event is detected and its duration is between T min and T max , it is recognized as a normal blink and counted into the blink count.

[0013] For the gradually changing trend of the blink frequency in the driver's fatigue state, time - series analysis is used to predict and correct the blink frequency. Using the blink frequency data in the past period of time, an LSTM model is constructed, and by calculating the deviation rate between the predicted value and the actual value, the blink frequency in the future period of time is predicted.

[0014] Before the driver enters the vehicle but before starting the driving action, the high-definition camera first captures a video sequence. The data processing unit processes each frame image in the sequence, converts it to a suitable color space, and uses the Gaussian mixture model to construct a background model. Each Gaussian distribution defines the mean μ k , covariance matrix Σk, and weight ω k , and the matching degree of each pixel point with each Gaussian distribution is output through the Mahalanobis distance D.

[0015] During the vehicle driving process, each frame image I(t) captured in real time is compared with the background model. Similarly, the current frame image is converted to the YUV color space, and for each pixel point P(x, y), its Mahalanobis distance from each Gaussian distribution in the background model is calculated.

[0016] If the Mahalanobis distances of all Gaussian distributions are greater than the set threshold, then this pixel point is considered to belong to the foreground (i.e., the driver's upper body area); otherwise, this pixel point is considered to belong to the background label.

[0017] For the foreground mask image sequence obtained by background subtraction, an improved Lucas-Kanade optical flow algorithm is used for optical flow calculation, introducing spatio-temporal context information. It not only considers the information of the current frame and the next frame but also combines the optical flow information of the previous few frames to constrain the displacement calculation of the current pixel point.

[0018] Different weights are assigned according to the position of the pixel point in the upper body area. First, the direction change amount between adjacent frames of each pixel point in the foreground area is statistically calculated, and the direction change amounts of all pixel points in the foreground area are weighted and averaged to obtain the overall action direction change.

[0019] In the action frequency, by more carefully defining the action events, when the optical flow vectors of the foreground pixel points in a series of adjacent frames meet the following conditions, it is considered to be part of an action event: the direction of the optical flow vector is within a certain angle range and the action amplitude is greater than the set amplitude threshold. The number of times that meet the action event definition in the analysis of consecutive frames is the action frequency.

[0020] The grip force sensor selects a high-precision piezoresistive grip force sensor (at least 6). The weak electrical signals output by the sensor are preprocessed, such as amplified and filtered, by a high-speed data acquisition circuit to remove noise interference, and then converted into digital signals and transmitted to the data processing unit.

[0021] The original data collected by the grip force sensor is linearly calibrated using the defined linear equation F = k s V + b s for calibration, where k s and b s are calibration coefficients determined through calibration experiments.

[0022] In order to obtain the overall grip strength of the driver's hand on the steering wheel, considering that different position sensors contribute differently to the overall grip strength, weighted averaging is adopted. Sensors near the top and bottom of the steering wheel may be more critical for controlling the direction and are given relatively higher weights.

[0023] To accurately identify grip strength change events, two thresholds are set, namely the rising threshold T r and the falling threshold T f . When the grip strength rapidly rises from below T r and exceeds T r , and then drops below T f , it is recognized as a complete grip strength change event. By continuously collecting grip strength data and performing real-time monitoring, the number of grip strength change events occurring per unit time is the grip strength frequency. Define a state variable to track the change state of the grip strength, such as "rising stage, falling stage, stable stage". When the grip strength enters the "rising stage" from the "stable stage" and exceeds T r , counting starts; when the grip strength enters the "falling stage" and is below T f , a count is completed, and the state variable is reset to the "stable stage" to prepare for the next event detection. The number of events obtained after 1 minute of statistics is the grip strength frequency.

[0024] The data processing unit uses the fused features as the input of the decision tree model, defines the driver's fatigue state as the output for training. The decision tree constructs a tree structure by continuously dividing the feature space, and based on the output of the decision tree model and the preset fatigue judgment rules, the fatigue state is judged.

[0025] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0026] 1. By using facial features, upper body features, and the feature of the hand grip strength on the steering wheel, it can reflect the driver's state from different angles, providing richer and more comprehensive data support for fatigue judgment. By combining multi-dimensional information such as the facial blink frequency, the upper body movement frequency, and the grip strength, the fatigue state can be judged more accurately.

[0027] 2. Deeply integrate the background subtraction technology with the improved Lucas-Kanade optical flow algorithm, adjust and improve the parameters for the driver's upper body monitoring scenario. In background subtraction, use the Gaussian mixture model to accurately model the complex and changeable driving environment background, and adapt to light changes by dynamically updating parameters; introduce spatio-temporal context information in optical flow calculation to improve the accuracy of upper body motion tracking under occlusion and light changes; moreover, assign different weights according to the position of pixel points in the upper body area, which can more accurately reflect the overall motion characteristics of the upper body. Compared with the traditional method that treats all pixel points equally, the weighted method can highlight the influence of the motion of key parts (such as near the body center) on the overall characteristics, and improve the accuracy and effectiveness of feature extraction.

[0028] 3. In the calculation of action frequency, by more carefully defining action events, comprehensively considering the direction and amplitude changes of optical flow vectors and their persistence within a certain number of frames, it avoids misjudgment caused by short-term and random small motions, improves the accuracy of action frequency statistics, and thus more accurately reflects the actual motion state of the driver's upper body.

[0029] 4. By assigning different weights to different position sensors according to their importance in driving operations, it can more accurately reflect the overall control force of the driver on the steering wheel. The weighted fusion can capture the force differences of the driver on different parts of the steering wheel in different driving scenarios. For example, when turning, the driver may exert more force on one side of the steering wheel, thus providing richer and more accurate grip force information for fatigue judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the driver fatigue state detection system of this application;

[0031] Figure 2 It is the flow of the driver fatigue state detection system of this application;

[0032] Figure 3 It is a schematic diagram of the hardware device of the driver fatigue state detection system of this application;

[0033] Figure 4 It is a schematic diagram of the layout of the grip force sensors on the steering wheel of this application. DETAILED DESCRIPTION OF THE INVENTION

[0034] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. The preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0036] Example 1

[0037] Please refer to Figures 1-3 , a driver fatigue state detection system based on computer vision processing, including a facial feature extraction unit, which selects a high-definition camera to capture the driver's facial image and records the eye closing time (unit: s), blink frequency (unit: times / minute), and head posture (deflection angle °);

[0038] An upper body feature monitoring unit, which collects upper body image data through a high-definition camera and uses a combination of background subtraction technology and optical flow method to accurately extract features such as the movement frequency (unit: times / minute), amplitude (unit: pixels), and direction (unit: °) changes of the upper body, so as to judge the driver's state;

[0039] A steering wheel grip feature monitoring unit, which sets a grip force sensor on the steering wheel and collects data at a frequency of 70 Hz, recording the grip force (unit: N) and frequency (unit: times / minute);

[0040] A data processing unit, which receives the data information collected from the facial feature extraction unit, the upper body feature monitoring unit, and the steering wheel grip feature monitoring unit, preprocesses this data, and then performs feature fusion and analysis;

[0041] A fatigue state judgment unit, which judges the driver's fatigue state based on the analysis results of the data processing unit according to the preset fatigue judgment rules.

[0042] Example 2

[0043] Using a convolutional neural network model to optimize the detection of facial feature points and output the coordinates of key facial feature points, accurately depicting the facial contour and the shape of key parts. First, determine the eye closing time according to the eye features. For the eyes, usually select specific feature points on the upper and lower eyelids (such as the feature points at the corners of the eyes and above and below the middle of the eyes), and judge whether the eyes are closed by calculating the distance between the specific feature points on the upper and lower eyelids. Secondly, detect the process of a complete blink, that is, the process of the eyes opening, closing, and then opening again. Define d as the Euclidean distance between the specific feature points on the upper and lower eyelids.

[0044]

[0045] where (x1, y1) and (x2, y2) are the coordinates of the corresponding feature points of the upper and lower eyelids respectively. When d is less than a pre-set threshold T d the eyes are considered to be in a closed state;

[0046] The high-definition camera records the start frame and end frame of each eye closure, defining the start frame as f s and the end frame as f e The eye closure time is T c ,

[0047]

[0048] where 30 is the assumed frame rate of the high-definition camera being 30fps.

[0049] Blink events can be identified by monitoring the alternating changes in the closed and open states of the eyes. However, in actual situations, the blink frequency may be interfered by various factors, such as the driver's emotional fluctuations, light changes, etc. Introduce dynamic blink frequency analysis to dynamically analyze and correct the blink frequency

[0050] By analyzing the eye states in a series of consecutive frames, normal blinks and abnormal blinks (such as rapid blinks caused by external stimuli or slow blinks caused by fatigue) are identified. For normal blinks, a standard blink cycle range T min to T max (in units of s) is set. When a blink event is detected and its duration is between T min and T max it is recognized as a normal blink and counted as the blink count;

[0051] For the gradually changing trend of the blink frequency in the driver's fatigued state, time series analysis is used to predict and correct the blink frequency. Using the blink frequency data over a past period (such as 3 minutes), an LSTM (Long Short-Term Memory Network) model is constructed to predict the blink frequency in the future period. By calculating the deviation rate between the predicted value and the actual value

[0052]

[0053] When δ > T δ (T δ is the deviation rate threshold), the actual blink frequency is adjusted according to the predicted value, where f p is the predicted value and f d is the actual value.

[0054] Example Three

[0055] Similarly, a high-definition camera is used to capture the upper body area of the driver. Before the driver enters the vehicle but has not started driving, the high-definition camera first captures a video sequence with a duration of T0 (such as 5 - 7s). For each frame image I t (t = 1, 2, 3, etc.) in this sequence is processed and converted to a suitable color space to reduce the impact of light changes on subsequent processing.

[0056] The Gaussian mixture model is used to construct the background model. For each pixel point in the image, at the initial stage, a mixture model of K Gaussian distributions is established for it. Each Gaussian distribution is defined by the mean μ k , covariance matrix Σk, and weight ω k denoted as. Specifically, for each pixel point P(x, y), at time t, according to the matching degree between the current pixel value I t (x, y) and each Gaussian distribution, the weight ω k (t), mean μ k (t), and covariance matrix Σ k (t) are updated. The matching degree is output through the Mahalanobis distance D.

[0057]

[0058] If D is less than a certain threshold, it is considered that this pixel point belongs to the corresponding Gaussian distribution, and its parameters are updated accordingly. After T0 seconds of learning, a stable background model B(x, y) is obtained;

[0059] During the driving process of the vehicle, each frame image I(t) captured in real time is compared with the background model. Similarly, the current frame image is converted to the YUV color space, and for each pixel point P(x, y), the Mahalanobis distance between it and each Gaussian distribution in the background model is calculated.

[0060] If the Mahalanobis distances of all Gaussian distributions are greater than the set threshold T d , then it is considered that this pixel point belongs to the foreground (i.e., the upper body area of the driver) and is marked as 1; otherwise, it is considered that this pixel point belongs to the background and is marked as 0. In this way, a binary foreground mask image M(t) is obtained, where the foreground area (the upper body of the driver) is white (value is 1), and the background area is black (value is 0).

[0061] For the foreground mask image sequence obtained through background subtraction, an improved Lucas-Kanade optical flow algorithm is used for optical flow calculation, introducing spatio-temporal context information. Not only the information of the current frame and the next frame is considered, but also the optical flow information of the previous few frames is combined to constrain the displacement calculation of the current pixel point.

[0062] Specifically, for the pixel P(x, y), when calculating its displacement (u, v), by performing weighted summation of the luminance change within a spatio-temporal window (e.g., including the current frame and the previous two frames in time, and a small neighborhood window centered at P(x, y) in space), a more accurate displacement estimation is obtained. Define the set of pixels within the spatio-temporal window as N(x, y), and the improved optical flow calculation equation is:

[0063]

[0064] where I x 、I y and I t are the partial derivatives of the image luminance I with respect to x, y, and t respectively, and ω(x', y', t') is the weight of each pixel within the spatio-temporal window, which is dynamically adjusted according to the spatio-temporal distance between the pixel and the current pixel P(x, y) and the consistency of the optical flow in the previous frames. The optical flow vector of each pixel is obtained by solving the above system of equations

[0065] Deeply fuse the background subtraction technique with the improved Lucas-Kanade optical flow algorithm, perform parameter adjustment and algorithm improvement for the driver's upper body monitoring scenario. In background subtraction, use the Gaussian mixture model to accurately model the complex and changeable driving environment background, and adapt to the illumination change by dynamically updating the parameters; introduce spatio-temporal context information in optical flow calculation, which improves the accuracy of upper body motion tracking under occlusion and illumination change.

[0066] To calculate the change in the action direction of the entire upper body, first, statistically calculate the change in direction Δθ(x, y) between adjacent frames for each pixel within the foreground region. Assume that at times t and t + 1, the directions of the pixel P(x, y) are defined as θ1(x, y) and θ2(x, y) respectively, then Δθ(x, y) = |θ2(x, y) - θ1(x, y)| (considering the periodicity of the angle, when |θ2 - θ1| > π, Δθ(x, y) = 2π - |θ2 - θ1|).

[0067] Perform weighted averaging on the change in direction of all pixels within the foreground region to obtain the overall change in action direction Δθ t and assign a weight ω w (x, y) to each pixel. The output Δθ t is:

[0068]

[0069] By assigning different weights according to the position of the pixel points in the upper body area, the overall motion characteristics of the upper body can be more accurately reflected. Compared with the traditional method of treating all pixel points equally, the weighted method can highlight the influence of the motion of key parts (such as those close to the body center) on the overall characteristics, improving the accuracy and effectiveness of feature extraction.

[0070] When the optical flow vectors of the foreground pixel points in a series of adjacent frames meet the following conditions, it is considered to be part of an action event: the direction of the optical flow vector is within a certain angular range (for example, the angle with the optical flow vector direction of the previous frame is less than θ d , such as θ d = 30°), and the action amplitude is greater than the set amplitude threshold A d ,

[0071] By analyzing consecutive frames and counting the number of times that meet the definition of an action event within a unit time (1 minute), it is the action frequency f d . Set a state machine. When a pixel point that meets the starting condition of the action event is detected, start the state machine to continuously monitor whether the continuous condition of the action event is met in the subsequent frames. If the condition is continuously met within a certain number of frames (such as 10 frames), it is considered a complete action event, and the counter is incremented by 1. At the end of 1 minute, read the value of the counter as the action frequency.

[0072] In the calculation of the action frequency, by more precisely defining the action event, comprehensively considering the direction and amplitude changes of the optical flow vector and its persistence within a certain number of frames, it avoids misjudgment caused by short-term and random small motions, improves the accuracy of action frequency statistics, and thus more accurately reflects the actual motion state of the driver's upper body.

[0073] When the driver is in a fatigued state, the action frequency of the upper body will decrease, the action amplitude will decrease, and the change in the action direction will become more disordered; for example, in the normal driving state, the action frequency may be between 30 - 50 times per minute. When the action frequency is lower than 20 times per minute, it may indicate a fatigued state; define the average value of the normal action amplitude as A n , when the action amplitude is less than 0.7×A n , it is also a sign of fatigue; the standard deviation of the change in the action direction has a certain range in the normal state. When the standard deviation exceeds 1.5 times the normal range, it indicates that the change in the action direction is more disordered, and it may also be a sign of fatigue. By comprehensively analyzing the changes in the upper body characteristics, the fatigue state of the driver can be more accurately judged.

[0074] Example 4

[0075] Please refer to Figure 4, put a protective cover on the steering wheel, with multiple high-precision piezoresistive grip sensors (at least 6) evenly distributed on its circumference, which can accurately sense the pressure exerted by the driver's hand. The sensors collect data at a frequency of 70 Hz, that is, 70 sets of grip data are obtained per second. The weak electrical signals output by the sensors are preprocessed such as amplified and filtered by a high-speed data acquisition circuit to remove noise interference, and then converted into digital signals and transmitted to the data processing unit.

[0076] First, perform linear calibration on the original data collected by the sensors. By pre-calibrating the sensors, obtain the relationship curve between the output electrical signal value V of the sensors and the actual grip force F. Define that this relationship can be approximately expressed as a linear equation F = k s V + b s , where k s and b s are calibration coefficients determined through calibration experiments. For the original electrical signal value V i (i = 1, 2,..., 6 represent different sensors) collected by each sensor, use this calibration equation to calculate the corresponding grip force value F i .

[0077] In order to obtain the overall grip force of the driver holding the steering wheel, considering that sensors at different positions contribute differently to the overall grip force, weighted average is adopted. Sensors near the top and bottom of the steering wheel may be more critical for controlling the direction, and relatively higher weights are assigned. Define that the weight of sensor i is ω i and then the overall grip force is defined as F t .

[0078]

[0079] Under normal driving conditions, the grip force of the driver will remain within a certain range and adjust with the change of driving tasks. When driving straight on a straight road, the grip force is relatively stable and small; when turning or dealing with complex road conditions, the grip force will increase accordingly; when the driver is gradually fatigued, the grip force may become unstable, manifested as large fluctuations in the grip force value in a short time or maintaining an abnormal high or low value for a long time; if the average grip force deviates from the normal range by more than a certain range and lasts for more than 1 minute, it may indicate that the driver is in a fatigued state.

[0080] Define a grip force change event as a process in which the grip force changes significantly in a short time. In order to accurately identify the grip force change event, set two thresholds, namely the rising threshold T r and the falling threshold T f . When the grip force rises rapidly from below T r and exceeds T r , and then drops below Tf When this occurs, it is recognized as a complete grip force change event.

[0081] By continuously monitoring the grip force data in real time, the number of grip force change events occurring within a unit time (1 minute) is statistically counted, and this is the grip force frequency f. w Define a state variable to track the change state of the grip force, such as "rising stage, falling stage, stable stage". When the grip force enters the "rising stage" from the "stable stage" and exceeds T r start counting; when the grip force enters the "falling stage" and is lower than T f complete one count, and reset the state variable to the "stable stage" to prepare for the next event detection. The number of events obtained after 1 minute is the grip force frequency f. w .

[0082] By assigning different weights to sensors at different positions according to their importance in driving operations, it can more accurately reflect the driver's overall control force on the steering wheel. Using weighted fusion can capture the differences in the driver's force on different parts of the steering wheel in different driving scenarios. For example, when turning, more force may be applied to one side of the steering wheel, thus providing richer and more accurate grip force information for fatigue judgment.

[0083] Embodiment Five

[0084] The data processing unit adopts a feature fusion architecture based on a deep neural network, constructs a deep neural network with multiple hidden layers, and takes the preprocessed features of the face, upper body, and hand holding the steering wheel as different input branches respectively.

[0085] Each branch performs feature extraction and transformation through a series of convolutional layers (for image-like features) or fully connected layers (for numerical-like features) to learn feature representations at different levels, and then fuses the features of these different branches in a certain hidden layer using a fusion method based on the attention mechanism. The fused feature F f is used as the input of the decision tree model, defines the driver's fatigue state as L (L = 0 means non-fatigued, L = 1 means fatigued) as the output for training, and the decision tree constructs a tree structure by continuously dividing the feature space. For example, at a certain node, it is divided according to the threshold of a certain feature (such as a certain dimension in the fused feature) so that the class purity of the samples in the sub-nodes after division is higher.

[0086] Based on the output of the decision tree model and combined with the preset fatigue judgment rules, the fatigue state is judged. Define the fatigue probability P p (L = 1), set a fatigue threshold T p (such as T p = 0.6); when P p (L = 1) ≥ Tp When it is determined that the driver is in a fatigued state, that is:

[0087]

[0088] Through the above steps, the determination of the driver's fatigue state is completed from the feature fusion and analysis of the data processing unit to the judgment by the fatigue state judgment unit according to the rules.

[0089] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A driver fatigue state detection system based on computer vision processing, characterized in that: It includes a facial feature extraction unit, which uses a high-definition camera to capture the driver's facial image; The upper body feature monitoring unit collects upper body image data through a high-definition camera and uses a combination of background subtraction technology and optical flow method to accurately extract the upper body movement frequency; The hand grip steering wheel characteristic monitoring unit uses a protective cover to set a grip force sensor on the steering wheel to sense the pressure applied by the driver's hand; The data processing unit adopts a feature fusion architecture based on a deep neural network, constructs a deep neural network with multiple hidden layers, receives data information collected from the facial feature extraction unit, the upper body feature monitoring unit, and the hand-holding steering wheel feature monitoring unit, and performs feature fusion using a fusion method based on an attention mechanism; A fatigue state judgment unit, based on the analysis result of the data processing unit, judges the fatigue state of the driver according to a preset fatigue judgment rule; Before the driver enters the vehicle but starts driving, the high-definition camera first collects a video sequence of duration T0 and then processes each frame of the sequence I t Processing is performed and a Gaussian mixture model is used to construct a background model. For each pixel in the image, a mixture model of K Gaussian distributions is established for it in the initial stage. Each Gaussian distribution defines a mean μ k , covariance matrix Σk and weight ω k express, Specifically, for each pixel point P(x, y), at time t, according to the current pixel value I t The matching degree between (x, y) and each Gaussian distribution, update the weight ω k (t), mean μ k (t) and the covariance matrix Σ k (t), the matching degree is output through the Mahalanobis distance D, If D is less than a certain threshold, the pixel is considered to belong to the corresponding Gaussian distribution and its parameters are updated accordingly. After T0 seconds of learning, a stable background model B(x, y) is obtained.

2. A driver fatigue state detection system based on computer vision processing as claimed in claim 1, characterized in that: For a pixel point P(x, y), when calculating its displacement (u, v), a more accurate displacement estimate is obtained by weighted summing of the brightness changes in the spatiotemporal window. The set of pixels in the spatiotemporal window is defined as N(x, y). The improved optical flow calculation equation is: Among them I x ,I y and I t are the partial derivatives of the image brightness I with respect to x, y and t respectively. ω(x', y', t') is the weight of each pixel in the spatiotemporal window. It is dynamically adjusted according to the spatiotemporal distance between the pixel and the current pixel P(x, y) and the consistency of the optical flow of the previous frames. The optical flow vector of each pixel is obtained by solving the above equations.

3. A driver fatigue state detection system based on computer vision processing as claimed in claim 1, characterized in that: The direction of the upper body movement changes. First, we count the direction change Δθ(x,y) between adjacent frames for each pixel in the foreground area. Assuming that at time t and t+1, the directions of the pixel points P(x, y) are defined as θ1(x, y) and θ2(x, y), respectively, then Δθ(x, y) = |θ2(x, y) - θ1(x, y)| (Taking into account the periodicity of the angle, when |θ2-θ1|>π, Δθ(x, y) = 2π-|θ2-θ1|). Take the weighted average of the direction changes of all pixels in the foreground area to get the overall action direction change Δθ t , assign weight ω to each pixel w (x,y), output Δθ t for:

4. A driver fatigue state detection system based on computer vision processing as claimed in claim 1, characterized in that: The grip force sensor uses a high-precision piezoresistive grip force sensor. The raw data collected by the grip force sensor is linearly calibrated. The relationship curve between the sensor output electrical signal value V and the actual grip force F is obtained by calibrating the sensor in advance. The relationship can be approximately expressed as a linear equation F = k s V+b s , where k s and b s is the calibration coefficient determined by the calibration experiment. For each sensor, the original electrical signal value V i , use the calibration equation to calculate the corresponding grip force value F i , In order to obtain the overall grip strength of the driver's hand on the steering wheel, considering the different contributions of sensors at different positions to the overall grip strength, a weighted average is used. Sensors near the top and bottom of the steering wheel may be more critical to control the direction and are given relatively high weights. The weight of sensor i is defined as ω i and The overall grip strength is defined as F t , 5. The driver fatigue state detection system based on computer vision processing as claimed in claim 1, characterized in that: The data processing unit will fused the feature F f As the input of the decision tree model, the driver's fatigue state is defined as L (L = 0 means non-fatigue, L = 1 means fatigue) and is used as the output for training. The fatigue state is judged based on the output of the decision tree model combined with the preset fatigue judgment rules, and the decision tree model output fatigue probability P is defined p (L=1), set a fatigue threshold T p (such as T p =0.6); when P p (L=1)≥T p When the driver is judged to be in a fatigue state, that is: From the feature fusion and analysis of the data processing unit to the fatigue status judgment unit, the driver's fatigue status is judged based on rules.

Citation Information

Patent Citations

  • Fatigue driving detection method

    CN117765684A

  • Driver fatigue detection method and system based on multi-modal feature fusion

    CN119274169A

  • Fatigue driving monitoring and early warning method based on multi-dimensional feature fusion

    CN119495164A

  • Method for assessing driver fatigue

    US20200320319A1

Cited By

  • Driver state detection method and system based on multi-dimensional data

    CN120840624A

  • Driver state detection method and system based on multi-dimensional data

    CN120840624B