Driver fatigue state real-time recognition system and method based on multi-modal deep learning

By introducing illumination data to evaluate eye-tracking reliability and dynamically adjusting multimodal fusion weights, the stability problem of multimodal fatigue detection under varying illumination environments is solved, achieving high accuracy and reliability of fatigue identification under complex illumination conditions.

CN121305532BActive Publication Date: 2026-03-31HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing multimodal fatigue detection methods are easily affected by changes in lighting conditions in complex driving environments, leading to decreased recognition accuracy and reliability, high false judgment rate, and impact on driving safety.

Method used

Ambient lighting data is introduced as a basis for credibility assessment. The eye-tracking credibility coefficient is calculated through a reliability mapping function. The weights of multimodal feature fusion are dynamically adjusted, and the weights are corrected in combination with the attention mechanism to generate brain-eye coordination fatigue value and trigger an early warning.

Benefits of technology

The robustness and reliability of fatigue state identification are improved under complex lighting conditions, the false positive rate is reduced, and the practicality and security of the system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305532B_ABST
    Figure CN121305532B_ABST
Patent Text Reader

Abstract

The application discloses a driver fatigue state real-time recognition system and method based on a multi-modal deep learning, and relates to the technical field of fatigue driving detection; firstly, brain waveforms and eye movement coordinates are subjected to feature extraction, and respective weights are calculated by using an attention mechanism, so as to realize dynamic distribution of different modal features. Then, in-vehicle illumination data are introduced to establish a credibility mapping function, so as to adaptively correct the eye movement weight, thereby effectively reducing the interference of a complex illumination environment on a recognition result. The corrected multi-modal features are subjected to weighted splicing, are mapped into a brain-eye collaborative fatigue value, and are combined with a continuous time length to be judged, so that stable and reliable early warning control is finally realized. The method has the advantages of high fusion precision, strong environmental adaptability and low false alarm rate while guaranteeing real-time performance, and can significantly improve driving safety guarantee capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fatigue driving detection technology, specifically to a real-time driver fatigue state recognition system and method based on multimodal deep learning. Background Technology

[0002] With the continuous development of intelligent driving and road safety, the identification of driver fatigue has gradually become an important research direction for vehicle monitoring systems. Existing methods are mostly based on the analysis of physiological signals or behavioral characteristics, such as using electroencephalogram (EEG) signals to reflect the level of neural activity, using eye movement signals to reflect attention and visual state, and improving the accuracy and real-time performance of detection through multimodal fusion.

[0003] However, in actual driving, the driving environment is often highly uncertain, with factors such as changes in lighting, differences in driving posture, or external disturbances significantly impacting the stability of some modal features. Existing methods generally lack mechanisms for effectively adapting to environmental changes, leading to distorted fusion results under complex conditions, thus affecting the accuracy and reliability of fatigue state identification. Therefore, improving the robustness of multimodal fatigue detection methods in complex environments has become an urgent technical problem to be solved. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a real-time driver fatigue state recognition system and method based on multimodal deep learning.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention discloses a method for real-time identification of driver fatigue state based on multimodal deep learning, comprising the following steps:

[0007] Simultaneously acquire the driver's raw EEG waveform data and raw eye movement coordinate data;

[0008] The original EEG waveform data is processed by time-frequency transformation to obtain EEG feature vectors, and the original eye movement coordinate data is processed by time-series statistical processing to obtain eye movement feature vectors.

[0009] The EEG feature vector and the eye movement feature vector are input into a weight calculation network based on an attention mechanism to calculate the EEG weight and the initial eye movement weight.

[0010] Ambient lighting data inside the vehicle is acquired, and the ambient lighting data is input into a preset reliability mapping function to calculate the eye-tracking reliability coefficient. The reliability mapping function is configured such that the greater the deviation of the ambient lighting data from the preset lighting range, the lower the eye-tracking reliability coefficient.

[0011] The initial eye movement weights are corrected using the eye movement reliability coefficient to obtain the final eye movement weights;

[0012] The EEG feature vector and the eye-tracking feature vector are weighted and concatenated using EEG weights and final eye-tracking weights to obtain a fused feature sequence.

[0013] The fused feature sequence is mapped to a brain-eye coordination fatigue value, and it is determined whether the brain-eye coordination fatigue value exceeds a preset fatigue threshold.

[0014] If the judgment result is yes, then the duration for which the brain-eye coordination fatigue value exceeds the preset fatigue threshold is recorded. When the duration exceeds the predetermined fatigue duration, an early warning control signal is generated and triggered.

[0015] Secondly, this invention discloses a real-time driver fatigue state recognition system based on multimodal deep learning, comprising:

[0016] The multimodal data acquisition module is used to simultaneously acquire the driver's raw EEG waveform data and raw eye movement coordinate data;

[0017] The feature extraction module is used to perform time-frequency transformation processing on the raw EEG waveform data to obtain EEG feature vectors, and to perform time-series statistical processing on the raw eye movement coordinate data to obtain eye movement feature vectors.

[0018] The weight calculation module is used to input the EEG feature vector and the eye movement feature vector into the attention mechanism-based weight calculation network to calculate the EEG weight and the initial eye movement weight.

[0019] The reliability assessment module is used to acquire ambient light data inside the vehicle, input the ambient light data into a preset reliability mapping function, and calculate the eye-tracking reliability coefficient. The reliability mapping function is configured such that the greater the deviation of the ambient light data from the preset light range, the lower the eye-tracking reliability coefficient.

[0020] The weight correction module is used to correct the initial eye movement weights using the eye movement credibility coefficient to obtain the final eye movement weights.

[0021] The feature fusion module is used to weight and concatenate the EEG feature vector and the eye-tracking feature vector using EEG weights and final eye-tracking weights to obtain a fused feature sequence.

[0022] The fatigue index mapping module is used to map the fused feature sequence to a brain-eye coordination fatigue value and determine whether the brain-eye coordination fatigue value exceeds a preset fatigue threshold.

[0023] The early warning execution module is used to record the duration for which the brain-eye coordination fatigue value exceeds a preset fatigue threshold when the judgment result is yes, and to generate and trigger an early warning control signal when the duration exceeds a predetermined fatigue duration.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] 1. A mapping function for illumination data is introduced to calculate the eye movement reliability coefficient, and the initial weight of eye movement is corrected based on the coefficient, thereby improving the reliability of illumination data. This allows the influence ratio of eye movement features in fusion to be automatically reduced when there are large changes in illumination or when the imaging quality deteriorates, thus effectively suppressing the interference of environmental noise on fatigue recognition accuracy.

[0026] 2. By integrating two typical physiological and behavioral data, EEG signals and eye movement signals, it is possible to jointly monitor the driver's neural activity level and visual attention state during driving. Compared with traditional fatigue monitoring methods that rely on a single modality, this solution can not only enhance the ability to capture fatigue state, but also achieve stable recognition under changing driving environment conditions. Attached Figure Description

[0027] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:

[0028] Figure 1 This is a flowchart of the method of the present invention;

[0029] Figure 2 This is a data flow diagram of the present invention;

[0030] Figure 3 This is a system architecture diagram of the present invention. Detailed Implementation

[0031] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0032] In traditional driver fatigue recognition methods, the reliability of multimodal feature fusion is easily affected by fluctuations in ambient lighting. Because the optical sensors in eye-tracking devices are sensitive to lighting conditions, when the ambient light intensity inside the vehicle exceeds a preset operating range, the accuracy of the raw eye-tracking coordinate data acquisition decreases significantly, leading to systematic biases in the eye-tracking feature vectors. If a fixed-weight multimodal fusion strategy is still used in this situation, the weight allocation of EEG and eye-tracking features will mismatch with the reliability of the actual data, thus causing output errors in the fatigue value mapping model.

[0033] For example, when a vehicle passes through a tunnel entrance or exit, or on a road with alternating periods of sunshine and rain, the ambient light intensity changes drastically within a short period. Due to excessively strong or weak lighting, eye-tracking devices experience increased pupil positioning errors, leading to abnormal fluctuations in the temporal statistics of the fixation point coordinates in the raw eye-tracking coordinate data. In this situation, the system still allocates attention weights based on distorted eye-tracking and EEG feature vectors, resulting in the fused feature sequence carrying incorrect multimodal association information. The brain-eye coordination fatigue value generated by the fatigue index mapping module based on this sequence will deviate from the true physiological state, potentially triggering false warnings or missing critical fatigue events.

[0034] If the aforementioned issues are not addressed, the false alarm rate of multimodal fatigue detection systems under dynamic lighting conditions will increase significantly. False triggering of warning signals may interfere with normal driver operation, while missed fatigue detection directly threatens driving safety. Prolonged environmental interference will lead to decreased system reliability, reduced user reliance on fatigue monitoring functions, and ultimately affect the overall effectiveness of the vehicle safety system.

[0035] To address the aforementioned challenges, this application first considers how to dynamically adjust the fusion weights of multimodal features to cope with environmental interference. Traditional methods with fixed weight allocation cannot adapt to eye-tracking data distortion caused by changes in illumination. Therefore, this application attempts to introduce ambient illumination data as a basis for reliability assessment. Analysis of the optical characteristics of eye-tracking devices reveals a non-linear relationship between illumination intensity and pupil positioning accuracy; illumination intensity deviating from the preset working range significantly reduces data reliability. Based on this, this application proposes establishing a mapping relationship between illumination deviation and reliability coefficients, transforming environmental parameters into a basis for weight correction. Simultaneously, to avoid excessive influence of a single environmental factor on the overall fusion strategy, this application employs an attention mechanism to generate initial weights, which are then combined with the reliability coefficients for secondary correction, forming a dynamic fusion mechanism that balances modal relevance and environmental adaptability.

[0036] In this regard, such as Figure 1 As shown, this application proposes a real-time driver fatigue state recognition method based on multimodal deep learning, including the following steps:

[0037] Simultaneous acquisition of the driver's raw EEG waveform data and raw eye movement coordinate data; This means simultaneously collecting the raw data of the driver's EEG signals and eye movement trajectories. Specifically, this can be achieved by using a multi-channel biosignal acquisition device and an eye tracker with a synchronous triggering mechanism to ensure the time alignment of the two types of data, thus providing a foundation for subsequent multimodal feature fusion.

[0038] EEG feature vectors are obtained by performing time-frequency transformation on the raw EEG waveform data, and eye movement feature vectors are obtained by performing time-series statistical processing on the raw eye movement coordinate data. Time-frequency transformation refers to converting the EEG waveform from the time domain to the frequency domain and extracting energy distribution features. This can be achieved using Fast Fourier Transform or Wavelet Transform algorithms to capture fatigue-related rhythmic changes in the EEG signal. Time-series statistical processing involves calculating the distribution characteristics of the eye movement coordinate data over time. Specifically, sliding window statistical methods can be used to calculate fixation dispersion or mean movement velocity to characterize the stability of the eye movement pattern.

[0039] The EEG feature vector and the eye movement feature vector are input into an attention-based weight calculation network to calculate the EEG weight and the initial eye movement weight. The attention-based weight calculation network refers to the automatic learning of the contribution of different modal features using a neural network. Specifically, it can be implemented using a multi-head attention mechanism or a gated attention module to dynamically allocate the fusion weights of EEG and eye movement features.

[0040] Ambient lighting data inside the vehicle is acquired and input into a preset reliability mapping function to calculate the eye-tracking reliability coefficient. The reliability mapping function is configured such that the greater the deviation of the ambient lighting data from the preset lighting range, the lower the eye-tracking reliability coefficient. Ambient lighting data refers to the light intensity information inside the vehicle, which can be acquired in real-time using a light sensor to assess the degree of interference from lighting on eye-tracking features. The reliability mapping function is a mathematical relationship that maps the degree of lighting deviation to the reliability of eye-tracking features; it can be implemented using a Gaussian function or a piecewise linear function to quantify the impact of environmental changes on the reliability of eye-tracking data.

[0041] The initial eye movement weights are corrected using the eye movement confidence coefficient to obtain the final eye movement weights. Correcting the initial eye movement weights refers to adjusting the fusion ratio of eye movement features based on the confidence coefficient. This can be achieved using linear scaling or nonlinear attenuation algorithms to reduce eye movement feature errors caused by illumination interference.

[0042] The EEG feature vector and eye-tracking feature vector are weighted and concatenated using EEG weights and final eye-tracking weights to obtain a fused feature sequence. Specifically, the EEG weights are multiplied by the EEG feature vector to obtain a weighted EEG vector, and the final eye-tracking weights are multiplied by the eye-tracking feature vector to obtain a weighted eye-tracking vector. The weighted EEG vector and the weighted eye-tracking vector are concatenated to form a one-dimensional fused feature vector. The fused feature vectors obtained from multiple consecutive time windows are arranged in chronological order to form the fused feature sequence.

[0043] The fused feature sequence is mapped to a brain-eye coordinated fatigue value, and it is determined whether the brain-eye coordinated fatigue value exceeds a preset fatigue threshold. The brain-eye coordinated fatigue value refers to a quantitative index of fatigue state generated by combining EEG and eye movement features. Specifically, it can be implemented through logistic regression or neural network classification models to determine whether the driver is in a fatigued state.

[0044] If the judgment result is yes, the duration for which the brain-eye coordination fatigue value exceeds the preset fatigue threshold is recorded. When the duration exceeds the predetermined fatigue duration, an early warning control signal is generated and triggered.

[0045] The core innovation of this application lies in dynamically evaluating the credibility of eye-movement features by introducing ambient lighting data, and making real-time corrections to the multimodal fusion weights based on the credibility coefficient, thereby improving the robustness of fatigue state recognition under complex lighting conditions.

[0046] like Figure 2 The diagram shown is a data flow chart of this application; as a preferred embodiment, the solution of this application is specifically implemented as follows:

[0047] First, raw EEG waveform data and raw eye movement coordinate data of the driver are simultaneously acquired using an EEG acquisition device and an eye-tracking device. The EEG acquisition device can be a dry electrode EEG cap, with a sampling frequency set to 1000Hz. The eye-tracking device can be an infrared camera, with a sampling frequency set to 60Hz.

[0048] The raw EEG waveform data is processed using time-frequency transformation, specifically short-time Fourier transform, with a time window length of 1 second, to obtain the EEG feature vector. The raw eye movement coordinate data is then processed using temporal statistical analysis, calculating the mean and variance of the fixation point coordinates within each 2-second time window to obtain the eye movement feature vector.

[0049] The EEG feature vectors and eye movement feature vectors are input into a pre-trained attention-based weight calculation network. This network consists of multiple fully connected layers and uses a softmax activation function to output the EEG weights and initial eye movement weights.

[0050] Ambient lighting data is acquired using an in-vehicle light sensor. This data is then input into a preset reliability mapping function to calculate the eye-tracking reliability coefficient. The reliability mapping function can be a Gaussian function; its output value decreases rapidly when the ambient light intensity deviates from a preset range.

[0051] The initial eye-tracking weights are corrected using the eye-tracking credibility coefficient to obtain the final eye-tracking weights.

[0052] The EEG feature vector and eye movement feature vector are weighted and concatenated using EEG weights and final eye movement weights to obtain a fused feature sequence, which is then mapped to brain-eye coordination fatigue value.

[0053] The system determines whether the brain-eye coordination fatigue value exceeds a preset fatigue threshold. If it does, a timer is started to record the duration. When the duration exceeds the predetermined fatigue duration, a warning control signal is generated and the vehicle's warning system is triggered, such as an audible and visual alarm or seat vibration.

[0054] The above-described scheme effectively addresses the stability issue of traditional multimodal fatigue detection methods under complex lighting conditions. By introducing ambient lighting data as a reliability assessment criterion and dynamically adjusting the fusion weights of multimodal features, the system can adapt to eye-tracking data distortion caused by changes in lighting. An attention mechanism combined with a secondary correction strategy for the reliability coefficient achieves dynamic fusion that balances modal correlation and environmental adaptability. This method effectively reduces the impact of environmental interference on fatigue detection accuracy and improves the system's reliability under dynamic lighting conditions. Furthermore, by recording the duration of fatigue states to trigger warnings, false alarms caused by instantaneous fluctuations are avoided, further enhancing the system's practicality.

[0055] This application further proposes the following calculation process for the eye-tracking reliability coefficient:

[0056] Set preset lighting range The range is set based on the lighting conditions under which the vehicle camera module can stably and clearly capture images of the driver's eyes under its nominal performance. and The specific value can be determined experimentally, for example... It can be set to 50 Lux (to avoid excessive light at night or in tunnels). It can be set to 10,000 Lux (to avoid overexposure scenarios such as direct sunlight at midday). This range ensures the basic reliability of eye-tracking data.

[0057] When the ambient light data In During the internal time, the eye movement reliability coefficient At this point, the lighting conditions are ideal, and the quality of the eye-tracking data is at its highest; therefore, the eye-tracking reliability coefficient is directly defined. .

[0058] when At that time, I recorded Light deviation ;

[0059] when At that time, I recorded Light deviation ;

[0060] Ambient lighting data Not in During this process, the reliability of the eye-tracking data needs to be discounted, and the obtained illumination deviation needs to be calculated. It quantifies the absolute value of how much the current lighting conditions deviate from the ideal.

[0061] Light deviation Substituting into the preset reliability mapping function, the eye-tracking reliability coefficient is obtained. The calculation formula is as follows:

[0062] ,in This is a width parameter that controls the width of the function curve; its value is greater than 0.

[0063] The reliability mapping function adopts an exponential decay form, in which... When the value is 0, it reaches its maximum value of 1, and then... The value increases and then smoothly and monotonically decreases to 0. This mathematical property well simulates the physical process of "data reliability gradually decreasing as environmental conditions deteriorate". Its change curve is continuous and differentiable, and this property can better reflect the gradual relationship in the real world.

[0064] And the function uses the width parameter Control the credibility coefficient as Sensitivity to change. For example, when When the value is large, the function curve is relatively flat, meaning that the system has a higher tolerance for illumination deviations, even if the illumination changes significantly. It will not drop sharply; when When the value is small, the curve is steep, meaning the system is very sensitive to changes in illumination; even slight deviations in illumination can cause problems. It decreased rapidly.

[0065] Specifically, eye-tracking data exhibits the highest reliability when ambient light levels are within a preset range, with a confidence coefficient remaining at 1; when the light level deviates from the range, the reliability is further reduced. Calculate the deviation and dynamically adjust the confidence coefficient using an exponential function. By adjusting... The value can be adapted to the reliability assessment needs under different lighting fluctuation scenarios. For example, in scenarios with frequent lighting fluctuations, by increasing... By reducing the sensitivity of the reliability coefficient to instantaneous deviations and avoiding instability in the fusion results due to frequent weight adjustments, this method can accurately reflect changes in the reliability of eye-tracking data under complex lighting conditions, thus improving the robustness of fatigue state recognition.

[0066] For example: preset lighting zone Set as When ambient light data In During the internal time, the eye movement reliability coefficient Set to 1.

[0067] When ambient light data Below At that time, calculate the illumination deviation. for .

[0068] When ambient light data Higher than At that time, calculate the illumination deviation. for .

[0069] Light deviation Substituting into the preset reliability mapping function, the eye-tracking reliability coefficient is obtained. ;

[0070] set up 50 lux The value is 100 lux. Substituting this into the formula, we obtain:

[0071] .

[0072] Through the above technical solution, this application achieves accurate calculation of the eye-tracking reliability coefficient. By setting a preset illumination range and introducing a reliability mapping function, the reliability of eye-tracking data can be dynamically adjusted according to the actual ambient illumination conditions. When the illumination conditions are within the ideal range, the eye-tracking data is assigned the highest reliability; while when the illumination conditions deviate from the ideal range, the reliability of the eye-tracking data decreases as the degree of deviation increases. This mechanism effectively improves the accuracy and reliability of eye-tracking data under different illumination environments, thereby enhancing the environmental adaptability and detection stability of the entire fatigue detection system.

[0073] This application further proposes a width parameter. The process of obtaining the value is as follows:

[0074] Obtain the preset illumination range And calculate the height value of the preset illumination zone. and will Half of the value is used as the base width parameter :

[0075] ;

[0076] height value Represents the range span of ideal lighting conditions, the reference width parameter Use height value Half of it provides a static baseline value, which ensures a reasonable and reliable default configuration even in the absence of additional information.

[0077] During real-time driving of the vehicle, recent ambient light data is periodically acquired, and the standard deviation of ambient light is calculated. Standard deviation of ambient light As a statistical measure, it can quantify the degree of fluctuation in recent ambient lighting conditions; when this value increases, it indicates drastic changes in lighting (e.g., frequently moving through the shade of trees or alternating between entering and exiting tunnels); conversely, a decrease in this value indicates a decrease in lighting conditions. A decrease in light level indicates stable ambient lighting conditions (e.g., consistently clear skies or driving on a highway at night).

[0078] Standard deviation of ambient light With the preset fluctuation threshold Comparison; preset fluctuation threshold This is a value used to define whether lighting conditions are stable or fluctuating; its value is derived from statistical analysis of a large amount of experimental data, for example, set to... .

[0079] Based on the comparison results, the width parameter is determined. :

[0080] like Then let This indicates that the recent lighting environment is stable. While the instantaneous value of the illumination may deviate from the range, its fluctuation is low, and these deviations are likely slow and continuous (such as at dusk). Therefore, the system does not require additional adjustments and can directly use the reference parameters, setting... In this configuration, the reliability mapping function is affected by illumination deviation. It is highly sensitive and can effectively identify and downweight low-quality eye-tracking data caused by continuous low or high light.

[0081] like Then let This indicates a recent drastic fluctuation in lighting conditions. In such cases, the momentary deviation in lighting is likely brief and sudden (e.g., a fleeting shadow or the high beams of an oncoming vehicle). To avoid these momentary disturbances affecting the reliability of eye-tracking... Rapid changes can lead to misjudgments; an adaptive mechanism can help address this. Parameters Proportional to This indicates that the greater the fluctuation, the larger the value of this item, and it dominates the overall trend. Growth; Parameter Items This is used to ensure the adjusted It will not fall below the baseline value to ensure baseline performance.

[0082] in, The value is configured by staff based on historical driving data and experimental results to control the impact of volatility. The intensity of the impact. For example, it can be set. .

[0083] For example, when 50 lux Set to 0.5. 60 lux and At 40 lux:

[0084] .

[0085] This adjustment method affects the width parameter. As the degree of illumination fluctuation increases linearly, the coverage of the reliability mapping function is expanded, the excessive suppression of eye-tracking reliability coefficient by illumination abrupt changes is reduced, thereby improving the stability of weight correction.

[0086] Through the above technical solution, this application achieves the width parameter The adaptive adjustment improves the reliability mapping function's adaptability to changes in ambient lighting. Consequently, the calculation of the eye-tracking reliability coefficient is more accurate, thus enhancing the correction effect of eye-tracking feature weights and strengthening the robustness of the multimodal fatigue detection method under complex lighting conditions.

[0087] This application further proposes a method for adjusting the final eye-tracking weight as follows:

[0088] Obtaining initial eye-tracking weights Eye movement reliability coefficient And calculate the correction deviation. :

[0089] ,in As a regulating factor;

[0090] The eye-tracking reliability coefficient is the primary modifier. The smaller the value, the worse the current lighting conditions, meaning the weaker the reliability of the eye-tracking data. The larger the value, the stronger the "force" or "motivation" for correction, and the greater the amount of correction required. The larger. The baseline value is used for correction and represents the importance of the eye movement feature itself, regardless of external lighting conditions. Correction should be based on its initial importance. It is a configurable parameter between 0 and 1, which acts as an adjustable damper in the formula; The smaller the value of , the more conservative the correction process. The more moderate the adjustment, the better; conversely, A larger value indicates a more aggressive correction; its specific value can be optimized and determined by the staff based on a large amount of experimental data, for example, set to 0.6 to achieve a balance between sensitivity and stability. Correction bias amount It is a non-negative value. It comprehensively reflects the amount that needs to be "deducted" from the initial weights due to poor ambient lighting. The worse the environment (…), the more weights are deducted. The lower the initial weight, the higher the initial weight. The larger the value, the smaller the damping. The larger the value, the greater the amount that needs to be deducted.

[0091] Based on the correction deviation Initial weights for eye movements Perform linear adjustments to obtain the final eye-tracking weights. :

[0092] .

[0093] The linear adjustment process uses subtraction to ensure that the direction of weight adjustment is consistent with the decreasing trend of credibility, thus avoiding abrupt errors caused by nonlinear operations.

[0094] For example:

[0095] when , , hour,

[0096] ;

[0097] .

[0098] In this way, the final weight of eye movements is appropriately reduced to reflect the impact of current ambient lighting conditions on the reliability of eye movement data.

[0099] Through the above technical solution, this application achieves dynamic adjustment of eye-tracking weights. Therefore, the system can adaptively adjust the contribution of eye-tracking features to fatigue state recognition based on real-time ambient lighting conditions, improving the accuracy and reliability of multimodal fusion results under complex lighting environments. Furthermore, by introducing an adjustment factor β, this solution provides a more refined control method for weight adjustment, enabling the system to more flexibly respond to different degrees of lighting changes, thereby enhancing the environmental adaptability of fatigue detection.

[0100] This application further proposes a regulatory factor. The process of obtaining the value is as follows:

[0101] Obtain the variance of the eye movement feature vector within a preset time window, and calculate its square root as the amplitude of eye movement fluctuations. ;

[0102] Eye movement fluctuation amplitude As a statistical measure, it quantifies the degree of dispersion or rate of change of the driver's gaze over a recent period. A higher value indicates that the driver is frequently and significantly shifting their gaze (e.g., checking the rearview mirror, side window, or multiple road signs), which is an active and conscious visual search behavior. Conversely, a lower value indicates a lower rate of change. The smaller the value, the more it indicates that the driver's gaze is fixed and changes very little. This could be a "staring" phenomenon caused by fatigue, or it could be a normal state with extremely low traffic volume.

[0103] Based on eye movement fluctuation amplitude With the preset benchmark fluctuation range The ratio of the two factors determines the adjustment factor. The calculation formula is as follows:

[0104] ,in, This is the preset minimum damping value.

[0105] The experience value set by staff represents the typical level of eye movement fluctuations in a normal, awake state, and can be determined by calculating its standard deviation through the collection of a large amount of eye movement data during normal driving.

[0106] when When the value is very small, the parameter item The value of is also small, which can be understood as approaching 0. At this time, The value is close to 1. This indicates that the eye movement pattern is abnormally stable, and the correction intensity should be increased, that is, let... Take a large value so that the correction deviation is This increases the weight of eye-tracking data, thus significantly reducing its final weight. This is because, in this situation, eye-tracking data is neither reliable (poor lighting) nor informative (small fluctuations), making it the least valuable reference.

[0107] when When the value increases, the parameter item The value also increases accordingly, at this time The change monotonically decreases from 1. This indicates that the driver is actively observing, and even in poor lighting conditions, their eye-tracking data contains rich information about driving behavior. At this point, the correction intensity should be suppressed, even... Take a small value so that the correction deviation is The weighting is reduced to preserve a higher weight for eye-tracking data. This is because the system needs this information to determine the driver's activity level and avoid misinterpreting normal visual searches as fatigue.

[0108] For example, a preset time window of 10 seconds is selected. Within this time window, the variance of the eye movement feature vector is calculated, and the square root of the variance is taken to obtain the amplitude of eye movement fluctuations. It is 0.3;

[0109] Set the benchmark fluctuation range It is 0.5. Set to 0.1;

[0110] Substituting into the formula, we get:

[0111] .

[0112] Therefore, the regulatory factor It was determined to be 0.549. This value will be used in the subsequent eye-tracking weight correction process to adapt to the fluctuations in eye-tracking characteristics under the current driving environment.

[0113] Through the above technical solution, this application can dynamically adjust the modulation factor according to the real-time fluctuations of eye movement characteristics. This adaptive mechanism improves the fatigue detection system's adaptability to complex driving environments and enhances the stability and reliability of multimodal fusion results.

[0114] This application further proposes mapping the fused feature sequence to brain-eye coordinated fatigue values, including:

[0115] The fused feature sequence is normalized within a preset time window to obtain a standardized feature sequence. The purpose of this process is to eliminate the differences in dimensions and numerical ranges between different feature dimensions, transform them to a unified scale, and form a standardized feature sequence. The specific normalization method can be Z-score standardization or Min-Max standardization.

[0116] Calculate the standardized feature sequence The weighted sum yields the initial fatigue value. :

[0117] ;

[0118] in, Represents the preset feature weight coefficients, and satisfies , For standardized feature sequences The number of fusion features, ;

[0119] As a pre-defined fixed vector, each coefficient represents the corresponding fused feature. The contribution and importance of fatigue are determined by training a machine learning model (such as linear regression or support vector machine) on a large amount of historical driving data, and the optimal value is obtained, satisfying the following conditions: The constraints make this It can be understood as a weighted average of the contributions of all features.

[0120] As a continuous scalar, its magnitude directly reflects the initial fatigue tendency calculated based on the current multimodal fusion characteristics. The higher the value, the higher the likelihood of fatigue.

[0121] For the initial fatigue value The brain-eye coordination fatigue value was obtained by amplification. :

[0122] , This is the center offset parameter.

[0123] This amplification process is non-linear, and a Sigmoid function is introduced. Enlarge it.

[0124] Center offset parameter The fatigue decision boundary is defined. It is actually an offset used to adjust the center point of the sigmoid function. When When the value is 0, A value of 0.5 is precisely the critical point between alertness and fatigue. By adjusting... You can set the strictness of the system's fatigue judgment, for example, increase the... This will make the system more "strict," requiring stronger evidence of fatigue to output a high value.

[0125] Preset sensitivity adjustment coefficient Controlled the function's center offset parameter The slope (steepness) of the surrounding terrain. The larger the value, the steeper the S-curve, meaning... Center offset parameter Tiny changes in the vicinity will be magnified to The abrupt changes in temperature make the system highly sensitive to early signs of fatigue. Conversely, The smaller the value, the flatter the curve, and the more conservative and smooth the system's judgment.

[0126] Through the above technical solution, this application achieves an effective mapping from fusion features to fatigue values. Standardization eliminates dimensional differences between different features, weighted summation reflects the importance of each feature, and nonlinear mapping enhances the discriminative power of fatigue values. This multi-step mapping method improves the accuracy and reliability of fatigue state identification, providing a more reliable basis for subsequent early warning judgments.

[0127] This application further proposes the following process for obtaining EEG feature vectors:

[0128] The raw EEG waveform data is divided according to a preset first time window;

[0129] Fast Fourier Transform was performed on the raw EEG waveform data within each time window, and its power spectral density was calculated.

[0130] The power spectral density values ​​of EEG were extracted in the frequency bands of 4-7Hz and 8-13Hz respectively. The ratio of the power spectral density values ​​of the 4-7Hz band and the 8-13Hz band within each time window was calculated. The ratio of multiple consecutive time windows was arranged in chronological order to form the EEG feature vector.

[0131] The preset first time window length can be set to 2 to 5 seconds to adapt to the physiological rhythm changes of EEG signals. The Fast Fourier Transform uses an overlapping window processing method, with the window overlap rate controlled between 30% and 50%. During power spectral density calculation, the Hanning window function is used to suppress spectral leakage. The 4–7 Hz frequency band corresponds to… Wave activity, corresponding to the 8-13Hz frequency band The power ratio of wave activity and its components was found to be significantly correlated with fatigue level. The ratio sequence of continuous time windows constitutes a two-dimensional temporal feature vector, with the dimension consistent with the number of time windows.

[0132] Specifically, the raw EEG data is first segmented into equal-length segments. Each segment undergoes windowing and frequency domain transformation to obtain a precise spectral distribution. This is achieved by extracting... Waves and The power ratio of the waves effectively captures the typical characteristic of the brain under fatigue: increased low-frequency rhythms and decreased high-frequency rhythms. The ratio sequence of continuous time windows forms a dynamic trajectory that can reflect the gradual process of fatigue.

[0133] For example, during driving, when Waves and A sustained wavy ratio exceeding 1.5 with an upward trend indicates that the driver has entered a state of mild fatigue. This feature vector construction method enhances the sensitivity of EEG features to fatigue states by quantifying the energy change relationship in specific frequency bands. At the same time, the temporal arrangement structure preserves the dynamic information of fatigue development, providing highly discriminative input features for subsequent multimodal fusion.

[0134] Through the above technical solution, this application achieves effective feature extraction of electroencephalogram (EEG) signals. The resulting EEG feature vector can reflect the dynamic changes in the driver's brain activity, providing reliable input data for subsequent fatigue state identification. Furthermore, by calculating the power spectral density ratio of theta waves to alpha waves, the EEG feature changes during driver fatigue can be effectively captured, improving the feature discrimination and sensitivity. In addition, constructing feature vectors using sequence data from multiple time windows helps capture the temporal evolution characteristics of fatigue states, enhancing the stability and reliability of identification.

[0135] This application further proposes the following process for obtaining eye-tracking feature vectors:

[0136] The raw eye-tracking coordinate data is divided according to a preset second time window;

[0137] Calculate the dispersion statistics of the eye-tracking fixation point coordinates on the two-dimensional plane within each time window to form an eye-tracking feature vector.

[0138] The length of the second time window is set to an adjustable parameter, dynamically adjusted according to the driving scenario. For example, a longer window is used in highway scenarios to smooth instantaneous fluctuations, while a shorter window is used in urban road scenarios to capture rapid changes. The dispersion statistic is calculated by determining the standard deviation or variance of the gaze point coordinates. Furthermore, the trace or determinant of the covariance matrix can be combined to characterize the two-dimensional spatial distribution characteristics. During the calculation of the dispersion statistic, outlier coordinate points exceeding the preset confidence interval are removed to improve the robustness of the statistical results.

[0139] Within the second time window, continuously acquired eye-tracking coordinate data are arranged chronologically and processed in segments using a sliding window mechanism. For the set of coordinate points within each window, the standard deviations in the horizontal and vertical directions are first calculated, and then the square root of the sum of their squares is used as a two-dimensional dispersion index. This index effectively characterizes the concentration of the driver's gaze within the window time; higher dispersion indicates more frequent gaze drift, which is associated with visual attention distraction under fatigue. By arranging the dispersion statistics of multiple windows chronologically, a time-continuous eye-tracking feature vector is formed. This vector reflects the dynamic changes in the driver's visual state, while statistical processing suppresses the influence of instantaneous noise.

[0140] Through the above technical solution, this application can effectively characterize the dynamic changes of driver's visual attention. By quantifying the spatial dispersion of the gaze point distribution, it overcomes the local deviation of eye movement coordinate data caused by sudden changes in light intensity, improves the stability of eye movement features in complex driving environments, and thus enhances the anti-interference ability of fatigue state detection.

[0141] This application further proposes the following computation process for a weight calculation network based on an attention mechanism:

[0142] The EEG feature vector and the eye-tracking feature vector are concatenated to form a multimodal fusion feature vector;

[0143] The multimodal fusion feature vector is input into the multilayer perceptron model to calculate the attention score vector;

[0144] A multilayer perceptron model, as a learnable feature transformer and evaluator, consists of an input layer, at least one hidden layer, and an output layer. Its computation process is as follows:

[0145] Forward propagation computation: The multimodal fusion feature vector is used as the input layer node value and transmitted to the hidden layer; the value of each node in the hidden layer is obtained by multiplying all its input node values ​​by the corresponding weight parameter, summing them, adding a bias parameter, and finally passing it through a non-linear activation function (such as the ReLU function);

[0146] Output layer computation: The output values ​​of the hidden layer nodes are transmitted to the output layer; the output layer contains two nodes, whose node values ​​are obtained by multiplying the hidden layer node values ​​by the corresponding weight parameters, summing the results, and adding the bias parameters to form the original attention score vector;

[0147] This model, through the aforementioned hierarchical structure and parameters, achieves nonlinear transformation and interactive computation of input features. Its internal weight and bias parameters are obtained by using historical driving data and optimizing through backpropagation during the training phase. The final output attention score vector has a dimension of 2, and its two components represent the "importance" or "score" of the EEG modality and eye-tracking modality, respectively, as initially determined by the model after computation.

[0148] The attention score vector is normalized so that the sum of all components is 1. A Softmax function is then used for normalization, converting each component's score into a probability value between 0 and 1, ensuring that the sum of any two components is 1. This makes the weights of the two modalities comparable and complementary. If one modality has a higher weight, the other modality's weight will inevitably be lower, forcing the model to make a clear trade-off and allocation between the two.

[0149] The first component of the normalized attention score vector is used as the EEG weight, and the second component is used as the initial eye-tracking weight. The EEG weight represents the contribution of EEG features to the fatigue recognition task, based on all available information; the initial eye-tracking weight represents the contribution of eye-tracking features to the information they contain, without considering the quality of external data (such as lighting).

[0150] Through the above technical solution, this application achieves dynamic adaptive allocation of multimodal feature weights. It automatically learns the correlation between EEG and eye-tracking modalities using a multilayer perceptron model, generating interpretable weight allocation results based on the attention mechanism. This solution effectively alleviates the insufficient adaptability of traditional static weight allocation methods in complex driving scenarios, enabling the contribution of different modal features to be automatically adjusted according to real-time data characteristics, thereby improving the environmental robustness of the fatigue state recognition system.

[0151] like Figure 3 The diagram shown is a system architecture diagram of this application. This application further proposes a real-time driver fatigue state recognition system based on multimodal deep learning, including a multimodal data acquisition module, a feature extraction module, a weight calculation module, a credibility evaluation module, a weight correction module, a feature fusion module, a fatigue index mapping module, and a warning execution module.

[0152] The multimodal data acquisition module synchronously acquires the driver's physiological signals through an EEG sensor and an eye-tracking device;

[0153] The feature extraction module performs time-frequency transformation on the EEG waveform to generate a frequency band power ratio feature vector, and performs time window statistics on the eye movement coordinates to generate a dispersion feature vector;

[0154] The weight calculation module uses a multilayer perceptron model to calculate attention scores on the spliced ​​multimodal features and outputs EEG weights and initial eye movement weights.

[0155] The credibility assessment module acquires ambient brightness data through a light sensor, calculates the eye-tracking credibility coefficient based on a Gaussian function, and automatically reduces the coefficient value when the ambient brightness deviates from the preset range.

[0156] The weight correction module dynamically adjusts the initial weights based on the confidence coefficient and the eye movement fluctuation amplitude to generate the final eye movement weights.

[0157] The feature fusion module performs weighted concatenation of dual-modal features to form a fused sequence;

[0158] The fatigue index mapping module maps the fused sequence to a fatigue index through normalization and logical functions;

[0159] The early warning execution module triggers an early warning signal based on the duration of fatigue index exceeding the threshold.

[0160] This application achieves accurate identification of driver fatigue under dynamic lighting conditions. By quantifying the impact of ambient light on the reliability of eye-tracking data in real time and dynamically adjusting the fusion weights of multimodal features, the interference of eye-tracking data noise on fatigue judgment under low-light conditions is effectively suppressed, improving the anti-interference capability and result consistency of the fatigue detection system in complex driving scenarios.

[0161] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. A method for real-time recognition of driver fatigue state based on multi-modal deep learning, characterized in that: The method comprises the following steps: synchronously acquiring raw electroencephalogram waveform data and raw eye movement coordinate data of a driver; performing time-frequency transformation on the raw electroencephalogram waveform data to obtain an electroencephalogram feature vector and performing time series statistical processing on the raw eye movement coordinate data to obtain an eye movement feature vector; inputting the electroencephalogram feature vector and the eye movement feature vector into a weight calculation network based on an attention mechanism to calculate an electroencephalogram weight and an initial eye movement weight; acquiring environmental illumination data inside a vehicle, inputting the environmental illumination data into a preset reliability mapping function, and calculating an eye movement reliability coefficient; the reliability mapping function is configured such that the greater the degree of deviation of the environmental illumination data from a preset illumination interval, the lower the eye movement reliability coefficient; correcting the initial eye movement weight by the eye movement reliability coefficient to obtain a final eye movement weight; performing weighted splicing on the electroencephalogram feature vector and the eye movement feature vector by using the electroencephalogram weight and the final eye movement weight to obtain a fusion feature sequence; mapping the fusion feature sequence into a brain-eye collaborative fatigue value and judging whether the brain-eye collaborative fatigue value exceeds a preset fatigue threshold; if the result of the judgment is yes, recording the duration for which the brain-eye collaborative fatigue value exceeds the preset fatigue threshold, and when the duration exceeds a predetermined fatigue duration, generating and triggering a warning control signal; the calculation process of the eye movement reliability coefficient is as follows: Setting preset light interval ; When the ambient lighting data is within the eye movement plausibility coefficient When the record for the degree of illumination deviation ; When the record for the degree of illumination deviation ; Deviation of light exposure The eye movement reliability coefficient is obtained by substituting the preset reliability mapping function The calculation formula is: wherein is a width parameter for the control function curve width, having a value greater than 0; The width parameter The value process is as follows: Obtaining a preset light interval , and calculating a height value of the preset light interval , and taking half of the value of as a reference width parameter : ; During real-time driving of the vehicle, periodically acquiring recent ambient light data and calculating an ambient light standard deviation ; Standard deviation of ambient light With the preset fluctuation threshold Compare; determining the width parameter based on the comparison result : If then let ; If then let ; wherein, is a scaling factor configured by the staff according to historical experience values, and .

2. The method of claim 1, wherein the method is based on multi-modal deep learning. the adjustment mode of the final eye movement weight is as follows: obtaining eye movement initial weights , eye movement confidence coefficient , and calculating a correction bias : wherein is a tuning factor, which takes a value in the range 0-1; based on the modified bias to the eye movement initial weight linear adjustment to obtain the eye movement final weight : 。 3. The method of claim 2, wherein the method is based on multi-modal deep learning. The adjustment factor The value process of the adjustment factor is: obtain the variance of the eye movement feature vector in the preset time window, and calculate the square root as the eye movement fluctuation amplitude ; According to the eye movement fluctuation amplitude The ratio of the preset reference fluctuation amplitude Determine the adjustment factor The formula is: wherein, is a preset minimum damping value. 4.The method of claim 1, wherein the method further comprises: determining a fatigue state of the driver in real time based on the multi-modal deep learning. the process of mapping the fusion feature sequence into the brain-eye collaborative fatigue value comprises: normalizing the fusion feature sequence in a preset time window to obtain a standardized feature sequence ; calculating a weighted sum of the normalized feature sequences to obtain an initial fatigue value : ; wherein, represents a preset feature weight coefficient, and satisfies , is a normalized feature sequence the number of fused features in ; For the initial fatigue value The brain-eye coordination fatigue value was obtained by amplification. : wherein is a preset sensitivity adjustment coefficient, is a center offset parameter. 5.The method of claim 1, wherein the method further comprises: determining a fatigue state of the driver in real time based on the multi-modal deep learning. the acquisition process of the electroencephalogram feature vector is as follows: dividing the raw electroencephalogram waveform data according to a preset first time window; performing fast Fourier transformation on the raw electroencephalogram waveform data in each time window and calculating the power spectral density thereof; extracting the power spectral density values of the electroencephalogram wave in the frequency bands of 4-7 Hz and 8-13 Hz respectively, calculating the power spectral density value ratio of the frequency bands of 4-7 Hz and 8-13 Hz in each time window, and arranging the ratio of a plurality of continuous time windows in time sequence to form the electroencephalogram feature vector.

6. The method of claim 1, wherein the method is based on multi-modal deep learning. the acquisition process of the eye movement feature vector is as follows: dividing the raw eye movement coordinate data according to a preset second time window; calculating the dispersion statistical quantity of the eye movement fixation point coordinates in a two-dimensional plane in each time window to form the eye movement feature vector.

7. The method of claim 1, wherein the method is based on multi-modal deep learning. the calculation process of the weight calculation network based on the attention mechanism is as follows: splicing the electroencephalogram feature vector and the eye movement feature vector to form a multi-modal fusion feature vector; inputting the multi-modal fusion feature vector into a multi-layer perception machine model to calculate an attention score vector; performing normalization processing on the attention score vector so that the sum of each component is 1; taking the first component of the normalized attention score vector as the electroencephalogram weight and the second component as the initial eye movement weight.

8. A real-time driver fatigue state recognition system based on multi-modal deep learning, characterized in that: comprise: a multi-modal data acquisition module for synchronously acquiring raw electroencephalogram waveform data and raw eye movement coordinate data of a driver; a feature extraction module for performing time-frequency transformation on the raw electroencephalogram waveform data to obtain an electroencephalogram feature vector and performing time series statistical processing on the raw eye movement coordinate data to obtain an eye movement feature vector; The weight calculation module is configured to input the electroencephalogram feature vector and the eye movement feature vector into a weight calculation network based on an attention mechanism, and calculate an electroencephalogram weight and an eye movement initial weight. The reliability evaluation module is configured to acquire environmental illumination data in a vehicle interior, input the environmental illumination data into a preset reliability mapping function, and calculate an eye movement reliability coefficient; the reliability mapping function is configured such that the greater the degree of deviation of the environmental illumination data from a preset illumination interval, the lower the eye movement reliability coefficient. The weight correction module is configured to correct the eye movement initial weight by using the eye movement reliability coefficient to obtain an eye movement final weight. The feature fusion module is configured to perform weighted splicing on the electroencephalogram feature vector and the eye movement feature vector by using the electroencephalogram weight and the eye movement final weight, and obtain a fusion feature sequence. The fatigue index mapping module is configured to map the fusion feature sequence into a brain-eye collaborative fatigue value, and determine whether the brain-eye collaborative fatigue value exceeds a preset fatigue threshold. The early warning execution module is configured to, when the determination result is yes, record a duration for which the brain-eye collaborative fatigue value exceeds the preset fatigue threshold, and when the duration exceeds a predetermined fatigue duration, generate and trigger an early warning control signal. The calculation process of the eye movement reliability coefficient is as follows: Setting a preset light interval ; when the ambient lighting data is within an eye movement plausibility coefficient When the record for the degree of illumination deviation ; When the lighting deviation degree ; The light deviation degree is calculated The eye movement reliability coefficient is obtained by substituting the reliability mapping function The calculation formula is: wherein is a width parameter for the control function curve width, having a value greater than 0; The width parameter The value process is as follows: Obtaining a preset light interval , and calculating a height value of the preset light interval , and taking half of the value of as a reference width parameter : ; During real-time driving of the vehicle, periodically acquiring recent ambient light data and calculating an ambient light standard deviation ; comparing the standard deviation of the ambient light illumination to a preset fluctuation threshold is performed;​ determining the width parameter based on the comparison result : If then let ; If then let ; wherein, is a scaling factor configured by the staff according to historical experience values, and .

Citation Information

Patent Citations

  • Fatigue driving monitoring method and device, storage medium and electronic equipment

    CN109840510A

  • Fatigue monitoring method based on multi-modal data fusion

    CN120154338A