A driving fatigue monitoring auxiliary system based on Beidou satellite positioning and multi-modal data fusion technology

The driver fatigue monitoring system, which integrates BeiDou satellite positioning and multimodal data fusion technology, solves the problems of insufficient accuracy and robustness in existing technologies, achieves efficient fatigue state identification and proactive early warning, and improves the real-time performance and accuracy of detection.

CN121287147BActive Publication Date: 2026-05-12HUNAN AUTOMOTIVE ENG VOCATIONAL COLLEGE +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN AUTOMOTIVE ENG VOCATIONAL COLLEGE
Filing Date
2025-11-18
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing driver fatigue monitoring technologies suffer from insufficient accuracy and poor robustness, and lack a multimodal data fusion mechanism, resulting in limited real-time detection and accuracy, and making it impossible to achieve proactive early warning.

Method used

The driving fatigue monitoring system adopts BeiDou satellite positioning and multimodal data fusion technology, including multimodal information perception and acquisition, data preprocessing, fatigue state identification and assessment, and intelligent early warning and adaptive optimization modules. It uses BeiDou satellite positioning to provide a unified time reference, performs spatiotemporal synchronization of multi-source data, and uses attention mechanism to enhance fatigue-related features, suppress noise interference, and achieve deep fusion of cross-modal information.

Benefits of technology

It significantly improves the accuracy and robustness of fatigue recognition, can dynamically adjust the modal data weights, realize the transformation from passive detection to active warning, trigger warning signals in advance, and improve the driver's reaction time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121287147B_ABST
    Figure CN121287147B_ABST
Patent Text Reader

Abstract

The application provides a driving fatigue monitoring auxiliary system based on Beidou satellite positioning and multi-modal data fusion technology, and relates to the field of digital data processing, and comprises a multi-modal information perception and acquisition module, a data preprocessing and quality guarantee module, a fatigue state recognition and evaluation module and an intelligent early warning and adaptive optimization module.The multi-modal information perception and acquisition module is responsible for real-time acquisition of driver state, driving behavior and environmental information.The data preprocessing and quality guarantee module is responsible for time-space alignment, quality evaluation and feature standardization of multi-source data.The fatigue state recognition and evaluation module is responsible for fusion of multi-dimensional features and identification of fatigue grade and risk trend.The intelligent early warning and adaptive optimization module is responsible for implementation of hierarchical intervention and continuous optimization of system performance.The system utilizes the accurate time-space reference provided by Beidou satellite positioning to realize effective fusion of multi-source heterogeneous data, and significantly improves the accuracy, real-time performance and individualization level of fatigue monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing, specifically to a driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology. Background Technology

[0002] Driver fatigue is one of the main causes of traffic accidents. When drivers are fatigued, their reaction time is prolonged, their attention is reduced, and their judgment is impaired, greatly increasing the risk of serious traffic accidents. To effectively prevent fatigued driving, researchers have conducted extensive research on driver fatigue monitoring technologies, mainly including detection methods based on physiological signals, facial visual features, driving behavior monitoring, and multimodal data fusion.

[0003] Physiological signal-based detection methods assess fatigue by collecting physiological parameters such as electroencephalogram (EEG), electrocardiogram (ECG), and electromyogram (EMG). While these methods offer high accuracy, they require contact sensors, causing discomfort and interference for the driver, thus limiting their practical application. Facial visual feature-based methods capture driver facial images using a camera and analyze features such as eyelid closure, blinking frequency, and yawning frequency to determine fatigue. This non-contact method is susceptible to changes in lighting and driver's glasses, resulting in inconsistent accuracy. Driving behavior-based methods indirectly infer driver fatigue by monitoring vehicle operating parameters such as steering wheel angle, lane departure, and speed changes. However, these methods are easily affected by road conditions and driving habits, leading to lower reliability.

[0004] Currently, three commonly used fatigue detection algorithms include the PERCLOS algorithm based on eye features, the HRV analysis algorithm based on heart rate variability, and deep learning algorithms based on multimodal feature fusion. The PERCLOS algorithm assesses fatigue level by calculating the proportion of eyelid closure time per unit time. This algorithm is simple to compute and has good real-time performance, but it relies solely on a single eye feature, resulting in a significant decrease in accuracy in low light conditions or when the driver wears glasses, and it cannot capture early signs of fatigue. The HRV analysis algorithm assesses the activity state of the autonomic nervous system by extracting time and frequency domain features from electrocardiogram signals, reflecting physiological fatigue changes. However, it requires wearable sensors to collect data, which can be restrictive for drivers, and heart rate changes are influenced by various factors such as emotions and exercise, resulting in insufficient specificity. Deep learning-based multimodal fusion algorithms input multi-source data such as physiological signals, facial images, and driving behavior into convolutional neural networks or recurrent neural networks for feature extraction and fusion, improving detection accuracy. However, most existing methods employ simple feature concatenation strategies, failing to fully exploit the complementarity and temporal correlation between different modalities. Furthermore, model training requires a large number of labeled samples, resulting in high computational complexity and difficulty meeting real-time requirements.

[0005] Foreign patent US8022831B1 discloses an interactive fatigue management system. This system detects driver fatigue and drowsiness, provides rest stop information to the driver using an in-vehicle GPS system, and contacts the driver's friends or family via mobile phone. The main shortcomings of this patent are its reliance on a single fatigue detection method, depending solely on facial monitoring or driving behavior analysis, lacking a multimodal data fusion mechanism, resulting in insufficient detection accuracy and robustness. After detecting fatigue, the system primarily intervenes through information prompts and external communication, failing to implement tiered warnings based on fatigue severity, and its intervention strategies lack specificity. Furthermore, this patent does not address spatiotemporal data synchronization technology based on BeiDou satellite positioning, failing to provide a unified time reference for multi-source sensor data, leading to timing misalignment issues during the fusion of different modal data, affecting the real-time performance and accuracy of fatigue assessment.

[0006] Domestic patent CN105652765A discloses a GPS monitoring and communication system for preventing fatigued driving through fingerprint recognition. This system uses a GPS positioning module to obtain vehicle location and speed information, records driver rest time through a timer, and triggers an alarm and sends information to a remote service center when the rest time is insufficient or fatigued driving is detected. The main drawback of this patent is its overly simplistic fatigue detection methodology, relying solely on driving duration and rest time to determine fatigue levels. It fails to collect multimodal data directly reflecting fatigue, such as the driver's physiological characteristics, facial expressions, and driving behavior, making it unable to accurately identify the actual fatigue state of different individuals. The system lacks real-time monitoring capabilities of the driver's physiological state and cannot capture the dynamic process of fatigue accumulation or provide early warning signals. GPS positioning data is only used for vehicle position and speed monitoring, failing to fully leverage the role of BeiDou satellite positioning in multimodal data spatiotemporal synchronization. The system as a whole is a passive monitoring system based on time rules, lacking intelligent fatigue recognition algorithms and adaptive optimization mechanisms, making it difficult to meet the needs of accurate fatigue state identification in complex driving environments. The positioning accuracy of the BeiDou system in China and surrounding areas is better than 2.5 meters, and with the adoption of a ground-based augmentation system, the accuracy can reach the centimeter level, providing high-precision position data support for vehicle trajectory analysis and driving behavior assessment. Summary of the Invention

[0007] The purpose of this invention is to address the shortcomings by proposing a driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology.

[0008] The present invention adopts the following technical solution:

[0009] A driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology includes a multimodal information perception and acquisition module, a data preprocessing and quality assurance module, a fatigue state identification and assessment module, and an intelligent early warning and adaptive optimization module.

[0010] The multimodal information perception and acquisition module is responsible for acquiring driver status, driving behavior and environmental information in real time; the data preprocessing and quality assurance module is responsible for spatiotemporal alignment, quality assessment and feature standardization of multi-source data; the fatigue state identification and assessment module is responsible for fusing multi-dimensional features and judging fatigue level and risk trend; and the intelligent early warning and adaptive optimization module is responsible for implementing graded intervention and continuously optimizing system performance.

[0011] The multimodal information perception and acquisition module includes a driver physiological state monitoring unit, a driving behavior feature extraction unit, and an environment and vehicle state perception unit. The driver physiological state monitoring unit is used to collect physiological and visual features in real time. The driving behavior feature extraction unit monitors driving behavior indicators and extracts behavioral features through the vehicle-mounted sensing system. The environment and vehicle state perception unit obtains the vehicle's geographical location, driving speed, acceleration, and route trajectory.

[0012] The data preprocessing and quality assurance module includes a multi-source data spatiotemporal synchronization unit, a data quality assessment and anomaly detection unit, and a feature extraction and standardization unit. The multi-source data spatiotemporal synchronization unit uses the BeiDou satellite positioning timestamp as a unified benchmark to synchronize the time and calibrate the spatial coordinates of data information from different sensors. The data quality assessment and anomaly detection unit performs integrity verification and reliability assessment on the raw data collected by the sensors. The feature extraction and standardization unit transforms the preprocessed multimodal data into a unified feature vector space.

[0013] The fatigue state identification and assessment module includes a multimodal feature fusion calculation unit, a fatigue level classification and discrimination unit, and a fatigue trend prediction and risk assessment unit. The multimodal feature fusion calculation unit is used to perform feature-level fusion of standardized physiological features, visual features, and driving behavior features. The fatigue level classification and discrimination unit classifies the driver's current state into three levels: alert, mild fatigue, and severe fatigue. The fatigue trend prediction and risk assessment unit is used to predict the fatigue evolution trend within a certain time window in the future.

[0014] The intelligent early warning and adaptive optimization module includes a graded early warning and multimodal intervention unit, a personalized threshold adaptive unit, and a system self-learning and performance optimization unit. The graded early warning and multimodal intervention unit implements differentiated warnings through multiple channels based on the fatigue level discrimination results. The personalized threshold adaptive unit establishes a dynamic fatigue threshold model based on individual driver differences and historical behavior data. The system self-learning and performance optimization unit records and evaluates the effects of fatigue events, intervention measures, and driver responses throughout the entire process.

[0015] Furthermore, the multimodal feature fusion computing unit includes a feature-level fusion processor, an attention mechanism processor, and a risk index calculator. The feature-level fusion processor is used to input standardized physiological feature vectors, visual feature vectors, and driving behavior feature vectors into a shared representation learning network, extract high-order semantic features of each modality through multi-layer nonlinear transformation, and achieve deep fusion of cross-modal information in the feature space. The attention mechanism processor is used to dynamically allocate the weights of different modal features in the fusion process, strengthen feature channels strongly correlated with fatigue state, suppress the interference of noise and redundant information, and explore the temporal dependencies and complementarities between multimodal data. The risk index calculator inputs the fused comprehensive feature vector into a fully connected layer and an activation function, and outputs a continuous comprehensive fatigue risk index through regression calculation.

[0016] Furthermore, the attention mechanism processor calculates the fatigue gradient attention score of each mode at time t according to the following formula:

[0017] ;

[0018] Where Q(t) is the query vector at time t, and K i (t) is the key vector of the i-th mode, d k To query the dimensions of the vector and key vector, softmax() is the normalization function. The gradient modulation intensity coefficient, Let F be the time gradient of the fatigue risk index at time t. i (t) refers to the eigenvector of the i-th mode. These are the mean eigenvectors of the three modes;

[0019] The feature-level fusion processor weights and fuses the feature vectors of the three modalities based on attention scores to obtain a fused feature vector F. fus (t).

[0020] Furthermore, the attention mechanism processor calculates the cooperative gain coefficient between the two modes according to the following formula:

[0021] ;

[0022] Among them, MI(F i F j |s) represents the mutual information between the eigenvectors of the i-th mode and the j-th mode under state s. ftg Indicates a state of fatigue, s awk Indicates a state of wakefulness. This represents the cross-covariance of the eigenvectors of the i-th mode and the j-th mode within the current time window. Let represent the autovariance of the eigenvector of the i-th modality.

[0023] Furthermore, the risk index calculator calculates the comprehensive fatigue risk index r(t) according to the following formula:

[0024] ;

[0025] in, For the sigmoid activation function, r ins (t) represents the instantaneous risk term, r tem (t) represents the time-series risk term. As a weighting factor, The modulation intensity coefficient is the cooperative gain. It is a cooperative gain modulation function;

[0026] The instantaneous risk term is obtained by mapping the fused feature vector, the temporal risk term is obtained by the changes in the historical risk index, and the cooperative gain modulation function is calculated according to the following formula:

[0027] ;

[0028] Where M is the number of modes, w ij Let be the weights of the mode pair (i, j).

[0029] The beneficial effects achieved by this invention are:

[0030] This system introduces an attention-based feature fusion method, enabling dynamic adjustment of the weights of different modalities. This strengthens feature channels strongly correlated with fatigue state, suppresses noise interference, and fully exploits the complementarity and temporal dependencies between multimodal data, significantly improving the accuracy and robustness of fatigue identification. The calculation of the cooperative gain coefficient quantifies the interaction between different modalities, allowing the system to maintain high detection performance even when the quality of data in a particular modality deteriorates. The system not only identifies the current fatigue state but also predicts the fatigue evolution trend within a certain time window through temporal modeling technology. This represents a shift from passive detection to proactive warning, triggering early warning signals before fatigue deteriorates, giving drivers more reaction and adjustment time.

[0031] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the overall structural framework of the present invention;

[0033] Figure 2 This is a schematic diagram of the multimodal information sensing and acquisition module of the present invention;

[0034] Figure 3 This is a schematic diagram of the data preprocessing and quality assurance module of the present invention;

[0035] Figure 4 This is a schematic diagram of the fatigue state identification and evaluation module of the present invention;

[0036] Figure 5 This is a schematic diagram of the intelligent early warning and adaptive optimization module of the present invention;

[0037] Figure 6 This is a schematic diagram comparing the fatigue detection accuracy of this invention with other systems;

[0038] Figure 7 This diagram illustrates the accuracy of the present invention under different environmental conditions compared to other systems. Detailed Implementation

[0039] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated beforehand. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.

[0040] Example 1: This example provides a driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology, combined with... Figure 1 It includes a multimodal information perception and acquisition module, a data preprocessing and quality assurance module, a fatigue state identification and assessment module, and an intelligent early warning and adaptive optimization module;

[0041] The multimodal information perception and acquisition module is responsible for acquiring driver status, driving behavior and environmental information in real time; the data preprocessing and quality assurance module is responsible for spatiotemporal alignment, quality assessment and feature standardization of multi-source data; the fatigue state identification and assessment module is responsible for fusing multi-dimensional features and judging fatigue level and risk trend; and the intelligent early warning and adaptive optimization module is responsible for implementing graded intervention and continuously optimizing system performance.

[0042] The multimodal information perception and acquisition module includes a driver physiological state monitoring unit, a driving behavior feature extraction unit, and an environment and vehicle state perception unit. The driver physiological state monitoring unit is used to collect physiological and visual features in real time. The driving behavior feature extraction unit monitors driving behavior indicators and extracts behavioral features through the vehicle-mounted sensing system. The environment and vehicle state perception unit obtains the vehicle's geographical location, driving speed, acceleration, and route trajectory.

[0043] The data preprocessing and quality assurance module includes a multi-source data spatiotemporal synchronization unit, a data quality assessment and anomaly detection unit, and a feature extraction and standardization unit. The multi-source data spatiotemporal synchronization unit uses the BeiDou satellite positioning timestamp as a unified benchmark to synchronize the time and calibrate the spatial coordinates of data information from different sensors. The data quality assessment and anomaly detection unit performs integrity verification and reliability assessment on the raw data collected by the sensors. The feature extraction and standardization unit transforms the preprocessed multimodal data into a unified feature vector space.

[0044] The fatigue state identification and assessment module includes a multimodal feature fusion calculation unit, a fatigue level classification and discrimination unit, and a fatigue trend prediction and risk assessment unit. The multimodal feature fusion calculation unit is used to perform feature-level fusion of standardized physiological features, visual features, and driving behavior features. The fatigue level classification and discrimination unit classifies the driver's current state into three levels: alert, mild fatigue, and severe fatigue. The fatigue trend prediction and risk assessment unit is used to predict the fatigue evolution trend within a certain time window in the future.

[0045] The intelligent early warning and adaptive optimization module includes a graded early warning and multimodal intervention unit, a personalized threshold adaptive unit, and a system self-learning and performance optimization unit. The graded early warning and multimodal intervention unit implements differentiated warnings through multiple channels based on the fatigue level discrimination results. The personalized threshold adaptive unit establishes a dynamic fatigue threshold model based on individual driver differences and historical behavior data. The system self-learning and performance optimization unit records and evaluates the effects of fatigue events, intervention measures, and driver responses throughout the entire process.

[0046] The multimodal feature fusion computing unit includes a feature-level fusion processor, an attention mechanism processor, and a risk index calculator. The feature-level fusion processor is used to input standardized physiological feature vectors, visual feature vectors, and driving behavior feature vectors into a shared representation learning network, extract high-order semantic features of each modality through multi-layer nonlinear transformation, and achieve deep fusion of cross-modal information in the feature space. The attention mechanism processor is used to dynamically allocate the weights of different modal features in the fusion process, strengthen feature channels strongly correlated with fatigue state, suppress the interference of noise and redundant information, and explore the temporal dependencies and complementarities between multimodal data. The risk index calculator inputs the fused comprehensive feature vector into a fully connected layer and an activation function, and outputs a continuous comprehensive fatigue risk index through regression calculation.

[0047] The attention mechanism processor calculates the fatigue gradient attention score of each modality at time t according to the following formula:

[0048] ;

[0049] Where Q(t) is the query vector at time t, and K i (t) is the key vector of the i-th mode, d k To query the dimensions of the vector and key vector, softmax() is the normalization function. The gradient modulation intensity coefficient, Let F be the time gradient of the fatigue risk index at time t. i (t) refers to the eigenvector of the i-th mode. These are the mean eigenvectors of the three modes;

[0050] The feature-level fusion processor weights and fuses the feature vectors of the three modalities based on attention scores to obtain a fused feature vector F. fus (t).

[0051] The attention mechanism processor calculates the cooperative gain coefficient between the two modes according to the following formula:

[0052] ;

[0053] Among them, MI(F i F j |s) represents the mutual information between the eigenvectors of the i-th mode and the j-th mode under state s. ftg Indicates a state of fatigue, s awk Indicates a state of wakefulness. This represents the cross-covariance of the eigenvectors of the i-th mode and the j-th mode within the current time window. Let represent the autovariance of the eigenvector of the i-th modality.

[0054] The risk index calculator calculates the comprehensive fatigue risk index r(t) according to the following formula:

[0055] ;

[0056] in, For the sigmoid activation function, r ins (t) represents the instantaneous risk term, r tem (t) represents the time-series risk term. As a weighting factor, The modulation intensity coefficient is the cooperative gain. It is a cooperative gain modulation function;

[0057] The instantaneous risk term is obtained by mapping the fused feature vector, the temporal risk term is obtained by the changes in the historical risk index, and the cooperative gain modulation function is calculated according to the following formula:

[0058] ;

[0059] Where M is the number of modes, w ij Let be the weights of the mode pair (i, j).

[0060] Example 2: A driving fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology, including a multimodal information perception and acquisition module, a data preprocessing and quality assurance module, a fatigue state identification and assessment module, and an intelligent early warning and adaptive optimization module;

[0061] The multimodal information perception and acquisition module is responsible for acquiring driver status, driving behavior and environmental information in real time; the data preprocessing and quality assurance module is responsible for spatiotemporal alignment, quality assessment and feature standardization of multi-source data; the fatigue state identification and assessment module is responsible for fusing multi-dimensional features and judging fatigue level and risk trend; and the intelligent early warning and adaptive optimization module is responsible for implementing graded intervention and continuously optimizing system performance.

[0062] Combination Figure 2 The multimodal information perception and acquisition module includes a driver physiological state monitoring unit, a driving behavior feature extraction unit, and an environment and vehicle state perception unit. The driver physiological state monitoring unit is used to collect physiological and visual features in real time and construct a driver fatigue physiological feature vector. The driving behavior feature extraction unit monitors driving behavior indicators through the vehicle-mounted sensing system and extracts behavioral features that reflect the stability of driving operation and the level of attention. The environment and vehicle state perception unit uses Beidou satellite positioning technology to obtain the vehicle's geographical location, driving speed, acceleration, and route trajectory, while also collecting time information and road environment parameters.

[0063] Combination Figure 3 The data preprocessing and quality assurance module includes a multi-source data spatiotemporal synchronization unit, a data quality assessment and anomaly detection unit, and a feature extraction and standardization unit. The multi-source data spatiotemporal synchronization unit uses the BeiDou satellite positioning timestamp as a unified benchmark to perform time synchronization and spatial coordinate calibration on physiological signals, visual data, driving behavior data, and environmental information from different sensors. The data quality assessment and anomaly detection unit performs integrity verification and reliability assessment on the raw data collected by the sensors, identifies sensor failures, data missing or abnormal fluctuations, and initiates data compensation or degradation processing mechanisms. The feature extraction and standardization unit transforms the preprocessed multimodal data into a unified feature vector space, forming a standardized feature set that can be used for fusion computing.

[0064] Combination Figure 4 The fatigue state identification and assessment module includes a multimodal feature fusion calculation unit, a fatigue level classification and discrimination unit, and a fatigue trend prediction and risk assessment unit. The multimodal feature fusion calculation unit is used to perform feature-level fusion of standardized physiological features, visual features, and driving behavior features, mine cross-modal correlations, and output a comprehensive fatigue risk index. The fatigue level classification and discrimination unit uses the fused features as input to the classification model to classify the driver's current state into three levels: alert, mild fatigue, and severe fatigue, generate fatigue level labels, and associate them with risk index thresholds. The fatigue trend prediction and risk assessment unit is used to model historical fatigue data and predict the fatigue evolution trend within a certain time window in the future.

[0065] Combination Figure 5 The intelligent early warning and adaptive optimization module includes a graded early warning and multimodal intervention unit, a personalized threshold adaptive unit, and a system self-learning and performance optimization unit. The graded early warning and multimodal intervention unit implements differentiated warnings through multiple channels based on the fatigue level discrimination results, and dynamically adjusts the warning intensity and frequency according to the severity of fatigue. The personalized threshold adaptive unit establishes a dynamic fatigue threshold model based on individual driver differences and historical behavior data, and adjusts the fatigue judgment criteria in real time. The system self-learning and performance optimization unit records and evaluates the effects of fatigue events, intervention measures, and driver responses throughout the process, and continuously optimizes the fusion model parameters and discrimination thresholds.

[0066] The driver physiological state monitoring unit includes a facial feature acquisition processor, a physiological signal acquisition processor, and a multimodal data preprocessor. The facial feature acquisition processor is used to capture driver facial images in real time and extract visual feature parameters to locate facial feature points. The physiological signal acquisition processor monitors physiological indicators in real time based on wearable sensors and extracts physiological parameters that reflect the activity state of the autonomic nervous system. The multimodal data preprocessor performs preliminary fusion and feature vectorization on the acquired facial visual features and physiological signals, removes redundant information, and extracts key feature combinations that have fatigue characterization significance to construct a driver fatigue physiological feature vector library.

[0067] The multimodal data and processor calculate the enhancement weight w of the i-th mode under the current fatigue state s according to the following formula. i (t):

[0068] ;

[0069] Where Qi(t) represents the basic signal quality index of the i-th mode at time t. Let be the discrimination sensitivity of the i-th mode under fatigue state s. T represents the deviation of the current feature of the i-th modality from the baseline feature. temp For deviation parameters;

[0070] The fatigue state is a discrete state label output by the fatigue level classification and discrimination unit at the previous moment. The fatigue sensitivity function is a lookup table that is pre-built through offline training. The specific construction method is as follows: During the system development phase, a large number of labeled samples are collected, and the contribution of each modal feature to fatigue discrimination under different fatigue states is statistically analyzed to form a sensitivity matrix.

[0071] The driving behavior feature extraction unit includes a steering operation monitoring processor, a pedal operation monitoring processor, and a lane keeping monitoring processor. The steering operation monitoring processor is used to collect parameters such as steering wheel angle, steering frequency, steering amplitude, and steering torque in real time, analyze the smoothness and stability of the driver's steering operation, and identify abnormal steering patterns and overcorrection behaviors. The pedal operation monitoring processor monitors the depth, frequency, and response time of the accelerator and brake pedals. By analyzing the coordination and reaction speed of the pedal operation, it extracts behavioral feature parameters that reflect the driver's concentration and reaction ability. The lane keeping monitoring processor uses the sensor data of the lane departure warning system to calculate the lateral deviation of the vehicle in the lane, the deviation frequency, and the correction time in real time, and evaluates the driver's lane keeping ability and attention persistence level.

[0072] The environment and vehicle status perception unit includes a Beidou satellite positioning processor, a vehicle motion status processor, and an environmental information collector. The Beidou satellite positioning processor acquires the vehicle's latitude and longitude coordinates, altitude, driving direction, and positioning accuracy information in real time through the Beidou satellite positioning system, and calculates the driving trajectory and route deviation based on continuous positioning data, providing an accurate spatiotemporal reference for fatigue monitoring. The vehicle motion status processor collects motion parameters such as instantaneous speed, acceleration, deceleration, and yaw rate of the vehicle through the vehicle CAN bus interface, and analyzes the stability of vehicle driving and the change law of motion status in combination with Beidou satellite positioning data. The environmental information collector collects environmental context information such as the current time period, light intensity, weather conditions, and road type through a light sensor, a rain sensor, and a clock module.

[0073] The multi-source data spatiotemporal synchronization unit includes a timestamp unification processor, a coordinate system calibration processor, and a data alignment buffer. The timestamp unification processor uses UTC time provided by the BeiDou satellite positioning system as the global time reference and adds a unified format timestamp identifier to the data streams from different sensors to compensate for clock deviations and sampling delays between sensors. The coordinate system calibration processor establishes a unified mapping relationship between the vehicle coordinate system and the geographic coordinate system and performs spatial transformation and calibration on the visual sensor coordinates, vehicle sensor coordinates, and BeiDou satellite positioning coordinates. The data alignment buffer caches sensor data with different sampling frequencies through a sliding time window mechanism and performs time registration on asynchronous data.

[0074] The data quality assessment and anomaly detection unit includes a sensor status diagnostic processor, a data integrity verification processor, and an outlier detection processor. The sensor status diagnostic processor monitors the working status and output signal quality of various sensors in real time, identifies sensor failures, poor contact, or performance degradation through self-testing protocols and redundancy checks, generates a sensor health assessment report, and triggers fault warnings. The data integrity verification processor verifies the integrity of the collected data stream, detects data packet loss, transmission errors, or sampling interruptions, calculates the statistical missing rate, and evaluates the missing patterns. The outlier detection processor identifies abnormal fluctuations and outliers in physiological signals, driving behavior, and environmental data, and removes noise interference by setting dynamic thresholds and multidimensional correlation analysis.

[0075] The feature extraction and standardization unit includes a feature engineering processor, a normalization processor, and a feature vector builder. The feature engineering processor performs deep feature mining on quality-assured multimodal data, extracting eye movement features and facial micro-movement features from facial images, frequency domain features and nonlinear dynamic features from physiological signals, and operational entropy and response delay features from driving behavior, forming a multidimensional original feature set. The normalization processor standardizes features with different physical dimensions and numerical ranges, eliminating dimensional differences and mapping various features to a unified numerical range for easy fusion calculation. The feature vector builder selects the most discriminative feature subset based on feature importance assessment and redundancy analysis results, and organizes the feature vector structure according to temporal relationships and logical associations, generating a standardized multimodal feature vector library for the fusion module to call.

[0076] The multimodal feature fusion computing unit includes a feature-level fusion processor, an attention mechanism processor, and a risk index calculator. The feature-level fusion processor is used to input standardized physiological feature vectors, visual feature vectors, and driving behavior feature vectors into a shared representation learning network, extract high-order semantic features of each modality through multi-layer nonlinear transformation, and achieve deep fusion of cross-modal information in the feature space. The attention mechanism processor is used to dynamically allocate the weights of different modal features in the fusion process, strengthen feature channels strongly correlated with fatigue state, suppress the interference of noise and redundant information, and explore the temporal dependencies and complementarities between multimodal data. The risk index calculator inputs the fused comprehensive feature vector into a fully connected layer and an activation function, and outputs a continuous comprehensive fatigue risk index through regression calculation.

[0077] The attention mechanism processor calculates the fatigue gradient attention score of each modality at time t according to the following formula:

[0078] ;

[0079] Where Q(t) is the query vector at time t, and K i (t) is the key vector of the i-th mode, d k To query the dimensions of the vector and key vector, softmax() is the normalization function. The gradient modulation intensity coefficient, Let F be the time gradient of the fatigue risk index at time t. i (t) refers to the eigenvector of the i-th mode. These are the mean eigenvectors of the three modes;

[0080] The feature-level fusion processor weights and fuses the feature vectors of the three modalities based on attention scores to obtain a fused feature vector F. fus (t);

[0081] The attention mechanism processor calculates the cooperative gain coefficient between the two modes according to the following formula:

[0082] ;

[0083] Among them, MI(F i F j |s) represents the mutual information between the eigenvectors of the i-th mode and the j-th mode under state s. ftg Indicates a state of fatigue, s awk Indicates a state of wakefulness. This represents the cross-covariance of the eigenvectors of the i-th mode and the j-th mode within the current time window. Let represent the autovariance of the eigenvector of the i-th modality;

[0084] The risk index calculator calculates the comprehensive fatigue risk index r(t) according to the following formula:

[0085] ;

[0086] in, For the sigmoid activation function, r ins (t) represents the instantaneous risk term, r tem (t) represents the time-series risk term. As a weighting factor, The modulation intensity coefficient is the cooperative gain. It is a cooperative gain modulation function;

[0087] The instantaneous risk term is obtained by mapping the fused feature vector, the temporal risk term is obtained by the changes in the historical risk index, and the cooperative gain modulation function is calculated according to the following formula:

[0088] ;

[0089] Where M is the number of modes, w ij Let (i, j) be the weights of the mode pair (i, j).

[0090] The fatigue level classification and discrimination unit includes a classification model processor, a threshold mapping processor, and a confidence evaluator. The classification model processor takes the fused feature vector and the comprehensive fatigue risk index as input, establishes a fatigue state classification model through supervised learning, and outputs three discrete level labels: awake, mild fatigue, and severe fatigue. The threshold mapping processor establishes a mapping relationship between fatigue level and risk index based on the classification results, sets a corresponding risk index threshold range for each fatigue level, and dynamically adjusts the threshold boundary in combination with historical statistical data to improve classification accuracy and robustness. The confidence evaluator calculates the posterior probability and confidence level of the classification decision, evaluates the reliability of the current classification result, and triggers secondary discrimination or extends the observation time window when the confidence level is lower than the set threshold.

[0091] The fatigue trend prediction and risk assessment unit includes a time series modeling processor, a trend prediction processor, and a risk level assessor. The time series modeling processor models historical fatigue data sequences, learns the dynamic patterns and periodic laws of fatigue state evolution over time, and captures nonlinear characteristics and precursors of sudden changes in the fatigue accumulation process. The trend prediction processor performs rolling predictions of the fatigue risk index within a certain time window based on the trained time series model, outputs fatigue trend curves for the next 5, 10, and 15 minutes, and identifies the time nodes and rate of change when the fatigue state is about to deteriorate. The risk level assessor integrates the current fatigue level, the predicted risk index value, and the rate of change of the trend to assess the short-term and medium-term risk levels. When it is predicted that the fatigue risk will increase significantly in a short period of time, an early warning signal is triggered in advance.

[0092] The trend prediction processor predicts the future based on the following formula. Fatigue risk index at a given time point :

[0093] ;

[0094] ;

[0095] ;

[0096] ;

[0097] Among them, W slow Extract the weight vector for the slow component, b slow W is the bias term. amp To extract weights for amplitude, Based on the oscillation frequency, c ij Let be the contribution coefficient of the combination of the i-th mode and the j-th mode to the fluctuation frequency. For frequency modulation amplitude, I is the cumulative fatigue growth rate coefficient.hist (t) represents the historical fatigue integral. For wave phase, The fast component attenuation coefficient;

[0098] The graded early warning and multimodal intervention unit includes an early warning intensity control processor, a multi-channel output controller, and an intervention effect feedback device. The early warning intensity control processor formulates differentiated warning strategies based on fatigue level classification results and risk assessment levels, setting a gentle prompt mode for mild fatigue and a mandatory warning mode for severe fatigue, and dynamically adjusts the warning intensity and repetition frequency according to the driver's response. The multi-channel output controller coordinates multiple output channels such as the voice broadcast system, seat vibration device, dashboard flashing light, and buzzer, and activates different warning methods synchronously or sequentially according to a preset combination scheme. Through multi-sensory stimulation, it ensures that the driver effectively perceives risk signals and attracts attention. The intervention effect feedback device monitors the changes in the driver's physiological state and the improvement in driving behavior after the intervention measures are activated in real time, evaluates the immediate effect and duration of the warning intervention, and inputs the feedback results into the self-learning module to optimize the intervention strategy and warning parameter configuration.

[0099] The personalized threshold adaptive unit includes an individual feature modeling processor, a dynamic threshold adjustment processor, and a context-aware processor. The individual feature modeling processor establishes an individualized fatigue feature database by long-term tracking and recording the driver's baseline physiological characteristics, driving habits, and fatigue response patterns, and identifies individual differences and specific indicators in fatigue performance among different drivers. The dynamic threshold adjustment processor adjusts the fatigue judgment threshold in real time based on the individual feature model and historical fatigue event data, and sets personalized risk index thresholds and classification boundaries for different drivers to improve the targeting and accuracy of fatigue identification. The context-aware processor combines current driving environment information and dynamically corrects the judgment threshold according to contextual factors such as time period, driving duration, road type, and weather conditions.

[0100] The system's self-learning and performance optimization unit includes an event recording and annotation processor, a model incremental learning processor, and a performance evaluation and iterator. The event recording and annotation processor fully records each fatigue warning event, the execution process of intervention measures, and the driver's response results, saving multimodal sensor data, fatigue judgment results, intervention strategy parameters, and post-event verification information. It also constructs an annotated fatigue event sample library for model optimization. The model incremental learning processor continuously optimizes feature fusion model parameters, classification model weights, and threshold configurations based on the accumulated fatigue event sample library. It achieves online model evolution through mini-batch gradient updates and regularization constraints. The performance evaluation and iterator periodically calculates key performance indicators such as the system's warning accuracy, false alarm rate, false negative rate, and intervention effectiveness, evaluates the model iteration effect, and adjusts the learning rate and optimization strategy based on performance feedback.

[0101] The 'i' and 'j' mentioned above are ordinal numbers used to represent sequence numbers and have no actual meaning.

[0102] To verify the beneficial effects of the driver fatigue monitoring and assistance system of this invention, the following comparative experiment was conducted. Thirty professional drivers aged 25 to 55 were selected as test subjects, and actual road tests were conducted on closed test roads for three months. Test scenarios included three typical driving environments: urban roads, highways, and mountain roads, covering different time periods such as daytime, nighttime, and early morning. Each driver wore standardized physiological monitoring equipment, and the vehicle was equipped with the system of this invention and three comparative systems: a single-vision monitoring system based on the PERCLOS algorithm, a physiological signal monitoring system based on HRV analysis, and an indirect detection system based on driving behavior monitoring. During the experiment, drivers were required to drive according to normal driving habits. The system automatically recorded fatigue detection results, while professional observers manually labeled the results based on the drivers' actual performance as the gold standard for true fatigue status. A total of 1850 valid fatigue event samples were collected during the test, including 1120 cases of mild fatigue and 730 cases of severe fatigue. The test data were then processed... Figure 6 and Figure 7 ;

[0103] Example 3: This example provides an implementation method for a driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology. The system also includes a multimodal information perception and acquisition module, a data preprocessing and quality assurance module, a fatigue state identification and assessment module, and an intelligent early warning and adaptive optimization module.

[0104] In the multimodal information perception and acquisition module, the driver's physiological state monitoring unit uses a binocular infrared camera to acquire facial images. Specifically, the Hikvision DS-2CD8626FWD camera has a resolution of 1920×1080 pixels, a frame rate of 30fps, and is equipped with an active infrared fill light with a working wavelength of 850nm, enabling clear imaging even in low-light environments of 0.01 lux. The facial feature acquisition processor locates 68 facial key points based on the Dlib face detection library and extracts 12 dimensions of visual feature parameters, including eye aspect ratio (EAR), blink frequency, eyelid closure duration (PERCLOS), yawn frequency (YAW), and head posture angle. When the EAR value is below 0.2 for three consecutive frames, it is determined to be a closed-eye state; when the PERCLOS value is greater than 0.15 and the duration exceeds 5 seconds, it is marked as a sign of fatigue. The physiological signal acquisition processor uses the Xiaomi Mi Band 7 wearable smart bracelet, which has a built-in PPG photoplethysmography sensor with a sampling frequency of 25Hz. It monitors heart rate (HR), heart rate variability (HRV) time-domain indices SDNN and RMSSD, and frequency-domain indices LF / HF ratio in real time. When the heart rate is lower than 85% of the driver's baseline heart rate and the SDNN value decreases by more than 30%, it indicates increased parasympathetic nerve activity. As an alternative, physiological signal acquisition can also use a steering wheel-integrated capacitive sensor array, which collects ECG and GSR signals through palm contact. This solution does not require the driver to wear additional equipment, but the signal quality depends on the stability of the grip. When the multimodal data preprocessor performs initial fusion of facial features and physiological signals, it dynamically allocates weights according to signal quality indices. When the infrared camera detects that the driver's head turning angle exceeds 45 degrees, making facial features invisible, the physiological signal weight is automatically increased to 0.75; conversely, when the physiological sensor has poor contact, the visual feature weight is increased.

[0105] In the driving behavior feature extraction unit, the steering operation monitoring processor reads steering wheel angle sensor data via the CAN bus at a sampling frequency of 100Hz, recording steering angular velocity, steering amplitude standard deviation, and steering reverse correction counts. Under normal driving conditions, the steering angular velocity standard deviation remains between 8 and 15 degrees per second. When the standard deviation drops below 5 degrees per second and persists for more than one minute, it indicates increased monotonicity in driving operations. The pedal operation monitoring processor collects signals from the accelerator pedal position sensor and brake pedal pressure sensor, extracting pedal response time, rate of change of pedal force, and accelerator-brake switching frequency. Under fatigue conditions, the frequency of minor accelerator pedal adjustments decreases from the normal 12 to 18 times per minute to below 6 times per minute, and the braking reaction time increases from an average of 0.6 seconds to more than 1.2 seconds. The lane keeping monitoring processor utilizes the lane line recognition function of the Mobileye 630 forward-facing camera to calculate the vehicle lateral displacement (TLC) and lane departure warning (LDW) trigger frequency in real time. When the TLC is less than 0.5 meters more than three times within 30 seconds, it is determined that the lane keeping capability has decreased. As a technological variation, lane keeping monitoring can also employ a relative positioning scheme that combines millimeter-wave radar ranging with high-precision maps, achieving an accuracy of 0.1 meters, but at an increased cost of approximately 1,500 yuan.

[0106] In the environmental and vehicle status perception unit, the BeiDou satellite positioning processor uses the BeiDou-3 model and the ChipCreate UM982, supporting multi-frequency reception of B1I, B2a, B3I, L1, and L5 frequencies, with a positioning accuracy better than 1.5 meters and a timing accuracy of 20 nanoseconds. The processor generates timestamps in the format of year, month, day, hour, minute, second, plus millisecond, based on Coordinated Universal Time (UTC), with timestamp errors controlled within 10 milliseconds to ensure time alignment of data from different sensors. The BeiDou satellite positioning data update frequency is 10Hz, and instantaneous speed is calculated with an accuracy of 0.1 km / h, acceleration with an accuracy of 0.05 m / s², and heading angle with an accuracy of 0.5 degrees through continuous positioning points. When a vehicle is detected traveling at a high speed of 80 to 120 km / h for more than 1 hour, the system automatically lowers the fatigue judgment threshold by 10% to improve early warning sensitivity. The vehicle motion status processor acquires data from the vehicle speed sensor, yaw rate sensor, and three-axis acceleration sensor from the CAN bus, analyzing the standard deviation of longitudinal acceleration, yaw rate fluctuation amplitude, and serpentine driving index. Under normal driving conditions, the standard deviation of longitudinal acceleration is less than 0.3 m / s². Under fatigue conditions, unstable acceleration control causes the standard deviation to increase to over 0.6 m / s². The environmental information acquisition unit integrates an APDS-9960 light sensor with a measurement range of 0 to 60,000 lux. When the illuminance is below 500 lux, it is considered nighttime or a tunnel environment, triggering the illuminance compensation mechanism. The rain sensor uses capacitive raindrop detection with a resolution of 0.1 mm of rainfall. Rainy conditions reduce the road surface adhesion coefficient, increasing driving risk by 20%, and the system accordingly activates a warning 2 minutes in advance.

[0107] In the data preprocessing and quality assurance module, the timestamp unification processor of the multi-source data spatiotemporal synchronization unit receives the UTC time from BeiDou satellite positioning as the master clock and adds unified timestamps to the data from the infrared camera, physiological sensor, and vehicle CAN bus. The internal clock of the infrared camera drifts by approximately 3 milliseconds per hour, and the processor synchronizes and calibrates the clock deviation every 10 minutes using the NTP network time protocol. The physiological sensor and the main controller are connected via Bluetooth 5.0 with a transmission delay of 12 to 18 milliseconds. The processor records the sending and receiving timestamps, calculates the transmission delay, and performs time compensation. The coordinate system calibration processor establishes the vehicle coordinate system with the origin at the rear axle center point, the X-axis pointing towards the front of the vehicle, the Y-axis pointing to the left, and the Z-axis vertically upward. The infrared camera is mounted above the steering wheel on the center console with an offset of X=0.85 meters, Y=0 meters, and Z=1.20 meters. The camera coordinates are transformed to vehicle coordinates using a homogeneous transformation matrix. The WGS84 latitude and longitude coordinates output by BeiDou satellite positioning are converted to a planar coordinate system using Mercator projection to facilitate the calculation of the curvature of the driving trajectory. The data alignment buffer sets the sliding window length to 2 seconds with a step size of 0.1 seconds. The infrared camera's 30fps data contains 60 frames per window, the physiological sensor's 25Hz data contains 50 sampling points, and the CAN bus's 100Hz data contains 200 sampling points. The data of different frequencies are aligned to a unified time grid through linear interpolation and resampling.

[0108] In the data quality assessment and anomaly detection unit, the sensor status diagnostic processor monitors the operating voltage, operating temperature, and signal-to-noise ratio of each sensor in real time. An out-of-range operating voltage (12V±0.5V) of the infrared camera triggers a power supply anomaly alarm; an image signal-to-noise ratio below 25dB indicates lens contamination requiring cleaning. Physiological sensor wear detection measures skin contact impedance; an impedance value greater than 500kΩ indicates loose wear and signal quality degradation to an unusable level. CAN bus data packet verification uses CRC cyclic redundancy check; an error rate exceeding 0.1% is considered bus interference, triggering a data filtering mechanism. The data integrity verification processor counts the number of data frames received per second; for a normal infrared camera receiving 30 frames, a frame rate below 25fps for more than 3 seconds is marked as a data loss event. Physiological sensor data missing rate is calculated using a sliding window; when more than 20% of samples are missing within a 10-second window (i.e., 5 seconds of data are unusable), a historical data extrapolation compensation mechanism is activated, filling the missing values ​​with the average of the previous 30 seconds of data. The outlier detection processor sets a dynamic threshold range of ±35% for heart rate data; outliers outside this range are removed using median filtering. Steering wheel angle data is detected using the 3-sigma criterion. Abrupt points where the rate of change in angle exceeds the mean plus three standard deviations are identified as false triggers or sensor malfunctions and are discarded. When multiple sensors malfunction simultaneously, the system enters a degraded operation mode, retaining only reliable data sources for continued monitoring.

[0109] In the feature extraction and standardization unit, the feature engineering processor extracts the eye region from the facial image and calculates five-dimensional eye features: eye opening angle (EAR), blink frequency (BF), eye closure duration (PERCLOS), pupil diameter change rate (PDC), and eye movement velocity (EMS). It also extracts two-dimensional mouth features from the mouth region: yawning frequency (YF) and mouth opening amplitude (MAO). Finally, it extracts three-dimensional head posture features from the overall face: head pitch angle (pitch), yaw angle (yaw), and roll angle (roll). Physiological signals are decomposed into five levels of detail coefficients and approximation coefficients using wavelet transform. From the heart rate time series, it extracts four-dimensional time-domain features: mean heart rate (MHR), heart rate standard deviation (SDHR), root mean square of the sum of squares of the differences between adjacent heartbeats (RMSSD), and the percentage of consecutive differences exceeding 50ms (pNN50). Finally, it extracts three-dimensional frequency-domain features using fast Fourier transform: low-frequency power (LF), high-frequency power (HF), and the low-to-high frequency ratio (LF / HF). Driving behavior features were extracted using three dimensions: steering wheel angle standard deviation (SWA), steering correction frequency (SCF), and maximum steering speed (MSS). Accelerator pedal features included three dimensions: accelerator pedal response time (ART), pedal change rate (APR), and average pedal depth (APD). Lane departure frequency (LDN), lateral position change rate (LPV), and lane center offset distance (LCD) were also extracted. The total feature dimension was 27, forming the original feature space. A normalization processor used Z-score normalization to convert each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1, eliminating the influence of dimensions. For eye features (EAR), the normalization was (EAR-0.25) / 0.08; for heart rate features, it was (HR-75) / 12; and for steering wheel angle, it was (SWA-10) / 5. The feature vector builder calculates the importance score of each feature through Lasso regression analysis, selects the key feature subset with a score greater than 0.05. In eye features, PERCLOS and BF scores are 0.18 and 0.12, respectively. In physiological features, RMSSD and LF / HF scores for HRV are 0.15 and 0.11, respectively. In driving behavior, lane departure frequency and steering standard deviation scores are 0.14 and 0.10, respectively. Finally, 15 high-discriminative features are retained to construct a standardized feature vector.

[0110] In the fatigue state recognition and assessment module, the feature-level fusion processor of the multimodal feature fusion computing unit adopts the ResNet18 deep residual network architecture, which includes an input layer, four residual blocks, a global average pooling layer, and a fully connected layer. The input layer receives a 15-dimensional normalized feature vector. The first residual block contains two convolutional layers with a kernel size of 3×3, a stride of 1, and padding of 1, expanding the number of feature map channels from 15 to 32. The second residual block expands the number of channels to 64, the third residual block to 128, and the fourth residual block to 256. Each residual block uses skip connections to directly add the input to the output to avoid gradient vanishing, and the activation function is ReLU. Global average pooling compresses the 256-dimensional feature map into a 256-dimensional feature vector as a high-order semantic feature. The network was trained using the Adam optimizer with a learning rate of 0.001 and a batch size of 32. The training sample consisted of 15,000 cases (6,000 conscious, 5,000 mildly fatigued, and 4,000 severely fatigued), with a validation set of 3,000 and a test set of 2,000. After 500 epochs of training, the validation set accuracy reached 92.6%, and the test set accuracy reached 91.8%. As a technical variation, the fusion processor can also use the Transformer architecture, capturing long-range dependencies between features through a self-attention mechanism. This reduces the number of model parameters by 40% but increases training time by 1.5 times.

[0111] The attention mechanism processor employs a CBAM module combining channel attention and spatial attention. Channel attention first performs global average pooling and global max pooling on the 256-dimensional fused features to obtain two 256-dimensional vectors. These are input into a shared multilayer perceptron (MLP) containing two fully connected layers: the first layer has 64 neurons activated using ReLU, and the second layer has 256 neurons. The outputs of the two MLPs are summed and then activated by a sigmoid function to obtain a 256-dimensional channel attention weight vector, with weights ranging from 0 to 1. Spatial attention performs average pooling and max pooling on the 256-dimensional feature map along the channel dimension to obtain two feature maps. These are concatenated along the channel dimension and input into a 7×7 convolutional layer, outputting a 1-channel feature map. This map is then activated by a sigmoid function to obtain a spatial attention weight map. The channel weights and spatial weights are multiplied sequentially by the fused features to enhance the features. Important features have channel weights close to 1, while less important features have weights close to 0. Weights are increased for fatigue-related spatial regions and decreased for irrelevant regions. The attention-processed feature vector is then input into a risk index calculator. The risk index calculator employs a three-layer fully connected neural network: 256 neurons in the first layer, 128 neurons in the second, and 64 neurons in the third. The output layer has one neuron using a sigmoid activation function to output a fatigue risk index between 0 and 1. A risk index less than 0.3 indicates a conscious state, 0.3 to 0.6 indicates mild fatigue, and greater than 0.6 indicates severe fatigue. The risk index calculation incorporates a temporal risk term, using a Long Short-Term Memory (LSTM) network to model the risk index sequence over the past two minutes. The LSTM has 128 hidden units and learns the changing trend of the fatigue index. When the current risk index is 0.55 but has risen from 0.4 to 0.55 in the past minute, the temporal term contributes an additional 0.08 risk value, bringing the overall risk index to 0.63, entering the severe fatigue range, thus providing an early warning of worsening fatigue trends.

[0112] The classification model processor of the fatigue level classification and discrimination unit adopts a Support Vector Machine (SVM) classifier with a Radial Basis Function (RBF) kernel function, a penalty coefficient C=10, and a kernel parameter gamma=0.1. The classifier inputs 256-dimensional fused features and a 1-dimensional risk index, totaling 257 dimensions, and outputs three labels: conscious, mild fatigue, and severe fatigue. Training samples are balanced using the SMOTE oversampling technique. The number of conscious samples (6000) is reduced to 4500 through downsampling, the number of mild fatigue samples (5000) remains unchanged, and the number of severe fatigue samples (4000) is increased to 5000 through SMOTE upsampling. The SVM classifier uses a one-to-one strategy to train three binary classifiers: 94.2% accuracy for conscious vs. mild fatigue, 97.8% accuracy for conscious vs. severe fatigue, and 89.5% accuracy for mild vs. severe fatigue. The final fatigue level label is determined by a voting decision. The threshold mapping processor, based on statistics from 1000 labeled samples, showed that the mean risk index for conscious state was 0.18 with a standard deviation of 0.09, the mean for mild fatigue was 0.45 with a standard deviation of 0.12, and the mean for severe fatigue was 0.73 with a standard deviation of 0.11. A threshold of 0.30 was set for the boundary between conscious state and mild fatigue, and 0.60 for the boundary between mild and severe fatigue. The thresholds were dynamically adjusted every 100 hours of operation based on the accumulated sample statistics, with the adjustment range limited to ±0.05 to avoid threshold drift. The confidence estimator calculated the distance from the SVM decision function to the classification hyperplane. A distance greater than 1.5 times the boundary width was marked as high confidence, a distance between 0.5 and 1.5 times the boundary width was marked as medium confidence, and a distance less than 0.5 times the boundary width was marked as low confidence. When the confidence was low, the system extended the observation window from 2 seconds to 5 seconds to accumulate more evidence before making a judgment, reducing the risk of misjudgment.

[0113] The time-series modeling processor for the fatigue trend prediction and risk assessment unit employs a gated recurrent unit (GRU) network, consisting of three GRU layers, each with a hidden dimension of 128. The input is a fatigue risk index sequence from the past 10 minutes, sampled at 10-second intervals for a total of 60 time steps. The output is a 128-dimensional vector encoding the historical fatigue evolution pattern from the last hidden state of the GRU. The GRU units control information flow through update and reset gates; the update gate determines how much historical information is retained, and the reset gate determines how much irrelevant information is forgotten. Compared to LSTM, this reduces the number of parameters by 25% while maintaining comparable performance. The time-series model is trained on 5000 fatigue evolution trajectories, including three modes: gradual fatigue deepening, fatigue relief, and fatigue fluctuation. Training uses a teacher forcing strategy with a learning rate of 0.001 for 300 epochs. The trend prediction processor inputs the 128-dimensional hidden vector from the GRU output into a three-layer fully connected network to predict the fatigue risk index for the next 5, 10, and 15 minutes. The number of neurons in the fully connected layers are 128, 64, and 3, respectively, using ReLU as the activation function. The output layer has three neurons corresponding to the predicted risk index values ​​at the three time points. On the test set, the mean squared error (MSE) for 5-minute predictions was 0.012, for 10-minute predictions it was 0.028, and for 15-minute predictions it was 0.045. The prediction accuracy decreased as the time window increased, which was expected. When the risk index was predicted to exceed 0.6 within the next 10 minutes, the system issued a trend warning 8 minutes in advance, giving the driver ample time to find a service area to rest. The risk level assessor calculated the comprehensive risk level by combining the current risk index, the predicted trend, and the rate of change. If the current risk was 0.55 (mild fatigue) but the prediction reached 0.72 after 10 minutes with a rate of change of 0.017 per minute, the comprehensive risk level was upgraded to medium-high risk, triggering a level two warning.

[0114] In the intelligent early warning and adaptive optimization module, the warning intensity control processor of the graded early warning and multimodal intervention unit sets up a three-level early warning scheme. Level 1 warning corresponds to a mild fatigue risk index of 0.30 to 0.45, using a soft female voice prompt "We have detected that you may be somewhat fatigued, and we recommend taking a break," at a volume of 60 decibels. The yellow fatigue icon on the instrument panel flashes at a frequency of 0.5 Hz for 3 seconds, and the lumbar airbag in the seat slightly inflates to provide a tactile reminder at an intensity of 20%. Level 2 warning corresponds to a moderate fatigue risk index of 0.45 to 0.60, with the voice prompt changed to a stronger male voice "Your fatigue level is high, please stop and rest as soon as possible," at a volume of 75 decibels. The orange icon on the instrument panel flashes at a frequency of 1 Hz for 5 seconds, the seat vibrates at 50% intensity at a frequency of 2 Hz for 2 seconds, and the windows automatically lower by 5 centimeters to increase ventilation. A Level 3 warning corresponds to a severe fatigue risk index greater than 0.60. The system triggers a voice alarm, "Danger! Severe fatigue, stop immediately," repeated three times at 90 decibels. A red icon on the instrument panel flashes rapidly at 2Hz for 10 seconds, the seat vibrates at 100% intensity at 5Hz for 5 seconds, a buzzer sounds an intermittent alarm at 80 decibels, and the air conditioning automatically adjusts to maximum fan speed and minimum temperature of 18 degrees Celsius for forced alertness. A multi-channel output controller coordinates the audio system, seat control module, instrument panel display, and air conditioning system, sending control commands via the CAN bus. The response delay of each device is controlled within 200 milliseconds for synchronized activation. If the driver does not respond within 30 seconds after the warning is activated (e.g., by slowing down or using the turn signal), the system automatically upgrades the warning level and reactivates it. The intervention effect feedback device assesses the intervention effect by monitoring physiological characteristics and driving behavior after the warning. Effective intervention should result in an 8% to 15% increase in heart rate, a blinking frequency increase of more than 20%, and a return to normal steering wheel operation frequency of more than 12 times per minute within one minute. Statistical analysis of 150 early warning events showed that the effectiveness rate of Level 1 early warning was 72%, Level 2 early warning was 89%, and Level 3 early warning was 97%, with an overall average effectiveness rate of 86%. For cases where early warnings were ineffective, the feedback device recorded the fatigue characteristics, environmental conditions, and early warning parameters at the time, and input them into the self-learning module to analyze the causes of failure and optimize the early warning strategy.

[0115] The personalized threshold adaptive unit's individual feature modeling processor creates a personal profile for each driver. After confirming the driver's identity via fingerprint or facial recognition, the corresponding profile is loaded. The profile records the driver's age, gender, driving experience, baseline heart rate, baseline blink frequency, and other physiological baseline parameters. For drivers under 30, the average baseline heart rate is 78 beats per minute with a blink frequency of 17 times per minute; for drivers over 50, the baseline heart rate is 72 beats per minute with a blink frequency of 14 times per minute. During the first week of system operation, the physiological and behavioral characteristics of drivers in a conscious state are recorded daily, and the individual baseline mean and standard deviation are calculated. Driver A, aged 30, has a baseline heart rate mean of 76 with a standard deviation of 5 and a PERCLOS baseline of 0.08; Driver B, aged 52, has a baseline heart rate of 70 with a standard deviation of 4 and a PERCLOS baseline of 0.12. The individual feature model stores accumulated fatigue event data over 30 days, including occurrence time, fatigue level, triggering characteristics, and intervention effects. Personalized modeling is initiated when the sample size reaches 50 or more cases. The dynamic threshold adjustment processor adjusts the fatigue judgment threshold based on individual models. For fatigue-sensitive drivers, the risk threshold is lowered by 10%, from the standard 0.30 to 0.27, allowing the system to issue warnings earlier. For fatigue-tolerant drivers, the threshold is raised by 8% to 0.32 to avoid excessive warnings and causing resentment. The age correction coefficient is set according to age groups; for drivers over 50 years old, the fatigue threshold is lowered by 15% overall to improve warning sensitivity. Gender difference correction found that female drivers blink more frequently on average than males, requiring gender normalization of blinking characteristics. The context perception processor acquires contextual parameters such as current time, driving duration, road type, weather, and lighting conditions to establish a context-threshold mapping table. Between 2 AM and 5 AM, the human physiological rhythm is at its lowest point, and the fatigue threshold is lowered by 20%, from 0.30 to 0.24. For continuous driving exceeding 2 hours, the threshold is lowered by 5% for every additional 30 minutes, reaching 0.25 after 4 hours. The threshold is lowered by 12% for monotonous highway environments and raised by 8% for complex mountain road environments. The visibility reduction threshold is lowered by 10% in rainy weather and by 15% in foggy weather. Scenario parameters are updated in real time, and the thresholds are dynamically adjusted to ensure optimal early warning performance under various conditions.

[0116] The event recording and labeling processor of the system's self-learning and performance optimization unit records 15 fields in the event database after each warning event is triggered, including event ID, timestamp, driver ID, fatigue level, risk index, trigger feature vector, warning method, driver response time, response operation type, and post-event verification results. Response time is measured by monitoring steering wheel operation, pedal operation, or turn signal activation; the average effective response time is 18 seconds. Post-event verification judges the accuracy of the warning by observing whether dangerous driving behaviors such as lane departure or sudden braking occur within 15 minutes after the warning. A true positive warning indicates the driver is indeed fatigued and has taken rest measures; a false positive warning indicates the driver is not actually fatigued and continues normal driving for 15 minutes without abnormalities. Out of 300 events, there were 256 true positives and 44 false positives, with an accuracy rate of 85.3%. The model incremental learning processor initiates incremental training every 100 newly labeled samples, using mini-batch gradient descent to update the fusion network parameters. The learning rate is set to 10% of the initial training rate (0.0001) to avoid catastrophic forgetting. During incremental training, 90% new samples and 10% historical samples are mixed to maintain the model's memory of old knowledge. After 20 epochs of training, accuracy improved by 3.2% on new samples and only decreased by 0.5% on historical samples, achieving stable evolution. Threshold parameters are updated every 200 hours of operation. The optimal threshold is calculated based on the most recent 1000 events to maximize accuracy. Optimal thresholds are found using a grid search within the range of 0.25 to 0.35 with a step size of 0.01. The current optimal thresholds are 0.29 for mild and 0.59 for severe cases. Performance evaluation and iterators generate monthly performance reports, statistically analyzing seven key metrics: accuracy, recall, precision, F1 score, false positive rate, false negative rate, and average alert time. Accuracy reached 89.6% in the first month, improved to 91.3% in the second month, and reached 94.1% in the third month, showing a continuous optimization trend. Recall improved from an initial 82.4% to 90.7%, and precision improved from 85.3% to 92.5%. The performance iterator automatically adjusts the optimization strategy based on the evaluation results. When the accuracy does not improve for two consecutive weeks, it increases the learning rate to accelerate convergence. When the false positive rate exceeds the 5% threshold, it increases the classification confidence threshold to reduce false positives. After six months of system operation, the accumulated event samples exceeded 2,000, and the model performance stabilized, achieving an excellent level of 96.2% accuracy, 2.8% false positive rate, and 3.1% false negative rate. Compared with the initial deployment, the accuracy improved by 6.6 percentage points, fully validating the effectiveness of the self-learning mechanism.

[0117] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of protection of the present invention. Furthermore, the elements therein can be updated as technology develops.

Claims

1. A driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology, characterized in that, It includes a multimodal information perception and acquisition module, a data preprocessing and quality assurance module, a fatigue state identification and assessment module, and an intelligent early warning and adaptive optimization module; The multimodal information perception and acquisition module is responsible for acquiring driver status, driving behavior and environmental information in real time; the data preprocessing and quality assurance module is responsible for spatiotemporal alignment, quality assessment and feature standardization of multi-source data; the fatigue state identification and assessment module is responsible for fusing multi-dimensional features and judging fatigue level and risk trend; and the intelligent early warning and adaptive optimization module is responsible for implementing graded intervention and continuously optimizing system performance. The multimodal information perception and acquisition module includes a driver physiological state monitoring unit, a driving behavior feature extraction unit, and an environment and vehicle state perception unit. The driver physiological state monitoring unit is used to collect physiological and visual features in real time. The driving behavior feature extraction unit monitors driving behavior indicators and extracts behavioral features through the vehicle-mounted sensing system. The environment and vehicle state perception unit obtains the vehicle's geographical location, driving speed, acceleration, and route trajectory. The data preprocessing and quality assurance module includes a multi-source data spatiotemporal synchronization unit, a data quality assessment and anomaly detection unit, and a feature extraction and standardization unit. The multi-source data spatiotemporal synchronization unit uses the BeiDou satellite positioning timestamp as a unified benchmark to synchronize the time and calibrate the spatial coordinates of data information from different sensors. The data quality assessment and anomaly detection unit performs integrity verification and reliability assessment on the raw data collected by the sensors. The feature extraction and standardization unit transforms the preprocessed multimodal data into a unified feature vector space. The fatigue state identification and assessment module includes a multimodal feature fusion calculation unit, a fatigue level classification and discrimination unit, and a fatigue trend prediction and risk assessment unit. The multimodal feature fusion calculation unit is used to perform feature-level fusion of standardized physiological features, visual features, and driving behavior features. The fatigue level classification and discrimination unit classifies the driver's current state into three levels: alert, mild fatigue, and severe fatigue. The fatigue trend prediction and risk assessment unit is used to predict the fatigue evolution trend within a certain time window in the future. The intelligent early warning and adaptive optimization module includes a graded early warning and multimodal intervention unit, a personalized threshold adaptive unit, and a system self-learning and performance optimization unit. The graded early warning and multimodal intervention unit implements differentiated warnings through multiple channels based on the fatigue level discrimination results. The personalized threshold adaptive unit establishes a dynamic fatigue threshold model based on individual driver differences and historical behavior data. The system self-learning and performance optimization unit records and evaluates the effects of fatigue events, intervention measures, and driver responses throughout the process. The multimodal feature fusion computing unit includes a feature-level fusion processor, an attention mechanism processor, and a risk index calculator. The feature-level fusion processor is used to input standardized physiological feature vectors, visual feature vectors, and driving behavior feature vectors into a shared representation learning network, extract high-order semantic features of each modality through multi-layer nonlinear transformation, and achieve deep fusion of cross-modal information in the feature space. The attention mechanism processor is used to dynamically allocate the weights of different modal features in the fusion process, strengthen feature channels strongly correlated with fatigue state, suppress the interference of noise and redundant information, and explore the temporal dependencies and complementarities between multimodal data. The risk index calculator inputs the fused comprehensive feature vector into a fully connected layer and an activation function, and outputs a continuous comprehensive fatigue risk index through regression calculation. The attention mechanism processor calculates the fatigue gradient attention score of each modality at time t according to the following formula: ; Where Q(t) is the query vector at time t, and K i (t) is the key vector of the i-th mode, d k To query the dimensions of the vector and key vector, softmax() is the normalization function. The gradient modulation intensity coefficient, Let F be the time gradient of the fatigue risk index at time t. i (t) refers to the eigenvector of the i-th mode. These are the mean eigenvectors of the three modes; The feature-level fusion processor weights and fuses the feature vectors of the three modalities based on attention scores to obtain a fused feature vector F. fus (t).

2. The driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology as described in claim 1, characterized in that, The attention mechanism processor calculates the cooperative gain coefficient between the two modes according to the following formula: ; Among them, MI(F i F j |s) represents the mutual information between the eigenvectors of the i-th mode and the j-th mode under state s. ftg Indicates a state of fatigue, s awk Indicates a state of wakefulness. This represents the cross-covariance of the eigenvectors of the i-th mode and the j-th mode within the current time window. Let represent the autovariance of the eigenvector of the i-th modality.

3. The driver fatigue monitoring and assistance system based on BeiDou satellite positioning and multimodal data fusion technology as described in claim 2, characterized in that, The risk index calculator calculates the comprehensive fatigue risk index r(t) according to the following formula: ; in, For the sigmoid activation function, r ins (t) represents the instantaneous risk term, r tem (t) represents the time-series risk term. As a weighting factor, The modulation intensity coefficient is the cooperative gain. It is a cooperative gain modulation function; The instantaneous risk term is obtained by mapping the fused feature vector, the temporal risk term is obtained by the changes in the historical risk index, and the cooperative gain modulation function is calculated according to the following formula: ; Where M is the number of modes, w ij Let be the weights of the mode pair (i, j).