HUD-based driver behavior-oriented prediction system and abnormity early warning system

By collecting multimodal data on the HUD and utilizing the cross-modal attention mechanism and the Informer model, the problem of insufficient multimodal data fusion in the driver monitoring system is solved, high-precision prediction and real-time warning of driver behavior are achieved, and driving safety is improved.

CN120635867APending Publication Date: 2025-09-12NANJING BOTUO VISION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510747444.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing driver monitoring system suffers from insufficient multimodal data fusion, resulting in low driver behavior prediction accuracy.

Method used

By installing a camera on the HUD to collect the driver's multimodal data in real time, including facial expressions, eye movement data, head posture and upper body posture, and using the cross-modal attention mechanism and Informer model for feature extraction and behavior prediction, efficient fusion and time series modeling of multimodal data can be achieved.

Benefits of technology

It improves the accuracy and reliability of driver behavior prediction, can promptly identify abnormal behaviors such as fatigue and distraction, reduce the risk of misjudgment and missed reports, and improve driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635867A_ABST
    Figure CN120635867A_ABST
Patent Text Reader

Abstract

The invention discloses an HUD-based driver-oriented behavior prediction system and an abnormal early warning system, and belongs to the technical field of intelligent driving safety, and the system comprises a data collection and preprocessing module, a feature extraction module, a cross-modal attention mechanism module and a behavior prediction module. The data acquisition and preprocessing module is used for acquiring multi-modal data of a driver in real time and preprocessing the multi-modal data to obtain preprocessed multi-modal data; the feature extraction module performs feature extraction based on the preprocessed multi-modal data; the cross-modal attention mechanism module is used for calculating an attention weight between modals and carrying out dynamic weighted fusion on features of each modal data in each time window to obtain cross-modal features; the behavior prediction module predicts a behavior of the driver based on the cross-modal features. According to the system, multi-modal information fusion and accurate time sequence modeling are realized, and the behavior prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent driving safety technology, and specifically relates to a driver behavior prediction system and an abnormality warning system based on HUD. Background Art

[0002] With the increasing complexity of road traffic and the continued growth of vehicle ownership, driving safety has become a global social issue. According to the U.S. National Highway Traffic Safety Administration, over 80% of traffic accidents are directly related to driver distraction, fatigue, or operational errors. Therefore, using intelligent technologies to accurately monitor driver behavior and predict risks has become a key breakthrough in improving active safety. Current driver state monitoring systems mostly rely on single-modal sensing technologies, such as vision-based facial expression analysis, eye movement data, head posture detection, or upper body posture detection. While these technologies provide basic early warning capabilities in specific scenarios, they still have significant limitations in complex real-world driving environments. Existing systems typically analyze and judge data from a single source, resulting in a lack of comprehensive assessment of driver state. For example, facial expressions can reflect a driver's emotional state, but they are not sufficient to assess their level of attention. Eye tracking can provide information on driver attention, but is less responsive to emotional changes. Head posture can reveal a driver's posture, but lacks sensitivity to emotion or attention. Furthermore, upper body posture can provide information on a driver's body position, which often changes when a driver is fatigued, distracted, or sitting in an awkward position. However, relying solely on upper-body posture is insufficient to fully assess a driver's condition. Therefore, effectively integrating data from multiple sources, such as facial expressions, eye movements, head posture, and upper-body posture, and fully tapping into the complementary information between these modalities, has become a key challenge in the field of intelligent driving safety technology. Summary of the Invention

[0003] This invention addresses the problem of insufficient multimodal data fusion and low driver behavior prediction accuracy in existing driver monitoring systems. Based on existing electronic, artificial intelligence, and computer vision technologies, this invention proposes a system that uses a driver-facing camera mounted on a HUD (Head-Up Display) to predict driver behavior within an intelligent driving system, thereby improving driver safety.

[0004] The present invention proposes a driver behavior prediction system, which includes a data acquisition and preprocessing module, a feature extraction module, a cross-modal attention mechanism module and a behavior prediction module; the data acquisition and preprocessing module is used to collect the driver's multimodal data in real time, including facial expression data, eye movement data, head posture data and upper body posture data, and preprocess the multimodal data to obtain preprocessed facial expression data, eye movement data, head posture data and upper body posture data; the feature extraction module performs feature extraction based on the preprocessed facial expression data, eye movement data, head posture data and upper body posture data to obtain facial expression features, eye movement timing features, head posture dynamic features and upper body posture features; the cross-modal attention mechanism module is used to calculate the attention weights between modalities, and dynamically weighted fuse the facial expression features, eye movement timing features, head posture dynamic features and upper body posture features to obtain cross-modal features; the behavior prediction module predicts the driver's behavior based on the cross-modal features.

[0005] Furthermore, the data acquisition and preprocessing module: HUDs are auxiliary tools for vehicle driving and are increasingly becoming standard features for vehicles such as cars and motorcycles. Their inward-facing position is often the optimal location for capturing frontal video of the driver's driving experience. Furthermore, in most cases, the HUD is large enough to accommodate a small GPU computer, providing additional analytical computing power. Therefore, the present invention utilizes a camera mounted on the driver-facing portion of the HUD to collect real-time data on the driver's facial expressions, eye movements, head posture, and upper body posture. A Jetson Nano microcomputer with a GPU is also added to provide data computing power.

[0006] The data acquisition and preprocessing module also uses video from the HUD's driver-facing camera to analyze the driver's multimodal data, including facial expressions, eye movements, head posture, and upper body posture. This data is then normalized and time-aligned to address issues arising from differences in data formats and sampling accuracy across different data sources. Normalization ensures that data from all sources are processed at the same scale, ensuring data consistency and comparability. Furthermore, by synchronizing time steps, the data acquisition and preprocessing module effectively eliminates temporal inconsistencies among different data sources, enabling accurate fusion of multimodal data. Furthermore, the data acquisition and preprocessing module partitions continuous time series data into multiple time windows to ensure accurate input for subsequent time series modeling and feature extraction. This windowing approach improves data quality, reduces false positives and missed detections, and enhances the accuracy and reliability of abnormal behavior prediction. Through these preprocessing steps, the system provides high-quality input data for driver status monitoring and abnormal behavior prediction.

[0007] Furthermore, the feature extraction module uses convolutional neural networks (CNNs), long short-term memory networks (LSTMs), 3D convolutional neural networks (3D-CNNs), and ResNet-18 to extract high-dimensional features from multimodal data, including facial expressions, eye tracking, head pose, and upper body posture. Specifically, the CNN is used to extract spatial features of facial expressions, the LSTM is used to analyze the temporal changes in eye movement data to obtain temporal features, and the 3D-CNN is used to process the spatial dynamic features of head pose data to obtain dynamic features. The ResNet-18 is used to extract features of upper body posture, particularly changes in shoulder, back, and head position, to help monitor driver fatigue or distraction. Through comprehensive analysis of this multimodal data, the feature extraction module not only accurately reflects the driver's immediate state but also captures potential changes in behavioral patterns. This provides high-quality feature input for subsequent multimodal fusion and abnormal behavior prediction, thereby enhancing the reliability and intelligence of the driver monitoring system.

[0008] Furthermore, the cross-modal attention module dynamically weights features from different modalities, such as facial expression data, eye movement data, head pose data, and upper body pose data, ensuring that the system effectively integrates relevant information across modalities. In the cross-modal attention module, the attention weights of each modality relative to the other modalities are first calculated. These weights are automatically adjusted through the attention mechanism to assign appropriate weights to the different modalities at each time step. Specifically, features from modalities such as facial expression, eye movement, head pose, and upper body pose are dynamically weighted at each time step to ensure that each modality's contribution to the driver's state at the current moment is accurately reflected. In this way, the cross-modal attention module integrates information such as facial expression, eye movement, head pose, and upper body pose, fully exploiting the interdependencies between modalities and thereby improving the accuracy of abnormal behavior prediction. The cross-modal attention module effectively addresses the potential limitations of a single modality, avoiding misjudgments caused by insufficient or inaccurate information from a particular modality, and ensuring the efficient integration of multimodal data in abnormal behavior prediction.

[0009] Furthermore, the behavior prediction module uses the Informer model to perform time series modeling and abnormal behavior prediction on the fused cross-modal feature features. Through the sparse attention mechanism of the Informer, the behavior prediction module can effectively process long time series data, capture long-term dependencies in driver behavior, and predict potential abnormal behaviors of the driver (such as fatigue, distraction, anxiety, etc.). Existing models often have difficulty capturing potential abnormal signals when dealing with gradual changes and long-term trends in driver behavior, resulting in low prediction accuracy. The Informer model, through its unique structure, can effectively capture the long-term dependencies of driver behavior and accurately predict the occurrence of abnormal behavior, thereby improving the accuracy of abnormal behavior prediction and the ability to provide early warning, thereby providing drivers with more reliable safety protection during driving.

[0010] Furthermore, the data acquisition and preprocessing module is responsible for collecting multimodal data of the driver from multiple data sources, and performing normalization and time alignment on it to ensure data quality and consistency. First, the system collects data from facial expressions, eye tracking, head posture, and upper body posture. Then, it normalizes this data and uses interpolation methods to align the data time. Finally, the continuous time series data is cut into multiple time windows. This module solves the problems of inconsistent data formats and asynchronous time steps, providing accurate and synchronized input data for subsequent modules. The specific processing flow is as follows:

[0011] Step 1, data collection

[0012] The system uses the video captured by the HUD's camera facing the driver to analyze and collect the driver's facial expressions, eye movement data, head posture and upper body posture data in real time.

[0013] Among them, facial expression image X face It has a fixed resolution and color channel. Its data can be obtained by referring to the facial expression detection module in the literature (Fang Daosheng. Development of driver fatigue recognition and warning system based on facial expressions [D]. Nanjing Forestry University, 2023. DOI: 10.27242 / d.cnki.gnjlu.2023.001110.).

[0014] Eye Tracking DataX eye The key features of eye movement, such as blinking frequency and eye position, are recorded. The specific data can be obtained by referring to the eye feature extraction part in the reference (Wang Haitong. Research on driver behavior recognition method based on eye movement prediction [D]. Xidian University, 2024.).

[0015] The head posture data X headIt contains information such as pitch angle and yaw angle. The specific data can be obtained by referring to the literature (Gao Qiang, Ding Bingru, Dong Changlin, et al. Head posture estimation algorithm based on real-time application [J]. Journal of Shenyang University (Natural Science Edition), 2025, 37(01): 25-33+99+93. DOI: 10.16103 / j.cnki.21-1583 / n.2025.01.008).

[0016] In addition, the upper body posture data X torso The Openpose tool, which converts human images into skeletal data (the principle and publicly available source code are available at https: / / github.com / CMU-Perceptual-Computing-Lab / openpose4), converts these images into skeletal data, which records the driver's upper body posture, such as changes in shoulder and back position, and any awkward sitting postures like slouching or leaning forward. This data can provide additional information on driver fatigue and distraction, and is crucial for accurately predicting potentially dangerous driver behavior.

[0017] These data provide the basis for subsequent driver status analysis and lay an important foundation for accurately monitoring driver behavior and predicting abnormal behavior.

[0018] Step 2: Data preprocessing

[0019] To ensure that data from different modalities can be effectively processed at the same scale, data normalization, interpolation and time window division are required. First, all collected data will be normalized to ensure that data from different data sources can be fused and analyzed at the same scale:

[0020]

[0021] Among them, X raw is the original data of a data source, X raw ∈{X face , X eye , X head , X torso}, X norm_raw represents the normalized data, μ and σ are the mean and standard deviation respectively.

[0022] Next, in order to solve the time alignment problem of different data sources, interpolation methods are used to fill in missing data to ensure that the data of each modality are comparable at the same time step. At the same time, in order to capture the driver's short-term behavior pattern and improve data processing efficiency, it is necessary to divide the time series data into windows. Assume that the total length of data collected by the system is T total, can be divided into fixed window sizes w, each window contains w consecutive time steps. The data of each time window is represented as follows:

[0023]

[0024] in:

[0025] X face (t) is the facial expression data at time step t;

[0026] X eye (t) is the eye movement data at time step t;

[0027] X head (t) is the head posture data at time step t;

[0028] X torso (t) is the data of the upper body posture at time step t;

[0029] t0 is the starting time step of the window;

[0030] Time window division helps to extract the driver's behavioral characteristics in different time periods and provide effective input for subsequent model training.

[0031] Step 3: Multimodal data fusion

[0032] In each time window, data from different data sources such as facial expressions, eye tracking, head posture and upper body posture are weighted fused to form a multimodal feature vector This vector effectively integrates the information of each mode and is convenient for capturing the driver's behavioral characteristics. Assume that the fusion feature vector of each time window is expressed by the following formula:

[0033]

[0034] in, is the fused multimodal vector, and Fusion(·) is a fusion operation. In the present invention, facial expression data, eye movement data, and head posture data are weightedly fused.

[0035] After completing the feature fusion of each time window, the vectors of all time windows are Integrate into a complete multimodal time series dataset X in chronological order multi , whose expression is:

[0036]

[0037] Right now:

[0038]

[0039] in:

[0040] X face (t)∈R H*W*C is the facial expression data at time step t;

[0041] is the eye movement data at time step t;

[0042] is the head posture data at time step t;

[0043] is the upper body posture data at time step t;

[0044] T is the total length of the time series, representing all time steps collected.

[0045] In this way, the final X multi Contains multimodal data of multiple time steps, each time step This reflects the comprehensive behavioral information from various data sources, providing analysis and processing for the subsequent cross-modal attention mechanism and behavior prediction module. Furthermore, to improve the accuracy of abnormal driver behavior detection, a feature extraction method based on multimodal data is proposed. This method integrates facial expressions, eye movement data, head posture, and upper body posture information, and utilizes a deep learning model to extract high-dimensional features, providing comprehensive input for abnormal behavior prediction.

[0046] In the feature extraction module, facial expression features are extracted through convolutional neural networks (CNN). CNN can recognize facial contours and expression changes, thereby reflecting the driver's fatigue, anxiety and other states. Eye movement data is modeled through long short-term memory networks (LSTM). LSTM can capture timing features such as blinking frequency and eye movement, helping to detect abnormal behaviors such as distraction. Head posture features are extracted through 3D convolutional neural networks (3D-CNN). 3D-CNN can analyze head movement trends and identify abnormal postures such as lowering the head and tilting the head. Upper body posture features are extracted through ResNet-18, especially changes in shoulder, back and head position, to help monitor whether the driver is in a state of fatigue or distraction.

[0047] At each time step t, facial expression, eye movement, and head pose features from different data sources are integrated into a multimodal feature vector:

[0048] F multi (t)={F face (t),F eye (t),F head (t),F torso (t)}

[0049] in: For facial expression features; is the eye movement feature; is the head posture feature, It is the posture characteristic of the upper body.

[0050] This multimodal feature vector serves as input for subsequent modal fusion and abnormal behavior prediction. These features comprehensively describe the driver's static facial expressions, eye movements, and spatial dynamic posture, providing high-quality data support for accurately assessing the driver's state and improving prediction accuracy.

[0051] Furthermore, the present invention proposes a cross-modal attention mechanism module. This module maps facial expression features, eye movement timing features, head posture dynamic features, and upper body posture features into corresponding query (Q), key (K), and value (V) representations, calculates attention weights between modalities, and implements dynamic weighted fusion. This method integrates key information from each modality, captures intermodal dependencies, ensures the integrity and accuracy of the fused features, and provides high-quality input for abnormal behavior prediction.

[0052] The features of each modality are mapped into three representations: query Q, key K, and value V through a specific linear transformation. The query vector is used to calculate the degree of attention the current modality pays to other modalities, the key vector represents the correlation between modalities, and the value vector contains the core information of the modal features. This step maps the features of different modalities into a unified space, enabling them to interact effectively in subsequent attention calculations and laying the foundation for dynamically adjusting the weights of modal features.

[0053] Calculate the attention weight w between modalities ij , expressed as:

[0054]

[0055] w ij represents the weighting coefficient of mode i to mode j, i, j∈{face,eye,head,torso};

[0056] The weighted fusion feature of each modality is expressed as

[0057] F fused,i (t)=∑ j w ij F j (t);

[0058] F fused,i (t) represents the weighted fusion feature of modality i, F j (t) represents the j-modal feature;

[0059] Cross-modal feature Ffused (t) is expressed as:

[0060] F fused (t)=∑ i F fused,i (t).

[0061] Furthermore, the behavior prediction module of the present invention uses the Informer model for time series modeling to predict the driver's behavior. The input of the behavior prediction module is the cross-modal fusion feature F fused (t), and outputs the predicted driver behavior (normal or abnormal) for each time step. Through an encoder-decoder architecture, Informer leverages a sparse attention mechanism to efficiently model temporal features, capturing the dynamic changes in driver status and optimizing them using a cross-entropy loss function. After training, Informer can predict driver status in real time, providing precise support for abnormal behavior detection and early warning, thereby improving driving safety.

[0062] The Informer model uses an encoder-decoder structure to process the input temporal cross-modal features F fused (t) to efficiently model long-term dependencies. The encoder part calculates the correlation between input features through a sparse attention mechanism and generates a weighted temporal feature representation H. This process is implemented by the following formula:

[0063] H=Encoder(F fused (t))=Attention(F fused (t))

[0064] Among them, Attention(·) calculates the correlation of features between different time steps, which improves the model's sensitivity to key state changes and thus enhances the ability to capture long-term dependencies.

[0065] The decoder combines the hidden features output by the encoder with the input of the current time step to generate the future driver behavior prediction Y pred (t) = Decoder(H), which outputs the prediction result. This output can be used for binary classification tasks (such as normal / abnormal), allowing the system to more accurately identify changes in the driver's state.

[0066] Furthermore, in order to optimize the parameters of the behavior prediction module, the cross-entropy loss function (Cross-EntropyLoss) is used to evaluate the error between the prediction result and the true label:

[0067]

[0068] Where: N is the total number of time steps; Y true(t) is the true behavior label (0: normal, 1: abnormal); Y pred (t) is the model's predicted value. Through backpropagation and parameter optimization, the Informer model gradually reduces prediction errors and improves its behavior prediction capabilities in complex driving scenarios.

[0069] The trained behavior prediction module can predict the driver's behavior status in real time:

[0070]

[0071] For anomaly detection tasks, the prediction results can serve as an early warning of abnormal driver behavior, issuing timely warnings to help drivers respond to potential risks early, thereby improving driving safety and system response capabilities.

[0072] The present invention also proposes a driver behavior abnormality warning system, which includes all modules in the driver behavior prediction system and an alarm mechanism module. The behavior prediction results of the driver behavior prediction system are input into the alarm mechanism module. The alarm mechanism module determines whether to trigger an alarm or take emergency intervention measures based on the prediction results output by the behavior prediction module. When the prediction result is abnormal behavior, the alarm mechanism will be triggered to remind the driver to pay attention to safety, thereby improving the reliability and practicality of the system and ensuring the safety of the driver and passengers.

[0073] Beneficial effects: By implementing the present invention, users can solve three problems in driver abnormal behavior prediction: insufficient multimodal information fusion, long-term dependence on time series data, and low accuracy in abnormal behavior prediction.

[0074] Efficient multimodal data fusion: Through a cross-modal attention mechanism, the system can adaptively fuse facial expression, eye tracking, head posture, and upper body posture information, dynamically adjusting the weight of each modality to avoid misjudgments caused by a single modality. For example, fatigue can be detected through facial expressions, distraction can be identified through eye movement data, head posture can determine whether the driver maintains a correct sitting posture, and upper body posture can provide additional information about whether the driver is sitting correctly and whether fatigue is present. In particular, when the driver exhibits poor posture such as lowering the head, drooping shoulders, or leaning the upper body forward, these changes are often signs of fatigue or distraction, effectively assisting the system in making accurate judgments. By integrating information from each modality, the cross-modal attention mechanism can dynamically adjust the weight of each modality, ensuring a comprehensive assessment of the driver's condition and improving the system's recognition accuracy and real-time warning capabilities for abnormal behavior.

[0075] Accurate Time Series Modeling: Utilizing the Informer model and a sparse attention mechanism, the system efficiently processes long-term dependencies and accurately captures changing trends in driver behavior. For example, fatigue and distraction typically develop gradually. Informer can effectively predict abnormal driver behavior based on historical data and accurately identify its changing trends, thereby improving the reliability of time series modeling. This mechanism enables the system to track long-term dependencies in driver behavior, promptly identify potential danger signals, and improve the accuracy of abnormal behavior prediction.

[0076] High-precision abnormal behavior prediction: Combining a cross-modal attention mechanism with the informer model ensures the effective integration of multimodal information, while improving the accuracy of anomaly detection and reducing the risk of missed and false positives. By comprehensively considering multimodal data such as facial expressions, eye tracking, head posture, and upper body posture, the system can more comprehensively assess the driver's condition and monitor the emergence of abnormal behaviors such as fatigue, distraction, and anxiety in real time. Based on the prediction results, the system can trigger timely warnings, prompting the driver to take appropriate safety measures, significantly improving driving safety and system reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 This is an overall flow chart of data processing using the driver behavior abnormal warning system of the present invention;

[0078] Figure 2 This is a flow chart of processing driver behavior data collected by the HUD using the system of the present invention;

[0079] Figure 3 It is a flow chart of data processing by the data acquisition and preprocessing module of the present invention;

[0080] Figure 4 It is a multimodal feature extraction flow chart;

[0081] Figure 5 This is a flowchart of how the cross-modal attention mechanism model processes data;

[0082] Figure 6 This is the prediction flow chart of the behavior prediction module. DETAILED DESCRIPTION

[0083] The present invention is further explained below with reference to the figures and specific implementation examples.

[0084] Drawing on existing artificial intelligence and computer vision technologies, this paper proposes a vehicle driver behavior prediction and anomaly warning system for intelligent driving safety technology based on a cross-modal attention mechanism and an informer model. This system integrates multimodal features from multiple data sources, such as facial expressions, eye movement data, head posture, and upper body posture, using a cross-modal attention mechanism. It also utilizes an informer model to model and predict time series data, effectively predicting driver behavior and accurately predicting abnormal behavior.

[0085] The present invention is a driver behavior prediction system based on multimodal fusion, which includes a data acquisition and preprocessing module, a feature extraction module, a cross-modal attention mechanism module and a behavior prediction module.

[0086] The driver behavior abnormality warning system of the present invention includes all the modules in the behavior prediction system and the alarm mechanism module. The overall flow chart of using the driver behavior abnormality warning system to sort out the data is as follows: Figure 1 shown.

[0087] Among them, the cross-modal attention mechanism module and the behavior prediction module are the main innovations of the present invention, which effectively improve the accuracy of predicting abnormal driver behavior.

[0088] The system of the present invention includes multiple modules working together to achieve efficient driver monitoring and abnormal behavior prediction. The data acquisition and preprocessing module processes facial expression, eye tracking and head posture data through normalization and window cutting to ensure the high quality and consistency of input data. The feature extraction module uses deep learning models such as CNN, LSTM, 3D-CNN, Resnet18, etc. to accurately extract high-dimensional features of multimodal data and capture changes in driver status. The driver behavior monitoring data processing process of its HUD is as follows: Figure 2 As shown. The cross-modal attention mechanism module dynamically weights and fuses these features to adaptively adjust the influence of each modality, enhance information complementarity, and avoid misjudgment of a single modality. Subsequently, the behavior prediction module adopts the Informer model and uses the sparse attention mechanism to efficiently model the driver's long-term behavioral changes, accurately predict abnormal states such as fatigue, distraction, and anxiety, and improve early warning capabilities. Finally, the alarm mechanism module determines whether to trigger the alarm mechanism based on the previous prediction results, ensuring that the alarm is accurate and reliable, and comprehensively improving driving safety. Specifically, it includes the following steps:

[0089] In the vehicle driver monitoring system, data acquisition and preprocessing are the key first steps, which directly affect the subsequent feature extraction, modal fusion and abnormal behavior prediction. In order to ensure the efficiency and accuracy of the entire system, the data acquisition and preprocessing module needs to perform multimodal data normalization, interpolation and filling, and time window division on the data from different data sources. Figure 3 The specific steps are as follows:

[0090] Step 1: On-the-go driver monitoring systems rely on multiple data sources to collect physiological and behavioral data from the driver, including facial expressions, eye movements, and head posture. However, these data sources vary significantly in data format, temporal resolution, and sampling accuracy. For example, eye trackers can be unstable in bright light, head posture data sources can be affected by ambient vibration, and facial recognition can be affected by lighting or angle. These factors highlight the limitations of single modalities and the need for enhanced multimodal fusion. Therefore, ensuring data collection accuracy while uniformly processing this data and ensuring its efficient use in multimodal fusion becomes a core issue in system design. To further enhance the system's comprehensive monitoring of driver behavior, incorporating a camera on the HUD to capture the driver's upper body posture is particularly important. Upper body posture provides information about the driver's body position, especially when the driver is fatigued, distracted, or sitting in an incorrect posture. These changes are often reflected in the dynamic performance of the upper body. Capturing this information through the HUD camera not only supplements facial expression and head posture data, but also provides additional clues about whether the driver is in the correct sitting posture and whether he or she is fatigued, thereby enhancing the comprehensive understanding of the driver's physiological and behavioral state and improving the accuracy and robustness of the driver monitoring system.

[0091] 1.1 Facial Expression Data Collection: Facial expressions reflect the driver's emotional state and can effectively reveal the driver's psychological and physiological reactions, such as fatigue, anxiety, anger, etc. To ensure that the system can capture this information in a timely manner, the on-board camera monitors the driver's facial expression changes in real time.

[0092] Facial expression data is collected using in-vehicle cameras mounted around the driver's seat. These cameras are typically positioned in front of the driver, enabling them to clearly capture the driver's facial expressions in various driving scenarios. These cameras must be adaptable to varying lighting conditions and require high resolution and wide-angle lenses to ensure the captured image captures the driver's entire face.

[0093] The camera captures the driver's facial expressions at a fixed frequency, usually as a continuous video stream or static image frames. Each frame records the driver's facial expression characteristics, including changes in areas such as the eyes, mouth, and eyebrows. Assume that the original facial expression captured by the on-board camera is:

[0094] X face ∈R H*W*C

[0095] Where: H is the height of the image (in pixels); W is the width of the image (in pixels); C is the number of channels of the image (RGB three color channels).

[0096] Assume that facial expression data is collected using a car camera and the sampling frequency is set to 10 frames per second. The data collection time is five minutes, so the time step number T of data collection will be:

[0097] T = 5 minutes × 60 seconds × 10 frames / second = 3000 frames

[0098] Therefore, the dimension of the collected data is X face ∈R 3000*H*W*C Each frame of the image records the changes in the driver's facial expression, and the data will be processed and analyzed in real time by the system within five minutes to identify whether the driver is fatigued or exhibiting other abnormal behaviors.

[0099] 1.2. Eye Tracking Data Collection: Eye tracking is a key technology for detecting driver attention. By monitoring eye movement, gaze position, and blink frequency, it is possible to determine whether the driver is fatigued, distracted, or experiencing other abnormalities. Therefore, eye trackers play a crucial role in data collection.

[0100] Eye trackers are typically mounted in front of the driver or on the interior mirror, using infrared technology to track eye position and movement. They must be highly accurate and responsive, capable of tracking the driver's head or eyes even when they rapidly move.

[0101] Eye trackers generate eye movement data by recording information such as eye movement trajectory, blink frequency, and gaze point shift. This data is collected in the form of a time series, recording changes in eye position and blink frequency.

[0102] Assume that the data collected by the eye tracker is:

[0103]

[0104] Where: T is the number of time steps (each time step corresponds to a sampling point, for example, if the sampling is 100 times per second, T is the total number of sampling time points); D eye The dimensions of the eye tracking data (such as blink frequency, eye position, gaze point, etc.).

[0105] The eye features are obtained using the eye feature extraction part in the reference (Wang Haitong. Research on driver behavior recognition method based on eye movement prediction [D]. Xidian University, 2024.), with a sampling frequency of 10 times per second and a collection time of five minutes. The time step number T = 5 × 60 × 10 = 3000, so the collected data dimension is During this five-minute period, the system can monitor the driver's attention state in real time to determine whether there are signs of fatigue or distraction.

[0106] 1.3 Head posture data collection: Head posture data provides important information about the driver's body posture, such as whether the driver is sitting properly, looking down, or turning their head. Abnormal head posture often indicates driver distraction or fatigue, so head posture collection is an integral part of the behavior prediction system.

[0107] The driver's head posture changes are recorded in real time, including pitch angle (angle of head tilt up and down), yaw angle (angle of head rotation left and right), and roll angle (angle of head rotation along its own axis). These data are recorded in time series to provide basic information for subsequent state prediction and abnormal behavior warning. Assume that the head posture data is:

[0108]

[0109] Where: T is the number of time steps, the same as the eye tracking data; D head are the characteristic dimensions of head posture (such as pitch angle, yaw angle, and roll angle).

[0110] 1.4. Upper body posture data collection: Upper body posture is a key indicator for assessing a driver's body posture and fatigue level. This is especially true when a driver has an improper sitting posture, lowers their head, or has an unstable posture. These changes are often reflected in the dynamic performance of the upper body. By monitoring the driver's upper body posture, it can help determine whether the driver is in the correct sitting position and whether there are signs of fatigue or distraction. To achieve this goal, the camera on the HUD is used to capture the dynamic changes of the driver's upper body in real time.

[0111] HUD cameras are typically mounted near the front window or dashboard inside the vehicle, ensuring full coverage of the driver's upper body movements. Using image recognition technology, the camera analyzes the driver's upper body posture, including the relative positions of the shoulders, back, and head, and detects any abnormal postures such as leaning, lowering the head, or hunching the back. This data is collected in real time as a video stream, and key features are extracted using computer vision algorithms.

[0112] Assume that the data collected by the HUD camera is:

[0113]

[0114] Where: T is the number of time steps (each time step corresponds to an image frame, for example, 30 frames are sampled per second, and T represents the total number of sampling time points); D torso It is the dimension of the upper body posture data (including the relative positions of the shoulders, back and head, and detecting whether there are abnormal postures such as tilting, lowering the head, and bending the back).

[0115] Assuming the HUD camera collects upper body posture data at a sampling rate of 10 frames per second for five minutes, the number of time steps is T = 3000. During these five minutes, the system captures each frame and uses image processing algorithms to analyze the driver's upper body posture, identifying any signs of poor posture or fatigue.

[0116] In this way, the HUD camera can capture and analyze the driver's upper body posture, generating a corresponding data stream that, along with other modal data (such as facial expressions, eye movements, and head posture), provides a comprehensive assessment of the driver's condition. Step 2: After data acquisition, the raw data from different modalities undergoes multi-step preprocessing to ensure data quality. These steps include normalization, alignment, and time window cutting.

[0117] 2.1, Data normalization

[0118] In order to ensure that different modal data can be fused and analyzed at the same scale, all collected data will be normalized. This not only helps to improve the system's processing efficiency for data from different data sources, but also avoids performance degradation caused by different dimensions. Assume that X raw is the original data of a data source, X raw ∈{X face , X eye , X head , X torso}, normalized data X norm_raw It can be expressed as:

[0119]

[0120] where μ and σ are the original data X raw Through this normalization process, data from different modalities will be unified to the same scale for subsequent analysis and processing.

[0121] 2.2, Interpolation Filling

[0122] Different data sources (such as facial expressions, eye tracking, head pose, and upper body pose) may have different sampling frequencies, resulting in data misalignment in time. For example, facial expression images may be sampled at 30 frames per second, while eye trackers, head pose, and upper body pose may be sampled at a higher frequency. Therefore, in order to synchronize processing, the data timestamps of all data sources need to be aligned to the same time step through interpolation methods.

[0123] The interpolation method can calculate the filling value by the following formula:

[0124]

[0125] Assume that the eye tracking data has missing values ​​at a certain time step t, that is, X eye If (t) = NAN, the data of the surrounding time steps can be used for interpolation filling. Specifically, the interpolation method for eye tracking data can calculate the filling value using the following formula:

[0126]

[0127] This interpolation method can effectively fill in missing values ​​in the data and ensure that the data from each data source are comparable at the same time step.

[0128] 2.3, Division of time windows

[0129] To facilitate subsequent feature extraction and model training, the continuous time series data is divided into multiple time windows. The choice of window size has a significant impact on system performance. Shorter time windows can help capture rapidly changing behavioral patterns (for example, momentary fatigue or distraction), while longer time windows help capture long-term driver behavioral trends. Therefore, choosing a reasonable time window length (for example, 5 seconds or 10 seconds) is crucial to accurately capture changes in driver behavior.

[0130] Assuming that a time window starts at t0 and the size of the window is w, the data of each time window can be expressed as:

[0131]

[0132] in:

[0133] X face (t) is the facial expression data at time step t;

[0134] X eye (t) is the eye movement data at time step t;

[0135] X head (t) is the head posture data at time step t;

[0136] X torso (t) is the data of the upper body posture at time step t;

[0137] t0 is the starting time step of the window;

[0138] w is the size of the window, and the present invention selects 5 seconds as the time length of the window.

[0139] The data in each time window contains data from multiple time steps, which helps to extract the driver's behavioral characteristics in different time periods and provide effective input for subsequent model training.

[0140] 3. Multimodal data splicing;

[0141] Within each time window, data from different data sources (facial expressions, eye movements, head posture, and upper body posture) are fused into a single feature vector. This process ensures that information from different modalities can be effectively combined within the same time window, fully exploring the interrelationships between the modalities.

[0142] Assume that for a certain time window, the fused data It can be calculated by the following formula:

[0143]

[0144] in:

[0145] is the fused multimodal data vector;

[0146] Fusion(·) means splicing and fusing facial expressions, eye movement data, head posture, and upper body posture.

[0147] After completing the feature fusion in each time window, the fusion vector of multiple time windows is obtained. These vectors describe the behavior of the system at successive time steps, but only in local time windows. Therefore, in order to perform global temporal modeling, all of these Integrate in time order to construct a multimodal time series dataset X containing multiple time steps multi .

[0148] Suppose there is a time series T (a total of T time steps), each time step corresponds to a Then the feature vectors of all windows can be concatenated into a long time series:

[0149]

[0150] Right now

[0151]

[0152] in:

[0153] x face (t)∈R H*W*C is the facial expression data at time step t;

[0154] is the eye movement data at time step t;

[0155] is the head posture data at time step t;

[0156] is the data of the upper body posture at time step t;

[0157] T is the total length of the time series, representing all time steps collected.

[0158] Among them, X multi is the final multimodal dataset, which contains data from multiple time steps. It can be a weighted fusion of features from different data sources, representing comprehensive behavioral information at that moment. This multimodal dataset will serve as input for subsequent feature extraction, model training, and behavior prediction.

[0159] The feature extraction module extracts the driver's behavioral characteristics from the collected unprocessed raw data. Through deep learning models (such as convolutional neural networks (CNN), long short-term memory networks (LSTM), 3D convolutional neural networks, Resnet18, etc.), the system can extract high-dimensional feature representations from different modalities (facial expressions, eye movement data, head posture). The feature representation of each modality will serve as input for subsequent multimodal fusion and abnormal behavior prediction, helping the system to better understand the driver's behavior patterns, thereby improving the accuracy and reliability of abnormal behavior detection. The overall flow chart is as follows: Figure 4 The specific steps are as follows:

[0160] 4. Use the feature extraction module to extract features from the multimodal raw data of driver status monitoring

[0161] 4.1 Facial Expression Feature Extraction

[0162] Facial expressions are important signals reflecting a driver's emotional state, revealing psychological and physiological reactions such as fatigue, anxiety, and anger. Convolutional neural networks (CNNs) are widely used in image processing, particularly in facial expression recognition. CNNs can extract spatially hierarchical features from facial expressions, particularly adept at capturing local information and complex nonlinear relationships within an image. These features are highly sensitive to dynamic changes in facial expressions, making them effective in identifying driver emotional fluctuations.

[0163] The input facial expression is X face (t) is captured by the vehicle camera or other sensors at time t and is represented as a three-dimensional tensor R H*W*C , H represents the height of the image (number of pixels), W represents the width of the image (number of pixels), and C is the number of color channels of the image (usually 3, representing RGB color channels).

[0164] A convolutional neural network processes an input image through a series of convolutional layers, pooling layers, and nonlinear activation functions, extracting key features from the image layer by layer. The following is a detailed description of the feature extraction process, clearly outlining the functions and roles of each layer.

[0165] Low-level features: The first three convolutional layers are primarily used to extract basic image features, such as edges and simple texture information. These layers use convolution operations to scan the input image using multiple convolution kernels, identifying facial contours and other basic information. Specifically, the first three convolutional layers are responsible for:

[0166] Extract the edges of facial features such as eyes, eyebrows, and mouth;

[0167] Extract basic texture features from images, such as the boundary information of facial contours.

[0168] These basic features provide the basis for subsequent layers to extract more detailed facial information.

[0169] The fourth through sixth convolutional layers are primarily used to extract higher-level features of facial expressions and identify changes in the driver's emotions. After processing through the first three convolutional layers, the network is already able to recognize basic facial features. Subsequent layers gradually focus on the details of emotional changes, such as subtle changes in expressions such as anger, anxiety, and fatigue. These layers detect subtle movement patterns of facial muscles, such as drooping mouth corners or furrowed brows, to help the network determine the driver's emotional state.

[0170] Pooling layers follow each convolutional layer to reduce the size of the feature map while retaining the most important information. Pooling operations (such as max pooling) help reduce computational overhead and improve the network's robustness to small-scale changes, such as slight shifts in facial position. Pooling layers not only reduce the computational burden but also maintain stability in facial expressions during recognition.

[0171] Non-linear activation function: A non-linear activation function (such as ReLU) is applied after each convolution layer to introduce non-linear features and help the network learn more complex relationships.

[0172] After multiple convolutional layers and pooling layers, the input image X face (t) is gradually converted into high-dimensional feature representation, namely:

[0173] F face (t) = CNN(X face (t))

[0174] F fact The (t) representation contains key information about facial expressions. These high-dimensional features, including basic facial contours, emotional fluctuations, and subtle changes in facial muscles, can be effectively used for driver behavior analysis and prediction. The resulting feature representation can be used for tasks such as emotion classification and fatigue detection, providing strong support for driving safety warning systems.

[0175] Through CNN's hierarchical feature extraction, the network gradually extracts key information about facial expressions, from simple low-level features (such as edges and contours) to complex high-level features (such as emotional expressions). This layered extraction approach helps the system accurately capture the driver's emotional changes, thereby improving the accuracy of driver monitoring and behavior prediction.

[0176] 4.2 Eye Tracking Feature Extraction

[0177] Eye movement data reflects a driver's attention state and is crucial for predicting driver behavior and providing early warning of abnormal behavior. Long Short-Term Memory (LSTM) networks are recurrent neural networks that excel at processing time series data and effectively capture long-term dependencies in the data. Eye tracking data often exhibits temporal characteristics, such as blink frequency and changes in eye position. These characteristics can accumulate over time, especially during fatigue or distraction. To better capture the temporal characteristics of eye movement data, a double-layer LSTM is employed in this system to enhance temporal modeling capabilities.

[0178] In the system, eye tracking data X eye (t) is the eye movement feature captured at time step t, expressed as:

[0179] X eye (t)∈R eye

[0180] Among them, X eye (t) contains multiple features of eye tracking data, such as eye position, blink frequency, gaze offset, etc. The two-layer LSTM model extracts and models the temporal features of these eye movement data through two LSTM layers to generate corresponding feature representations.

[0181] The processing of LSTM can be expressed as:

[0182] F eye (t) = LSTM2(LSTM1(X eye (t)))

[0183] in, It is a feature representation after double-layer LSTM processing, which contains the high-dimensional time series features of eye movement data.

[0184] In a two-layer LSTM, two stacked LSTM layers work together to model the temporal sequence of eye movement data. The functions of each layer are as follows:

[0185] The first LSTM layer is used to extract the original eye movement data X eye (t) extracts preliminary temporal features. This layer learns short-term temporal dependencies in the input data, such as changes in blink frequency and short-term fluctuations in eye position. Through this layer, the network can identify relatively simple temporal patterns in eye movement data, providing rich temporal information for the input of the second LSTM layer.

[0186] Second LSTM layer: The second LSTM layer further processes the output of the first LSTM layer to extract deeper temporal features. This layer is designed to capture long-term dependencies, such as the gradual accumulation of fatigue or prolonged periods of inattention. The second LSTM layer abstracts more complex temporal patterns from the features extracted by the first layer, allowing it to identify driver fatigue or distraction.

[0187] Each LSTM layer processes and updates its state through its gating mechanism (input gate, forget gate, and output gate). These gating mechanisms enable LSTM to dynamically adjust the degree of attention paid to different information at each moment, thereby capturing long-term dependencies in eye movement data.

[0188] Input gate: The input gate determines which input information should be written into the memory cell. It filters the input data through the sigmoid activation function and generates a "gated" output with a value between 0 and 1, indicating the "importance" of the current input.

[0189] Forget Gate: The forget gate determines which memories should be forgotten. At each time step, the forget gate performs a "forget" operation on the information in the memory unit, ensuring that the network can dynamically and selectively discard information that is no longer useful.

[0190] Output gate: The output gate determines the output at the current moment. It calculates the information in the memory unit and generates the output features of the current moment based on the input and the hidden state of the previous moment.

[0191] Through these gating mechanisms, LSTM can capture correlations between different time steps in eye movement data. For example, when a driver blinks frequently or shifts their eyes excessively, LSTM can identify these abnormal patterns by updating its internal state, thereby determining whether the driver is fatigued or distracted.

[0192] The two-layer LSTM effectively captures long-term dependencies in eye movement data through multi-level temporal modeling. The first LSTM layer extracts short-term temporal features, while the second layer further captures complex long-term dependencies. This hierarchical structure enables the system to more accurately determine whether the driver is experiencing abnormal conditions such as fatigue or distraction. The LSTM gating mechanism further enhances the model's ability to recognize eye movement data by dynamically adjusting its focus on different information, providing powerful temporal analysis capabilities for driver behavior monitoring systems.

[0193] 4.3 Head posture feature extraction

[0194] Head pose data can reflect the driver's body posture. This can be particularly noticeable when the driver is fatigued or distracted, as evidenced by changes in head posture, such as prolonged head bowing or frequent head tilts. These changes often distract the driver's attention. Therefore, extracting dynamic head pose characteristics using a 3D convolutional neural network (3D-CNN) is a crucial component of driver behavior monitoring in this system. 3D-CNN effectively captures both spatial and temporal head pose information, enabling identification of improper driver posture.

[0195] Head posture data X head (t) represents the head posture information captured at time step t, which usually includes the pitch angle, yaw angle and roll angle of the head. It is a three-dimensional tensor expressed as:

[0196] X head (t)∈R head

[0197] Among them, X head (t) contains the head posture features captured at each time step t. In order to extract the dynamic changes of head posture, 3D-CNN extracts spatiotemporal features from this data. The processing process is as follows:

[0198] F head (t) = 3DCNN(X head (t))

[0199] in, It is the head posture feature after 3D-CNN processing, which represents the abstract feature of the head posture at time step t.

[0200] 3D-CNNs typically consist of multiple convolutional layers, pooling layers, and activation function layers. Each layer performs a different function, gradually extracting and abstracting the spatiotemporal features of head pose.

[0201] The first convolutional layer: This layer uses a 3D convolution kernel to convolutionally transform the input data X head(t) Perform preliminary feature extraction, mainly capturing the basic shape and simple spatiotemporal features of the head. The first convolution layer usually extracts features including basic angular changes of the head, such as preliminary changes in pitch and yaw angles.

[0202] Second convolutional layer: Based on the preliminary features extracted by the first layer, the second convolutional layer further processes these features to identify more complex spatiotemporal patterns. For example, the second convolutional layer may capture the trend of head posture changes over different time periods, such as when a driver lowers or tilts their head for a long period of time.

[0203] Third and deeper convolutional layers: As the number of network layers increases, the 3D convolutional layers can gradually capture more complex dynamic changes. For example, the third and subsequent convolutional layers can identify subtle changes in the driver's head posture caused by fatigue or inattention, further improving the accuracy of feature extraction.

[0204] Each convolutional layer uses convolution kernels to perform convolution operations on the spatial and temporal dimensions, extracting spatiotemporal features at different levels. The feature map after the convolution operation is passed to the next convolutional layer. Through multiple layers of convolution and pooling, higher-dimensional and more complex features are gradually abstracted.

[0205] Through multiple layers of 3D convolution operations, 3D-CNN is able to extract increasingly abstract features layer by layer. Each convolution kernel learns different levels of spatiotemporal features, including spatial features of head posture (such as head position and angle) and temporal features (such as the changing trend of head posture). As the network depth increases, the convolutional layers are able to capture more complex dynamic changes, such as whether the driver's head is shaking due to fatigue or other unfavorable posture.

[0206] 4.4 Upper Body Posture Feature Extraction

[0207] Upper body posture data can reflect the driver's body position, especially when fatigued, distracted, or sitting in an awkward position. These changes often manifest as dynamic shifts in the position of the shoulders, back, and head, such as drooping shoulders, a rounded back, or a lowered head. To effectively extract these upper body posture features, the ResNet-18 network was used. It utilizes residual connections to improve the network's feature extraction capabilities, making it particularly well-suited for processing upper body posture image data.

[0208] In the process of upper body posture feature extraction, the input data is the image data of the upper body, which is expressed as:

[0209]

[0210] Among them, X torso (t) is the upper body image data captured at time step t, which is normalized and resized before being passed into the ResNet-18 network as input for processing. T represents the number of time steps, D torso is the depth of the upper body posture image (the number of channels, usually 3, representing the RGB color channels). These image data capture the posture information of the driver's upper body, such as the relative position changes of the shoulders, back, and head.

[0211] ResNet-18 is a deep residual network. Its core concept is to address the vanishing gradient and information loss problems in deep network training by introducing residual connections. ResNet-18's feature extraction process is primarily accomplished through multiple convolutional layers and residual blocks. Each residual block incorporates convolution operations, batch normalization, ReLU activation functions, and residual connections.

[0212] The specific processing process is as follows:

[0213] In ResNet-18, the first 6 convolution layers (including the initial convolution layer and the first few residual blocks) are mainly used to transform the input image X torso (t) Extract the underlying features of the upper body image. The main task of the first six layers is to capture the basic morphology, texture, and positional relationship of the face and upper body. These layers have the following functions:

[0214] Edge features: Extract basic contours from upper body images, such as edge information of shoulders, back, and head.

[0215] Local texture: Local texture features are extracted through convolution kernels, such as simple shape changes of shoulders and back.

[0216] Low-level morphological information: Identify general shapes in the image, such as the position of shoulders, the curve of the back, etc.

[0217] In these layers, the convolution kernel processes the input image through a sliding window, gradually extracting high-level visual information from local features. After processing through these layers, the network is able to capture basic information about the upper body posture, which serves as input for further processing in subsequent layers.

[0218] The deep convolutional layers in ResNet-18 (including layers 7 and above) are used to extract more complex upper body posture changes from the underlying features. Through deep residual blocks, the network is able to learn more complex and abstract features, especially small posture changes. Specifically, the deep convolutional layers have the following functions:

[0219] Complex postural changes: Changes in mood or posture such as drooping shoulders, arching the back, or lowering the head.

[0220] Subtle body dynamics: These layers learn through multiple layers of convolution to extract more complex dynamic features, such as the changing patterns of the upper body when tired, and the details of lowering the head or poor sitting posture.

[0221] In these deep residual blocks, the convolution operation extracts more complex posture information through the deeper convolution kernel. The residual connection ensures that the network can retain low-level features and avoid the gradient disappearance problem in deep networks. Assume that the input feature of the kth residual block is X k , the output after the convolution operation is F k , then the output Y of this block k for:

[0222] Y k =F k +X k

[0223] Among them, X k is the input of the residual block, F k is the output of the convolution operation. The residual connection adds the input features to the features after the convolution operation to ensure that information can be effectively transmitted and avoid the gradient vanishing problem.

[0224] The residual blocks in the network continuously perform deep feature learning, and the output Y k It is passed to the next layer. As multiple residual blocks are stacked, the network is able to extract increasingly abstract posture features. After several residual blocks, the network aggregates the feature maps through the global average pooling layer. The global average pooling layer averages all the feature values ​​of each channel to generate a fixed-length feature vector that represents the overall information in the image. The purpose of global average pooling is:

[0225] Reduce the number of parameters: thereby avoiding overfitting;

[0226] Provide concise feature representation: compress high-dimensional features into a fixed-length vector for easy subsequent processing.

[0227] After processing by the residual block and the global average pooling layer, the network generates a high-dimensional feature vector F torso (t) represents the abstract features of the driver's upper body posture. This feature vector contains key information about the upper body, such as behavioral changes such as fatigue, lowering the head, or poor sitting posture. It is specifically expressed as:

[0228] F torso (t) = Resnet18(X torso (t))

[0229] in, It is the output upper body posture feature, which represents the high-dimensional feature of the upper body posture. torso is the dimension of the feature vector, which is usually large to fully capture the diversity of upper body postures.

[0230] By using the ResNet-18 network, the network is able to effectively extract high-level feature information from upper-body posture images. The introduction of residual connections allows the network to retain more low-level features during training and avoid the gradient vanishing problem in deeper networks. This enables the network to maintain high accuracy and robustness when extracting complex upper-body posture features, and performs particularly well in identifying behavioral changes such as fatigue, lowering the head, or poor sitting posture. ResNet-18, through its multi-layer convolution and residual connection mechanism, effectively extracts high-level features from upper-body posture images and is particularly suitable for handling posture changes and complex body dynamics. The network performs deep feature learning through multiple residual blocks and generates fixed-length feature vectors through global average pooling. These features can accurately reflect changes in the driver's upper-body posture, such as dynamic changes in the position of the shoulders, back, and head, providing strong support for subsequent driver behavior prediction.

[0231] These features from different sources provide multi-dimensional information about the driver's state. By fusing them at each time step t, the system can more comprehensively understand the driver's behavioral state. This multimodal fusion can help the system make more accurate judgments in complex driving environments. For example, when the driver's facial expression shows fatigue, changes in eye movement characteristics and head posture may further confirm this fatigue state, thereby improving the accuracy of abnormal behavior prediction. By fusing multimodal features such as facial expressions, eye movements, head posture, and upper body posture, the system can more comprehensively and accurately assess the driver's behavioral state, providing more reliable driver prediction and abnormal behavior warning capabilities. This combination of multimodal information not only reduces the potential misjudgment caused by a single modality, but also provides stronger judgment in more complex situations.

[0232] The cross-modal attention mechanism module proposed in this paper combines the key information of each modality through weighted summation, dynamically adjusts the weights of different modalities, and ensures the contribution of each modality in the final feature representation. This fusion method can effectively capture the dependencies between modalities, allowing the model to dynamically adjust the attention paid to different modal features at each time step, thereby more accurately reflecting the driver's actual state. The flowchart of the cross-modal attention mechanism module processing data is shown in the figure. Figure 5 shown.

[0233] 5.1, Query, Key and Value Representation

[0234] The features of each modality are mapped into three representations: query Q, key K, and value V through a specific linear transformation. The query vector is used to calculate the degree of attention the current modality pays to other modalities, the key vector represents the correlation between modalities, and the value vector contains the core information of the modal features. This step maps the features of different modalities into a unified space, enabling them to interact effectively in subsequent attention calculations and laying the foundation for dynamically adjusting the weights of modal features.

[0235] Through the cross-modal attention mechanism, the system can dynamically calculate the weight of each modality at the current time step based on the similarity between the modalities. When each modal feature is weighted and summed, the weight is adjusted according to its importance, so that more critical modal features have a greater influence on the fusion result. The weight calculation process is as described above. Based on the dot product calculation of the query (Q) and the key (K), the degree of attention of each modality to other modalities can be quantified. The specific formula is:

[0236]

[0237] Q i is the query vector of the i-th modality, representing the characteristics of the modality;

[0238] K j is the key vector of the jth mode, representing the characteristics of another mode;

[0239] N j It is a scaling factor of the feature dimension, which is used to avoid computational instability caused by feature dimension differences.

[0240] For example: the weight of facial modality to eye movement modality:

[0241]

[0242] These weights are used to measure the degree of dependence between different modalities, so that the fusion features can dynamically adjust the influence of each modality, thereby improving the accuracy of abnormal behavior prediction.

[0243] 5.2, Weighted Sum among Multimodal

[0244] In the multimodal fusion process, we calculate the attention value Attention(Q i ,K j ) to dynamically adjust the contribution of modal features. Specifically, the attention score Attention(Q i ,K j ) reflects the similarity between mode i and mode j, so it can be used as the weighting coefficient of mode i to mode j characteristics, that is, weight w ij .Right now:

[0245] wij =Attention(Q i ,K j )

[0246] Because data sources from different modalities (such as facial expressions, eye movements, and head posture) provide information from different perspectives, a single modality may not fully reflect the driver's behavioral state. For example, facial expressions can provide information about emotional state, eye movement data can reveal whether the driver is focused, and head posture reflects whether the driver's posture is appropriate. Therefore, when fusion is performed, it is necessary to dynamically adjust the weights of each modality so that the information from different modalities can complement each other and form a more complete representation of driver behavior.

[0247] After obtaining the weighted weights of each modality, the system uses these weights to perform a weighted summation of the features of each modality to obtain the fusion features of each modality. The specific formula is as follows:

[0248] F fused,eye (t) = w eye,face F face (t)+w eye,eye F eye (t)+w eye,head F head (t)+w eye,torso F torso (t)

[0249] F fused,head (t) = w head,face F face (t)+w head,eye F eye (t)+w head,head F head (t)+w head,torso F torso (t)

[0250] F fused,torso (t) = w torso,face F face (t)+w torso,eye F eye (t)+w torso,head F head (t)+w torso, torso F torso (t)

[0251] F fused,eye (t) = w eye,face F face (t)+w eye,eye F eye (t)+w eye,head F head (t)+weye,torso F torso (t)

[0252] Weighted fusion feature F fused,face (t), F fused,eye (t), F fused,head (t) and F fused,torso (t) represents the comprehensive information of each modality and other modalities at the current moment, providing rich input for subsequent behavior prediction.

[0253] After weighted summation, the fusion features of the four modalities of facial expression data, eye movement data, head posture and upper body posture will be integrated into a unified cross-modal feature representation. This representation integrates the key information of different modalities and maintains the information complementarity between the modalities.

[0254] F fused (t) = F (fused,face) (t)+F (fused,eye) (t)+F (fused,head) (t)+F (fused,torso) (t)

[0255] in:

[0256] F fused,face (t) is the weighted fusion feature of facial expression modality features.

[0257] F (fused,eye) (t) is the weighted fusion feature of eye movement data.

[0258] F (fused,head) (t) is the weighted fusion feature of the head posture modal features.

[0259] F (fused,torso) (t) is the weighted fusion feature of the upper body posture modality features.

[0260] At this time, the cross-modal feature representation F cross (t) By integrating key information from each modality, the system comprehensively captures the driver's behavioral state, providing strong input support for subsequent behavioral prediction. Through this weighted summation and dynamic adjustment between modalities, the system can more accurately understand the driver's state and improve the accuracy of behavior prediction.

[0261] 6. The behavior prediction module is a core component of the autonomous driving system. It is responsible for predicting whether the driver has abnormal behaviors such as fatigue, distraction, anxiety, anger, etc. based on the fusion features from the cross-modal attention mechanism module. This module uses the Informer model, an efficient time series prediction model designed for long time series data and high-dimensional features. The Informer model performs time series modeling based on the self-attention mechanism, which can capture the long-term and short-term dependencies in the driver's behavior and identify abnormal patterns. Its overall flow chart is as follows Figure 6 As shown:

[0262] This module inputs the multimodal fusion features from the cross-modal attention mechanism module for further processing and analysis to predict whether the driver has abnormal behaviors such as fatigue, distraction, anxiety, and anger.

[0263] The core advantage of the Informer model lies in its ability to process long time series data and efficiently capture dependencies between long time steps. Traditional self-attention models (such as the Transformer) have high computational complexity, especially when processing long time series data, where the computational complexity increases rapidly with the length of the sequence. To address this issue, the Informer model introduces two innovations:

[0264] Sparse attention mechanism: By selectively paying attention to time steps, the computational overhead is reduced, the attention is retained on important time steps, and unimportant time steps are discarded, thereby improving computational efficiency.

[0265] Generative Decoder: By using a generative model to predict future time series data, the system can not only capture current dependencies when processing time series data, but also effectively predict future driver behavior, thereby enhancing prediction accuracy.

[0266] 6.1 Time Series Modeling

[0267] The core of the Informer model is the encoder-decoder structure based on the sparse attention mechanism. Input feature F fused (t) is a multimodal fusion feature representing the driver’s state at each time step. The model first models the time series data through the encoder part, and then generates abnormal behavior predictions for future moments through the decoder part.

[0268] Input sequence length: Input sequence F fused The length of (t) is T in , indicating the past T inThe driver's state information is collected in time steps. Each time step t corresponds to a multimodal fusion feature that describes the driver's eye movements, facial expressions, head posture, and other information at that moment. Assuming that the length of the input sequence is the past 5 minutes (300 time steps), and each time step corresponds to 1 second of sampled data, then:

[0269] T in = 300 time steps

[0270] This means that the input sequence will contain data from the past 300 time steps, with each time step corresponding to the driver's state characteristics within 1 second.

[0271] Prediction time length: based on input sequence F fused (t), the Informer model will output the future T out Behavior prediction results of time steps. That is, the model predicts the future T out Is there any possibility of fatigue, distraction or other abnormal behavior at this moment? Assuming the prediction time length is 3 seconds (3 time steps), then:

[0272] T out = 3 time steps

[0273] This means that based on the driver’s behavior data from the past 5 minutes, the Informer model will predict the behavior (e.g. fatigue, distraction, etc.) within the next 3 seconds.

[0274] 6.2, Encoder part

[0275] The encoder's task is to learn the context of the input sequence and capture long-term dependencies in the time series. Traditional self-attention mechanisms calculate the similarity between all elements in the input sequence, but this computational complexity increases significantly when processing long time series data. To improve efficiency, Informer uses a sparse attention mechanism to optimize this process.

[0276] Specifically, the encoder transforms the input features of each time step into a higher-level feature representation through the following process

[0277] H=Encoder(F fused (t))=Attention(F fused (t))

[0278] Among them, Attention(·) represents the sparse attention mechanism, which can calculate the dependencies between input features and generate weighted feature representations.

[0279] To improve computational efficiency, Informer uses a sparse attention mechanism to selectively calculate the similarity between important time steps. This mechanism only calculates the similarity between a fixed number of key time steps, avoiding the need for full connection calculations for all time steps, thus significantly reducing the amount of computation.

[0280] 6.3, Decoder

[0281] The decoder's task is to predict the driver's behavior at future time steps based on the hidden state sequence generated by the encoder. The core of the decoder is to use a generative model to predict future behavior.

[0282] The decoder outputs the prediction Y pred (t) is the predicted value of the driver's behavior at the future moment. Assume that the output of the decoder is:

[0283] Y pred (t) = Decoder(H)

[0284] Among them, Y pred (t) is the behavior prediction result generated by the decoder. For the behavior prediction task, the output value Y pred (t) can be 0 (normal) or 1 (abnormal).

[0285] The generative decoder combines feature information from the current moment with data from past time steps to predict future driver behavior. If the driver exhibits persistent abnormal eye movements or fluctuating facial expressions over a period of time, the decoder predicts possible fatigue or distraction in the future.

[0286] 7. Model training and optimization

[0287] The training goal of the Informer model is to optimize the model parameters by minimizing the error between the predicted results and the true labels. During the training process, the cross-entropy loss function is used to evaluate the difference between the model's predicted results and the true labels, and then adjust the model parameters. For a binary classification problem, the cross-entropy loss function can be expressed as:

[0288]

[0289] Where N is the total number of time steps, Y true (t) is the true label, Y pred (t) is the predicted value of the model.

[0290] Through the backpropagation algorithm, the gradient of the loss function is used to update the model parameters, gradually reducing the error between the predicted results and the true value, thereby improving the model's predictive ability for complex driving scenarios. This process continuously optimizes the model, making its performance in driver behavior detection tasks more accurate.

[0291] After training and optimization, the Informer model will output the behavior prediction value Y at each time step pred (t), and ultimately provide a prediction of the driver's state. For example:

[0292]

[0293] The system will monitor the driver's behavior in real time based on these predictions and issue an alarm when necessary to alert the driver of potential abnormal behavior.

[0294] Summary: The behavior prediction module uses the Informer model, which utilizes a sparse attention mechanism and a generative decoder to efficiently perform time series modeling and behavior prediction on the fused features from the cross-modal attention module. Through its encoder and decoder architecture, the Informer model captures long-term dependencies in driver behavior and accurately predicts whether the driver exhibits abnormal behavior (such as fatigue, anxiety, and distraction). This module can effectively enhance the safety and reliability of autonomous driving systems, provide comprehensive monitoring and prediction of driver behavior, and further promote the development of intelligent driving technology.

[0295] Furthermore, the alarm mechanism module, the final step in the autonomous driving system, is responsible for determining whether to trigger the alarm mechanism based on the predictions from the behavior prediction module (i.e., the informer model). This module's purpose is to issue timely warnings or trigger automated intervention measures when abnormal driver behavior is detected, based on real-time predictions of driver behavior. This ensures the safety of the driver and passengers and prevents accidents.

[0296] 8. After the behavior prediction module completes processing, the system will output the behavior prediction result Y for each time step t pred (t), its value is:

[0297] Y pred (t)∈{0,1}(normal or abnormal)

[0298] Among them, 0 indicates that the driver's behavior is normal, and 1 indicates that abnormal behavior (such as fatigue, distraction, etc.) is detected. These prediction results will be passed as input to the system's alarm mechanism module.

[0299] The core task of the alarm mechanism module is to determine whether to activate the alarm based on the prediction results. Specifically, the system will analyze the prediction results of each time step t. If the prediction value is 1 (indicating abnormal behavior), an alarm is triggered; if the prediction value is 0 (indicating normal behavior), no alarm is required. The output of the alarm mechanism is the system's alarm status A. alarm (t), indicating whether an alarm needs to be triggered, the output is:

[0300] A alarm (t)∈{0,1}(no alarm or alarm)

[0301] Among them, 0 means no alarm, and 1 means the alarm has been triggered. When the system triggers the alarm, the system will emit a continuous alarm or prompt sound to remind the driver of the current driving status, such as "Driver fatigue, please take a rest" or "Driver distracted, please pay attention".

[0302] The alarm mechanism module analyzes the behavior prediction module's output in real time to ensure the accuracy of alarm signals and minimize the impact of false alarms. Whenever the system predicts abnormal driver behavior, it triggers an alarm and takes appropriate precautionary measures, ensuring the safety of drivers and passengers and reducing the risk of traffic accidents.

[0303] The driver behavior monitoring system and abnormal warning system of the present invention solve three problems in the prediction of driver abnormal behavior: insufficient multimodal information fusion, long-term dependence on time series data, and low accuracy in abnormal behavior prediction.

[0304] (1) Multimodal data fusion: The present invention adopts a cross-modal attention mechanism to dynamically weighted fuse data from different modalities such as facial expressions, eye tracking, head posture, and upper body posture. This mechanism can automatically adjust the weight of each modality according to its relevance, thereby overcoming the misjudgment problem that may be caused by single modality data. For example, when facing changes in lighting, the weight of eye movement data can be increased accordingly to compensate for the shortcomings of facial expressions; when emotional changes are more obvious, the weight of facial expression data can be increased, while the weight of eye movement data can be appropriately reduced. In this way, the system can effectively integrate information from different data sources, comprehensively evaluate the driver's status, and improve the accuracy of the system.

[0305] (2) Capturing long-term dependencies of time series data: The driver's behavior and state are dynamic and often accompanied by long-term dependencies. For example, the driver's fatigue and distraction are often gradual processes, not instantaneous changes. However, existing systems often fail to effectively capture these long-term dependencies when processing time series data. Although traditional time series modeling methods (such as RNN and LSTM) can process time series data to a certain extent, they often find it difficult to effectively capture long-term dependencies when facing long time series data, resulting in false positives or missed positives when the system predicts the driver's potential dangerous behavior. Especially when monitoring the driver's upper body posture, the driver's posture changes, such as lowering the head, bending the back, or shoulder fatigue, are usually gradual, and these changes are often closely related to the driver's fatigue or distraction. : Existing systems fail to fully consider the changes in upper body posture and long-term dependencies, and therefore cannot accurately identify the driver's gradually occurring dangerous behavior, resulting in significant deficiencies in the accuracy and timeliness of the system in complex driving environments, and it is difficult to cope with subtle changes in driver behavior and the challenges of complex environments; in order to solve the problem that traditional methods are difficult to capture long-term dependencies in long time series, the present invention adopts the Informer model. The Informer model efficiently processes long time series data through a sparse attention mechanism and effectively captures long-term dependencies in driver behavior. Compared to traditional LSTM or RNN models, the Informer model is able to better model long-term behavioral trends, thereby accurately predicting potentially dangerous driver behavior. This innovative approach enables the system to promptly detect gradual changes in driver behavior, such as fatigue and distraction, avoiding both false positives and false negatives and improving the accuracy of abnormal behavior prediction.

[0306] (3) High-precision prediction of abnormal behavior: Although some existing driver monitoring systems are able to predict abnormal driver behavior, due to insufficient multimodal data fusion and limited time series modeling capabilities, existing prediction systems often have high false positives and false negatives, resulting in the failure to accurately predict the driver's possible dangerous behavior in a timely manner. For example, when the driver exhibits abnormal behaviors such as fatigue, anxiety, or distraction, the existing system has low prediction accuracy and is difficult to accurately identify these potential dangerous signals, which in turn affects the protection of traffic safety. Especially in complex driving environments, the driver's behavior patterns often show strong individual differences and time series changes. Existing methods are difficult to make accurate behavior predictions based on these characteristics; the present invention combines the cross-modal attention mechanism and the informer model to accurately capture the driver's abnormal behavior based on multimodal data fusion. The cross-modal attention mechanism provides dynamic weighting of different modal information, while the informer model improves the sensitivity to long-term changes in driver behavior through time series modeling. Through this combined method, the system can accurately predict the driver's potential dangerous actions, such as fatigue, distraction, anxiety, etc., significantly improving the prediction accuracy of abnormal behavior and providing more reliable safety protection for the autonomous driving system.

[0307] Therefore, the system of the present invention effectively solves problems such as insufficient multimodal data fusion, inaccurate long-term dependency capture of time series data, and low accuracy in abnormal behavior prediction, significantly improving the accuracy, reliability, and real-time performance of driver abnormal behavior prediction, and further promoting the development of the field of intelligent driving safety technology.

Claims

1. A driver behavior prediction system, characterized in that: It includes data acquisition and preprocessing module, feature extraction module, cross-modal attention mechanism module and behavior prediction module; The data acquisition and preprocessing module is used to collect multimodal data of the driver in real time, including facial expression data, eye movement data, head posture data and upper body posture data, and preprocess the multimodal data to obtain preprocessed facial expression data, eye movement data, head posture data and upper body posture data; The feature extraction module extracts features based on the pre-processed facial expression data, eye movement data, head posture data and upper body posture data to obtain facial expression features, eye movement timing features, head posture dynamic features and upper body posture features; The cross-modal attention mechanism module is used to calculate the attention weights between modalities, and dynamically weighted fuse facial expression features, eye movement timing features, head posture dynamic features, and upper body posture features to obtain cross-modal features; The behavior prediction module predicts the driver's behavior based on the cross-modal features.

2. A driver behavior monitoring system according to claim 1, characterized in that: The data acquisition and preprocessing module is further used to divide the continuous facial expression data, eye movement data, head posture data and upper body posture data into multiple time windows, each time window including w time steps of data.

3. The driver behavior prediction system according to claim 1, characterized in that: The preprocessing includes normalizing the facial expression data, eye movement data, head posture data and upper body posture data respectively to obtain normalized data of each modality, and time aligning the normalized data of different modalities to obtain preprocessed facial expression data, eye movement data, head posture data and upper body posture data.

4. A driver behavior prediction system according to claim 3, characterized in that: In the process of time alignment of normalized data of different modalities, interpolation method is used to fill in missing values.

5. The driver behavior prediction system according to claim 1, characterized in that: The feature extraction module includes a convolutional neural network, a two-layer long short-term memory network, a 3D convolutional neural network and a ResNet-18 network; The convolutional neural network is used to process the preprocessed facial expression data X face (t) Perform feature extraction, expressed as: F face (t)=CNN(X face (t)) F face (t) represents facial expression features; The two-layer long short-term memory network is used to process the preprocessed eye data X eye (t) Perform feature extraction, expressed as: F eye (t)=LSTM2(LSTM1(X eye (t))) F eye (t) represents the temporal characteristics of eye movements; The 3D convolutional neural network is used to process the preprocessed head posture data X head (t) Perform feature extraction, expressed as: F head (t)=3DCNN(X head (t)) F head (t) represents the dynamic characteristics of head posture; The ResNet-18 network is based on the upper body posture data X torso (t) Extract the upper body posture features, expressed as: F torso (t)=Resnet18(X torso (t)) F torso (t) represents the upper body posture feature.

6. A driver behavior prediction system according to claim 1, characterized in that: In the cross-modal attention mechanism module, facial expression features, eye movement timing features, head posture dynamic features, and upper body posture features are mapped to corresponding query, key, and value representations, and the attention weight w between modalities is calculated. ij , expressed as: w ij represents the weighting coefficient of mode i to mode j, i, j∈{face,eye,head,torso}; The weighted fusion feature of each modality is expressed as F fused,i (t)=∑ j w ij F j (t); F fused,i (t) represents the weighted fusion feature of modality i, F j (t) represents the j-modal feature; Cross-modal feature F fused (t) is expressed as: F fused (t)=∑ i F fused,i (t)。 7. The driver behavior prediction system according to claim 1, characterized in that: T in Cross-modal features F of time steps fused (t) is input to the behavior prediction module and outputs T out Behavior prediction for each time step; The behavior prediction module includes an encoder-decoder structure based on a sparse attention mechanism; The encoder transforms the input cross-modal features F fused (t) is transformed into a high-level feature H, expressed as: H=Encoder(F fused (t))=Attention(F fused (t)) Among them, Attention(·) represents the sparse attention mechanism; The decoder outputs the predicted behavior, expressed as: Y pred (t)=Decoder(H) Y pred (t) is the behavior prediction result generated by the decoder, and the behavior prediction result is normal or abnormal.

8. The driver behavior prediction system according to claim 1, characterized in that: The behavior prediction module is trained using a cross entropy loss function, which is expressed as: Among them, L ce represents the cross entropy loss function, N is the total number of time steps, Y true (t) is the true label, Y pred (t) is the predicted value of the model.

9. The driver behavior monitoring system according to claim 1, characterized in that: The head posture data includes the pitch angle, yaw angle and roll angle of the head; The upper body posture data includes the relative positions of the shoulders, back and head, whether the upper body posture is tilted, whether the head is lowered, and whether the back is bent.

10. A driver behavior abnormal warning system, characterized in that: It includes the driver behavior prediction system as described in claim 5, and also includes an alarm mechanism module. The behavior prediction result of the driver behavior prediction system is input into the alarm mechanism module. When the input behavior prediction result is abnormal, the alarm mechanism module outputs an alarm signal.