Driver workload prediction method based on multiple physiological signals under multiple auditory disturbances
By using a multi-physiological signal fusion and stacking model to predict driver workload, this approach solves the reliability and interpretability issues of single indicators in existing technologies, achieving high accuracy and stability prediction under various auditory interferences and supporting the dynamic adjustment of intelligent driving assistance systems.
Patent Information
- Application Number
- CN202511311209.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing technologies for predicting driver workload suffer from low reliability and high uncertainty of single indicators, making it difficult to reflect individual differences. Furthermore, machine learning models lack interpretability and generalization ability, and are particularly unstable in complex traffic environments.
A multi-physiological signal fusion method was adopted, combining skin conductance, electrocardiogram and electroencephalogram signals. Driver workload was predicted using stacked models (including random forest, extreme gradient boosting tree, lightweight gradient booster and classification boosting tree), and the SHAP framework was used for interpretation and analysis to integrate the impact of multiple auditory interferences on driver workload.
It improves the accuracy, universality, and robustness of driver workload prediction, significantly enhances the accuracy of workload prediction, provides interpretability and stability under complex traffic conditions, and supports real-time adjustments of intelligent driving assistance systems.
Smart Images

Figure CN121171550B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent transportation technology, specifically involving a driver workload prediction method based on multiple physiological signals under various auditory interferences. Background Technology
[0002] Music is one of the most common auditory stimuli in vehicles, and its impact on driving behavior, cognitive workload, and psychological state has been extensively studied. With the development of driver assistance systems, voice navigation has become a major source of auditory input in vehicles and is an important factor affecting driving performance. Although previous studies have investigated music and navigation voice independently, research examining them simultaneously in driving situations remains limited, and the mechanisms by which they jointly influence driver workload, cognition, and behavior are still unclear.
[0003] Driver workload is typically assessed using three types of methods: physiological measurements, performance-based measurements, and subjective assessments. Physiological methods characterize workload through changes in the driver's physiological signals during driving. Widely used indicators include electrocardiogram (ECG), electroencephalogram (EEG), electrical skin activity (EDA), eye-tracking data, facial expressions, functional near-infrared spectroscopy, and respiratory parameters. Performance-based methods assess workload based on vehicle control behavior, primarily using longitudinal and lateral driving indicators. Common measurements include speed variability, lane position, acceleration and deceleration frequency, steering angle, pedal position, and lateral deviation. Subjective methods rely on self-reported workload and typically use standardized rating scales. NASA-TLX, MCH, and SWAT are widely used, but current technology has not yet combined them for comprehensive assessment.
[0004] Many workload metrics are sensitive to changes in the traffic environment; therefore, relying on a single metric can lead to low reliability, high uncertainty, and difficulty in reflecting individual differences. Recently, an increasing number of studies have employed multi-source data fusion to construct more comprehensive workload assessment systems, achieving promising results. For example, combinations of eye-tracking, electrocardiography (ECG), and vehicle control; combinations of electroskin activity, ECG, and subjective assessment; combinations of eye-tracking and facial temperature; and combinations of electroencephalography (EEG) and driving behavior have all performed well in characterizing workload. Angka et al. further confirmed that multi-source methods generally outperform single-metric methods in terms of accuracy. While multi-source fusion methods offer better adaptability and robustness, standardized evaluation criteria for different metrics remain lacking. Some studies directly apply previously used metrics without considering their applicability in different traffic environments, leading to inconsistent performance. Furthermore, some studies have excluded physiological metrics due to cost constraints, improving system feasibility but sacrificing accuracy. In addition, different workload metrics vary in terms of measurability, feasibility, and reliability. Despite significant progress in workload prediction using machine learning models, two major limitations remain. First, many models are "black box" models, lacking interpretability and making it difficult to understand how input features affect predictions—a problem particularly important in high-risk applications such as driving safety. Second, the generalization and accuracy of single models in complex traffic environments remain unstable. Therefore, employing a combination of dynamic and task-specific strategies to develop more adaptive and representative multi-source workload assessment models specifically for diverse traffic environments holds great market potential. Summary of the Invention
[0005] The purpose of this application is to address the problems of the prior art and to provide a method for predicting driver workload based on multiple physiological signals under various auditory interferences.
[0006] To address the technical problem, the technical solution of this application is: a driver workload prediction method based on multiple physiological signals under various auditory interferences, comprising the following steps:
[0007] Step 1: Scene Setup; Combine various auditory disturbances with traffic scenes to create multiple driving scenarios;
[0008] Step 2: Data Acquisition; Collect driving behavior data from multiple drivers operating driving simulators in various driving scenarios, and simultaneously collect skin conductance signals, electrocardiogram signals, and electroencephalogram signals from multiple drivers, and statistically analyze the subjective workload assessments of multiple drivers for multiple driving scenarios.
[0009] Step 3: Data processing;
[0010] Step 3-1: Perform anomaly removal and filtering on the collected driving behavior data, and then process the driving behavior data by mean and standard deviation to obtain driving behavior features;
[0011] Step 3-2: The collected skin conductance signals, electrocardiogram signals, and electroencephalogram signals are processed to obtain physiological characteristics;
[0012] Step 3-3: Classify the subjective workload assessment to obtain workload categories, and assign driving behavior features and physiological features to the corresponding workload categories to obtain the processed dataset;
[0013] Step 4: Establish a stacking model;
[0014] Step 4-1: Basic learner selection: Select four tree models as basic learners: Random Forest, Extreme Gradient Boosting Tree, Lightweight Gradient Boosting Machine, and Classification Boosting Tree.
[0015] Step 4-2: Basic learner training: Optimize each tree model through grid search to determine the optimal hyperparameters; then input the processed dataset into each optimized tree model to generate preliminary prediction results;
[0016] Step 4-3: Meta-learner training: The preliminary prediction results are used to construct the meta-feature matrix Z, which serves as the input to the meta-learner, and the workload category serves as the output of the meta-learner.
[0017] Step 4-4: Train the meta-learner to obtain the prediction model for the workload category and the final prediction result;
[0018] Step 5: Model Result Interpretation; Perform SHAP interpretation analysis on the trained prediction model and the final prediction results. Use the meta-feature matrix Z as the interpretation feature, and quantify the contribution of each interpretation feature to the driver's workload prediction through the SHAP framework. This is used to predict the impact of driving behavior features and physiological features on workload under various auditory interferences.
[0019] Preferably, in step 1, the various auditory interferences include navigation modes and music rhythms. The navigation modes include high-frequency navigation and low-frequency navigation, the music rhythms include no music, slow-paced music and fast-paced music, and the traffic scenarios include regular roads, school roads and construction site roads. By arranging and combining various auditory interferences with traffic scenarios, 18 driving scenarios are obtained.
[0020] Preferably, in step 2, the driving simulator is a fixed-base driving simulator, which includes a steering wheel, brake pedal, accelerator pedal, three displays, and a main controller. The main controller is equipped with UC-Win / Road software to simulate the driving environment and record driving behavior data at a sampling rate of 25 Hz. The skin conductance signal is collected through a skin resistance sensor, with the two electrodes of the skin resistance sensor attached to the driver's index and middle fingers respectively, and the skin conductance signal is collected at a frequency of 50 Hz. The electrocardiogram (ECG) signal is collected through an ECG sensor, which is attached to the driver's chest area and records the raw heart signal at a frequency of 200 Hz. The electroencephalogram (EEG) signal is collected through the EPOC X portable EEG headset at a sampling rate of 128 Hz. The X portable EEG headset includes 14 EEG sensors and reference sensors P3 and P4. EEG sensors AF3, AF4, F3, F4, F7, and F8 correspond to the frontal lobe, EEG sensors P7 and P8 correspond to the parietal lobe, EEG sensors O1 and O2 correspond to the occipital lobe, and EEG sensors T7, T8, FC5, and FC6 correspond to the temporal lobe. Each channel of the EPOC X portable EEG headset records 16-bit high-resolution data with a sensitivity of 0.1275µV, accurately capturing weak brainwave fluctuations.
[0021] Preferably, step 3-1 specifically involves: identifying and removing outliers in the physiological time series using a standard deviation-based outlier detection method, filtering with a Savitzky-Golay filter, and then processing the driving behavior data by mean and standard deviation to obtain driving behavior features; the driving behavior data includes driving speed, lateral acceleration, longitudinal acceleration, lane departure, road centerline departure, accelerator pedal force, and brake pedal force; the driving behavior features include average speed and speed standard deviation, average lateral acceleration and lateral acceleration standard deviation, average longitudinal acceleration and longitudinal acceleration standard deviation, mean and standard deviation of lane departure, mean and standard deviation of road centerline departure distance, average accelerator pedal force and accelerator pedal force standard deviation, and average brake pedal force and brake pedal force standard deviation.
[0022] Preferably, step 3-2 includes the following steps:
[0023] Step 3-2-1: Baseline standardization is performed by subtracting the average resting state of each driver from the collected skin conductance, electrocardiogram, and electroencephalogram signals;
[0024] Step 3-2-2: Use the three sigma rule to identify and remove standardized outliers;
[0025] Step 3-2-3: Use a combination of high-pass and low-pass filters to filter signal artifacts, with the frequency range set according to the signal characteristics;
[0026] Step 3-2-4: Smooth the signal using Savitzky-Golay filtering;
[0027] Step 3-2-5: Apply spline interpolation to fill the small gaps caused by missing values to obtain physiological characteristics.
[0028] Preferably, step 3-3 specifically includes:
[0029] The NASA-TLX scale was used to collect subjective workload assessments from multiple drivers for multiple driving scenarios. The subjective workload assessment was based on a 6-point scale. Based on the subjective workload assessment, the multiple driving scenarios were classified into workload categories, namely low workload, medium workload, and high workload. Driving behavior characteristics and physiological characteristics were then assigned to the corresponding low workload, medium workload, and high workload categories to obtain the processed dataset.
[0030] Preferably, step 4-2 specifically involves: using a random forest model, an extreme gradient boosting tree model, a lightweight gradient boosting machine model, and a classification boosting tree model as base learners; performing K-fold cross-validation on the processed dataset; recording the predicted probability of each sample in the dataset at the folds not involved in training, output by each base learner; and obtaining preliminary prediction results.
[0031] Preferably, step 4-3 specifically involves: concatenating the preliminary prediction results of the four basic learners column by column to form... n The meta-feature matrix Z is 4×4, and each column of the meta-feature matrix Z corresponds to the prediction probability of a base learner on all samples; the meta-learner is a multilayer perceptron model (MLP), with the meta-feature matrix Z as the input of the multilayer perceptron model MLP and the workload category as the output of the multilayer perceptron model MLP.
[0032] Step 4-4 specifically involves the following: The multilayer perceptron model (MLP) consists of two fully connected layers, using the ReLU activation function and dropout regularization, trained using the cross-entropy loss function, and employing an early stopping strategy to control overfitting, thereby obtaining a prediction model for the workload category and the final prediction result.
[0033] Preferably, step 5 specifically involves: performing SHAP interpretation analysis on the trained prediction model and the final prediction result, taking the four columns of the meta-feature matrix Z as interpretation features, and quantifying the contribution of each interpretation feature to the driver's workload prediction through the SHAP framework.
[0034] The SHAP value of the preliminary prediction results of each basic learner to the final prediction result is calculated to obtain the marginal contribution of each basic learner, which is used to quantify the role of the basic learner in the final prediction result.
[0035] Preferably, in step 5, quantifying the contribution of each explanatory feature to driver workload prediction using the SHAP framework specifically involves:
[0036] make x i This indicates the first element in the processed dataset. i One sample, x ij Indicates the first in this sample j One explanatory feature; within the framework of the stacked model, x ij The SHAP value is represented as The baseline value of the stacked model is represented as y b Then the sample x i Predicted workload y i Represented as:
[0037] ;
[0038] ;
[0039] In the formula:
[0040] M To explain the number of features;
[0041] N The set of all explanatory features;
[0042] S For those that do not contain explanatory features j of N Any subset of;
[0043] v(S) For containing only subsets S The prediction results when interpreting features;
[0044] v(SU{j}) To explain the features j The prediction results after adding to the subset.
[0045] Compared with the prior art, the advantages of this application are:
[0046] (1) This application provides a driver workload prediction method based on multiple physiological signals under multiple auditory interferences. It uses a multimodal method to integrate physiological characteristics, driving behavior characteristics and subjective assessments, and explores the combined effect of multiple auditory interferences on driver workload, making the prediction results more accurate.
[0047] (2) This application adopts a stacked model, combining RF, XGBoost, LightGBM and CatBoost models, achieving an accuracy of over 90% in workload classification, outperforming individual models and significantly improving the accuracy, universality and robustness of workload prediction; the SHAP framework analysis identifies factors with consistent predictive ability under different workload categories and has a non-linear threshold, which is crucial for assessing collision risk.
[0048] (3) This application adopts a stacking integration method to integrate the complementary advantages of RF, XGBoost, LightGBM and CatBoost models, because RF, XGBoost, LightGBM and CatBoost are all tree-based models, maintaining structural consistency, facilitating the coherence of the output structure and the seamless application of SHAP in terms of interpretability.
[0049] (4) The integration of multiple variables in this application can effectively make up for the limitations of the univariate model, especially when dealing with complex traffic conditions or unstable features. These prediction results provide practical insights for the design of driver monitoring systems.
[0050] (5) The SHAP framework of this application reveals the impact of various auditory disturbances and traffic scenarios on driver workload. Based on these findings, intelligent driving assistance systems should incorporate real-time workload assessment to dynamically adjust auditory stimuli, which will help reduce cognitive burden in complex driving scenarios. Attached Figure Description
[0051] Figure 1 The flowchart of the driver workload prediction method based on multiple physiological signals under various auditory interferences is shown in this application.
[0052] Figure 2 The original and processed electrodermal signals of a driver in Example 1 of this application;
[0053] Figure 3 The original and processed electrocardiogram (ECG) signals of a driver in Example 1 of this application;
[0054] Figure 4 The original and processed EEG signals of a driver in Example 1 of this application;
[0055] Figure 5 This is a flowchart illustrating the creation of the stacking model in Embodiment 1 of this application;
[0056] Figures 6-8 The importance ranking of the top ten explanatory features under different workload categories in Example 2 of this application;
[0057] Figure 9 The SHAP values of driving behavior characteristics under different workload categories in Example 2 of this application;
[0058] Figure 10 The SHAP values of electrodermal and electrocardiographic signals under different workload categories in Example 2 of this application;
[0059] Figure 11 The SHAP values of EEG signals under different workload categories in Example 2 of this application;
[0060] Figure 12 The SHAP value is the explanatory feature that affects a specific workload category in Example 2 of this application. Detailed Implementation
[0061] The present application is described in detail below with reference to the accompanying drawings and specific embodiments, but the present application is not limited to these embodiments. The present application covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present application. To provide the public with a thorough understanding of the present application, specific details are described in detail in the following embodiments, but those skilled in the art will fully understand the present application even without these detailed descriptions.
[0062] Example 1
[0063] like Figure 1 As shown, this application discloses a method for predicting driver workload based on multiple physiological signals under various auditory interferences, including the following steps:
[0064] Step 1: Scene Setup; Combine various auditory disturbances with traffic scenes to create multiple driving scenarios;
[0065] Step 2: Data Acquisition; Collect driving behavior data from multiple drivers operating driving simulators in various driving scenarios, and simultaneously collect skin conductance signals, electrocardiogram signals, and electroencephalogram signals from multiple drivers, and statistically analyze the subjective workload assessments of multiple drivers for multiple driving scenarios.
[0066] Step 3: Data processing;
[0067] Step 3-1: Perform anomaly removal and filtering on the collected driving behavior data, and then process the driving behavior data by mean and standard deviation to obtain driving behavior features;
[0068] Step 3-2: The collected skin conductance signals, electrocardiogram signals, and electroencephalogram signals are processed to obtain physiological characteristics;
[0069] Step 3-3: Classify the subjective workload assessment to obtain workload categories, and assign driving behavior features and physiological features to the corresponding workload categories to obtain the processed dataset;
[0070] Step 4: Establish a stacking model;
[0071] Step 4-1: Basic learner selection: Select four tree models as basic learners: Random Forest, Extreme Gradient Boosting Tree, Lightweight Gradient Boosting Machine, and Classification Boosting Tree.
[0072] Step 4-2: Basic learner training: Optimize each tree model through grid search to determine the optimal hyperparameters; then input the processed dataset into each optimized tree model to generate preliminary prediction results;
[0073] Step 4-3: Meta-learner training: The preliminary prediction results are used to construct the meta-feature matrix Z, which serves as the input to the meta-learner, and the workload category serves as the output of the meta-learner.
[0074] Step 4-4: Train the meta-learner to obtain the prediction model for the workload category and the final prediction result;
[0075] Step 5: Model Result Interpretation; Perform SHAP interpretation analysis on the trained prediction model and the final prediction results. Use the meta-feature matrix Z as the interpretation feature, and quantify the contribution of each interpretation feature to the driver's workload prediction through the SHAP framework. This is used to predict the impact of driving behavior features and physiological features on workload under various auditory interferences.
[0076] Example 2
[0077] Preferably, in step 1, the various auditory interferences include navigation modes and music rhythms. The navigation modes include high-frequency navigation and low-frequency navigation, the music rhythms include no music, slow-paced music and fast-paced music, and the traffic scenarios include regular roads, school roads and construction site roads. By arranging and combining various auditory interferences with traffic scenarios, 18 driving scenarios are obtained.
[0078] Preferably, in step 2, the driving simulator is a fixed-base driving simulator, which includes a steering wheel, brake pedal, accelerator pedal, three displays, and a main controller. The main controller is equipped with UC-Win / Road software to simulate the driving environment and record driving behavior data at a sampling rate of 25 Hz. The skin conductance signal is collected through a skin resistance sensor, with the two electrodes of the skin resistance sensor attached to the driver's index and middle fingers respectively, and the skin conductance signal is collected at a frequency of 50 Hz. The electrocardiogram (ECG) signal is collected through an ECG sensor, which is attached to the driver's chest area and records the raw heart signal at a frequency of 200 Hz. The electroencephalogram (EEG) signal is collected through the EPOC X portable EEG headset at a sampling rate of 128 Hz. The X portable EEG headset includes 14 EEG sensors and reference sensors P3 and P4. EEG sensors AF3, AF4, F3, F4, F7, and F8 correspond to the frontal lobe, EEG sensors P7 and P8 correspond to the parietal lobe, EEG sensors O1 and O2 correspond to the occipital lobe, and EEG sensors T7, T8, FC5, and FC6 correspond to the temporal lobe. Each channel of the EPOC X portable EEG headset records 16-bit high-resolution data with a sensitivity of 0.1275µV, accurately capturing weak brainwave fluctuations.
[0079] Example 3
[0080] Preferably, step 3-1 specifically involves: identifying and removing outliers in the physiological time series using a standard deviation-based outlier detection method, filtering with a Savitzky-Golay filter, and then processing the driving behavior data by mean and standard deviation to obtain driving behavior features; the driving behavior data includes driving speed, lateral acceleration, longitudinal acceleration, lane departure, road centerline departure, accelerator pedal force, and brake pedal force; the driving behavior features include average speed and speed standard deviation, average lateral acceleration and lateral acceleration standard deviation, average longitudinal acceleration and longitudinal acceleration standard deviation, mean and standard deviation of lane departure, mean and standard deviation of road centerline departure distance, average accelerator pedal force and accelerator pedal force standard deviation, and average brake pedal force and brake pedal force standard deviation.
[0081] Preferably, step 3-2 includes the following steps:
[0082] Step 3-2-1: Baseline standardization is performed by subtracting the average resting state of each driver from the collected skin conductance, electrocardiogram, and electroencephalogram signals;
[0083] Step 3-2-2: Use the three sigma rule to identify and remove standardized outliers;
[0084] Step 3-2-3: Use a combination of high-pass and low-pass filters to filter signal artifacts, with the frequency range set according to the signal characteristics;
[0085] Step 3-2-4: Smooth the signal using Savitzky-Golay filtering;
[0086] Step 3-2-5: Apply spline interpolation to fill the small gaps caused by missing values to obtain physiological characteristics.
[0087] Preferably, step 3-3 specifically includes:
[0088] The NASA-TLX scale was used to collect subjective workload assessments from multiple drivers for multiple driving scenarios. The subjective workload assessment was based on a 6-point scale. Based on the subjective workload assessment, the multiple driving scenarios were classified into workload categories, namely low workload, medium workload, and high workload. Driving behavior characteristics and physiological characteristics were then assigned to the corresponding low workload, medium workload, and high workload categories to obtain the processed dataset.
[0089] Example 4
[0090] Preferably, step 4-2 specifically involves: using a random forest model, an extreme gradient boosting tree model, a lightweight gradient boosting machine model, and a classification boosting tree model as base learners; performing K-fold cross-validation on the processed dataset; recording the predicted probability of each sample in the dataset at the folds not involved in training, output by each base learner; and obtaining preliminary prediction results.
[0091] Preferably, step 4-3 specifically involves: concatenating the preliminary prediction results of the four basic learners column by column to form... n The meta-feature matrix Z is 4×4, and each column of the meta-feature matrix Z corresponds to the prediction probability of a base learner on all samples; the meta-learner is a multilayer perceptron model (MLP), with the meta-feature matrix Z as the input of the multilayer perceptron model MLP and the workload category as the output of the multilayer perceptron model MLP.
[0092] Step 4-4 specifically involves the following steps: The multilayer perceptron model (MLP) consists of two fully connected layers, using the ReLU activation function and dropout regularization, trained using the cross-entropy loss function, and employing an early stopping strategy to control overfitting, thereby obtaining a prediction model for the workload category and the final prediction result.
[0093] Preferably, step 5 specifically involves: performing SHAP interpretation analysis on the trained prediction model and the final prediction result, taking the four columns of the meta-feature matrix Z as interpretation features, and quantifying the contribution of each interpretation feature to the driver's workload prediction through the SHAP framework.
[0094] The SHAP value of the preliminary prediction results of each basic learner to the final prediction result is calculated to obtain the marginal contribution of each basic learner, which is used to quantify the role of the basic learner in the final prediction result.
[0095] This application uses a multilayer perceptron (MLP) model to nonlinearly combine the preliminary prediction results of four basic learners, which can model the interaction and non-additive effects among the basic learners. The SHAP method is used to interpret and analyze the meta-learners, which can output the contribution of each basic learner in the ensemble prediction (final prediction result), thus realizing the interpretability of the ensemble model. At the same time, the OOF prediction (preliminary prediction result) is used to construct the meta-feature matrix Z, which avoids information leakage and improves the generalization ability of the model.
[0096] The main hyperparameters of each model in the basic learner are automatically tuned based on cross-validation performance. The number of hidden layer nodes in the meta-learner MLP can be selected within the range of 16 to 128 depending on the task size, and the dropout rate is set between 0.2 and 0.5. An early stopping strategy on the validation set is employed during training to prevent overfitting. In the model interpretation module, TreeExplainer is used to perform SHAP analysis on the trained meta-learner MLP. Each column in the meta-feature matrix Z (corresponding to the preliminary prediction results of each basic learner) is used as the explanatory feature. The feature contribution value of each sample is calculated, and the global contribution ranking is obtained by averaging the sample-level SHAP values. Explanatory feature importance analysis, nonlinear analysis, and threshold analysis can be added to the interpretation to deeply analyze the behavioral differences of drivers under different workload levels.
[0097] Preferably, in step 5, quantifying the contribution of each explanatory feature to driver workload prediction using the SHAP framework specifically involves:
[0098] make x i This indicates the first element in the processed dataset. i One sample, x ij Indicates the first in this sample j One explanatory feature; within the framework of the stacked model, x ij The SHAP value is represented as The baseline value of the stacked model is represented as y b Then the sample x i Predicted workload y i Represented as:
[0099] ;
[0100] ;
[0101] In the formula:
[0102] M To explain the number of features;
[0103] N The set of all explanatory features;
[0104] S For those that do not contain explanatory features j of N Any subset of;
[0105] v(S) For containing only subsets S The prediction results when interpreting features;
[0106] v(SU{j}) To explain the features j The prediction results after adding to the subset.
[0107] Application Example 1
[0108] (1) Scene setting:
[0109] The experimental route was designed based on urban roads in a certain city. The area consists of three transverse roads and three longitudinal roads, each with six lanes in any direction. The lane width is 3.75 meters, the central median is 1.5 meters wide, forming nine intersections. To ensure consistency between scenarios, the connecting roads between intersections are 1.3 kilometers long, the traffic flow is 800 vehicles per hour, and the maximum speed limit is 60 kilometers per hour.
[0110] To ensure drivers experience different workload conditions, a 2x3x3 hybrid experimental design was employed. For navigation, since drivers typically remain active in navigation throughout the driving process, and the frequency of navigation prompts can affect workload in different ways, two driving routes with the same direction and distance were designed. One route had high-frequency prompts, and the other had low-frequency prompts. The navigation frequency design was based on typical modes of in-vehicle navigation systems: standard mode and simplified mode. Both modes provide the same key information, but at different frequencies. For example, the former provides route-by-route guidance 200 meters, 100 meters, and 50 meters before intersections, while the latter only provides guidance at 100 meters. Pre-experiments were conducted to ensure that both navigation modes could smoothly guide drivers to their destinations. All navigation prompts used the default voice of the Amap navigation system, which uses a standard female Mandarin voice with a neutral tone. The command language remained consistent under all conditions to avoid linguistic complexity or intonation variations.
[0111] Regarding the music, this application designed three musical conditions: no music, slow-tempo music, and fast-tempo music. The slow-tempo music had a tempo of approximately 60 beats per minute (bpm), while the fast-tempo music had a tempo of approximately 120 beats per minute. In the formal experiment, all drivers listened to the same musical pieces under each tempo condition. The slow-tempo music consisted of instrumental pieces, while the fast-tempo music contained English lyrics unfamiliar to most drivers. The structure, length, and volume of the pieces were matched to ensure consistency, with tempo being the primary variable.
[0112] To enhance navigation prompts and create task scenarios of varying difficulty, three typical traffic environments were designed: regular roads, school zones, and construction zones.
[0113] (2) Data collection:
[0114] A total of 74 drivers (46 men and 28 women) aged between 22 and 40 participated in the experiment. Their driving experience ranged from 4 to 10 years. All drivers held a valid driver's license for at least one year and had experience driving in real-world cities. Their vision was normal or corrected, and they had no chronic or acute respiratory or cardiovascular diseases, and were well-rested the night before the experiment. The experiment used a fixed-base driving simulator equipped with a Logitech G923 steering wheel, brake and accelerator pedals, and three 27-inch displays with a resolution of 2560×1440 pixels. The simulation software used was UC-Win / Road.
[0115] Driving behavior data acquisition: The simulator records driver operation data in real time at a frequency of 25 Hz, including speed, acceleration / deceleration, lateral acceleration, longitudinal acceleration, lane departure, and the degree to which the accelerator and brake pedals are pressed.
[0116] Skin conductance signal (EDA) and electrocardiogram (ECG) acquisition: These data were collected using EDA and ECG sensors. The EDA sensor has two electrodes fixed to the index and middle fingers of the driver's left hand, recording skin conductance level (SCL) at a sampling rate of 50Hz. The ECG sensor is attached to the chest area, recording raw cardiac signals at a frequency of 200Hz. The raw signals from both sensors were exported using custom software.
[0117] EEG signal acquisition: EEG data was collected using the EPOC X portable EEG headset at a sampling rate of 128 Hz, following the international 10-20 system. Power spectral density features were extracted every 16 data points (i.e., one feature was generated every 8 Hz). The device includes 14 EEG sensors and 2 reference sensors (P3 and P4). Specifically, AF3, AF4, F3, F4, F7, and F8 correspond to the frontal lobe; P7 and P8 to the parietal lobe; O1 and O2 to the occipital lobe; and T7, T8, FC5, and FC6 to the temporal lobe. Each channel of the device records 16 bits of high-resolution data with a sensitivity of 0.1275µV, enabling precise capture of subtle EEG fluctuations. The EEG signals were exported using customized software.
[0118] Driver workload assessment: This application employs a simplified version of the NASA-TLX, where drivers assessed their perceived workload using a 1-6 verbal rating scale after each driving scenario. This approach was chosen because the experiment involved 18 consecutive scenarios, making it impractical to repeatedly implement the full NASA-TLX with pairwise weighting. The simplified procedure efficiently reports workload while still capturing meaningful variations under different conditions. Standardized briefings were provided prior to the formal experiment to ensure drivers clearly understood the meaning of each scale. Detailed explanations and concrete examples were used to illustrate the state differences in the 1-6 ratings across different workload categories, helping drivers distinguish between different levels of demand. Brief familiarization sessions were also conducted to improve rating consistency and minimize subjective bias. During the experiment, drivers verbally reported their perceived workload immediately after each scenario. Each driver completed two routes, each containing nine scenarios, generating a total of 18 workload reports. These verbal ratings were conducted at intersections connecting adjacent scenarios, specifically referring to the workload experienced in the previous scenario. The intersections themselves were excluded in subsequent analysis to maintain the integrity of the workload-scenario mapping.
[0119] (3) Data processing:
[0120] The driving behavior data processing involves anomaly removal and filtering of the collected driving behavior data, and calculation of the mean and standard deviation of these data to obtain a total of 14 driving behavior features.
[0121] Processing of EDA, ECG, and EEG signals: Physiological signals (EDA, ECG, and EEG) are processed in five steps. First, baseline normalization is performed by subtracting the resting mean of each driver from the dynamic measurements to reduce individual variability. Second, outliers are identified and removed using the three-sigma rule. Third, common signal artifacts, such as muscle activity, electrical noise, and respiratory fluctuations, are filtered using a combination of high-pass and low-pass filters, with the frequency range set according to the signal characteristics. Fourth, signal smoothing is performed using Kalman filtering, exponential moving average, median filtering, and Savitzky-Golay filtering. Among these, Savitzky-Golay filtering yielded the best results. Finally, spline interpolation is applied to fill small gaps caused by missing values.
[0122] Electrodermal conductance (EDA) signals are represented by skin conductance level (SCL), while electrocardiogram (ECG) signals are converted into RR intervals to reflect heart rate variability (HRV). Electroencephalogram (EEG) signals are processed according to brain regions and wave types and converted into power values for five frequency bands (Theta, Alpha, low Beta, high Beta, and Gamma) in four major cortical regions. Figures 2-4 The figure presents raw and processed physiological data of a representative driver. Blue represents raw physiological data, and red represents processed physiological data. During the experiment, physiological signals and driving behavior data were recorded synchronously with the activation of the driving simulator. Any data that could not be aligned in time was discarded. Due to the different sampling rates of different modalities, all signals were resampled to match the lowest sampling frequency, thus corresponding to driving behavior characteristics.
[0123] (4) Establish a stacking model:
[0124] A comparative evaluation of the current dataset revealed that four tree-based models (RF, XGBoost, LightGBM, and CatBoost) offered superior prediction performance, stronger generalization ability, and better interpretability and framework compatibility. Therefore, these four models were selected as the base classifiers for further ensemble. A stacked ensemble approach was employed to improve prediction accuracy and leverage the strengths of each model.
[0125] Given that the experimental design involves various auditory interference scenarios, ensuring the robustness of model predictions under different conditions is crucial. To this end, a stacked ensemble approach is employed to integrate the complementary strengths of the selected models. This framework also benefits from structural consistency, as the four base learners are all tree-based models, facilitating the coherence of the output structure and the seamless application of SHAP for interpretability. Furthermore, each model possesses unique methodological advantages: RF enhances generalization by introducing bootstrap aggregation and random feature selection; XGBoost combines regularized gradient-based boosting to capture complex decision boundaries; LightGBM improves computational efficiency through histogram-based learning and leaf growth; and CatBoost processes categorical feature encoding through ordered boosting techniques while reducing overfitting. Integrating these models within a stacked architecture aims to enhance overall prediction robustness and support reliable interpretation of workload-related patterns under different experimental conditions.
[0126] like Figure 5 The flowchart shown is a flowchart of the stacked model construction process. The main steps of stacked model construction are: (1) Basic learner selection: Select these four tree-based models as basic learners; (2) Basic learner training: Optimize each model through grid search to determine the best hyperparameters; The processed dataset is input into each optimized model to generate preliminary prediction results; (3) Meta-learner training: The preliminary prediction results are used to construct the meta-feature matrix Z, which is used as the input of the meta-learner, and the workload category is used as the output of the meta-learner; (4) Obtain the prediction model of the workload category and the final prediction result through meta-learner training. Typical choices of meta-learners include linear regression, logistic regression, or multilayer perceptron (MLP) neural networks. In this application, these three choices were evaluated, and the MLP-based meta-learner achieved the best performance and was therefore adopted.
[0127] Model evaluation metrics:
[0128] The prediction of low, medium, and high workload labels is framed as a classification task. The main evaluation metrics for the classification model include:
[0129] Accuracy: Reflects the overall proportion of correct predictions made by the model.
[0130] Precision: Represents the proportion of actual samples that are correctly identified as the predicted category; a higher precision can reduce the number of false positives.
[0131] Recall: Measures the ability of a model to correctly identify samples of a specific class. A high recall ensures that as many true positive examples as possible are detected.
[0132] F1 score: The harmonic mean of precision and recall, providing a more balanced assessment and mitigating the bias that may arise from relying on a single metric.
[0133] To ensure the reliability of the evaluation and enhance the robustness of the model, 5-fold cross-validation was employed. This method reduces the risk of overfitting and improves the broad applicability of the results through repeated dataset splitting and validation.
[0134] (5) Interpretation of model results:
[0135] SHAP is an additivity interpretation framework based on cooperative game theory that can be applied to any machine learning model. In this application, SHAP values are used to quantify the contribution of each physiological and behavioral indicator to driver workload prediction.
[0136] make x i Represents the first in the dataset i One sample, x ij Indicates the first in this sample j Each explanatory feature (i.e., physiological signals or driving behavior indicators). Within the stacked model framework, x ij The SHAP value is represented as The model baseline value is represented as y b Therefore, the sample x i Predicted workload y i It can be represented as:
[0137] ;
[0138] The core idea of SHAP is to compute each feature x ij The marginal contribution is calculated across all possible feature subsets, and their incremental impact on the prediction is weighted and averaged. This provides... x ij The attribution in the final prediction provides a fair allocation. Formally, the SHAP value... Represented as:
[0139] ;
[0140] In the formula:
[0141] M To explain the number of features;
[0142] N The set of all explanatory features;
[0143] S For those that do not contain explanatory features j of N Any subset of;
[0144] v(S) For containing only subsets S The prediction results when interpreting features;
[0145] v(SU{j}) To explain the features j The prediction results after adding to the subset.
[0146] Application Example 2
[0147] The following data was obtained by conducting experiments using the steps designed in Application Example 1:
[0148] Table 1 presents descriptive statistics of driving behavior data and physiological indicators across all driving scenarios. Regarding driving behavior data, speed (42.2 ± 18.4 km / h), acceleration (lateral 0.01 ± 0.62 m / s², longitudinal -0.02 ± 0.79 m / s²), accelerator pedal force (13.65% ± 16.35%), and brake pedal force (2.49% ± 9.14%) exhibited significant variability, indicating that drivers adjusted their control strategies according to different task requirements. These fluctuations may reflect dynamic changes in perceived motion and workload under different scenarios.
[0149] Regarding physiological parameters, the mean skin conductance level (SCL) was 1.82 ± 2.84 µS, showing high inter-individual variability, which may reflect fluctuations in arousal and task-related workload. The RR interval was 0.79 ± 0.59 seconds, corresponding to approximately 75.6 bpm, showing significant heart rate variability. This pattern may be partly attributed to the limited number of scenarios without auditory stimulation (e.g., silence or baseline conditions), suggesting that drivers are typically in a state of high arousal or cognitive engagement in most scenarios.
[0150] For EEG signals, band power varied across different brain regions and frequency bands. Among all EEG signals, the frontal lobe Theta power (10.09 ± 7.10 µV²) was the highest, especially compared to Theta activity in the occipital lobe (2.59 ± 2.39 µV²) and temporal lobe (2.54 ± 5.85 µV²), supporting its association with executive processing and cognitive workload. In contrast, the frontal lobe Alpha power (2.43 ± 3.49 µV²) was relatively low, consistent with Alpha activity in the parietal and temporal lobes, indicating sustained attentional involvement. Although the absolute amplitude of the temporal lobe Gamma (0.36 ± 0.30 µV²) was low, it was significantly higher than the Gamma power in other regions such as the occipital lobe (0.12 ± 0.15 µV²) and parietal lobe (0.14 ± 0.22 µV²), and has been associated with emotional stimulation and cognitive integration under high task demands.
[0151] To improve the interpretability and effectiveness of subsequent modeling, two preparatory analyses were performed. Initially, the Mann-Whitney U test was used to perform paired comparisons of four representative variables (spdKPH_mean, SG_Scl, SG_RR, and T_Gamma) in driving scenarios, revealing significant differences in distribution (Tables 1A–4A, Appendix). Subsequently, the variance inflation factor (VIF) was used to assess multicollinearity of all candidate indices. The results showed that six variables (P_Alpha, T_Alpha, O_BetaH, P_BetaL, O_Alpha, and T_BetaH) had VIF values exceeding 10 and were excluded, retaining the remaining 30 explanatory features for model training and interpretation.
[0152] Table 1. Descriptive statistics explaining the features
[0153]
[0154]
[0155] Table 2 shows the drivers' subjective ratings for each scenario, which form the basis for classifying workload categories. Specifically, ratings of 0–2 are categorized as low workload, 2–4 as medium workload, and 4–6 as high workload. This scheme represents a balanced and symmetrical division of the NASA-TLX scale. Furthermore, this classification shows strong consistency with the expected difficulty level of driving scenarios and the actual distribution of driver responses. According to this definition, each workload category corresponds to six scenarios: low workload includes driving scenarios 1–4, 10, and 11; medium workload includes driving scenarios 5–8, 12, and 13; and high workload includes driving scenarios 9, 14, 15, 16, 17, and 18.
[0156] Table 2 Descriptive statistics of subjective workload assessment
[0157]
[0158] As shown in Table 3, this application compares the classification performance of Random Forest (RF), Extreme Gradient Boosting Tree (XGBoost), Lightweight Gradient Boosting Machine (LightGBM), CatBoost, Multilayer Perceptron (MLP), and Stacking models under low, medium, and high workload categories. The Stacking model achieved the highest accuracy for each workload category, at 0.9328, 0.9184, and 0.9232, respectively, surpassing all single models. Furthermore, the Stacking model maintained a good balance between precision and recall, indicating higher classification stability. CatBoost and XGBoost performed competitively, but exhibited an imbalance between precision and recall, increasing the risk of misclassification. LightGBM also performed quite well, but its recall was slightly lower, while RF lagged behind other models in both accuracy and generalization ability. Overall, the Stacking model, by combining the advantages of the base learner, proved to be the most effective model, providing superior prediction performance and robustness.
[0159] Table 3. Prediction and evaluation metrics for various models
[0160]
[0161] The impact of different explanatory features on workload:
[0162] To further interpret the output of the stacked model, the SHAP interpreter was applied. For example... Figures 6-8 As shown, the SHAP summary plots for low, medium, and high workloads are presented, along with the top 10 explanatory features that have the greatest impact on model predictions under different workload categories, and their corresponding SHAP value distributions. In the plot, blue dots represent lower feature values, while red dots represent higher values.
[0163] Figures 6-8The results indicate that while some explanatory features rank differently across workload categories, several key metrics, including SG_Scl, ofstRoad_mean, SG_RR, P_Gamma, and T_Gamma, consistently remain highly important. These results underscore the importance of integrating physiological and behavioral features for accurate workload prediction. Furthermore, the differences in the ranking of explanatory features across workload categories suggest that different types of features contribute differently under varying task demands. For example, SG_Scl and P_Gamma rank highest under low workload conditions, indicating that emotional arousal and higher cognitive activity play a more significant role when driving demands are lower. Conversely, ofstRoad_mean and SG_RR are more important under medium and high workload conditions, suggesting that biases in physiological burden and driving behavior become more pronounced as task difficulty increases. Particularly in high workload scenarios, reduced attention and weakened vehicle control may further impact workload assessment.
[0164] Furthermore, some explanatory features exhibited relatively stable rankings but displayed different SHAP value distributions across different workload categories. For example, under medium workloads, lower ofstRoad_mean values were generally associated with positive SHAP values, suggesting that minimal road deviation corresponds to medium workloads. However, under high workloads, higher ofstRoad_mean values were associated with positive SHAP values, reflecting greater vehicle deviation from lane center during cognitively demanding driving conditions. Similarly, SG_Scl showed differentiated SHAP patterns across different workload categories, exhibiting greater variability under high workloads, potentially reflecting increased emotional responsiveness and physiological stress. These findings highlight the need to consider not only the overall importance of features but also the variations in feature behavior under different workload conditions, which can significantly impact model interpretation and predictive accuracy.
[0165] Explain the impact of characteristics on different workload categories:
[0166] To further explore how specific explanatory features affect different levels of driver workload, scatter plots were generated to show the relationship between selected explanatory features and their corresponding SHAP values. To help the system identify thresholds associated with different workload categories, a locally weighted scatter smoothing (LOWESS) curve was applied to the SHAP value of each explanatory feature, revealing its non-linear marginal contribution. Subsequently, threshold points were determined based on the intersection of the smoothed SHAP curve with the zero line, indicating a reversal of the contribution direction. These data-driven inflection points were used to delineate value ranges that might correspond to specific workload categories, such as... Figures 9-12As shown in the figures, a positive SHAP value indicates that the explanatory feature contributes positively to the prediction of a specific workload category, while a negative value indicates an inhibitory effect. In all the charts, the relationship between features and SHAP values exhibits a clear non-linear pattern, highlighting the complex and context-sensitive impact of individual explanatory features on workload prediction.
[0167] Figures 9-11 The chart shows SHAP value patterns for the same explanatory features across different workload categories. Among driving behavior metrics, ofstRoad_mean and spdKPH_mean significantly influence workload prediction. It's important to note that to maintain consistency in scenario design, some driving directions are reversed between scenarios. Therefore, a negative ofstRoad_mean (i.e., to the left of the lane centerline) may correspond to the opposite driving direction indicated by a positive value (i.e., to the right of the centerline). For low workload prediction, a positive SHAP value when ofstRoad_mean is less than -7.8 meters or between 6.1 and 8.8 meters indicates that the driver is more likely to be in a low workload state when in the outermost left lane or the second lane closest to the right of the centerline. This may reflect that these lanes are associated with a more relaxed state of attention due to driving habits.
[0168] For medium workload prediction, when the ofstRoad_mean value is between -8.2 meters and 4.8 meters, the SHAP value fluctuates significantly but is generally positive, indicating an increased likelihood of a driver being under medium workload when approaching the first or second lane. This highlights the complexity of the relationship between lateral position and workload, particularly around the second lane, where the workload category appears to shift. For high workload prediction, when ofstRoad_mean exceeds 5.1 meters, the SHAP value is positive, indicating a significantly increased probability of high workload when a driver occupies the second or third lane to the right of the centerline. Notably, the second lane exhibits overlapping areas between different workload categories, potentially representing a transitional region requiring additional features to differentiate workload states.
[0169] In contrast, spdKPH_mean has a more significant impact on workload prediction. The probability of high workload increases sharply when driving speeds are below 17.6 km / h or above 67.8 km / h. Conversely, speeds in the intermediate range are associated with low and medium workload categories.
[0170] Regarding physiological indicators, both EDA and ECG played important roles in predicting different workload categories. For EDA, SG_Scl showed differentiated effects under different workload conditions. When SG_Scl ranged from 0.8 µS to 2.4 µS, low workload was more likely; when it ranged from 2.3 µS to 6.0 µS, medium workload was indicated; and when SG_Scl exceeded 3.6 µS, it was associated with high workload. This pattern suggests that stronger skin conductance responses are often associated with higher workloads, reflecting the sensitivity of SCL to physiological stress. It also indicates a certain degree of overlap in SG_Scl values between workload categories.
[0171] For ECG, SG_RR has a more significant impact on predicting high workloads, while its impact is relatively limited under low and medium workloads. In both cases, the SHAP value of SG_RR remains relatively flat, indicating that HRV contributes less to the prediction, possibly due to the stability of the physiological rhythm. Conversely, during high workloads, the SHAP value increases rapidly as SG_RR rises, indicating that longer RR intervals (i.e., slower heart rates) significantly increase the probability of high workloads. This trend may reflect stress-related physiological responses under high-pressure conditions, where drivers are simultaneously exposed to fast-paced music, sparse navigation cues, and complex traffic environments. The ensuing increase in parasympathetic activity and decrease in heart rate may indicate increased anxiety or psychological stress.
[0172] Regarding EEG metrics, P_Gamma, T_Gamma, and P_Theta play significant roles in predicting different workload categories. For low workloads, the SHAP value is positive when P_Gamma is greater than 0.16 µV² and T_Gamma is greater than 0.26 µV², suggesting that moderately enhanced Gamma power may represent mild cognitive activation, helping drivers maintain focus without leading to high workload. Furthermore, the SHAP value is also positive when P_Theta is less than 0.55 µV², indicating that lower Theta power may correspond to reduced allocation of cognitive resources, thus decreasing the need for sustained alertness.
[0173] For medium workloads, the SHAP value is positive when P_Gamma ranges from 0.04 µV² to 0.12 µV². These values fluctuate significantly, indicating that Gamma power within this range makes a significant contribution to medium workload prediction, possibly related to fluctuations in driver information processing needs and attention regulation. Simultaneously, the SHAP value is also positive when T_Gamma is less than 0.26 µV², indicating that temporal lobe Gamma power makes a significant contribution to medium workload prediction within this range. When P_Theta is between 0.56 µV² and 2.03 µV², the SHAP value remains positive but lower, suggesting that enhanced Theta power may slightly increase the probability of medium workloads, possibly related to the need for continuous attention monitoring.
[0174] For high workloads, the SHAP value is positive when P_Gamma is between 0.13 µV² and 0.24 µV², indicating that increased parietal Gamma power contributes to high workload prediction, reflecting that drivers allocate more cognitive resources under high workloads. The SHAP value is also positive when T_Gamma is greater than 1.22 µV², suggesting that higher temporal Gamma power may be associated with high-intensity information integration needs, such as auditory processing or coping with complex environments. Furthermore, the SHAP value is positive when P_Theta is greater than 1.18 µV², indicating that higher Theta power may increase the likelihood of high workload prediction, suggesting the need for sustained cognitive control and attention maintenance under high workload conditions.
[0175] Overall, moderately enhanced gamma waves under low workloads may correspond to a relaxed yet focused cognitive state, while significantly enhanced gamma waves under high workloads indicate higher cognitive demands. Theta waves play a more subtle role in medium workloads, while enhancement under high workloads may be associated with increased cognitive control. These findings suggest that different EEG bands play different roles across different workload categories.
[0176] Figure 12 Explanatory features that demonstrate predictive effectiveness only under specific workload conditions are highlighted. For example, when ofsLane_mean is less than zero, the probability of low workload increases, indicating that the vehicle is veering to the left of the lane. Medium workload is more likely when F_Gamma ranges from 0.25 µV² to 0.47 µV². Similarly, high workload increases when accZ_mean is greater than or less than zero (corresponding to sudden acceleration or deceleration). These explanatory features under specific conditions can serve as complementary features to enhance workload predictions at different levels.
[0177] The preferred embodiments of this application have been described in detail above. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.
[0178] Many other changes and modifications can be made without departing from the concept and scope of this application. It should be understood that this application is not limited to the specific embodiments, and the scope of this application is defined by the appended claims.
Claims
1. A method for predicting driver workload based on multiple physiological signals under various auditory interferences, characterized in that, Includes the following steps: Step 1: Scene Setup; Various auditory disturbances are combined with traffic scenarios to form a variety of driving scenarios; Step 2: Data Acquisition; Collect driving behavior data from multiple drivers operating driving simulators in various driving scenarios, and simultaneously collect skin conductance signals, electrocardiogram signals, and electroencephalogram signals from multiple drivers, and statistically analyze the subjective workload assessments of multiple drivers for multiple driving scenarios. Step 3: Data processing; Step 3-1: Perform anomaly removal and filtering on the collected driving behavior data, and then process the driving behavior data by mean and standard deviation to obtain driving behavior features; Step 3-2: The collected skin conductance signals, electrocardiogram signals, and electroencephalogram signals are processed to obtain physiological characteristics; Step 3-3: Classify the subjective workload assessment to obtain workload categories, and assign driving behavior features and physiological features to the corresponding workload categories to obtain the processed dataset; Step 4: Establish a stacking model; Step 4-1: Basic learner selection: Select four tree models as basic learners: Random Forest, Extreme Gradient Boosting Tree, Lightweight Gradient Boosting Machine, and Classification Boosting Tree. Step 4-2: Basic learner training: Optimize each tree model through grid search to determine the optimal hyperparameters; then input the processed dataset into each optimized tree model to generate preliminary prediction results; Step 4-3: Meta-learner training: The preliminary prediction results are used to construct the meta-feature matrix Z, which serves as the input to the meta-learner, and the workload category serves as the output of the meta-learner. Step 4-4: Train the meta-learner to obtain the prediction model for the workload category and the final prediction result; Step 5: Model Result Interpretation; Perform SHAP interpretation analysis on the trained prediction model and the final prediction results. Use the meta-feature matrix Z as the interpretation feature, and quantify the contribution of each interpretation feature to the driver's workload prediction through the SHAP framework. This is used to predict the impact of driving behavior features and physiological features on workload under various auditory interferences.
2. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 1, characterized in that, In step 1, the various auditory interferences include navigation modes and music rhythms. The navigation modes include high-frequency navigation and low-frequency navigation, the music rhythms include no music, slow-paced music and fast-paced music, and the traffic scenarios include regular roads, school roads and construction site roads. Combining these various auditory interferences with the traffic scenarios yields 18 driving scenarios.
3. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 1, characterized in that, In step 2, the driving simulator is a fixed-base driving simulator. The fixed-base driving simulator includes a steering wheel, brake pedal, accelerator pedal, three displays and a main controller. The main controller is equipped with UC-Win / Road software to simulate the driving environment of the driving scenario and record driving behavior data at a sampling rate of 25 Hz. The electrodermal (ED) signal is acquired using a skin resistance sensor. The two electrodes of the skin resistance sensor are attached to the driver's index and middle fingers, respectively, and the EED signal is acquired at a frequency of 50 Hz. The electrocardiogram (ECG) signal is acquired using an ECG sensor. The ECG sensor is attached to the driver's chest area, and the raw heart signal is recorded at a frequency of 200 Hz. The electroencephalogram (EEG) signal is acquired using the EPOC X portable EEG headset at a sampling rate of 128 Hz. The EPOC X portable EEG headset includes 14 EEG sensors and reference sensors P3 and P4. EEG sensors AF3, AF4, F3, F4, F7, and F8 correspond to the frontal lobe; EEG sensors P7 and P8 correspond to the parietal lobe; EEG sensors O1 and O2 correspond to the occipital lobe; and EEG sensors T7, T8, FC5, and FC6 correspond to the temporal lobe. Each channel of the EPOC X portable EEG headset records 16-bit high-resolution data with a sensitivity of 0.1275 µV, accurately capturing weak brainwave fluctuations.
4. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 1, characterized in that, Step 3-1 specifically involves: using a standard deviation-based outlier detection method to identify and remove outliers in the physiological time series, filtering with a Savitzky-Golay filter, and then processing the driving behavior data by mean and standard deviation to obtain driving behavior features. The driving behavior data includes driving speed, lateral acceleration, longitudinal acceleration, lane departure, road centerline deviation, accelerator pedal force, and brake pedal force; driving behavior characteristics include average speed and speed standard deviation, average lateral acceleration and lateral acceleration standard deviation, average longitudinal acceleration and longitudinal acceleration standard deviation, mean and standard deviation of lane departure, mean and standard deviation of road centerline deviation distance, average accelerator pedal force and accelerator pedal force standard deviation, and average brake pedal force and brake pedal force standard deviation.
5. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 1, characterized in that, Step 3-2 includes the following steps: Step 3-2-1: Baseline standardization is performed by subtracting the average resting state of each driver from the collected skin conductance, electrocardiogram, and electroencephalogram signals; Step 3-2-2: Use the three sigma rule to identify and remove standardized outliers; Step 3-2-3: Use a combination of high-pass and low-pass filters to filter signal artifacts, with the frequency range set according to the signal characteristics; Step 3-2-4: Smooth the signal using Savitzky-Golay filtering; Step 3-2-5: Apply spline interpolation to fill the small gaps caused by missing values to obtain physiological characteristics.
6. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 1, characterized in that, Step 3-3 specifically involves: The NASA-TLX scale was used to collect subjective workload assessments from multiple drivers for multiple driving scenarios. The subjective workload assessment was based on a 6-point scale. Based on the subjective workload assessment, the multiple driving scenarios were classified into workload categories, namely low workload, medium workload, and high workload. Driving behavior characteristics and physiological characteristics were then assigned to the corresponding low workload, medium workload, and high workload categories to obtain the processed dataset.
7. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 1, characterized in that, Step 4-2 specifically involves using a random forest model, an extreme gradient boosting tree model, a lightweight gradient boosting machine model, and a classification boosting tree model as base learners. K-fold cross-validation is performed on the processed dataset, and the predicted probability of each sample in the dataset is recorded by the output of each base learner on the folds that are not involved in training, thus obtaining preliminary prediction results.
8. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 7, characterized in that, Step 4-3 specifically involves concatenating the preliminary prediction results of the four basic learners column-wise to form... n The meta-feature matrix Z is 4×4, and each column of the meta-feature matrix Z corresponds to the prediction probability of a base learner on all samples; the meta-learner is a multilayer perceptron model (MLP), with the meta-feature matrix Z as the input of the multilayer perceptron model MLP and the workload category as the output of the multilayer perceptron model MLP. Step 4-4 specifically involves the following steps: The multilayer perceptron model (MLP) consists of two fully connected layers, using the ReLU activation function and dropout regularization, trained using the cross-entropy loss function, and employing an early stopping strategy to control overfitting, thereby obtaining a prediction model for the workload category and the final prediction result.
9. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 8, characterized in that, Step 5 specifically involves: performing SHAP interpretation analysis on the trained prediction model and the final prediction results, taking the four columns of the meta-feature matrix Z as interpretation features, and quantifying the contribution of each interpretation feature to the driver's workload prediction through the SHAP framework. The SHAP value of the preliminary prediction results of each basic learner to the final prediction result is calculated to obtain the marginal contribution of each basic learner, which is used to quantify the role of the basic learner in the final prediction result.
10. The driver workload prediction method based on multiple physiological signals under various auditory interferences according to claim 9, characterized in that, In step 5, the contribution of each explanatory feature to driver workload prediction is quantified using the SHAP framework as follows: make x i This indicates the first element in the processed dataset. i One sample, x ij Indicates the first in this sample j One explanatory feature; in Within the framework of the stacked model, x ij The SHAP value is represented as The baseline value of the stacked model is represented as y b Then the sample x i Predicted workload y i Represented as: ; ; In the formula: M To explain the number of features; N The set of all explanatory features; S For those that do not contain explanatory features j of N Any subset of; v(S) For containing only subsets S The prediction results when interpreting features; v(SU{j}) To explain the features j The prediction results after adding to the subset.
Citation Information
Patent Citations
Intelligent driving fatigue evaluation system and method
CN119498853A
Control device, system and method for determining the perceptual load of a visual and dynamic driving scene
US20190272450A1