A real-time evaluation and adaptive early warning method for cognitive load based on multi-modal physiological signals

By integrating multimodal physiological signals and employing an adaptive early warning mechanism, the limitations of single indicators and real-time issues in existing cognitive load assessments are addressed. This enables real-time, accurate assessment and early warning of cognitive load, making it suitable for critical scenarios such as autonomous driving and air traffic control.

CN122096732APending Publication Date: 2026-05-29JIANGSU HOPERUN SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU HOPERUN SOFTWARE CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing cognitive load assessment methods suffer from limitations such as single-indicator assessment, lack of multimodal fusion mechanisms, insufficient real-time performance, poor individual adaptability, coarse cognitive load grading, and imperfect early warning mechanisms, making it difficult to meet the needs for real-time, accurate, and individualized assessment.

Method used

By integrating heart rate variability, skin conductance response, surface electromyography (EMG) signals, and electroencephalography (EEG) signals, a multidimensional cognitive load quantification model is constructed. An adaptive baseline calibration mechanism and a multimodal feature fusion network are adopted, combined with a cognitive load dynamic prediction model, to achieve real-time assessment and early warning.

Benefits of technology

It enables real-time and accurate assessment of cognitive load, improves the stability and individual adaptability of the assessment, can complete the assessment within 5 seconds and provide early warning, and reduces the risk of errors caused by cognitive overload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122096732A_ABST
    Figure CN122096732A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal physiological signal's cognitive load real-time evaluation and adaptive early warning method, specifically including the following steps: S1, multi-modal physiological signal acquisition and pretreatment;S2, adaptive baseline calibration mechanism;S3, multi-modal feature fusion network;S4, cognitive load quantitative evaluation;S5, cognitive load dynamic prediction;S6, adaptive early warning mechanism;S7, model training strategy;Multi-stage training strategy is used to optimize model performance.The application realizes the real-time monitoring, dynamic evaluation and early warning of the degree of individual cognitive resource occupation by fusing heart rate variability, skin conductance response, surface electromyography signal and electroencephalogram four physiological modalities, constructs multi-dimensional cognitive load quantitative model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cognitive science, biosignal processing, and human-computer interaction, specifically to a real-time cognitive load assessment and adaptive early warning method based on multimodal physiological signals. It is applicable to various scenarios that require real-time monitoring of cognitive status, such as intelligent driving, air traffic control, medical surgery, online education, and human-computer collaboration, providing technical support for preventing cognitive overload, optimizing task allocation, and improving the safety of human-computer interaction. Background Technology

[0002] Cognitive load refers to the degree of mental resource consumption borne by an individual's working memory system when performing cognitive tasks. Accurate assessment of cognitive load is crucial for optimizing human-computer interaction, preventing operational errors, and improving work efficiency. However, traditional cognitive load assessment methods have many limitations and struggle to meet the application requirements of real-time performance, accuracy, and universality.

[0003] The main problems with existing technologies:

[0004] 1. Limitations of Single-Indicator Assessment: Existing studies often rely on a single physiological signal (such as using only EEG or only heart rate) for cognitive load assessment. Single indicators are easily affected by individual differences, environmental interference, and changes in physiological state, resulting in unstable assessment results and limited accuracy. For example, when assessing based solely on heart rate variability, factors such as exercise and emotional fluctuations can significantly interfere with the results, leading to misjudgments.

[0005] 2. Lack of multimodal fusion mechanisms: Although existing studies have attempted to combine multiple physiological signals, they mostly employ simple linear weighting or feature splicing methods, failing to fully explore the complementary and synergistic relationships between different physiological indicators. Different physiological signals reflect different aspects of cognitive load: heart rate variability mainly reflects autonomic nervous system activity, skin conductance reflects emotional arousal, electromyography (EMG) signals reflect muscle tension, and electroencephalography (EEG) signals directly reflect neural activity. Existing methods have failed to establish deep correlation models between these signals.

[0006] 3. Insufficient real-time performance: Traditional methods require a relatively long signal acquisition window (usually more than 30 seconds) to obtain stable evaluation results, which cannot meet the needs of real-time application scenarios requiring second-level response. In critical scenarios such as autonomous driving and air traffic control, rapid changes in cognitive load require the system to complete the evaluation and issue warnings within seconds, and the time delay of existing methods is too long.

[0007] 4. Poor individual adaptability: Significant differences exist in the physiological baselines of different individuals, and the same cognitive load level may manifest as completely different physiological signal patterns in different individuals. Existing methods mostly employ population average models and lack adaptive calibration mechanisms for individuals, leading to systematic biases in the assessment results.

[0008] 5. Coarse Cognitive Load Grading: Existing methods typically only distinguish between 3-5 coarse-grained levels such as "low load," "medium load," and "high load," failing to provide finer-grained quantification of cognitive resource usage. In practical applications, it is necessary to differentiate between different states such as "mild overload," "moderate overload," and "severe overload" in order to implement differentiated intervention strategies.

[0009] 6. Inadequate early warning mechanism: Existing systems mostly use fixed thresholds for early warning, failing to dynamically adjust early warning strategies based on individual characteristics, task type, and historical trends. They lack the ability to predict trends in cognitive load changes, making it impossible to provide early warnings before cognitive overload occurs.

[0010] 7. Difficulty in assessing multi-task scenarios: In complex multi-task scenarios, individuals need to process multiple tasks simultaneously, and different tasks may generate cognitive loads of varying intensities. Existing methods struggle to distinguish and quantify the cognitive load contribution from different task sources, thus failing to provide a basis for adjusting task priorities.

[0011] Therefore, there is an urgent need for a novel cognitive load assessment system that can integrate multimodal physiological signals, achieve real-time and accurate assessment, and possess individual adaptive capabilities and a dynamic early warning mechanism. This invention addresses the aforementioned technical challenges by constructing a multimodal physiological signal fusion network, designing an adaptive baseline calibration algorithm, and establishing a dynamic prediction model for cognitive load, thus providing an efficient and reliable technical solution for cognitive load assessment. Summary of the Invention

[0012] To address the aforementioned issues, this invention provides a real-time cognitive load assessment and adaptive early warning system based on multimodal physiological signals. By integrating four physiological modalities—heart rate variability, skin conductance response, surface electromyography (EMG) signals, and electroencephalography (EEG) signals—a multidimensional cognitive load quantification model is constructed, enabling real-time monitoring, dynamic assessment, and early warning of an individual's cognitive resource occupancy.

[0013] The specific plan is as follows:

[0014] A real-time cognitive load assessment and adaptive early warning method based on multimodal physiological signals, specifically including the following steps:

[0015] S1. Multimodal physiological signal acquisition and preprocessing: Four physiological signals are acquired in real time: heart rate variability (HRV), skin conductance response (SCR), surface electromyography (sEMG), and electroencephalography (EEG). These four signals reflect an individual's cognitive state from different perspectives: HRV reflects autonomic nervous system activity, skin conductance reflects emotional arousal, SCR reflects muscle tension, and EEG directly reflects neural activity. By fusing these complementary physiological information, cognitive load can be assessed more comprehensively and accurately. After preprocessing (including denoising, filtering, and normalization), temporal feature sequences are extracted for each signal, providing a foundation for subsequent multimodal fusion analysis.

[0016] S2. Adaptive Baseline Calibration Mechanism: After extracting multimodal physiological signal features, a key issue to address is the difference in physiological baselines between individuals. The same cognitive load level may manifest as completely different physiological signal patterns in different individuals, and directly using the original features can lead to systematic bias in the assessment results. To eliminate the impact of individual differences on the assessment results, this invention designs an adaptive baseline calibration mechanism to establish an individual physiological baseline model. This mechanism collects physiological signals of individuals in a resting state before real-time assessment, calculates baseline statistical features, and then standardizes the real-time features relative to the baseline, thereby eliminating individual differences and improving the accuracy and universality of the assessment.

[0017] S3. Multimodal Feature Fusion Network: Although the modal features after baseline calibration eliminate individual differences, they still exist in different feature spaces (different dimensions, different physical meanings). To fully utilize the complementarity and synergistic relationship of different physiological signals, this invention constructs a multimodal feature fusion network that maps the features of four physiological modalities to a unified cognitive load representation space. This fusion network first maps the features of each modality to a unified dimensional space through a projection matrix, then uses a cross-modal attention mechanism to learn the correlation between different modalities, and finally obtains a comprehensive multimodal cognitive load feature through weighted fusion, providing input for subsequent cognitive load quantitative assessment.

[0018] S4. Cognitive Load Quantitative Assessment: After multimodal feature fusion, a unified multimodal cognitive load feature vector was obtained. Based on this fusion feature, a cognitive load quantification model is constructed, which maps high-dimensional features to one-dimensional cognitive load scores. The model adopts a multilayer perceptron structure, learns the complex mapping relationship between features and cognitive load through nonlinear transformation, and outputs a continuous cognitive load score (range of 0-1), which is the corresponding cognitive load assessment result, thus realizing fine-grained quantification of cognitive resource occupancy.

[0019] S5. Dynamic prediction of cognitive load; obtaining the cognitive load assessment results at the current moment. Subsequently, the system not only needs to understand the current state but also needs to predict future trends in order to provide early warnings before cognitive overload occurs. To achieve early warnings before cognitive overload occurs, this invention constructs a time-series prediction model that predicts future trends based on historical cognitive load sequences. This model extracts the statistical characteristics (mean, standard deviation, and trend) of historical cognitive load sequences and uses linear regression to establish the mapping relationship between these characteristics and future cognitive load, thereby predicting the cognitive load level over a future period and obtaining future load prediction results, providing a basis for early warning decisions.

[0020] S6. Adaptive early warning mechanism; upon obtaining the current cognitive load assessment results... and future load forecast results Afterwards, the system needs to determine whether to trigger an early warning based on this information. An adaptive early warning mechanism is designed based on the current cognitive load assessment and future load forecast. This mechanism first dynamically calculates the early warning threshold based on individual historical data and current status, then checks whether the current load, predicted load, and changing trends exceed the threshold, and triggers different levels of early warning (Level 1 alert, Level 2 warning, Level 3 severe) according to the degree of exceedance, providing decision support for timely intervention.

[0021] S7. Model training strategy: Employ a multi-stage training strategy to optimize model performance.

[0022] Further, step S1 includes:

[0023] S11, Heart Rate Variability Feature Extraction

[0024] Heart rate variability reflects the activity state of the autonomic nervous system and is an important indicator for assessing cognitive load. Assume that the acquired electrocardiogram signal, after R-wave detection, yields the RR interval sequence. ,in Indicates the first RR intervals (unit: milliseconds). The sequence length;

[0025] The sliding window method is used to calculate the time-domain and frequency-domain characteristics; for the time window ,in For the current moment, For the window length (typically 30 seconds), calculate the following characteristics:

[0026] Time-domain features include the standard deviation of the RR interval (SDNN) and the root mean square (RMSSD) of the difference between adjacent RR intervals.

[0027]

[0028]

[0029] in This represents the number of RR intervals within the window. The mean RR interval, and Representing time respectively SDNN and RMSSD values;

[0030] The power spectral density is calculated using the Fast Fourier Transform (FFT) to represent the frequency domain characteristics. Specifically, the RR interval sequence is first interpolated and resampled to convert the non-uniformly sampled RR intervals into a uniform time series (typically resampled to 4Hz). Then, the resampled sequence is subjected to an FFT to obtain its frequency domain representation, and the power spectral density function is calculated. The power spectral density function reflects the energy distribution of heart rate variability across different frequency components; low-frequency power (LF), 0.04–0.15 Hz, and high-frequency power (HF), 0.15–0.4 Hz, are extracted.

[0031]

[0032] in The power spectral density function is obtained by FFT calculation. and Representing time respectively The low-frequency and high-frequency power is calculated by numerical integration methods (such as the trapezoidal method or Simpson's method) to calculate the power integral within the frequency band;

[0033] The heart rate variability eigenvector is defined as:

[0034]

[0035] in For a moment The heart rate variability feature vector contains 5 feature dimensions;

[0036] S12, Extraction of Skin Conductivity Response Features

[0037] Skin conductance reflects the degree of activation of the sympathetic nervous system and is closely related to cognitive load and emotional arousal. Let the collected skin conductance signal be... ,in Indicates the first Conductivity values ​​at each sampling point (unit: microSiemens). This represents the number of sampling points; the sampling frequency is typically 10Hz.

[0038] For time windows Extract the following features:

[0039] Baseline conductivity level (SCL) and conductivity response amplitude (SCR):

[0040]

[0041]

[0042] in This represents the number of sampling points within the window. Indicates time The basic conductivity level, Indicates the magnitude of the electrical conductivity response;

[0043] The conductance response frequency (SCR frequency) is calculated by detecting the number of conductance rising edges.

[0044]

[0045] in The difference between adjacent sampling points. The rise threshold is typically 0.05 microSiemens per second. For counting functions, For window length, Indicates the number of electrical reactions per unit time;

[0046] The skin conductance eigenvector is defined as:

[0047]

[0048] in For a moment The skin conductance feature vector;

[0049] S13, Extraction of surface electromyography signal features

[0050] Surface electromyography (EMG) signals reflect muscle tension and can be used to assess physical tension induced by cognitive tasks. Let the acquired EMG signals be... ,in Indicates the first Electromyography amplitude at each sampling point (unit: microvolts). This represents the number of sampling points; the sampling frequency is typically 1000Hz.

[0051] For time windows The root mean square (RMS) and average power frequency (MPF) are calculated. The RMS value is directly calculated from the time-domain signal and reflects the amplitude intensity of the electromyographic signal.

[0052]

[0053] in This represents the number of sampling points within the window. For the first Electromyography amplitude at each sampling point Indicates time The root mean square value;

[0054] The calculation of the average power frequency (MPF) requires first obtaining the power spectral density; specifically, this is achieved by performing an FFT transform on the electromyographic signal and calculating the power spectral density function. The FFT transform converts a time-domain signal into a frequency-domain representation, including the power spectral density. Indicates the signal at different frequencies The power distribution on the surface; then calculate the average power frequency:

[0055]

[0056] in The power spectral density function of the electromyographic signal is calculated by FFT transformation. This is the maximum analysis frequency (typically 500Hz). Indicates time The average power frequency is obtained by calculating the frequency domain weighted average using numerical integration methods (such as the trapezoidal method), which reflects the main frequency components of the electromyographic signal.

[0057] The surface electromyography feature vector is defined as:

[0058]

[0059] in For a moment Surface electromyographic feature vectors;

[0060] S14, EEG signal feature extraction

[0061] Electroencephalogram (EEG) signals directly reflect brain neural activity and are a core indicator for cognitive load assessment. Let the acquired multi-channel EEG signals be... ,in Indicates the first Multichannel EEG data at various time points This refers to the number of channels (usually 32 or 64). The sampling frequency is typically 250Hz or 500Hz, representing the number of time points.

[0062] For time windows The power features of different frequency bands are extracted using wavelet transform. Specifically, the discrete wavelet transform (DWT) or continuous wavelet transform (CWT) is used, and appropriate wavelet basis functions (such as Daubechies wavelet, Morlet wavelet, etc.) are selected to perform multi-scale decomposition of the EEG signal for each channel. By setting different wavelet scale parameters, the EEG signal is decomposed into five frequency bands: δ, θ, α, β, and γ, corresponding to 0.5-4Hz, 4-8Hz, 8-13Hz, 13-30Hz, and 30-100Hz, respectively. For each frequency band, wavelet coefficients at the corresponding scale are extracted. ,in Indicates the frequency band type. Indicates the channel index. Indicates the time point index; calculates the average power for each frequency band:

[0063]

[0064]

[0065] in and Channels At any moment The wavelet coefficients of the δ and θ frequency bands (obtained through wavelet transform). For the number of channels, The number of time points within the window. and Representing time respectively The average power in the δ and θ bands is obtained by averaging the squares of the wavelet coefficients over all channels and time points; similarly, it is calculated using the same wavelet transform method. , and .

[0066] The frequency band power ratio is used as a sensitive indicator of cognitive load.

[0067]

[0068] in and These represent the α / θ power ratio and the β / α power ratio, respectively, and these ratios are closely related to cognitive load levels.

[0069] The EEG feature vector is defined as:

[0070]

[0071] in For a moment The EEG feature vector.

[0072] Further, step S2 includes:

[0073] S21, Baseline Feature Extraction

[0074] Physiological signals were collected for 5-10 minutes as baseline data while the individual was at rest (without cognitive task); for each physiological modality, the mean and covariance matrix of the baseline eigenvectors were calculated.

[0075]

[0076]

[0077] in The number of time points for the baseline data. This is the baseline mean vector of heart rate variability. For the baseline covariance matrix, calculate the same. , , , , and ;

[0078] S22, Feature Standardization

[0079] During real-time evaluation, the current eigenvector is standardized relative to the baseline to eliminate individual differences. Specifically, this is achieved by first calculating the difference between the eigenvector and the baseline mean, and then performing a whitening transformation using the square root inverse of the covariance matrix. The calculation method is as follows: for the covariance matrix Perform eigenvalue decomposition ,in The eigenvector matrix, If the matrix is ​​a diagonal eigenvalue matrix, then ,in Let be the diagonal matrix of the inverses of the square roots of the eigenvalues. The standardized formula is:

[0080]

[0081] in This is the standardized heart rate variability feature vector. The square root inverse of the covariance matrix is ​​obtained through eigenvalue decomposition. This is the original feature vector at the current moment. This is the baseline mean vector; this standardization method transforms the feature vector into a zero-mean, unit-covariance distribution, effectively eliminating individual differences. Similarly, the same method is used to calculate... , and .

[0082] Further, step S3 includes:

[0083] S31, Feature Projection

[0084] Project the standardized features of each modality onto a unified... Dimensional cognitive load representation space:

[0085]

[0086]

[0087]

[0088]

[0089] in , , , The projected weight matrix is ​​a learnable matrix. , , , For bias vectors, To standardize the representation of spatial dimensions (usually set to 64 or 128). , , , These are the projected feature vectors;

[0090] S32, Cross-modal attention fusion

[0091] A multi-head cross-attention mechanism is employed to calculate the correlation between different modalities, achieving deep fusion; EEG modalities are used as queries, and other modalities are used as keys.

[0092]

[0093]

[0094] in For the learnable attention parameter matrix, This represents a vector concatenation operation;

[0095] Calculate attention weights and fusion features:

[0096]

[0097]

[0098] in Here is the attention weight matrix. The feature vector after attention fusion;

[0099] S33, Weighted Fusion

[0100] The final multimodal cognitive load features were obtained by weighted summation:

[0101]

[0102] in For learnable fusion weight coefficients, satisfying , This is the final multimodal cognitive load feature vector.

[0103] Further, step S4 includes:

[0104] S41, Cognitive Load Regression

[0105] Multilayer perceptrons are used to map multimodal features to cognitive load scores:

[0106]

[0107]

[0108] in This is the first layer weight matrix. This is the first layer bias vector. This is the dimension of the hidden layer (usually 128 or 256). To modify the activation function of the linear unit, For hidden layer feature vectors, This is the weight vector for the second layer. For the second-level bias scalar, For a moment The cognitive load score is 0, which indicates no cognitive load and 1 indicates cognitive overload.

[0109] S42, Cognitive Load Classification

[0110] Based on cognitive load scores, five levels are defined:

[0111] Extremely low load, They have sufficient cognitive resources and can undertake additional tasks.

[0112] Low load, Cognitive resources are relatively abundant;

[0113] Medium load, : Use cognitive resources appropriately;

[0114] High load, Cognitive resources are nearing saturation and require attention.

[0115] Extremely high load, Cognitive overload requires immediate intervention.

[0116] Further, step S5 includes:

[0117] S5. Temporal Feature Extraction

[0118] Extract the past Cognitive load sequence at each time step ,in The length of the history window (usually 10-20 time steps);

[0119] Calculate the time series statistical characteristics:

[0120]

[0121]

[0122]

[0123] in For historical average cognitive load, For historical standard deviation, To show the trend of cognitive load changes;

[0124] S52, Future Load Forecast

[0125] Predicting the future using a linear regression model The cognitive load is calculated after a certain time step. The specific implementation method is to train a linear regression model on historical data using the least squares method or gradient descent method. The training data includes historical cognitive load sequences and their corresponding time-series characteristics (mean, standard deviation, trend, current value), with the target variable being the future... The actual cognitive load after each time step; the regression coefficients are solved by minimizing the sum of squared prediction errors; the prediction formula is:

[0126]

[0127] in These are the regression coefficients obtained by training on historical data using the least squares method or gradient descent method. This represents the predicted future cognitive load score. The linear regression model establishes a linear mapping relationship between time-series statistical characteristics and future cognitive load, enabling prediction of future trends based on current and historical conditions.

[0128] Further, step S6 includes:

[0129] S61. Calculation of Early Warning Threshold

[0130] The warning threshold is dynamically adjusted based on individual historical data and current status:

[0131]

[0132] in The basic warning threshold is (usually 0.7). The baseline average cognitive load for individuals, and For adjustment factors (usually set to 0.3 and 0.2), For a moment Dynamic early warning threshold;

[0133] S62. Warning Triggering Conditions

[0134] An alert is triggered when any of the following conditions are met:

[0135] 1. Current cognitive load exceeds the warning threshold: ;

[0136] 2. Predicted future load exceeds threshold: ;

[0137] 3. Cognitive load increases too rapidly: ;

[0138] in The trend threshold is typically 0.05 per time step; warning levels are divided into three levels based on cognitive load: Level 1 warning (alert): Or it is predicted that it will reach this range; Level 2 warning: Or it is predicted that it will reach this range; Level 3 warning (severe): Or it is predicted that it will reach this range.

[0139] Further, step S7 includes:

[0140] Phase 1: Single-modal pre-training

[0141] Feature extractors and projection networks are pre-trained on each modality-specific dataset to enable the model to learn the mapping relationship between each modality and cognitive load. The specific training method involves updating model parameters using backpropagation and gradient descent optimizers (such as Adam and SGD). For each modality, single-modality features are input into the projection and regression networks to obtain the predicted cognitive load score, which is then compared with the true label to calculate the loss. The mean squared error loss function is used.

[0142]

[0143] in The number of training samples, Labels for actual cognitive load. The cognitive load score is predicted for a single modality. By minimizing this loss function, the projection matrix and regression network parameters are iteratively updated using gradient descent, enabling the model to learn the mapping relationship of cognitive load from single-modal features.

[0144] Phase Two: Multimodal Fusion Training

[0145] Train the fusion network and cognitive load regression model on paired multimodal datasets. Specifically, the training method is as follows: fix the parameters of the single-modal feature extractor trained in stage one, and optimize only the parameters of the multimodal fusion network (including projection matrix, attention mechanism, and weighted fusion weights) and the cognitive load regression network; use the backpropagation algorithm to calculate the gradient of the total loss with respect to each parameter, and use a gradient descent optimizer (such as Adam) to update all trainable parameters simultaneously; the total loss function is:

[0146]

[0147] in The mean squared error loss for cognitive load regression measures the difference between the multimodal fusion prediction results and the true labels. To mitigate the consistency loss of prediction results across different modalities, we encourage that prediction results from different modalities remain consistent. To compensate for the mean square error loss in future load forecasting, optimize the accuracy of the forecasting model; These are the loss weighting coefficients (typically set to 1.0, 0.5, and 0.3), used to balance the contributions of different loss terms. By jointly optimizing these three loss terms, the model simultaneously learns accurate cognitive load assessment and prediction capabilities.

[0148] Through the above technical solutions, this invention achieves deep fusion and real-time processing of multimodal physiological signals, and constructs an accurate, fast, and adaptive cognitive load assessment and early warning system, providing strong technical support for application scenarios that require real-time monitoring of cognitive status.

[0149] The beneficial effects of this invention are as follows: by integrating multiple physiological indicators such as heart rate variability (HRV), skin conductance response (SCR), surface electromyography (sEMG), and electroencephalography (EEG), a multi-dimensional cognitive load quantification model is constructed, enabling real-time monitoring, dynamic assessment, and early warning of individual cognitive resource occupancy. Specific advantages are as follows:

[0150] 1. By deeply fusing information from four physiological modalities, this invention reduces the root mean square error of cognitive load assessment by 30-40% and increases the correlation coefficient to over 0.85 compared to single-indicator assessment methods. Multimodal fusion can fully utilize the complementarity of different physiological signals, overcome the limitations of single indicators, and significantly improve the accuracy and stability of the assessment.

[0151] 2. Through optimized algorithms and parallel processing technology, the system can complete the entire process of multimodal signal acquisition, feature extraction, fusion evaluation, and early warning output within 5 seconds, meeting the real-time requirements of key scenarios such as intelligent driving and air traffic control. Compared to traditional methods that require an evaluation window of more than 30 seconds, the response speed is improved by more than 6 times.

[0152] 3. Through an adaptive baseline calibration mechanism, the system can adapt to the differences in physiological characteristics among individuals, eliminating the influence of individual baselines on the assessment results. Experiments show that after baseline calibration, the consistency of assessments among different individuals is improved by 25%, and the system's universality is significantly enhanced.

[0153] 4. Through a dynamic cognitive load prediction model, the system can predict cognitive load trends over the next 25-50 seconds, providing early warnings before cognitive overload occurs and offering sufficient time for timely intervention. Compared to warning methods based solely on the current state, this increases the warning lead time by 5-10 seconds, effectively reducing the risk of errors caused by cognitive overload.

[0154] 5. Employing continuous value output (range 0-1) and a five-level hierarchical system, it can provide more refined quantification of cognitive resource usage, supporting more accurate state judgment and decision-making. Compared with the traditional 3-5 level discrete classification, the information content is increased by more than 50%, providing richer state information for application systems.

[0155] 6. The multimodal fusion mechanism enables the system to maintain evaluation accuracy even when the quality of a certain modal signal is poor or missing, relying on information from other modalities. Experiments show that the system performance degrades by less than 15% when a single modal signal is missing, demonstrating good robustness.

[0156] 7. This invention can be applied to multiple fields such as intelligent driving, air traffic control, medical surgery, online education, human-computer collaboration, and virtual reality interaction, and has broad market application prospects and commercial value. Especially in critical scenarios requiring real-time monitoring of cognitive status and prevention of operational errors, this invention possesses irreplaceable technical advantages.

[0157] 8. The technical framework of this invention has good scalability, can easily integrate more physiological signals (such as respiration, body temperature, etc.), can be extended to more application scenarios, and can be adapted to hardware platforms of different sizes (from embedded devices to cloud servers), providing a solid foundation for subsequent technology iteration and functional enhancement.

[0158] 9. Through real-time monitoring and early warning, the system can promptly detect cognitive overload and remind users to take mitigation measures, effectively reducing the risk of misoperation, accidents, and losses caused by cognitive overload. In intelligent driving scenarios, it is expected to reduce the accident rate caused by driver cognitive overload by more than 30%. Attached Figure Description

[0159] Figure 1 This is a flowchart of the method of the present invention.

[0160] Figure 2 This is the network structure diagram in this invention. Detailed Implementation

[0161] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0162] As shown in the figure, this embodiment provides a real-time cognitive load assessment and adaptive early warning method based on multimodal physiological signals, taking driver cognitive load monitoring in intelligent driving scenarios as the application background. The specific solution is as follows:

[0163] I. System Architecture and Hardware Configuration

[0164] Construct a complete cognitive load monitoring system, including a data acquisition module, a signal preprocessing module, a feature extraction module, a fusion evaluation module, and an early warning output module.

[0165] Hardware Configuration: ECG Acquisition: Wearable ECG sensor (250Hz sampling rate) is used to acquire ECG signals via chest strap or wristband; Skin Conductivity Acquisition: Two-finger electrode skin conductivity sensor (10Hz sampling rate) is worn on the middle and ring fingers of the non-driving hand; Electromyography (EMG) Acquisition: Surface EMG sensor (1000Hz sampling rate) is attached to the driver's forearm muscles (e.g., flexor carpi radialis); Electroencephalography (EEG) Acquisition: Portable multi-channel EEG cap (32 channels, 250Hz sampling rate) is used, employing dry electrode technology to reduce discomfort; Data Processing: Embedded processing unit (e.g., NVIDIA Jetson Xavier) is used for real-time signal processing and model inference.

[0166] The system performs real-time assessments with a 5-second time window, updating the cognitive load assessment results and warning status every 5 seconds.

[0167] II. Implementation of Multimodal Physiological Signal Acquisition and Preprocessing

[0168] The system performs real-time evaluation in a 5-second time window (i.e.) Compared to the 30-second window mentioned in the invention description, the 5-second window provides a faster response time, meeting real-time requirements. In practical applications, the window length can be adjusted according to specific scenario needs. The following details the specific implementation process of feature extraction for each modality.

[0169] (1) Implementation of Heart Rate Variability Feature Extraction

[0170] The acquired electrocardiogram (ECG) signal was subjected to R-wave detection to obtain the RR interval sequence. Assuming the interval is within a 5-second window (…), (seconds) detected RR intervals: milliseconds, of which Indicates the first RR intervals (unit: milliseconds).

[0171] The time-domain characteristics are calculated according to the formula in the invention. First, the average RR interval is calculated:

[0172]

[0173] in The mean RR interval, This represents the number of RR intervals within the window.

[0174] Then calculate the standard deviation of the RR interval (SDNN):

[0175]

[0176] in Indicates time The SDNN value.

[0177] Calculate the root mean square (RMSSD) of the difference between adjacent RR intervals:

[0178]

[0179] in Indicates time The RMSSD value.

[0180] The power spectral density is calculated by performing an FFT transform on the RR interval sequence according to the method described in the invention. The specific implementation steps are as follows: First, the non-uniformly sampled RR interval sequence is interpolated and resampled (using spline interpolation or linear interpolation methods) to convert the RR intervals into a uniform time series (the resampling frequency is typically 4Hz). Then, the FFT algorithm (using the Cooley-Tukey Fast Fourier Transform algorithm) is applied to the resampled sequence to obtain its frequency domain representation. Calculate the power spectral density function According to the formulas in the invention, low-frequency power (LF, 0.04-0.15Hz) and high-frequency power (HF, 0.15-0.4Hz) are calculated using numerical integration methods (such as the trapezoidal rule or Simpson's rule).

[0181]

[0182]

[0183] in and Representing time respectively Low-frequency and high-frequency power, The power spectral density function obtained by FFT is used to calculate the power integral within the frequency band using numerical integration methods (such as the trapezoidal method or Simpson's method).

[0184] Calculate the LF / HF power ratio:

[0185]

[0186] According to the definition in the invention, the heart rate variability feature vector is:

[0187]

[0188] in For a moment The heart rate variability feature vector contains 5 feature dimensions.

[0189] (2) Implementation of Skin Conductivity Response Feature Extraction

[0190] Within a 5-second window ( (seconds), the skin conductance signal sampling frequency is 10Hz, therefore the number of sampling points is Assume the collected conductivity value sequence is as follows: ,in Indicates the first Conductivity values ​​at each sampling point (unit: microSiemens). The range is between 2.5 and 4.2 micro Siemens.

[0191] Calculate the baseline conductivity level (SCL) according to the formula in the invention:

[0192]

[0193] in Indicates time The basic conductivity level, This represents the number of sampling points within the window.

[0194] Calculate the conductivity response amplitude (SCR amplitude):

[0195]

[0196] in This indicates the magnitude of the electrical conductivity response.

[0197] Detect the rising edge of the conductance and calculate the conductance response frequency. Assume that three distinct rising edges are detected ( ),but:

[0198]

[0199] in The difference between adjacent sampling points. The rise threshold is microSiemens per second. For counting functions, This indicates the number of electrical reactions per unit time. The window length is in seconds.

[0200] According to the definition in the invention, the skin conductance feature vector is:

[0201]

[0202] in For a moment The skin conductance feature vector.

[0203] (3) Implementation of surface electromyography signal feature extraction

[0204] Within a 5-second window ( (seconds), the electromyography signal sampling frequency is 1000Hz, therefore the number of sampling points is Assume the acquired electromyography amplitude sequence is as follows: ,in Indicates the first The electromyography amplitude (unit: microvolt) at each sampling point ranges from -200 to +200 microvolts.

[0205] Calculate the root mean square (RMS) value according to the formula in the invention description:

[0206]

[0207] in Indicates time The root mean square value, This represents the number of sampling points within the window.

[0208] The power spectral density is calculated by performing an FFT transform on the electromyography (EMG) signal. Specifically, the FFT algorithm (using the Cooley-Tukey Fast Fourier Transform algorithm) is applied to the time-domain EMG signal sequence to obtain its frequency-domain representation. Then calculate the power spectral density. According to the formula in the invention, the average power frequency (MPF) is calculated using a numerical integration method (such as the trapezoidal rule).

[0209]

[0210] in The power spectral density function of the electromyographic signal is calculated by FFT transformation. Hz is the maximum analysis frequency. Indicates time The average power frequency is obtained by calculating the frequency domain weighted average using the numerical integration method.

[0211] According to the definition in the invention, the surface electromyography feature vector is:

[0212]

[0213] in For a moment Surface electromyography feature vectors.

[0214] (4) Implementation of EEG signal feature extraction

[0215] Within a 5-second window ( (seconds), the sampling frequency of the 32-channel EEG signal is 250Hz, therefore the number of time points is Let the acquired multi-channel EEG signals be... ,in Indicates the first 32-channel EEG data at various time points, This represents the number of channels.

[0216] Wavelet transform is performed on the EEG signal of each channel, decomposing it into five frequency bands according to the method described in the invention. The specific implementation steps are as follows: Discrete Wavelet Transform (DWT) or Continuous Wavelet Transform (CWT) is used, selecting Daubechies wavelet (db4) or Morlet wavelet as the wavelet basis function. For each channel's EEG signal, multi-scale wavelet decomposition is performed, setting different wavelet scale parameters to correspond to different frequency bands: the δ band (0.5-4Hz) corresponds to a larger scale, while the θ band (4-8Hz), α band (8-13Hz), β band (13-30Hz), and γ band (30-100Hz) correspond to progressively decreasing scales. Wavelet coefficients corresponding to the scale are extracted for each frequency band. The average power of each frequency band is calculated according to the formula in the invention description:

[0217]

[0218]

[0219] in and Channels At any moment The wavelet coefficients in the δ and θ frequency bands, For the number of channels, This represents the number of time points within the window. Similarly, the calculation yields... , and .

[0220] Calculate the frequency band power ratio according to the formula in the invention:

[0221]

[0222]

[0223] in and These represent the α / θ power ratio and the β / α power ratio, respectively.

[0224] According to the definition in the invention, the EEG feature vector is:

[0225]

[0226] in For a moment The EEG feature vector.

[0227] III. Adaptive Baseline Calibration Implementation

[0228] After extracting features from each modality, an individual baseline model needs to be established to eliminate individual differences. Before the driver begins driving, physiological signals are collected for 10 minutes at rest to establish the baseline model. This baseline data is used to calculate the mean and covariance matrix of each modality feature, providing a reference for subsequent real-time feature standardization.

[0229] Baseline feature calculation:

[0230] Assuming that the system samples in 5-second time windows during a 10-minute baseline period, a total of [number] samples are obtained. A time window. Based on the formula in the invention, for heart rate variability, the baseline mean and covariance are calculated:

[0231]

[0232]

[0233] Similarly, the baseline mean values ​​of other modalities (SCR, sEMG, EEG) are calculated according to the formulas in the invention. , , Covariance Matrix , , Establish a complete baseline model.

[0234] During real-time evaluation, the current eigenvector is standardized relative to the baseline according to the formulas and methods described in the invention. The specific steps are as follows: first, calculate the difference between the eigenvector and the baseline mean; then, calculate the square root inverse of the covariance matrix. The calculation method is as follows: for the covariance matrix Perform eigenvalue decomposition (Using the QR algorithm or the Jacobi method), where The eigenvector matrix, If the matrix is ​​a diagonal eigenvalue matrix, then ,in This is a diagonal matrix of the inverses of the square roots of the eigenvalues ​​(each eigenvalue is the square root and then its reciprocal). The standardized formula is:

[0235]

[0236] in This is the standardized heart rate variability feature vector. Let be the square root inverse of the covariance matrix obtained through eigenvalue decomposition. Assume the standardized eigenvectors are calculated as follows:

[0237]

[0238] Similarly, normalized features for other modalities are calculated. , and These standardized features eliminate individual differences and provide a unified basis for subsequent multimodal fusion.

[0239] IV. Implementation of Multimodal Feature Fusion

[0240] After baseline calibration, although the standardized features of each modality eliminated individual differences, they still resided in different feature spaces (dimensions of 5, 3, 2, and 7, respectively). To fully utilize the complementarity of different physiological signals, these features need to be mapped to a unified cognitive load representation space and deeply fused.

[0241] (1) Implementation of feature projection

[0242] According to the formula in the invention, the standardized features of each modality are projected onto a unified 64-dimensional cognitive load representation space. ):

[0243]

[0244] in The projected weight matrix is ​​a learnable matrix. For bias vectors, To unify the representation of spatial dimensions, This is the projected heart rate variability feature vector. Assume the projection yields... (The specific value is determined by the parameters of the trained model).

[0245] Similarly, the projected features of other modalities are calculated according to the formulas in the invention:

[0246]

[0247]

[0248]

[0249] in , , The projected weight matrix is ​​a learnable matrix. , , For bias vectors, , , is the projected feature vector.

[0250] (2) Implementation of cross-modal attention fusion

[0251] After feature projection, the features of each modality are mapped to a unified 64-dimensional space. To fully utilize the complementarity and synergy between different modalities, a cross-modal attention mechanism is employed for deep fusion, based on the formula described in the invention. Using the EEG modality as the query and the other three modalities as the keys, cross-attention is calculated as follows:

[0252]

[0253]

[0254] in For the learnable attention parameter matrix, This represents a vector concatenation operation. For querying the matrix, and It is a key-value matrix.

[0255] Calculate the attention weight according to the formula in the invention:

[0256]

[0257] in This is the attention weight matrix (because the dimension of the concatenated key-value matrices is 192). To standardize the representation of spatial dimensions, we assume the attention weights are: HRV modality weight 0.35, SCR modality weight 0.42, and sEMG modality weight 0.23, indicating that the skin conductance response contributes the most to cognitive load at the current moment.

[0258] Calculate the fusion features according to the formula in the invention:

[0259]

[0260] in This is the feature vector after attention fusion.

[0261] (3) Weighted integration implementation

[0262] After completing cross-modal attention fusion, attention fusion features were obtained. To synthesize the original features and attention fusion features of each modality, the final multimodal cognitive load features are obtained through weighted summation according to the formula in the invention. The fusion weights are assumed to be: (HRV) (SCR) (sEMG) (EEG) (Fusion features), satisfying :

[0263]

[0264] in For learnable fusion weight coefficients, This results in the final multimodal cognitive load feature vector, which integrates information from four physiological modalities, providing input for subsequent quantitative assessment of cognitive load.

[0265] V. Implementation of Cognitive Load Quantitative Assessment

[0266] After completing the multimodal feature fusion, a unified multimodal cognitive load feature vector was obtained. Based on this fusion feature, it needs to be mapped to a one-dimensional cognitive load score to achieve a quantitative assessment of cognitive resource occupancy.

[0267] (1) Implementation of cognitive load regression

[0268] According to the formula in the invention, the fused multimodal features are input into a multilayer perceptron to calculate the cognitive load score. Assume the hidden layer dimension... :

[0269]

[0270] in This is the first layer weight matrix. This is the first layer bias vector. For the hidden layer dimension, To unify the representation of spatial dimensions, To modify the activation function of the linear unit, These are the hidden layer feature vectors.

[0271] Based on the formula in the invention, output the cognitive load score:

[0272]

[0273] in This is the weight vector for the second layer. For the second-level bias scalar, For a moment The cognitive load score is 0, which indicates no cognitive load and 1 indicates cognitive overload.

[0274] Assuming the calculation yields According to the classification criteria in the invention description, it belongs to the "high load" level. This indicates that the driver's cognitive resources are nearing saturation and require attention.

[0275] (2) Implementation of cognitive load classification

[0276] Based on the calculated cognitive load score The system determines that the cognitive load is "high" and outputs the corresponding status description and suggestions: "Current cognitive load is high. It is recommended to reduce driving speed or reduce additional tasks to avoid cognitive overload."

[0277] VI. Implementation of Dynamic Prediction of Cognitive Load

[0278] Obtain the cognitive load assessment results at the current moment. Subsequently, the system not only needs to understand the current state, but also needs to predict future trends in order to provide early warnings before cognitive overload occurs.

[0279] (1) Implementation of temporal feature extraction

[0280] Based on the formula in the invention, extract the past A historical sequence of cognitive load over 50 time steps (each time step is 5 seconds):

[0281]

[0282] in For the length of the history window, This is a historical cognitive load sequence.

[0283] Based on the formula in the invention description, calculate the time-series statistical characteristics. First, calculate the historical average cognitive load:

[0284]

[0285] in The historical average cognitive load.

[0286] Calculate historical standard deviation:

[0287]

[0288] in The standard deviation is the historical value.

[0289] Calculate the trend of cognitive load change:

[0290]

[0291] in To understand the trend of load changes, it represents the average change at each time step.

[0292] (2) Implementation of future load forecasting

[0293] After completing the extraction of time-series features, a trained linear regression model is used to predict the future, based on the formulas and methods described in the invention. Cognitive load after 25 seconds (time steps). This linear regression model is trained on historical data using the least squares method. The training process involves collecting historical cognitive load sequences and their corresponding time-series features (mean, standard deviation, trend, current value) as input features, and then... Using the actual cognitive load after each time step as the target variable, a linear regression equation is constructed, and the regression coefficients are solved by minimizing the sum of squared prediction errors. Assume the regression coefficients obtained after training are: , , , , :

[0294]

[0295] in These are the regression coefficients obtained by training on historical data using the least squares method or gradient descent method. To predict the number of time steps, This represents the predicted future cognitive load score. The linear regression model establishes a linear mapping relationship between time-series statistical characteristics and future cognitive load, enabling prediction of future trends based on current and historical conditions.

[0296] The forecast results show that if the current trend continues, the cognitive load will reach 0.75 after 25 seconds, still within the "high load" range. However, it is close to the "extremely high load" threshold (0.8), which requires close monitoring.

[0297] VII. Implementation of Adaptive Early Warning Mechanism

[0298] To obtain the current cognitive load assessment results and future load forecast results Then, the system needs to determine whether an alert needs to be triggered based on this information. The alert mechanism first dynamically calculates the alert threshold, and then checks whether the alert triggering conditions are met.

[0299] (1) Implementation of early warning threshold calculation

[0300] According to the formula in the invention, the warning threshold is dynamically adjusted based on individual historical data and current status. It is assumed that the individual's baseline average cognitive load... (This value is calculated by collecting resting state data during the baseline calibration phase), basic warning threshold. Adjustment coefficient , :

[0301]

[0302] in The basic warning threshold is (usually 0.7). The baseline average cognitive load for individuals, Based on the current historical average cognitive load, For historical standard deviation, and For adjustment factors (usually set to 0.3 and 0.2), For a moment The dynamic early warning threshold.

[0303] The dynamic warning threshold is 0.762, which is slightly higher than the baseline threshold. This reflects that the current level of individual cognitive load is slightly higher than the baseline. Therefore, the warning threshold is raised accordingly to reduce false alarms.

[0304] (2) Implementation of early warning trigger

[0305] Based on the warning triggering conditions in the invention description, check the following three conditions:

[0306] 1. Current cognitive load check: It did not exceed the warning threshold.

[0307] 2. Predict future load checks: It did not exceed the warning threshold.

[0308] 3. Check for increasing cognitive load: (Trend threshold, typically 0.05 / time step), the rate of increase is within an acceptable range.

[0309] No warning is triggered at the current moment, but the system continues to monitor. Assume that at a subsequent moment ( Cognitive load continued to rise to After recalculating the warning threshold, at this point:

[0310] Current load exceeds warning threshold: This triggers a Level 1 warning (alert).

[0311] According to the definition of warning levels in the invention description. It falls within the scope of a Level 1 warning.

[0312] The system outputs a warning message: "Increased cognitive load detected; currently in a high-load state (cognitive load 0.78). Immediate mitigation measures are recommended, such as reducing driving speed or minimizing additional tasks."

[0313] VIII. Model Training Implementation

[0314] (1) Training dataset

[0315] Constructing a multimodal cognitive load dataset: Data source: In a simulated driving environment, 30 drivers performed driving tasks of varying difficulty (urban roads, highways, complex intersections, etc.); Task difficulty labeling: Based on task complexity (traffic density, road type, weather conditions, etc.) and expert evaluation, cognitive load was labeled with 5 levels (0.0-1.0); Data scale: 60 minutes of data were collected from each driver, totaling 1800 samples in 5-second windows, of which 1440 were used for training and 360 for testing.

[0316] Phase 1: Single-modal pre-training

[0317] According to the training strategy described in the invention, feature extraction and projection networks are pre-trained on data from each modality, enabling the model to learn the mapping relationship between each modality and cognitive load. Taking heart rate variability as an example, the mean squared error loss function is adopted according to the formula in the invention:

[0318]

[0319] For the HRV modality, specifically:

[0320]

[0321] in The number of training samples, Labels for actual cognitive load. The cognitive load score is the prediction of HRV monomodal.

[0322] Using the Adam optimizer, the learning rate The batch size was 32, and the training lasted for 20 epochs. After pre-training, the root mean square error (RMSE) of the single-modal prediction was 0.18. Similarly, other modalities (SCR, sEMG, EEG) were pre-trained separately to enable each modal encoder to acquire a strong single-modal cognitive load representation ability.

[0323] Phase Two: Multimodal Fusion Training

[0324] After completing the unimodal pre-training, the fusion network and cognitive load regression model are trained on a paired multimodal dataset according to the training strategy described in the invention. According to the formula in the invention, the total loss function is:

[0325]

[0326] The loss weight coefficient is set to , , (Typically set to 1.0, 0.5, or 0.3), the total loss function is as follows:

[0327]

[0328] in:

[0329] The mean squared error loss of cognitive load regression, where The number of training samples, Labels for actual cognitive load. The cognitive load score predicted by multimodal fusion;

[0330] The consistency loss is the sum of the prediction results for different modalities. and These are the prediction results for different modalities;

[0331] The mean squared error loss for future load forecasting, where To predict the number of time steps, For future real cognitive load, For predicted cognitive load;

[0332] Using the Adam optimizer, the learning rate The batch size was 16, and the training lasted for 30 epochs. After training, the RMSE of the multimodal fusion model decreased to 0.12, which is 33% higher than that of the single-modal method, demonstrating the effectiveness of multimodal fusion.

[0333] IX. System Deployment and Real-Time Optimization

[0334] To meet the real-time requirement (evaluation to be completed within 5 seconds), the following optimization measures are adopted:

[0335] Parallelized Feature Extraction: Feature extraction of four physiological signals is performed in parallel, reducing processing time.

[0336] Model quantization: Quantizing model weights from FP32 to INT8 improves inference speed by 2.5 times.

[0337] Caching mechanism: Cache baseline model parameters and projection matrices to avoid redundant calculations.

[0338] Sliding window optimization: Incremental update method is adopted to avoid repeated calculation of historical window features.

[0339] After optimization, the complete multimodal evaluation process takes approximately 3.2 seconds, meeting the real-time requirements.

[0340] This invention innovatively integrates four physiological modalities: heart rate variability, skin conductance response, surface electromyography (EMG) signals, and electroencephalography (EEG) signals. It establishes a deep correlation model between these different signals through a cross-modal attention mechanism. Unlike simple feature splicing, this invention employs a cross-attention mechanism with EEG as the query and other modalities as the key, enabling adaptive learning of the contribution weights of different modalities under different cognitive states, achieving true deep fusion.

[0341] First, this invention designs an individual physiological baseline model. By calculating the baseline mean and covariance matrix, the real-time feature vector is standardized relative to the baseline, effectively eliminating the influence of individual differences on the assessment results. This mechanism enables the system to adapt to the physiological characteristics of different individuals, improving the accuracy and universality of the assessment.

[0342] Secondly, this invention extracts time-series statistical features (mean, standard deviation, and trend) based on historical cognitive load sequences and uses a linear regression model to predict future cognitive load levels. This predictive capability enables the system to provide early warnings before cognitive overload occurs, offering a window of opportunity for timely intervention.

[0343] Furthermore, the warning threshold in this invention is not a fixed value, but is dynamically adjusted based on individual historical data and current status. By considering the fluctuations in individual baseline levels and current cognitive load, personalized warning thresholds are calculated, improving the accuracy and relevance of warnings and reducing false alarms and missed alarms.

[0344] Furthermore, this invention uses continuous value output (0-1 range) instead of discrete classification, which provides a more refined quantification of cognitive resource occupancy. Combined with a five-level classification system (very low, low, medium, high, very high), it ensures both the accuracy of the assessment and facilitates state judgment and decision-making in practical applications.

[0345] This invention employs a two-stage strategy of single-modal pre-training and multimodal fusion training. First, each modal encoder learns the mapping relationship between its single modality and cognitive load, then deep fusion training is performed. This strategy effectively solves the problems of high cost and difficulty in multimodal data annotation, improving the model's convergence speed and final performance.

[0346] This invention uses techniques such as parallelized feature extraction, model quantization, caching mechanisms, and sliding window incremental updates to keep the processing time within 5 seconds while maintaining high accuracy, thus meeting the response speed requirements of real-time application scenarios.

[0347] The above embodiments are merely typical illustrative methods of the present invention, and the scope of protection of the present invention is not limited thereto. All equivalent substitutions and improvements made under the concept of the present invention should fall within the scope of protection. It should be emphasized that any modifications or minor adjustments made by those skilled in the art without departing from the basic principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals, characterized in that, Includes the following steps: S1. Real-time acquisition of four physiological signals: heart rate variability (HRV), skin conductance response (SCR), surface electromyography (sEMG), and electroencephalography (EEG); each signal is preprocessed and its temporal feature sequence is extracted. S2. After completing the extraction of physiological signal features, design an adaptive baseline calibration mechanism and establish an individual physiological baseline model. This mechanism collects physiological signals of individuals at rest before real-time assessment, calculates baseline statistical characteristics, and then standardizes the real-time characteristics relative to the baseline to eliminate individual differences. S3. Construct a multimodal feature fusion network to map the features of the four physiological modalities to a unified cognitive load representation space. The fusion network first maps the features of each modality to a unified dimensional space through a projection matrix, then uses a cross-modal attention mechanism to learn the correlation between different modalities, and finally obtains a comprehensive multimodal cognitive load feature through weighted fusion. S4. Based on the fusion features, a cognitive load quantification model is constructed, which maps high-dimensional features to one-dimensional cognitive load scores. This model adopts a multilayer perceptron structure, learns the complex mapping relationship between features and cognitive load through nonlinear transformation, and outputs a continuous cognitive load score, which is the corresponding cognitive load assessment result. S5. After obtaining the cognitive load assessment results at the current moment, a time series prediction model is constructed to predict future trends based on historical cognitive load sequences. The model extracts the statistical features of historical cognitive load sequences, uses linear regression to establish the mapping relationship between these features and future cognitive load, thereby predicting the cognitive load level in the future and obtaining the future load prediction results. S6. After obtaining the current cognitive load assessment results and future load forecast results, design an adaptive early warning mechanism; The mechanism first dynamically calculates the warning threshold based on individual historical data and current status, then checks whether the current load, predicted load, and trend exceed the threshold, and triggers different levels of warnings based on the degree of exceedance.

2. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Step S1 includes: S11, Heart Rate Variability Feature Extraction Suppose that the acquired electrocardiogram signal is used to obtain the RR interval sequence after R-wave detection. ,in Indicates the first RR interval, The sequence length; The sliding window method is used to calculate the time-domain and frequency-domain characteristics; for the time window ,in For the current moment, Given the window length, calculate the following characteristics: Time-domain features include the standard deviation of the RR interval (SDNN) and the root mean square (RMSSD) of the difference between adjacent RR intervals. in This represents the number of RR intervals within the window. The mean RR interval, and Representing time respectively SDNN and RMSSD values; The power spectral density is calculated using the Fast Fourier Transform (FFT) to represent the frequency domain features. Specifically, the RR interval sequence is first interpolated and resampled to convert the non-uniformly sampled RR intervals into a uniform time series. Then, the resampled sequence is subjected to an FFT to obtain its frequency domain representation, and the power spectral density function is calculated. The power spectral density function reflects the energy distribution of heart rate variability across different frequency components; low-frequency power (LF), 0.04–0.15 Hz, and high-frequency power (HF), 0.15–0.4 Hz, are extracted. in The power spectral density function is obtained by FFT calculation. and Representing time respectively The low-frequency and high-frequency power is calculated by numerical integration method to calculate the power integral within the frequency band; The heart rate variability eigenvector is defined as: in For a moment The heart rate variability feature vector contains 5 feature dimensions; S12, Extraction of Skin Conductivity Response Features Let the collected skin conductance signal be ,in Indicates the first The conductivity value at each sampling point This represents the number of sampling points; the sampling frequency is typically 10Hz. For time windows Extract the following features: Baseline conductivity level (SCL) and conductivity response amplitude (SCR): in This represents the number of sampling points within the window. Indicates time The basic conductivity level, Indicates the magnitude of the electrical conductivity response; The conductance response frequency (SCR frequency) is calculated by detecting the number of conductance rising edges. in The difference between adjacent sampling points. The threshold for the rise, For counting functions, For window length, Indicates the number of electrical reactions per unit time; The skin conductance eigenvector is defined as: in For a moment The skin conductance feature vector; S13, Extraction of surface electromyography signal features Let the collected electromyographic signals be ,in Indicates the first Electromyography amplitude at each sampling point This represents the number of sampling points; the sampling frequency is typically 1000Hz. For time windows The root mean square (RMS) and average power frequency (MPF) are calculated. The RMS value is directly calculated from the time-domain signal and reflects the amplitude intensity of the electromyographic signal. in This represents the number of sampling points within the window. For the first Electromyography amplitude at each sampling point Indicates time The root mean square value; The calculation of the average power frequency (MPF) requires first obtaining the power spectral density; specifically, this is achieved by performing an FFT transform on the electromyographic signal and calculating the power spectral density function. The FFT transform converts a time-domain signal into a frequency-domain representation, including the power spectral density. Indicates the signal at different frequencies The power distribution on the surface; then calculate the average power frequency: in The power spectral density function of the electromyographic signal is calculated by FFT transformation. For maximum analysis frequency, Indicates time The average power frequency was obtained by calculating the frequency domain weighted average using the numerical integration method, which reflects the main frequency components of the electromyographic signal. The surface electromyography feature vector is defined as: in For a moment Surface electromyographic feature vectors; S14, EEG signal feature extraction Let the acquired multi-channel EEG signals be ,in Indicates the first Multichannel EEG data at various time points For the number of channels, The sampling frequency is typically 250Hz or 500Hz, representing the number of time points. For time windows The power features of different frequency bands are extracted using wavelet transform. Specifically, the discrete wavelet transform (DWT) or continuous wavelet transform (CWT) is used, and appropriate wavelet basis functions are selected to perform multi-scale decomposition of the EEG signal for each channel. By setting different wavelet scale parameters, the EEG signal is decomposed into five frequency bands: δ, θ, α, β, and γ, corresponding to 0.5-4Hz, 4-8Hz, 8-13Hz, 13-30Hz, and 30-100Hz, respectively. For each frequency band, wavelet coefficients at the corresponding scale are extracted. ,in Indicates the frequency band type. Indicates the channel index. Indicates the time point index; calculates the average power for each frequency band: in and Channels At any moment The wavelet coefficients in the δ and θ frequency bands, For the number of channels, The number of time points within the window. and Representing time respectively The average power in the δ and θ frequency bands; calculated using the same wavelet transform method. , and . The frequency band power ratio is used as a sensitive indicator of cognitive load. in and These represent the α / θ power ratio and the β / α power ratio, respectively. The EEG feature vector is defined as: in For a moment The EEG feature vector.

3. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Step S2 includes: S21, Baseline Feature Extraction Physiological signals were collected for 5-10 minutes while the individual was at rest as baseline data; for each physiological modality, the mean and covariance matrix of the baseline eigenvectors were calculated. in The number of time points for the baseline data. This is the baseline mean vector of heart rate variability. For the baseline covariance matrix, calculate the same. , , , , and ; S22, Feature Standardization During real-time evaluation, the current eigenvector is standardized relative to the baseline to eliminate individual differences. Specifically, this is achieved by first calculating the difference between the eigenvector and the baseline mean, and then performing a whitening transformation using the square root inverse of the covariance matrix. The calculation method is as follows: for the covariance matrix Perform eigenvalue decomposition ,in The eigenvector matrix, If the matrix is ​​a diagonal eigenvalue matrix, then ,in Let be the diagonal matrix of the inverses of the square roots of the eigenvalues. The standardized formula is: in This is the standardized heart rate variability feature vector. The square root inverse of the covariance matrix is ​​obtained through eigenvalue decomposition. This is the original feature vector at the current moment. For the baseline mean vector; calculate using the same method. , and .

4. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Step S3 includes: S31, Feature Projection Project the standardized features of each modality onto a unified... Dimensional cognitive load representation space: in , , , The projected weight matrix is ​​a learnable matrix. , , , For bias vectors, To unify the representation of spatial dimensions, , , , These are the projected feature vectors; S32, Cross-modal attention fusion A multi-head cross-attention mechanism is employed to calculate the correlation between different modalities, achieving deep fusion; EEG modalities are used as queries, and other modalities are used as keys. in This is a learnable attention parameter matrix. This represents a vector concatenation operation; Calculate attention weights and fusion features: in This is the attention weight matrix. The feature vector after attention fusion; S33, Weighted Fusion The final multimodal cognitive load features were obtained by weighted summation: in For learnable fusion weight coefficients, satisfying , This is the final multimodal cognitive load feature vector.

5. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Step S4 includes: S41, Cognitive Load Regression Multilayer perceptrons are used to map multimodal features to cognitive load scores: in This is the first layer weight matrix. This is the first layer bias vector. For the hidden layer dimension, To modify the activation function of the linear unit, For hidden layer feature vectors, This is the weight vector for the second layer. For the second-level bias scalar, For a moment The cognitive load score is 0, where 0 indicates no cognitive load and 1 indicates cognitive overload. S42, Cognitive Load Classification Based on cognitive load scores, five levels are defined: Extremely low load, They have sufficient cognitive resources and can undertake additional tasks. Low load, Cognitive resources are relatively abundant; Medium load, : Use cognitive resources appropriately; High load, Cognitive resources are nearing saturation and require attention. Extremely high load, Cognitive overload requires immediate intervention.

6. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Step S5 includes: S5. Temporal Feature Extraction Extract the past Cognitive load sequence at each time step ,in The length of the history window; Calculate the time series statistical characteristics: in For historical average cognitive load, For historical standard deviation, To show the trend of cognitive load changes; S52, Future Load Forecast Predicting the future using a linear regression model The cognitive load is calculated after a certain time step; the specific implementation method is to train a linear regression model on historical data using the least squares method or gradient descent method; the training data includes historical cognitive load sequences and their corresponding temporal characteristics, and the target variable is the future cognitive load. The actual cognitive load after each time step; the regression coefficients are solved by minimizing the sum of squared prediction errors; the prediction formula is: in These are the regression coefficients obtained by training on historical data using the least squares method or gradient descent method. The predicted future cognitive load score.

7. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Step S6 includes: S61. Calculation of Early Warning Threshold The warning threshold is dynamically adjusted based on individual historical data and current status: in Basic warning threshold, The baseline average cognitive load for individuals, and To adjust the coefficient, For a moment Dynamic early warning threshold; S62. Warning Triggering Conditions An alert is triggered when any of the following conditions are met: 1) Current cognitive load exceeds the warning threshold: ; 2) Forecast future load exceeds the threshold: ; 3) Cognitive load increases too rapidly: ; in The trend threshold is used; the warning level is divided into three levels based on the cognitive load level: Level 1 warning: Or it is predicted that it will reach this range; Level 2 warning: Or it is predicted that it will reach this range; Level 3 warning: Or it is predicted that it will reach this range.

8. The method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 1, characterized in that, Also includes: S7, Model Training Strategy; A multi-stage training strategy is adopted to optimize model performance.

9. A method for real-time assessment and adaptive early warning of cognitive load based on multimodal physiological signals according to claim 8, characterized in that, Step S7 includes: Phase 1: Single-modal pre-training Feature extractors and projection networks are pre-trained on each modality-specific dataset to enable the model to learn the mapping relationship between each modality and cognitive load. The specific training method involves updating model parameters using backpropagation and gradient descent optimizers. For each modality, single-modality features are input into the projection and regression networks to obtain the predicted cognitive load score, which is then compared with the true label to calculate the loss. The mean squared error loss function is used. in The number of training samples, Labels for actual cognitive load. The cognitive load score is predicted for a single modality. By minimizing this loss function, the projection matrix and regression network parameters are iteratively updated using gradient descent, enabling the model to learn the mapping relationship of cognitive load from single-modal features. Phase Two: Multimodal Fusion Training The fusion network and cognitive load regression model are trained on paired multimodal datasets. Specifically, the training method is as follows: the parameters of the single-modal feature extractor trained in stage one are fixed, and only the parameters of the multimodal fusion network and cognitive load regression network are optimized; the gradient of the total loss with respect to each parameter is calculated using the backpropagation algorithm, and all trainable parameters are updated simultaneously using the gradient descent optimizer; the total loss function is: in The mean squared error loss for cognitive load regression measures the difference between the multimodal fusion prediction results and the true labels. To mitigate the consistency loss of prediction results across different modalities, we encourage that prediction results from different modalities remain consistent. To compensate for the mean square error loss in future load forecasting, optimize the accuracy of the forecasting model; These are the loss weighting coefficients, used to balance the contributions of different loss terms. By jointly optimizing these three loss terms, the model can simultaneously learn accurate cognitive load assessment and prediction capabilities.