A hierarchical situational awareness prediction method and system based on multimodal fusion

CN122548660APending Publication Date: 2026-08-11SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本发明的目的是为了解决列车驾驶员分级态势感知的实时识别与前瞻预测的问题,提出了一种多模态融合的分级态势感知预测方法及系统

Benefits of technology

本发明的多模态融合的分级态势感知预测方法,利用HPMF-SANET模型实现了列车驾驶员分级态势感知的实时识别与前瞻预测。本发明的HPMF-SANET模型针对态势感知不同层级的认知异质性,自适应分配脑电与眼动的模态权重,克服了固定融合策略导致的僵化问题,在列车驾驶分级态势感知任务中,显著提升了各层级的识别准确率能够为风险预警提供更可靠的高层级判别依据。本发明的HPMF-SANET模型显式建模“感知→理解→预测”的认知传递逻辑,使高层级状态判别建立在低层级表征的条件化支持之上,得到的三个层级的感知和预测结果具有内在逻辑一致性,避免了输出结果自相矛盾的问题;同时,低层级信息可辅助高层级决策,异常状态可逐级溯源,便于分析驾驶员认知退化的起始环节。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548660A_ABST
    Figure CN122548660A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of intelligent driving assistance technology, specifically disclosing a multimodal fusion-based hierarchical situational awareness prediction method and system. The method first collects and preprocesses the EEG and eye-tracking signals of a train driver; then, it inputs these signals into a trained HPMF-SANET model. This model enhances cross-modal interaction and performs hierarchical adaptive fusion of EEG and eye-tracking features. It simulates the cognitive progression of perception, understanding, and prediction through hierarchical progressive submodules, and predicts future states through a temporal transfer submodule. Finally, it outputs the current recognition and forward-looking prediction results for each level. This invention can adaptively adjust modal weights, suppress modal suppression, ensure logical consistency of hierarchical results, and support forward-looking prediction, significantly improving the reliability of hierarchical situational awareness prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent driving assistance technology, specifically relating to a hierarchical situational awareness prediction method and system based on multimodal fusion. Background Technology

[0002] Train driving is a typical safety-critical task. Drivers need to continuously perceive environmental information such as external signals, speed limits, and operational phases to accurately understand the current operational situation and predict subsequent risk evolution. This cognitive ability—from information perception to understanding to prediction—is called situational awareness (SA). Research shows that a decline in train drivers' situational awareness is one of the main reasons for increased operational risks and weakened ability to handle operational errors and anomalies. Therefore, how to objectively and in real-time monitor the driver's graded situational awareness status and provide early warnings before the status deteriorates has become an important research topic in intelligent train auxiliary systems.

[0003] Situational awareness is typically divided into three levels: the perception level (SA1) for acquiring key environmental information, the understanding level (SA2) for comprehensive judgment of the current situation, and the prediction level (SA3) for extrapolating future states. These three levels have a cognitive progression from low to high; biases in lower-level information are propagated and amplified at each level, ultimately affecting higher-level decisions. Traditional situational awareness assessment methods primarily rely on single-modal physiological signals, such as using only electroencephalography (EEG) or only eye-tracking signals. While EEG can capture the driver's internal neural activity with high temporal resolution, its spatial resolution is limited and it is susceptible to motion artifacts. Eye-tracking signals (fixation point, pupil diameter, etc.) can directly reflect the driver's visual sampling behavior and instantaneous cognitive load, but they are difficult to reveal the deeper understanding and prediction processes. The inherent limitations of single-modal approaches restrict their reliability and accuracy in complex driving environments.

[0004] To address these issues, several multimodal assessment models integrating EEG and eye-tracking signals have emerged in recent years. However, existing multimodal assessment models still suffer from the following technical shortcomings: They employ a fixed fusion method for different cognitive levels, failing to adaptively adjust modal weights, resulting in low accuracy in high-level situational awareness recognition and an inability to provide reliable high-level discrimination criteria for risk warning. Existing research tends to treat the three levels as independent classification problems, neglecting to model the progressive relationship between perception, understanding, and prediction. This can lead to logical contradictions in the model outputs, making them difficult to interpret and trust in decision support systems. Summary of the Invention

[0005] The purpose of this invention is to solve the problem of real-time identification and forward prediction of hierarchical situational awareness for train drivers, and to propose a multimodal fusion hierarchical situational awareness prediction method and system.

[0006] The technical solution of this invention is as follows: This invention provides a hierarchical situational awareness prediction method based on multimodal fusion, comprising: The brainwave signals and eye movement signals of the train driver during the driving process are collected, and the brainwave signals and eye movement signals are preprocessed to obtain brainwave time series data and eye movement time series data. The EEG time series data and the eye movement time series data are input into the pre-trained HPMF-SANET model to obtain the current state recognition results and future state prediction results corresponding to each level of situational awareness; The HPMF-SANET model is used to sequentially extract features, enhance cross-modal interaction, and perform hierarchical adaptive fusion on the EEG time-series data and the eye-tracking time-series data to obtain hierarchical perception fusion features corresponding to each level in the situational awareness. The hierarchical perception fusion features are then sequentially simulated to advance the cognitive hierarchy and evolve the state time-series to obtain the current state recognition result and the future state prediction result.

[0007] Preferably, the electroencephalogram (EEG) signal and the eye movement signal are preprocessed, including: The EEG signal is subjected to bandpass filtering, artifact detection and removal, and baseline correction to obtain the processed EEG signal. The processed EEG signal and the eye movement signal are aligned on the time axis and then cropped according to a preset window to obtain the EEG timing data and the eye movement timing data.

[0008] Preferably, the HPMF-SANET model includes a cascaded feature extraction module, a two-branch residual cross-modal attention module, a hierarchical perceptual dynamic fusion module, and a hierarchical progressive-temporal transfer module. The feature extraction module is used to extract features from the EEG time-series data and the eye-tracking time-series data to obtain EEG features and eye-tracking features. The dual-branch residual cross-modal attention module employs a bidirectional cross-attention mechanism to perform bidirectional cross-modal interactive enhancement of the EEG features and the eye-tracking features, thereby obtaining enhanced EEG features and enhanced eye-tracking features. The hierarchical perception dynamic fusion module is used to generate adaptive fusion weights for EEG modalities and eye movement modalities corresponding to each level, and to perform weighted fusion of the EEG enhancement features and the eye movement enhancement features based on the adaptive fusion weights to obtain the hierarchical perception fusion features corresponding to each level. The hierarchical progression-temporal migration module is used to simulate the cognitive progression relationship from the perception level to the understanding level and from the understanding level to the prediction level in the situational awareness based on the hierarchical perception fusion features, to obtain the current state representation of each level, and to output the current state recognition result of each level based on the current state representation of each level; to generate the future state representation of each level based on the current state representation of each level, and to output the future state prediction result of each level based on the future state representation of each level.

[0009] Preferably, the feature extraction module includes a parallel EEG feature extraction branch and an eye movement feature extraction branch; The EEG feature extraction branch is used to extract the EEG temporal features of the EEG temporal data at different scales, and obtain the EEG features by splicing and pooling; the EEG feature extraction branch includes a cascaded first multi-scale parallel convolutional unit, a depthwise separable convolutional unit, and a second multi-scale parallel convolutional unit. The eye-tracking feature extraction branch is used to extract local temporal changes in the eye-tracking time-series data and aggregate short-time dynamic features to obtain the eye-tracking features; the eye-tracking feature extraction branch includes two cascaded one-dimensional convolutional units.

[0010] Preferably, the dual-branch residual cross-modal attention module includes a first attention branch and a second attention branch in parallel; The first attention branch is used to use the EEG feature as a query, the eye movement feature as a key and value, calculate the first cross-attention output, and add the first cross-attention output to the EEG feature to obtain the EEG enhancement feature; The second attention branch is used to calculate the second cross-attention output by taking the eye movement feature as a query and the EEG feature as a key and value, and then adding the second cross-attention output to the eye movement feature to obtain the eye movement enhancement feature.

[0011] Preferably, the hierarchical perception dynamic fusion module includes a compact vector generation unit, a weight prediction unit, and a fusion unit; The compact vector generation unit is used to perform global average pooling on the EEG enhancement features and the eye movement enhancement features respectively to obtain EEG compact vectors and eye movement compact vectors. The weight prediction unit is used to concatenate the EEG compact vector, the eye movement compact vector, and the learnable hierarchical embedding vector corresponding to the current level, and then process them through a multilayer perceptron and a softmax function to obtain the adaptive fusion weights of the EEG modality and the eye movement modality corresponding to the current level. The fusion unit is used to perform weighted fusion of the EEG compact vector and the eye-tracking compact vector according to the adaptive fusion weights of the EEG modality and eye-tracking modality corresponding to each level, so as to obtain the hierarchical perception fusion features corresponding to each level.

[0012] Preferably, the hierarchical progression-time migration module includes a hierarchical progression submodule and a time migration submodule; In the hierarchical progressive submodule, the current state representation of the perception level is obtained by mapping the hierarchical perception fusion feature of the perception level; the current state representation of the understanding level is obtained by fusing the hierarchical perception fusion feature of the understanding level and the current state representation of the perception level through a first gated weighted fusion; the current state representation of the prediction level is obtained by fusing the hierarchical perception fusion feature of the prediction level and the current state representation of the understanding level through a second gated weighted fusion. In the temporal transition submodule, for each level, a state transition increment is generated based on the current state representation. The current state representation is concatenated with the score of the current state recognition result and then a temporal gating weight is generated through temporal transition gating. The future state representation is obtained by weighted fusion of the state transition increment and the current state representation according to the temporal gating weight.

[0013] Preferably, the training process of the HPMF-SANET model employs a joint loss function, which includes current identification loss, future prediction loss, single-modal auxiliary loss, and contribution guidance loss.

[0014] Preferably, the single-mode auxiliary loss is expressed as: ; In the formula, This represents a single-modal auxiliary supervision term representing a branch of the electroencephalogram (EEG). Single-modal auxiliary supervision term for the eye-movement branch; The contribution-guided loss is expressed as: ; in, express Norm, Indicates the number of training samples. This indicates the samples generated by the hierarchical perception dynamic fusion module. Adaptive fusion weights of EEG modalities This indicates the samples generated by the hierarchical perception dynamic fusion module. Adaptive fusion weights for eye-tracking modalities Indicates sample The contribution rate of EEG modalities Indicates sample The contribution rate of eye movement modality.

[0015] This invention provides a multimodal fusion hierarchical situation awareness prediction system, applicable to the multimodal fusion hierarchical situation awareness prediction method described in any of the above embodiments, comprising: The data acquisition module is used to collect the electroencephalogram (EEG) and eye movement signals of the train driver during the driving process; The preprocessing module is used to preprocess the EEG signals and the eye movement signals to obtain EEG timing data and eye movement timing data. The hierarchical situational awareness prediction module is used to input the EEG time series data and the eye movement time series data into the pre-trained HPMF-SANET model to obtain the current state recognition results and future state prediction results corresponding to each level of situational awareness. The HPMF-SANET model is used to sequentially extract features, enhance cross-modal interaction, and perform hierarchical adaptive fusion on the EEG time-series data and the eye-tracking time-series data to obtain hierarchical perception fusion features corresponding to each level in the situational awareness. The hierarchical perception fusion features are then sequentially simulated to advance the cognitive hierarchy and evolve the state time-series to obtain the current state recognition result and the future state prediction result.

[0016] The beneficial effects of this invention are: This invention presents a multimodal fusion-based hierarchical situational awareness prediction method that utilizes the HPMF-SANET model to achieve real-time identification and forward prediction of hierarchical situational awareness for train drivers. The HPMF-SANET model of this invention adaptively allocates modal weights for EEG and eye movements to address the cognitive heterogeneity at different levels of situational awareness, overcoming the rigidity problem caused by fixed fusion strategies. In the hierarchical situational awareness task for train drivers, it significantly improves the identification accuracy of each level, providing a more reliable high-level discrimination basis for risk warning. The HPMF-SANET model of this invention explicitly models the cognitive transmission logic of "perception → understanding → prediction," enabling high-level state discrimination to be built upon the conditional support of low-level representations. The resulting perception and prediction results at the three levels have inherent logical consistency, avoiding the problem of contradictory output results. Simultaneously, low-level information can assist high-level decision-making, and abnormal states can be traced step-by-step, facilitating the analysis of the initiation points of driver cognitive decline. Attached Figure Description

[0017] Figure 1 The diagram shows a flowchart of a hierarchical situational awareness prediction method based on multimodal fusion. Figure 2 The diagram shown is a schematic representation of the overall process of the HPMF-SANET model. Figure 3 The diagram shows the confusion matrix for the six task groups. Figure 4 The diagram shows the structural block diagram of a hierarchical situational awareness prediction system that integrates multiple modalities. Detailed Implementation

[0018] Exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary and are intended to illustrate the principles and spirit of the invention, and are not intended to limit the scope of the invention.

[0019] In a first aspect, embodiments of the present invention provide a hierarchical situational awareness prediction method based on multimodal fusion. Please refer to... Figure 1 , Figure 1 The diagram shows a flowchart of a hierarchical situational awareness prediction method based on multimodal fusion.

[0020] like Figure 1 As shown, the hierarchical situational awareness prediction method based on multimodal fusion in this embodiment includes the following steps: Step 1: Collect EEG and eye movement signals from the train driver during the driving process, preprocess the EEG and eye movement signals to obtain EEG time series data and eye movement time series data.

[0021] In this embodiment, electroencephalogram (EEG) signals and eye movement (EM) signals are acquired simultaneously. The EEG signal sampling rate is 256Hz, and the EM signal is acquired using a head-mounted eye tracker with a sampling rate of 100Hz, mainly including the pupil pixel area and the horizontal and vertical coordinates of the gaze. Before the actual acquisition, eye movement calibration is performed using a 9-point grid.

[0022] Since EEG and eye movement data are easily affected by motion, electromyography (EMG) interference, and equipment status, preprocessing of the acquired EEG and eye movement signals is necessary. This mainly includes two parts: EEG signal purification and multimodal time alignment.

[0023] In this embodiment, the EEG signal and eye movement signal are preprocessed, including: bandpass filtering, artifact detection and removal, and baseline correction of the EEG signal to obtain the processed EEG signal; the processed EEG signal and eye movement signal are aligned on the time axis and then truncated according to a preset window to obtain EEG time series data and eye movement time series data.

[0024] Specifically, before data acquisition, the impedance of all EEG channels was controlled below 10kΩ. The raw EEG signal was first subjected to a fourth-order zero-phase Butterworth bandpass filter with a passband of 1–40Hz to preserve cognitively relevant frequencies and suppress low-frequency drift and high-frequency noise. Subsequently, an automated artifact detection method was used to scan continuous EEG data channel by channel, with a detection window of 500ms and a step size of 250ms. If the potential jump between adjacent sampling points exceeded 90μV / ms, it was identified as an EMG artifact; if the signal standard deviation was less than 0.8μV within 200ms, it was identified as an electrode contact artifact. For detected artifacts, samples with a duration of 500ms or more were directly discarded; artifact segments with a duration of less than 500ms were extended by 250ms before and after them and then deleted.

[0025] To facilitate the fusion of EEG and eye-tracking data, the EEG signal was downsampled to 100Hz after filtering. Subsequently, the two modalities were synchronously segmented based on probe events. Using the moment of the probe event as the zero point, the 0-5s reaction period data was extracted as the analysis window. An additional baseline segment of -1.5-0s was extracted from the EEG signal, and baseline correction was performed by subtracting the baseline mean. Within a total time range of 6.5s, if data loss occurred in a particular EEG channel due to artifact removal, it was handled according to the duration of the loss: samples with consecutive missing values ​​exceeding 500ms were discarded; samples with consecutive missing values ​​not exceeding 500ms were completed using natural cubic spline interpolation.

[0026] Step 2: Input the EEG time series data and eye movement time series data into the pre-trained HPMF-SANET model to obtain the current state recognition results and future state prediction results corresponding to each level of situational awareness.

[0027] In this embodiment, the HPMF-SANET (Hierarchical Progressive Multimodal FusionSituation Awareness Network) model is used to sequentially extract features, enhance cross-modal interaction, and perform hierarchical adaptive fusion on EEG temporal data and eye movement temporal data to obtain hierarchical perception fusion features corresponding to each level in situational awareness. The hierarchical perception fusion features are then sequentially simulated to advance cognitive levels and evolve state temporally to obtain the current state recognition result and the future state prediction result.

[0028] In this embodiment, the current state identification result is determined in real time by the HPMF-SANET model to determine whether the driver is in a high situational awareness state (high SA), a low situational awareness state (low SA), or a high situational awareness state (low SA) at the three situational awareness levels of perception (SA1), understanding (SA2), and prediction (SA3). For example, three sets of binary labels are output: SA1 is high SA, SA2 is low SA, and SA3 is low SA, reflecting the quality of the driver's performance on each level of cognitive tasks at the current moment.

[0029] In this embodiment, the future state prediction result is a binary classification label (high SA / low SA) inferred by the HPMF-SANET model for the driver's situational awareness state at each of the three levels five minutes from now. This prediction result is used for early warning, enabling the assistance system to take intervention measures before the driver's cognitive state actually deteriorates, such as adjusting the information display method or issuing risk warnings.

[0030] Furthermore, the construction and training process of the HPMF-SANET model in this embodiment will be described in detail.

[0031] Please see Figure 2 , Figure 2 The diagram shown is a schematic representation of the overall flow of the HPMF-SANET model. Figure 2 As shown, the HPMF-SANET model includes a cascaded feature extraction module, a two-branch residual cross-modal attention module (RCAM), a hierarchical perceptual dynamic fusion module (HDFM), and a hierarchical progressive-temporal transfer module (HPTM).

[0032] The HPMF-SANET model takes EEG and eye-tracking time-series data as input. It then learns the temporal representations of EEG and eye-tracking respectively through a feature extraction module, and models the interaction between the two modalities using a two-branch residual cross-modal attention module. Further, a hierarchical perception dynamic fusion module adaptively assigns fusion weights for EEG and eye-tracking modalities at different SA levels, thus obtaining hierarchically related perception fusion features. These hierarchical perception fusion features are then input into a hierarchical progression-temporal transfer module. The hierarchical progression branch explicitly models the progression from perception to understanding to prediction (SA1→SA2→SA3), while the temporal transfer branch generates future state representations based on the current state representation, achieving joint modeling of current recognition and future prediction.

[0033] Specifically, the feature extraction module extracts features from EEG and eye-tracking time-series data to obtain EEG and eye-tracking features. The bi-branch residual cross-modal attention module employs a bi-directional cross-attention mechanism to enhance EEG and eye-tracking features through bi-directional cross-modal interaction, resulting in enhanced EEG and eye-tracking features. The hierarchical perception dynamic fusion module generates adaptive fusion weights for EEG and eye-tracking modalities at each level, and performs weighted fusion of the enhanced EEG and eye-tracking features based on these adaptive weights to obtain hierarchical perception fusion features for each level. The hierarchical progression-temporal transfer module simulates the cognitive progression from the perception level to the understanding level and from the understanding level to the prediction level in situational awareness based on the hierarchical perception fusion features, obtaining the current state representation of each level, and outputting the current state recognition result for each level based on the current state representation; it also generates the future state representation for each level based on the current state representation, and outputs the future state prediction result for each level based on the future state representation.

[0034] In this embodiment, to fully utilize the complementary information of EEG signals and eye movement signals in train driving situational awareness and recognition, a feature extraction module is constructed, comprising parallel EEG feature extraction branches and eye movement feature extraction branches, to perform representation learning on the two modalities respectively. Let the EEG time-series data input to the EEG feature extraction branch be... The eye-tracking time-series data input to the eye-tracking feature extraction branch is: ,in, Indicates the number of brainwave channels. Indicates the number of eye movement signal components. This represents the number of sampling points within the time window. Considering that EEG signals have more complex multi-timescale dynamic patterns, while eye movement signals mainly reflect short-term fixation and pupillary changes, multi-scale convolutional EEG feature extraction branches and lightweight convolutional eye movement feature extraction branches are used for modeling, respectively. To facilitate subsequent multimodal fusion, the outputs of the two branches are mapped to a consistent feature dimension.

[0035] Specifically, the EEG feature extraction branch is used to extract EEG temporal features from EEG temporal data at different scales, and obtains EEG features through concatenation and pooling. The EEG feature extraction branch includes a cascaded first multi-scale parallel convolutional unit, a depthwise separable convolutional unit, and a second multi-scale parallel convolutional unit. Considering that EEG signals simultaneously contain short-term transient changes and dynamic patterns over a longer time range, this embodiment uses a multi-scale one-dimensional convolutional structure for feature extraction to simultaneously capture local patterns and longer-term dependencies.

[0036] Specifically, the first multi-scale parallel convolutional unit first applies four sets of parallel one-dimensional convolutions to the input EEG temporal data, with the kernel length decreasing proportionally to extract multi-scale temporal features under different receptive fields. Subsequently, each set of features undergoes batch normalization, ELU activation, and Dropout to improve training stability and alleviate overfitting. The depthwise separable convolutional unit further uses depthwise separable convolution to perform channel compression and recombination on the features at each scale to extract inter-channel correlation information and reduce the number of parameters. After concatenating the four sets of output features along the channel dimension, average pooling is used to complete temporal downsampling. The second multi-scale parallel convolutional unit applies multi-scale convolutions again to the dimensionality-reduced features to further extract higher-level temporal representations; its output is also concatenated and average pooled to finally obtain the EEG feature representation. , This indicates the feature channel dimension of the output of the EEG feature extraction branch. This represents the length of the feature sequence after time-dimension downsampling.

[0037] In this embodiment, the final output dimension of the EEG feature extraction branch is: This design can preserve multi-scale temporal information in EEG signals that is discriminative for SA state recognition while maintaining low computational complexity.

[0038] Compared to EEG signals, eye movement signals have a relatively simple temporal structure, mainly reflecting short-term fixation changes, saccade shifts, and pupillary changes. Therefore, in this embodiment, a lightweight convolutional network is used for feature extraction in the eye movement feature extraction branch to extract effective temporal patterns while ensuring computational efficiency.

[0039] Specifically, the eye-tracking feature extraction branch is used to extract local temporal changes in eye-tracking time-series data and aggregate short-term dynamic features to obtain eye-tracking features; the eye-tracking feature extraction branch includes two cascaded one-dimensional convolutional units.

[0040] Specifically, the first one-dimensional convolutional unit is used to extract local temporal variations, and the second one-dimensional convolutional unit further aggregates short-term dynamic patterns. Each one-dimensional convolutional unit is followed by non-linear activation and Dropout to improve the model's expressive power and generalization performance. Finally, average pooling is used to obtain a compact eye-tracking feature representation. To facilitate subsequent fusion with EEG features, the output dimension of the eye-tracking feature extraction branch was set to be consistent with that of the EEG feature extraction branch, i.e. .

[0041] Through the above design, the eye-tracking feature extraction branch can effectively extract short-term dynamic features related to the driver's visual sampling behavior and complement the EEG feature extraction branch, providing multimodal input for subsequent SA state recognition and prediction.

[0042] For example, the specific structural parameters of the feature extraction module in this embodiment are shown in Table 1.

[0043] Table 1

[0044] It is understandable that EEG and eye movements reflect a driver's internal cognitive processing and external visual sampling behavior, respectively, and are naturally complementary in situational awareness modeling. However, if a direct splicing method is used for fusion, it is difficult to fully model the information interaction between the two modalities, and it is also easy for the modality with stronger discriminative ability to dominate the training process. To this end, this embodiment proposes a Dual-Branch Residual Cross-Attention Module (RCAM), which explicitly models the interaction relationship between EEG and eye movements before fusion to enhance the discriminative power of single-modal features.

[0045] In this embodiment, the dual-branch residual cross-modal attention module includes a first attention branch and a second attention branch operating in parallel. The first attention branch uses EEG features as a query and eye movement features as a key and value to calculate a first cross-attention output and adds it to the EEG features to obtain EEG enhancement features. The second attention branch uses eye movement features as a query and EEG features as a key and value to calculate a second cross-attention output and adds it to the eye movement features to obtain eye movement enhancement features.

[0046] Specifically, let the EEG features and eye movement features output by the feature extraction module be respectively... To facilitate cross-attention along the time dimension, it is first transposed into a token representation. In this embodiment, a bidirectional cross-attention mechanism is employed, enabling the first attention branch (EEG branch) to extract supplementary information from eye-tracking features, while the second attention branch (eye-tracking branch) simultaneously extracts supplementary information from EEG features. The following presents the single-head attention pattern and the representation of EEG enhancement features: ; ; In the formula, The query matrix represents the branches of the EEG. This represents the linear projection weight matrix of the EEG query. The key matrix representing the eye-tracking branches. This represents the linear projection weight matrix of the eye-tracking keys. The value matrix representing the eye-tracking branch. This represents the linear projection weight matrix of eye-tracking values. This indicates the dimension across the attention projection, used to scale the attention score. This is a characteristic of enhanced brainwave activity.

[0047] Similarly, the characteristics of eye movement enhancement in the eye movement branches can be obtained: ; ; In the formula, The query matrix representing the eye-tracking branches. This represents the linear projection weight matrix for eye-tracking queries. The key matrix representing the branches of EEG, This represents the linear projection weight matrix of the EEG synapses. The value matrix representing the branches of the EEG, The linear projection weight matrix represents the EEG values. This indicates the dimension across the attention projection, used to scale the attention score. This is an eye-movement enhancement feature.

[0048] The above processes represent how EEG and eye movement, while preserving the original representations, introduce supplementary information from another modality. The resulting enhanced EEG and eye movement features will then serve as inputs to the subsequent motion-level perception dynamic fusion module.

[0049] Understandably, if fixed weights are still used for fusion after obtaining cross-modal complementary EEG and eye-tracking enhancement features, the differences in contributions between the two modalities at different SA levels cannot be reflected. Considering that SA1, SA2, and SA3 correspond to the perception, understanding, and prediction levels, respectively, the dependence of different levels on EEG and eye-tracking may not be the same. Therefore, this embodiment proposes a Hierarchy-Aware Dynamic Fusion Module (HDFM), which enables the model to adaptively adjust the fusion weights of EEG and eye-tracking for different SA levels.

[0050] In this embodiment, the hierarchical perception dynamic fusion module includes a compact vector generation unit, a weight prediction unit, and a fusion unit. The compact vector generation unit performs global average pooling on the EEG enhancement features and eye-tracking enhancement features respectively to obtain EEG compact vectors and eye-tracking compact vectors. The weight prediction unit concatenates the EEG compact vector, the eye-tracking compact vector, and the learnable hierarchical embedding vector corresponding to the current level, and then processes them through a multilayer perceptron and a softmax function to obtain the adaptive fusion weights for the EEG modality and eye-tracking modality corresponding to the current level. The fusion unit performs weighted fusion of the EEG compact vector and the eye-tracking compact vector according to the adaptive fusion weights for the EEG modality and eye-tracking modality corresponding to each level to obtain the hierarchical perception fusion features corresponding to each level.

[0051] Specifically, the compact vector generation unit performs global average pooling on the EEG enhancement features and eye movement enhancement features to obtain compact vectors for both EEG and eye movement modalities: ; in, This represents the pooling operation. Let the first... The learnable layer embedding vectors of each SA level are ,in These correspond to SA1, SA2, and SA3, respectively. The hierarchical information and the bimodal compact vector are input into the weight prediction unit to obtain the adaptive fusion weights of EEG and eye movement at the current level. ; in, and They represent the first Adaptive fusion weights of EEG and eye movement at each SA level, and satisfying Subsequently, the fusion unit constructs hierarchical perceptual fusion features: ; Through the above design, the model no longer uses uniform modal weights for different levels, but adaptively adjusts the contributions of the two modes according to the task attributes of SA1, SA2 and SA3, thereby improving the pertinence of hierarchical SA state modeling.

[0052] Because the formation of hierarchical SAs exhibits a cognitive progression from low to high, modeling the three levels completely independently may make it difficult to utilize lower-level information to assist in the judgment of higher-level states. This embodiment constructs a Hierarchical Progressive-Temporal Transfer module (HPTM) based on hierarchical perceptual fusion features to achieve current state recognition and future state prediction. The Hierarchical Progressive-Temporal Transfer module includes: a hierarchical progressive submodule for modeling the progressive relationship of SA1→SA2→SA3, and a hierarchical progressive branch (TTB) for future state prediction.

[0053] In this embodiment, within the hierarchical progression submodule, the current state representation of the perception level is obtained by mapping the hierarchical perception fusion feature of the perception level. The current state representation of the understanding level is obtained by fusing the hierarchical perception fusion feature of the understanding level with the current state representation of the perception level through a first-gated weighted fusion. The current state representation of the prediction level is obtained by fusing the hierarchical perception fusion feature of the prediction level with the current state representation of the understanding level through a second-gated weighted fusion.

[0054] Specifically, let the first The hierarchical perceptual fusion features of each SA level are represented as follows: First, a lightweight mapping layer is used to map it to a unified latent space, resulting in an initial hierarchical representation: ; in, Represents a mapping function. This represents the dimension of the unified latent space. For SA1, its initial hierarchical representation is directly taken as the current state representation: ; For SA2 and SA3, low-level auxiliary information is introduced through a step-by-step gating method. Specifically, the current state representation of SA2 is defined as: ; in, This represents the initial hierarchy of SA1. This represents the initial hierarchy of SA2. This indicates element-wise multiplication. The gating weight, determined by both SA1 and SA2 features, is defined as follows: ; In the formula, This represents the multilayer perceptron used to generate the SA1 to SA2 gating weights. This represents the Sigmoid activation function, used to constrain the gate weights between 0 and 1. This represents vector concatenation. express 3D real space.

[0055] Similarly, the current state representation of SA3 is defined as: ; in, This represents the initial hierarchy of SA3. The gating weight, determined by both SA2 and SA3 features, is defined as follows: ; In the formula, This represents the multilayer perceptron used to generate the SA2 to SA3 gating weights. This represents the Sigmoid activation function, used to constrain the gate weights between 0 and 1.

[0056] Through the above design, the current state representation of the high-level SA can adaptively absorb the auxiliary information of the lower level while retaining its own information, thereby reflecting the progressive relationship between SA1, SA2 and SA3.

[0057] Based on this, the current state is identified using the hierarchical current state representation. For the first... Each SA level has a classification score for its current state defined as: ; In the formula, For the first The classification score of the current state at each SA level. The weight matrix of the current state recognition head. For the first The current state representation of each SA level, This is the bias term for the current state recognition head.

[0058] Furthermore, the binary classification output at the current time is obtained according to the threshold rule, which is the current state recognition result: ; in, This indicates an indicator function that takes the value 1 when a condition is met and 0 otherwise. In this embodiment, Indicates the first The current state identification result of each SA level at the current moment.

[0059] In this embodiment, in the temporal transition submodule, for each level, a state transition increment is generated based on the current state representation. The current state representation is concatenated with the score of the current state recognition result and then a temporal gating weight is generated through temporal transition gating. The future state representation is obtained by weighted fusion of the state transition increment and the current state representation according to the temporal gating weight.

[0060] Specifically, the prediction of the next 5 minutes is considered as the evolution of the current state over time. First, the current state is represented based on the current level. Incremental learning state transition: ; in, This represents the transition mapping function. Then, the current state representation of the current level and the classification score at the current time step are input into the temporal transition gating to construct the future state representation: ; ; in, This represents the temporal gating weights, determined jointly by the current state representation at the current level and the classification score at the current moment. Finally, the... The classification score of each SA level in the next 5 minutes, i.e., the classification score of the future state, is defined as: ; In the formula, For the first The classification score of the future state at each SA level The weight matrix for the future state prediction head. The bias term for the future state prediction head. For the first The future state representation of each SA level.

[0061] The corresponding binary classification output, i.e. the future state prediction result, is: ; in, Indicates the first The SA level provides future state prediction results at future moments. This design prevents future prediction from being viewed as a parallel task independent of current identification, but rather as an evolutionary modeling process built upon the current state representation.

[0062] In this embodiment, the HPMF-SANET model is trained using the AdamW optimizer with standard parameter settings (β1=0.9, β2=0.999). =1×10 -8 The initial learning rate is set to 1×10. -3 The batch size was set to 32, and the weight decay coefficient was set to 0.01. An early stopping strategy was employed during training: training was stopped if the validation set loss did not decrease for 20 consecutive epochs. Network parameters were initialized using He to improve training stability under non-linear activation conditions.

[0063] For the dataset used in training the HPMF-SANET model, physiological data was collected from 31 participants with train driving qualifications and complete skills training experience through a highly realistic CR400AF high-speed train simulator built on the Open Rails platform. Specifically, a 32-channel EEG system conforming to the international 10-10 standard layout was used, with saline electrodes to rapidly reduce scalp impedance, and a sampling rate of 256Hz. Eye-tracking data was collected using TOBII Pro Glasses 3 at a sampling rate of 100Hz. The experimental scenario was set as a real-world railway line, employing a within-subject design, ensuring all participants completed the simulated driving task under identical weather and route conditions.

[0064] A tiered auditory detection task was designed for train driving scenarios. The train driving simulation lasted approximately 180 minutes. The system triggered an auditory probe task at pseudo-random time intervals (approximately 15-25 seconds), with each probe trigger considered as a sample. Each sample included current operating context information, a 5-second continuous driving and physiological signal record captured from the moment the probe event occurred (time zero), and a subsequent auditory judgment task. After hearing the voice question, the driver needed to respond as quickly as possible using physical buttons on the control panel. To avoid additional eye-tracking and EEG artifacts caused by visual pop-ups, all situational awareness detection tasks were presented auditorily. Three types of probe tasks were designed, corresponding to the three levels of situational awareness; each type of task had two replaceable task templates that could be randomly selected during the experiment to reduce the likelihood of participants forming fixed response strategies.

[0065] During the train driving simulation, the system triggered auditory probe tasks at pseudo-random time intervals, with probe types randomly selected from three task categories: SA1, SA2, and SA3. To support SA prediction 5 minutes in advance, the experiment additionally included some paired probes at the same level to establish the correspondence between the current moment and the SA state at the same level 5 minutes later. Participants responded via physical buttons on the control panel based on voice prompts while continuously driving. EEG, eye movements, train speed, signal status, and station entry / exit phases were recorded simultaneously throughout the experiment. Data collection ceased after all driving tasks were completed. The EEG and eye movement data for each participant were preprocessed to obtain a sample set. The specific preprocessing process was consistent with the data preprocessing process described in the HPMF-SANET model application phase and will not be elaborated here.

[0066] To label each probe sample as either high situational awareness (high SA) or low situational awareness (low SA), a two-stage label inference method is constructed based on behavioral performance metrics accuracy (ACC) and reaction time (RT). Considering that SA1, SA2, and SA3 correspond to different levels of situational awareness processing, label inference is performed for each of the three levels separately. Specifically, correctness of the response is used as the initial criterion: incorrect responses are directly labeled as low SA. For correct response samples, unsupervised clustering is further performed in conjunction with reaction time to distinguish between high SA and low SA states.

[0067] Let the first Each SA level ( ) The accuracy per sample is Its initial label definition is: ; in, This indicates that the driver answered the current probe incorrectly, meaning they failed to make a correct judgment at the corresponding SA level; therefore, this sample is directly labeled as low SA. For The sample only indicates that the driver's final answer was correct, but their situational awareness level may still vary. Therefore, it is temporarily recorded as an unclassified sample and will be further determined in the next stage in conjunction with the reaction.

[0068] For the initial unclassified samples (i.e.) The sample was further analyzed to infer the SA status based on reaction time. Considering the differences in baseline reaction speed among different participants, and the possible differences in RT distribution among different task templates under the same SA level, RT was first standardized within participants and within tasks to reduce the impact of individual and task differences on the clustering results.

[0069] Let the first The number of subjects was in the first SA level, task template The initial reaction was as follows: The standardized result is expressed as: ; in, and These represent the mean and standard deviation of the subject's response time under the corresponding SA level and task template, respectively. To prevent extremely small constants with a denominator of zero, this treatment effectively reduces the differences in baseline velocity among different participants and the overall duration shift caused by different task templates, allowing the RT of each sample to be compared on the same scale.

[0070] After standardization is completed, all SA levels will be included. The samples were aggregated and clustered using a two-component Gaussian Mixture Model (GMM): ; in, For the first The mixing coefficient of the Gaussian components, and These are the mean and variance of the corresponding components, respectively. This represents all parameters of the GMM at this level. Since the SA state is divided into high and low categories, the default number of GMM components is 2.

[0071] With ACC=1, the correct response samples are further differentiated based on standardized reaction time. For the same SA level, after GMM clustering, the cluster with the smaller mean is denoted as... Clusters with larger means are denoted as Subsequently, The samples in the data are marked as high SA, and the data will be... The samples in the data are labeled with low SA, i.e.: Marked as high SA; : Marked as low SA.

[0072] Combining the clustering results of ACC and GMM, the first In the SA level, the first The final label for each sample is defined as: .

[0073] To ensure the effectiveness of current state recognition, future state prediction, and multimodal fusion weight learning simultaneously, the training process of the HPMF-SANET model in this embodiment adopts a joint loss function, which includes current recognition loss, future prediction loss, single-modal auxiliary loss, and contribution guidance loss.

[0074] set up and They represent the first Given the current true label and the true label for the next 5 minutes at each SA level, the current recognition loss is defined as: ; in, This represents the binary cross-entropy loss calculated based on the classification score. Indicates the first Number of samples at each SA level.

[0075] For future predictions, since not all samples can effectively pair with probes at the same level 5 minutes later, a mask variable is introduced. The loss is defined as the loss that indicates whether a sample has a valid future label: .

[0076] Furthermore, a single-modal auxiliary classification head was set up for both the EEG and eye-tracking branches, and a single-modal auxiliary loss was added during the training phase, denoted as... This loss is used to maintain the discriminative power of each mode and to provide a basis for estimating the contributions of subsequent modes. Let... and They represent the first Samples in each SA level The unimodal auxiliary classification scores of the EEG and OMG branches are used to define the unimodal auxiliary loss as follows: ; ; ; in, and These represent the single-modal auxiliary supervision terms for the EEG and eye-tracking branches, respectively.

[0077] Furthermore, to address the lack of explicit guidance for dynamic fusion weights, this embodiment introduces a contribution-guided loss. Let the EEG and eye-tracking adaptive fusion weights predicted by the hierarchical perception dynamic fusion module be respectively... and Based on the classification score output by the single-modal auxiliary classification head, the sample is defined. The modal contribution rate is: ; in, and Representing samples respectively EEG and eye-tracking monomodal auxiliary classification scores, To prevent extremely small constants with a denominator of zero, the above definition normalizes the absolute value of the single-modal score to characterize the two modes in the sample. The relative discriminative contribution. Based on this, the contribution-guided loss is defined as: ; in, express Norm, Indicates the number of training samples. This indicates the samples generated by the hierarchical perception dynamic fusion module. Adaptive fusion weights of EEG modalities This indicates the samples generated by the hierarchical perception dynamic fusion module. Adaptive fusion weights for eye-tracking modalities Indicates sample The contribution rate of EEG modalities Indicates sample The eye-tracking modal contribution rate. This loss constraint dynamically fusion weights are consistent with the actual modal contribution.

[0078] In summary, the joint loss function in this embodiment is expressed as: ; in, , and These are the weighting coefficients for each loss term. For example, , and Set them to 1.0, 0.3 and 0.1 respectively.

[0079] By jointly optimizing the HPMF-SANET model, it is possible to simultaneously learn the current state recognition, future state prediction, and multimodal adaptive fusion of hierarchical SA, while maintaining the discriminative ability of the single-modal branch and the stability of the fusion weight learning during training.

[0080] To more clearly illustrate the overall execution process of the HPMF-SANET model during the training and application phases, this embodiment further provides pseudocode for the model training and hierarchical situational awareness prediction process, as shown in Algorithm 1. This pseudocode is only used to illustrate the data flow relationships between modules and does not constitute a limitation on specific programming languages, deep learning frameworks, or parameter update methods.

[0081] Furthermore, the HPMF-SANET model of this invention possesses excellent causal analysis capabilities and interpretability at the structural level, effectively overcoming the black-box decision-making problem commonly found in existing deep learning models. Specifically: First, the hierarchical progressive submodule explicitly models the cognitive transmission link of perception → understanding → prediction through gating weights. The gating weights can quantify the dependence of higher-level discrimination on lower-level states, enabling the hierarchical situational awareness results to have a causal path that can be traced step by step. Second, the hierarchical perception dynamic fusion module independently outputs adaptive fusion weights for EEG and eye movement at each SA level, which can directly read and quantify the relative contributions of the two modalities at different cognitive levels. Combined with contribution-guided loss, the fusion weights are strictly aligned with the actual discrimination contribution of the modality, avoiding the uninterpretability of weight learning. Third, the temporal transfer submodule controls the fusion ratio of state transfer increment and current state representation through explicit temporal gating weights, making the evolutionary basis of future prediction results directly readable. Fourth, the EEG and eye movement single-modality auxiliary classification heads can independently provide the discrimination output of each modality, thereby supporting modality-level attribution analysis of fusion decisions. The above mechanism enables each step of the model's decision-making to be explained by explicit weights and hierarchical state representations, significantly improving the model's transparency and credibility in driving assistance decision-making scenarios.

[0082] The multimodal fusion-based hierarchical situational awareness prediction method of this invention utilizes the HPMF-SANET model to achieve real-time identification and forward prediction of hierarchical situational awareness for train drivers. The HPMF-SANET model adaptively allocates modal weights for EEG and eye movements to address the cognitive heterogeneity at different levels of situational awareness, overcoming the rigidity problem caused by fixed fusion strategies. In the hierarchical situational awareness task for train drivers, it significantly improves the identification accuracy of each level, providing a more reliable high-level discrimination basis for risk warning. The HPMF-SANET model explicitly models the cognitive transmission logic of "perception → understanding → prediction," enabling high-level state discrimination to be based on the conditional support of low-level representations. The resulting perception and prediction results at the three levels have inherent logical consistency, avoiding the problem of contradictory output results. Simultaneously, low-level information can assist high-level decision-making, and abnormal states can be traced step-by-step, facilitating the analysis of the initiation points of driver cognitive degradation.

[0083] Furthermore, the effectiveness of the multimodal fusion hierarchical situational awareness prediction method of the present invention is illustrated through simulation comparison experiments.

[0084] Leave-one-subject-out cross-validation (LOSO-CV) was used to evaluate the model on data from 31 participants. Specifically, one participant was selected as the test set in each round, and the remaining 30 participants were used as training and validation data to evaluate the model's generalization ability under cross-participant conditions. The main hyperparameters of the HPMF-SANET model are shown in Table 2. Among them, the number of EEG channels ( ), EEG sampling rate and number of EM components ( The number of EM components is determined by the input data. The input EEG signal contains 32 channels with a sampling rate of 256Hz. The EM signal has a sampling rate of 100Hz, and its input includes the pupil area of ​​both eyes and the horizontal and vertical coordinates of the fixation point; therefore, the number of EM components ( The value is 6. To achieve multimodal fusion, the EEG signal is further resampled and aligned to 100Hz in the time dimension to match the EM mode.

[0085] Dropout ratio ( The number of attention heads is set to 0.2. Since the bi-branch residual cross-modal attention module uses a single-head cross-attention mechanism, the number of cross-attention heads ( ) is set to 1. (Across attention projection dimensions) The hierarchical embedding dimension is set to 64. ) and the hidden layer dimension of the fusion weight prediction MLP (Multilayer Perceptron) The initial learning rate (%) was set to 16 to provide sufficient representational power while maintaining the model's lightweight design. During training, the initial learning rate (%) was set to 16. ) set to Batch size ( The value is set to 32. In the loss function, , and The weight coefficients are set to 1.0, 0.3 and 0.1 respectively, so as to ensure that prediction supervision plays a dominant role in the overall training, while using unimodal assisted learning and contribution-guided regularization to improve training stability without weakening the optimization effect of the main task.

[0086] Table 2

[0087] To quantitatively evaluate the recognition and 5-minute advance prediction performance of the HPMF-SANET model on SA1, SA2, and SA3, high-SA was used as the positive class, and accuracy (Acc), recall (Re), precision (Pre), and F1 score were calculated. The results are shown in Table 3. Overall, the HPMF-SANET model achieved good classification performance on all three SA layers, and its recognition performance was generally better than its 5-minute advance prediction performance.

[0088] Looking at different SA levels, the models exhibit a consistent trend across both types of tasks: SA1 is the best, followed by SA2, and SA3 is relatively the worst. This pattern is consistent across the four metrics: Acc, Re, Pre, and F1. SA1-Rec achieves the highest performance across the six task groups, SA2-Rec is slightly lower but still maintains a high level, while SA3-Rec further declines. The same trend is observed in the prediction task, with SA1-Pred performing best and SA3-Pred performing worst. This indicates that as the task level progresses from SA1 to SA3, the overall model performance decreases, and this trend is consistent across both recognition and prediction tasks.

[0089] Table 3

[0090] Please see Figure 3 , Figure 3The diagram shows the confusion matrices for the six tasks. As can be seen, the diagonal elements of each group's results are significantly higher than the off-diagonal elements, indicating that the HPMF-SANET model has good discriminative ability for both high-SA and low-SA samples. Meanwhile, the off-diagonal elements in the prediction task are generally higher than those in the recognition task, consistent with the results in Table 3 where the recognition task outperforms the prediction task. Stratified, the confusion matrix for SA1 has the most concentrated diagonal elements, while SA3, especially in the prediction task, has a relatively larger number of off-diagonal elements, corresponding to its lower Acc and F1 scores.

[0091] To verify the effectiveness of the method of this invention, several representative baseline models were selected for comparison, including methods based solely on EEG signals (EEGNet, MDCNet, ConTraNet, GATransformer, HCANN) and multimodal methods fusing EEG and EM signals (CMGFNet, FGFRNet). Two monomodal versions of the HPMF-SANET model of this invention were also included: EEGBaseline and EMBaseline, corresponding to models retaining only the EEG branch and only the eye-tracking branch, respectively. HPMF-SANET represents the complete model fusing EEG and eye-tracking information.

[0092] Table 4 shows the experimental results of each method under Rec-ACC and Pred-ACC in three scenarios (SA1, SA2, SA3). To systematically evaluate the impact of different methods and scenarios on model performance, a two-way repeated measures ANOVA was conducted with Rec-ACC and Pred-ACC as dependent variables, respectively, where Methods and different SAs were considered within-subjects factors. Statistical results show that the main effect of the method was significant under Rec-ACC. The main effect of context was significant. Furthermore, the interaction effect between Methods and SA is also significant. Significant main effects of methods were also observed on Pred-ACC. ), SA main effect ( Furthermore, post-hoc pairwise comparisons were employed, and paired t-tests combined with multiple comparison correction were used to analyze the differences between the HPMF-SANET model and each baseline method. The asterisks in the table indicate the significance level of the HPMF-SANET model relative to the corresponding comparison methods: * p < 0.05, ** p < 0.01, *** p < 0.001. Table 4

[0093] Overall, the HPMF-SANET model achieved state-of-the-art results across all six evaluation settings. In the unimodal EEG methods, ConTraNet, GATransformer, and MDCNet performed relatively well. In the multimodal methods, CMGFNet and FGFRNet were the most competitive models, with CMGFNet performing best on SA1 and SA3 of Rec-ACC and SA1 and SA3 of Pred-ACC, while FGFRNet performed best on SA2 of Rec-ACC and SA2 of Pred-ACC. Results from the EEGBaseline and EMB aseline indicate that EEG is the dominant modality for hierarchical situational awareness modeling, while eye-tracking provides stable supplementary information. The HPMF-SANET model outperformed both unimodal versions on all tasks, further validating the complementarity of the two modalities and the effectiveness of the proposed fusion strategy. The complete HPMF-SANET model consistently outperformed all unimodal and multimodal baselines across all evaluation settings, with most improvements reaching statistical significance.

[0094] Furthermore, to verify the contribution of each key component in the HPMF-SANET model to the hierarchical situational awareness recognition and prediction performance, this embodiment conducted an ablation experiment under the same leave-one-out cross-validation setting. The ablation experiment used the same dataset partitioning, training strategy, and evaluation metrics as the full model, and constructed the following ablation model: w / o RCAM indicates the removal of the Dual-Branch Residual Cross-Attention Module, directly inputting EEG and eye-tracking features into the subsequent fusion module; w / o HDFM indicates the removal of the Hierarchy-Aware Dynamic Weighted Fusion Module, replacing it with a fixed fusion method; w / o : indicates that the hierarchical perceptual dynamic fusion structure is retained but the contributing guiding loss is removed; w / o HPTM indicates that the Hierarchical Progressive-Temporal Transfer Module is removed, and only independent classifiers at each level are used for current recognition and future prediction; w / o TTB indicates that the hierarchical progressive structure (Temporal Transfer Branch) is retained but the temporal transfer branch in future prediction is removed; w / o This indicates the loss of single-modal assistance after removing the EEG and eye-tracking branches. The results of the HPMF-SANET ablation experiments on six tasks are shown in Table 5. The asterisks in the table indicate the significance level of HPMF-SANET relative to the corresponding ablation model, * p < 0.05, ** p < 0.01, *** p < 0.001.

[0095] Table 5

[0096] Table 5 shows that the complete HPMF-SANET model achieved the best overall results across the six tasks. Repeated measures statistical analysis was conducted with Rec-ACC and Pred-ACC as dependent variables, respectively, with model type and SA level considered as within-subjects factors. The results indicate that the main effect of the model is significant on Rec-ACC (…). The main effect of SA hierarchy is significant. Significant model main effects, hierarchical main effects, and interaction effects were also observed on Pred-ACC. Further analysis using paired models... Post-hoc analysis was conducted using multiple comparison corrections. Removing the bi-branch residual cross-modal attention module or the hierarchical perception dynamic fusion module resulted in a decrease in model accuracy on most recognition and prediction tasks, indicating that cross-modal interaction enhancement and hierarchical adaptive fusion effectively utilize the complementarity of EEG and eye-tracking information. Removing the contribution-guided loss also led to a decrease in model performance, suggesting that this loss can constrain the fusion weights to maintain consistency with the actual discriminative contribution of the single modality, thereby alleviating the modality suppression problem in multimodal training. Removing the hierarchical progression-temporal transfer module resulted in a more significant decrease in SA2, SA3, and prediction tasks, indicating that the progressive relationship between the perception, understanding, and prediction levels plays a crucial role in high-level state modeling. Retaining the hierarchical progression structure but removing the temporal transfer branch resulted in a more significant decrease in future prediction tasks, indicating that the temporal transfer branch enhances the modeling ability for the state evolution in the next 5 minutes. Removing the single-modal auxiliary loss resulted in an overall performance decrease, suggesting that single-modal auxiliary supervision helps maintain the discriminative capabilities of the EEG and eye-tracking branches.

[0097] Secondly, embodiments of the present invention provide a multimodal fusion hierarchical situational awareness prediction system, which can be used to implement the multimodal fusion hierarchical situational awareness prediction method provided in the first aspect above.

[0098] Please see Figure 4 , Figure 4 The diagram shown is a structural block diagram of a hierarchical situational awareness and prediction system based on multimodal fusion. Figure 4As shown, the multimodal fusion hierarchical situational awareness prediction system of this embodiment includes: a data acquisition module, a preprocessing module, and a hierarchical situational awareness prediction module. The data acquisition module is used to acquire the EEG and eye movement signals of the train driver during the driving process. The preprocessing module is used to preprocess the EEG and eye movement signals to obtain EEG time-series data and eye movement time-series data. The hierarchical situational awareness prediction module is used to input the EEG time-series data and eye movement time-series data into a pre-trained HPMF-SANET model to obtain the current state recognition results and future state prediction results corresponding to each level of situational awareness.

[0099] Among them, the HPMF-SANET model is used to extract features, enhance cross-modal interaction, and perform hierarchical adaptive fusion on EEG time-series data and eye-tracking time-series data in sequence to obtain hierarchical perception fusion features corresponding to each level in situational awareness. The hierarchical perception fusion features are then used to simulate cognitive hierarchical progression and state temporal evolution in sequence to obtain the current state recognition result and the future state prediction result.

[0100] For details regarding the hierarchical situational awareness prediction system based on multimodal fusion and its corresponding beneficial effects, please refer to the relevant content on the hierarchical situational awareness prediction method based on multimodal fusion provided in the first aspect; it will not be elaborated upon here.

[0101] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or apparatus comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or apparatus that includes said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0102] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0103] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A multi-modal fusion hierarchical situation awareness prediction method, characterized in that, include: The brainwave signals and eye movement signals of the train driver during the driving process are collected, and the brainwave signals and eye movement signals are preprocessed to obtain brainwave time series data and eye movement time series data; The EEG time series data and the eye movement time series data are input into the pre-trained HPMF-SANET model to obtain the current state recognition results and future state prediction results corresponding to each level of situational awareness; The HPMF-SANET model is used to sequentially extract features, enhance cross-modal interaction, and perform hierarchical adaptive fusion on the EEG time-series data and the eye-tracking time-series data to obtain hierarchical perception fusion features corresponding to each level in the situational awareness. The hierarchical perception fusion features are then sequentially simulated to advance the cognitive hierarchy and evolve the state time-series to obtain the current state recognition result and the future state prediction result.

2. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 1, characterized in that, Preprocessing of the EEG signals and the eye movement signals includes: The EEG signal is subjected to bandpass filtering, artifact detection and removal, and baseline correction to obtain the processed EEG signal. The processed EEG signal and the eye movement signal are aligned on the time axis and then cropped according to a preset window to obtain the EEG timing data and the eye movement timing data.

3. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 1, characterized in that, The HPMF-SANET model includes a cascaded feature extraction module, a two-branch residual cross-modal attention module, a hierarchical perceptual dynamic fusion module, and a hierarchical progressive-temporal transfer module. The feature extraction module is used to extract features from the EEG time-series data and the eye-tracking time-series data to obtain EEG features and eye-tracking features. The dual-branch residual cross-modal attention module employs a bidirectional cross-attention mechanism to perform bidirectional cross-modal interactive enhancement of the EEG features and the eye-tracking features, thereby obtaining enhanced EEG features and enhanced eye-tracking features. The hierarchical perception dynamic fusion module is used to generate adaptive fusion weights for EEG modalities and eye movement modalities corresponding to each level, and to perform weighted fusion of the EEG enhancement features and the eye movement enhancement features based on the adaptive fusion weights to obtain the hierarchical perception fusion features corresponding to each level. The hierarchical progression-temporal migration module is used to simulate the cognitive progression relationship from the perception level to the understanding level and from the understanding level to the prediction level in the situational awareness based on the hierarchical perception fusion features, to obtain the current state representation of each level, and to output the current state recognition result of each level based on the current state representation of each level; to generate the future state representation of each level based on the current state representation of each level, and to output the future state prediction result of each level based on the future state representation of each level.

4. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 3, characterized in that, The feature extraction module includes parallel EEG feature extraction branches and eye movement feature extraction branches; The EEG feature extraction branch is used to extract the EEG temporal features of the EEG temporal data at different scales, and obtain the EEG features by splicing and pooling; the EEG feature extraction branch includes a cascaded first multi-scale parallel convolutional unit, a depthwise separable convolutional unit, and a second multi-scale parallel convolutional unit. The eye-tracking feature extraction branch is used to extract local temporal changes in the eye-tracking time-series data and aggregate short-time dynamic features to obtain the eye-tracking features; the eye-tracking feature extraction branch includes two cascaded one-dimensional convolutional units.

5. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 3, characterized in that, The dual-branch residual cross-modal attention module includes a first attention branch and a second attention branch in parallel. The first attention branch is used to use the EEG feature as a query, the eye movement feature as a key and value, calculate the first cross-attention output, and add the first cross-attention output to the EEG feature to obtain the EEG enhancement feature; The second attention branch is used to calculate the second cross-attention output by taking the eye movement feature as a query and the EEG feature as a key and value, and then adding the second cross-attention output to the eye movement feature to obtain the eye movement enhancement feature.

6. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 3, characterized in that, The hierarchical perception dynamic fusion module includes a compact vector generation unit, a weight prediction unit, and a fusion unit. The compact vector generation unit is used to perform global average pooling on the EEG enhancement features and the eye movement enhancement features respectively to obtain EEG compact vectors and eye movement compact vectors. The weight prediction unit is used to concatenate the EEG compact vector, the eye movement compact vector, and the learnable hierarchical embedding vector corresponding to the current level, and then process them through a multilayer perceptron and a softmax function to obtain the adaptive fusion weights of the EEG modality and the eye movement modality corresponding to the current level. The fusion unit is used to perform weighted fusion of the EEG compact vector and the eye-tracking compact vector according to the adaptive fusion weights of the EEG modality and eye-tracking modality corresponding to each level, so as to obtain the hierarchical perception fusion features corresponding to each level.

7. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 3, characterized in that, The hierarchical progression-time migration module includes a hierarchical progression submodule and a time migration submodule; In the hierarchical progressive submodule, the current state representation of the perception level is obtained by mapping the hierarchical perception fusion feature of the perception level; The current state representation of the understanding level is obtained by fusing the hierarchical perception fusion feature of the understanding level with the current state representation of the perception level through a first gated weighted fusion; the current state representation of the prediction level is obtained by fusing the hierarchical perception fusion feature of the prediction level with the current state representation of the understanding level through a second gated weighted fusion. In the temporal transition submodule, for each level, a state transition increment is generated based on the current state representation. The current state representation is concatenated with the score of the current state recognition result and then a temporal gating weight is generated through temporal transition gating. The future state representation is obtained by weighted fusion of the state transition increment and the current state representation according to the temporal gating weight.

8. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 1, characterized in that, The training process of the HPMF-SANET model adopts a joint loss function, which includes current identification loss, future prediction loss, single-modal auxiliary loss, and contribution guidance loss.

9. The hierarchical situational awareness prediction method based on multimodal fusion according to claim 8, characterized in that, The single-mode auxiliary loss is expressed as: ; In the formula, This represents a single-modal auxiliary supervision term representing a branch of the electroencephalogram (EEG). Single-modal auxiliary supervision term for the eye-movement branch; The contribution-guided loss is expressed as: ; in, express Norm, Indicates the number of training samples. This indicates the samples generated by the hierarchical perception dynamic fusion module. Adaptive fusion weights of EEG modalities This indicates the samples generated by the hierarchical perception dynamic fusion module. Adaptive fusion weights for eye-tracking modalities Indicates sample The contribution rate of EEG modalities Indicates sample The contribution rate of eye movement modality.

10. A hierarchical situational awareness and prediction system based on multimodal fusion, characterized in that, The hierarchical situational awareness prediction method applicable to the multimodal fusion of any one of claims 1-9 includes: The data acquisition module is used to collect the electroencephalogram (EEG) and eye movement signals of the train driver during the driving process; The preprocessing module is used to preprocess the EEG signals and the eye movement signals to obtain EEG timing data and eye movement timing data. The hierarchical situational awareness prediction module is used to input the EEG time series data and the eye movement time series data into the pre-trained HPMF-SANET model to obtain the current state recognition results and future state prediction results corresponding to each level of situational awareness. The HPMF-SANET model is used to sequentially extract features, enhance cross-modal interaction, and perform hierarchical adaptive fusion on the EEG time-series data and the eye-tracking time-series data to obtain hierarchical perception fusion features corresponding to each level in the situational awareness. The hierarchical perception fusion features are then sequentially simulated to advance the cognitive hierarchy and evolve the state time-series to obtain the current state recognition result and the future state prediction result.