Pilot situational awareness state determination method

CN122471023BActive Publication Date: 2026-09-29NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610980231.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29
Estimated Expiration
2046-07-02

AI Technical Summary

Technical Problem

[0005]本发明提供了一种飞行员情景意识状态确定方法,以解决如何在飞行员飞行过程中准确确定飞行员的情景意识的问题

Benefits of technology

[0025]本申请实施例提供的飞行员情景意识状态确定方法,分层提取目标融合特征中的三层特征。按照感知、生理、认知功能拆分特征,实现信息分层隔离,避免不同维度信息互相干扰,方便分项评估。分层输出三项独立得分,分维度量化评估,可单独定位飞行员感知不足、生理疲劳或认知下降的具体诱因,评价可溯源。三项得分融合生成综合评分,互补各维度优劣,规避单一指标片面评判缺陷,量化得到整体情景意识水平。依靠综合评分判定情景意识等级,连续分数转化为分级结果,结果直观易懂,便于机载系统开展状态预警与管控。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122471023B_ABST
    Figure CN122471023B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of situational awareness state determination, in particular to a pilot situational awareness state determination method. The method comprises: obtaining original eye movement data, original electroencephalogram data and original face image corresponding to a target pilot; performing feature extraction on the original eye movement data to obtain target comprehensive eye movement features; performing feature extraction on the original electroencephalogram data to obtain target comprehensive electroencephalogram features; performing feature extraction on the original face image to obtain target comprehensive face features; performing fusion processing on the target comprehensive eye movement features, the target comprehensive electroencephalogram features and the target comprehensive face features to generate target fusion features; and determining a target situational awareness state grade corresponding to the target pilot based on the target fusion features. Based on complete fusion feature hierarchical evaluation, the grading result takes into account the multi-dimensional state of perception, physiology and cognition, and the accuracy of the determination result is high, which can be directly used for pilot state monitoring and risk warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of situational awareness state determination technology, and more specifically to a method for determining the situational awareness state of a pilot. Background Technology

[0002] Flight safety is the core bottom line for the development of the civil aviation industry. With the technological iteration and upgrading of aircraft structural design, airborne flight control systems, and overall manufacturing processes, the proportion of flight accidents caused by aircraft hardware failures has been declining. However, flight safety accidents cannot be completely eliminated. Existing accident statistics and civil aviation safety research data show that over 70% of unsafe flight incidents are caused by pilot human error. The decline in pilot situational awareness and cognitive state during flight are key factors inducing operational errors and improper handling. Accurately identifying the pilot's situational awareness level in real time has become a key technical direction for managing human risk and reducing the flight accident rate.

[0003] Currently, the assessment methods for pilots' situational awareness mostly rely on ground-based theoretical examinations, post-flight simulation scoring, and subjective evaluations by flight instructors, which have significant technical shortcomings: First, traditional assessment methods are offline and non-real-time evaluations, which can only be completed after training or after landing. They cannot continuously monitor changes in pilots' instantaneous situational awareness in real-time during actual flight or dynamic flight in simulators, and it is difficult to capture the dynamic process of sudden decline in situational awareness under special flight situations and high-load conditions. Second, relying on subjective scoring by humans is greatly influenced by the experience and subjective preferences of the evaluators, making it difficult to unify assessment standards, resulting in low quantification of situational awareness classification results and insufficient objectivity in judgment.

[0004] In conclusion, accurately determining a pilot's situational awareness during flight has become an urgent problem to be solved. Summary of the Invention

[0005] This invention provides a method for determining a pilot's situational awareness state, in order to solve the problem of how to accurately determine a pilot's situational awareness during flight.

[0006] In a first aspect, the present invention provides a method for determining the situational awareness state of a pilot, the method comprising: acquiring raw eye-tracking data, raw electroencephalogram (EEG) data, and raw facial image corresponding to the target pilot; extracting features from the raw eye-tracking data to obtain target comprehensive eye-tracking features; extracting features from the raw EEG data to obtain target comprehensive EEG features; extracting features from the raw facial image to obtain target comprehensive facial features; fusing the target comprehensive eye-tracking features, target comprehensive EEG features, and target comprehensive facial features to generate target fusion features; and determining the target situational awareness state level corresponding to the target pilot based on the target fusion features.

[0007] The method for determining the situational awareness state of a pilot provided in this application acquires the original eye-tracking data, original electroencephalogram (EEG) data, and original facial images corresponding to the target pilot. Multi-source raw data is collected simultaneously, comprehensively covering the pilot's physiological and behavioral information from three dimensions: visual, neurological, and facial, avoiding information gaps from single data sources and providing full-dimensional raw materials for subsequent feature extraction. Feature extraction is performed on the original eye-tracking data to obtain the target's comprehensive eye-tracking features. Invalid noise data in the eye-tracking is removed, and effective features representing environmental perception, such as fixation and saccades, are refined, compressing the data volume and accurately quantifying the pilot's ability to capture external information. Feature extraction is performed on the original EEG data to obtain the target's comprehensive EEG features. EMG artifacts and power frequency interference are filtered out, and features related to attention and cognitive load are extracted to accurately reflect deep cognitive activities of the brain, achieving quantification of internal cognitive state. Feature extraction is performed on the original facial images to obtain the target's comprehensive facial features. Based on key points, EAR, MAR, and FAR indicators are converted to quantify fatigue and facial state, intuitively representing the physiological support state and compensating for the shortcomings of EEG and eye-tracking in intuitively reflecting physical fatigue. The target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features are fused to generate a target fusion feature. Adaptive weighting is used to rationally allocate the contribution of each modality, complementing the advantages and disadvantages of the three feature types and eliminating the one-sidedness of single-modal evaluation. The fusion feature incorporates all information from external perception, physiology, and deep cognition, resulting in stronger representational capabilities. Based on the target fusion feature, the target pilot's corresponding target situational awareness level is determined. A hierarchical assessment based on the complete fusion feature is conducted, with the grading results taking into account multiple dimensions of perception, physiology, and cognition, resulting in high accuracy and direct applicability for pilot status monitoring and risk warning.

[0008] In one optional implementation, feature extraction is performed on the raw eye-tracking data to obtain comprehensive target eye-tracking features, including: preprocessing the raw eye-tracking data to obtain target eye-tracking data; acquiring multiple regions of interest (ROIs) based on the flight scenario; the ROIs are at least one of the following: main flight area, instrument area, left-side environment area, and right-side environment area; extracting features from the target eye-tracking data to obtain spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye-tracking data; the spatial allocation dimension features include the proportion of access time corresponding to each ROI and the number of times each ROI is accessed; the temporal fixation dimension features include the total fixation time, average fixation duration, and number of fixations corresponding to each ROI; the physiological load dimension features include the average pupil diameter, maximum pupil diameter, and minimum pupil diameter; the gaze deviation dimension features include the average horizontal distance, average vertical distance, and average absolute distance; and obtaining comprehensive target eye-tracking features based on the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye-tracking data.

[0009] The pilot situational awareness state determination method provided in this application preprocesses raw eye-tracking data to obtain target eye-tracking data. Abnormal noise and invalid data caused by blinking and equipment vibration are removed, the data format is standardized, and interference from abnormal data is reduced to ensure the reliability of subsequent feature extraction data. Regions of interest (ROIs) for various flight scenarios, such as the main flight area and instrument area, are delineated. Adhering to the logic of real flight observation, the method focuses on the pilot's key observation positions, avoids statistical analysis of invalid areas across the entire screen, and ensures a strong correlation between eye-tracking features and flight missions. Corresponding subdivided eye-tracking features are extracted according to four dimensions. Indicators are broken down from multiple angles, including region access patterns, fixation timing, pupil physiology, and line-of-sight deviation, to comprehensively quantify visual observation behavior and cover various representational elements of environmental perception. The four-dimensional features are summarized to generate comprehensive target eye-tracking features. Integrating spatial, temporal, physiological, and offset information, a complete eye-tracking representation vector is formed, reflecting the pilot's situational awareness capabilities and providing a complete eye-tracking data source for multimodal fusion.

[0010] In one optional implementation, based on the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data, a comprehensive target eye movement feature is obtained, including: inputting the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data into a preset gradient boosting decision tree model, and outputting the importance score corresponding to each eye movement sub-feature included in the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features; and determining the comprehensive target eye movement feature from each eye movement sub-feature based on the importance score corresponding to each eye movement sub-feature.

[0011] The pilot situational awareness state determination method provided in this application inputs the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data into a preset gradient boosting decision tree model, and outputs the importance scores corresponding to each eye movement sub-feature included in the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features. The model automatically quantifies the contribution of each subdivided eye movement indicator to state identification, avoiding the subjectivity of manual feature selection based on experience, and accurately distinguishing between effective and redundant features. Based on the importance scores corresponding to each eye movement sub-feature, the comprehensive target eye movement features are determined from each eye movement sub-feature. Low-importance redundant sub-features are eliminated, simplifying feature dimensions and reducing computational overhead, while retaining high-contribution key indicators, improving the efficiency and accuracy of subsequent multimodal fusion and state assessment.

[0012] In one optional implementation, feature extraction is performed on the raw EEG data to obtain target comprehensive EEG features, including: preprocessing the raw EEG data to obtain target EEG data; calculating the power spectral density and energy spectral density corresponding to the target EEG data; calculating the global field power corresponding to the target EEG data; determining the target EEG state features corresponding to the target EEG data based on the global field power; extracting source features from the target EEG data to obtain target EEG source features corresponding to the target EEG data; and fusing the power spectral density, energy spectral density, target EEG state features, and target EEG source features to obtain target comprehensive EEG features.

[0013] The method for determining the situational consciousness state of pilots provided in this application preprocesses the raw EEG data to obtain the target EEG data. It removes power frequency interference and electrooculography / electromyography artifacts, optimizes data quality, and avoids noise interference in subsequent spectrum and source tracing calculations. Power spectral density and energy spectral density are calculated. This allows for the breakdown of energy distribution across frequency bands, quantifying the activity of different EEG waves, and reflecting changes in brain excitation and load in the frequency domain. Global field power is calculated. This quantifies the overall brain activation intensity at a holistic level, intuitively reflecting the overall level of mental exertion. EEG state features are generated based on the global field power. Global mental exertion indicators are refined, simplifying massive spectral data and rapidly representing the overall cognitive state. Then, EEG source tracing features are extracted. The location of brain region discharges is identified, associated with corresponding cognitive function partitions, and details of deep brain region activity are supplemented. Multiple types of EEG features are fused to obtain comprehensive EEG features. This approach considers frequency domain energy, global brain power, and brain region spatial information, representing internal cognition from multiple angles with stronger feature completeness.

[0014] In one optional implementation, the target EEG state features corresponding to the target EEG data are determined based on global field power, including: selecting target brain topographic maps corresponding to the target EEG data based on global field power; performing iterative clustering on the target brain topographic maps to obtain multiple initial clustering results; calculating the global explained variance and cross-validation criterion corresponding to each initial clustering result; updating each initial clustering result based on the principle of maximizing global explained variance and minimizing cross-validation criterion until clustering converges, determining the optimal number of clusters and the target clustering results corresponding to the optimal number of clusters; determining the target state category corresponding to each target clustering result based on the similarity between each target clustering result and the preset state category; extracting the clustering time series features corresponding to each target clustering result according to the target state category corresponding to each target clustering result; and fusing the clustering time series features to obtain the target EEG state features.

[0015] The pilot situational awareness state determination method provided in this application uses global field power to screen target brain topographic maps. Invalid artifact maps are removed, retaining only valid brain topographic maps that accurately reflect brain activation changes, reducing interference from invalid samples. Iterative clustering of the brain topographic maps generates multiple initial clustering results, automatically grouping them according to the spatial distribution patterns of EEG, eliminating the need for manual delineation of classification boundaries and achieving unsupervised grouping of EEG states. Global explained variance and cross-validation index are calculated for each clustering result. Quantitative indicators are used to measure the quality of clustering, overcoming the drawbacks of subjective human evaluation. Iterative optimization is performed until convergence, using the maximum variance and minimum cross-validation index to determine the optimal number of clusters and clustering results. This approach balances data information retention and model generalization ability, avoiding over- or under-segmentation of clusters and improving grouping rationality. The similarity between the clustering results and preset state categories is compared to match the corresponding target state category. Unlabeled clustering results are mapped to actual cognitive states, achieving automatic correspondence between EEG atlases and cognitive types. Based on classification results, cluster temporal features are extracted to record the switching patterns of different brain states over time, supplementing temporal dimension information and reflecting the dynamic changes in cognition. The target EEG state features are obtained by fusing the temporal features of each cluster. By integrating spatial distribution and temporal evolution information, the system comprehensively represents the dynamic changes in cognitive load in the brain, enhancing feature representation capabilities.

[0016] In one optional implementation, source feature extraction is performed on the target EEG data to obtain the target EEG source features corresponding to the target EEG data, including: using a standard human brain template as a benchmark, the cerebral cortex is finely divided and evenly divided into multiple independent voxel units; and a standardized digital brain model is established based on the correspondence between each independent voxel unit and the fixed spatial coordinates of the brain and the functional areas of the cortex. The target EEG data is input into a standardized digital brain model. Based on the sLORETA low-resolution electromagnetic tomography algorithm, the spatial distribution of whole-brain current density corresponding to all independent voxel units is calculated and output point by point. From the spatial distribution of whole-brain current density, the spatial distribution of regional current density corresponding to the Brodman partition is selected. Feature extraction is performed on the spatial distribution of regional current density to obtain the average current density, maximum activation intensity, and spatial activation range of the brain region. Based on the average current density, maximum activation intensity, and spatial activation range of the brain region, the source features of the target EEG are obtained.

[0017] The method for determining the situational awareness state of pilots provided in this application relies on a standard human brain template to obtain multiple sets of independent voxel units. It unifies the brain spatial division scale, refines brain region computational units, and provides a regular computational grid for subsequent inverse problem solving. A standardized digital brain model is constructed by combining voxel coordinates and cortical functional areas. Spatial location and physiological functional partitions are bound together, unifying the individual brain data mapping benchmark and eliminating computational biases caused by individual brain morphological differences. Target EEG data is input into the model, and sLORETA is used to solve the whole-brain voxel current density. Intracranial neural source activity is inverted from scalp EEG, overcoming the limitation of raw EEG only providing surface signals and obtaining true intracranial activation information. The current density data corresponding to the Broadman partition is screened. Focusing on cognitively relevant functional brain regions, irrelevant voxel data is eliminated, and effective activation areas related to situational awareness are anchored. Three types of indicators are extracted: average current density, peak activation, and activation range. Brain region activity is quantified from multiple perspectives, including activation strength, activation extremes, and diffusion range, enriching the dimensions of source traceability features. The three types of indicators are summarized to generate target EEG source traceability features. To realize the quantitative application of brain-derived spatial information, accurately represent deep cognitive activities, and improve the intrinsic physiological basis of comprehensive EEG characteristics.

[0018] In one optional implementation, feature extraction is performed on the original face image to obtain the target comprehensive face features, including: image preprocessing of the original face image to obtain the target face image; recognition of the target face image to determine the effective region of interest; recognition of the effective region of interest to determine multiple standard facial key points; selection of multiple core feature points from the multiple standard facial key points; calculation of core features of eye aspect ratio, mouth aspect ratio, and face contour aspect ratio based on each core feature point; and fusion processing of the core features of eye aspect ratio, mouth aspect ratio, and face contour aspect ratio to obtain the target comprehensive face features.

[0019] The pilot situational awareness state determination method provided in this application preprocesses the original face image to obtain the target face image, removing interference from lighting, noise, and image distortion to optimize image quality and ensure the accuracy of subsequent face localization and key point extraction. It identifies and delineates the effective region of interest (ROI) of the face. Invalid background areas are cropped and removed to reduce the computational range and decrease interference from irrelevant pixels on feature extraction. Multiple standard facial key points are detected and output. The positions of the eyes, mouth, and cheeks are located based on these points, providing coordinate basis for subsequent quantitative geometric indicators. Core feature points are selected, redundant key points are discarded, and key points related to fatigue and mental state are retained to simplify computation. Three types of proportional features (EAR, MAR, FAR) are solved using these points. The facial morphology is converted into quantitative values, intuitively reflecting physiological changes such as drowsiness and facial tension. The three types of proportional values ​​are fused to obtain the comprehensive facial features of the target. Multi-dimensional summarization of facial physiological representations fully reflects the pilot's real-time physiological support state and improves the multimodal data source.

[0020] In one optional implementation, the target integrated eye-tracking features, target integrated EEG features, and target integrated facial features are fused to generate target fusion features. This includes: performing temporal alignment processing on the target integrated eye-tracking features, target integrated EEG features, and target integrated facial features to obtain temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features; inputting the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features into a preset weight determination model, and outputting target weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features respectively; performing weighted fusion on the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features based on the target weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features respectively, to obtain an intermodal correlation matrix; and obtaining the target fusion features based on the intermodal correlation matrix.

[0021] The pilot situational awareness state determination method provided in this application performs temporal alignment processing on the target's integrated eye movement features, target integrated EEG features, and target integrated facial features to obtain temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned facial features. This unifies the timestamps of the three types of data, eliminates temporal misalignment caused by different acquisition frame rates, and ensures that eye movement, EEG, and facial data are matched one-to-one at the same time. The temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned facial features are input into a preset weight determination model, which outputs target weight information corresponding to each of the temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned facial features. Modal weights are adaptively allocated based on the recognition effect, unlike fixed weights, to adapt to the dynamic changes in the contribution of each modality under different flight conditions. Based on the target weight information corresponding to the temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned facial features, the temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned facial features are weighted and fused to obtain an inter-modal correlation matrix. Effective modal information is amplified based on contribution proportion, and redundant interference is suppressed. Modal correlations and temporal correlations are retained in matrix form, resulting in a well-structured data structure. Target fusion features are obtained based on the inter-modal correlation matrix. Effective matrix information is refined, and redundant dimensions are eliminated to form an integrated feature that takes into account perception, physiology, and cognition, facilitating subsequent hierarchical extraction and scoring calculation.

[0022] In one optional implementation, temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features are input into a preset weight determination model, and the model outputs target weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features, respectively. This includes: inputting the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features into the preset weight determination model; the preset weight determination model generates initial weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features, respectively; based on the initial weight information, the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features are weighted and fused, and forward inference is completed, outputting the initial contextual awareness level identification corresponding to the initial weight information. Results: The global loss value between the initial situational awareness level identification result and the real cognitive label in the model determined by the preset weights is calculated. Based on the global loss value, the initial weight information is updated to obtain the updated weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features, respectively. Based on the updated weight information, the initial weight information is updated to obtain the modality weight fluctuation amount corresponding to each modality in the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features. Based on the modality weight fluctuation amount, the average fluctuation amplitude is calculated. Based on the global loss value and the average fluctuation amplitude, the updated weight information is updated to obtain the target weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features, respectively.

[0023] The pilot situational awareness state determination method provided in this application involves feeding three types of aligned features into a weight determination model. A unified input data source is used, and the model enables data-driven weight initialization, eliminating the subjectivity of manually setting weights. The model outputs initial weights for each modality. A set of baseline weights is quickly generated, providing a basic matching scheme for initial feature fusion. Initial weight weighted fusion + forward inference outputs initial identification results. This determines the actual effect of the weights, obtaining prediction results that can be used for error calculation and quantifying the current weight identification performance. The prediction is compared with the true value to calculate the global loss, quantifying the overall identification bias from a full-sample perspective, providing a quantitative monitoring indicator for weight iterative optimization. Updated weights are obtained through back-correction based on the global loss. The proportion of each modality is adaptively adjusted to reduce identification error, optimizing the rationality of modality contribution allocation. The fluctuation of each modality weight is calculated by comparing the old and new weights. The change amplitude of each modality weight in a single iteration is quantified to distinguish the optimization sensitivity of each modality. The average fluctuation amplitude is calculated from the single-modality fluctuation. This macroscopically characterizes the overall change of the entire weight set, used to evaluate the convergence and stability of the weights. The target weights are obtained through a secondary optimization combining loss and volatility. This approach balances identification accuracy and weight stability, avoiding weight oscillations caused by relying solely on loss updates, and ultimately yields the optimal steady-state weights.

[0024] In one optional implementation, the target situational awareness level of the target pilot is determined based on the target fusion features, including: extracting environmental perception layer features, physiological state support layer features, and deep cognitive core layer features from the target fusion features; outputting the real-time situational awareness score, physiological state score, and core cognitive score of the target pilot based on the environmental perception layer features, physiological state support layer features, and deep cognitive core layer features; fusing the situational awareness score, physiological state score, and core cognitive score to obtain the situational awareness level score of the target pilot; and determining the target situational awareness level of the target pilot based on the situational awareness level score.

[0025] The pilot situational awareness state determination method provided in this application extracts three layers of features from the target fusion features. Features are split according to perception, physiological, and cognitive functions, achieving hierarchical isolation of information and avoiding interference between different dimensions, facilitating itemized evaluation. Three independent scores are output layer by layer, providing dimensional quantitative evaluation, allowing for the identification of specific causes of insufficient pilot perception, physiological fatigue, or cognitive decline, making the evaluation traceable. The three scores are fused to generate a comprehensive score, complementing the strengths and weaknesses of each dimension, avoiding the shortcomings of single-indicator one-sided evaluation, and quantifying the overall situational awareness level. The situational awareness level is determined based on the comprehensive score, with continuous scores converted into graded results, making the results intuitive and easy to understand, facilitating status warning and control by airborne systems. Attached Figure Description

[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0027] Figure 1 This is a schematic flowchart of a first method for determining a pilot's situational awareness state according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a second flowchart of a method for determining a pilot's situational awareness state according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the principle of microstate analysis according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the core aspect ratio feature EAR of the eye according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the mouth aspect ratio core feature (MAR) according to an embodiment of the present invention; Figure 6This is a schematic diagram of the facial contour aspect ratio (FAR) core feature according to an embodiment of the present invention; Figure 7 This is a structural block diagram of a pilot situational awareness state determination device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0030] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0031] According to an embodiment of the present invention, a method for determining a pilot's situational awareness state is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0032] This embodiment provides a method for determining a pilot's situational awareness state, which can be used in electronic devices. The electronic devices can be control devices in a target aircraft or flight simulator, or terminal devices or server devices. This application embodiment does not specifically limit the electronic devices. Figure 1 This is a flowchart of a method for determining a pilot's situational awareness state according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain the raw eye movement data, raw electroencephalogram data, and raw facial image corresponding to the target pilot.

[0033] Specifically, eye-tracking, electroencephalography (EEG), and high-definition image acquisition devices are installed in the flight cockpit. The eye-tracking device is aimed at the pilot's eye area to ensure complete coverage of the eye movement range. The EEG device is worn on the pilot's head according to standard lead configuration to check electrode contact impedance and signal continuity. The image acquisition device is positioned in an unobstructed viewing angle and aimed at the pilot's face. The electronic equipment can sequentially activate the three types of acquisition devices to perform hardware self-tests, signal link self-tests, and storage module self-tests, troubleshooting equipment malfunctions, loose wiring, signal disconnections, and other issues to ensure all devices are functioning properly.

[0034] Then, the electronic devices can determine a consistent sampling frame rate, sampling resolution, and sampling range. A global clock synchronization calibration is performed on the entire acquisition system, using a unified timestamp as the data timing reference to eliminate clock deviations between different devices and ensure a one-to-one correspondence in the timing of eye movement, EEG, and facial image data from the source. Simultaneously, to address complex operating conditions such as changes in cockpit lighting, slight head movements, and fuselage vibrations, the devices' anti-interference, image stabilization, and dynamic exposure adaptation functions are activated to reduce the interference of environmental factors on the raw data.

[0035] Once the flight mission enters the formal monitoring phase, a synchronous data acquisition initiation command is issued, and all three types of equipment start operating simultaneously upon receiving the command. The eye-tracking acquisition device captures raw, time-series eye-tracking data in real time, including the pilot's eye position, fixation point, and eye movement trajectory. The electroencephalogram (EEG) acquisition device continuously acquires raw potential signals from various leads in the cerebral cortex, outputting time-series EEG waveform data. The image acquisition device continuously captures images at a preset frame rate, outputting continuous frames of raw facial images of the pilot. This uninterrupted, continuous data acquisition follows the entire flight process.

[0036] During the data acquisition process, the raw eye-tracking data, raw EEG data, and raw facial images corresponding to each time-series node are all appended with a unified global timestamp and synchronously written to the local cache. They are stored sequentially according to the time sequence, without disrupting the data order or losing any single-frame samples. At the same time, independent data partitions are used to classify and store the three types of raw data to avoid data mixing and file disorder.

[0037] Step S102: Extract features from the original eye-tracking data to obtain the target comprehensive eye-tracking features.

[0038] Specifically, the electronic device can first preprocess the raw eye movement data, and then extract features from the preprocessed eye movement data to obtain the target comprehensive eye movement features.

[0039] This step will be explained in detail below.

[0040] Step S103: Extract features from the raw EEG data to obtain the target comprehensive EEG features.

[0041] Specifically, the electronic device can first preprocess the raw EEG data, and then extract features from the preprocessed EEG data to obtain the target comprehensive EEG features.

[0042] This step will be explained in detail below.

[0043] Step S104: Extract features from the original face image to obtain the target comprehensive face features.

[0044] Specifically, the electronic device can first preprocess the original face image, and then extract features from the preprocessed face image to obtain the target comprehensive face features.

[0045] This step will be explained in detail below.

[0046] Step S105: The target's comprehensive eye movement features, comprehensive EEG features, and comprehensive facial features are fused to generate the target fusion features.

[0047] Specifically, electronic devices can stitch together the target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features to generate target fusion features.

[0048] This step will be explained in detail below.

[0049] Step S106: Based on the target fusion features, determine the target situational awareness level corresponding to the target pilot.

[0050] Specifically, the electronic device can input the target fusion features into a preset situational awareness recognition model and output the target situational awareness state level corresponding to the target pilot.

[0051] This step will be explained in detail below.

[0052] This embodiment provides a method for determining the situational awareness state of pilots, acquiring raw eye-tracking data, raw electroencephalogram (EEG) data, and raw facial images corresponding to the target pilot. Simultaneous acquisition of multi-source raw data comprehensively covers the pilot's physiological and behavioral information from three dimensions: visual, neurological, and facial, avoiding information gaps from single data sources and providing full-dimensional raw material for subsequent feature extraction. Feature extraction is performed on the raw eye-tracking data to obtain the target's comprehensive eye-tracking features. Invalid noise data in eye-tracking is removed, and effective features representing environmental perception, such as fixation and saccades, are refined, compressing the data volume and accurately quantifying the pilot's ability to capture external information. Feature extraction is performed on the raw EEG data to obtain the target's comprehensive EEG features. EMG artifacts and power frequency interference are filtered out, and features related to attention and cognitive load are extracted to accurately reflect deep cognitive activities of the brain, achieving quantification of internal cognitive state. Feature extraction is performed on the raw facial images to obtain the target's comprehensive facial features. Based on key points, EAR, MAR, and FAR indicators are converted to quantify fatigue and facial state, intuitively representing physiological support state and compensating for the shortcomings of EEG and eye-tracking in intuitively reflecting physical fatigue. The target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features are fused to generate a target fusion feature. Adaptive weighting is used to rationally allocate the contribution of each modality, complementing the advantages and disadvantages of the three feature types and eliminating the one-sidedness of single-modal evaluation. The fusion feature incorporates external perception, physiological, and deep cognitive information, resulting in stronger representational capabilities. Based on the target fusion feature, the target pilot's corresponding target situational awareness level is determined. A hierarchical assessment based on the comprehensive fusion feature is conducted, with the grading results taking into account multiple dimensions of perception, physiology, and cognition, resulting in high accuracy and direct applicability for pilot status monitoring and risk warning.

[0053] This embodiment provides a method for determining a pilot's situational awareness state, which can be used in the aforementioned electronic equipment. Figure 2 This is a flowchart of a method for determining a pilot's situational awareness state according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the raw eye movement data, raw electroencephalogram data, and raw facial image corresponding to the target pilot.

[0054] For details on this step, please refer to step S101 above, which will not be repeated here.

[0055] Step S202: Extract features from the original eye-tracking data to obtain the target comprehensive eye-tracking features.

[0056] Specifically, step S202 above may include the following steps: Step S2021: Perform data preprocessing on the raw eye-tracking data to obtain the target eye-tracking data.

[0057] Specifically, affected by characteristics of eye movement and device sampling, the original eye movement data is prone to short-term data loss. The entire segment of time-series eye movement data is traversed, and data missing segments are detected point by point. For local data vacancy regions with a duration not exceeding 60 ms, the electronic device can use the linear interpolation method to complete missing values, ensuring the eye movement time-series sequence is continuous and complete. For long-term data loss segments exceeding 60 ms, they are marked as invalid data segments and eliminated separately, so as to avoid interpolation distortion affecting subsequent analysis.

[0058] Next, the electronic device can use a sliding mean filtering algorithm to smooth the eye movement time-series data completed by linear interpolation, so as to suppress high-frequency noise introduced during the acquisition process. The specific implementation rules are as follows: The electronic device presets a filtering half-window width k, and determines that the total length of the sliding window is 2k+1; the window size is selected according to actual requirements. The larger the window, the stronger the smoothing effect, but the signal delay will be increased synchronously. The smoothing effect and signal real-time performance are balanced as needed.

[0059] The electronic device sets the original eye movement time-series data as x i (i=1,2,…,N), where N is the total number of data points. When a data point satisfies k<i<N-k, a total of 2k+1 adjacent data points on the left and right of the current point are taken to calculate the arithmetic mean, which is used as the filtered output value y i , and the calculation formula is: . When a data point satisfies i≤k (the left boundary of the sequence), the sliding window is dynamically reduced, and all data from the first point to the current point are taken to calculate the mean value: . When a data point satisfies i>N-k (the right boundary of the sequence), the sliding window is dynamically reduced, and all data from the current point to the last point are taken to calculate the mean value: . The electronic device traverses all data points to complete filtering, and obtains denoised and smoothed basic eye movement data.

[0060] Then, based on the filtered data, the electronic device can use the I-VT algorithm to calculate the angular velocity of eye movement, divide fixation state and saccade state in combination with a preset threshold, and extract effective fixation points. The electronic device can determine the distance D (in mm) between the pilot and the display screen as a fixed parameter for angular velocity calculation. According to adjacent eye movement coordinates, time interval and observation distance, the eye movement angular velocity ω at each moment is calculated, and the angular velocity of the whole sequence is solved. A preset eye movement speed threshold ω0 is set, and a binary classification rule is implemented: if ω≤ω0, the current state is determined as fixation, and the corresponding data point is retained as an initial fixation point; if ω>ω0, the current state is determined as saccade, and this segment of data is marked.

[0061] The electronic device can traverse all identified initial gaze points, performing correlation verification and merging processing on adjacent gaze points one by one. It calculates the time interval and spatial angle difference between two adjacent gaze points. If the time interval between adjacent gaze points is less than 75ms and the spatial angle difference is less than 0.5°, they are determined to be discrete points generated by the same continuous gaze behavior, and a merging operation is performed to integrate them into a single valid gaze point. After completing the traversal and merging of the entire sequence of gaze points, redundant discrete points are removed, and the temporal and spatial distribution of gaze points is normalized.

[0062] Finally, after a full-process processing including missing data compensation, moving mean filtering, motion state segmentation, and gaze point merging, standardized data with low noise, continuous temporal sequence, and accurate gaze region segmentation is obtained, which is the target eye movement data. This data can be used for subsequent eye movement feature extraction and analysis of pilot attention and situational awareness capabilities.

[0063] Step S2022: Obtain multiple regions of interest for targets based on the flight scenario settings.

[0064] The target region of interest is at least one of the following: the main flight area, the instrument area, the left-side environmental area, and the right-side environmental area.

[0065] Specifically, electronic devices can pre-divide the cockpit visual area based on the conventional flight pilot's field of vision layout, defining four types of standard target regions of interest: the main flight area, the instrument area, the left-side environment area, and the right-side environment area. The pixel coordinate boundaries, spatial range, and visual coverage of each area are clearly defined, forming a region coordinate mapping table, which serves as the benchmark for subsequent eye-tracking point determination.

[0066] The electronic device can select at least one region of interest from the four categories mentioned above as the target region of interest for this analysis, based on the current flight mission type and monitoring requirements. A single region can be selected, or multiple regions can be selected in combination. Once selected, the region boundary parameters are locked to ensure a consistent and unbiased range for subsequent eye-tracking data statistics. The selected region of interest number, boundary coordinates, and region type are permanently stored, establishing a correspondence between region identifiers and coordinates, providing a basis for judgment in the next step of eye-tracking point matching and feature statistics.

[0067] Step S2023: Extract features from the target eye movement data to obtain the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data.

[0068] The spatial allocation dimension features include the proportion of access time corresponding to each target's region of interest and the number of times each target's region of interest is accessed. The temporal fixation dimension features include the total fixation time corresponding to each target's region of interest, the average fixation duration, and the number of fixations. The physiological load dimension features include the average pupil diameter, the maximum pupil diameter, and the minimum pupil diameter. The gaze deviation dimension features include the average horizontal distance of the gaze, the average vertical distance, and the average absolute distance.

[0069] Specifically, considering the spatial allocation dimension characteristics, the electronic device can traverse the target eye-tracking data frame by frame. Based on the gaze coordinates and preset region boundaries, it determines the target region of interest (ROI) to which the gaze point belongs at each moment, and classifies and labels all valid eye-tracking points into regions. The electronic device can then accumulate the total time the pilot's gaze remains within each ROI according to region classification. Then, using the total effective eye-tracking time as the base, it calculates the ratio of the access time to each individual ROI to the total time, obtaining the access time percentage for each ROI. The electronic device can recognize the switching behavior of the gaze between different regions. Each time the pilot enters the current ROI from another region, it is counted as a valid access. The cumulative access frequency is calculated for each region, yielding the access count for each ROI.

[0070] For the temporal gaze dimension, electronic devices can combine the valid gaze points identified by the I-VT algorithm described earlier to match each gaze behavior to the corresponding target region of interest (ROI), distinguishing independent gaze segments within different regions. The duration of all gaze segments is accumulated across regions to obtain the total gaze duration for each ROI. The electronic device can count independent, non-continuous gaze behaviors within a single ROI to obtain the gaze count for each ROI. The average gaze duration for each ROI can be calculated by dividing the total gaze duration of a single region by the number of gazes in that region.

[0071] Based on the physiological load dimension, the electronic device can extract the raw pupil diameter values ​​from the target eye-tracking data over the entire time series, removing abnormal extreme values ​​and invalid data caused by device jitter, occlusion, and blinking, retaining the valid sample set. The electronic device can calculate the arithmetic mean of all valid pupil diameter values ​​to obtain the average pupil diameter. The electronic device then iterates through the valid pupil diameter sample set, retrieving the maximum and minimum values ​​in the sequence to obtain the maximum and minimum pupil diameters, respectively.

[0072] To address the dimensional characteristics of line-of-sight offset, the electronic equipment, in conjunction with the flight scenario, defines a standard cockpit-facing reference point and determines its horizontal and vertical coordinates as references for offset calculation. For each valid line-of-sight coordinate, the horizontal, vertical, and spatial absolute offset distances relative to the reference point are calculated. The electronic equipment can calculate the arithmetic mean of the horizontal offset distances across the entire sequence to obtain the average horizontal distance; calculate the arithmetic mean of the vertical offset distances across the entire sequence to obtain the average vertical distance; and calculate the arithmetic mean of the spatial absolute offset distances across the entire sequence to obtain the average absolute distance.

[0073] Step S2024: Based on the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data, the comprehensive eye movement features of the target are obtained.

[0074] Specifically, step S2024 above may include the following steps: Step a1: Input the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data into the preset gradient boosting decision tree model, and output the importance score corresponding to each eye movement sub-feature included in the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features.

[0075] Specifically, electronic devices can input all eye-movement sub-features under spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features into a pre-trained Gradient Boosting Decision Tree (GBDT) model. This GBDT model automatically completes feature contribution evaluation and outputs a quantitative importance score for each eye-movement sub-feature, achieving accurate identification of features with high discriminative power and high contribution.

[0076] Electronic devices can structure all eye movement sub-features corresponding to the target eye movement data to form a standardized feature vector. The input eye movement sub-features include: spatial allocation dimension: the percentage of time spent visiting each region of interest (ROI) and the number of visits to each ROI; temporal fixation dimension: the total fixation time for each ROI, the average fixation duration, and the number of fixations; physiological load dimension: the average pupil diameter, the maximum pupil diameter, and the minimum pupil diameter; and gaze deviation dimension: the average horizontal distance, the average vertical distance, and the average absolute distance of the gaze.

[0077] Electronic devices can normalize and impute missing values ​​for all eye-tracking sub-features, eliminating dimensional differences and outlier interference to form a standardized feature matrix that meets the model input format requirements. Then, the standardized eye-tracking feature matrix is ​​input into a gradient boosting decision tree model pre-trained based on a pilot eye-tracking dataset. The gradient boosting decision tree model consists of multiple decision trees serially integrated. Each tree uses feature splitting gain as the core criterion for node partitioning, evaluating the effectiveness of each eye-tracking sub-feature one by one. It traverses all eye-tracking sub-features under the current tree node, attempting to split the node's sample set into left and right branches based on a single feature. The information gain or Gini coefficient change corresponding to each feature split is calculated to quantify the improvement in node purity. A higher improvement in node purity after splitting indicates a stronger ability of that feature to distinguish the pilot's situational awareness level. The gradient boosting decision tree model also statistically analyzes the reduction in the overall prediction error before and after splitting; a more significant decrease in error indicates a greater contribution of that feature to reducing classification bias. Select the feature with the largest gain as the optimal splitting feature for the current node, complete the single node split, and repeat this process until the single decision tree grows to the preset depth or meets the stopping condition.

[0078] The gradient boosting decision tree model employs a residual iteration mechanism. Each subsequent tree fits the residual between the true label and the predicted value of a sample, and the next tree continues to fit new residuals. Multiple decision trees are processed sequentially. The first decision tree performs preliminary sample classification based on the original feature matrix and calculates the global prediction residual. Each subsequent decision tree uses the residual from the previous round as the fitting target, traversing all eye-tracking sub-features again, performing node splitting and gain calculation. During the splitting process of each tree, the frequency of each eye-tracking sub-feature being selected as the optimal splitting feature, the corresponding total splitting gain, and the cumulative error reduction are continuously recorded. After all decision trees have iterated and grown, all quantitative indicators from single-tree and multi-tree operations are summarized to form a comprehensive contribution statistic for each feature.

[0079] Electronic devices can utilize three core quantitative indicators accumulated throughout the entire tree computation process: total feature splitting gain, total improvement in node purity, and cumulative error reduction, to uniformly convert these into standardized feature importance scores. For each eye-tracking sub-feature, the splitting gain and error reduction generated across all decision trees and all tree nodes are summed. Using the sum of the cumulative indicators for all eye-tracking sub-features as the base, normalization is performed, mapping the indicator values ​​to an importance score in the range of 0-1 or as a percentage. The greater the cumulative feature gain and error reduction, the higher the final importance score, indicating a stronger ability of that eye-tracking sub-feature to distinguish the level of situational awareness.

[0080] Following the order of the original feature matrix, the quantitative importance score corresponding to each eye-tracking sub-feature is output one by one, generating a feature-score correspondence list. This result objectively reflects the actual contribution of each sub-feature to the contextual awareness recognition task, serving as the core basis for subsequent selection of target comprehensive eye-tracking features.

[0081] Step a2: Based on the importance score corresponding to each eye movement sub-feature, determine the target comprehensive eye movement feature from each eye movement sub-feature.

[0082] Specifically, electronic devices can preset feature importance score screening thresholds based on the needs of flight scenario cognitive monitoring and model accuracy requirements. These thresholds can be fixed score thresholds, percentage thresholds (e.g., retaining the top 80% of important features), or cumulative contribution thresholds (e.g., retaining features with a cumulative importance ratio of 95%), ensuring that the screening rules are objective, consistent, and reproducible.

[0083] Electronic devices can iterate through all eye-tracking sub-features and their importance scores. If the importance score of an eye-tracking sub-feature is greater than or equal to a preset feature importance score screening threshold, it is determined to be a high-contribution core feature and retained. If the importance score of an eye-tracking sub-feature is less than the preset feature importance score screening threshold, it is determined to be a low-contribution redundant feature and discarded. Simultaneously, the retained features are sorted from high to low importance scores to form a core feature sequence with clear priorities, which is then integrated to determine the target comprehensive eye-tracking feature set. This target comprehensive eye-tracking feature set, while retaining the original eye-tracking behavior representation capabilities, achieves dimensionality reduction, noise removal, and information enhancement. This reduces the computational complexity of subsequent models and significantly improves the accuracy and generalization ability of eye-tracking features in representing contextual awareness levels.

[0084] Step S203: Extract features from the raw EEG data to obtain the target comprehensive EEG features.

[0085] Specifically, step S203 above may include the following steps: Step S2031: Perform data preprocessing on the raw EEG data to obtain the target EEG data.

[0086] Specifically, the raw EEG signal acquisition output format is PROJ, which cannot be directly read in EEGLAB. Therefore, the electronic device first uses ErgoLAB's dedicated data analysis software to export the raw PROJ format EEG data into the universal EDF format. Then, the converted EDF format file is imported into the BIOSIGtoolbox module in the EEGLAB toolbox to complete the reading, loading, and time-series construction of the EEG signals. Basic information such as the number of channels, sampling rate, and signal duration is analyzed to form a multi-channel EEG time-series data structure that can be further processed.

[0087] Next, the electronic device can use the international 10-20 system standard to spatially locate the EEG electrodes, ensuring that the electrode positions strictly correspond to the functional areas of the brain, providing a spatial basis for subsequent brain region activation analysis. The channellocation function of the EEGLAB toolbox is used to import and spatially register the electrode coordinates, compensating for the phase delay caused by signal transmission through the scalp. Amplitude normalization is then performed on the signals from each channel to improve the spatiotemporal consistency of the whole-brain channel data, ensuring the accuracy of subsequent topographic mapping and microstate analysis.

[0088] Electronic devices can employ finite impulse response (FIR) filters to bandpass filter EEG signals, utilizing their linear phase characteristics and low computational noise to retain the effective signal while filtering out interference. A 1Hz high-pass filter is applied to remove DC offset and low-frequency baseline drift. A 30Hz low-pass filter is applied to retain the effective frequency band components of the EEG signal while filtering out high-frequency environmental noise and equipment interference. A notch filter is used to remove 50Hz power frequency interference and eliminate periodic noise from the power supply system. The filtered signal yields a pre-purified EEG signal that is frequency-clean, baseline-drift-free, and free from power frequency interference.

[0089] To further remove non-stationary noise and mode aliasing interference, electronic devices can employ a combined denoising strategy of ICEEMDAN and wavelet thresholding. The electronic device performs adaptive noise-assisted decomposition on the preprocessed EEG signal, obtaining a series of Intrinsic Mode Components (IMFs) arranged from high to low frequency, and one residual component. After decomposition, a total of 9 IMF components + 1 residual component are obtained, where IMF1–IMF5 contain high-frequency noise and interference components. Wavelet thresholding is then used to denoise the high-frequency components IMF1–IMF5, removing residual noise while preserving signal abrupt changes and effective details. The denoised effective mode components are then reconstructed with the unprocessed stationary components to obtain an EEG signal with smooth waveform, eliminated extreme peaks, and a significantly improved signal-to-noise ratio.

[0090] Finally, to address physiological artifacts such as those from electrooculography (EOG), electromyography (EMG), and electrocardiography (ECG) that overlap with the EEG spectrum, the FastICA (Fast Independent Component Analysis) algorithm was employed for blind source separation and artifact removal. The EEG signal was centered to satisfy the zero-mean condition. A whitening transformation was then performed on the centered signal to remove inter-component correlations and improve the efficiency of independent component separation. Through a fixed-point iterative algorithm, orthogonalization, and standardization steps, the independent components (ICs) of each channel were quickly converged. The ICLabel toolbox was used to automatically classify all independent components, identifying non-EEG components such as EOG, EMG, and ECG. A judgment threshold was set: when the probability of a non-EEG component was >0.8, the artifact component was marked and removed. The remaining neurogenic independent components were retained for signal reconstruction, ultimately yielding artifact-free, high-purity, and high signal-to-noise ratio target EEG data, providing a reliable data foundation for subsequent frequency domain analysis, microstate analysis, and feature extraction.

[0091] After the above five-step standardization process—data import, electrode localization, FIR filtering, ICEEMDAN-wavelet joint denoising, and FastICA artifact removal—the obtained target EEG data has the following characteristics: no power frequency interference, no baseline drift; no physiological artifacts such as electrooculography, electromyography, and electrocardiography; smooth and stable waveforms, and complete preservation of effective neural components; and can be directly used for subsequent feature calculations such as power spectrum, energy spectrum, global field power, and EEG microstates.

[0092] Step S2032: Calculate the power spectral density and energy spectral density corresponding to the target EEG data.

[0093] Specifically, the theoretical basis of power spectral density is the Wiener-Khinchin theorem, which states that power spectral density is the Fourier transform of the signal's autocorrelation function. To avoid spectral distortion and large variance caused by traditional methods, electronic devices can use the Welch method for PSD estimation.

[0094] The electronic device can divide the preprocessed target EEG data into several data segments based on 512 sampling points. A 256-point overlap is set between segments to improve data utilization. A Hamming window is applied to each signal segment to suppress spectral leakage.

[0095] Perform a 1024-point Fast Fourier Transform (FFT) on each windowed signal segment to obtain the frequency domain sequence, and then calculate the power spectrum of a single segment using the modulus squared. ; Where: x i w(n) represents the i-th segment of the target EEG signal, w(n) is the Hamming window function, M is the length of each signal segment, and U is the window function energy normalization factor.

[0096] Electronic devices can average the power spectrum of all segments to obtain a smooth and stable global power spectral density. Where K is the total number of segments.

[0097] The electronic device calculates the average power of the four classic EEG frequency bands δ, θ, α, and β, respectively, as the characteristics of the EEG power spectrum (PSD).

[0098] Energy spectral density (ESD) is based on Passevar's theorem, ensuring that the total energy of a signal remains conserved in both the time and frequency domains. ESD characterizes the total energy per unit frequency, directly preserving the energy accumulation characteristic without averaging. Electronic devices can employ the same segmentation, windowing, and FFT processing flow as PSD, but without performing averaging; they simply superimpose the energy spectra of each segment.

[0099] The formula for calculating a single-segment energy spectrum is: Where U is the window function correction factor: .

[0100] The electronic device directly sums all segmented energy spectra to obtain the final global energy spectral density. .

[0101] Electronic devices statistically analyze the energy values ​​of the δ, θ, α, and β frequency bands to form the characteristics of the brainwave energy spectrum.

[0102] Finally, the electronic device can calculate the whole-brain average of PSD and ESD for all EEG channels, improving feature stability. Independent samples t-tests were used to verify differences in frequency bands under different contextual awareness: PSD and ESD in the δ, θ, α, and β bands all showed significant differences (P<0.05); under low contextual awareness, θ, α, and β power / energy significantly increased, while δ significantly decreased. To further enhance representational ability, three types of cross-frequency band ratios were calculated: θ / β, α / β, and θ / α. These ratios can effectively reflect changes in brain cognitive load and contextual awareness.

[0103] After completing PSD and ESD calculations, the following EEG frequency domain features are output: whole-brain average power spectral density (PSD) in the δ, θ, α, and β bands; whole-brain average energy spectral density (ESD) in the δ, θ, α, and β bands; and cross-band ratios of θ / β, α / β, and θ / α. These features can be directly used as input for subsequent EEG state analysis, microstate calculations, and contextual awareness assessment models.

[0104] Step S2033: Calculate the global field power corresponding to the target EEG data.

[0105] Specifically, the electronic device reads the preprocessed target EEG data sequentially, node by node, and extracts the potential values ​​of all electrode channels at the current moment.

[0106] Let the total number of electrode channels be C, V i (t) represents the potential of the i-th channel at time t. The electronic device calculates the arithmetic mean of all electrode potentials at time t: .

[0107] The electronic device calculates the current global field power (GFP) using the standard deviation formula, as follows: .

[0108] The electronic device can repeat the above calculation point by point along the time axis to obtain a global field power time series that corresponds one-to-one with the duration of the EEG signal, thus completing the output of this step. The GFP peak corresponds to the most typical moment of brain activation characteristics and is the core basis for subsequent brain topography screening.

[0109] Step S2034: Based on the global field power, determine the target EEG state characteristics corresponding to the target EEG data.

[0110] Specifically, step S2034 above may include the following steps: Step b1: Based on global field power, the target brain topography map corresponding to the target EEG data is obtained by filtering.

[0111] Specifically, the electronic device can draw a topographic map of the brain at each moment based on the spatial distribution of multi-channel potentials and the spatial location of electrodes, visually presenting the spatial distribution of brain activity. Then, the electronic device uses the global field power time series as a screening criterion, prioritizing the topographic map corresponding to the peak time of GFP, as this type of map has a high signal-to-noise ratio and typical brain activation characteristics; invalid topographic maps corresponding to abnormally high / low GFP, waveform distortion, and artifact residues are removed. After screening, a set of maps with normal morphology that can truly reflect neural activity is retained, which is the target brain topographic map, used as input samples for cluster analysis.

[0112] Step b2 involves iteratively clustering the target brain topography map to obtain multiple initial clustering results.

[0113] Specifically, the electronic device can select the target brain topography map corresponding to the GFP peak as the initial cluster center; set multiple different cluster numbers (such as 2, 3, 4, 5, etc.) to form multiple clustering experimental schemes. The electronic device can convert all target brain topography maps into feature vectors and classify the samples according to the spatial morphological similarity of the maps; after each round of classification, the cluster centers of each category are recalculated, and the process is iteratively updated repeatedly. When the cluster centers no longer change significantly and the sample categories no longer jump, the single-group clustering is considered to have converged. Different cluster number schemes generate multiple sets of initial clustering results, each set of results including the number of cluster categories and the set of topography map samples under each category.

[0114] Step b3: For each initial clustering result, calculate the global explained variance and cross-validation criterion corresponding to the initial clustering result.

[0115] Specifically, Global Explained Variance (GEV) is the core evaluation metric for microstate clustering. It quantifies the proportion of typical microstate topological patterns obtained from clustering that can explain the spatial features of the original global brain topography. The larger the GEV value, the higher the degree of restoration of the original EEG spatial distribution information by the current cluster category, the clearer the distinction between various microstate features, and the higher the effectiveness of clustering. Conversely, a lower GEV value indicates that the clustering has failed to effectively extract typical EEG spatial patterns, resulting in significant information loss.

[0116] Electronic devices can be treated as a single initial clustering result, containing several cluster categories (microstates), each corresponding to a standard topological template; the entire target brain topography map serves as the original sample set to be interpreted. For each original brain topography map, its spatial similarity or residual with its respective cluster template is calculated, measuring the degree to which a single map is represented by its corresponding microstate template. The smaller the difference between the sample and the template, the better the fit.

[0117] Electronic devices can decompose the total variance of all samples into within-class variance and between-class variance. The total variance represents the dispersion of all brain topographic maps relative to the overall mean, representing the total information content of the original data; the within-class variance represents the dispersion of samples within the same cluster relative to that cluster template, representing residual information not explained by clustering; and the between-class variance represents the dispersion between different cluster templates, representing the effective feature information distinguished by clustering.

[0118] The Global Explained Variance (GEV) formula is: GEV = (Inter-class Variance / Total Variance) × 100%. The GEV value ranges from 0 to 100%. In this experimental scenario, the optimal clustering scheme achieves a GEV of over 80%, indicating that the vast majority of EEG spatial features can be explained by the extracted microstates. This embodiment compares GEV with different numbers of clusters (2, 3, 4, 5, etc.): when the number of clusters is 4, the GEV reaches a higher level, proving that 4 types of microstates can maximally explain the spatial distribution features of the pilot's brain topography.

[0119] Cross-validation (CV) is used to evaluate the generalization ability, stability, and robustness of clustering results. Clustering is an unsupervised learning process, which is greatly affected by the random selection of samples and the initial cluster centers. The CV metric quantifies the degree of fluctuation in the results through cross-validation: the smaller the CV value, the less the clustering results are affected by random samples and initial parameters, the stronger the consistency of the clustering patterns obtained from different subsets of samples, and the more stable and generalizable the clustering results are.

[0120] Electronic devices can randomly divide all target brain topographic map samples corresponding to a current cluster into multiple non-overlapping subsets (commonly using k-fold cross-validation). Iterative clustering is performed, selecting one subset as the test set and the remaining subset as the training set. Improved K-means clustering with the same parameters is repeated on the training set to obtain multiple sets of sub-clustering results. The deviation of the results is calculated by comparing the differences in category topology, sample partitioning, and cluster centers across the multiple sets of sub-clustering results. The overall deviation and error value are statistically analyzed and integrated to obtain the cross-validation criterion (CV) for that cluster. The CV is a comprehensive indicator that quantifies deviation: the smaller the deviation, the lower the CV value, and the stronger the clustering robustness; if the CV is large, it indicates that the clustering scheme depends on specific samples, and the results will change significantly after changing the samples, lacking general applicability.

[0121] In this embodiment, when the number of clusters is set to 4, the CV reaches its minimum value, indicating that the micro-state clustering result is the most stable under this grouping method and is not affected by random sample partitioning.

[0122] GEV (Generative Encephalogram) focuses on the explanatory power of clustering for the original features (effectiveness), while CV (Consciousness of Clustering) focuses on the stability and reliability of the clustering results (credibility). The combination of these two factors forms a dual evaluation system. High GEV and high CV alone: ​​The clustering appears to explain many features, but the results are highly volatile and unstable, making it practically unusable. Low CV and low GEV alone: ​​The clustering results are stable, but they cannot effectively extract core EEG features, resulting in low information utilization. Maximizing GEV and minimizing CV: This simultaneously satisfies strong feature explanatory power and stable, reliable results, representing the optimal clustering scheme and the selection principle adopted in this study.

[0123] Step b4: Based on the principle of maximizing the global explained variance and minimizing the cross-validation criterion, update each initial clustering result until the clustering converges, and determine the optimal number of clusters and the target clustering results corresponding to the optimal number of clusters.

[0124] Specifically, the electronic device follows a joint optimality principle of maximizing the global explained variance and minimizing the cross-validation criterion, comparing the index data of all initial clustering schemes. For clustering schemes that do not reach the optimal index, parameters such as cluster centers and number of clusters are adjusted, and steps b2 and b3 are re-executed for iterative optimization. After multiple rounds of iteration, the GEV and CV indices no longer fluctuate, the clustering results remain stable, and overall convergence is determined. In this embodiment, the optimal number of clusters is finally determined to be 4, and 4 sets of target clustering results under this number are output, corresponding to 4 typical EEG microstates (MSA, MSB, MSC, MSD). For example, as shown... Figure 3The diagram illustrates the principle of microstate analysis. This diagram fully demonstrates the standardized processing chain for EEG microstate extraction and parameter analysis, starting from the raw EEG signal and progressively completing GFP calculation, EEG topographic map clustering, microstate classification, time series generation, and feature parameter extraction. Specifically, the input is an artifact-free EEG signal: the dense red waveform at the top represents the multi-channel raw EEG time series data after removing interference from electrooculography (EOG), electromyography (EMG), etc.; the global field power (GFP) is calculated from the full-channel EEG, resulting in the single-channel GFP time series curve below; the GFP peak corresponds to the instantaneous maximum value of the whole-brain electric field energy and is a marker point for selecting key EEG topographic snapshots; the segmented intervals marked in the diagram are the GFP peak sampling windows. The whole brain electric field distribution corresponding to each GFP peak moment is picked up to generate a GFP peak topographic map (color brain potential distribution map); unsupervised clustering operation is performed on massive instantaneous brain topographic maps, and the electric field spatial morphology is automatically grouped to obtain the topographic map of each category (the color tree-like hierarchy in the figure is the clustering hierarchy, and different color branches represent brain topographic categories with different spatial topologies). Four classic EEG microstate topographic maps were selected: MSA, MSB, MSC, and MSD (four types of spherical color EEG topographic maps, corresponding to four basic whole-brain electric field configurations in the human brain). The full-time EEG was spatially matched frame by frame with the four types of microstate topographic maps, and the continuous EEG time sequence was encoded into a microstate sequence in which MSA / MSB / MSC / MSD alternated (bottom left bar time sequence plot: different color blocks correspond to the continuous distribution of different microstates on the time axis, the horizontal axis is the sampling point, and the vertical axis is the GFP amplitude). Based on the time sequence, microstate parameters (including quantitative indicators such as the average duration, frequency of occurrence, coverage ratio, and transition probability of each type of microstate) were statistically calculated.

[0125] Step b5: Based on the similarity between each target clustering result and the preset state category, determine the target state category corresponding to each target clustering result.

[0126] Specifically, electronic devices can pre-establish four types of microstate standard topological templates, corresponding to different brain functional activity patterns, by combining prior EEG knowledge of different situational awareness in flight scenarios. Then, the spatial morphological similarity between each type of target cluster topographic map and the standard template is calculated. According to the principle of maximum similarity, each group of target cluster results is bound to a preset standard state category, clarifying the target state category (MSA, MSB, MSC, MSD) corresponding to each type of map.

[0127] Step b6: Extract the clustering time series features corresponding to the target clustering results based on the target state category corresponding to each target clustering result.

[0128] Specifically, electronic devices can statistically analyze the duration of each consecutive occurrence of a single type of microstate, calculate the average value, and reflect the stability and persistence of that state. The formula is as follows: In the formula, N k T represents the number of occurrences of the k-th type of microstate. k,i Let be the duration of the i-th instance of this type.

[0129] Electronic devices can calculate the ratio of the total duration of a single microstate to the total duration of the entire analysis, representing the overall active proportion of that neural activity pattern. The formula is: In the formula, T total This represents the total signal analysis time.

[0130] Electronic devices can count the total number of occurrences of a single type of microstate throughout the entire time series, reflecting the frequency of state occurrence.

[0131] Electronic devices can also count the number of transitions between different microstates, calculate the probability of state transitions, and characterize the switching patterns of different functional states of the brain. The formula is: In the formula: Numm→n is the total number of jumps from state m to state n.

[0132] Finally, the electronic devices compile the duration, coverage, frequency of occurrence, and conversion probability corresponding to all categories to form a complete clustering time series feature set.

[0133] Step b7: The temporal features of each cluster are fused to obtain the target EEG state features.

[0134] Specifically, electronic devices can perform outlier removal and dimension normalization on all temporal features to eliminate differences in the numerical ranges of different indicators. Then, using feature splicing or weighted fusion, multi-dimensional temporal features such as average duration, coverage, frequency of occurrence, and state transition probability are integrated, preserving all effective information on the spatial patterns, temporal evolution, and state transitions of microstates. After fusion, the target EEG state features are obtained with reduced dimensions and complete representational capabilities, which can be directly used for subsequent multimodal feature fusion, pilot situational awareness assessment, and analysis.

[0135] Step S2035: Extract source features from the target EEG data to obtain the target EEG source features corresponding to the target EEG data.

[0136] Specifically, step S2035 above may include the following steps: Step c1 involves refining the cerebral cortex using a standard human brain template, dividing it into multiple independent voxel units.

[0137] Specifically, the electronic device can use the MNI152 standard human brain template as the segmentation benchmark. This template is a universal standardized human brain anatomy template that is compatible with international 10-20 electrode systems and EEG source analysis requirements.

[0138] The electronic device delineates the gray layer of the cerebral cortex as the object of segmentation, performing fine segmentation only on the main areas where neural electrical activity occurs, excluding non-functional tissues such as the skull, scalp, and cerebrospinal fluid. Then, the target region is segmented into a three-dimensional mesh according to a uniform spatial scale, dividing the complete gray layer of the cerebral cortex into 6239 non-overlapping independent voxel units. All voxels have the same spatial size and are uniformly arranged, forming a whole-brain discretized computational unit, providing a foundation for subsequent point-by-point current density calculation.

[0139] Step c2: Establish a standardized digital brain model based on the correspondence between each independent voxel unit and the fixed spatial coordinates of the brain and the functional areas of the cortex.

[0140] Specifically, the electronic device can assign a unique three-dimensional spatial coordinate to each independent voxel unit, completing a one-to-one mapping between voxels and three-dimensional coordinates, and determining the positional information of all voxels in the MNI standard brain space. Combining knowledge of brain anatomy and the Broadman partitioning system, a correspondence between voxels and cortical functional areas and Broadman partitions is established, clarifying the brain lobe, functional brain region, and corresponding partition number of each voxel. Finally, by integrating four types of information—head model structure, voxel grid, three-dimensional coordinates, and anatomical partitions—a standardized digital brain model is constructed. This model unifies spatial benchmarks, anatomical benchmarks, and computational units, ensuring that the source tracing results of different subjects and under different experimental conditions are comparable and analyzable.

[0141] Step c3: Input the target EEG data into the standardized digital brain model, and calculate and output the spatial distribution of whole brain current density corresponding to all independent voxel units based on the sLORETA low-resolution electromagnetic tomography algorithm.

[0142] Specifically, electronic devices can connect the preprocessed, frequency domain analyzed, and microstate analyzed target EEG data to a standardized digital brain model, complete the coordinate registration of the international 10-20 electrode system and the MNI152 template, and achieve alignment of scalp potential signals with the model space.

[0143] Then, the inverse problem is solved by relying on the standardized low-resolution electromagnetic tomography (sLORETA) algorithm. The algorithm is based on the current density normalization theory, introduces forward model coefficients, regularization constraints, cost equations and inverse operator matrices, and combines the head conduction characteristics to reverse the brain endogenous activity from the scalp observation potential.

[0144] Specifically, electronic devices can construct linear mapping equations according to the sLORETA model: Where V is the observation signal matrix, representing the target EEG potential signal acquired by the scalp electrodes; A is the forward model coefficient matrix, determined by the head anatomy, tissue conductivity, electrode position, and voxel spatial position, describing the attenuation and mapping law of the endogenous electrical signal in the brain transmitted to the scalp via volume conduction; J is the current density source signal matrix, i.e., the endogenous activity of each voxel in the cerebral cortex to be solved; c is a constant term used for model baseline correction. The scalp observation potential is the result obtained by weighting and superimposing the current sources of each voxel in the brain through the head conduction matrix, and then adding the baseline constant. This formula is the expression for the forward problem of EEG; given the endogenous J, the scalp potential V can be calculated; the source tracing task is to know V and A, and then calculate J.

[0145] Then, the electronic device introduces a cost equation and a regularization constraint to optimize the solution objective. Let... To actually observe scalp potential signals, with λ as the normalization coefficient, the target cost equation is constructed as follows: Among them, the first item The first term is the fitting error term, which characterizes the deviation between the model-calculated potential and the actual observed potential. The smaller the deviation, the closer the intrinsic solution is to the actual observation. The second term... R is the regularization constraint term, and R is the regularization constraint matrix, designed in conjunction with the sLORETA current density normalization rule; λ controls the constraint strength, used to suppress noise and limit the range of solutions. By minimizing the cost equation, a balance is achieved between "fitting the observed data" and "the stability and rationality of the solution": on the one hand, the error between the model and the measured signal is reduced, and on the other hand, the occurrence of extreme intrinsic solutions without physical meaning is avoided, thus solving the multi-solution problem of EEG inverse, while attenuating the computational bias caused by signal noise.

[0146] Next, the electronic device finds the minimum value of the cost equation, derives the inverse operator through matrix operations, and obtains a unique and stable solution for the intrinsic signal.

[0147] Electronic devices define intermediate matrices sequentially: I is the identity matrix. u is a unit vector. The electronic device inverts matrix M to obtain the inverse operator matrix M. -1 This matrix is ​​the core conversion operator for inferring the intrinsic current from the scalp potential. Then, combined with the inverse operator matrix, the final solution of the current density source signal is derived: This solution simultaneously satisfies the minimum fitting error and optimal regularization constraint, possessing three major characteristics: high accuracy, noise resistance, and uniqueness.

[0148] Finally, the electronic device can decompose the global source signal into a single spatial computing unit to complete the solution of the current density distribution of the whole cerebral cortex.

[0149] Based on the already defined MNI152 brain model, the electronic device binds the global current source matrix J obtained by solving the formula to 6239 independent voxels one by one, with each voxel corresponding to a set of independent components in J.

[0150] Finally, the first voxel is selected sequentially, and its corresponding spatial location, forward coefficient component, and inverse operator component are retrieved. These are then substituted into the derived source signal calculation formula to calculate the current density at that voxel location. The current voxel's 3D coordinates and corresponding current density are recorded to form the calculation result for a single voxel. The electronic device iterates through all remaining voxels, repeating the above calculation process until all 6239 voxels have been calculated. The 3D spatial coordinates and current density values ​​of all voxels are then summarized to construct a whole-brain current density spatial distribution dataset, fully presenting the intensity and spatial distribution of neural currents in various parts of the cerebral cortex, thus completing the sLORETA source tracing calculation.

[0151] Step c4: Select the spatial distribution of partition current density corresponding to the Brodman partition from the spatial distribution of whole brain current density.

[0152] Specifically, the electronic device can call the voxel-Brodman partition association mapping table pre-established in step c2. This table records the three-dimensional coordinates, the lobe to which each voxel belongs, and the corresponding Brodman partition number. Using the experimental target partitions BA2, BA6, BA44, BA45, BA46, and BA47 as search keywords, all 6239 voxels are traversed and searched, and the target partition to which each voxel belongs is marked, thus completing the partition index matching.

[0153] Electronic devices can segment the whole-brain current density dataset based on the matching results. Voxels are grouped by partition number, and all voxels belonging to the same Broadman partition, along with their corresponding 3D coordinates and current density values, are grouped together. The electronic device removes voxels and their corresponding data that do not belong to the target partition, eliminating interference from irrelevant regions. Data collection for all target partitions is completed sequentially, with each partition forming an independent dataset containing the location and current density information of all voxels within that partition. Finally, the independent spatial distribution of current density for each target Broadman partition is obtained, serving as a dedicated data source for subsequent feature extraction.

[0154] Step c5 involves extracting features from the spatial distribution of the brain region's current density to obtain the average current density, maximum activation intensity, and spatial activation range of the brain region.

[0155] Specifically, for a single Broadman partition dataset, quantitative indicators are extracted from three dimensions: overall strength, local extrema, and spatial coverage, thereby transforming the original distributed data into computable features.

[0156] Specifically, suppose there are n voxels in a certain partition, and the current density of the i-th voxel is J.i Calculation formula: The electronic device calculates the arithmetic mean of the current density of all voxels within the partition. This index characterizes the overall neural activation intensity of the entire functional brain region and directly reflects the average activity level of the brain region.

[0157] The electronic device iterates through the current density values ​​of all voxels within the partition and takes the maximum value: Jmax=max(J1,J2,…,Jn); This index represents the strongest neural discharge intensity at a local site within the region, used to characterize the activity level of core activation points in the brain region.

[0158] Then, the electronic device can combine the experimental noise level and signal characteristics to set a fixed current density activation threshold J. th The number of effective activated voxels, Nactive, within the electronic device's statistical partition that satisfies Ji>Jth is calculated. Combining this with the standard spatial size of a single voxel, the actual area of ​​the activated region is calculated from the number of effective voxels, thereby quantifying the spatial coverage of neural activity in that brain region.

[0159] Finally, the electronic device records the average current density, maximum activation intensity, and spatial activation range of the current partition, forming a feature subset for that partition. This process is repeated for all target Brodman partitions to complete feature extraction for all partitions.

[0160] Step c6: Based on the average current density, maximum activation intensity, and spatial activation range of the brain region, the target EEG source characteristics are obtained.

[0161] Specifically, the electronic device unifies and summarizes the three features of all target regions, including BA2, BA6, BA44, BA45, BA46, and BA47, to construct a multi-dimensional feature set. The set simultaneously contains activation information from different brain regions and different quantification dimensions, comprehensively covering the source tracing results of the entire target brain region.

[0162] The set of multidimensional quantitative indicators integrated by electronic devices constitutes the target EEG source feature. This feature can be directly used to analyze the neural activity patterns of various brain regions under different situational awareness conditions. At the same time, as an independent feature dimension, it participates in the subsequent multimodal feature fusion and the construction of the pilot situational awareness assessment model.

[0163] Step S2036: The power spectral density, energy spectral density, target EEG state characteristics, and target EEG source characteristics are fused to obtain the target comprehensive EEG characteristics.

[0164] Specifically, electronic devices can splice or weightedly fuse power spectral density, energy spectral density, target EEG state characteristics, and target EEG source characteristics to obtain comprehensive target EEG characteristics.

[0165] Step S204: Extract features from the original face image to obtain the target comprehensive face features.

[0166] Specifically, step S204 above may include the following steps: Step S2041: Perform image preprocessing on the original face image to obtain the target face image.

[0167] Specifically, electronic devices can convert the input original color RGB face image into a grayscale image, eliminating computational redundancy caused by color channels and avoiding interference from color information in gradient calculation.

[0168] Then, the electronic device can perform a Gamma transformation on the grayscale image to adjust the overall brightness of the image and reduce the impact of uneven lighting, shadows, and reflections on feature extraction. Formula: I out =I in γ, where I in I represents the original pixel value. out γ represents the corrected pixel value, and γ is the Gamma correction coefficient.

[0169] The image after grayscale conversion and Gamma correction is the target face image, which is used for subsequent face region detection and key point localization.

[0170] Step S2042: Recognize the target face image and determine the effective region of interest of the face.

[0171] Specifically, electronic devices can calculate gradients pixel by pixel to capture image edge and texture information.

[0172] Electronic devices can use the horizontal difference operator to calculate the intensity of pixel variation in the horizontal direction (x-direction), i.e., to calculate the horizontal gradient, as shown in the formula: .

[0173] Electronic devices can use the vertical difference operator to calculate the intensity of pixel changes in the vertical direction, i.e., to calculate the vertical gradient (y-direction), as shown in the formula: Among them, G x G reflects the intensity of horizontal edges in an image. y It reflects the intensity of vertical edges in an image. The larger the gradient, the more drastic the pixel change at that location, and the more likely it is to be the outline of a face or the edges of facial features.

[0174] Then, the electronic device further synthesizes the gradient intensity and gradient direction, forming the basic unit of the HOG feature. The formula for calculating the gradient magnitude (intensity) is: The larger the amplitude, the more distinct the edge and the more prominent the texture at that point. The formula for calculating the gradient direction (angle) is: The gradient direction is used to describe the edge orientation and is key to constructing a histogram.

[0175] Electronic devices can divide a target face image into small units, statistically analyze the gradient direction distribution within each unit, and form a local texture description. The electronic device uniformly divides the target face image into several small units (e.g., 8×8 pixels / unit). The electronic device divides the 0°~180° range into nine equal intervals, each interval occupying 20°: 0°~20°; 20°~40°; ...; 160°~180°.

[0176] The electronic device can determine the angle range within each pixel of a cell based on its gradient direction, and then perform weighted voting based on the gradient magnitude. This results in a 9-dimensional gradient direction histogram for each cell. Each cell outputs a 9-dimensional feature vector describing the texture orientation distribution of that local region.

[0177] To eliminate the effects of variations in lighting, shadows, and contrast, multiple units are grouped into blocks and normalized. Electronic devices can combine multiple adjacent units into a single block (e.g., 2×2 units / block), with blocks overlapping each other. Then, the histogram vectors of all units within the block are jointly normalized. This makes the features invariant to illumination and shadow.

[0178] Electronic devices can concatenate the normalized features of all blocks in a target face image to form the HOG feature vector of the entire image. HOG features can stably describe the contour, facial features, and texture structure of a face, and are classic features for face detection.

[0179] Next, the electronic device can input the HOG features into the SVM classifier to determine whether the current region is a face or background. The core of SVM is finding the optimal classification hyperplane: w·x + b = 0, where w is the hyperplane normal vector, x is the HOG feature vector, and b is the intercept. The goal of SVM is to maximize the margin between face and background samples, thereby improving generalization ability. Constraints: .

[0180] Electronic devices can transform the problem of maximizing the margin into a standard convex optimization problem: Satisfying the constraints: .

[0181] Electronic devices can introduce Lagrange multipliers to transform the original problem into a dual problem, obtaining support vectors and ultimately determining the decision function. For each sliding window of the target face image, HOG features are calculated, input into an SVM, and the result is output as a face / background binary classification.

[0182] The electronic device can locate the face position and output a standard rectangular region based on the SVM classification results. SVM outputs a high-confidence score for the face region, while background regions with low scores are used to remove duplicate and overlapping detection boxes, retaining the optimal face box. The electronic device outputs the minimum bounding rectangle (Rect) of the face region, including the coordinates of the top-left corner, width, and height. This final output rectangle is the effective Region of Interest (ROI) for the face, serving as the dedicated input area for subsequent facial landmark localization.

[0183] Step S2043: Identify the effective regions of interest of the face and determine multiple standard facial key points.

[0184] Specifically, within the detected effective Region of Interest (ROI) of the face, the electronic device first performs initial position estimation of the facial key points. The input is the determined face region bounding box (Rect). Based on the position and size of the face bounding box, the electronic device generates a set of standardized and averaged initial positions of the facial key points and outputs an initial key point position vector S0.

[0185] The electronic device uses a first-level regression model (global coarse regression) to quickly and roughly approximate the actual key point location, completing a rough alignment from the initial position to the actual position.

[0186] Specifically, the electronic device inputs the target face image I and the current key point positions S. t-1 This establishes a mapping relationship between image features and keypoint offsets. Electronic devices can train a regressor to learn a mapping function from image features to the true offset increments of keypoints: r t (I,S t-1 After completing the first coordinate correction, a coarse set of key points is output, which quickly brings the key points closer to their true positions, laying the foundation for fine regression.

[0187] Next, the electronic device uses a regression tree ensemble model based on the second-level regression model (regression tree fine fitting) to perform point-by-point fine iterative optimization of the key point coordinates.

[0188] Specifically, the electronic device can extract local image features within the neighborhood of the current keypoint, capturing facial texture and structural information. The feature space is divided into several independent subspaces, and local linear fitting is performed within each subspace to achieve high-precision offset prediction. The regression tree model outputs the optimal offset increment, which is used to correct the current keypoint coordinates. The electronic device uses a cascaded regression iteration formula to complete the coordinate update: In the formula: S t Let S be the keypoint location vector after the t-th layer regression. t-1 The key point location for the regression output of the previous layer, r t Let I be the t-th layer regressor (regression tree model), and let I be the input target face image.

[0189] The electronic device employs a cascaded regression system consisting of K layers of regressors connected in series, continuously reducing prediction errors through multiple iterations. Each regression layer further corrects coordinate deviations based on the output of the previous layer. This progressively reduces keypoint localization errors, achieving a transition from coarse to fine localization. Iteration continues until the change in keypoint coordinates is less than a threshold, reaching convergence. This ultimately yields high-precision, highly robust facial keypoint coordinates. This process, from coarse to fine, gradually optimizes and stably approximates the true positions of facial features and contours. After convergence, the model outputs 68 two-dimensional coordinate points covering the entire facial structure, encompassing complete facial feature regions: facial contour points, left and right eyebrow keypoints, left and right eye keypoints, nose region keypoints, and mouth region keypoints. Each keypoint is output as (x, y) two-dimensional coordinates, collectively forming the foundation set for subsequent feature model calculations.

[0190] Step S2044: Select multiple core feature points from multiple standard facial key points.

[0191] Specifically, electronic devices can identify the eyes, mouth, and facial contours as the three main effective feature regions by combining the characteristics of facial changes during flight. The morphological deformation of these three regions can intuitively reflect differences in facial state. Other points, such as eyebrows and nose, contribute less to the state representation and are therefore excluded. Target points are then selected from the total of 68 facial key points, grouped by region.

[0192] The electronic device focuses on the eye area and extracts 6 feature points from 68 standard key points, uniformly numbered p1-p6. This set of points covers key positions of the corner of the eye and the upper and lower eyelids, and is specifically used to calculate the aspect ratio (EAR) of the eye, quantify the degree of eye opening and closing, and complete the extraction and separate storage of eye point data.

[0193] The electronic device focuses on the mouth area and extracts six feature points from the remaining key points, uniformly numbered p7-p12. This set of points includes the corners of the mouth and the contours of the upper and lower lips, and is specifically used to calculate the mouth aspect ratio (MAR), quantify the mouth opening and closing amplitude, and complete the extraction and separate storage of mouth points.

[0194] The electronic device focuses on the outer contour region of the face and extracts 6 feature points, uniformly numbered p13-p18. These points are distributed on both sides of the facial contour and are specifically used to calculate the facial contour aspect ratio (FAR), characterizing the overall facial morphological changes, and completing the extraction and separate storage of facial contour points.

[0195] Step S2045: Based on each core feature point, calculate the core features of the eye aspect ratio, the core features of the mouth aspect ratio, and the core features of the face contour aspect ratio.

[0196] Specifically, electronic devices can use six feature points of the eye to calculate the core aspect ratio feature (EAR) of the eye, using the following formula: The smaller the EAR value, the more closed the eyes (indicating fatigue or drowsiness). For example, ... Figure 4 The image shown is a schematic diagram of the core aspect ratio feature EAR of the eye.

[0197] Electronic devices can use six feature points of the mouth to calculate the degree of mouth opening, i.e., the core feature of the mouth's aspect ratio, using the following formula: The larger the MAR value, the wider the mouth is opened (when yawning or talking). For example, ... Figure 5 The image shown is a schematic diagram of the mouth's aspect ratio core feature (MAR).

[0198] Electronic devices can use six feature points of the facial contour to calculate changes in the facial contour, namely the core feature of the facial contour's aspect ratio, using the following formula: FAR reflects facial muscle tension and expression intensity. For example, ... Figure 6 The image shown is a schematic diagram of the facial contour aspect ratio (FAR), a core feature of facial contours.

[0199] Step S2046: The core features of the eye aspect ratio, the core features of the mouth aspect ratio, and the core features of the face contour aspect ratio are fused to obtain the target comprehensive face features.

[0200] Specifically, electronic devices can normalize EAR, MAR, and FAR to eliminate dimensional differences.

[0201] Then, the three independent features are concatenated into a one-dimensional feature vector in a fixed order to obtain the target comprehensive face feature F. face =[EAR,MAR,FAR].

[0202] Step S205: The target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features are fused to generate the target's fused features.

[0203] Specifically, step S205 above may include the following steps: Step S2051: Perform temporal alignment processing on the target integrated eye movement features, target integrated EEG features, and target integrated face features to obtain temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned face features.

[0204] Specifically, the electronic device can first statistically analyze the original acquisition parameters of the three types of features to clarify the sources of temporal differences. The electronic device reads the original sampling frequency, single-frame timestamp, data duration, and frame sequence number of the three types of data: eye movement, electroencephalography (EEG), and facial video. The electronic device records the start-up time difference and sampling interval deviation of each acquisition device, distinguishes the time axis offset and frame rate inconsistency of different data sources, and determines the global reference time axis for the entire experiment. Usually, the data with the highest stability and optimal sampling accuracy is selected as the reference (such as EEG signals).

[0205] Electronic devices can use a selected baseline timeline to calibrate the start time of all data to the same global zero point, eliminating the overall time offset caused by the order of device startup. For each data sample, a globally unified timestamp is established, linking each set of data from eye movements, EEG, and facial features to the same absolute time.

[0206] Because the three types of data have different original sampling rates, resampling is necessary to ensure consistent sampling intervals. Electronic devices can use a baseline sampling interval as a standard, downsampling data with frequencies higher than the baseline, extracting valid data points at fixed intervals, and removing redundant frames. For data with frequencies lower than the baseline, upsampling or interpolation is used to fill in missing data points using linear interpolation, nearest neighbor filling, or other methods, ensuring that the sampling intervals of all data sources are completely equal. After processing, eye-tracking, EEG, and facial feature data achieve the same frame rate and sampling step size.

[0207] Then, the electronic device iterates along the global timeline moment by moment, matching eye-tracking features, EEG features, and facial features at the same time point according to a unified timestamp. For a small number of segments with short-term frame drops or signal interruptions, repairs are made by combining preceding and following time-series data to ensure the continuity and integrity of the time sequence. The time-series binding of the three types of features is completed using a combination of frame number and timestamp, ensuring that at any given time point, the three types of features correspond one-to-one.

[0208] Finally, the electronic device checks the entire time sequence, removing invalid segments due to acquisition failures, excessive signal noise, or incorrect timestamps. Multiple sets of paired data are sampled and verified to confirm the absence of frame misalignment, time jumps, or reversed order. The duration and total number of frames of the entire time sequence are verified to be consistent. After benchmark calibration, resampling, frame-by-frame matching and verification, the final output includes: time-aligned eye-tracking features, time-aligned EEG features, and time-aligned facial features. These three types of features have completely unified timelines, timestamps, and frame sequences, with synchronized temporal relationships, and can be directly used for subsequent multimodal feature fusion, temporal modeling, and state recognition analysis.

[0209] Step S2052: Input the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features into the preset weight determination model, and output the target weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features respectively.

[0210] Specifically, step S2052 above may include the following steps: Step d1: Input the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features into the preset weight determination model.

[0211] Specifically, the electronic device can sequentially align eye-tracking features, EEG features, and facial features, determine the input interface defined by the model according to preset weights, and import the feature dataset in batches. At the same time, it loads the corresponding real cognitive labels (human-annotated standard situational awareness levels) to provide truth values ​​for subsequent loss calculations. Once the data loading is complete, the model enters the weight allocation and inference preparation stage.

[0212] Step d2: Preset weights to determine the initial weight information corresponding to the time-aligned eye-tracking features, time-aligned EEG features, and time-aligned face features generated by the model.

[0213] Optionally, the preset weight determination model can retrieve and summarize historical experimental data, feature correlation analysis, and discrimination ability assessment results to analyze the correlation strength and classification performance of eye-tracking, EEG, and facial features with the pilot's situational awareness state, clarifying the actual contribution level of each modality. The preset weight determination model combines modality discrimination ability and correlation degree to complete the initial assignment: larger coefficients are assigned to modalities with excellent recognition performance and close correlation with the target state; smaller coefficients are assigned to modalities with weaker discrimination ability and lower correlation, resulting in three initial sets of weight values. The electronic device checks individual coefficients to ensure that each weight falls within the (0,1) interval; the sum of the three sets of coefficients is calculated, and if the sum is not equal to 1, the size of each coefficient is fine-tuned proportionally; during the fine-tuning process, the relative weight ratio of each modality remains unchanged to ensure that the initial allocation logic does not change. After multiple rounds of verification and fine-tuning, a weight combination that simultaneously satisfies interval constraints and sum constraints is obtained as the final empirical initial weights.

[0214] Optionally, the preset weight determination model can call a random function to independently generate three temporary values, all of which must fall strictly within the open interval (0,1), denoted as k1, k2, and k3. The preset weight determination model calculates the sum of the three temporary coefficients: K = k1 + k2 + k3. The final weights are then obtained by normalizing the results using the normalization formula. After conversion, the three sets of weights naturally satisfy w. eye +w eeg +w face =1.

[0215] Check the converted weights one by one. Since the value of the original temporary coefficient is in the (0, 1) interval, the normalization result must satisfy 0 < w < 1. Outlier values and extreme values are removed to ensure data validity. After the check is passed, three sets of random weight coefficients are output to form random initial weight information, which is used for subsequent weighted fusion and iterative optimization.

[0216] Step d3: based on the initial weight information, perform weighted fusion on the time-aligned eye movement features, time-aligned electroencephalogram features and time-aligned face features, complete forward inference, and output the initial situation awareness level identification result corresponding to the initial weight information.

[0217] Specifically, the electronic device may adopt a feature-by-feature weighting method to complete the fusion, amplify the influence of high-contribution modalities and weaken the influence of low-contribution modalities. The electronic device can multiply the three groups of features under the same time sequence node by the corresponding weights respectively: ; wherein: , , are the original time sequence aligned feature vectors, respectively; , , are the weighted feature vectors. The electronic device accumulates the weighted features according to dimensions to generate a single comprehensive feature vector: .

[0218] The electronic device can input the fused comprehensive feature vector F all into the built-in inference module of the model, keep the data format and dimension consistent with the model input requirements, and complete the data loading before inference.

[0219] According to its own network hierarchical structure, the model processes the input comprehensive feature vector F all to carry out linear operation and activation function transformation layer by layer. The fixed network parameters of the trained model are called to perform deep semantic and feature mapping on the features, and mine the correlation between features and situation awareness levels. Only forward data flow and calculation are performed in this stage, and back propagation, parameter correction and weight update are not performed, so as to ensure that the network parameters and modal weights maintain the current state.

[0220] After the operation of all layers in forward propagation is completed, the inference module outputs the original result in two forms: one is continuous prediction score, and the other is discrete classification label. The electronic device can analyze and convert the original output of the model in combination with the preset level judgment rules: if the output is a prediction score, match the corresponding situation awareness level according to the score interval division standard; if the output is a classification label, it is directly mapped to the preset situation awareness level.

[0221] Finally, the electronic device defines the converted level result as the initial situational awareness level identification result, which represents the prediction effect of the model under the current initial weight combination. The result is stored and passed to the next link for subsequent global loss calculation and weight iterative optimization.

[0222] Step d4, calculating the global loss value between the initial situational awareness level identification result and the real cognitive label in the preset weight determination model.

[0223] Specifically, the electronic device may align the initial situational awareness level identification results of all samples obtained by forward reasoning of the model with the manually calibrated real cognitive labels in the data set sample by sample, so as to construct a complete prediction-ground truth pairing sequence and avoid error calculation distortion caused by time series misalignment. Then, a unified loss function is used to calculate the error in batches for all samples, and the global loss value is obtained by accumulation. Different from single-sample loss, the global loss can objectively reflect the overall adaptation effect of the current eye movement, electroencephalogram (EEG) and face weight proportions in the whole scene. A larger global loss value indicates that the current modal weight distribution is unreasonable, and there is a larger deviation between the multi-modal fusion feature and the real situational awareness state; a smaller global loss indicates that the current weight fusion effect is more consistent with the real cognitive law. Finally, the global loss value of the current iteration is output as the only supervision for subsequent weight gradient update.

[0224] Step d5, updating the initial weight information according to the global loss value to obtain the updated weight information respectively corresponding to the time-series aligned eye movement features, the time-series aligned EEG features and the time-series aligned face features.

[0225] Specifically, the electronic device can reversely solve the gradients corresponding to the weights of the three types of modes of eye movement, EEG and face according to the chain rule, accurately judge the error contribution direction of each type of weight, and clarify the adjustment trend that the weight needs to increase or decrease. Then, the initial weights are iteratively updated in combination with the preset learning rate: appropriately increase the weights for modes with large identification error contribution and insufficient representation, and reduce the weights for redundant interference and over-fitting modes, so as to realize dynamic balance of modal contribution.

[0226] Finally, the updated weights are subjected to constrained normalization processing, which forcibly satisfies 0<w<1 and the sum of the three modal weights is 1, ensuring that the weights always have physical meaning and fusion effectiveness. Finally, the updated weight information corresponding to the three modes is output, and the first round of data-driven weight correction is completed.

[0227] Step d6, updating the initial weight information based on the updated weight information to obtain the modal weight fluctuation amount corresponding to each modality among the time-series aligned eye movement features, the time-series aligned EEG features and the time-series aligned face features.

[0228] Specifically, the electronic device can perform a difference calculation between the initial weights in step d2 and the updated weights in step d5, and calculate the weight fluctuation amount modally: ;in, This represents the unimodal weight fluctuation; positive or negative indicates the direction of weight adjustment, and the absolute value represents the intensity of the adjustment. To update the weights, The initial weights are used. The fluctuations of the three modalities—eye movement, EEG, and face—are calculated sequentially to obtain a complete set of modal fluctuations. This step can precisely distinguish which modalities are in a rapid optimization phase and which have stabilized, providing fine-grained evidence for subsequent convergence determination.

[0229] Step d7: Calculate the average fluctuation amplitude based on the fluctuation of each modal weight.

[0230] Specifically, the electronic device can first take the absolute value of each modal fluctuation to cancel out the mutual interference between positive and negative adjustments and retain the true intensity of change; then, it can calculate the arithmetic mean of the absolute fluctuations of the three types of modes: The average fluctuation amplitude directly reflects the overall update intensity of the multimodal weights: a large amplitude indicates that the weights are still actively seeking optimization and the model has not converged; a smaller amplitude indicates that the weights are approaching a steady state. This indicator compensates for the deficiency that relying solely on loss cannot determine parameter stability.

[0231] Step d8: Based on the global loss value and the average fluctuation amplitude, update the weight information to obtain the target weight information corresponding to the time-aligned eye-tracking features, time-aligned EEG features, and time-aligned face features, respectively.

[0232] Specifically, the electronic device can perform a secondary fine-tuning update on the updated weights obtained from d5 based on the optimization logic of "accuracy first, stability constraint". If the global loss value is too large, it indicates that the current weight ratio fitting accuracy is insufficient. The weights of each modality are then adjusted slightly according to the loss gradient to improve the model's ability to identify differences in contextual awareness. If the average fluctuation amplitude is too large, it indicates that the single-round weight update amplitude is too large and there is a risk of oscillation. The gradient update step size is appropriately weakened to suppress drastic weight fluctuations and ensure smooth convergence of the iteration. This dual-indicator linkage adjustment ensures identification accuracy through the global loss and constrains the stability of weight iteration through fluctuation amplitude, avoiding overfitting and weight oscillation problems caused by a single loss iteration.

[0233] The weights, after the second update, are once again uniformly adjusted for physical constraints: ensuring that the weights for the eye-tracking, EEG, and face modes satisfy 0. <w<1、 This ensures that all weights always have a true physical contribution percentage and comply with multimodal fusion rules.

[0234] The electronic device can be configured with dual termination conditions: a loss convergence threshold and a fluctuation convergence threshold. When the global loss value is lower than a preset accuracy threshold, it indicates that the model's recognition accuracy has reached its optimal level. When the average fluctuation amplitude is lower than a preset stability threshold, it indicates that the weights of each modality are almost no longer changing, and the system is approaching convergence. When both conditions are met simultaneously, the weight iteration optimization is considered complete, and the current weights are no longer updated. After convergence, the target weight information corresponding to time-aligned eye-tracking features, time-aligned EEG features, and time-aligned facial features is output. This set of target weights is the optimal fusion parameter obtained through error-supervised iteration and fluctuation-constrained voltage stabilization. It can accurately adapt to the differentiated feature patterns of EEG, eye-tracking, and facial features in different situational awareness states of pilots, and also has extremely strong stability and generalization ability, providing optimal weight support for subsequent high-precision multimodal comprehensive feature fusion and situational awareness level recognition.

[0235] Step S2053: Based on the target weight information corresponding to the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features, the temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned face features are weighted and fused to obtain the intermodal correlation matrix.

[0236] Specifically, the electronic device can perform weighted calculations on the time-aligned eye-tracking features, time-aligned EEG features, and time-aligned face features based on the target weight information corresponding to each. This yields weighted single-modal features. Then, using an organization format with time-series frames as rows, feature dimensions as columns, and modality as sub-dimensions, the weighted features from the entire sample, across all time series, and across all modalities are integrated to generate an inter-modal correlation matrix.

[0237] Step S2054: Based on the intermodal correlation matrix, obtain the target fusion features.

[0238] Specifically, electronic devices can traverse the intermodal correlation matrix to check for problems such as null values, abnormal extreme values, and dimensional errors, and correct or remove invalid elements to ensure that the matrix data is complete and valid.

[0239] Then, feature integration is performed according to task requirements. Following a time-series row-by-row approach, all column features within each row are concatenated, summed, or mapped to transform the two-dimensional matrix structure into a one-dimensional feature vector sequence. Based on feature distribution patterns, low-discrimination and highly redundant column dimensions are removed, retaining core dimensions that contribute significantly to contextual awareness recognition. The matrix values ​​are normalized / standardized as needed to unify the numerical distribution range. The processed matrix undergoes a structural transformation: for time-series recognition, the time-series structure is preserved, and time-series target fusion features are output; for single-sample classification, samples are aggregated, and sample-level one-dimensional target fusion feature vectors are output. This feature integrates effective information from eye-tracking, EEG, and face modalities, while relying on optimal weights to ensure a reasonable contribution ratio for each modality.

[0240] Step S206: Based on the target fusion features, determine the target situational awareness level corresponding to the target pilot.

[0241] Specifically, step S206 above may include the following steps: Step S2061: Extract environmental perception layer features, physiological state support layer features, and deep cognitive core layer features from the target fusion features.

[0242] Specifically, electronic devices can unify all feature dimensions originating from eye-tracking modalities within the target fusion features into environmental perception layer features, specifically responsible for representing external information reception, visual search, and situational awareness capabilities. Electronic devices can separately extract feature dimensions corresponding to all facial modalities and categorize them into physiological state support layer features, used to represent fatigue levels, facial state, and mental stability. Electronic devices can separately extract feature dimensions corresponding to all EEG modalities and categorize them into deep cognitive core layer features, used to represent attention, cognitive load, and brain processing state. The decomposition process ensures that the three layers of feature dimensions do not overlap, completely cover all fusion features, do not lose information, and do not have redundancy, ultimately resulting in three independent, functionally defined hierarchical feature sets.

[0243] Step S2062: Based on the features of the environmental perception layer, the features of the physiological state support layer, and the features of the deep cognitive core layer, output the real-time situational awareness score, physiological state score, and core cognitive score corresponding to the target pilot, respectively.

[0244] Specifically, based on the characteristics of the environmental perception layer, electronic devices can extract effective information such as visual scanning speed, gaze concentration, and eye movement stability. Through a preset feature-score mapping function, a real-time situational awareness score is calculated, and the score directly reflects the pilot's efficiency in capturing the external situation.

[0245] Based on the characteristics of the physiological support layer, electronic devices can comprehensively calculate the degree of fatigue and mental activity according to the changes in facial features of EAR, MAR, and FAR, and output a physiological state score to characterize whether the pilot's current physical support state is stable.

[0246] Based on the core features of deep cognition, electronic devices can quantify the pilot's thinking and decision-making state and output a core cognitive score based on the attention intensity, cognitive load fluctuations, and brain activation level of EEG characteristics.

[0247] Step S2063: The situational awareness score, physiological state score, and core cognitive score are fused to obtain the situational awareness state score corresponding to the target pilot.

[0248] Specifically, the electronic device can acquire the weights of situational awareness score, physiological state score, and core cognition score. Then, it performs a weighted summation of these three scores to obtain a unique, comprehensive situational awareness state score. The core cognition score has the largest weight, followed by physiological state, with environmental awareness as a supplementary factor, perfectly aligning with real-world cognitive logic. The electronic device can perform interval constraint calibration on the fused total score, fixing the value within a unified scoring range to prevent score deviations and out-of-bounds errors, ensuring consistent scoring standards for each frame and each moment, and finally outputting a real-time comprehensive situational awareness state score in a single temporal dimension.

[0249] Step S2064: Determine the target pilot's target situational awareness level based on the situational awareness state score.

[0250] Specifically, the electronic device can be pre-set with multiple fixed scoring threshold intervals, each interval uniquely corresponding to a situational awareness level. The thresholds are determined through sample statistics and training calibration and remain fixed throughout. The electronic device can match the real-time situational awareness score with each threshold interval one by one to determine the current score's position within that interval. Based on the matching results, it outputs the corresponding target situational awareness level, achieving automatic conversion from quantitative scores to qualitative levels. The final output is a real-time situational awareness level result for the pilot, which can be used for system display, evaluation statistics, and status warnings.

[0251] The pilot situational awareness state determination method provided in this application preprocesses the raw eye movement data to obtain target eye movement data. Abnormal noise and invalid data caused by blinking and equipment vibration are removed, the data format is standardized, and interference from abnormal data is reduced to ensure the reliability of subsequent feature extraction data. Regions of interest (ROIs) for various flight scenarios, such as the main flight area and instrument area, are delineated. Adhering to the logic of real flight observation, the method focuses on the pilot's key observation positions, avoids statistical analysis of invalid areas across the entire screen, and ensures a strong correlation between eye movement features and flight missions. Corresponding subdivided eye movement features are extracted across four dimensions. Indicators are broken down from multiple angles, including region access patterns, fixation timing, pupil physiology, and gaze deviation, to comprehensively quantify visual observation behavior and cover various representational elements of environmental perception. The spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data are input into a preset gradient boosting decision tree model, which outputs importance scores for each eye movement sub-feature included in the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features. The model automatically quantifies the contribution of each sub-eye-tracking indicator to state identification, avoiding the subjectivity of manual feature selection based on experience and accurately distinguishing between effective and redundant features. Based on the importance score corresponding to each eye-tracking sub-feature, the target comprehensive eye-tracking feature is determined from these sub-features. Low-importance redundant sub-features are eliminated, simplifying feature dimensions and reducing computational overhead, while retaining high-contribution key indicators to improve the efficiency and accuracy of subsequent multimodal fusion and state assessment.

[0252] Then, the raw EEG data is preprocessed to obtain the target EEG data. Power frequency interference and EEG / EMG artifacts are removed to optimize data quality and avoid noise interference in subsequent spectrum and source tracing calculations. Power spectral density and energy spectral density are calculated. This allows for the breakdown of energy distribution across frequency bands, quantifying the activity of different EEG waves and reflecting changes in brain excitation and workload in the frequency domain. Global field power is calculated to quantify the overall brain activation intensity and intuitively reflect the overall level of brainpower consumption. Target brain topography maps are selected based on global field power. Invalid artifact maps are removed, retaining valid brain topography maps that truly reflect changes in brain activation, reducing interference from invalid samples. Iterative clustering of the brain topography maps generates multiple initial clustering results, automatically grouping them according to the spatial distribution of EEG data without requiring manual delineation of classification boundaries, achieving unsupervised grouping of EEG states. Global explained variance and cross-validation index are calculated for each clustering result. Quantitative indicators are used to measure the quality of clustering, eliminating the drawbacks of subjective human evaluation. Iterative optimization is performed until convergence, using the clusters with the maximum variance and minimum cross-validation index to determine the optimal number of clusters and clustering results. Balancing data information retention with model generalization ability, this approach avoids over- or under-segmentation in clustering, improving grouping rationality. It compares the similarity between clustering results and preset state categories to match corresponding target state categories. Unlabeled clustering results are mapped to actual cognitive states, achieving automatic correspondence between EEG atlases and cognitive types. Based on classification results, temporal features of clustering are extracted, recording the switching patterns of different brain states over time, supplementing temporal dimension information, and reflecting the dynamic changes in cognition. The temporal features of each cluster are fused to obtain the target EEG state features. Integrating spatial distribution and temporal evolution information, it comprehensively represents the dynamic changes in cognitive load in the brain, improving feature representation capabilities. Then, multiple independent voxel units are obtained by partitioning using a standard human brain template. A unified brain spatial partitioning scale is established, refining brain region computational units and providing a regular computational grid for subsequent inverse problem solving. A standardized digital brain model is built by combining voxel coordinates and cortical functional areas. Spatial location and physiological functional partitions are bound together, unifying the mapping benchmark for individual brain data and eliminating computational biases caused by individual brain morphological differences. The target EEG data is integrated into a model, and sLORETA is used to solve for the voxel current density of the whole brain. Intracranial neural activity is inverted from scalp EEG, overcoming the limitation of raw EEG which only shows surface signals, and obtaining true intracranial activation information. Current density data corresponding to Broadman partitions are screened. Focusing on cognitively relevant brain regions, irrelevant voxel data is removed, and effective activation areas related to contextual awareness are anchored. Three types of indicators are extracted: average current density, peak activation, and activation range. Brain region activity is quantified from multiple perspectives, including activation strength, activation extremes, and diffusion range, enriching the dimensions of source-tracing features. The three types of indicators are summarized to generate target EEG source-tracing features. This achieves the quantitative implementation of brain source spatial information, accurately represents deep cognitive activity, and improves the intrinsic physiological basis of comprehensive EEG features.

[0253] Next, the original face image is preprocessed to obtain the target face image. Illumination, noise, and image distortion interference are removed to optimize image quality and ensure the accuracy of subsequent face localization and key point extraction. The effective region of interest (ROI) of the face is identified and defined. This eliminates invalid background areas, narrows the computational range, and reduces interference from irrelevant pixels on feature extraction. Multiple standard facial key points are detected and output. Based on these points, the positions of the eyes, mouth, and cheek contours are located, providing coordinate basis for subsequent quantitative geometric indicators. Core feature points are selected, redundant key points are discarded, and key points related to fatigue and mental state are retained to simplify computation. The EAR, MAR, and FAR ratio features are calculated using these points. The facial morphology is converted into quantitative values, intuitively reflecting physiological changes such as drowsiness and facial tension. The three ratios are fused to obtain the comprehensive facial features of the target. Multi-dimensional summarization of facial physiological representations fully reflects the pilot's real-time physiological support status and improves the multimodal data source.

[0254] Temporal alignment processing is performed on the target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features to obtain temporally aligned eye-tracking features, temporally aligned EEG features, and temporally aligned facial features. This unifies the timestamps of the three types of data, eliminates temporal misalignment caused by different acquisition frame rates, and ensures that eye-tracking, EEG, and facial data are matched one-to-one at the same time. The three types of aligned features are fed into the weight determination model. A unified input data source is used, and the model achieves data-driven weight initialization, eliminating the subjectivity of manually setting weights. The model outputs initial weights for each modality. A set of baseline weights is quickly generated, providing a basic matching scheme for the initial feature fusion. The initial weights are weighted and fused, and forward inference is performed to output the initial identification results. This demonstrates the actual effect of the weights, obtaining prediction results that can be used for error calculation and quantifying the identification performance of the current weights. The global loss is calculated by comparing the prediction with the ground truth, quantifying the overall identification bias from the full sample dimension, and providing a quantitative supervision indicator for weight iterative optimization. The updated weights are obtained by back-correction based on the global loss. To minimize identification errors, the proportion of each modality is adaptively adjusted to optimize the rationality of modal contribution allocation. The fluctuation of each modal weight is calculated by comparing the old and new weights. The change amplitude of each modal weight in a single iteration is quantified to distinguish the optimization sensitivity of each modality. The average fluctuation amplitude is calculated from the single-modal fluctuation. This macroscopically characterizes the overall change in the weight set and is used to evaluate the convergence and stability of the weights. The target weights are obtained through secondary optimization combining loss and fluctuation amplitude. Balancing identification accuracy and weight stability, and avoiding weight oscillations caused by solely relying on loss updates, the optimal steady-state weights are ultimately obtained. Modal weights are adaptively allocated based on the identification effect, unlike fixed weights, to adapt to the dynamic changes in modal contributions under different flight conditions. Based on the target weight information corresponding to time-aligned eye-tracking features, time-aligned EEG features, and time-aligned face features, a weighted fusion of these features is performed to obtain the inter-modal correlation matrix. Effective modal information is amplified based on contribution proportion, and redundant interference is suppressed. Modal correlations and temporal correlations are retained in matrix form, resulting in a well-structured data structure. Target fusion features are obtained based on the inter-modal correlation matrix. Effective matrix information is refined, and redundant dimensions are eliminated to form an integrated feature that takes into account perception, physiology, and cognition, facilitating subsequent hierarchical extraction and scoring calculation.

[0255] Finally, three layers of features are extracted from the target fusion features. Features are broken down according to perception, physiological function, and cognitive function, achieving hierarchical isolation of information and avoiding interference between different dimensions, facilitating itemized evaluation. Three independent scores are output hierarchically, providing dimensional quantitative evaluation and allowing for the identification of specific causes of pilots' insufficient perception, physiological fatigue, or cognitive decline, making the evaluation traceable. The three scores are fused to generate a comprehensive score, complementing the strengths and weaknesses of each dimension, avoiding the shortcomings of single-indicator one-sided evaluation, and quantifying the overall situational awareness level. The situational awareness level is determined based on the comprehensive score; continuous scores are converted into graded results, making the results intuitive and easy to understand, facilitating status warning and control by airborne systems.

[0256] This embodiment also provides a pilot situational awareness state determination device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0257] This embodiment provides a device for determining a pilot's situational awareness state, such as... Figure 7 As shown, it includes: The acquisition module 301 is used to acquire the raw eye movement data, raw electroencephalogram data and raw facial image of the target pilot. The eye movement feature extraction module 302 is used to extract features from the raw eye movement data to obtain the target comprehensive eye movement features; The EEG feature extraction module 303 is used to extract features from the raw EEG data to obtain the target comprehensive EEG features; The face feature extraction module 304 is used to extract features from the original face image to obtain the target comprehensive face features; The fusion module 305 is used to fuse the target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features to generate target fusion features; The determination module 306 is used to determine the target situational awareness level of the target pilot based on the target fusion features.

[0258] The pilot situational awareness state determination device provided in this embodiment of the invention can execute the pilot situational awareness state determination method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.

[0259] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0260] The following is a detailed reference. Figure 8 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 01, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 02 or a program loaded from a memory 08 into a random access memory (RAM) 03. The RAM 03 also stores various programs and data required for the operation of the electronic device. The processor 01, ROM 02, and RAM 03 are interconnected via a bus 04. An input / output (I / O) interface 05 is also connected to the bus 04.

[0261] Typically, the following devices can be connected to I / O interface 05: input devices 06 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 07 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 08 including, for example, magnetic tapes, hard disks, etc.; and communication devices 09. Communication device 09 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0262] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 09, or installed from memory 08, or installed from ROM 02. When the computer program is executed by processor 01, it performs the functions defined in the pilot situational awareness state determination method of the embodiments of the present invention.

[0263] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0264] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the pilot situational awareness state determination method shown in the above embodiments is implemented.

[0265] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0266] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for determining a pilot's situational awareness state, characterized in that, The method includes: Acquire raw eye-tracking data, raw electroencephalogram data, and raw facial images of the target pilot; Feature extraction is performed on the raw eye movement data to obtain the target comprehensive eye movement features; Feature extraction is performed on the raw EEG data to obtain the target comprehensive EEG features; Feature extraction is performed on the original face image to obtain the target comprehensive face features; The target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features are fused to generate target fusion features; Based on the target fusion features, the target situational awareness level of the target pilot is determined; The step of extracting features from the original eye-tracking data to obtain the target comprehensive eye-tracking features includes: The raw eye movement data is preprocessed to obtain the target eye movement data; Obtain multiple regions of interest (ROIs) based on the flight scenario; the ROI is at least one of the following: main flight area, instrument area, left-side environment area, and right-side environment area. Feature extraction is performed on the target eye-tracking data to obtain spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye-tracking data. The spatial allocation dimension features include the proportion of access time corresponding to each of the target regions of interest and the number of times each of the target regions of interest is accessed. The temporal fixation dimension features include the total fixation time, average fixation duration, and number of fixations corresponding to each of the target regions of interest. The physiological load dimension features include the average pupil diameter, maximum pupil diameter, and minimum pupil diameter. The gaze deviation dimension features include the average horizontal distance of the gaze, the average vertical distance, and the average absolute distance. Based on the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data, the comprehensive eye movement features of the target are obtained.

2. The method according to claim 1, characterized in that, The comprehensive eye movement features of the target are obtained based on the spatial allocation dimension features, temporal fixation dimension features, physiological load dimension features, and gaze deviation dimension features corresponding to the target eye movement data, including: The spatial allocation dimension feature, temporal fixation dimension feature, physiological load dimension feature, and gaze deviation dimension feature corresponding to the target eye movement data are input into a preset gradient boosting decision tree model, and the importance score corresponding to each eye movement sub-feature included in the spatial allocation dimension feature, temporal fixation dimension feature, physiological load dimension feature, and gaze deviation dimension feature is output. Based on the importance score corresponding to each of the described eye movement sub-features, the target comprehensive eye movement feature is determined from each of the described eye movement sub-features.

3. The method according to claim 1, characterized in that, The step of extracting features from the raw EEG data to obtain the target comprehensive EEG features includes: The raw EEG data is preprocessed to obtain the target EEG data; Calculate the power spectral density and energy spectral density corresponding to the target EEG data; Calculate the global field power corresponding to the target EEG data; Based on the global field power, the target EEG state characteristics corresponding to the target EEG data are determined; Source feature extraction is performed on the target EEG data to obtain the target EEG source feature corresponding to the target EEG data; The power spectral density, the energy spectral density, the target EEG state features, and the target EEG source features are fused to obtain the target comprehensive EEG features.

4. The method according to claim 3, characterized in that, The step of determining the target EEG state features corresponding to the target EEG data based on the global field power includes: Based on the global field power, the target brain topography map corresponding to the target EEG data is obtained by filtering. Iterative clustering of the target brain topography map yields multiple initial clustering results; For each initial clustering result, calculate the global explained variance and cross-validation criterion corresponding to the initial clustering result; Based on the principle of maximizing the global explained variance and minimizing the cross-validation criteria, the initial clustering results are updated until the clustering converges, and the optimal number of clusters and the target clustering results corresponding to the optimal number of clusters are determined. Based on the similarity between each target clustering result and the preset state category, the target state category corresponding to each target clustering result is determined; Based on the target state category corresponding to each target clustering result, extract the clustering time series features corresponding to the target clustering results; The clustered temporal features are fused to obtain the target EEG state features.

5. The method according to claim 3, characterized in that, The step of extracting source features from the target EEG data to obtain the target EEG source features corresponding to the target EEG data includes: Based on the standard human brain template, the cerebral cortex is finely divided and evenly divided into multiple independent voxel units; A standardized digital brain model is established based on the correspondence between each independent voxel unit and the fixed spatial coordinates of the brain and the functional areas of the cortex. The target EEG data is input into the standardized digital brain model, and the spatial distribution of whole brain current density corresponding to all independent voxel units is calculated and output point by point based on the sLORETA low-resolution electromagnetic tomography algorithm. From the spatial distribution of the whole brain current density, the spatial distribution of the partition current density corresponding to the Brodman partition is selected. Feature extraction is performed on the spatial distribution of the current density in the aforementioned regions to obtain the average current density, maximum activation intensity, and spatial activation range of the brain regions. The target EEG source characteristics are obtained based on the average current density of the brain region, the maximum activation intensity, and the spatial activation range.

6. The method according to claim 1, characterized in that, The step of extracting features from the original face image to obtain the target comprehensive face features includes: The original face image is preprocessed to obtain the target face image; The target face image is identified to determine the effective region of interest (ROI). The effective region of interest of the face is identified, and multiple standard facial key points are determined; Multiple core feature points were selected from the aforementioned multiple standard facial key points; Based on each of the aforementioned core feature points, calculate the core features of the eye aspect ratio, the core features of the mouth aspect ratio, and the core features of the face contour aspect ratio. The core features of the eye aspect ratio, the core features of the mouth aspect ratio, and the core features of the facial contour aspect ratio are fused to obtain the target comprehensive facial features.

7. The method according to claim 1, characterized in that, The process of fusing the target's comprehensive eye-tracking features, comprehensive EEG features, and comprehensive facial features to generate target fusion features includes: The target's integrated eye movement features, integrated EEG features, and integrated facial features are subjected to temporal alignment processing to obtain temporally aligned eye movement features, temporally aligned EEG features, and temporally aligned facial features. The temporally aligned eye-tracking features, the temporally aligned EEG features, and the temporally aligned face features are input into a preset weight determination model, and the target weight information corresponding to the temporally aligned eye-tracking features, the temporally aligned EEG features, and the temporally aligned face features is output. Based on the target weight information corresponding to the time-aligned eye-tracking features, the time-aligned EEG features, and the time-aligned face features, the time-aligned eye-tracking features, the time-aligned EEG features, and the time-aligned face features are weighted and fused to obtain the intermodal correlation matrix. The target fusion features are obtained based on the intermodal correlation matrix.

8. The method according to claim 7, characterized in that, The step of inputting the temporally aligned eye-tracking features, the temporally aligned EEG features, and the temporally aligned face features into a preset weight determination model, and outputting the target weight information corresponding to the temporally aligned eye-tracking features, the temporally aligned EEG features, and the temporally aligned face features respectively, includes: The time-aligned eye-tracking features, the time-aligned EEG features, and the time-aligned face features are input into a preset weight determination model; The preset weight determination model generates initial weight information corresponding to the time-aligned eye-tracking features, the time-aligned EEG features, and the time-aligned face features, respectively. Based on the initial weight information, the time-aligned eye-tracking features, the time-aligned EEG features, and the time-aligned face features are weighted and fused, and forward inference is completed to output the initial situational awareness level identification result corresponding to the initial weight information. Calculate the global loss value between the initial situational awareness level identification result and the real cognitive label in the preset weight determination model; Based on the global loss value, the initial weight information is updated to obtain the updated weight information corresponding to the temporally aligned eye-tracking feature, the temporally aligned EEG feature, and the temporally aligned face feature, respectively. Based on the updated weight information, the initial weight information is updated to obtain the modality weight fluctuation amount corresponding to each modality in the temporally aligned eye-tracking feature, the temporally aligned EEG feature, and the temporally aligned face feature. Calculate the average fluctuation amplitude based on the fluctuation of each modal weight; Based on the global loss value and the average fluctuation amplitude, the updated weight information is updated to obtain the target weight information corresponding to the temporally aligned eye-tracking feature, the temporally aligned EEG feature, and the temporally aligned face feature, respectively.

9. The method according to claim 1, characterized in that, The step of determining the target situational awareness level of the target pilot based on the target fusion features includes: From the target fusion features, extract the features of the environmental perception layer, the features of the physiological state support layer, and the features of the deep cognition core layer; Based on the features of the environmental perception layer, the features of the physiological state support layer, and the features of the deep cognitive core layer, the real-time situational awareness score, physiological state score, and core cognitive score corresponding to the target pilot are output respectively. The situational awareness score, the physiological state score, and the core cognitive score are fused to obtain the situational awareness state score corresponding to the target pilot. Based on the situational awareness score, the target pilot's situational awareness level is determined.

Citation Information

Patent Citations

  • Pilot workload identification method and device

    CN120732423A

  • Flight trainee pressure assessment system based on fusion of heart rate variability and multi-modal data

    CN121512526A