Cognitive load recognition model training method

CN122508509BActive Publication Date: 2026-09-29NANJING VOCATIONAL UNIV OF IND TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611001071.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29
Estimated Expiration
2046-07-07

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明提供了一种认知负荷辨识模型训练方法,以解决训练一种可以准确对飞行员负荷进行辨识的模型的问题

Benefits of technology

[0020]本实施例提供的认知负荷辨识模型训练方法,获取训练飞行员对应的初始训练数据集。能够全面反映飞行员在不同飞行场景下的生理状态与认知负荷的关联关系,为后续模型训练提供了丰富的原始依据。然后,基于预设数据扩充模型对所述初始训练数据集进行扩充,生成目标训练数据集。通过数据扩充,可生成大量符合生理规律的模拟数据,补充稀缺样本,平衡不同负荷等级、不同飞行阶段的样本比例,避免模型因数据偏向性而产生过拟合。数据扩充不仅增加样本数量,还能通过特征变换生成更多样化的生理特征组合,帮助模型挖掘更全面的认知负荷关联模式,提升泛化能力。最后,基于所述目标训练数据集,对初始认知负荷辨识网络进行训练,得到目标认知负荷辨识模型。目标训练数据集的丰富性和均衡性,为初始认知负荷辨识网络提供了更充分的学习素材。网络能通过训练不断优化参数,更精准地捕捉认知负荷与生理特征、飞行场景的内在联系,从而提高对认知负荷等级的预测准确性,保证了训练得到的目标认知负荷辨识模型的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122508509B_ABST
    Figure CN122508509B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer, in particular to a cognitive load recognition model training method. An initial training data set corresponding to a training pilot is obtained, which includes initial electrocardiogram data, initial respiration data, initial eye movement data, training flight phase and training cognitive load level corresponding to the training pilot of the training pilot in the initial training data set; the initial training data set is expanded based on a preset data expansion model to generate a target training data set; the initial cognitive load recognition network is trained based on the target training data set to obtain a target cognitive load recognition model. The accuracy of the target cognitive load recognition model obtained by training is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically to a method for training a cognitive load identification model. Background Technology

[0002] In a human-machine interface closed-loop system, pilots constantly engage in cognitive activities such as information perception, processing, and decision-making while performing flight missions. With the rapid development of civil aviation and artificial intelligence technologies, modern aircraft cockpit systems are highly information-intensive, making human-machine systems increasingly complex and placing higher demands on pilots' cognitive abilities. The assessment of cognitive load began in the early 1970s, and current methods for assessing cognitive load can be divided into subjective and objective assessments. Subjective assessment methods primarily use questionnaires or scales to reflect the pilot's level of tension, stress, and mission difficulty during mission execution. Commonly used subjective rating scales include the NASA-TXL (National Aeronautics and Space Administration Traffic Load Index), the SWAT (Subject Workload Assessment Technique), and the WP (Workload Profile Index Ratings). Objective assessment methods primarily reflect the pilot's cognitive load based on objective data sources such as mission performance or physiological signals.

[0003] Currently, research on pilot cognitive load has yielded many results, but some problems still exist, such as research subjects not being pilots or flight trainees, experimental tasks being unrelated to flight tasks, high data acquisition costs, incomplete characteristic indicators, and a single evaluation method.

[0004] Therefore, training a model that can accurately identify pilot load has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the present invention provides a cognitive load identification model training method to solve the problem of training a model that can accurately identify pilot load.

[0006] In a first aspect, the present invention provides a method for training a cognitive load identification model, the method comprising: Obtain the initial training dataset corresponding to the trainees. The initial training dataset includes the initial electrocardiogram data, initial respiratory data, initial eye movement data, training flight phase, and training cognitive load level corresponding to the trainees. Expand the initial training dataset based on the preset data augmentation model to generate the target training dataset. Train the initial cognitive load identification network based on the target training dataset to obtain the target cognitive load identification model.

[0007] In one optional implementation, the initial training dataset is expanded based on a preset data augmentation model to generate a target training dataset, including: extracting features from the initial electrocardiogram (ECG) data in the initial training dataset to obtain total initial ECG features corresponding to the initial ECG data; the total initial ECG features include initial ECG time-domain features, initial ECG frequency-domain features, and initial ECG nonlinear features; extracting features from the initial respiratory data in the initial training dataset to obtain total initial respiratory features corresponding to the initial respiratory data; the total initial respiratory features include initial respiratory time-domain features and initial respiratory frequency-domain features; extracting features from the initial eye-tracking data in the initial training dataset to obtain initial eye-tracking features; and then expanding the total initial ECG data into a target training dataset. The initial physiological features are generated by fusing electrical features, total initial respiratory features, and initial eye-tracking features. Task complexity features are generated based on operational elements at each stage of the five-sided flight. Initial conditional features are generated by fusing training flight stages, training cognitive load levels, and task complexity features. Initial fusion features are generated by fusing initial physiological features and initial conditional features. These initial fusion features are then input into a pre-defined generation network within a pre-defined data augmentation model to generate noise data corresponding to the initial training dataset. The noise data is then input into a pre-defined discrimination network within the pre-defined data augmentation model to discriminate the noise data. Based on the discrimination results, the initial training dataset is augmented to generate the target training dataset.

[0008] In one optional implementation, the initial fusion features are input into a preset generator network based on a preset data augmentation model to generate noise data corresponding to the initial training dataset. This includes: inputting the initial fusion features into a first hidden layer in the preset generator network to expand the initial fusion features and generate a first hidden feature; inputting the first hidden feature into a second hidden layer in the preset generator network to scale the first hidden feature and generate a second hidden feature; inputting the second hidden feature into an ECG constraint branch, an eye-tracking constraint branch, and a respiratory constraint branch to perform modal constraint processing on the second hidden feature and generate a target constraint feature; and activating the second hidden feature to generate noise data.

[0009] In one optional implementation, initial fused features are input into a first hidden layer of a preset generator network, and the initial fused features are expanded to generate first hidden features, including: determining the training flight stage and training cognitive load level from the initial fused features; determining the target expanded dimension corresponding to the initial fused features from a preset mapping table according to the training flight stage and training cognitive load level; generating a first initial weight matrix according to the initial dimension corresponding to the initial fused features and the target expanded dimension; correcting the first initial weight matrix according to the training flight stage and training cognitive load level to obtain a first target weight matrix; expanding the initial fused features based on the first target weight matrix to generate a target expanded dimension feature vector; activating the target expanded dimension feature vector based on a preset activation function to output the first hidden features.

[0010] In one optional implementation, the first hidden feature is input into the second hidden layer of a preset generator network, and the first hidden feature is scaled to generate the second hidden feature, including: calculating the first mutual information value between each initial sub-fusion feature in the initial fusion feature; constructing a modal correlation matrix based on each first mutual information value; labeling each first hidden sub-feature according to the mapping relationship between each first hidden sub-feature in the first hidden feature and each initial sub-fusion feature in the initial fusion feature; and calculating the mean value corresponding to each first hidden sub-feature dominated by ECG based on the labeling results. The mean vector of each first hidden feature dominated by respiration is used as the first cluster center; the mean vector of each first hidden feature dominated by eye movement is used as the second cluster center; the mean vector of each first hidden feature dominated by eye movement is used as the third cluster center; the second mutual information value between each first hidden feature is calculated based on the modal correlation matrix; the first hidden features are clustered according to the second mutual information value to obtain the ECG cluster group, the respiration cluster group, the eye movement cluster group, and the conditional cluster group; the second hidden feature is output according to the ECG cluster group, the respiration cluster group, the eye movement cluster group, and the conditional cluster group.

[0011] In one optional implementation, outputting a second hidden feature based on the ECG cluster group, respiratory cluster group, eye-tracking cluster group, and conditional cluster group includes: calculating the average correlation degree among the ECG cluster group, respiratory cluster group, eye-tracking cluster group, and conditional cluster group based on the modal correlation matrix; determining the connection weights among the ECG cluster group, respiratory cluster group, eye-tracking cluster group, and conditional cluster group based on each average correlation degree; generating a second initial weight matrix based on the initial dimension corresponding to the initial fusion feature and the target extended dimension; and arranging the second initial weight matrix according to the groups. The system divides the region into intra-group and inter-group regions. Based on the connection weights between the ECG cluster, respiratory cluster, eye-tracking cluster, and the respiratory cluster, the first target weight value corresponding to the inter-group region in the second initial weight matrix is ​​determined. Based on the criticality of each cluster sub-feature in the ECG cluster, respiratory cluster, eye-tracking cluster, and conditional cluster, the second target weight value corresponding to the intra-group region is determined. Based on the first and second target weight values, a second target weight matrix is ​​generated. The first hidden feature is multiplied by the second target weight matrix to obtain the second hidden feature.

[0012] In one optional implementation, noise data is input into a preset discriminant network in a preset data augmentation model to discriminate the noise data, including: calculating the load change rate based on initial ECG data and initial eye movement data; fusing the noise data, training flight phase, training cognitive load level, and load change rate to generate target fusion features; compressing the target fusion features based on a first branch to obtain first branch features; compressing the target fusion features based on a second branch to obtain second branch features; and dividing and labeling the target fusion features into four specific groups according to physiological modalities and functional attributes; the four specific groups are the ECG group, ... The system includes a respiratory group, an eye-tracking group, and a load change rate label group. The first and second branch features are validated respectively, yielding ECG, respiratory, and eye-tracking scores for each feature. The first and second branch features, along with their respective ECG, respiratory, and eye-tracking scores, are fused to obtain a discriminative fusion feature. This discriminative fusion feature is then input into the authenticity discrimination branch, outputting the probability that noisy data is real data. The discriminative fusion feature is also input into the load level discrimination branch, outputting the load level probability corresponding to the noisy data. Finally, the noisy data is discriminated based on a preset loss function.

[0013] In one optional implementation, the noise data is judged based on a preset loss function, including: calculating authenticity loss based on the probability that the noise data is real data; calculating load level loss based on the probability of the load level corresponding to the noise data; calculating compliance loss based on ECG dimension score, respiratory dimension score, and eye movement dimension score; generating a preset loss function based on authenticity loss, load level loss, and compliance loss; and judging the noise data based on the preset loss function.

[0014] In one optional implementation, an initial cognitive load identification network is trained based on a target training dataset to obtain a target cognitive load identification model. This includes: extracting features from the target training dataset to generate training fusion features; inputting the training fusion features into the input layer of the initial cognitive load identification network; adjusting the weight information corresponding to each sub-training fusion feature in the training fusion features according to the target flight stage in the target training dataset to obtain weighted fusion features; inputting the weighted fusion features into the third hidden layer of the initial cognitive load identification network; extracting features from the weighted fusion features in the third hidden layer to obtain third hidden features; inputting the third hidden features into the output layer of the initial cognitive load identification network; outputting the predicted cognitive load level corresponding to the target training dataset based on the third hidden features; and training the initial cognitive load identification network based on the predicted cognitive load level to obtain the target cognitive load identification model.

[0015] In one optional implementation, the third hidden layer extracts features from the weighted fusion features to obtain third hidden features, including: the residual network in the third hidden layer extracts features from the weighted fusion features and outputs target residual features; the ECG-respiratory cross feature pairs and the eye-tracking-ECG cross feature pairs are selected from the target residual features; the third mutual information value corresponding to the ECG-respiratory cross feature pairs and the fourth mutual information value corresponding to the eye-tracking-ECG cross feature pairs are calculated; the attention weights corresponding to the ECG-respiratory cross feature pairs and the eye-tracking-ECG cross feature pairs are calculated based on the third and fourth mutual information values; and the third hidden features are generated based on the attention weights corresponding to the ECG-respiratory cross feature pairs and the eye-tracking-ECG cross feature pairs.

[0016] In one optional implementation, the residual network includes a first residual network, a second residual network, and a third residual network. The residual network in the third hidden layer extracts features from the weight fusion features and outputs target residual features, including: the first residual network in the third hidden layer extracts multimodal features from the weight fusion features to obtain multimodal features; the second residual network in the third hidden layer compresses the multimodal features to obtain compressed features, and performs a convolution operation on the compressed features to obtain convolutional features; and calculates the corresponding features of each modality in the weight fusion features. The signal-to-noise ratio (SNR) of each modality feature is calculated. Based on the SNR of each modality feature, the compressed features and convolutional features are subjected to residual superposition processing to generate superimposed features. The superimposed features are divided into ECG time domain subgroup, ECG frequency domain subgroup, respiratory rhythm subgroup, respiratory depth subgroup, eye movement pupil subgroup, and eye movement fixation subgroup based on the third residual network in the third hidden layer. Each subgroup is subjected to convolution processing, and the convolution results are compressed by max pooling and regularization to obtain the preserved features corresponding to each subgroup. The preserved features corresponding to each subgroup are fused to obtain the target residual features.

[0017] In one optional implementation, the output layer outputs the predicted cognitive load level corresponding to the target training dataset based on the third hidden feature, including: extracting features from the third hidden feature to obtain load-sensitive features and stability features; inputting the load-sensitive features and stability features into the main branch in the output layer; the main branch outputs the first initial probability of each predicted cognitive load level corresponding to the load-sensitive features and stability features through a fully connected network; and outputting the predicted cognitive load level corresponding to the target training dataset according to each first initial probability.

[0018] In one optional implementation, the predicted cognitive load level corresponding to the target training dataset is output based on each first initial probability, including: inputting load-sensitive features and stability features into an auxiliary branch in the output layer; the auxiliary branch extracting dynamic features from the load-sensitive features and stability features; calculating the change trend based on the dynamic features; outputting second initial probabilities of load increase, load stabilization, and load decrease corresponding to the load-sensitive features and stability features based on the change trend; and outputting the predicted cognitive load level corresponding to the target training dataset based on the first initial probability and the second initial probability.

[0019] In one optional implementation, the initial cognitive load identification network is trained based on the predicted cognitive load level to obtain the target cognitive load identification model, including: obtaining the current training round number; substituting the predicted cognitive load level into the target loss function corresponding to the current training round number based on the current training round number; and training the initial cognitive load identification network based on the target loss function to obtain the target cognitive load identification model.

[0020] The cognitive load identification model training method provided in this embodiment obtains an initial training dataset corresponding to the pilots. This dataset comprehensively reflects the correlation between the pilots' physiological states and cognitive load under different flight scenarios, providing rich original data for subsequent model training. Then, the initial training dataset is expanded based on a preset data augmentation model to generate a target training dataset. Data augmentation generates a large amount of simulated data that conforms to physiological laws, supplements scarce samples, balances the sample ratios of different load levels and flight stages, and avoids overfitting due to data bias. Data augmentation not only increases the number of samples but also generates more diverse combinations of physiological features through feature transformation, helping the model to discover more comprehensive cognitive load correlation patterns and improve generalization ability. Finally, the initial cognitive load identification network is trained based on the target training dataset to obtain the target cognitive load identification model. The richness and balance of the target training dataset provide more sufficient learning materials for the initial cognitive load identification network. The network can continuously optimize its parameters through training, more accurately capturing the intrinsic relationship between cognitive load and physiological characteristics and flight scenarios, thereby improving the accuracy of predicting cognitive load levels and ensuring the accuracy of the target cognitive load identification model obtained through training. Attached Figure Description

[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating the cognitive load identification model training method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating another cognitive load identification model training method according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating another cognitive load identification model training method according to an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] According to an embodiment of the present invention, a method for training a cognitive load identification model is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0025] This embodiment provides a method for training a cognitive load identification model, which can be used in electronic devices. Figure 1 This is a flowchart of a cognitive load identification model training method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain the initial training dataset corresponding to the pilots being trained.

[0026] The initial training dataset includes initial electrocardiogram data, initial respiratory data, initial eye movement data, training flight phases, and training cognitive load levels corresponding to the training pilots.

[0027] Specifically, electronic devices can collect various types of data from pilots during flight training, including: physiological data: initial electrocardiogram data (such as heart rate and heart rate variability), initial respiratory data (such as respiratory rate and depth), and initial eye movement data (such as pupil diameter and fixation point); scenario and label data: training flight phases (such as takeoff, cruise, and landing), and training cognitive load levels (such as load assessment results from 0 to 5). These data collectively constitute the basic dataset for model training, reflecting the pilot's cognitive load status in different scenarios.

[0028] Step S102: Expand the initial training dataset based on the preset data augmentation model to generate the target training dataset.

[0029] Specifically, electronic devices use a pre-defined data augmentation model to process the initial training dataset, thereby expanding the data scale and enriching data diversity by generating noise data that conforms to physiological laws, supplementing scarce scenario samples such as high load / low load, and strengthening cross-modal feature association.

[0030] The resulting target training dataset can more comprehensively cover the cognitive load scenarios that pilots may encounter, reducing model bias caused by insufficient or imbalanced data.

[0031] This step will be explained in detail below.

[0032] Step S103: Based on the target training dataset, train the initial cognitive load identification network to obtain the target cognitive load identification model.

[0033] Specifically, the electronic device inputs the target training dataset into the initial cognitive load identification network, and continuously adjusts the network parameters through mechanisms such as feature extraction (e.g., multimodal physiological feature fusion, cross-modal correlation analysis), dynamic prediction (combining the current load level and the trend of change), and phased loss function optimization.

[0034] After training, the initial network was gradually optimized, and finally a target cognitive load identification model that can accurately identify the cognitive load level of pilots was obtained, which can be used for cognitive load assessment of pilots in real-world scenarios.

[0035] This step will be explained in detail below.

[0036] The cognitive load identification model training method provided in this embodiment obtains an initial training dataset corresponding to the pilots. This dataset comprehensively reflects the correlation between the pilots' physiological states and cognitive load under different flight scenarios, providing rich original data for subsequent model training. Then, the initial training dataset is expanded based on a preset data augmentation model to generate a target training dataset. Data augmentation generates a large amount of simulated data that conforms to physiological laws, supplements scarce samples, balances the sample ratios of different load levels and flight stages, and avoids overfitting due to data bias. Data augmentation not only increases the number of samples but also generates more diverse combinations of physiological features through feature transformation, helping the model to discover more comprehensive cognitive load correlation patterns and improve generalization ability. Finally, the initial cognitive load identification network is trained based on the target training dataset to obtain the target cognitive load identification model. The richness and balance of the target training dataset provide more sufficient learning materials for the initial cognitive load identification network. The network can continuously optimize its parameters through training, more accurately capturing the intrinsic relationship between cognitive load and physiological characteristics and flight scenarios, thereby improving the accuracy of predicting cognitive load levels and ensuring the accuracy of the target cognitive load identification model obtained through training.

[0037] This embodiment provides a method for training a cognitive load identification model, which can be used in electronic devices. Figure 2 This is a flowchart of a cognitive load identification model training method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the initial training dataset corresponding to the pilots being trained.

[0038] The initial training dataset includes initial electrocardiogram data, initial respiratory data, initial eye movement data, training flight phases, and training cognitive load levels corresponding to the training pilots.

[0039] Please refer to the above description of step S101 for details on this step, which will not be repeated here.

[0040] Step S202: Expand the initial training dataset based on the preset data augmentation model to generate the target training dataset.

[0041] Specifically, step S202 above may include the following steps: Step S2021: Extract features from the initial ECG data in the initial training dataset to obtain the total initial ECG features corresponding to the initial ECG data.

[0042] The total initial ECG characteristics include initial ECG time-domain characteristics, initial ECG frequency-domain characteristics, and initial ECG nonlinear characteristics.

[0043] Specifically, electronic devices can extract initial ECG time-domain features, initial ECG frequency-domain features, and initial ECG nonlinear features from initial ECG data.

[0044] For example, as shown in Table 1, the extracted initial ECG time-domain features, initial ECG frequency-domain features, and initial ECG nonlinear features are shown.

[0045] Table 1 Initial ECG time-domain characteristics, initial ECG frequency-domain characteristics, and initial ECG nonlinear characteristics

[0046] For the initial ECG time-domain characteristics (corresponding to items 1-4 in Table 1): the electronic device can extract the RR interval sequence from the initial ECG data and calculate: average heart rate (HR), count the number of R wave peaks per unit time, and convert it to bpm; RR interval standard deviation (SDNN): calculate the standard deviation of continuous RR interval values, in ms; RR interval standard deviation mean (SDANN): calculate the standard deviation of the RR sequence segment and take the mean; root mean square of adjacent RR interval differences (RMSSD): take the square root of the mean of the squares of the differences between adjacent RR intervals.

[0047] Regarding the initial ECG frequency domain characteristics (corresponding to items 5-8 in Table 1): Electronic devices can perform Fourier transform on the RR interval sequence to decompose the following: High-frequency power (HF): power value in the 0.15-0.4Hz frequency band, in ms²; High-frequency power percentage (HFPowerPercent): the proportion of HF to the total power; Low-frequency power (LF): power value in the 0.04-0.15Hz frequency band, in ms²; Low-frequency power percentage (LFPowerPercent): the proportion of LF to the total power.

[0048] For the initial ECG nonlinear characteristics (corresponding to items 9-12 in Table 1): the electronic device can draw a Poincaré scatter plot and extract: the standard deviation of the major axis of the ellipse (SD1) and the standard deviation of the minor axis (SD2); the number of points in the first quadrant (A++) and the number of points in the third quadrant (B--) of the difference scatter plot, expressed as a percentage.

[0049] Step S2022: Extract features from the initial respiratory data in the initial training dataset to obtain the total initial respiratory features corresponding to the initial respiratory data.

[0050] The total initial respiratory characteristics include the initial respiratory time-domain characteristics and the initial respiratory frequency-domain characteristics; Specifically, the electronic device can extract initial respiratory time-domain features and initial respiratory frequency-domain features from initial respiratory data. For example, Table 2 shows the initial respiratory time-domain features and initial respiratory frequency-domain features.

[0051] Table 2. Initial respiratory temporal and frequency domain characteristics.

[0052] Specifically, for the initial respiratory time-domain characteristics (corresponding to items 1-5 in Table 2), the electronic device can perform peak and trough detection on the initial respiratory data and calculate: the average value of the respiratory peak and trough interval (AVRESP): in rpm; the standard deviation of the respiratory peak interval (Std), maximum value (Max), minimum value (Min), and range (Range), all in rpm.

[0053] Based on the initial respiratory frequency domain characteristics (corresponding to items 6-7 in Table 2), electronic devices can perform spectral analysis on the initial respiratory data to extract: respiratory signal energy (Power): unit %²; respiratory peak (Peak): the frequency value with the highest energy in the spectrum, unit Hz.

[0054] Step S2023: Extract features from the initial eye-tracking data in the initial training dataset to obtain the initial eye-tracking features.

[0055] Specifically, electronic devices can extract initial eye movement features from initial eye movement data.

[0056] For example, as shown in Table 3, are the initial eye movement features.

[0057] Table 3 Initial eye movement characteristics

[0058] Specifically, electronic devices can extract features from initial eye-tracking data based on the I-VT algorithm to obtain initial eye-tracking features (corresponding to items 1-5 in Table 3), including: average pupil diameter (PupilMean): the average pupil diameter during fixation, in mm; fixation count per second (FixationCount): in N / s; total fixation duration (FixationTotalduration) and average fixation duration (Fixation Meanduration): in s; and average saccades mean variances (Saccades Mean Variances): in mm².

[0059] Step S2024: The total initial electrocardiogram features, total initial respiratory features, and initial eye movement features are fused to generate initial physiological features.

[0060] Specifically, the electronic device can normalize the total initial electrocardiogram features (12 items), total initial respiratory features (7 items), and initial eye movement features (5 items), and then splice the normalized total initial respiratory features and initial eye movement features to generate initial physiological features.

[0061] Step S2025: Based on the operational elements of each stage of the five-sided flight, generate mission complexity features.

[0062] Specifically, electronic equipment can generate mission complexity characteristics by setting key operational parameters for each stage of the five-way flight (e.g., controlling the climb rate to 500 ft / min in the first leg and the touchdown speed to 40 ft / min in the fifth leg). For example, a difficulty coefficient (1-5 points) can be set for four key operational parameters in each stage. These four key operational parameters can be heading maintenance, altitude control, airspeed adjustment, and position calibration, which are extended to 20 dimensions through one-hot encoding. Example: When the "heading maintenance" parameter in the third leg (downwind) has a difficulty of 3 points, the corresponding encoding is [0,0,1,0,0]; when the altitude control parameter has a difficulty of 4 points, the corresponding encoding is [0,0,0,1,0]; when the airspeed adjustment parameter has a difficulty of 2 points, the corresponding encoding is [0,1,0,0,0]; and when the position calibration parameter has a difficulty of 5 points, the corresponding encoding is [0,0,0,0,1], which are then concatenated into 20 dimensions.

[0063] Step S2026: The training flight phase, training cognitive load level, and task complexity features are fused to generate initial condition features.

[0064] Specifically, the electronic device can encode the training flight phase and the training cognitive load level, and then splice the encoded training flight phase, training cognitive load level and task complexity features to generate initial condition features.

[0065] For example, 3D one-hot encoding is used for the training flight phases (first / third / fifth phases), and 2D one-hot encoding is used for the training cognitive load levels (low / high). The five-phase flight in this field is a classic task in pilot training, which includes key phases such as takeoff (first phase), leeward (third phase), and landing (fifth phase). The operational complexity and cognitive load of different phases are different (for example, the fifth phase is the landing phase, which has high accuracy requirements and usually higher cognitive load).

[0066] One-hot encoding is a method for converting categorical variables (here, the "flight phase") into binary vectors: when a sample belongs to the "first side", it is represented by the vector [1,0,0] (the first bit is 1, and the rest are 0); when a sample belongs to the "third side", it is represented by the vector [0,1,0] (the second bit is 1, and the rest are 0); when a sample belongs to the "fifth side", it is represented by the vector [0,0,1] (the third bit is 1, and the rest are 0).

[0067] Step S2027: The initial physiological features and initial conditional features are fused to generate initial fused features.

[0068] Specifically, the electronic device fuses initial physiological features and initial conditional features to generate initial fused features. For example, the electronic device splices 24-dimensional initial physiological features with 25-dimensional initial conditional features to ultimately form 49-dimensional initial fused features.

[0069] Step S2028: Input the initial fusion features into the preset generation network based on the preset data augmentation model to generate noisy data corresponding to the initial training dataset.

[0070] Specifically, step S2028 above may include the following steps: Step a1: Input the initial fused features into the first hidden layer of the preset generator network, and perform expansion processing on the initial fused features to generate the first hidden features.

[0071] Specifically, step a1 above may include the following steps: Step a11: Determine the training flight phase and training cognitive load level from the initial fusion features.

[0072] Specifically, electronic devices can determine the training flight phase and the training cognitive load level from the initial fusion characteristics.

[0073] Step a12: Based on the training flight phase and the training cognitive load level, determine the target expansion dimension corresponding to the initial fusion feature from the preset mapping table.

[0074] Specifically, the electronic device can determine the target expansion dimension corresponding to the initial fusion feature from a preset mapping table based on the combination of training cognitive load level and training flight phase.

[0075] For example, low load + first / third side: target expansion dimension 88 (expansion ratio 1.8), because the operational complexity of this scenario is low and the physiological signal fluctuations are gentle, so there is no need for too many feature dimensions; high load + fifth side (landing phase): target expansion dimension 108 (expansion ratio 2.2), because the cognitive load is the highest during the landing phase and the physiological signals (such as pupil fluctuations and heart rate changes) are complex, so more dimensions are needed to capture details. Other scenarios (such as low load + fifth side, high load + first side): default target expansion dimension 97 (expansion ratio 1.98).

[0076] Step a13: Generate the first initial weight matrix based on the initial dimension corresponding to the initial fusion feature and the target extended dimension.

[0077] Specifically, the electronic device can initialize the weight matrix using a Xavier normal distribution to generate the first initial weight matrix, ensuring that the variances of the input and output features are consistent (avoiding gradient vanishing or exploding), as shown in the formula: Where in_dim is the dimension of the initial fused feature, for example, 49, and out_dim is the target expanded dimension (e.g., 97).

[0078] Step a14: Based on the training flight phase and the training cognitive load level, the first initial weight matrix is ​​modified to obtain the first target weight matrix.

[0079] Specifically, the electronic device can initially determine the current scene features by fusing the training flight phase and training cognitive load level from the initial fusion features. Then, the initial weight matrix is ​​modified based on the current scene features to obtain the first target weight matrix. For example, if the training flight phase and training cognitive load level are "high load + fifth side", then the current scene features are determined to be the first scene. The first scene is a highly stressful scene, so the weight values ​​related to eye movement features (the last 5 dimensions) are enhanced (multiplied by a factor of 1.2), because eye movement features such as pupil sway during the landing phase are more sensitive to cognitive load.

[0080] If the training flight phase and the training cognitive load level are "low load + first side", then the current scenario characteristics are determined to be the second scenario. The second scenario is a relatively easy scenario, so the weight value related to the high-frequency characteristics of ECG (such as HF, LF) is reduced (multiplied by a coefficient of 0.8), because the ECG frequency domain characteristics fluctuate less under low load.

[0081] Optionally, the electronic device can also differentiate and strengthen the weights of corresponding regions in the initial weight matrix based on the importance differences of various physiological feature modalities in the initial fusion features. For example, for the region corresponding to the total initial ECG features (the first 12 dimensions of the input vector), the standard deviation of the weights is reduced by 20% (compared to other regions) to suppress the excessive amplification of high-frequency noise in the ECG signal (meeting the low activation requirement of the subsequent LeakyReLU slope of 0.1). For the region corresponding to the initial eye-tracking features (the last 5 dimensions of the input vector), the standard deviation of the weights is increased by 20% to enhance the ability to capture dynamic features such as pupil diameter and fixation duration (matching the high activation requirement of the subsequent LeakyReLU slope of 0.3). For the region corresponding to the total initial respiratory features (the middle 7 dimensions of the input vector), the weight values ​​are positively correlated with the respiratory rate (AVRESP) feature. When the conditional encoding indicates a "high respiratory rate scenario," the corresponding weights are multiplied by a factor of 1.1 to obtain the final first target weight matrix.

[0082] Step a15: Expand the initial fused features based on the first target weight matrix to generate a target extended dimension feature vector.

[0083] Specifically, the electronic device uses the initial fusion features to right-multiply the first target weight matrix to generate a target extended dimension feature vector.

[0084] Step a16: Activate the target extended dimension feature vector based on the preset activation function and output the first hidden feature.

[0085] Specifically, the electronic device can determine the LeakyReLU activation function and, based on the different modalities corresponding to the sub-features in the target extended dimension feature vector, determine the slope in the LeakyReLU activation function, thereby activating the target extended dimension feature vector based on LeakyReLU activation functions with different slopes and outputting the first hidden feature.

[0086] The expression for the LeakyReLU activation function is: Where α is the slope parameter. Different slopes are used for different modal features, based on the physiological differences in ECG, respiration, and eye movement signals. The slope of the ECG feature neurons (generated from the first 12 physiological features in the initial physiological features) in the target extended dimension feature vector is set to 0.1 to suppress excessive amplification of high-frequency noise in the ECG signal. The slope of the eye movement feature neurons (generated from the last 5 physiological features in the initial physiological features) in the target extended dimension feature vector is set to 0.3 to retain key dynamic features such as pupillary fluctuations under high load. An adaptive slope is added for the respiration feature neurons (based on the corresponding middle 7 physiological features in the initial physiological features) in the target extended dimension feature vector. a. Extract AVRESP values ​​from the middle 7 features and determine if "AVRESP > 18 rpm" (rapid breathing state) is satisfied. b. If satisfied, apply LeakyReLU with α = 0.4 to the 7-dimensional respiration features to enhance sensitivity to negative features during rapid breathing (such as a brief decrease in respiratory depth). c. If the condition is not met (AVRESP≤18rpm), apply the default slope α=0.2 to balance the characteristic activation intensity under normal breathing conditions.

[0087] Step a2: Input the first hidden feature into the second hidden layer of the preset generation network, scale the first hidden feature, and generate the second hidden feature.

[0088] Specifically, step a2 above may include the following steps: Step a21: Calculate the first mutual information value between each initial sub-fusion feature in the initial fusion feature.

[0089] Then, for each initial sub-fusion feature among the total initial ECG features, total initial respiratory features, initial eye movement features, and initial conditional features, local mutual information is calculated using a sliding window of a preset time length (the window size can be adjusted according to the feature dynamics, such as a 3-second window for physiological features and a 5-second window for conditional features). The average value is then taken to obtain the sub-feature mutual information value corresponding to each initial sub-fusion feature, thus avoiding the obscuring of dynamic correlations by a single global value.

[0090] Optionally, electronic devices can assign a 1.2x weight to high-load samples and a 0.8x weight to low-load samples, thereby strengthening the association of features related to cognitive load (such as the association between heart rate and pupil diameter under high load).

[0091] Then, the electronic device calculates the first mutual information value between each initial sub-fusion feature, using the following formula (after discretization): Here, X and Y represent two initial sub-fusion features to be analyzed (e.g., "heart rate" in the total initial ECG features and "respiratory rate" in the total initial respiratory features). P(X) and P(Y) represent the marginal probability distributions of X and Y, respectively, that is, the probability of a single initial sub-fusion feature taking all possible values ​​(e.g., the probability of a heart rate of 70 bpm and the probability of a respiratory rate of 15 rpm). P(X,Y) is the joint probability distribution of initial sub-fusion feature X and initial sub-fusion feature Y, representing the probability of both initial sub-fusion features taking a certain combination of values ​​(e.g., the probability of a heart rate of 70 bpm and a respiratory rate of 15 rpm).

[0092] Step a22: Construct a modal correlation matrix based on each first mutual information value.

[0093] Specifically, the electronic device can construct a modal correlation matrix using all initial sub-fusion features as rows and columns. Each element in the matrix corresponds to the first mutual information value of two initial sub-fusion features. This matrix clearly presents the degree of correlation between different modal sub-features (ECG, respiration, eye movement, and conditional features), such as the correlation strength between ECG features and conditional features, and the correlation strength between respiration features and eye movement features, providing a basic correlation basis for subsequent calculations.

[0094] For example, an electronic device can construct a 49×49 modal correlation matrix M based on the normalized first mutual information values, where M[i,j] represents the normalized mutual information value between the i-th initial sub-fusion feature and the j-th initial sub-fusion feature. For example, the correlation between ECG HR (E1) and respiratory AVRESP (R1) may be 0.65 (strong correlation); the correlation between ECG LF (E7) and eye movement O1 (pupil diameter) may be 0.2 (weak correlation).

[0095] Step a23: Mark each first hidden sub-feature according to the mapping relationship between each first hidden sub-feature in the first hidden features and each initial sub-fusion feature in the initial fusion features.

[0096] Specifically, the electronic device can determine the mapping relationship between each first hidden sub-feature in the first hidden features and each initial sub-fusion feature in the initial fusion features. By calculating the weight contribution of the first hidden sub-feature to each initial sub-fusion feature (such as the feature importance value obtained through gradient backpropagation), the type of initial sub-fusion feature it mainly depends on is determined.

[0097] Labeling rule: If a certain first hidden sub-feature has the highest total weight contribution to the initial ECG sub-fusion feature, it is labeled as "ECG dominant"; similarly, the first hidden sub-features of "respiration dominant", "eye movement dominant" and "conditional dominant" are labeled respectively.

[0098] Step a24: Based on the labeling results, calculate the mean vector corresponding to each first hidden sub-feature dominated by ECG, and use it as the first cluster center.

[0099] Specifically, the electronic device can determine the first hidden features dominated by ECG, the first hidden features dominated by respiration, the first hidden features dominated by eye movement, and the first hidden features dominated by condition based on the labeling results. Then, the electronic device calculates the mean vector corresponding to each first hidden feature dominated by ECG, which serves as the first cluster center.

[0100] Step a25: Calculate the mean vector of each first hidden sub-feature dominated by breathing, and use it as the second cluster center.

[0101] Specifically, the electronic device calculates the mean vector of each of the first hidden sub-features dominated by breathing, and uses it as the second cluster center.

[0102] Step a26: Calculate the mean vector of each first hidden sub-feature dominated by eye movement, and use it as the third cluster center.

[0103] Specifically, the electronic device calculates the mean vector of each of the first hidden sub-features dominated by eye movement, which serves as the third cluster center.

[0104] Step a27: Calculate the mean vector of each conditionally dominant first hidden sub-feature, and use it as the fourth cluster center.

[0105] Specifically, the mean vector of each of the first hidden sub-features dominated by the computational conditions of the electronic device is used as the fourth cluster center.

[0106] Step a28: Calculate the second mutual information value between each first hidden sub-feature based on the modal correlation matrix.

[0107] Specifically, the electronic device derives the second mutual information value between each first hidden sub-feature based on the constructed modal correlation matrix and the mapping relationship between the first hidden sub-feature and the initial sub-fusion feature.

[0108] For example, for two first hidden sub-features A and B, firstly, the core source tracing initial sub-fusion features corresponding to the first hidden sub-features A and B are determined. Then, the first mutual information values ​​of these two core source tracing initial sub-fusion features in the modal correlation matrix are calculated. The first mutual information values ​​of the secondary source tracing initial sub-fusion features corresponding to the first hidden sub-features A and B are also determined. Then, the proportions of the core source tracing initial sub-fusion features and secondary source tracing initial sub-fusion features in the first hidden sub-feature A of the electronic device, and the proportions of the core source tracing initial sub-fusion features and secondary source tracing initial sub-fusion features in the first hidden sub-feature B, are used to calculate the second mutual information value between the first hidden sub-features A and B.

[0109] For example, the core source fusion features of A (e.g., ECG SDNN, weighted at 80%) and the core source fusion features of B (e.g., respiratory AVRESP, weighted at 75%) are determined. Then, the electronic device extracts the mutual information values ​​of these two core source features from the modal correlation matrix (e.g., MI = 0.52 for SDNN and AVRESP). The mutual information of the secondary source features in A and B is weighted and summed (the weights are the contribution percentages during feature mapping), ultimately yielding MI(A,B). Example: If A has 70% weight from ECG HR (heart rate) and 30% from respiratory AVRESP; and B has 60% weight from respiratory AVRESP and 40% from eye movement FixationCount, then MI(X,Y)=0.7×0.6×MI(HR,AVRESP)+0.7×0.4×MI(HR,FixationCount)+0.3×0.6×MI(AVRESP,AVRESP)+0.3×0.4×MI(AVRESP,FixationCount) (where MI(AVRESP,AVRESP)=1, because the mutual information of the same first hidden sub-feature is 1).

[0110] Step a29: Based on the second mutual information value, cluster each first hidden sub-feature to obtain the ECG cluster group, the respiration cluster group, the eye movement cluster group, and the conditional cluster group.

[0111] Specifically, the electronic device can employ hierarchical clustering. Starting with four cluster centers, the similarity between each first hidden feature and each cluster center is calculated (similarity = 1 - Euclidean distance of the second mutual information value). Then, each first hidden feature is assigned to the cluster group with the highest similarity, forming ECG cluster, respiration cluster, eye movement cluster, and conditional cluster. For "mixed dominant" features, they are classified according to their maximum similarity with the four cluster centers. After classification, the electronic device can recalculate the mean vector of each cluster and update the cluster centers. The above clustering process is then repeated until the change in the cluster centers is <0.01 (convergence).

[0112] Step a210: Output the second hidden feature based on the ECG cluster group, respiratory cluster group, eye movement cluster group, and conditional cluster group.

[0113] Specifically, step a210 above may include the following steps: Step a2101: Based on the modal correlation matrix, calculate the average correlation degree between the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the conditional cluster group.

[0114] Specifically, the first mutual information value corresponding to the initial sub-fusion feature of each combination is extracted from the modal correlation matrix. Then, the arithmetic mean of the extracted first mutual information values ​​is calculated to obtain the average correlation degree between each two cluster groups, forming a 4×4 average correlation degree matrix.

[0115] Example: Correlation between ECG and respiratory groups (A) er ): Extract the intersection region (a 12×7 submatrix) of the modal correlation matrix M, which consists of the central electrical features (E1-E12) and the respiratory features (R1-R7). Calculate the average of all mutual information values ​​in this submatrix using the formula: A er =∑i / 12×7. If the sum of mutual information values ​​in the submatrix is ​​42, then A er =42 / (12×7)=0.5.

[0116] Step a2102: Determine the connection weights between the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the conditional cluster group based on the average correlation. Specifically, electronic devices can determine the connection weights between ECG clusters, respiratory clusters, eye-tracking clusters, and conditional clusters based on the correspondence between average correlation and connection weights.

[0117] For example, the connection weight = average correlation × (1 + average correlation) to increase the weight of high correlation, with the weight range controlled within [0,2]. For example, if the average correlation between ECG and respiration is 0.6, then the connection weight = 0.6 × (1 + 0.6) = 0.96; if the average correlation between respiration and eye movement is 0.3, then the connection weight = 0.3 × (1 + 0.3) = 0.39.

[0118] Step a2103: Generate a second initial weight matrix based on the initial dimension corresponding to the initial fusion feature and the target extended dimension.

[0119] Specifically, the electronic device generates a second initial weight matrix (i.e., a 97×49 matrix) with dimensions of "target expanded dimension × initial dimension" based on the initial dimension of the initial fusion features (e.g., 49 dimensions, including 24 physiological features and 25 conditional features) and the target expanded dimension (e.g., 128 dimensions). The electronic device can use Kaiming initialization to make the matrix elements follow a normal distribution with a mean of 0 and a variance of 2 / initial dimension, adapting to the gradient characteristics of ReLU-type activation functions.

[0120] Step a2104: Divide the second initial weight matrix into intra-group and inter-group regions.

[0121] Specifically, the electronic device can define the region in the second initial weight matrix that corresponds to the same cluster group feature mapped to its own exclusive dimension as the intra-group region, and the region in the second initial weight matrix that corresponds to the cross-mapping of features from different cluster groups as the inter-group region.

[0122] Step a2105: Determine the first target weight value corresponding to the inter-group region in the second initial weight matrix based on the connection weights between the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the respiratory cluster group.

[0123] Specifically, the electronic device can calculate the first target weight value corresponding to the inter-group region in the second initial weight matrix by using the first target weight value = the element value of the second initial weight matrix × the connection weight.

[0124] Step a2106: Determine the second target weight value corresponding to the region within the group based on the criticality of each cluster sub-feature in the ECG cluster group, respiratory cluster group, eye movement cluster group, and conditional cluster group.

[0125] Specifically, electronic devices can determine the criticality of each cluster sub-feature by feature contribution (such as the feature importance score of random forest). The higher the criticality (e.g., ≥0.6), the greater the impact of the feature on cognitive load identification.

[0126] The weight calculation formula is: Second target weight value = Second initial weight matrix element value × (1 + criticality). That is, the sub-feature with a criticality of 0.8 will have its weight increased by 80% within the group (e.g., initial value 0.3 → 0.3 × 1.8 = 0.54). Among them, the criticality of the conditional clustering sub-features in the conditional clustering group is dynamically adjusted according to the scenario (e.g., the criticality of "flight phase coding" in the landing phase is set to 0.9).

[0127] Step a2107: Generate a second target weight matrix based on the first target weight and the second target weight values.

[0128] Specifically, the electronic device fills the first target weight value of the inter-group region and the second target weight value of the intra-group region into the corresponding positions, covering the original elements of the second initial weight matrix, and normalizes the integrated matrix so that the sum of the elements in each row is 1, so as to avoid the weight of a certain type of feature being too high, and finally obtains the second target weight matrix.

[0129] Step a2108: Multiply the first hidden feature by the second target weight matrix to obtain the second hidden feature.

[0130] Specifically, the first hidden feature is multiplied by the second target weight matrix to obtain the second hidden feature with the initial dimension.

[0131] Step a3: Activate the second hidden feature to generate noisy data.

[0132] Specifically, the electronic device splits the second hidden feature by modality and maps it to the target dimension: the first 12 dimensions: ECG noise features; the middle 7 dimensions: respiratory noise features; the last 5 dimensions: eye movement noise features; finally forming a 24-dimensional noise feature vector.

[0133] For ECG noise feature processing: OutputECG = Tanh(x) × 0.6 + 0.2, mapping the features to the [-0.4, 0.8] interval (matching the amplitude distribution of real ECG noise). For respiratory noise feature processing: OutputResp = Tanh(x) × 0.5 - 0.1, mapping the features to the [-0.6, 0.4] interval (consistent with the baseline offset characteristics of respiratory signals). For eye movement noise feature processing: directly using the Tanh(x) output, mapping to the [-1, 1] interval (eye movement signals have a wide dynamic range).

[0134] The electronic device checks whether each noise feature falls within a physiologically reasonable range (e.g., ECG noise amplitude ≤ 0.8mV, respiratory rate 12-30rpm). Then, it checks whether the correlation of noise data across modal features conforms to a pattern (e.g., when HR increases by > 10bpm, does the respiratory rate increase synchronously?). Finally, only noise data with a comprehensive score ≥ 80 points are retained, and abnormal samples that do not conform to physiological patterns are removed.

[0135] Step S2029: Input the noise data into the preset discrimination network in the preset data augmentation model to discriminate the noise data.

[0136] Specifically, step S2029 above may include the following steps: Step b1: Calculate the rate of load change based on the initial ECG and initial eye movement data.

[0137] Specifically, electronic devices can extract heart rate (HR) data and RR interval standard deviation (SDNN) data from initial electrocardiogram (ECG) data, and extract pupil diameter change rate and saccade speed from initial eye movement (EMG) data.

[0138] For ECG load values: the formula ECG_load = 0.6 × (current HR - resting HR) / resting HR + 0.4 × (baseline SDNN - current SDNN) / baseline SDNN is used to quantify the load increase caused by increased heart rate and decreased heart rate variability (range 0-1). Here, resting HR and baseline SDNN can be obtained by testing the initial ECG data and initial eye movement data of the trained pilots.

[0139] For eye movement load (EOG) values: the formula EOG_load = 0.7 × (current pupil diameter - resting pupil diameter) / resting pupil diameter + 0.3 × current saccade velocity / maximum saccade velocity is used to reflect pupil dilation and visual scanning intensity (range 0-1). The resting pupil diameter and maximum saccade velocity can be obtained by testing trained pilots using initial ECG and initial eye movement data. The overall load value is fused: Overall Load = 0.5 × ECG_load + 0.5 × EOG_load, balancing the contributions of ECG and eye movement characteristics.

[0140] Load change rate calculation: with a time window of 10 seconds, the rate = (current total load - total load 10 seconds ago) / 10, the unit is "load units / second", a positive value indicates that the load is increasing, and a negative value indicates that it is decreasing.

[0141] Step b2 involves fusing noise data, training flight phases, training cognitive load levels, and load change rates to generate target fusion features.

[0142] Specifically, electronic devices can stitch together noise data, training flight phases, training cognitive load levels, and load change rates to generate target fusion features.

[0143] For example, the noise data is 24-dimensional, including 12-dimensional ECG, 7-dimensional respiration, and 5-dimensional eye-tracking noise features; the training flight phase is converted into a 3-dimensional vector through one-hot encoding, such as takeoff=[1,0,0], cruise=[0,1,0], landing=[0,0,1]; the training cognitive load level is a 1-dimensional numerical value, such as integer encoding of levels 0-5; and the load change rate is a 1-dimensional continuous value, output in step b1. The fusion method is to concatenate the data in the order of "noise data → flight phase → load level → change rate" to form a 24+3+1+1=29-dimensional target fusion feature, preserving the original correlation between each element.

[0144] Step b3: Based on the first branch, perform feature compression on the target fusion features to obtain the first branch features.

[0145] Specifically, electronic devices can construct a first weight matrix (e.g., a 50×32 matrix) and, through training, select features strongly correlated with the steady state of cognitive load. Higher weight values ​​indicate a stronger correlation between the feature and the steady state. Key stable features to retain include: ECG features: resting heart rate (average heart rate at baseline), baseline value of heart rate variability SDNN (standard deviation of normal sinus intervals); respiratory features: baseline respiratory rate (average respiratory rate per minute at rest), baseline tidal volume; eye movement features: resting pupil diameter, baseline blink rate. The fluctuation characteristics of stable features are as follows: under a steady cognitive load state (e.g., at rest or with low load), the fluctuations in physiological features are mostly small positive changes (e.g., a slight increase in heart rate near the baseline).

[0146] Matrix multiplication operation: Multiply the target fused features with the first weight matrix to obtain the first compressed intermediate features, thus completing the first feature compression.

[0147] Then, the electronic device constructs a second weight matrix (e.g., a 32×16 matrix) to compress the first compressed intermediate features. Compression objective: Focus on the most core stable physiological features, eliminate secondary correlated features, and improve feature discrimination efficiency. Core feature selection: ECG core features: SDNN baseline value, 24-hour mean resting heart rate, LF / HF baseline ratio (low-frequency to high-frequency power ratio); Respiratory core features: baseline values ​​of basic respiratory rate and respiratory depth coefficient of variation; Eye movement core features: resting pupil diameter level, fixation stability baseline value; Conditional correlation features: flight phase coding under low-load conditions, stable load level label.

[0148] Finally, through the mechanism of the ReLU activation function, expressed as f(x) = max(0,x), the input feature values ​​are thresholded, retaining only positive features and setting negative features directly to 0, thus obtaining the first branch features. Noise suppression indicates that negative fluctuations (such as sudden drops in heart rate or abnormal pupil constriction) are usually noise signals in an unstable state (such as measurement errors or sudden interference). ReLU can effectively filter these outliers and purify stable features.

[0149] Example: The baseline resting heart rate is 65 bpm, and the normal fluctuation range is 60-70 bpm (positive deviation). A sudden drop to 50 bpm with a negative deviation (-15) will be suppressed by ReLU to avoid interfering with the judgment of the steady state.

[0150] Step b4: Based on the second branch, perform feature compression on the target fusion features to obtain the second branch features.

[0151] Specifically, electronic devices can construct a third weight matrix (e.g., a 50×32 matrix), with the core objective of filtering features strongly correlated with dynamic changes in cognitive load. Feature types emphasized include: ECG dynamic features: heart rate increase (heart rate difference between adjacent 5-second intervals), LF / HF ratio change rate (sympathetic nerve activity indicator), RR interval shortening rate; Eye movement dynamic features: pupil dilation rate (diameter change per unit time), fixation duration shortening magnitude, scan rate increase; Conditional correlation dynamic features: task complexity increase rate, load level change slope. The filtering logic is that the matrix weight values ​​are positively correlated with the dynamic correlation of the features; for example, the weight of heart rate increase is higher than that of resting heart rate, ensuring that dynamic features are preserved after compression, resulting in 32-dimensional second compressed intermediate features.

[0152] In the 32-dimensional second compressed intermediate features output, the key focus is on retaining the detailed changes in high-load associated features: such as the instantaneous fluctuation of the LF / HF ratio rapidly increasing from 1.2 to 2.5 within 10 seconds; and the sudden change signal of fixation duration abruptly shortening from 0.5 seconds to 0.3 seconds. These details directly reflect the rapid increase in cognitive load and are the core basis for dynamic analysis.

[0153] Electronic devices can construct a 32×16 fourth weight matrix to further refine the changing trends of features rather than instantaneous values. Focus on trend types: ECG trends: 5-minute moving average slope of the LF / HF ratio, cumulative increment of heart rate increase; Eye movement trends: 30-second moving average of pupil diameter, deceleration rate of fixation duration; Comprehensive trends: synergy of multimodal feature changes (e.g., synchronicity between heart rate increase and pupil dilation).

[0154] Compression results: The 16-dimensional features fully focus on dynamic trends, such as regular signals like "continuous increase in the LF / HF ratio" and "accelerated pupil diameter dilation." Then, the compressed 16-dimensional features are activated using the ELU activation function to obtain the second branch features.

[0155] The expression for the ELU activation function is as follows: It directly preserves positive features (such as increased heart rate and pupil dilation, which are signals of increased workload), and retains a certain response to negative features (such as a slow decrease in heart rate and pupil constriction, which are signals of decreased workload), rather than completely suppressing them like ReLU. It can capture feature changes when the workload decreases (such as a negative change in heart rate from 100 bpm to 70 bpm after the task difficulty decreases). Such trends are as important as those when the workload increases in dynamic analysis (such as reflecting the recovery process after the task difficulty decreases).

[0156] Step b5: Divide the target fusion features into 4 specific groups according to physiological modality and functional attributes and label them.

[0157] The four dedicated groups are ECG group, respiratory group, eye movement group, and load change rate label group.

[0158] Specifically, for the ECG group (labeled "ECG"): Core features: 12 ECG features from 24-dimensional noisy data, including time-domain, frequency-domain, and nonlinear features such as heart rate (HR), RR interval standard deviation (SDNN), high-frequency power (HF), and low-frequency power (LF). Corresponding conditional encoding: The part of the 25-dimensional conditional encoding related to ECG features, mainly task complexity-assisted encoding related to changes in ECG features (such as encoding related to ECG feature weight adjustment in emergency scenarios). Labeling method: All feature vector elements in this group are labeled "ECG". In subsequent processing, features with this label will be given priority for ECG modality-specific verification (such as frequency band range verification for LF and HF).

[0159] For the respiratory group (labeled "RESP"): Core features: 7 respiratory features from the 24-dimensional noise data, including the average respiratory rate (AVRESP), standard deviation of respiratory peak interval (Std), and respiratory signal energy (Power). Corresponding conditional coding: The part of the 25-dimensional conditional coding related to respiratory features, such as the flight phase coding related to changes in respiratory rate (the coding related to changes in respiratory features during the fifth-side landing phase). Labeling method: A "RESP" label is added. Subsequent processing will focus on verifying the matching between respiratory rate and load level for this group of features (e.g., whether AVRESP > 16 rpm under high load).

[0160] For the eye-tracking group (labeled "EYE"): Core features: Eye-tracking features (5 items) in the 24-dimensional noisy data, including mean pupil diameter, fixation count, and fixation duration. Corresponding conditional encoding: The part of the 25-dimensional conditional encoding related to eye-tracking features, such as the encoding related to the weighting of eye-tracking features in visually intensive tasks (landing phase). Labeling method: Add the "EYE" label; subsequent verification will focus on the physiological laws unique to the eye-tracking modality, such as the negative correlation between pupil diameter and fixation duration.

[0161] The load change rate label group (labeled "RATE") features are constructed as follows: A separate 1D load change rate label is generated by calculating the average of the heart rate increase and pupil dilation rate over adjacent 5 seconds, ranging from -1 to 1. Labeling method: The "RATE" label is attached. Since it reflects the dynamic trend of cognitive load changes, it will participate as an independent feature in the triggering judgment of the "load difference amplifier" mechanism in subsequent processing (e.g., increasing the weight of sensitive path features when the load level changes from low to high).

[0162] Step b6: Verify the first branch feature and the second branch feature respectively, and obtain the ECG dimension score, respiratory dimension score and eye movement dimension score corresponding to the first branch feature and the second branch feature respectively.

[0163] Specifically, (1) For the calculation of ECG dimension scores: verification indicators and physiological range: LF (low frequency power): the effective physiological range is limited to 0.04-0.15Hz (reflecting sympathetic nerve activity), and exceeding this range (such as 0.03Hz or 0.16Hz) is considered abnormal; HF (high frequency power): the effective physiological range is limited to 0.15-0.4Hz (reflecting parasympathetic nerve activity), and exceeding this range (such as 0.14Hz or 0.41Hz) is considered data distortion.

[0164] The electronic device extracts LF and HF values ​​from the first branch feature (normal path, focusing on static features) and the second branch feature (sensitive path, focusing on dynamic features), respectively. If LF or HF exceeds the physiological range, the weight of the corresponding feature is reduced by "exceeding the limit by 0.5". For example: if HF = 0.45Hz (exceeding the upper limit by 0.05Hz), the weight is reduced by 0.05 × 0.5 = 0.025 times; if LF = 0.03Hz (below the lower limit by 0.01Hz), the weight is reduced by 0.01 × 0.5 = 0.005 times. ECG dimension score generation: Scoring rules: 100 points for both LF and HF within the physiological range; 50 points deducted for one exceeding the limit (50 points gained); 0 points for both exceeding the limit.

[0165] (2) Calculation of respiratory dimension score: Verification index and matching standard: AVRESP (average respiratory rate) must match the load level label and conform to physiological laws. For example, in a high load scenario (load level label is "high"), AVRESP is required to be >16 rpm (breathing speed increases under high load); in a low load scenario (load level label is "low"), AVRESP is required to be 12-16 rpm (breathing is stable).

[0166] The electronic device extracts AVRESP values ​​from the first and second branch features respectively, and combines them with the load level label to determine the matching. If AVRESP does not match the scene (e.g., AVRESP=15rpm under high load), it is marked as a "feature to be corrected", and a penalty weight of 1.2 times is added to the subsequent loss calculation (to strengthen the correction requirement).

[0167] Respiratory dimension score generation: Scoring rules: AVRESP matching the load level gets 100 points; mismatch is marked "to be corrected" and gets 50 points (leaving room for correction).

[0168] (3) Calculation of eye movement dimension score: Verification index and physiological relationship: The physiological effective range of pupil diameter is limited to 2-8mm. Data outside the range (such as 1.5mm or 8.5mm) are considered invalid. The correlation between pupil diameter and fixation duration should be negative (Pearson correlation coefficient < -0.2), that is, fixation duration is shortened when the pupil dilates (which is consistent with the rule that visual scanning speeds up under high load).

[0169] Pupil diameter and fixation duration are extracted from the first and second branch features, respectively, and their Pearson correlation coefficients are calculated. If the pupil diameter exceeds the limit, or the correlation coefficient is ≥-0.2 (if it is only -0.1, the negative correlation is insufficient), the connection weight of the corresponding feature is reduced by 30% (to weaken the influence of invalid features).

[0170] Eye movement dimension score generation: Scoring rules: 100 points are awarded if the pupil diameter is within the range and the correlation coefficient is <-0.2; 50 points are deducted if one item fails to meet the standard; 0 points are awarded if two items fail to meet the standard.

[0171] Finally, the initial ECG dimension scores, initial respiratory dimension scores, and initial eye movement dimension scores corresponding to the first-branch feature and the second-branch feature are standardized to the [0,1] interval (e.g., 100 points equals 1.0, 50 points equals 0.5, and 0 points equals 0.0), respectively, to obtain the ECG dimension scores, respiratory dimension scores, and eye movement dimension scores corresponding to the first-branch feature and the second-branch feature.

[0172] Step b7: The first branch features, the second branch features, and the scores of the ECG dimension, the respiratory dimension, and the eye movement dimension are fused to obtain the discriminative fusion features.

[0173] Optionally, the electronic device can concatenate the first branch features, the second branch features, and the scores of the electrocardiogram dimension, the respiratory dimension, and the eye movement dimension to obtain discriminative fusion features.

[0174] Optionally, the electronic device can also perform weighted fusion of the ECG dimension scores corresponding to the first branch features and the second branch features respectively to obtain the target ECG dimension score; similarly, the target respiratory dimension score and the target eye movement dimension score are obtained. Then, the first branch features, the second branch features, and the target ECG dimension score, the target respiratory dimension score, and the target eye movement dimension score are concatenated.

[0175] For example, the electronic device appends the standardized 3D compliance score (target ECG dimension score, target respiration dimension score, and target eye movement dimension score) to the 32-dimensional feature vector formed by the intersection of the first and second branch features, creating a 35-dimensional feature vector. This vector contains both original feature information and compliance indicators. For instance, the 32-dimensional features reflect the numerical characteristics of physiological signals, while the 3D score reflects whether these features conform to physiological laws and scenario logic, providing a comprehensive basis for subsequent authenticity judgment, load level assessment, and loss calculation.

[0176] Step b8: Input the discriminative fusion features into the authenticity discrimination branch and output the probability that the noisy data is real data.

[0177] Specifically, the authenticity discrimination branch employs a two-layer fully connected network (35-dimensional → 20-dimensional → 1-dimensional), with ReLU activation in the intermediate layers. The output layer processes the discriminative fusion features using a Sigmoid activation function, outputting the probability (P) that noisy data is real data. 真实 ), with a value range of 0-1.

[0178] Judgment criteria: Set the threshold to 0.5, when P 真实 > When a threshold is set, the noisy data is judged to be "close to reality," indicating that its feature distribution differs little from the real physiological characteristics in the original training data; when P 真实 If the value is less than or equal to the set threshold, it is considered "deviation from reality" and requires further analysis in conjunction with other indicators.

[0179] Step b9: Input the discriminative fusion features into the load level discriminative branch and output the load level probability corresponding to the noise data.

[0180] Specifically, the load level discrimination branch uses a two-layer fully connected network (35-dimensional → 24-dimensional → 6-dimensional), with Mish activation in the intermediate layers. The output layer outputs two probability values ​​(summing to 1) through the Softmax activation function, corresponding to the predicted probabilities of high cognitive load and low cognitive load.

[0181] Step b10: Based on the preset loss function, the noisy data is judged.

[0182] Step b101: Calculate the authenticity loss based on the probability that the noisy data is real data; Specifically, the authenticity loss uses the binary classification cross-entropy loss function, with the following formula:

[0183] in, For data type labels (1 represents real samples, 0 represents generated noise data); This represents the probability that noisy data is real data.

[0184] Step b102: Calculate the load level loss based on the load level probability corresponding to the noise data.

[0185] Specifically, the cross-entropy loss function is used to calculate the loss for low and high load levels, and the formula is as follows:

[0186] in, and For load level labels (when under low load) =1、 =0; under high load =0、 =1); and This refers to the probability values ​​for low load and high load in the "Probability of Load Level Corresponding to Noise Data" output in step b9.

[0187] Step b103: Calculate compliance loss based on ECG, respiration, and eye movement scores.

[0188] Specifically, compliance losses are calculated based on the following formula:

[0189] Among them, s i For the i-th dimension score (e.g., ECG dimension score s1 = 80 points); the lower the score, the lower the s1 score. i The larger the value, the greater the compliance, thus penalizing noisy data that does not conform to physiological characteristics.

[0190] Step b104: Generate a preset loss function based on authenticity loss, load level loss, and compliance loss; Specifically, the preset loss function is:

[0191] Step b105: Based on the preset loss function, the noisy data is judged.

[0192] Specifically, electronic devices can identify noisy data based on a preset loss function.

[0193] Step S20210: Based on the discrimination results, the initial training dataset is expanded to generate the target training dataset.

[0194] Specifically, the electronic device can supplement the initial training dataset with noise data that is judged to be true based on the discrimination result, thereby expanding the initial training dataset and generating the target training dataset.

[0195] Step S203: Based on the target training dataset, train the initial cognitive load identification network to obtain the target cognitive load identification model.

[0196] Please refer to the above description of step S103 for details on this step, which will not be repeated here.

[0197] The cognitive load identification model training method provided in this application extracts and fuses features from initial electrocardiogram data, initial respiratory data, and initial eye movement data to generate initial physiological features. These features are then fused with features related to the training flight phase and task complexity to generate initial fused features. This integrates multimodal physiological data and task-related information to obtain richer and more comprehensive feature representations, improving the model's accuracy in describing the pilot's cognitive load state and reducing the risk of overfitting. Simultaneously, the complementary nature of different modal data helps capture more complex physiological and task-related relationships, enhancing the model's generalization ability. The target expansion dimension is determined based on the training flight phase and training cognitive load level, generating a first initial weight matrix. This first initial weight matrix is ​​then corrected to obtain a first target weight matrix. The initial fused features are expanded based on the first target weight matrix to obtain a target expansion dimension feature vector. This enables the target cognitive load identification model to adaptively adjust feature dimensions and weights according to the characteristics of different flight phases and load levels, highlighting key features, better adapting to dynamic data changes, and improving the model's accuracy in assessing cognitive load in specific scenarios. The first hidden feature is obtained by activating the expanded feature vector of the target dimension by a preset activation function. This provides a feature representation that has been preliminarily transformed and is more suitable for model learning, which helps the model to further extract high-level features and discover potential patterns in the data.

[0198] Then, the first mutual information value of each sub-feature in the initial fused features is calculated to construct the modality association matrix. First mutual information measures the dependency between features, helping the model understand the correlation between features of different modalities, discover potential connections between features, and provide a basis for subsequent clustering and weight determination. This helps reduce feature redundancy and enables the model to utilize feature information more efficiently. Each first hidden sub-feature in the first hidden features is labeled and clustered to determine the average correlation and connection weights between each cluster group. Grouping related features through clustering allows the target cognitive load identification model to process features of different modalities and functions in a targeted manner, improving the compactness and discriminativeness of feature representation. Determining connection weights helps the model consider the interaction between different cluster groups and better capture the collaborative relationships of multimodal data.

[0199] Next, a second objective weight matrix is ​​generated based on the clustering results and the keyity of features. This second objective weight matrix is ​​then multiplied by the first hidden feature to obtain the second hidden feature. This second objective weight matrix comprehensively considers the inter-group correlation and the importance of features within groups, enabling the target cognitive load identification model to allocate weights more reasonably during feature transformation, highlighting important features and suppressing redundant features, thereby obtaining a more representative and discriminative second hidden feature.

[0200] Noise data is generated by activating the second hidden feature, and the rate of change of cognitive load is calculated. The generation of noise data can simulate noise or uncertainty in real-world data, enhancing the robustness of the model. The calculation of the rate of change of cognitive load reflects the dynamic changes in cognitive load, providing the model with more comprehensive load information and helping to more accurately assess the real-time state of cognitive load. The noise data, training flight phase, training cognitive load level, and rate of change of cognitive load are fused to generate the target fusion feature. Based on the first branch, the target fusion feature is compressed to obtain the first branch feature. Based on the second branch, the target fusion feature is compressed to obtain the second branch feature. The target fusion feature is then divided into four specific groups according to physiological modality and functional attributes and labeled. The first branch feature and the second branch feature are validated separately, obtaining the corresponding ECG dimension score, respiratory dimension score, and eye movement dimension score. Feature compression reduces data dimensionality and computational complexity while retaining key information. The validated dimensional scores can evaluate the rationality and effectiveness of the features from different modal perspectives, providing a more reliable basis for subsequent fusion. The features from the first and second branches, along with scores from the ECG, respiration, and eye-tracking dimensions, are fused to obtain discriminative fusion features. These features are then input into the authenticity discrimination branch, outputting the probability that noisy data is real data. Similarly, the discriminative fusion features are input into the load level discrimination branch, outputting the load level probability corresponding to the noisy data. This fusion method integrates multiple aspects of information, enabling the target cognitive load identification model to evaluate data from both authenticity and load level perspectives. This provides a more comprehensive judgment for cognitive load assessment, contributing to improved accuracy and reliability of the assessment results. Authenticity loss, load level loss, and compliance loss are calculated based on different probabilities and dimensional scores, generating a preset loss function to discriminate noisy data. Through the calculation of multi-dimensional loss functions, the model can measure the difference between predicted results and reality from multiple perspectives, guiding the model to learn more accurate feature representations and prediction rules, continuously optimizing model parameters, and improving the accuracy of the target cognitive load identification model in cognitive load assessment and its ability to discriminate noisy data.

[0201] This embodiment provides a method for training a cognitive load identification model, which can be used in electronic devices. Figure 3 This is a flowchart of a cognitive load identification model training method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain the initial training dataset corresponding to the pilots being trained.

[0202] The initial training dataset includes initial electrocardiogram data, initial respiratory data, initial eye movement data, training flight phases, and training cognitive load levels corresponding to the training pilots.

[0203] Step S302: Expand the initial training dataset based on the preset data augmentation model to generate the target training dataset.

[0204] Step S303: Based on the target training dataset, train the initial cognitive load identification network to obtain the target cognitive load identification model.

[0205] Specifically, step S303 above may include the following steps: Step S3031: Extract features from the target training dataset to generate training fusion features.

[0206] Specifically, feature extraction is performed on the target ECG data in the target training dataset to obtain the total target ECG features. The total target ECG features include time-domain features, frequency-domain features, and nonlinear features. Feature extraction is performed on the target respiration data in the target training dataset to obtain the total target respiration features; these include time-domain and frequency-domain features. Feature extraction is performed on the target eye-tracking data in the target training dataset to obtain the target eye-tracking features. The total target ECG features, total target respiration features, and target eye-tracking features are fused to generate target physiological features. Based on the operational elements of each stage of the five-sided flight, task complexity features are generated. The training flight stage, training cognitive load level, and task complexity features are fused to generate target conditional features.

[0207] For a detailed description of the above content, please refer to the detailed description of steps S2021-S2026, which will not be repeated here.

[0208] Step S3032: Input the training fusion features into the input layer of the initial cognitive load identification network.

[0209] Specifically, the electronic device inputs the training fused features into the input layer of the initial cognitive load identification network.

[0210] For example, the input layer simultaneously receives 24-dimensional training physiological features (12 ECG + 7 respiration + 5 eye movement) and 25-dimensional training conditional codes, generating 49-dimensional training fusion features. The training conditional codes include 3-dimensional one-hot encoding for flight phases, 2-dimensional load level encoding, and 20-dimensional task complexity encoding. All features are standardized (Z-score transform) before entering the input layer.

[0211] In step S3033, the input layer adjusts the weight information corresponding to each sub-training fusion feature in the training fusion feature according to the target flight stage in the target training dataset to obtain the weighted fusion feature.

[0212] Specifically, the input layer can adjust the weight information of each sub-training fusion feature in the training fusion feature according to the target flight stage in the target training dataset, through the flight stage-feature importance mapping table, to obtain the weighted fusion feature. The parameters of the flight stage-feature importance mapping table are dynamically optimized through scene feature importance analysis (Gini coefficient) of the target training dataset.

[0213] For example, if the target flight phase is the fifth side landing phase, the weight of eye movement features (such as pupil diameter and fixation duration) is automatically increased by 30%, because the visual scan intensity during landing is highly correlated with cognitive load; if the target flight phase is the first side takeoff phase, the weight of electrocardiogram features (such as heart rate and LF / HF ratio) is increased by 20%, prioritizing the capture of sympathetic nerve activation state.

[0214] Step S3034: Input the weighted fusion features into the third hidden layer of the initial cognitive load identification network.

[0215] Specifically, electronic devices can input weighted fusion features into the third hidden layer of the initial cognitive load identification network.

[0216] In step S3035, the third hidden layer extracts features from the weighted fusion features to obtain the third hidden features.

[0217] Specifically, step S3035 above may include the following steps: Step c1: The residual network in the third hidden layer extracts features from the weighted fusion features and outputs the target residual features.

[0218] Specifically, the residual network includes a first residual network, a second residual network, and a third residual network. Step c1 above may include the following steps: In step c11, the first residual network in the third hidden layer extracts multimodal features from the weighted fusion features to obtain multimodal features.

[0219] Specifically, the electronic device can be divided into three groups based on the convolution kernels in the first residual network in the third hidden layer: the ECG group convolution kernel, the eye movement and conditional group convolution kernel, and the respiratory group convolution kernel.

[0220] Then, based on the ECG group convolutional kernel, feature extraction is performed on the ECG feature data in the weighted fusion feature; eye movement and conditional group convolutional kernels are used to extract features from the eye movement feature data and associated conditional coding feature data in the weighted fusion feature; and respiratory group convolutional kernels are used to extract features from the respiratory feature data in the weighted fusion feature. Finally, the features extracted by each group of convolutional kernels are concatenated to generate multimodal features.

[0221] For example, an electronic device can divide 64 convolutional kernels into 3 groups according to modality: 24 convolutional kernels for ECG, 16 convolutional kernels for respiration, and 24 convolutional kernels for eye movement and conditional features. Each group processes ECG feature data (12 dimensions), respiration feature data (7 dimensions), eye movement feature data (5 dimensions), and associated conditional coding feature data (corresponding part of 25 dimensions) in the input features. Then, group convolution is performed on the 49-dimensional input features (24-dimensional physiological features + 25-dimensional conditional coding): F1modal=Conv(Xmodal,Kmodal)+bmodal; where Xmodal is the input feature, Kmodal is the corresponding group convolutional kernel, and bmodal is the bias term; the output dimension of each group convolution is doubled (e.g., 24 ECG convolutional kernels output 48-dimensional features) and concatenated to generate 128-dimensional multimodal features F1: 24×2+16×2+24×2=128 dimensions.

[0222] In step c12, the second residual network in the third hidden layer compresses the multimodal features to obtain compressed features, and then performs a convolution operation on the compressed features to obtain convolutional features.

[0223] Specifically, the first convolutional layer in the second residual network of the third hidden layer can compress multimodal features using a 1×1 convolutional kernel to generate compressed features. For example, a 128-dimensional feature F1 is compressed to 64 dimensions using a 1×1 convolutional kernel: F1proj=Conv(F1,K1×1), with the convolutional kernel allocated according to modality proportions (ECG 24-dimensional → 12-dimensional, respiration 16-dimensional → 8-dimensional, eye movement 24-dimensional → 12-dimensional).

[0224] The second convolutional layer in the second residual network of the third hidden layer performs a second convolution on the first compressed feature F1proj (64 dimensions) after 1×1 convolution compression to obtain the convolutional feature. The formula is: F2=Conv(F1proj,K3×64). Where K3×64 represents 64 convolutional kernels of size 3×3 (matching the dimensions of the input feature), and each convolutional kernel is responsible for extracting local features of a specific pattern.

[0225] Step c13: Calculate the signal-to-noise ratio corresponding to each modal feature in the weighted fusion features.

[0226] Specifically, electronic devices can extract segments of feature data from modalities such as ECG, respiration, and eye movement in the weighted fusion features by using a sliding window (window size of 2 seconds, step size of 0.5 seconds).

[0227] Signal component (S): Feature data x within each window i (i=1,2,...,N, where N is the number of data points within the window), calculate the effective value of the signal according to the formula, i.e. (x) i (where N is the feature value within the window and N is the number of data points within the window).

[0228] Noise component (N): The db4 wavelet is selected (it has good time-frequency localization characteristics, suitable for denoising physiological signals). The feature data within the window is decomposed into three levels of wavelets to obtain high-frequency coefficients (containing noise) and low-frequency coefficients (containing signal) at each level. Soft thresholding is applied to the high-frequency coefficients (threshold = 0.6745 × median(|high-frequency coefficient|)). After filtering the noise component, the denoised signal is reconstructed. Noise signal = original window features - denoised signal, i.e., the filtered noise component is separated. The noise signal is extracted, and its effective value is calculated as a measure of noise energy. The effective value of the extracted noise signal is calculated using the same formula as the signal component: , where n i This represents the value at each point of the noise signal.

[0229] Then, calculate the signal-to-noise ratio (SNR) per decibel (dB): S 2 N is the signal power (S squared). 2 This is the noise power (N squared).

[0230] Example: If the effective value of the ECG signal is 50 and the effective value of the noise is 5, then .

[0231] Each window corresponds to an SNR value, forming a time series sequence of signal-to-noise ratio for each mode.

[0232] Example outputs: ECG SNR sequence [22dB, 21dB, 18dB, ...], respiratory SNR sequence [19dB, 16dB, 14dB, ...], eye movement SNR sequence [17dB, 13dB, 16dB, ...], reflecting the signal quality of each modality in real time.

[0233] Step c14: Based on the signal-to-noise ratio corresponding to each modality feature, perform residual superposition processing on the compressed feature and the convolutional feature to generate superimposed features.

[0234] Specifically, the electronic device can generate a time series of signal-to-noise ratios (SNRs) for each modality based on a corresponding SNR value for each window. Then, the SNRs in each time series corresponding to each modality are compared with a preset SNR threshold. Based on the comparison results, a superposition vector is generated. Specifically, when each SNR is greater than the preset SNR threshold, the elements in the superposition vector are 1; when each SNR is less than or equal to the preset SNR threshold, the elements in the superposition vector are 0. For example, assume that the element values ​​of the superposition vector G are updated in real time according to the SNR of each modality: when the SNR of a certain modality is greater than or equal to the threshold, the corresponding element in G is set to 1 (residual connection enabled); when the SNR of a certain modality is less than the threshold, the corresponding element in G is set to 0 (residual connection disabled).

[0235] Then, the electronic device performs residual superposition processing on the compressed features and convolutional features based on the superposition vector to generate superimposed features. The residual superposition formula is: F2res = F2 + F1proj × G; where F2res is the superimposed feature, F2 is the convolutional feature, and F1proj is the compressed feature.

[0236] High SNR mode (G=1): Residual connections are enabled, and the original compressed features (F1proj) and convolutional features (F2) are added to achieve the fusion of "original features + higher-order features" and enhance feature stability (e.g., ECG mode retains basic heart rate information while superimposing higher-order features of heart rate variability).

[0237] Low SNR mode (G=0): Residual connections are closed, only convolutional features (F2) are retained, and the original features (F1proj) that are contaminated by noise are avoided from participating in the superposition, thus preventing noise amplification (e.g., when the breathing mode is disturbed, only the breathing rhythm features extracted by convolution are retained).

[0238] Step c15: Based on the third residual network in the third hidden layer, the superimposed features are divided into ECG time domain subgroup, ECG frequency domain subgroup, respiratory rhythm subgroup, respiratory depth subgroup, eye movement pupil subgroup, and eye movement fixation subgroup.

[0239] Specifically, electronic devices can classify superimposed features based on the third residual network in the third hidden layer into ECG time domain subgroup, ECG frequency domain subgroup, respiratory rhythm subgroup, respiratory depth subgroup, eye movement pupil subgroup, and eye movement fixation subgroup.

[0240] For example, the ECG subgroup is further subdivided into: a time-domain subgroup (10 dimensions): including time-domain features such as SDNN (standard deviation of RR intervals), RMSSD (root mean square of the difference between adjacent RR intervals), and NN50 (the number of adjacent RR interval differences > 50ms), reflecting the overall level of heart rate variability. A frequency-domain subgroup (14 dimensions): including frequency-domain features such as LF (low-frequency power), HF (high-frequency power), and the LF / HF ratio, focusing on the sympathetic-parasympathetic balance of the autonomic nervous system.

[0241] The respiratory group is further subdivided into: Respiratory rhythm subgroup (8 dimensions): including AVRESP (mean respiratory rate), standard deviation of respiratory cycle, respiratory peak interval, etc., to characterize the stability of respiratory rhythm. Respiratory depth subgroup (8 dimensions): including tidal volume, maximum respiratory depth, inspiratory / expiratory time ratio, etc., to reflect changes in respiratory intensity.

[0242] The eye-tracking subgroup is further divided into: the eye-tracking pupil subgroup (10 dimensions): including average pupil diameter, dilation rate, and constriction rate, capturing the dynamic response of the pupil to cognitive load. and the eye-tracking fixation subgroup (14 dimensions): including fixation duration, fixation frequency, and fixation point distribution entropy, reflecting the distribution pattern of visual attention.

[0243] Step c16: Perform convolution processing on each subgroup, and compress the convolution results through max pooling and regularization to obtain the preserved features corresponding to each subgroup.

[0244] Specifically, for the six subgroups of ECG, respiration, and eye movement (ECG time / frequency domain, respiratory rhythm / depth, and eye movement pupil / fixation), dedicated convolutional kernels are configured to focus on extracting core physiological features within each subgroup. For example, for the ECG time domain subgroup (10 dimensions): 16 1×3 convolutional kernels are used to extract time-domain features such as the fluctuation pattern of the standard deviation of RR intervals (SDNN) and the changing trend of adjacent RR interval differences (RMSSD). For the ECG frequency domain subgroup (14 dimensions): 24 1×5 convolutional kernels are used to capture the spectral peak positions of low-frequency power (LF) and high-frequency power (HF) and the dynamic changes in the LF / HF ratio. For the respiratory rhythm subgroup (8 dimensions): 12 1×4 convolutional kernels are used to extract rhythmic features such as the periodicity of the average respiratory rate (AVRESP) and the stability of the standard deviation of the respiratory cycle. For the respiratory depth subgroup (8 dimensions): 12 1×4 convolutional kernels were used to focus on intensity features such as peak tidal volume and abrupt changes in the inspiratory / expiratory time ratio. For the eye-tracking pupil subgroup (10 dimensions): 16 1×3 convolutional kernels were used to extract dynamic features such as the dilation / contraction rate of pupil diameter and step changes in the mean diameter. For the eye-tracking fixation subgroup (14 dimensions): 24 1×5 convolutional kernels were used to capture the distribution patterns of fixation duration and the correlation between fixation frequency and task difficulty.

[0245] The convolutional layers for each subgroup are computed independently, using the formula: Fsub=Conv(Xsub,Ksub)+bsub; where Xsub is the original feature of the subgroup, Ksub is the corresponding convolutional kernel, and bsub is the bias term. Local correlation features within the subgroup (such as the relationship between changes in adjacent RR intervals in the ECG time domain) are extracted through local receptive fields (e.g., a 1×3 convolutional kernel covering 3 consecutive time points), enhancing the discriminative power of the features.

[0246] Then, 1×2 max pooling (window size 1×2, stride 2) was used to reduce the dimensionality of the convolutional features. For the ECG time domain subgroup: 20 dimensions after convolution → 10 dimensions after pooling → further filtering to 4 dimensions (retaining core metrics such as SDNN and RMSSD). For the ECG frequency domain subgroup: 28 dimensions after convolution → 14 dimensions after pooling → further filtering to 6 dimensions (retaining key spectral features such as LF, HF, and LF / HF). For the respiratory rhythm subgroup: 16 dimensions after convolution → 8 dimensions after pooling → further filtering to 5 dimensions (retaining features such as AVRESP and respiratory cycle standard deviation). For the respiratory depth subgroup: 16 dimensions after convolution → 8 dimensions after pooling → further filtering to 3 dimensions (retaining features such as peak tidal volume and inspiratory / expiratory ratio). For the oculomotor pupil subgroup: 20 dimensions after convolution → 10 dimensions after pooling → further filtering to 8 dimensions (retaining features such as mean diameter and dilation rate). For the eye-tracking fixation subgroup: 28 dimensions after convolution → 14 dimensions after pooling → further filtering down to 6 dimensions (preserving fixation duration, distribution entropy, etc.).

[0247] This allows for the selection of the maximum value within a window (e.g., the maximum pupil dilation rate at two time points), strengthening strong response features within subgroups. Reducing the number of features (e.g., decreasing the ECG frequency domain from 14 dimensions to 6 dimensions) lowers subsequent computational complexity and avoids overfitting. Aggregating local features (e.g., merging features from two time points) makes the features less sensitive to minor noise (e.g., short-term pupil measurement errors).

[0248] Next, the electronic device can perform batch normalization on the pooled features, using the following formula: ;in, and Let γ and β be the mean and variance of the features within the batch, γ and β be the learnable parameters, and ϵ be a small value to prevent division by zero. This stabilizes the distribution of features in each subgroup around a mean of 0 and a variance of 1, accelerating model convergence and preventing any single feature from dominating training due to excessively large values.

[0249] Add an L2 regularization term to the parameter updates of convolutional and pooling layers, as shown in the formula: Where λ is the regularization coefficient (usually taken as 0.001), and w is the network weight. For the original task loss, This represents the total loss after adding L2 regularization constraints. This helps to suppress excessively large weights, prevent the model from overfitting noise in the training data (such as an abnormal respiratory depth measurement), and improve generalization ability.

[0250] Finally, after convolution, max pooling and regularization, the preserved features with fixed output dimensions for each subgroup are obtained: ECG group: 4-dimensional time domain + 6-dimensional frequency domain = 10-dimensional; Respiratory group: 5-dimensional rhythm + 3-dimensional depth = 8-dimensional; Eye movement group: 8-dimensional pupil + 6-dimensional fixation = 14-dimensional.

[0251] Step c17: The retained features corresponding to each subgroup are fused to obtain the target residual features.

[0252] Specifically, the electronic device can stitch together the retained features corresponding to each subgroup to obtain the target residual features.

[0253] For example, the modal features are spliced ​​together: 10-dimensional ECG group (6-dimensional frequency domain + 4-dimensional time domain) + 8-dimensional respiratory group (5-dimensional rhythm + 3-dimensional depth) + 14-dimensional eye movement group (8-dimensional pupil + 6-dimensional fixation) to form a 32-dimensional target residual feature F3.

[0254] Step c2: Select ECG-respiratory crossover feature pairs and eye movement-ECG crossover feature pairs from the target residual features.

[0255] Specifically, electronic devices can extract features from target residual features and, based on the inherent correlation of physiological signals, select ECG-respiratory cross feature pairs and eye-movement-ECG cross feature pairs.

[0256] Among them, the ECG-respiratory crossover features are: the two types of signals are synergistically regulated by the autonomic nervous system, and typical combinations include: heart rate (ECG) and respiratory rate (respiration): both increase synchronously under high load (heart rate↑→respiratory rate↑); LF / HF ratio (ECG, reflecting sympathetic-parasympathetic balance) and respiratory depth (respiration): respiratory depth increases when the sympathetic nervous system is activated (LF / HF↑); RR interval standard deviation (ECG) and respiratory cycle stability (respiration): both show high volatility in the relaxed state.

[0257] Eye-ECG crossover features: There is a neural connection between visual attention and cardiovascular response. Typical combinations include: pupil diameter (eye movement) and LF / HF ratio (ECG): pupil dilation (diameter ↑) is accompanied by sympathetic activation (LF / HF ↑) under high load; fixation duration (eye movement) and heart rate (ECG): fixation time shortens (attention distraction) and heart rate increases as task difficulty increases; saccade rate (eye movement) and heart rate variability (ECG): rapid saccades (frequency ↑) correspond to decreased heart rate variability (sympathetic dominance).

[0258] Step c3: Calculate the third mutual information value corresponding to the ECG-respiratory crossover feature pair and the fourth mutual information value corresponding to the eye movement-ECG crossover feature pair.

[0259] Specifically, electronic devices can calculate the third mutual information value corresponding to the ECG-respiratory crossover feature pair and the fourth mutual information value corresponding to the eye-tracking-ECG crossover feature pair based on the KL divergence (relative entropy) of the joint probability distribution and the marginal probability distribution, using the following formula: .

[0260] Where X and Y are two feature variables (such as heart rate and respiratory rate) in the ECG-respiration cross feature pair or the eye movement-ECG cross feature pair; P(x,y) is the joint probability distribution of X and Y; P(x) and P(y) are the marginal probability distributions of X and Y, respectively.

[0261] The specific calculation steps may include: (1) dividing continuous feature values ​​(such as heart rate 50-120 bpm) into 10 bins (intervals) and converting them into discrete probability distributions. (2) counting the frequency of occurrence of two features in each bin combination to obtain P(x,y). (3) counting the frequency of occurrence of a single feature in each bin to obtain P(x) and P(y). KL divergence solution: substituting into the formula to calculate the mutual information value and quantify the correlation strength.

[0262] Step c4: Calculate the attention weights corresponding to the ECG-respiration cross feature pairs and the eye-movement-ECG cross feature pairs based on the third and fourth mutual information values.

[0263] Specifically, electronic devices can calculate the attention weights corresponding to the ECG-respiration cross-feature pairs and the eye-movement-ECG cross-feature pairs based on the third and fourth mutual information values.

[0264] For example, suppose a certain type of feature pair contains n strongly correlated pairs, and the mutual information of the i-th feature pair is MI. i Then its weight is: .

[0265] Example: The ECG-respiratory feature pair contains two strongly correlated pairs: heart rate-respiratory rate (MI=0.55). LF / HF - respiratory depth (MI=0.65). The weights are: w1=0.55 / (0.55+0.65)=0.458, w2=0.65 / 1.2=0.542.

[0266] Step c5: Generate the third hidden feature based on the attention weights corresponding to the ECG-respiration cross feature pairs and the eye-movement-ECG cross feature pairs, respectively.

[0267] Specifically, the electronic device can perform a weighted summation of the two feature values ​​of each cross feature pair: let w be the weight of the feature pair (X,Y), then the fused feature is F=w•X+(1-w)•Y (X and Y have been standardized to [0,1]).

[0268] Then, the fused features of all the cross feature pairs are spliced ​​together to form the third hidden feature.

[0269] Step S3036: Input the third hidden feature into the output layer of the initial cognitive load identification network.

[0270] Specifically, the electronic device inputs the third hidden feature into the output layer of the initial cognitive load recognition network.

[0271] Step S3037: The output layer outputs the predicted cognitive load level corresponding to the target training dataset based on the third hidden feature.

[0272] Specifically, step S3037 above may include the following steps: Step d1 involves extracting features from the third hidden feature to obtain the load-sensitive feature and the stability feature.

[0273] Specifically, electronic devices can extract features from the third hidden feature to obtain load-sensitive features and stability features, laying the foundation for multi-branch output.

[0274] Among them, load-sensitive features refer to features that respond significantly to changes in cognitive load, such as a rapid increase in the LF / HF ratio, the rate of pupil dilation, and the rate of heart rate increase. Electronic devices can use a bandpass filter (high-frequency cutoff frequency of 0.5Hz) to retain the rapidly changing components in the features, and filter out features with variance greater than a preset variance threshold (reflecting high volatility), thus obtaining the load-sensitive features.

[0275] Stability characteristics refer to baseline features that remain relatively stable under varying loads, such as baseline resting heart rate, mean respiratory rate, and resting pupil diameter. A low-pass filter (low-frequency cutoff of 0.1Hz) is used to retain slowly changing components within the features, and features with variances less than or equal to a preset variance threshold (reflecting low volatility) are selected, ultimately yielding the stability characteristics.

[0276] Step d2: Input the load sensitivity feature and stability feature into the main branch of the output layer.

[0277] Specifically, the electronic device can splice load-sensitive features and stability features, and input the spliced ​​features into the main branch of the output layer.

[0278] Step d3: The main branch outputs the first initial probability of each predicted cognitive load level corresponding to the load-sensitive features and stability features through a fully connected network.

[0279] Specifically, the main branch contains two fully connected layers. The first layer compresses the concatenated load-sensitive and stability features. The second layer in the fully connected network uses a Softmax activation function to convert the output vector into the first initial probability for each predicted cognitive load level, as shown in the formula: Where zk is the output value of the second fully connected layer network. This represents the probability of being predicted as the k-th level of load, where k is the predicted cognitive load level.

[0280] Step d4: Based on each initial probability, output the predicted cognitive load level corresponding to the target training dataset.

[0281] Specifically, step d4 above may include the following steps: Step d41: Input the load sensitivity feature and stability feature into the auxiliary branch in the output layer.

[0282] Specifically, the electronic device splices load-sensitive features and stability features, and inputs the spliced ​​features into the auxiliary branch in the output layer.

[0283] Step d42: The auxiliary branch extracts dynamic features from the load sensitivity features and stability features.

[0284] Among them, the dynamic characteristic definition can reflect the trend of load characteristics changing over time, such as the 5-second moving average slope of the LF / HF ratio and the cumulative change in pupil diameter.

[0285] The auxiliary branch can use a 1-layer LSTM network (10 hidden units) to perform time-series modeling of the concatenated load-sensitive features and stability features of the input, capture the feature change patterns of continuous time steps (such as the first 3 time points), and output dynamic features.

[0286] Step d43: Calculate the trend of change based on dynamic features.

[0287] Specifically, electronic devices can calculate three types of trend indicators of load changes based on dynamic characteristics. These three types of trend indicators can include load increase trend, load stability trend, and load decrease trend.

[0288] For an increasing load trend, the electronic device can calculate the proportion of positively changing features in the dynamic characteristics (such as the proportion of time when the heart rate increase is greater than 0); for a stable load trend, the electronic device can calculate the proportion of features whose change amplitude is less than a preset change amplitude in the dynamic characteristics (such as the proportion of time when the respiratory rate fluctuates within ±1 rpm); for a decreasing load trend, the electronic device can calculate the proportion of negatively changing features in the dynamic characteristics (such as the proportion of time when the pupil diameter constricts).

[0289] Step d44: Based on the changing trend, output the second initial probabilities of load increase, load stabilization, and load decrease corresponding to the load sensitivity and stability characteristics; Specifically, the fully connected layer of the auxiliary branch can convert dynamic features into a 3D vector based on the calculated trend of change, and activate the second initial probability of load increase, load stabilization and load decrease through Softmax activation.

[0290] Step d45: Based on the first initial probability and the second initial probability, output the predicted cognitive load level corresponding to the target training dataset.

[0291] Specifically, the electronic device can adjust the first initial probability of the main branch according to the trend, and output the predicted cognitive load level corresponding to the target training dataset based on the adjustment result.

[0292] For example, assuming the predicted cognitive load level is 1-5, if the probability of load increase is greater than a preset increase probability threshold (which can be 0.5 or 0.6), then the probability of load level 4-5 is multiplied by a first weight (strengthening the weight of higher levels), and the probability of load level 1-2 is multiplied by a second weight (weakening the weight of lower levels), while the probability of load level 3 remains unchanged. Here, the first weight is greater than 1, and the second weight is less than 1. If the probability of load decrease is greater than a preset decrease probability threshold (which can be 0.5 or 0.6), then the probability of load level 1-2 is multiplied by a first weight, and the probability of load level 4-5 is multiplied by a second weight, while the probability of load level 3 remains unchanged. Here, the first weight is greater than 1, and the second weight is less than 1. If the probability of load stabilization is greater than a preset stabilization probability threshold (which can be 0.5 or 0.6), then the first initial probability remains unchanged.

[0293] For example, the first initial probability output by the main branch is [0.03, 0.05, 0.7, 0.15, 0.05]. In the second initial probability output by the auxiliary branch, the probability of predicted load increase (P=0.8) is adjusted to 0.7×1.0=0.7 for level 3 and 0.15×1.2=0.18 for level 4, ultimately still predicted as level 3 (highest probability). If the probability of predicted load increase output by the auxiliary branch is P=0.8 and the initial probability of level 4 is 0.3, then the adjusted probability of level 4 is 0.3×1.2=0.36, which may overtake the previous probability and become the final result.

[0294] Step S3038: Based on the predicted cognitive load level, train the initial cognitive load identification network to obtain the target cognitive load identification model.

[0295] Specifically, step S3038 above may include the following steps: Step e1: Obtain the current training round number.

[0296] Specifically, the electronic device can obtain the current training round number.

[0297] Step e2: Based on the current training round number, substitute the predicted cognitive load level into the target loss function corresponding to the current training round number.

[0298] Specifically, if the current training round number is less than the first training round number threshold, then the target loss function is the first loss function; The first loss function is: ; Among them, w i For weight information; y i The actual cognitive load level in the target training set; To predict cognitive load levels.

[0299] If the current training round number is greater than or equal to the first training round number threshold and less than the second training round number threshold, then the target loss function is the second loss function; the first training round number threshold is less than the second training round number threshold.

[0300] The second loss function is: L = a1L1 + a2L2; where... .

[0301] Where i and j represent cognitive load levels, and i=j indicates samples belonging to the same cognitive load level. This indicates that for samples of the same cognitive load level, the first difference in similarity between 0.8 and samples of the same cognitive load level is calculated, and the maximum value is selected from the calculated differences and 0, and then summed; i≠j indicates samples belonging to different cognitive load levels; This means that for samples with different cognitive load levels, the second difference between the similarity between the samples with different cognitive load levels and 0.3 is calculated, and the maximum value is selected from the calculated second difference and 0, and then summed; a1 and a2 are both coefficients, and a1+a2=1.

[0302] If the current training epoch is greater than or equal to the second training epoch threshold, then the objective loss function is the third loss function. The third loss function is: L = a1L1 + a2L2 + a3L3; where, Among them, the predicted feature value is the feature value corresponding to each sub-feature in the third hidden feature; the physiologically reasonable value is the physiologically reasonable range corresponding to each sub-feature in the third hidden feature.

[0303] Step e3: Based on the target loss function, train the initial cognitive load identification network to obtain the target cognitive load identification model.

[0304] Specifically, electronic devices can train an initial cognitive load identification network based on a target loss function to obtain a target cognitive load identification model.

[0305] The cognitive load identification model training method provided in this application extracts features from the target training dataset and generates training fusion features. The input layer dynamically adjusts the weights of sub-features according to the target flight stage to obtain weighted fusion features. By adaptively adjusting the weights according to the flight stage, the model automatically strengthens features strongly related to the current stage (such as increasing the weight of eye movement features during landing) and weakens irrelevant features in different scenarios such as takeoff, cruise, and landing, thereby improving feature specificity and scenario adaptability. The first residual network extracts multimodal features in groups to ensure the independence and integrity of ECG, respiration, and eye movement features and avoids modal information confusion. The second residual network dynamically adjusts the residual connections in combination with the signal-to-noise ratio. When there is noise interference (such as ECG signals being interfered with by EMG), it closes the residuals of low-quality modes to prevent noise amplification, while maintaining feature continuity through a feature repair mechanism. The third residual network divides the features into 6 subgroups and extracts sub-features in a targeted manner (such as the LF / HF ratio in the ECG frequency domain and the pupil dilation rate in eye movement). Multi-dimensional feature decomposition and recombination not only preserve single-modal details, but also enhance robustness through residual mechanisms, enabling the model to capture subtle physiological responses to cognitive load (such as the synergistic changes in LF / HF increase and pupil dilation caused by sympathetic nerve activation under high load).

[0306] Then, convolution is performed on each subgroup to extract key features. Max pooling is used to retain significant features (such as the peak value of respiratory rhythm) and dimensionality is reduced. Regularization is combined to prevent overfitting. Feature redundancy is reduced (e.g., from 14-dimensional ECG frequency domain features to 6-dimensional core features), while feature distribution is standardized to make model training more stable and improve the generalization ability to small sample data.

[0307] Next, cross-feature pairs of ECG-respiration and eye-tracking-ECG are selected from the target residual features, and the correlation strength is quantified using mutual information. Focusing on cross-modal physiological correlation patterns overcomes the limitations of single-modality models, providing a more comprehensive physiological basis for cognitive load assessment. Attention weights are assigned based on mutual information values ​​(higher mutual information results in greater weight), and cross-feature pairs are weighted and fused to generate a third hidden feature. The contribution of strongly correlated feature pairs is automatically strengthened (e.g., increased weighting of pupil dilation and LF / HF), while weakly correlated information is weakened, making the model focus more on the core physiological markers of cognitive load and improving feature discriminative power.

[0308] Finally, the main branch outputs the first initial probability of each predicted cognitive load level through a fully connected network, directly quantifying the absolute level of the current load. This provides a clear level classification, meeting the intuitive requirements of cognitive load assessment. The auxiliary branch extracts dynamic features and calculates the second initial probabilities of load increase, load stabilization, and load decrease, supplementing the temporal change information of the load. Combining trend information to correct level predictions improves prediction accuracy in rapidly changing scenarios. The current training epoch is obtained; based on the current training epoch, the predicted cognitive load level is substituted into the target loss function corresponding to the current training epoch; based on the target loss function, the initial cognitive load identification network is trained to obtain the target cognitive load identification model. This ensures the accuracy of the obtained target cognitive load identification model. The above method combines static levels and dynamic trends to maintain stable prediction accuracy in complex flight scenarios; it resists noise interference and data fluctuations through signal-to-noise ratio filtering, feature repair, and regularization; it constrains features within physiological ranges to ensure that the prediction results conform to the physiological mechanism of human cognitive load; and it adapts to the needs of different task complexities and training stages through flight phase weight adjustment and phased loss.

[0309] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for training a cognitive load identification model, characterized in that, The method includes: Obtain the initial training dataset corresponding to the pilots being trained. The initial training dataset includes the initial electrocardiogram data, initial respiratory data, initial eye movement data, training flight phases, and training cognitive load levels corresponding to the pilots being trained. Feature extraction is performed on the initial ECG data in the initial training dataset to obtain the total initial ECG features corresponding to the initial ECG data; the total initial ECG features include the initial ECG time domain features, the initial ECG frequency domain features, and the initial ECG nonlinear features. Feature extraction is performed on the initial respiratory data in the initial training dataset to obtain the total initial respiratory features corresponding to the initial respiratory data; the total initial respiratory features include the initial respiratory time domain features and the initial respiratory frequency domain features. Feature extraction is performed on the initial eye-tracking data in the initial training dataset to obtain the initial eye-tracking features; The initial electrocardiogram features, initial respiratory features, and initial eye movement features are fused to generate initial physiological features; Based on the operational elements of each stage of the five-plane flight, task complexity characteristics are generated. The initial condition features are generated by fusing the characteristics of the training flight phase, the training cognitive load level, and the task complexity. The initial physiological features and initial conditional features are fused to generate initial fused features; The initial fused features are input into the first hidden layer of the preset generator network, and the initial fused features are expanded to generate the first hidden features; Calculate the first mutual information value among the initial sub-fusion features in the initial fusion feature; Based on the first mutual information values, construct the modal correlation matrix; Based on the mapping relationship between each first hidden sub-feature in the first hidden feature and each initial sub-fusion feature in the initial fusion feature, each first hidden sub-feature is labeled; Based on the labeling results, the mean vectors corresponding to each first hidden sub-feature dominated by ECG are calculated and used as the first cluster centers. Calculate the mean vector of each first hidden sub-feature dominated by breathing, and use it as the second cluster center; Calculate the mean vector of each first hidden sub-feature dominated by eye movement, and use it as the third cluster center; The mean vector of each of the first hidden sub-features dominated by the calculation conditions is used as the fourth cluster center; Based on the modal correlation matrix, the second mutual information value between each first hidden sub-feature is calculated; Based on the second mutual information value, each first hidden sub-feature is clustered to obtain the ECG cluster group, the respiration cluster group, the eye movement cluster group, and the conditional cluster group; Output the second hidden feature based on the ECG cluster, respiratory cluster, eye movement cluster, and conditional cluster; The second hidden feature is activated to generate noisy data. The noise data is input into the preset discrimination network in the preset data augmentation model to discriminate the noise data; Based on the discrimination results, the initial training dataset is expanded to generate the target training dataset; Based on the target training dataset, the initial cognitive load identification network is trained to obtain the target cognitive load identification model.

2. The method according to claim 1, characterized in that, The step of inputting the initial fused features into the first hidden layer of the preset generation network, and expanding the initial fused features to generate the first hidden features includes: The training flight phase and the training cognitive load level are determined from the initial fusion features; Based on the training flight phase and the training cognitive load level, determine the target expansion dimension corresponding to the initial fusion feature from a preset mapping table; A first initial weight matrix is ​​generated based on the initial dimension corresponding to the initial fusion feature and the target extended dimension; Based on the training flight phase and the training cognitive load level, the first initial weight matrix is ​​modified to obtain the first target weight matrix; Based on the first target weight matrix, the initial fusion features are expanded to generate a target extended dimension feature vector. The target extended dimension feature vector is activated based on a preset activation function, and the first hidden feature is output.

3. The method according to claim 1, characterized in that, The step of outputting the second hidden feature based on the ECG cluster group, the respiration cluster group, the eye movement cluster group, and the conditional cluster group includes: Based on the modal correlation matrix, the average correlation degree among the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the conditional cluster group is calculated; Based on the average correlation degree, the connection weights between the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the conditional cluster group are determined; A second initial weight matrix is ​​generated based on the initial dimension corresponding to the initial fusion feature and the target extended dimension; The second initial weight matrix is ​​divided into intra-group and inter-group regions; Based on the connection weights between the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the respiratory cluster group, determine the first target weight value corresponding to the inter-group region in the second initial weight matrix; Based on the criticality of each cluster sub-feature in the ECG cluster group, the respiratory cluster group, the eye movement cluster group, and the conditional cluster group, the second target weight value corresponding to the region within the group is determined; Based on the first target weight and the second target weight value, a second target weight matrix is ​​generated; The second hidden feature is obtained by multiplying the first hidden feature by the second target weight matrix.

4. The method according to claim 1, characterized in that, The step of inputting the noise data into the preset discrimination network in the preset data augmentation model and discriminating the noise data includes: Based on the initial electrocardiogram data and the initial eye movement data, calculate the rate of load change; The noise data, the training flight phase, the training cognitive load level, and the load change rate are fused to generate target fusion features; Based on the first branch, the target fusion features are compressed to obtain the first branch features; Based on the second branch, the target fusion features are compressed to obtain the second branch features; The target fusion features are divided into four specific groups and labeled according to physiological modalities and functional attributes; the four specific groups are the electrocardiogram group, the respiratory group, the eye movement group, and the load change rate label group. The first branch feature and the second branch feature are verified respectively to obtain the ECG dimension score, respiratory dimension score and eye movement dimension score corresponding to the first branch feature and the second branch feature respectively; The first branch features, the second branch features, the ECG dimension score, the respiratory dimension score, and the eye movement dimension score are fused together to obtain the discriminative fusion features; The discriminative fusion features are input into the authenticity discrimination branch, and the probability that the noisy data is real data is output. The discriminative fusion features are input into the load level discrimination branch, and the load level probability corresponding to the noise data is output. The noise data is judged based on a preset loss function.

5. The method according to claim 1, characterized in that, The step of training the initial cognitive load identification network based on the target training dataset to obtain the target cognitive load identification model includes: Feature extraction is performed on the target training dataset to generate training fusion features; The trained fusion features are input into the input layer of the initial cognitive load identification network; The input layer adjusts the weight information corresponding to each sub-training fusion feature in the training fusion feature according to the target flight stage in the target training dataset to obtain the weighted fusion feature; The weighted fusion features are input into the third hidden layer of the initial cognitive load identification network; The third hidden layer extracts features from the weighted fusion features to obtain the third hidden features; The third hidden feature is input into the output layer of the initial cognitive load identification network; The output layer outputs the predicted cognitive load level corresponding to the target training dataset based on the third hidden feature; Based on the predicted cognitive load level, the initial cognitive load identification network is trained to obtain the target cognitive load identification model.

6. The method according to claim 5, characterized in that, The third hidden layer extracts features from the weighted fusion features to obtain the third hidden features, including: The residual network in the third hidden layer extracts features from the weighted fusion features and outputs the target residual features. From the target residual features, ECG-respiration crossover feature pairs and eye movement-ECG crossover feature pairs are selected; Calculate the third mutual information value corresponding to the ECG-respiratory crossover feature pair and the fourth mutual information value corresponding to the eye movement-ECG crossover feature pair; Based on the third mutual information value and the fourth mutual information value, calculate the attention weights corresponding to the ECG-respiratory cross feature pair and the eye movement-ECG cross feature pair, respectively. The third hidden feature is generated based on the attention weights corresponding to the ECG-respiration cross feature pairs and the eye-movement-ECG cross feature pairs, respectively.

7. The method according to claim 6, characterized in that, The residual network includes a first residual network, a second residual network, and a third residual network. The residual network in the third hidden layer extracts features from the weighted fusion features and outputs target residual features, including: The first residual network in the third hidden layer performs multimodal feature extraction on the weighted fusion features to obtain multimodal features; The second residual network in the third hidden layer compresses the multimodal features to obtain compressed features, and then performs a convolution operation on the compressed features to obtain convolutional features. Calculate the signal-to-noise ratio corresponding to each modal feature in the weighted fusion features; Based on the signal-to-noise ratio corresponding to each modal feature, the compressed feature and the convolutional feature are subjected to residual superposition processing to generate superimposed features; The superimposed features are divided based on the third residual network in the third hidden layer into ECG time domain subgroup, ECG frequency domain subgroup, respiratory rhythm subgroup, respiratory depth subgroup, eye movement pupil subgroup, and eye movement fixation subgroup. Each subgroup is processed by convolution, and the convolution results are compressed by max pooling and regularization to obtain the preserved features corresponding to each subgroup; The retained features corresponding to each subgroup are fused to obtain the target residual features.

8. The method according to claim 5, characterized in that, The output layer outputs the predicted cognitive load level corresponding to the target training dataset based on the third hidden feature, including: Feature extraction is performed on the third hidden feature to obtain load-sensitive features and stability features; The load sensitivity feature and the stability feature are input into the main branch of the output layer; The main branch outputs the first initial probability of each predicted cognitive load level corresponding to the load sensitivity feature and the stability feature through a fully connected network; Based on each of the first initial probabilities, output the predicted cognitive load level corresponding to the target training dataset.

9. The method according to claim 8, characterized in that, The step of outputting the predicted cognitive load level corresponding to the target training dataset based on each of the first initial probabilities includes: The load sensitivity feature and the stability feature are input into the auxiliary branch of the output layer; The auxiliary branch extracts dynamic features from the load-sensitive features and the stability features; The changing trend is calculated based on the aforementioned dynamic characteristics; Based on the changing trend, output the second initial probabilities of load increase, load stabilization, and load decrease corresponding to the load sensitivity feature and the stability feature; Based on the first initial probability and the second initial probability, the predicted cognitive load level corresponding to the target training dataset is output.

10. The method according to claim 5, characterized in that, The step of training the initial cognitive load identification network based on the predicted cognitive load level to obtain the target cognitive load identification model includes: Get the current training round number; Based on the current training round number, the predicted cognitive load level is substituted into the target loss function corresponding to the current training round number; Based on the target loss function, the initial cognitive load identification network is trained to obtain the target cognitive load identification model.