Method and device for evaluating flight risk index
Patent Information
- Application Number
- CN202610980315.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本发明提供了一种飞行风险指数评估方法及装置,以解决如何准确评估飞行风险的问题
Smart Images

Figure CN122677166A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of flight risk assessment technology, specifically to a method and apparatus for assessing flight risk index. Background Technology
[0002] In the modern, highly automated civil aviation system, the iterative upgrades of aircraft automation equipment and human-machine interface systems have effectively reduced flight hazards caused by hardware failures, making human factors increasingly the core constraint on flight safety. Throughout the entire flight mission, pilots continuously face multiple stressors, including high workloads, complex operating environments, and unexpected flight events, making them highly susceptible to negative emotions such as anxiety, fear, and anger. These negative emotions directly lead to an imbalance in the allocation of cognitive resources, biased decision-making, and decreased operational accuracy, significantly increasing the probability of human error and seriously threatening civil aviation flight safety.
[0003] Currently, a large number of studies have been conducted both domestically and internationally on pilot emotion monitoring and assessment. Most mainstream solutions rely on single-modal data to classify pilot emotions and assess stress, resulting in low accuracy of assessment results and thus failing to accurately assess flight risks.
[0004] Therefore, how to accurately assess flight risks has become an urgent technical problem to be solved. Summary of the Invention
[0005] This invention provides a method and apparatus for assessing flight risk index, in order to solve the problem of how to accurately assess flight risk.
[0006] In a first aspect, the present invention provides a method for assessing a flight risk index, the method comprising: Acquire multi-frame facial image data and electrocardiogram signals corresponding to the target pilot; Identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot; Identify the target's electrocardiogram signal and extract the target's physiological characteristics corresponding to the pilot; Based on comprehensive facial features and target physiological features, determine the emotion probability vector corresponding to the target pilot; Based on the emotion probability vector, the current flight risk index corresponding to the target pilot is determined.
[0007] The flight risk index assessment method provided in this application acquires multi-frame target facial image data and target electrocardiogram (ECG) signals. It employs simultaneous dual-modal acquisition of facial vision and ECG / physiological data. Facial images capture millisecond-level instantaneous stress expressions, while ECG signals capture latent long-term stress that surface expressions cannot reveal, providing a reliable data source foundation for accurate feature extraction. The method identifies target facial images and extracts comprehensive facial features, fully preserving the details of the pilot's overt emotional dynamics. ECG signals are identified to extract target physiological features, accurately representing autonomic nervous system fluctuations and latent suppressed emotions, compensating for the limitation of facial recognition in detecting internal high pressure. An emotion probability vector is calculated by combining comprehensive facial features and physiological features. This approach takes into account both overt expressions and latent physiological states, significantly improving the accuracy and anti-interference capabilities of emotion recognition compared to single-method identification. The current flight risk index is calculated based on the emotion probability vector, enabling personalized, condition-based, and quantitative risk grading. A standardized risk index of 0-1 is output, which can be directly connected to the airborne early warning system to intuitively quantify the safety impact of emotions on flight operations, supporting crew status monitoring and safety intervention.
[0008] In one optional implementation, the facial image data of each target are identified, and the comprehensive facial features corresponding to the target pilot are extracted, including: The facial image data of each target is identified, and the static facial motion unit feature vector corresponding to each facial image data is determined. The activation intensity difference of the same facial action unit between the feature vectors of static facial action units in two adjacent frames is calculated to obtain the inter-frame temporal difference feature. The static facial motion unit feature vector is fused with the inter-frame temporal difference feature to obtain the comprehensive facial features.
[0009] The flight risk index assessment method provided in this application determines the feature vector of a single-frame static facial action unit. It quantifies the activation level of nine key AUs based on the FACS standard, accurately quantifying the instantaneous basic facial expression state; inheriting the weighted parameters of previous occlusion, it weakens the interference from flight equipment occlusion areas, retaining only the core emotional information of the eyebrows, eyes, and lips, establishing a stable and quantifiable static emotional benchmark. It calculates the difference in activation intensity of the same AU in adjacent frames to obtain inter-frame temporal difference features. It captures the rate and magnitude of dynamic changes in facial expressions, distinguishing between stable emotions and instantaneous fright or rage caused by bird strikes or malfunctions; it overcomes the deficiency of static AUs, which can only reflect the current state and cannot reflect the speed of emotional fluctuations, significantly improving sensitivity to sudden stress conditions. The static vector and temporal difference features are fused to obtain comprehensive facial features. It takes into account both long-term calm / anxiety states and millisecond-level emotional abrupt changes, providing more complete feature information and offering a complete facial emotional representation for subsequent multimodal attention matching and emotion classification, improving recognition and discrimination.
[0010] In one optional implementation, the target electrocardiogram signal is identified, and the target physiological characteristics corresponding to the target pilot are extracted, including: Identify the target ECG signal and locate the R wave from the target ECG signal; Obtain a reasonable range for the preset RR interval; If the detected peak exceeds the reasonable range of the preset RR interval, it is determined to be a noise spurious peak and is removed to obtain the ECG signal to be extracted. Feature extraction is performed on the ECG signal to be extracted to obtain the time-domain features, frequency-domain features, basic nonlinear features, and preset nonlinear features corresponding to the ECG signal to be extracted. The time-domain features include the average value of all RR intervals and the root mean square of the difference between adjacent RR intervals. The frequency-domain features include the first frequency power, the second frequency power, and the ratio between the first frequency power and the second frequency power; the first frequency is less than the second frequency. The basic nonlinear features include the standard deviation of the minor axis of the Poincaré plot and the standard deviation of the major axis of the Poincaré plot. The preset nonlinear features include the eccentricity of the Poincaré plot ellipse and the sample entropy of the RR interval. By fusing time-domain features, frequency-domain features, basic nonlinear features, and preset nonlinear features, the target physiological features corresponding to the target pilot are obtained.
[0011] The flight risk index assessment method provided in this application identifies electrocardiogram (ECG) signals and locates the R wave. Utilizing a modified Pan-Tompkins algorithm combined with neighborhood verification, it accurately marks heartbeat peaks, mitigating false positives and false negatives caused by cockpit motion artifacts, and providing a precise temporal coordinate basis for RR interval calculation. It obtains reasonable RR interval ranges. By dynamically setting thresholds based on pilot resting and high-stress physiological states, it conforms to the physiological limits of human heartbeat, establishing a reliable true / false interval judgment scale suitable for different flight load conditions. False peaks are removed to obtain the ECG signal to be extracted. Abnormal peaks are iteratively screened out, purifying the RR interval time series and eliminating noise distortion interference with subsequent heart rate variability calculations, ensuring all feature calculations are based on the true heartbeat rhythm. ECG features are extracted in multiple dimensions, including the time domain, frequency domain, and two types of nonlinearity. This system comprehensively characterizes the autonomic nervous system state from multiple dimensions: the time domain reflects average heart rate and instantaneous heart rate fluctuations; the frequency domain quantifies the balance between the sympathetic and parasympathetic nervous systems; Poincaré double standard deviation distinguishes between long-term and short-term heart rate variability; eccentricity and sample entropy accurately differentiate between long-term anxiety disorders and instantaneous startle fluctuations, with multiple complementary indicators covering implicit emotional physiological representations. Multiple feature classes are fused to generate target physiological features. Z-score standardization eliminates dimensional differences, and orderly splicing forms a unified physiological vector; weights are simultaneously adjusted based on the signal-to-noise ratio, allowing high-quality signals to fully play their role, resulting in stable and regular overall features that can be directly matched and fused with facial visual features across modalities.
[0012] In one optional implementation, based on comprehensive facial features and target physiological features, an emotion probability vector corresponding to the target pilot is determined, including: The comprehensive facial features are input into the facial feature recognition branch, and the target visual feature vector is output. The target physiological characteristics are input into the ECG feature recognition branch, and the ECG feature vector is output. The target visual feature vector and electrocardiogram feature vector are fused to obtain an initial fused feature vector; Obtain the current flight phase; The current flight phase is concatenated with the initial fused feature vector to obtain the fused phase feature vector; Based on the feature vectors from the fusion stage, the emotional probability vector corresponding to the target pilot is determined.
[0013] The flight risk index assessment method provided in this application provides a visual feature vector as the input of comprehensive facial features. The branch incorporates dynamic-static separation encoding, ROI occlusion weight inheritance, AU emotion combination enhancement, and a transient mutation amplification mechanism to deeply refine explicit transient facial emotion information, uniformly map to fixed dimensions, strengthen the ability to recognize sudden expression changes under dangerous situations, and shield against noise from flight equipment occlusion. The branch also outputs an ECG feature vector as the input of physiological features. It employs grouped 1D-CNN stationary convolution, SNR adaptive correction, and implicit anxiety steady-state enhancement to specifically mine long-term suppressed emotions and autonomic nervous system steady-state states that cannot be reflected in the face, forming a dynamic-static complementarity with visual features. Alignment of dimensions with the visual vector facilitates fusion calculation. The visual and ECG vectors are fused to obtain an initial fused feature vector. Fusion weights are adaptively allocated based on facial dynamic intensity and modal confidence. In cases of sudden danger, transient facial features are prioritized, while in stable conditions, implicit stress signals from the ECG are emphasized, balancing the reliability of both modalities and outputting a multimodal initial fused feature vector. The current flight phase is then obtained. By capturing prior information on climb, cruise, and approach conditions, and recognizing the significant differences in control tolerance and emotional interference levels across different phases, the model provides a basis for scenario determination in subsequent attention constraints and risk weighting. The flight phase encoding is concatenated with the initial fusion features to obtain a fused phase feature vector. Embedding scenario information from the flight phase, the model can distinguish differences in emotional expression under different control loads, accurately identifying subtle negative emotions deliberately suppressed during the approach phase, thus improving scenario adaptability. An emotion probability vector is output based on the fused phase feature vector. Combined with scene masking, cross-modal attention, temporal compensation correction, and multi-layer regularization fitting, six standard probability distributions of emotions are output, simultaneously considering external facial expressions, internal physiological factors, and flight conditions, significantly improving emotion recognition accuracy and anti-interference capabilities.
[0014] In one optional implementation, based on the feature vector from the fusion phase, the emotion probability vector corresponding to the target pilot is determined, including: The target visual feature vector in the feature vector of the fusion stage is used as the query vector, and the electrocardiogram feature vector in the feature vector of the fusion stage is used as the key vector and value vector. Based on the query vector, key vector, and value vector, the cross-modal attention weights are calculated to obtain the target cross-modal attention weight matrix; Based on the current flight stage in the feature vector of the fusion stage, the target cross-modal attention weight matrix is corrected to obtain the initial attention weight heatmap; the horizontal axis of the attention weight heatmap corresponds to the visual feature region, the vertical axis corresponds to the electrocardiogram feature region, and the color intensity represents the magnitude of the attention weight. Temporal weight smoothing is performed on the initial attention weight heatmap to obtain the target attention weight heatmap; Based on the target attention weight heatmap, the target visual feature vector and electrocardiogram feature vector are fused to obtain a global fused feature vector; The global fusion feature vector is regularized to obtain a unified fusion feature vector. Temporal feature processing is performed on the unified fusion feature vector to obtain the temporal feature vector; Based on the temporal feature vector, the emotional probability vector corresponding to the target pilot is determined.
[0015] The flight risk index assessment method provided in this application allocates Q-visual queries and K / V ECG key value vectors. It uses external facial expressions as query anchors to match internal ECG physiological signals, conforming to the correspondence between "external facial expressions and endogenous neural fluctuations" in human emotions, thus constructing a reasonable cross-modal matching logic. A target cross-modal attention weight matrix is calculated. First, the original similarity is calculated, then invalid features are masked based on the flight scenario, and the weights are adjusted by combining image and ECG confidence scores. Dual constraints on operating conditions and acquisition quality reduce interference from irrelevant features and improve modal matching accuracy. The weight matrix is adjusted according to the flight stage, and an initial attention heatmap is generated. The overall matching strength is fine-tuned according to climb / cruise / approach load, amplifying weak emotional correlations during high-pressure approaches. The heatmap intuitively presents the correlation strength between visual AU and ECG indicators, improving model interpretability. The initial heatmap is time-smoothed to obtain the target heatmap. Multi-frame weighted averaging is used to suppress weight jumps caused by single-frame occlusion and ECG pulse noise, ensuring a stable and continuous heatmap and preventing drastic inter-frame jitter in emotion recognition results. Global fusion features are obtained by weighted fusion of visual and ECG data using a smoothed heatmap. Instead of simple addition, attention-based selection of highly correlated feature pairs is used. In dangerous situations, the combination of fear / surprise (AU) and high-stress ECG is emphasized; during cruise, a pleasant and stable heart rate is prioritized, resulting in higher information utilization. Global fusion features are regularized to obtain unified fusion features. Layer normalization balances the differences in visual and ECG scales, and Dropout suppresses overfitting, enhancing the model's generalization stability under cockpit occlusion and motion artifact environments. Temporal segmentation compensation is applied to the unified fusion features to obtain temporal feature vectors. Multiple subsequences are split to select high-contribution core vectors, and the level of control anomaly is determined by comparing it with the pilot's personal stability benchmark. Adaptive temporal compensation is enabled between frames to smooth out single-frame feature distortion, balancing detail and temporal continuity. The temporal feature vector outputs an emotion probability vector. By fusing multi-modal, flight condition, temporal smoothing, and individual control benchmark information, six standard probabilities—calm, pleasant, surprised, anxious, angry, and fear—are output, with recognition accuracy and anti-interference capabilities far exceeding those of the single-modal approach.
[0016] In one optional implementation, cross-modal attention weights are calculated based on the query vector, key vector, and value vector to obtain the target cross-modal attention weight matrix, including: Calculate the original attention score matrix based on the query vector and key vector; Obtain current flight parameters; Based on the current flight parameters, determine the current scenario type; Based on the current scene type, the original attention score matrix is masked to obtain the masked attention score matrix; The mask attention score matrix is normalized to obtain the basic cross-modal attention weight matrix; The target cross-modal attention weight matrix is obtained by modifying the modal confidence scores corresponding to the target visual feature vector and ECG feature vector, respectively.
[0017] The flight risk index assessment method provided in this application calculates the original attention score matrix using query vectors and key vectors. It quantifies the matching similarity between visual and ECG features in each dimension by scaling the dot product, accurately quantifying the intrinsic correlation between facial and physiological emotional representations, and establishing a basic numerical system for cross-modal matching. It acquires current flight parameters by collecting multi-dimensional quantitative data on attitude, power, control, and alarms, providing objective quantitative basis for scenario classification and avoiding operational condition judgment biases caused by relying solely on stage labels. The flight parameters determine the current scenario type. It automatically distinguishes between three scenarios: stable cruise, moderate load operation, and sudden high-stress failure, determining the pilot's flight scenario or operational load state and providing scenario labels for attention differentiation constraints. The scenario type is used to mask the original score matrix to obtain a masked score matrix. Invalid matching channels are specifically masked for different scenarios; cruise suppresses ECG stress index interference, and failure conditions mask pleasant facial features, reducing the pairing weights of features with low correlation to the current flight scenario and reducing noise matching interference. The masked scores are normalized to generate a basic attention weight matrix. Softmax normalization transforms the scores into a standard weight distribution of 0-1 with a row sum of 1, unifying the attention scale across dimensions and ensuring that the weight values can stably participate in subsequent feature weighting calculations. Modal confidence scores are used to correct the base weights, resulting in the target attention weight matrix. Weights are adaptively adjusted based on the degree of face occlusion and the signal-to-noise ratio; image occlusion reduces the contribution weight of the visual modality to the fusion result, and large ECG artifacts weaken the contribution of ECG matching. This adaptive balance of the differences in acquisition quality between the two modalities improves robustness under harsh acquisition environments.
[0018] In one optional implementation, temporal feature processing is performed on the unified fused feature vector to obtain a temporal feature vector, including: The unified fused feature vector is divided into multiple sub-sequences; For each subsequence, the distribution difference of all query vectors and key vectors within the subsequence is calculated, and the core vectors with a contribution greater than the preset contribution threshold are selected. Weighted calculations are performed on each core vector to obtain the preliminary coding features corresponding to the subsequences; For each preliminary coding feature, the baseline flight parameters corresponding to the target pilot are obtained; the baseline flight parameters are flight parameters collected when the target pilot is calm and the target aircraft is flying steadily. Compare the current flight parameters corresponding to each preliminary coding feature with the baseline flight parameters; Based on the comparison results, determine the anomaly level corresponding to the current flight parameters; By binding the anomaly level with the initial coding features, the first backup coding features corresponding to each subsequence are obtained; Based on the anomaly level corresponding to each first backup coding feature, the compensation interval corresponding to each first backup coding feature is determined, and the second backup coding feature corresponding to the compensation interval is obtained; the compensation interval represents the number of preset frames before and after each subsequence. Weighted fusion of each second backup coding feature yields the weighted compensation feature corresponding to the subsequence; The first backup coding feature and the weighted compensation feature are weighted and fused to obtain the sub-time series feature vector corresponding to the sub-sequence; The sub-time series feature vectors are fused together to obtain the time series feature vector.
[0019] The flight risk index assessment method provided in this application uniformly integrates features and splits them into multiple subsequences. This enables refined analysis of feature blocks, processing visual and ECG feature segments separately, facilitating local filtering, correction, and temporal compensation, and avoiding the loss of local difference information through single-vector processing. Q and K distribution differences within subsequences are calculated to filter high-contribution core vectors. The feature discrimination capability is quantified based on distribution differences, filtering out low-contribution redundant vectors and retaining only matching features with high emotional recognition accuracy, reducing invalid calculations and noise interference. Weighted core vectors generate preliminary encoded features for the subsequences. Visual and ECG core vectors are fused using contribution as weight, adaptively balancing the proportion of local modalities to generate basic encoding with local emotional semantics. Baseline flight parameters for the pilot's calm and stable state are retrieved. A personalized steady-state reference standard is established for the pilot, abandoning a uniform group baseline and adapting to differences in pilot operating habits and physiological tolerance. Current parameters are compared item by item with baseline parameters. The actual deviation amplitude of control, attitude, and power is quantified to objectively measure the magnitude of operational stress, rather than relying on rough stage labeling. Parameter anomaly levels are classified based on deviations. Quantitative differentiation distinguishes between no abnormality, mild, moderate, and severe stress conditions, providing a grading basis for subsequent compensation amplitude and feature correction coefficients. Abnormality levels are bound to initial codes, generating the first backup code feature. The code amplitude is fine-tuned according to the stress level; high-pressure conditions moderately amplify emotional feature signals, while mild conditions undergo minor corrections, aligning with the influence of stress on emotional expression. Compensation intervals are matched with preceding and following frames according to the abnormality level, extracting the second backup code set. Higher stress levels result in a larger compensation frame range, relying on multi-frame temporal information to compensate for feature distortion caused by single-frame occlusion and ECG artifacts. The second backup codes within the interval are temporally weighted and fused to obtain weighted compensation features. Nearer frames have higher weights, while farther frames have lower weights, ensuring real-time response to sudden emotional changes while utilizing historical frames to smooth and suppress severe jitter in single frames. The primary and backup codes are weighted and fused to obtain a sub-temporal feature vector. A balance is struck between the true features of the current frame and the temporal smoothing compensation strength; the greater the abnormality, the higher the compensation ratio, while stable conditions are dominated by the current frame, balancing real-time performance and stability. All sub-temporal features are fused to obtain the overall temporal feature vector. By stitching together and restoring the complete feature dimensions, integrating all optimization information such as local filtering, personalized benchmark correction, and hierarchical temporal compensation, a continuous, stable, and more accurate global temporal sentiment representation is output.
[0020] In one optional implementation, the emotion probability vector corresponding to the target pilot is determined based on the temporal feature vector, including: The temporal feature vector is subjected to dimensionality reduction, nonlinear transformation, and regularization to obtain fully connected output features; The fully connected output features are fused with the temporal feature vector to obtain the hidden layer feature vector; The hidden layer feature vector is mapped to the emotion probability vector; the emotion probability vector includes at least one emotion type among calm, joy, surprise, anxiety, anger, and fear.
[0021] The flight risk index assessment method provided in this application obtains fully connected output features from temporal feature vectors through dimensionality reduction, nonlinear transformation, and regularization. Dimensionality reduction compresses high-dimensional redundant information; nonlinear transformation fits the complex coupling relationship between emotion, working conditions, and physiology; layer normalization and Dropout dual regularization suppress overfitting, filter acquisition noise, and extract highly recognizable deep emotional semantics. The fully connected output features are fused with the temporal feature vectors to obtain hidden layer feature vectors. A residual fusion structure is adopted, which retains both the deep features refined by dimensionality reduction and the details of the original temporal compensation and personalized correction, avoiding the loss of key emotional details during dimensionality reduction and resulting in more complete feature expression. The hidden layer feature vectors are mapped to six types of emotion probability vectors. By combining a classification layer with Softmax to output a normalized probability distribution, not only are single emotion labels output, but the confidence ratio of mixed emotions is fully presented, quantifying the strength of each emotion and providing accurate numerical input for subsequent staged risk weighting calculations.
[0022] In one optional implementation, the current flight risk index corresponding to the target pilot is determined based on the emotion probability vector, including: Obtain the current flight phase; Based on the current flight phase, determine the emotion weights corresponding to each emotion type in the emotion probability vector; The basic flight risk index is obtained by weighting the probability values of each emotion type in the emotion probability vector based on the emotion weight. Obtain the current operational deviation corresponding to the target pilot; the current operational deviation includes the standard deviation of yaw angle and the standard deviation of altitude. Based on the current operational deviation, determine the deviation correction parameters; Obtain the individual correction parameters corresponding to the target pilot; Based on the deviation correction parameters and individual correction parameters, the basic flight risk index is corrected to obtain the current flight risk index.
[0023] The flight risk index assessment method provided in this application obtains the current flight phase. It distinguishes between different control load scenarios such as climb, cruise, and approach, recognizing that the degree of harm of emotions to flight safety varies at different stages, providing a scenario-based basis for differentiated risk weighting. Weights are set for each emotion type based on the current flight phase. Under high-tolerance-pressure conditions such as approach, the weights of negative emotions such as fear, anger, and anxiety are amplified, while the negative impact is weakened during cruise, aligning with the actual impact patterns on flight safety and making the basic risk assessment more consistent with practical aviation logic. The basic flight risk index is calculated by weighted summation of emotion probabilities. The theoretical safety risk resulting from the superposition of various emotions is quantified, transforming discrete emotion probabilities into a single, comparable risk benchmark value. Yaw angle and altitude data are collected, and the corresponding standard deviations are calculated as operational deviations. The stability of heading and altitude control objectively reflects the actual interference of emotions on control actions, compensating for the one-sidedness of relying solely on emotions to determine risk. Deviation correction parameters are calculated based on operational deviations. The greater the control fluctuation, the higher the risk coefficient, achieving a linkage correction between emotional risk and actual control performance, ensuring the risk assessment aligns with real control conditions. Pilot-specific correction parameters are retrieved. Personalized calibration is performed based on individual baseline information such as pilots' flight experience, eliminating individual assessment errors caused by uniform group standards. The final flight risk index is obtained by correcting the base index using dual parameters. Integrating three dimensions—emotional hazard of the scenario, actual operational deviation, and individual personnel differences—it outputs a standardized, predictable risk value of 0-1, providing accurate and reliable results that can be directly integrated with airborne safety warning systems.
[0024] Secondly, the present invention provides a flight risk index assessment device, the device comprising: The acquisition module is used to acquire multi-frame facial image data and electrocardiogram signals of the target pilot. The first extraction module is used to identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot. The second extraction module is used to identify the target's electrocardiogram signal and extract the target's physiological characteristics corresponding to the pilot. The first determining module is used to determine the emotion probability vector corresponding to the target pilot based on comprehensive facial features and target physiological features; The second determination module is used to determine the current flight risk index corresponding to the target pilot based on the emotion probability vector. Attached Figure Description
[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the first process of the flight risk index assessment method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the second process of the flight risk index assessment method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the spectrum distribution according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a Poincaré diagram according to an embodiment of the present invention; Figure 5 This is a structural block diagram of a flight risk index assessment device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0029] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0030] According to an embodiment of the present invention, a method for assessing flight risk index is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0031] This embodiment provides a method for assessing flight risk index, which can be used in electronic devices, such as control devices in a target aircraft or flight simulator. Figure 1 This is a flowchart of a flight risk index assessment method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Acquire multi-frame target facial image data and target electrocardiogram signal corresponding to the target pilot.
[0032] Specifically, the electronic device continuously receives the continuous video stream output from the camera, timestamps it, and ensures that each frame is linked to the acquisition time. Then, it delineates the effective shooting area, locks onto the pilot's face, and removes irrelevant background to reduce subsequent computation. The electronic device can utilize temporal optical flow to calculate the pixel displacement trajectory between adjacent frames based on the pixel motion correlation between video frames. It performs continuous cross-frame tracking of key emotional features such as eyebrows, eyes, nose, and lips, locking onto the effective facial area throughout. It continuously outputs the tracked facial region image, ensuring stable facial position across consecutive frames and achieving spatiotemporal consistency of facial features.
[0033] Electronic devices can combine tracked facial key points to divide the face into core ROI regions and invalid occlusion regions. Core ROI regions include areas directly associated with emotional expression, such as eyebrows, the area around the eyes, and lips; invalid occlusion regions are facial areas obscured by hats, headsets, etc. Electronic devices can perform local pixel magnification and feature weight enhancement on core ROI regions to strengthen effective emotional features. For invalid occlusion regions, weight decay is applied to reduce noise interference from occlusion. All information from the original image is preserved, and effective features are highlighted only through weight optimization.
[0034] Finally, the electronic equipment can read flight parameters such as aircraft altitude, yaw angle, and engine status in real time, classifying flight scenarios into two main categories: stable flight and sudden failure / high-stress conditions. If the flight scenario is stable flight (normal climb, cruise), a low frame rate of 120Hz is used for frame subtraction. During this phase, the pilot's expression is stable, reducing the number of frames and lowering storage and computational costs. If the flight scenario is a sudden failure (bird strike, engine failure, runway incursion), it automatically switches to a high frame rate of 480Hz for frame subtraction. The high frame rate captures millisecond-level facial movement changes, avoiding the loss of stress-related emotional features. The electronic equipment can then extract image frames from the optimized video stream at the set frame rate to form the original frame sequence.
[0035] The electronic device can group discrete image frames obtained from frame extraction into temporal sequences, dividing continuous images into temporal frame groups with a fixed duration, each group containing a fixed 16 frames. This preserves the evolutionary patterns of facial expressions over a short period while matching the format requirements of subsequent convolutional networks for sequential input. The electronic device can uniformly crop and scale all single-frame images within each group, ultimately outputting 224×224 pixel target facial image data.
[0036] Electronic devices can acquire the pilot's real-time raw electrocardiogram (ECG) signals via electrodes and synchronously bind them with timestamps consistent with facial video, achieving temporal alignment between facial video and ECG signals. The electronic devices can also combine flight control data to label the current operational condition in real time. The current operational condition label can be static stable flight / high-load control (for subsequent differentiated noise reduction).
[0037] The electronic equipment can sequentially perform three layers of fixed filtering on all raw ECG (electrocardiogram) signals to eliminate common cockpit noise. Specifically, for the 0.5~100Hz fourth-order Butterworth bandpass filter: filtering out low-frequency baseline drift and high-frequency noise, retaining the effective ECG waveform frequency band; 50Hz IIR notch filter: specifically filtering out 50Hz mains frequency interference generated by cockpit electrical equipment, eliminating regular harmonics; adaptive median filter: eliminating spike pulse noise caused by poor equipment contact and transient electromagnetic interference, smoothing signal glitches. After completing the basic noise reduction, the signal has removed most of the common interference.
[0038] The electronic device can switch the wavelet decomposition level and soft threshold according to the flight conditions marked in advance, and add a QRS group protection mechanism to avoid motion artifacts and waveform distortion.
[0039] For static, stable flight conditions (with weak motion artifacts), electronic equipment can employ 5-level decomposition of db4 wavelets plus standard soft threshold coefficients for deep denoising. While filtering out residual noise, this process preserves the details of the original ECG waveform to the greatest extent possible, ensuring the integrity of basic characteristics such as heart rate and heart rate variability.
[0040] To address high-load piloting conditions (strong motion artifacts), where frequent stick and rudder operations and limb movements generate broadband non-stationary noise, enhanced noise reduction is implemented. Electronic equipment can increase the number of db4 wavelet decomposition layers to 6, simultaneously raising the soft threshold coefficient to enhance the filtering capability of broadband motion artifacts. A QRS complex waveform protection threshold is added to lock the amplitude range of the core ECG waveform and limit the filtering intensity. While filtering out strong artifacts, this prevents the QRS complex from being excessively smoothed, protecting the core ECG characteristics.
[0041] Next, the electronic device can use the signal-to-noise ratio (SNR) as the sole verification metric to calculate the SNR of a single ECG signal segment. If the SNR is greater than or equal to a preset SNR threshold (e.g., 25dB), it is considered a qualified target ECG signal and proceeds to the next stage (R-wave localization and feature extraction); if the SNR is less than the preset SNR threshold (e.g., 25dB), it is considered to have excessive residual noise. The entire noise reduction process is then repeated for the segments exceeding the standard; if the signal still fails to meet the standard after the second processing, the signal segment is directly discarded.
[0042] Finally, the electronic device can obtain a signal that passes the SNR check, which is a low-noise, waveform-complete target ECG signal.
[0043] Step S102: Identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot.
[0044] Specifically, the electronic device can identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot based on the identification results.
[0045] This step will be explained in detail below.
[0046] Step S103: Identify the target electrocardiogram signal and extract the target physiological characteristics corresponding to the target pilot.
[0047] Specifically, electronic devices can identify the target's electrocardiogram signal and extract the target's physiological characteristics corresponding to the pilot based on the identification results.
[0048] This step will be explained in detail below.
[0049] Step S104: Based on the comprehensive facial features and the target's physiological features, determine the emotion probability vector corresponding to the target pilot.
[0050] Specifically, the electronic device can fuse facial features and target physiological features to obtain target fusion features. Then, based on the target fusion features, the corresponding emotion probability vector of the target pilot is determined.
[0051] This step will be explained in detail below.
[0052] Step S105: Based on the emotion probability vector, determine the current flight risk index corresponding to the target pilot.
[0053] Specifically, electronic devices can determine the current emotion type of the target pilot based on an emotion probability vector. Then, based on the current emotion type, they can determine the current flight risk index for the target pilot.
[0054] This step will be explained in detail below.
[0055] The flight risk index assessment method provided in this embodiment acquires multi-frame target facial image data and target electrocardiogram (ECG) signals. It employs simultaneous dual-modal acquisition of facial vision and ECG / physiological data. Facial images capture millisecond-level instantaneous stress expressions, while ECG signals capture latent long-term stress that surface expressions cannot reveal, providing a reliable data source foundation for accurate feature extraction. The method identifies target facial images and extracts comprehensive facial features, fully preserving the details of the pilot's overt emotional dynamics. ECG signals are identified to extract target physiological features, accurately representing autonomic nervous system fluctuations and latent suppressed emotions, compensating for the limitation of facial recognition in detecting internal high pressure. An emotion probability vector is calculated by combining comprehensive facial features and physiological features. This approach takes into account both overt expressions and latent physiological states, significantly improving the accuracy and anti-interference capabilities of emotion recognition compared to single-method identification. The current flight risk index is calculated based on the emotion probability vector. This enables personalized, condition-based, and quantitative risk grading; the output is a standardized risk index of 0-1, which can be directly connected to the airborne early warning system to intuitively quantify the safety impact of emotions on flight operations, supporting crew status monitoring and safety intervention.
[0056] This embodiment provides a method for assessing flight risk index, which can be used in electronic devices, such as control devices in a target aircraft or flight simulator. Figure 2 This is a flowchart of a flight risk index assessment method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Acquire multi-frame target facial image data and target electrocardiogram signal corresponding to the target pilot.
[0057] Please refer to the above description of step S101 for details on this step, which will not be repeated here.
[0058] Step S202: Identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot.
[0059] Specifically, step S202 above may include the following steps: Step S2021: Identify each target facial image data and determine the static facial action unit feature vector corresponding to each target facial image data.
[0060] Specifically, the electronic device can use the FACS standard to identify the nine core facial action units (AUs) of this application: AU1, AU2, AU4, AU5, AU6, AU10, AU20, AU24, and AU26. AU1 represents the Inner Brow Raiser, with the action pattern being: the inner eyebrows of both eyes are lifted upwards, stretching the skin between the eyebrows; the associated emotions are: surprise, slight tension, and high concentration; in flight scenarios, it manifests as: when encountering a sudden bird strike or abnormal posture, the pilot instantly raises the inner eyebrows, which is the initial stress signal; the recognition area is: the inner ROI of the eyebrows and eyes. AU2 represents the Outer Brow Raiser, with the action pattern being: the outer eyebrows of both eyes are raised upwards, with the eyebrow tails lifted; the associated emotions are: strong surprise, astonishment, and unexpectedness; it often appears in conjunction with AU1; in flight scenarios, it manifests as: major emergencies such as abnormal engine noise or runway incursion, with AU1 and AU2 simultaneously highly activated; the recognition area is: the outer region of the eyebrows and eyes. AU4 is characterized by downward-pressed eyebrows (BrowLowerer); the movement pattern is: both eyebrows are squeezed downwards, the brow is furrowed, and vertical lines are formed between the eyebrows; the emotional associations are: anger, irritability, anxiety, and stress; the flight scenario manifestations are: prolonged high-load approach, repeated handling of malfunctions, and continuous activation during irritability due to operational errors; it is a core marker of negative emotions; key combination: elevated AU4 alone = irritability and anxiety; AU4 + AU5 combination = fear and rage. AU5 is characterized by: raised upper eyelid (UpperLidRaiser); the movement pattern is: the upper eyelid is raised significantly, exposing more of the upper sclera and increasing the exposed area of the eyeball; the emotional associations are: fear, panic, high alertness, and fright; the flight scenario manifestations are: pupil dilation and staring during dangerous malfunctions, with AU5 instantly surging; the core pairing is: high activation of AU4 + AU5 together, which is the key feature combination for this system to determine fear and rage. AU6 represents Cheek Raiser; Action: Cheekbone muscles lift, cheeks puff out, and smile lines form around the eyes; Emotional association: Pleasant, relaxed, and easy positive emotions; Flight scenario manifestation: Stably activated during smooth cruise and low-load downwind phases when the pilot is in a relaxed, stable, or pleasant state; The only core AU for positive emotions, a key indicator distinguishing between calm and pleasant states. AU10 represents Upper Lip Raiser; Action: Upper lip lifts upward, nostrils slightly expand; Emotional association: Disgust, aversion, resistance, and mild irritability; Flight scenario manifestation: Occurs when there is resistance and irritability to tedious operations or equipment malfunctions; Often accompanied by AU4. AU20 represents Lip Stretcher; Action: Lips are pulled horizontally to the left and right, corners of the mouth extend horizontally, and lips become thinner; Emotional association: Fear, tension, anxiety, and suppressed unease; Flight scenario manifestation: Suppressed tension during major emergencies, lips are tightly stretched horizontally, often occurring simultaneously with AU4 and AU5.AU24 represents lip pressing; action: upper and lower lips are pressed inward and pursed tightly, without opening or closing; emotional association: suppressed anger, high tension, forced calmness, high-pressure restraint; flight scenario representation: pilots deliberately suppress negative emotions, appearing calm on the surface but under immense internal pressure, a hallmark of implicit high-pressure emotions. AU26 represents lip opening (Jaw Drop); action: jaw drops, mouth opens naturally, lips are relaxed and separated; emotional association: extreme surprise, sudden shock, astonishment; flight scenario representation: a sudden, extreme malfunction causing the mouth to open in astonishment, with extremely strong instantaneous pulse characteristics, requiring a high frame rate of 480Hz to capture completely.
[0061] The electronic device can identify the facial image data of each target based on the coordinates of key points such as eyebrows, eyes, corners of lips, and nostrils obtained from previous tracking, and divide the local pixel region of the face corresponding to each AU. The electronic device can quantify the muscle movement amplitude of each type of AU and output an activation intensity value in the range of 0 to 1: a value of 0 represents no muscle movement, and a value of 1 represents that the AU has reached the maximum range of movement. Among them, AU6 high activation corresponds to pleasant emotions; AU4+AU5 simultaneous high activation corresponds to negative emotions such as anger and fear.
[0062] The electronic device can arrange the activation intensity values of 9 AUs in a fixed order to form a one-dimensional array, which is the feature vector of a single-frame static facial motion unit. For example, the vector format is: F static =[AU1,AU2,AU4,AU5,AU6,AU10,AU20,AU24,AU26], with a fixed dimension of 9 dimensions.
[0063] Step S2022: Calculate the activation intensity difference of the same facial action unit between the feature vectors of static facial action units in two adjacent frames to obtain the inter-frame temporal difference feature.
[0064] Specifically, for each pair of adjacent frames, the electronic device can perform a subtraction operation on the activation intensity of the same facial motion unit, using the following formula: .in, This is the activation value of the k-th action unit in frame t; This represents the difference in changes between two frames for the AU. The positive or negative value indicates the trend of muscle movement: a positive value indicates that the AU's muscle movement becomes larger and the facial expression amplitude increases; a negative value indicates that the AU's muscle movement contracts and the facial expression amplitude decreases; and a value close to 0 indicates that the facial state is stable with no significant change.
[0065] The electronic device can sequentially arrange the differences corresponding to the 9 AUs to generate a single set of 9-dimensional inter-frame temporal difference features: ΔF(t)=[ΔAU1,ΔAU2,ΔAU4,ΔAU5,ΔAU6,ΔAU... 10 ,ΔAU 20 ,ΔAU24 ,ΔAU 26 This difference feature can amplify the magnitude of sudden changes in instantaneous stress emotions such as surprise and fear, making up for the shortcoming that a single static AU cannot characterize dynamic changes in emotions.
[0066] Step S2023: The static facial motion unit feature vector is fused with the inter-frame temporal difference feature to obtain the comprehensive facial features.
[0067] Specifically, electronic devices can directly concatenate and fuse static facial motion unit feature vectors with inter-frame temporal difference features without dimensionality compression loss to obtain comprehensive facial features, as shown in the formula: .in, This is a 9-dimensional static facial motion unit feature vector. It is a 9-dimensional inter-frame temporal difference feature.
[0068] Step S203: Identify the target's electrocardiogram signal and extract the target's physiological characteristics corresponding to the target pilot.
[0069] Specifically, step S203 above may include the following steps: Step S2031: Identify the target ECG signal and locate the R wave from the target ECG signal.
[0070] Specifically, the electronic device can load the classic Pan-Tompkins detection algorithm to sequentially perform bandpass filtering preprocessing on the target ECG signal, signal differentiation to amplify the waveform slope, squaring to amplify the QRS peak difference, and sliding window integration to integrate the waveform energy, thereby initially screening all suspected peak points.
[0071] Then, since cockpit limb movements easily generate a large number of false peaks, the electronic equipment can cache the time interval between three consecutive adjacent peaks to temporarily generate a preliminary RR interval. The electronic equipment can use the range of an adult's physiological heart rate during flight as a coarse filter to directly remove false points with extremely abnormal intervals, retaining peaks whose shape and intervals are relatively consistent with the heartbeat rhythm, and obtaining a raw R-wave point sequence with timestamps, which serves as the basic data source for calculating the RR interval.
[0072] Step S2032: Obtain the reasonable range of the preset RR interval.
[0073] Specifically, a healthy adult's resting heart rate is 60-100 beats per minute. The interval can be calculated as: single heartbeat interval = 60 / heart rate, resulting in a reasonable baseline range of 0.6s to 1.0s.
[0074] Electronic devices can switch the lower limit of the operating range according to the operating conditions indicated by the synchronization marker. For stable, low-load operating conditions: the preset reasonable range for the RR interval is 0.6s to 1.0s; for high-stress, high-intensity operating conditions: when a person is tense and their heart rate increases, the lower limit is lowered to 0.4s, while the upper limit remains at 1.0s, and the preset reasonable range for the RR interval becomes 0.4s to 1.0s; under extreme panic, the heart rate will not exceed 150 beats per minute, and 0.4s is the safe minimum threshold.
[0075] In step S2033, if the detected peak exceeds the reasonable range of the preset RR interval, it is determined to be a noise spurious peak and removed to obtain the ECG signal to be extracted.
[0076] Specifically, the electronic device can traverse all the R-wave time coordinates output by S2031, and subtract the previous peak time from the next peak time in turn to generate a complete array of original RR interval values.
[0077] The electronic device matches the interval values one by one with the reasonable range of the preset RR interval of S2032. If the original RR interval value array is within the reasonable range of the preset RR interval, it is determined to be the true heartbeat interval, and all the preceding and following R waves are retained. If the original RR interval value array is higher than the upper limit of the preset reasonable range for RR intervals or lower than the lower limit, it is determined that there are noise spurious peaks within that interval, and the abnormal peak points are directly deleted. After the electronic device deletes the abnormal peaks, the combination relationship of adjacent R waves changes, the updated RR interval array is recalculated, and the interval threshold is compared with all intervals again; this process is repeated until all RR intervals in the sequence fall within the reasonable physiological range, thus obtaining the ECG signal to be extracted.
[0078] Step S2034: Perform feature extraction on the ECG signal to be extracted to obtain the time-domain features, frequency-domain features, basic nonlinear features, and preset nonlinear features corresponding to the ECG signal to be extracted.
[0079] The time-domain features include the average of all RR intervals and the root mean square of the difference between adjacent RR intervals. The frequency-domain features include the power at the first frequency, the power at the second frequency, and the ratio between the first and second frequency powers. The first frequency is less than the second frequency. The basic nonlinear features include the standard deviation of the minor axis and the standard deviation of the major axis of the Poincaré plot. The preset nonlinear features include the eccentricity of the Poincaré plot ellipse and the sample entropy of the RR intervals.
[0080] Specifically, electronic devices can extract the mean RR interval (MeanNN) and the root mean square difference (RMSSD) of adjacent RR intervals. MeanNN aims to define the pilot's baseline heart rate reference; while RMSSD focuses on quantifying instantaneous, beat-by-beat heart rate fluctuations driven by the parasympathetic nervous system, and is highly sensitive to capturing emotional changes caused by stress states such as surprise. The quantification process of these two indicators is shown in formula (1).
[0081] ;in, Representing the The duration of each RR interval, and It is the total number of RR intervals in the sequence.
[0082] Specifically, electronic devices can estimate the power spectral density (PSD) of a sequence using Welch's method, with the specific spectral distribution as follows: Figure 3 As shown.
[0083] Electronic devices can quantify key frequency domain characteristics by integrating energy within a specific frequency range based on the obtained power spectrum. The first frequency power, the low-frequency band (LF, ranging from 0.04Hz to 0.15Hz), is driven by both the sympathetic and parasympathetic nervous systems and is considered a representative signal of sympathetic nerve excitability; its fluctuations are highly correlated with the pilot's emotional state, such as anxiety and tension, during a mission. Meanwhile, the second frequency power, the high-frequency band (HF, ranging from 0.15Hz to 0.4Hz), primarily originates from respiratory regulation and is a reliable carrier for assessing parasympathetic (vagus) nerve activity. Typically, an increase in HF power reflects a pilot's relaxed state, while its decrease during fear or anxiety reveals a decline in emotional regulation ability, which is of significant value in identifying whether a pilot is in a negative emotional state. Electronic devices can also use the ratio of the low-frequency power to the high-frequency power to obtain the ratio between the first and second frequency powers. A larger ratio indicates higher sympathetic nerve excitability and stronger negative emotions such as fear, anger, and anxiety.
[0084] In addition, electronic devices employ, for example Figure 4 The Poincaré plot shown is used to analyze the nonlinear characteristics of the ECG. The horizontal axis X = RR. n The ordinate Y = RR n+1 Each group of adjacent intervals generates a scatter point, forming an overall scatter point cluster, and a Poincaré diagram is generated based on the overall scatter point cluster.
[0085] Specifically, the electronic device can quantify the geometric features of the elliptical scatter set in the Poincaré plot, and calculate two key non-linear indicators: the standard deviation of the short axis of the Poincaré plot (SD1) and the standard deviation of the long axis of the Poincaré plot (SD2). Wherein, the standard deviation of the short axis of the Poincaré plot SD1 represents the standard deviation of the scatter points in the direction perpendicular to the diagonal line (y=x), which is used to quantify short-term, beat-to-beat heart rate variability, and is highly correlated with high-frequency (HF) components and RMSSD, mainly reflecting the regulation ability of the parasympathetic nerve. A decrease in its value is usually associated with increased mental stress of pilots or the occurrence of negative emotions such as anxiety and fear; SD2 represents the standard deviation of scatter points along the diagonal direction, which is used to quantify long-term heart rate variability and is correlated with low-frequency (LF) components. Since SD2 can indirectly reflect the activation level of the sympathetic nerve, it is often used as an effective reference for evaluating the emotional states such as anxiety and sustained stress of pilots during task execution. Illustratively, the calculation formulas of SD1 and SD2 are shown in formula (2).
[0086] ; wherein, represents the duration of the -th R-R interval.
[0087] The electronic device can perform minimum enclosing ellipse fitting on the Poincaré scatter cluster, and substitute the major and minor axes of the ellipse to calculate the elliptic eccentricity of the Poincaré plot. The formula is: , where a is the length of the semi-major axis and b is the length of the semi-minor axis; emotions such as fear and rage will pull the distribution shape of the scatter points, causing the ellipse to stretch and deform, and the eccentricity to deviate significantly from the normal value.
[0088] Finally, the electronic device can perform sliding interception on the electrocardiographic signal to be extracted with a preset window length to obtain a plurality of sub-vectors.
[0089] Illustratively, the electronic device can slide to intercept sub-vectors with a window length m=2: , and the total number of sub-vectors is: M=N-m.
[0090] The electronic device can calculate the maximum difference (Chebyshev distance) between any two sub-vectors X i and X j : , if d[X i ,X j <r (similarity tolerance threshold), the two sub-vector patterns are determined to be similar.
[0091] For each X i , count the number of j that satisfy , which is recorded as B i ; calculate the average matching proportion: : Let be the average similarity matching probability of all subvectors in m dimensions.
[0092] The electronic device can reconstruct a subvector of length 3 by setting m′=m+1=3: Similarly, the electronic device calculates all matches to obtain the average match probability. .
[0093] The formula for calculating the sample entropy during the RR interval in electronic devices is: .
[0094] Step S2035: The time domain features, frequency domain features, basic nonlinear features and preset nonlinear features are fused to obtain the target physiological features corresponding to the target pilot.
[0095] Specifically, the electronic device can normalize the time-domain features, frequency-domain features, fundamental nonlinear features, and preset nonlinear features, and then strictly concatenate them horizontally in a preset fixed order to form a one-dimensional long vector. The concatenation order is as follows: The target physiological features with a dimension of 9 are obtained by splicing them together.
[0096] Step S204: Based on the comprehensive facial features and the target's physiological features, determine the emotion probability vector corresponding to the target pilot.
[0097] Specifically, step S204 above may include the following steps: Step S2041: Input the comprehensive facial features into the facial feature recognition branch and output the target visual feature vector.
[0098] Specifically, the static processing branch in the facial feature recognition branch acquires the occlusion status of the local pixel region of the face corresponding to each AU. Based on the occlusion status of each local pixel region of the face, the weight of each AU in the static facial action unit feature vector in the facial comprehensive feature is determined. Based on the weights corresponding to each AU, weighted processing is performed on each AU in the static facial action unit feature vector.
[0099] Then, the electronic device can apply feature association constraints to three sets of pilot core emotion AU combinations, with independent coupling enhancement in three groups. For the fear combination (AU4+AU5), element-wise multiplication coupling is performed on the AU4 and AU5 channels within the feature tensor, amplifying the synergistic features of their synchronous increase; minor fluctuations in individual AUs and weak perturbations that are not activated simultaneously are naturally suppressed. For the momentary surprise combination (AU1+AU2+AU26), joint feature superposition enhancement is performed on the three channels, and the combined feature amplitude is significantly improved only when all three reach a synchronous peak; meaningless small movements such as raising eyebrows or opening the mouth alone are filtered out. For the pleasure combination (AU6), an independent gain baseline is set for the AU6 channel, and a small increase in AU6 under stable flight can amplify the pleasure representation, avoiding the positive emotion being masked by the weak fluctuations of other AUs. For isolated small AU tremors (unconscious facial movements) that do not belong to the three emotion combinations, no coupling enhancement is applied, and the feature amplitude remains at the original low level, without interfering with the emotion encoding.
[0100] The dynamic processing branch in the facial feature recognition branch calculates the global mean (mean) of the nine AU differences within the inter-frame temporal difference feature of the facial comprehensive features. If the mean value is higher than the preset mutation threshold, the current frame is determined to be a high-stress sudden condition such as bird strike or engine failure; if the mean value is lower than the preset mutation threshold, the flight is determined to be stable and the dynamic characteristics remain unchanged.
[0101] After triggering the mutation determination, the electronic device can multiply all feature tensors of the entire dynamic difference branch by the dynamic amplification factor k. dyn (e.g., 1.2-1.5); short-term stress facial expressions such as raising eyebrows, widening eyes, and stretching lips captured at a high frame rate of 480Hz are significantly amplified.
[0102] Then, the electronic device can concatenate the outputs of the static processing branch and the dynamic processing branch to obtain the initial visual features. A 1×1 convolution layer is then applied for channel equalization to balance the numerical scale of the static baseline amplitude and the dynamic abrupt change amplitude. ReLU activation and a small dropout (0.1) are interspersed to suppress overfitting and retain all the effective feature information after enhancement.
[0103] Global average pooling is performed on the fused initial visual features to compress redundant spatial dimensions, preserve global emotional semantics, and eliminate fluctuations caused by minor local jitter in single frames. The pooled features are then fed into the final fully connected layer and uniformly mapped to a fixed 512-dimensional length. The vector dimension remains constant throughout the process to ensure strict alignment with the output dimension of the ECG branch, ultimately yielding the target visual feature vector.
[0104] Step S2042: Input the target physiological features into the ECG feature recognition branch and output the ECG feature vector.
[0105] Specifically, the ECG feature recognition branch can decompose the target physiological features into three independent feature sub-vectors according to the physiological regulation mechanism. After grouping, the sub-vectors are input into the first layer of the 1D-CNN through different channels to match the hierarchical pattern of neural regulation. Among them, the steady-state heart rate group (2-dimensional): [MeanNN, RMSSD], reflects the basic rhythm of the heartbeat and short-term small fluctuations; the frequency domain neural regulation group (3-dimensional): [LF, HF, LF / HF], represents the long-term balance level of the sympathetic and parasympathetic systems; and the higher-order nonlinear steady-state group (4-dimensional): elliptic eccentricity and sample entropy, representing the degree of long-term autonomic nervous system disorder and rhythm complexity.
[0106] Then, the ECG feature recognition branch adopts a three-layer one-dimensional convolution stack with low-intensity smooth convolution kernels throughout, without amplifying abrupt feature changes, which fits the inherent properties of ECG emotions being delayed and slowly changing.
[0107] The convolutional layer consists of three layers: The first layer has a kernel size of k=3 and a stride of 1. It performs basic correlation extraction on the three-group channel features, capturing the linkage of indicators within each group (e.g., LF increases synchronously with SD2 increase, and anxiety is simultaneously manifested). The second layer has a kernel size of k=5, expanding the receptive field and mining deep coupling relationships across groups (a consistently high LF / HF and sample entropy can serve as auxiliary representations of persistent stress). The third layer uses small-scale compression convolution to converge global steady-state semantics. Each layer is equipped with ReLU activation and minimal mean pooling: mean pooling replaces max pooling, weakening single-point impulse noise and strengthening the overall average steady-state trend, completely different from the logic of facial branch amplification of peak abrupt changes.
[0108] Next, the electronic device retrieves the SNR (signal-to-noise ratio) value stored in the preprocessed ECG segment and performs global adaptive scaling on the intermediate feature tensor of the convolution to resolve accuracy fluctuations caused by motion artifacts in ECG acquisition. For high-quality signals (SNR... 25dB): Scaling factor w ecg =1.0, steady-state characteristics are fully preserved, original feature weights are maintained, and implicit long-term pressure is reliably supplemented; for the critical qualified signal (SNR=25dB): scaling factor w ecg =0.85, slightly reducing the overall characteristic amplitude and weakening the steady-state deviation caused by trace residual noise.
[0109] Finally, for restrained, latent negative emotions that cannot be recognized by facial expressions (pilots deliberately suppress their expressions during high-pressure approaches, with no obvious fluctuations in facial AU, but persistent tension in the autonomic nervous system), a steady-state feature anchoring gain is set. If the three conditions of high LF / HF, high sample entropy, and high SD2 are simultaneously detected (long-term sympathetic excitation), a small steady-state gain k is applied to the convolutional feature tensor. stable=1.15. This amplifies implicit emotional features such as high stress and anxiety, which are difficult to express on the face, filling a significant blind spot in single visual recognition. Finally, the entire convolutional temporal features are aggregated to output a one-dimensional hidden layer vector, maintaining a steady-state average trend throughout. The LayerNorm layer unifies the feature distribution scale, ensuring matching with the 512-dimensional visual feature distribution range. The output dimension is fixed at 512 dimensions, strictly consistent with the output dimension of the facial branch, satisfying the requirements of subsequent weighted fusion operations to obtain the ECG feature vector.
[0110] Step S2043: The target visual feature vector and electrocardiogram feature vector are fused to obtain the initial fused feature vector.
[0111] Specifically, the electronic device can obtain the global mean (mean) of the nine AU differences within the inter-frame temporal difference feature. If the mean is higher than the preset mutation threshold, the weight corresponding to the target visual feature vector is increased. If the mean is lower than the preset mutation threshold, the weight corresponding to the ECG feature vector is increased.
[0112] Then, based on the weights corresponding to the target visual feature vector and the ECG feature vector, the target visual feature vector and the ECG feature vector are fused to obtain the initial fused feature vector.
[0113] Step S2044: Obtain the current flight phase.
[0114] Specifically, the electronic equipment can read flight bus data in real time, classify the flight into three phases, and obtain phase labels. Phase 1: Upwind / crosswind climb (medium load); Phase 2: Downwind cruise (low load, high relaxation); Phase 3: Approach and landing (high load, high precision control).
[0115] Electronic devices can convert discrete phase labels into 3D one-hot encoded vectors as prior information about the scene. The freedom and explicitness of pilots' facial emotional expressions vary greatly in different flight phases; their expressions are more restrained during approach and more relaxed during cruise, providing scene constraints for subsequent feature correction.
[0116] Step S2045: Concatenate the current flight phase with the initial fused feature vector to obtain the fused phase feature vector.
[0117] Specifically, the electronic device can use a lossless dimensional concatenation method to directly concatenate the 3D flight phase vector with the 512-dimensional initial fusion feature vector to obtain a 515-dimensional fusion phase feature vector, as shown in the formula: .
[0118] Step S2046: Based on the feature vector from the fusion stage, determine the emotion probability vector corresponding to the target pilot.
[0119] Specifically, step S2046 above may include the following steps: Step a1: Use the target visual feature vector in the feature vector of the fusion stage as the query vector, and the electrocardiogram feature vector in the feature vector of the fusion stage as the key vector and value vector.
[0120] Specifically, electronic devices can extract the operating condition code from the 515-dimensional fusion stage feature vector and retrieve the preceding 512-dimensional multimodal basic features; these features are then directionally allocated according to cross-modal matching rules. The query vector Q = target visual feature vector (512-dimensional), using facial visual emotion as the query subject; the key vector K = ECG feature vector (512-dimensional); and the value vector V = ECG feature vector (512-dimensional). Visual explicit emotion is used as the query anchor point to match the ECG implicit physiological emotion correlation; this aligns with the emotional expression patterns of pilots: the face is the intuitive externalization of emotion, and the ECG serves as an endogenous physiological reference, achieving cross-modal alignment of "external expression matching internal neural state."
[0121] Step a2: Based on the query vector, key vector, and value vector, calculate the cross-modal attention weights to obtain the target cross-modal attention weight matrix.
[0122] Specifically, step a2 above may include the following steps:
[0123] Step a21: Calculate the original attention score matrix based on the query vector and key vector.
[0124] Specifically, the electronic device can calculate the original attention score matrix based on the query vector and the key vector using the following formula: Among them, d k =512 represents the unified feature dimension for Q and K. To scale and suppress gradient saturation, the original attention score matrix is: Matrix element s ij This represents the original matching similarity between the i-th visual feature and the j-th electrocardiogram feature; a higher value indicates a stronger emotional correlation between the two sets of features.
[0125] Step a22: Obtain the current flight parameters.
[0126] Specifically, electronic devices can synchronously read multiple types of quantified current flight parameters from the airborne bus as a quantitative basis for scenario judgment. The current flight parameter acquisition indicators include: attitude parameters: yaw angle, altitude deviation, pitch error; power parameters: engine speed, thrust status; control load parameters: stick and rudder offset amplitude, correction frequency.
[0127] Step a23: Determine the current scenario type based on the current flight parameters.
[0128] Specifically, the electronic device can compare the current flight parameters with the baseline flight parameters corresponding to each scenario type. Then, based on the comparison results, the current scenario type is determined.
[0129] Among them, electronic devices can be pre-set with three standardized aviation scenario types, which are automatically identified and classified based on parameter thresholds. Scenario 1: Smooth low-load cruise, with minimal altitude and attitude deviations, no fault alarms, and low frequency of stick and rudder control corrections; the pilot's emotions are minimal, and facial and ECG are generally stable; Scenario 2: Routine climb / approach with moderate load, involving minor attitude corrections, no major faults, and stable workload for controls; mild stress and slight anxiety are present; Scenario 3: High-stress sudden fault condition, with fault markers, large attitude deviations, and emergency control actions; high incidence of intense negative emotions such as instantaneous fright, fear, and anger.
[0130] Step a24: Based on the current scene type, perform a masking operation on the original attention score matrix to obtain the masked attention score matrix.
[0131] Specifically, electronic devices can generate a mask matrix based on the current scene type.
[0132] For example, if the current scene type is smooth low-load cruise, the electronic device traverses all visual rows i∈[1,512] in the mask matrix: M[i,j]=-10 9 ,∀j∈E stress ECG baseline steady-state channel E base All columns remain at 0. If the current scenario type is a normal climb / approach to moderate load, the electronic device only marks a small number of isolated ECG columns that are highly susceptible to slight motion artifacts and were calibrated offline. noise Electronic devices perform local single-point assignment: The majority of elements in the matrix remain 0, and no large-scale channel blocking is implemented. If the current scenario is a high-stress, sudden failure condition, the electronic device acquires all elements belonging to the visually pleasing partition V. joy The row index i; the electronic device traverses all ECG columns j∈[1,512]: The reserved area remains at 0.
[0133] Then, the electronic device can superimpose the mask matrix with the original attention score matrix to obtain the masked attention score matrix.
[0134] Step a25: Normalize the mask attention score matrix to obtain the basic cross-modal attention weight matrix.
[0135] Specifically, the electronic device can perform Softmax normalization on the mask attention score matrix: Each row sums all weights to 1, with a value range of [0,1]; outputs the basic cross-modal attention weight matrix. At this point, the basic cross-modal attention weight matrix has been incorporated into the prior constraints of the flight scenario, but the differences in the acquisition quality of images and electrocardiograms have not been considered.
[0136] Step a26: Based on the modal confidence scores corresponding to the target visual feature vector and ECG feature vector, the basic cross-modal attention weight matrix is corrected to obtain the target cross-modal attention weight matrix.
[0137] Specifically, the electronic device can calculate the visual confidence score corresponding to the target visual feature vector based on the face occlusion area and image sharpness in the target facial image data, within the range (0,1); the more severe the occlusion, the higher the confidence score. vis The lower the signal-to-noise ratio, the lower the confidence score. Electronic devices can calculate the ECG confidence score corresponding to the ECG feature vector based on the SNR of the target ECG signal, in the range (0,1]. The lower the signal-to-noise ratio, the lower the confidence score.
[0138] Then, the electronic device can calculate the visual confidence balance coefficient and the electrocardiogram confidence balance coefficient corresponding to the target visual feature vector and electrocardiogram feature vector, respectively, based on the visual confidence score and the electrocardiogram confidence score.
[0139] The electronic device multiplies the visual query side (matrix row direction) in the basic cross-modal attention weight matrix by the visual confidence balance coefficient, and multiplies the ECG key value side (matrix column direction) in the basic cross-modal attention weight matrix by the ECG confidence balance coefficient. After scaling, the row sum is calibrated again with Softmax row-by-row to 1, and finally the target cross-modal attention weight matrix is obtained.
[0140] Step a3: Based on the current flight stage in the feature vector of the fusion stage, the target cross-modal attention weight matrix is corrected to obtain the initial attention weight heatmap.
[0141] In the attention weight heatmap, the horizontal axis corresponds to the visual feature area, the vertical axis corresponds to the electrocardiogram feature area, and the color intensity represents the magnitude of the attention weight.
[0142] Specifically, the electronic device can determine the phase-specific global correction scaling factor λ corresponding to the current flight phase based on the correspondence between flight phases and global correction scaling factors. stage Then, the electronic device can adjust the scaling factor λ according to the stage-specific global correction factor. stage The target cross-modal attention weight matrix is modified to obtain the modified target cross-modal attention weight matrix.
[0143] Then, the electronic device establishes a horizontal axis (X) with 512 columns, corresponding one-to-one with ECG feature partitions (baseline heart rate region, LF / HF frequency domain region, Poincaré / sample entropy nonlinear region); and a vertical axis (Y) with 512 rows, corresponding one-to-one with visual feature partitions (AU6 pleasure region, AU4+AU5 fear region, AU1+AU2+AU26 surprise region, and other ordinary AU regions). The color mapping rule uses a blue→yellow→red gradient: blue = weight approaching 0 (no matching attention), red = weight peak (highly correlated matching). The electronic device then maps each value in the corrected target cross-modal attention weight matrix to the corresponding pixel color brightness, generating a two-dimensional color pixel map, which is the initial attention weight heatmap; the map visually shows which set of visual AUs and ECG physiological indicators the model establishes a strong correlation match with.
[0144] Step a4: Perform temporal weight smoothing on the initial attention weight heatmap to obtain the target attention weight heatmap.
[0145] Specifically, the electronic device can store the corrected target cross-modal attention weight matrix for three consecutive frames: the current frame, the previous frame, and the two frames before that, denoted as W. stage (t),W stage (t-1),W stage (t-2). Pilot emotions do not change between frames, and single-frame heatmaps are susceptible to interference from instantaneous acquisition noise, so time-series smoothing is necessary to suppress jitter.
[0146] Then, the electronic device can use a time-weighted smoothing formula, with closer frames having higher weights, to obtain a smoothed attention weight matrix. The smoothing formula is as follows: .
[0147] Finally, the electronic device can perform Softmax normalization row by row on the smoothed attention weight matrix to restore the sum of weights in each row to 1; and use the smoothed attention weight matrix W smooth (t) Remap the color pixels, refresh the spectrum, and obtain a stable, non-jumping target attention weight heatmap. The heatmap is simultaneously labeled with: current flight stage, modality confidence score, and peak weight coordinates (strongest visual-ECG matching position), supporting interpretable retrospective analysis of the model.
[0148] Step a5: Based on the target attention weight heatmap, the target visual feature vector and ECG feature vector are fused to obtain a global fused feature vector.
[0149] Specifically, the electronic device can use a smooth weight matrix W smooth As a fusion coefficient, the ECG V is weighted and aggregated, and then combined with the query visual Q.
[0150] The weighted ECG aggregation vector is calculated using the following formula: V att These are ECG features filtered by visual attention, retaining only physiological information that highly matches facial emotions. Then, the two channels are concatenated and fused to generate a global fused feature vector, using the following formula: Where Q is a 512-dimensional target visual feature vector; Vatt is a 512-dimensional attention-filtered ECG feature vector; and the concatenated global fusion feature vector F is... global Dimensions = 1024.
[0151] Step a6: Regularize the global fusion feature vector to obtain a unified fusion feature vector.
[0152] Specifically, electronic devices can process a 1024-dimensional globally fused feature vector F global Calculate the mean μ and standard deviation σ of the vector dimension by dimension; normalization transformation: ε is set to a minimum value of 10⁻⁶ to prevent division by zero; after normalization, the mean of the integer vector is 0 and the variance is 1, eliminating the original scale differences between the visual and electrocardiogram branches.
[0153] Then, electronic devices can be configured with a Dropout layer with a dropout rate of 0.1, randomly invalidating 10% of the feature neurons in each dimension, weakening redundant dependencies between features, and improving the model's generalization ability to occlusion and ECG noise; in flight scenarios, equipment occlusion and motion artifacts occur frequently, and regularization can significantly improve the stability of online inference.
[0154] Finally, the 1024-dimensional vector after LayerNorm+Dropout regularization is defined as the unified fused feature vector.
[0155] Step a7: Perform temporal feature processing on the unified fusion feature vector to obtain the temporal feature vector.
[0156] Specifically, step a7 above may include the following steps: Step a71: Divide the unified fused feature vector into multiple subsequences.
[0157] Specifically, electronic devices can unify and fuse a 1024-dimensional feature vector, dividing it according to a fixed, equal dimension to achieve segmented, fine-grained verification and compensation. For example, an electronic device can set the dimension of a single sub-sequence to a fixed 128 dimensions: Sub-sequence count: 1024 ÷ 128 = 8 sub-sequences, denoted as S1, S2, S3, S4, S5, S6, S7, S8. Among these, S1~S4 carry the visually dominant fusion dimension (AU static, inter-frame dynamic mutation features); S5~S8 carry the electrophysiologically dominant fusion dimension (time domain, frequency domain, nonlinear heart rate variability features).
[0158] Step a72: For each subsequence, calculate the distribution difference of all query vectors and key vectors within the subsequence, and filter out the core vectors whose contribution is greater than the preset contribution threshold.
[0159] Specifically, for any subsequence S p The electronic device obtains the local query vector block Q corresponding to this dimension interval. p Local bond vector block K p Electronic devices can use the KL divergence to measure Q. p With K p Differences in probability distribution: The larger the KL divergence value, the more fragmented the distribution of visual and electrocardiogram features in the group, the greater the matching differences, and the higher the discrimination and contribution to emotion recognition; the smaller the value, the more similar the features, the more redundant and indistinguishable.
[0160] Electronic devices can normalize the KL divergence to obtain a contribution score in the 0-1 interval. p Then, the electronic device retrieves the preset contribution threshold Th stored in the electronic device. score (Engineering calibration fixed value). Then, the electronic device can compare the contribution score with the preset contribution threshold. If Scorep > Thscore, then this group is marked as a core vector and retained for encoding; if Scorep ≤ Thscore: it is determined to be a low contribution redundant vector, directly discarded, and not included in subsequent weighted calculations.
[0161] Step a73: Perform weighted calculations on each core vector to obtain the preliminary coding features corresponding to the subsequences.
[0162] Specifically, the electronic device can be expressed as the normalized KL divergence fraction S. corep As a weighting coefficient, the higher the score, the greater the weight. Then, the weighted calculation is performed on each core vector to obtain the preliminary encoding features corresponding to the subsequence, as shown in the formula: Each subsequence generates a unique initial encoded feature. p,init .
[0163] Step a74: For each preliminary coding feature, obtain the baseline flight parameters corresponding to the target pilot.
[0164] Among them, the baseline flight parameters are flight parameters collected when the target pilot is calm and the target aircraft is flying steadily.
[0165] Specifically, the electronic equipment can retrieve from the database a pre-stored steady-state benchmark dataset Para, collected when the target pilot is completely calm, the aircraft is in stable flight with low operational load, and under fault-free conditions. baseThe baseline parameters include: altitude baseline, yaw angle baseline, pitch deviation baseline, stick / rudder correction baseline, and engine stability thrust baseline. The electronic equipment can map these baseline flight parameters one-to-one with the real-time acquired flight parameters Para of the current frame. curr The metrics and dimensions are perfectly matched.
[0166] Step a75: Compare the current flight parameters corresponding to each preliminary coding feature with the baseline flight parameters.
[0167] Specifically, the electronic device can compare the current flight parameters corresponding to each preliminary encoded feature with the baseline flight parameters. Based on the current flight parameters and the baseline flight parameters, it calculates the relative offset ratio for each flight indicator, using the following formula: Then, the electronic equipment can perform a weighted summation of four types of deviations: altitude, attitude, control, and power, to obtain a global comprehensive deviation value, Dev. total Among them, the deviation of the control action has the highest weight, followed by the attitude, and the power has the lowest weight.
[0168] Step a76: Based on the comparison results, determine the anomaly level corresponding to the current flight parameters.
[0169] Specifically, electronic equipment can determine the anomaly level corresponding to the current flight parameters based on the calculated global comprehensive deviation value.
[0170] For example, Level 1 (no anomalies): 0 ≤ Dev total <10%, flight status is almost identical to the baseline stable state; Level 2 (minor anomaly): 10% ≤ Dev total <25%, minor adjustment, light load; Level 3 (moderate anomaly): 25% ≤ Dev total <50%, continuous high-load operation, significant pressure; Level 4 (Severe Anomaly): Dev total ≥50%, malfunction, emergency response, severe attitude deviation, high-stress conditions. Electronic equipment can bind a global anomaly level Lp∈{1,2,3,4} to each subsequence.
[0171] Step a77: Bind the anomaly level to the preliminary coding features to obtain the first backup coding features corresponding to each subsequence.
[0172] Specifically, electronic devices can bind anomaly levels with preliminary coding features to obtain the first backup coding features corresponding to each subsequence.
[0173] Step a78: Determine the compensation interval corresponding to each first backup coding feature based on the anomaly level corresponding to each first backup coding feature, and obtain the second backup coding feature corresponding to the compensation interval.
[0174] The compensation interval represents the number of preset frames before and after each subsequence.
[0175] Specifically, the compensation interval represents the number of consecutive frames that need to participate in the compensation weighting. The higher the anomaly level, the larger the range of compensation frames, relying on the timing information of multiple frames to offset the noise jitter of a single frame. For example, the mapping rules between anomaly level and compensation interval are as follows: Level 1 L=1: compensation interval ±1 frame (1 previous frame, 1 next frame); Level 2 L=2: compensation interval ±2 frames; Level 3 L=3: compensation interval ±3 frames; Level 4 L=4: compensation interval ±4 frames.
[0176] The electronic device can retrieve the second backup coding features of all subsequences at the same position within the compensation interval before and after the current subsequence from the buffered time sequence queue, and uniformly collect them into interval feature groups. All vectors in this group are collectively referred to as the second backup coding feature set.
[0177] Step a79: Weighted fusion of each second backup coding feature is performed to obtain the weighted compensation feature corresponding to the subsequence.
[0178] Specifically, the electronic device can determine the weights corresponding to each second backup coding feature based on temporal distance. Specifically, the closer to the current frame, the higher the weight; the farther away, the lower the weight. The electronic device can use a linear decreasing method to determine the weight information corresponding to each second backup coding feature.
[0179] Then, the electronic device can perform weighted fusion of each second backup coding feature based on the weight information corresponding to each second backup coding feature to obtain the weighted compensation feature corresponding to the subsequence.
[0180] For example, the formula is: , This refers to the weight information corresponding to each second backup coding feature.
[0181] Step a710: The first backup coding feature and the weighted compensation feature are weighted and fused to obtain the sub-temporal feature vector corresponding to the sub-sequence.
[0182] Specifically, the electronic device can determine the fusion balance coefficient β based on the anomaly level corresponding to the first backup coding feature. Lp For example, level one has no anomalies: β Lp =0.1, with the current first backup code as the primary code; Secondary minor: β Lp =0.2; Grade III moderate β Lp =0.35; Grade IV severe β Lp =0.5; Fusion formula: Output SubSeq p This is the stable sub-temporal feature vector corresponding to the p-th subsequence.
[0183] Step a711: Fuse the sub-time series feature vectors to obtain the time series feature vector.
[0184] Specifically, electronic devices can concatenate the sub-time series feature vectors to obtain a time series feature vector.
[0185] Step a8: Determine the emotion probability vector corresponding to the target pilot based on the temporal feature vector.
[0186] Specifically, step a8 above may include the following steps: Step a81 involves performing dimensionality reduction, nonlinear transformation, and regularization on the temporal feature vector to obtain the fully connected output features.
[0187] Specifically, the electronic device can input a 1024-dimensional temporal feature vector, set the first FC layer mapping dimension to 256 dimensions, and complete the compression of high-dimensional redundant information, as shown in the formula: ,in, Let W1 be a time-series feature vector, ∈ R. 256×1024 Let b1 ∈ R be the weight matrix. 256 As a bias term, weakly contributing dimensions are eliminated through linear projection, while core emotional representation dimensions are retained.
[0188] Then, electronic devices can employ the ReLU nonlinear activation function to introduce nonlinear fitting capabilities, overcoming the limitations of linear mapping representation and learning the complex coupling relationships between AU combinations, ECG indicators, and flight conditions. For example, the formula is: Negative features are set to zero, amplifying the differences in positive sentiment features.
[0189] Next, the electronic device can normalize the 256-dimensional vector to eliminate internal scale shifts, using the following formula: μ is the vector mean, σ is the standard deviation, and ε = 10. -6 To prevent the denominator from being zero, electronic devices can be set to a dropout rate of 0.15, randomly blocking 15% of neurons to suppress overfitting and improve generalization in occlusion and ECG noise scenarios. .
[0190] Finally, the second fully connected layer maps to 128 dimensions as a deep refined feature: This refers to the fully connected output features (128 dimensions), which carry deep emotional semantics after compression, nonlinear fitting, and regularization denoising.
[0191] Step a82: The fully connected output features are fused with the temporal feature vector to obtain the hidden layer feature vector.
[0192] Specifically, the 1024-dimensional temporal feature vector is much larger than the 128-dimensional fully connected output feature F.fc Therefore, electronic devices can add a lightweight 1×1 convolutional projection layer to compress the temporal feature vector TimeSeq to a 128-dimensional matching size, as shown in the formula: .
[0193] The electronic device can be configured with a balance coefficient γ=0.7, favoring the refinement of deep features while preserving original temporal details. The fully connected output features are then fused with the temporal feature vector to obtain the hidden layer feature vector, as shown in the formula: After fusion, a nonlinear transformation via ReLU activation is applied again: H hidden The hidden layer feature vector (fixed at 128 dimensions).
[0194] Step a83: Map the hidden layer feature vector to the emotion probability vector.
[0195] The emotion probability vector includes at least one emotion type among calm, joy, surprise, anxiety, anger, and fear.
[0196] Specifically, electronic devices can map 128-dimensional hidden layer vectors to 6-dimensional original score vectors based on the terminal classification FC layer, corresponding to six emotion categories: , The score has no range constraint and can be positive or negative.
[0197] Then, the electronic device can perform a softmax operation on the 6-dimensional raw score element by element, converting it into a standard probability vector with a value range of [0,1] and a sum of 1: Finally, the electronic device can output an emotion probability vector. Each term in the vector represents the confidence probability of the pilot's corresponding emotion at the current moment.
[0198] Step S205: Based on the emotion probability vector, determine the current flight risk index corresponding to the target pilot.
[0199] Specifically, step S205 above may include the following steps: Step S2051: Obtain the current flight phase.
[0200] Specifically, the electronic equipment can read flight bus data in real time, classify the flight into three phases, and obtain phase labels. Phase 1: Upwind / crosswind climb (medium load); Phase 2: Downwind cruise (low load, high relaxation); Phase 3: Approach and landing (high load, high precision control).
[0201] Electronic devices can convert discrete phase labels into 3D one-hot encoded vectors as prior information about the scene. Pilots exhibit varying degrees of freedom and explicitness in facial emotional expression across different flight phases; their expressions are more restrained during approach and more relaxed during cruise, providing scene constraints for subsequent feature correction.
[0202] Step S2052: Determine the emotion weights corresponding to each emotion type in the emotion probability vector based on the current flight phase.
[0203] Specifically, electronic devices can determine the emotion weights corresponding to each emotion type in the emotion probability vector based on the correspondence between the current flight phase and each emotion type.
[0204] For example, if the current flight phase is Phase 2 (cruise low load), then the emotion weights corresponding to each emotion type in the emotion probability vector are: If the current flight phase is Phase 1 (climbing with moderate load), then the emotion weights corresponding to each emotion type in the emotion probability vector are: If the current flight phase is Phase 3 (high approach load), then the emotion weights corresponding to each emotion type in the emotion probability vector are: .
[0205] Step S2053: Based on the emotion weight, the probability values corresponding to each emotion type in the emotion probability vector are weighted and calculated to obtain the basic flight risk index.
[0206] Specifically, electronic devices can use the probabilities of various emotions as coefficients, multiply them by the corresponding emotional weights for each stage, and sum them up to obtain a basic flight risk index. The formula is as follows: ,in, Let k be the probability of the k-th emotion. The basic flight risk index naturally falls below... The maximum value does not exceed 0.85; it only represents the theoretical risk of emotion itself, without taking into account actual operational performance and individual differences among pilots.
[0207] Step S2054: Obtain the current operational deviation corresponding to the target pilot.
[0208] The current operational deviations include the standard deviation of yaw angle and the standard deviation of altitude.
[0209] Specifically, electronic devices can extract two core control deviation indicators from the real-time statistics of the airborne inertial navigation and flight control systems: σ yaw Yaw angle standard deviation represents heading stability; the larger the value, the more frequent the heading corrections and the greater the maneuvering fluctuations; σ altThe standard deviation of altitude represents the accuracy of altitude maintenance; the larger the value, the more obvious the altitude fluctuations and the less stable the control; the two deviations are matched synchronously with the timestamp of the time sequence characteristics of this segment, and the time sequence is completely aligned.
[0210] Step S2055: Determine the deviation correction parameters based on the current operational deviation.
[0211] Specifically, the electronic equipment can pre-store the reference yaw angle deviation σ of the target pilot in a calm and stable flight state. yaw_base and the standard deviation of the reference height σ alt_base .
[0212] Electronic equipment can calculate the relative deviation of yaw angle and relative deviation of altitude based on the standard deviation of yaw angle, the standard deviation of altitude, and the standard deviation of reference yaw angle and reference altitude. The calculation formula is as follows: , Among them, Dev yaw For the relative deviation of the yaw angle, Dev alt This represents the relative deviation in height.
[0213] Then, the electronic equipment can calculate the average of the relative deviation of yaw angle and the relative deviation of altitude as the overall deviation correction parameter: .
[0214] Step S2056: Obtain the personal correction parameters corresponding to the target pilot.
[0215] Specifically, the electronic device can receive the user-inputted personal correction parameters corresponding to the target pilot.
[0216] Step S2057: Based on the deviation correction parameters and the individual correction parameters, the basic flight risk index is corrected to obtain the current flight risk index.
[0217] Specifically, electronic devices can multiply the base flight risk index by the deviation correction parameter and the personal correction parameter to obtain the current flight risk index.
[0218] The flight risk index assessment method provided in this application determines the feature vector of a single-frame static facial action unit. It quantifies the activation level of nine key AUs based on the FACS standard, accurately quantifying the instantaneous basic facial expression state; inheriting the weighted parameters of previous occlusion, it weakens the interference from flight equipment occlusion areas, retaining only the core emotional information of the eyebrows, eyes, and lips, establishing a stable and quantifiable static emotional benchmark. It calculates the difference in activation intensity of the same AU in adjacent frames to obtain inter-frame temporal difference features. It captures the rate and magnitude of dynamic changes in facial expressions, distinguishing between stable emotions and instantaneous fright or rage caused by bird strikes or malfunctions; it overcomes the deficiency of static AUs, which can only reflect the current state and cannot reflect the speed of emotional fluctuations, significantly improving sensitivity to sudden stress conditions. The static vector and temporal difference features are fused to obtain comprehensive facial features. It takes into account both long-term calm / anxiety states and millisecond-level emotional abrupt changes, providing more complete feature information and offering a complete facial emotional representation for subsequent multimodal attention matching and emotion classification, improving recognition and discrimination.
[0219] Next, the ECG signal is identified and the R wave is located. Using a modified Pan-Tompkins algorithm combined with neighborhood verification, heartbeat peaks are accurately marked, mitigating false positives and false negatives caused by cockpit motion artifacts, and providing a precise temporal coordinate basis for RR interval calculation. Reasonable RR interval ranges are obtained. Thresholds are dynamically set based on the pilot's resting and high-stress physiological states, conforming to the physiological limits of the human heartbeat, establishing a reliable true / false interval judgment scale suitable for different flight load conditions. False peaks are removed to obtain the ECG signal to be extracted. Abnormal peaks are iteratively screened out, the RR interval time series is purified, and noise distortion is eliminated to prevent interference with subsequent heart rate variability calculations, ensuring that all feature calculations are based on the true heartbeat rhythm. ECG features are extracted in multiple dimensions, including the time domain, frequency domain, and two types of nonlinearity. This system comprehensively characterizes the autonomic nervous system state from multiple dimensions: the time domain reflects average heart rate and instantaneous heart rate fluctuations; the frequency domain quantifies the balance between the sympathetic and parasympathetic nervous systems; Poincaré double standard deviation distinguishes between long-term and short-term heart rate variability; eccentricity and sample entropy accurately differentiate between long-term anxiety disorders and instantaneous startle fluctuations, with multiple complementary indicators covering implicit emotional physiological representations. Multiple feature classes are fused to generate target physiological features. Z-score standardization eliminates dimensional differences, and orderly splicing forms a unified physiological vector; weights are simultaneously adjusted based on the signal-to-noise ratio, allowing high-quality signals to fully play their role, resulting in stable and regular overall features that can be directly matched and fused with facial visual features across modalities.
[0220] Next, the facial feature input is processed by the facial branch, which outputs a visual feature vector. This branch incorporates dynamic-static separation encoding, ROI occlusion weight inheritance, AU emotion combination enhancement, and a transient mutation amplification mechanism to deeply refine explicit transient facial emotion information, uniformly map to fixed dimensions, strengthen the ability to recognize sudden facial expression changes under dangerous conditions, and shield against noise from flight equipment occlusion. The physiological feature input is processed by the ECG branch, which outputs an ECG feature vector. Grouped 1D-CNN stationary convolution, SNR adaptive correction, and implicit anxiety steady-state enhancement are employed to specifically mine long-term suppressed emotions and autonomic nervous system steady-state states that cannot be reflected in the face. This forms a complementary dynamic-static dynamic-static relationship with the visual features, and the alignment of the dimensions with the visual vector facilitates fusion calculation. The visual and ECG vectors are fused to obtain the initial fused feature vector. Fusion weights are adaptively assigned based on facial dynamic intensity and modal confidence. In cases of sudden danger, transient facial features are prioritized, while in stable conditions, implicit stress signals from the ECG are emphasized, balancing the reliability of both modalities and outputting the initial multimodal fused features. The current flight phase is then obtained. Prior information on climb, cruise, and approach conditions is captured. Significant differences in control tolerance and emotional interference levels exist across different phases, providing a basis for scenario determination in subsequent attention constraints and risk weighting. Flight phase encoding is concatenated with initial fusion features to obtain a fused phase feature vector. Embedding scenario information from the flight phase allows the model to distinguish differences in emotional expression under different control loads, accurately identifying subtle negative emotions deliberately suppressed during the approach phase, thus improving scenario adaptability. Q-visual queries and K / V ECG key-value vectors are assigned. The original attention score matrix is calculated based on the query vector and key vector. The similarity between each dimension of visual and ECG features is quantified by scaling the dot product, accurately quantifying the intrinsic correlation between facial and physiological emotional representations, and establishing a basic numerical system for cross-modal matching. Current flight parameters are acquired. Multi-dimensional quantitative data on attitude, power, control, and alarms are collected to provide objective quantitative basis for scenario classification, avoiding biases in condition determination caused by relying solely on phase labels. The current scenario type is determined based on flight parameters. The system automatically distinguishes between three scenarios: smooth cruise, moderate-load operation, and sudden high-stress failure, determining the pilot's flight scenario or operational load state and providing scenario labels for differentiated attention constraints. The scenario type is used to mask the original score matrix, resulting in a masked score matrix. Invalid matching channels are selectively masked for different scenarios; cruise mode suppresses interference from ECG stress indicators, and failure conditions mask pleasant facial features, reducing the weight of feature pairings with low relevance to the current flight scenario and minimizing noise matching interference. Masked scores are normalized to generate a basic attention weight matrix. Softmax normalization converts the scores into a standard weight distribution of 0-1 with a row sum of 1, unifying the attention scale across dimensions and ensuring that weight values can stably participate in subsequent feature weighting calculations. Modal confidence scores are used to correct the basic weights, resulting in the target attention weight matrix.The system adaptively adjusts weights based on the degree of facial occlusion and signal-to-noise ratio. Image occlusion reduces the contribution of the visual modality to the fusion result, while large ECG artifacts weaken the ECG matching contribution. This adaptive balance of dual-modal acquisition quality differences enhances robustness under harsh acquisition conditions. The system fine-tunes the volumetric matching strength according to climb / cruise / approach loads, amplifying weak emotional correlations during high-pressure approaches. Heatmaps visually represent the correlation strength between visual AU and ECG indicators, improving model interpretability. A time-series smoothing of the initial heatmap yields the target heatmap. Multi-frame weighted averaging suppresses weight jumps caused by single-frame occlusion and ECG impulse noise, ensuring a stable and continuous heatmap and preventing drastic inter-frame jitter in emotion recognition results. Global fusion features are obtained by weighted fusion of visual and ECG data based on the smoothed heatmap. Instead of simple equal addition, attention autonomously selects highly correlated feature pairs: dangerous situations emphasize the combination of fear / surprise AU and high-stress ECG, while cruising emphasizes pleasant and stable heart rate matching, resulting in higher information utilization. Global fusion features are regularized to obtain unified fusion features. Layer normalization balances visual and ECG scale differences, and Dropout suppresses overfitting, enhancing the model's generalization stability under cockpit occlusion and motion artifact environments. Unified fusion features are split into multiple subsequences. Fine-grained parsing of feature blocks is achieved, processing visual and ECG feature segments separately, facilitating local filtering, correction, and temporal compensation, avoiding the loss of local difference information by processing the entire vector singularly. Q and K distribution differences within subsequences are calculated to select high-contribution core vectors. The feature discrimination capability is quantified based on distribution differences, filtering out low-contribution redundant vectors and retaining only matching features with high emotional recognition accuracy, reducing invalid computation and noise interference. Weighted core vectors generate preliminary subsequence encoding features. Visual and ECG core vectors are fused using contribution as weight, adaptively balancing local modality proportions to generate basic encoding with local emotional semantics. Baseline flight parameters for the pilot's calm and stable state are retrieved. A personalized steady-state reference standard for the pilot is established, abandoning a uniform group baseline and adapting to differences in pilot operating habits and physiological tolerance. Current parameters are compared with baseline parameters item by item. The system quantifies the actual deviations in manipulation, posture, and power to objectively measure the magnitude of operational stress, rather than relying on rough stage labels. Parameter anomaly levels are categorized based on deviations. A quantitative distinction is made between no-anomaly, mild, moderate, and severe stress conditions, providing a grading basis for subsequent compensation amplitude and feature correction coefficients. Anomaly levels are bound to initial codes to generate the first backup code feature. The code amplitude is fine-tuned according to the stress level; high-pressure conditions are moderately amplified to amplify emotional feature signals, while mild conditions are slightly corrected, aligning with the influence of stress on emotional expression. Compensation intervals are matched with preceding and following frames according to the anomaly level to extract the second backup code set. Higher stress results in a larger compensation frame range, relying on multi-frame temporal information to compensate for feature distortions caused by single-frame occlusion and ECG artifacts. The second backup codes within the interval are temporally weighted and fused to obtain weighted compensation features. Nearer frames have higher weights, while farther frames have lower weights, ensuring real-time response to sudden emotional changes while utilizing historical frames to smooth and suppress severe jitter in single frames.The primary and backup encodings are weighted and fused to obtain the sub-temporal feature vector. The balance between the real features of the current frame and the temporal smoothing compensation is achieved; the greater the anomaly, the higher the compensation ratio. For stable conditions, the current frame takes precedence, balancing real-time performance and stability. All sub-temporal features are fused to obtain the overall temporal feature vector. The complete feature dimensions are then concatenated to restore the overall feature vector, integrating all optimization information from local filtering, personalized baseline correction, and hierarchical temporal compensation, outputting a continuous, stable, and more accurate global temporal emotion representation. The temporal feature vector undergoes dimensionality reduction, nonlinear transformation, and regularization to obtain the fully connected output features. Dimensionality reduction compresses high-dimensional redundant information; nonlinear transformation fits the complex coupling relationship between emotion, condition, and physiology; layer normalization and Dropout dual regularization suppress overfitting, filter acquisition noise, and extract highly discriminative deep emotional semantics. The fully connected output features are fused with the temporal feature vector to obtain the hidden layer feature vector. A residual fusion structure is adopted, preserving both the deep features refined by dimensionality reduction and the detailed information of the original temporal compensation and personalized correction, avoiding the loss of key emotional details during dimensionality reduction, resulting in a more complete feature representation. The hidden layer feature vectors are mapped to six emotion probability vectors. By combining the classification layer with the Softmax output, a normalized probability distribution is generated, which not only outputs a single emotion label, but also fully presents the confidence ratio of mixed emotions, quantifies the strength of each emotion, and provides accurate numerical input for subsequent staged risk weighting calculations.
[0221] Finally, the current flight phase is determined. Differentiating between climb, cruise, and approach control load scenarios reveals varying degrees of harm to flight safety from different emotions, providing a scenario-based basis for differentiated risk weighting. Weights are assigned to each emotion type based on the current flight phase. Under high-tolerance-pressure conditions like approach, negative emotions such as fear, anger, and anxiety are amplified, while negative impacts are mitigated during cruise, aligning with real-world flight safety impact patterns and ensuring a more consistent basic risk assessment with practical aviation logic. A basic flight risk index is calculated by weighted summation of emotion probabilities. The theoretical safety risk resulting from the superposition of various emotions is quantified, transforming discrete emotion probabilities into a single, comparable risk benchmark value. Yaw angle and altitude data are collected, and their standard deviations are calculated as operational deviations. Heading and altitude control stability objectively reflects the actual interference of emotions on control actions, compensating for the one-sidedness of relying solely on emotions to determine risk. Deviation correction parameters are calculated based on operational deviations. Greater control fluctuations result in an increased risk coefficient, achieving a linkage correction between emotional risk and actual control performance, ensuring risk assessment aligns with real-world control conditions. Pilot-specific correction parameters are retrieved. Personalized calibration is performed based on individual baseline information such as pilots' flight experience, eliminating individual assessment errors caused by uniform group standards. The final flight risk index is obtained by correcting the base index using dual parameters. Integrating three dimensions—emotional hazard of the scenario, actual operational deviation, and individual personnel differences—it outputs a standardized, predictable risk value of 0-1, providing accurate and reliable results that can be directly integrated with airborne safety warning systems.
[0222] This embodiment also provides a flight risk index assessment device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0223] This embodiment provides a flight risk index assessment device, such as... Figure 5 As shown, it includes: The acquisition module 301 is used to acquire multi-frame target facial image data and target electrocardiogram signal corresponding to the target pilot; The first extraction module 302 is used to identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot. The second extraction module 303 is used to identify the target electrocardiogram signal and extract the target physiological characteristics corresponding to the target pilot; The first determining module 304 is used to determine the emotion probability vector corresponding to the target pilot based on the comprehensive facial features and the target physiological features; The second determination module 305 is used to determine the current flight risk index corresponding to the target pilot based on the emotion probability vector.
[0224] The flight risk index assessment device provided in this embodiment of the invention can execute the flight risk index assessment method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0225] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0226] The following is a detailed reference. Figure 6 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 01, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 02 or a program loaded from a memory 08 into a random access memory (RAM) 03. The RAM 03 also stores various programs and data required for the operation of the electronic device. The processor 01, ROM 02, and RAM 03 are interconnected via a bus 04. An input / output (I / O) interface 05 is also connected to the bus 04.
[0227] Typically, the following devices can be connected to I / O interface 05: input devices 06 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 07 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 08 including, for example, magnetic tapes, hard disks, etc.; and communication devices 09. Communication device 09 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0228] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 09, or installed from memory 08, or installed from ROM 02. When the computer program is executed by processor 01, it performs the functions defined in the flight risk index assessment method of the embodiments of the present invention.
[0229] Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0230] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded via a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the flight risk index assessment method shown in the above embodiments is implemented.
[0231] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0232] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for assessing flight risk index, characterized in that, The method includes: Acquire multi-frame facial image data and electrocardiogram signals corresponding to the target pilot; The facial image data of each target are identified, and the comprehensive facial features corresponding to the target pilot are extracted. The target's electrocardiogram signal is identified, and the target's physiological characteristics corresponding to the target pilot are extracted; Based on the facial features and the target physiological features, the emotion probability vector corresponding to the target pilot is determined; Based on the emotion probability vector, the current flight risk index corresponding to the target pilot is determined.
2. The method according to claim 1, characterized in that, The step of identifying each of the target facial image data and extracting the comprehensive facial features corresponding to the target pilot includes: The target facial image data is identified to determine the static facial motion unit feature vector corresponding to each target facial image data. The activation intensity difference of the same facial action unit between the feature vectors of the static facial action unit in two adjacent frames is calculated to obtain the inter-frame temporal difference feature. The static facial motion unit feature vector is fused with the inter-frame temporal difference feature to obtain the comprehensive facial feature.
3. The method according to claim 1, characterized in that, The process of identifying the target electrocardiogram signal and extracting the target physiological characteristics corresponding to the target pilot includes: The target electrocardiogram signal is identified, and the R wave is located from the target electrocardiogram signal; Obtain a reasonable range for the preset RR interval; If the detected peak exceeds the reasonable range of the preset RR interval, it is determined to be a noise spurious peak and removed to obtain the ECG signal to be extracted. Feature extraction is performed on the ECG signal to be extracted to obtain the time-domain features, frequency-domain features, basic nonlinear features, and preset nonlinear features corresponding to the ECG signal to be extracted; the time-domain features include the average value of all RR intervals and the root mean square of the difference between adjacent RR intervals; the frequency-domain features include a first frequency power, a second frequency power, and the ratio between the first frequency power and the second frequency power; the first frequency is less than the second frequency; the basic nonlinear features include the standard deviation of the minor axis of the Poincaré plot and the standard deviation of the major axis of the Poincaré plot; the preset nonlinear features include the eccentricity of the Poincaré plot ellipse and the sample entropy of the RR interval; The time-domain features, frequency-domain features, basic nonlinear features, and preset nonlinear features are fused to obtain the target physiological features corresponding to the target pilot.
4. The method according to claim 1, characterized in that, The step of determining the emotion probability vector corresponding to the target pilot based on the facial comprehensive features and the target physiological features includes: The comprehensive facial features are input into the facial feature recognition branch, and the target visual feature vector is output. The target physiological characteristics are input into the ECG feature recognition branch, and the ECG feature vector is output. The target visual feature vector and the electrocardiogram feature vector are fused to obtain an initial fused feature vector; Obtain the current flight phase; The current flight phase is concatenated with the initial fused feature vector to obtain the fused phase feature vector; Based on the feature vector of the fusion stage, the emotion probability vector corresponding to the target pilot is determined.
5. The method according to claim 4, characterized in that, The step of determining the emotion probability vector corresponding to the target pilot based on the feature vector of the fusion stage includes: The target visual feature vector in the feature vector of the fusion stage is used as the query vector, and the electrocardiogram feature vector in the feature vector of the fusion stage is used as the key vector and value vector. Based on the query vector, the key vector, and the value vector, the cross-modal attention weights are calculated to obtain the target cross-modal attention weight matrix; Based on the current flight stage in the feature vector of the fusion stage, the target cross-modal attention weight matrix is corrected to obtain an initial attention weight heatmap; the horizontal axis of the attention weight heatmap corresponds to the visual feature region, the vertical axis corresponds to the electrocardiogram feature region, and the color intensity represents the magnitude of the attention weight. The initial attention weight heatmap is subjected to temporal weight smoothing to obtain the target attention weight heatmap; Based on the target attention weight heatmap, the target visual feature vector and the electrocardiogram feature vector are fused to obtain a global fused feature vector. The global fusion feature vector is regularized to obtain a unified fusion feature vector; Temporal feature processing is performed on the unified fusion feature vector to obtain a temporal feature vector; Based on the temporal feature vector, the emotion probability vector corresponding to the target pilot is determined.
6. The method according to claim 5, characterized in that, The step of calculating cross-modal attention weights based on the query vector, the key vector, and the value vector to obtain the target cross-modal attention weight matrix includes: Based on the query vector and the key vector, calculate the original attention score matrix; Obtain current flight parameters; Based on the current flight parameters, determine the current scenario type; Based on the current scene type, a masking operation is performed on the original attention score matrix to obtain a masked attention score matrix; The mask attention score matrix is normalized to obtain the basic cross-modal attention weight matrix; Based on the modal confidence scores corresponding to the target visual feature vector and the electrocardiogram feature vector, the basic cross-modal attention weight matrix is modified to obtain the target cross-modal attention weight matrix.
7. The method according to claim 5, characterized in that, The step of performing temporal feature processing on the unified fused feature vector to obtain a temporal feature vector includes: The unified fused feature vector is divided into multiple sub-sequences; For each of the subsequences, the distribution difference of all the query vectors and key vectors within the subsequences is calculated, and the core vectors with a contribution greater than a preset contribution threshold are selected. The weighted calculation of each core vector is performed to obtain the preliminary coding features corresponding to the subsequence; For each of the aforementioned preliminary coding features, the reference flight parameters corresponding to the target pilot are obtained; the reference flight parameters are flight parameters collected when the target pilot is calm and the target aircraft is flying steadily. Compare the current flight parameters corresponding to each of the preliminary encoded features with the baseline flight parameters; Based on the comparison results, the anomaly level corresponding to the current flight parameters is determined; The anomaly level is bound to the preliminary coding feature to obtain the first backup coding feature corresponding to each subsequence; Based on the anomaly level corresponding to each of the first backup coding features, a compensation interval corresponding to each of the first backup coding features is determined, and a second backup coding feature corresponding to the compensation interval is obtained; the compensation interval represents the number of preset frames before and after each of the sub-sequences. The second backup coding features are weighted and fused to obtain the weighted compensation features corresponding to the subsequence; The first backup coding feature and the weighted compensation feature are weighted and fused to obtain the sub-temporal feature vector corresponding to the sub-sequence; The sub-time series feature vectors are fused together to obtain the time series feature vector.
8. The method according to claim 5, characterized in that, Determining the emotion probability vector corresponding to the target pilot based on the temporal feature vector includes: The time-series feature vector is subjected to dimensionality reduction, nonlinear transformation, and regularization to obtain fully connected output features; The fully connected output features are fused with the temporal feature vector to obtain the hidden layer feature vector; The hidden layer feature vector is mapped to the emotion probability vector; the emotion probability vector includes at least one emotion type among calm, joy, surprise, anxiety, anger, and fear.
9. The method according to claim 1, characterized in that, The step of determining the current flight risk index corresponding to the target pilot based on the emotion probability vector includes: Obtain the current flight phase; Based on the current flight phase, determine the emotion weight corresponding to each emotion type in the emotion probability vector; Based on the emotion weights, the probability values corresponding to each emotion type in the emotion probability vector are weighted and calculated to obtain the basic flight risk index. Obtain the current operational deviation corresponding to the target pilot; the current operational deviation includes the standard deviation of yaw angle and the standard deviation of altitude. Based on the current operational deviation, determine the deviation correction parameters; Obtain the individual correction parameters corresponding to the target pilot; Based on the deviation correction parameters and the personal correction parameters, the basic flight risk index is corrected to obtain the current flight risk index.
10. A flight risk index assessment device, characterized in that, The device includes: The acquisition module is used to acquire multi-frame facial image data and electrocardiogram signals of the target pilot. The first extraction module is used to identify the facial image data of each target and extract the comprehensive facial features corresponding to the target pilot; The second extraction module is used to identify the target electrocardiogram signal and extract the target physiological features corresponding to the target pilot; The first determining module is used to determine the emotion probability vector corresponding to the target pilot based on the facial comprehensive features and the target physiological features; The second determining module is used to determine the current flight risk index corresponding to the target pilot based on the emotion probability vector.