Method and system for identifying unsafe behaviors based on expressions and physiological indexes of pilot

The unsafe behavior identification system, which utilizes multidimensional data acquisition, intelligent preprocessing, and feature optimization, solves the problems of poor timeliness, insufficient multimodal data correlation, and unbalanced data processing in existing technologies. It enables real-time identification and accurate early warning of unsafe pilot behaviors, thereby improving aviation safety.

CN121743992AActive Publication Date: 2026-03-27CHINA ACAD OF CIVIL AVIATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing pilot behavior pattern recognition technologies suffer from poor timeliness, insufficient multimodal data correlation, and limited ability to process imbalanced data, making it impossible to provide early or real-time warnings. Furthermore, existing algorithms struggle to effectively address class imbalance issues, resulting in low accuracy in minority class pattern recognition.

Method used

By employing multi-dimensional data acquisition, intelligent preprocessing, feature optimization, differentiated category balancing enhancement, and parallel competitive training of multiple algorithms, an unsafe behavior recognition system based on pilot facial expressions and physiological indicators is constructed to achieve end-to-end real-time monitoring and early warning.

Benefits of technology

It significantly improves the response speed of flight safety monitoring, enhances identification accuracy and interpretability, reduces the risk of false alarms, and strengthens the robustness and applicability of the system in real flight scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743992A_ABST
    Figure CN121743992A_ABST
Patent Text Reader

Abstract

The invention provides an unsafe behavior identification method and system based on pilot expressions and physiological indexes, and relates to the technical field of modern aviation safety management, and the method comprises the steps: multi-dimensional data collection: collecting pilot facial expression data through a camera and expression calculation software, and collecting pilot physiological index data through wearable equipment; flight parameter data are collected through a flight quality monitoring system, pilot behavior mode labels are generated, and pilot facial expression data, pilot physiological index data and the flight parameter data are summarized to form an original data set. The problems of poor timeliness and lag of the identification window in the prior art are effectively solved. Specifically, the system utilizes a camera and a wearable device to realize non-invasive real-time monitoring of facial expressions and physiological indexes, and synchronously generates a behavior mode label through a flight quality monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of modern aviation safety management, in particular to an unsafe behavior identification method and system based on pilot expressions and physiological indicators. BACKGROUND

[0002] With the continuous development of aviation safety management, flight quality monitoring (QAR) system has become one of the core technologies of modern aviation. This system can effectively represent the operation behavior pattern of pilots by recording and analyzing various parameters in the flight process in real time, such as altitude, speed and attitude change, etc. At present, professional flight quality analysis algorithm has been widely used in practice, which can mark and evaluate the pilot behavior pattern based on QAR parameter characteristics after the flight, supporting the safety management and training optimization of airlines.

[0003] However, the existing pilot behavior pattern identification technology faces the following key technical problems in actual application:

[0004] Poor timeliness and recognition window lag, the existing QAR-based monitoring mode mainly relies on post-data processing, which cannot realize pre-warning or in-process warning, so that potential unsafe behavior cannot be intervened in time;

[0005] Insufficient multi-modal data correlation, although the collection means of facial expressions and physiological indicators (such as heart rate and blood oxygen) are relatively mature, there is a lack of effective data processing and analysis mechanism, which cannot form a reliable correlation between these behavior precursors and flight patterns;

[0006] Limited imbalance data processing capability, the number of unsafe mode samples in real flight scenarios is small, and the existing algorithm is difficult to effectively handle the class imbalance problem, resulting in low recognition accuracy of minority class mode.

[0007] Therefore, an unsafe behavior identification method and system based on pilot expressions and physiological indicators are needed to solve the above problems. SUMMARY

[0008] Technical problems to be solved

[0009] In view of the deficiencies of the prior art, the present application provides an unsafe behavior identification method and system based on pilot expressions and physiological indicators, which solves the problems of the prior art.

[0010] Technical scheme

[0011] To achieve the above purpose, the present application is implemented by the following technical scheme: an unsafe behavior identification method and system based on pilot expressions and physiological indicators, comprising the following steps:

[0012] Sp1: Multidimensional data acquisition, pilot facial expression data is collected through a camera and expression calculation software, pilot physiological indicator data is collected through a wearable device, flight parameter data is collected through a flight quality monitoring system and pilot behavior pattern labels are generated, the pilot facial expression data, pilot physiological indicator data and flight parameter data are summarized to form an original data set; the facial expression data includes time series data of valence and arousal, the physiological indicator data includes time domain feature data of heart rate variability, heart rate, blood oxygen saturation and body temperature; the camera and expression calculation software, wearable device are in data interaction with the flight quality monitoring system to ensure time synchronization of the collected data;

[0013] Sp2: Original data preprocessing, the original data set is screened, quality screened and reorganized, missing values are processed using a multi-level intelligent filling strategy, features are standardized, data is reorganized into a variable length time series sample data set, and a hierarchical cross-validation method is used to divide the training set and the test set; the multi-level intelligent filling strategy is constructed based on individual baseline and group baseline, combined with the physiological meaning of the features, and the missing feature filling is realized by using correlation calculation, multivariate derivation, baseline quantile sampling and physiological constraint bottom-up scheme; the Sp2 receives the original data set summarized in Sp1, and the processed data is transmitted to Sp3;

[0014] Sp3: Feature optimization, a two-stage progressive optimization architecture is adopted, the first stage expands to generate a candidate feature pool through frequency domain feature extraction, statistical feature extraction and physiological feature extraction, the second stage performs intelligent feature selection based on a multi-dimensional feature importance evaluation system to obtain an optimal feature subset; the multi-dimensional feature importance evaluation system integrates RandomForest importance evaluation, correlation analysis, variance analysis, statistical significance test and physiological prior knowledge dimension; Sp3 interacts with Sp2 data, receives preprocessed data and outputs the optimal feature subset to Sp4;

[0015] Sp4: Differentiated class balance enhancement, based on the class imbalance degree of each flight mode label, a balanced class enhancer or an extreme class balancer is intelligently selected to perform data enhancement on the training set; the balanced class enhancer adopts a mild balancing strategy, integrates time warping, amplitude warping and Gaussian noise injection technology and applies physiological constraints; the extreme class enhancer adopts an aggressive balancing strategy, integrates synthetic sample generation and frequency domain offset technology, and cooperates with negative sample downsampling mechanism; Sp4 receives the optimal feature subset output by Sp3, and the enhanced training set is transmitted to Sp5;

[0016] Sp5: model training, construct an algorithm pool containing traditional machine learning algorithms and deep learning algorithms, dynamically select candidate algorithm combinations based on the distribution characteristics of the enhanced data, and complete model training using a parallel competitive training mechanism; the deep learning algorithm contains an Enhanced GRU model that integrates FocalLoss and SMOTE technology and supports variable-length time series processing; Sp5 interacts with Sp4 data, receives the enhanced training set, and transmits the trained model to Sp6;

[0017] Sp6: intelligent model selection, using a four-stage selection strategy to determine the optimal model, realizing the identification and early warning of pilot unsafe behavior patterns; the four-stage selection strategy is quality screening stage, anomaly detection stage, business scoring stage and final selection stage in turn, and the business scoring stage uses a pre-set weight configuration to calculate the comprehensive business score; Sp6 receives the trained model output by Sp5, and outputs the optimal model for unsafe behavior pattern identification and early warning after selection.

[0018] Preferably, the screening of the original data set in Sp2 includes flight phase screening and operator screening, and data records marked as landing phase and actual execution operators are retained; the evaluation standard of the quality screening is that each feature has at least 3 valid values, and at least 50% of the features are in the valid state.

[0019] Preferably, the process of reorganizing the data into a variable-length time series sample data set in Sp2 uses a variable-length sequence storage method to maintain the original time length of each flight sample; the stratified cross-validation uses a five-fold multi-label stratified K-fold algorithm to ensure that the proportion of various label combinations in each fold is consistent with the overall data set.

[0020] Preferably, the first stage of frequency domain feature extraction in Sp3 uses Welch power spectral density estimation method for heart rate variability signal, extracts power values and derived indicators in very low frequency band, low frequency band and high frequency band, and actively excludes total power frequency band to reduce redundancy; the first stage of physiological feature extraction calculates the abnormal value proportion and range utilization of physiological features relative to the normal range.

[0021] Preferably, the class imbalance degree division in Sp4 is divided into four levels, namely super-minority class, high-minority class, moderate-minority class and super-extreme-minority class; the balanced class enhancer configures different enhancement multiples and technology combinations for super-minority class, high-minority class and moderate-minority class; the extreme class enhancer configures aggressive enhancement parameters for super-extreme-minority class, and the synthesized sample generation is achieved by mixing two real samples of the same class and injecting noise.

[0022] Preferably, the process of dynamically selecting candidate algorithm combination based on distribution characteristics of enhanced data in Sp5 is to generate data distribution portrait, build algorithm characteristic profile, implement hierarchical algorithm selection strategy based on data balance state, data complexity and enhancement multiple, and calculate algorithm comprehensive priority through multi-dimensional scoring function.

[0023] Preferably, the quality screening stage of the four-stage selection strategy in Sp6 sets performance threshold value adaptive to unbalanced data; the abnormality detection stage of the four-stage selection strategy identifies over-prediction mode of the algorithm; the weight configuration of the business scoring stage of the four-stage selection strategy is accuracy weight 40%, recall rate weight 35%, F1 score weight 15% and efficiency weight 10%; the final selection stage of the four-stage selection strategy selects the model with the highest comprehensive score, and starts the alternative selection mechanism when the candidate model pool is empty.

[0024] Preferably, the system comprises a data acquisition module, a data preprocessing module, a feature optimization engineering module, a data enhancement module, a model training module and an intelligent selection module, and the system is used to execute the method.

[0025] The data acquisition module is used to collect facial expression data, physiological index data and flight parameter data of pilots and generate behavior pattern labels, and to summarize to form an original data set; the data acquisition module comprises a camera and expression calculation software, a wearable device and a flight quality monitoring submodule; the camera and expression calculation software, the wearable device are in data interaction with the flight quality monitoring submodule to ensure the time synchronization of the collected data; the data acquisition module is in communication connection with the data preprocessing module, and transmits the original data set to the data preprocessing module;

[0026] The data preprocessing module is used to filter, quality screen and recombine the original data set, process missing values, standardize the features, recombine the data and divide the training set and the test set; the data preprocessing module is built-in a multi-level intelligent filling unit which constructs individual baseline and group baseline and selects corresponding filling scheme in combination with the physiological meaning of the features; the data preprocessing module is in data interaction with the feature optimization engineering module, and transmits the processed data to the feature optimization engineering module;

[0027] The feature optimization engineering module is used to realize feature expansion and intelligent feature selection to obtain an optimal feature subset; the feature optimization engineering module comprises an intelligent feature engineer and an intelligent feature selector, the intelligent feature engineer integrates frequency domain, statistical and physiological feature extraction technology, and the intelligent feature selector adopts a multi-dimensional feature importance evaluation system; the feature optimization engineering module is in communication connection with the data enhancement module, and transmits the optimal feature subset to the data enhancement module;

[0028] The data enhancement module is used to solve the class imbalance problem of training data, and contains a balanced class enhancer and an extreme class balancer; the balanced class enhancer adopts a mild balancing strategy and applies physiological constraints, and the extreme class enhancer adopts an aggressive balancing strategy and supports synthetic sample generation and negative sample downsampling; the data enhancement module interacts with the model training module, and transmits the enhanced training set to the model training module;

[0029] The model training module is used to realize multi-algorithm parallel competition training, and contains an algorithm pool, a dynamic algorithm selector and an algorithm factory; the algorithm pool contains traditional machine learning algorithms and deep learning algorithms, the dynamic algorithm selector selects candidate algorithm combinations based on data distribution characteristics, and the algorithm factory provides a unified algorithm instance creation and calling interface; the model training module is in communication connection with the intelligent selection module, and transmits the trained model to the intelligent selection module;

[0030] The intelligent selection module is used to determine the optimal model through a four-stage selection strategy, and realizes identification and early warning of pilot unsafe behavior patterns; the intelligent selection module contains a quality screening unit, an anomaly detection unit, a business scoring unit and a final selection unit; the intelligent selection module is in instruction transmission with the data acquisition module, and can adjust data acquisition parameters according to the identification result.

[0031] Preferably, the wearable device contained in the data acquisition module is a smart watch, which supports the collection of PPG-based heart rate time interval, heart rate mean, body temperature and blood oxygen value data; the expression calculation software contained in the data acquisition module is a facial expression analysis system, which can analyze seven basic emotions, valence and arousal indicators.

[0032] Preferably, the deep learning algorithm in the algorithm pool contained in the model training module contains Enhanced_GRU, GRU, LSTM_Attention and CNN-LSTM, and the traditional machine learning algorithm contains RandomForest, XGBoost, LightGBM, SVM, LogisticRegression and GradientBoosting; the Enhanced_GRU model integrates FocalLoss and SMOTE technology, and contains GRU stacking layer, global average pooling layer and full connection classification head.

[0033] Beneficial effects

[0034] The application provides a pilot expression and physiological index-based unsafe behavior identification method and system.

[0035] 1. The present application effectively solves the problems of poor timeliness and recognition window lag in the prior art by integrating multi-modal real-time data acquisition and intelligent processing mechanisms. Specifically, the system uses cameras and wearable devices to achieve non-invasive real-time monitoring of facial expressions and physiological indicators, and generates behavior pattern labels synchronously through a flight quality monitoring system. In the continuous process of Sp1 to Sp6, end-to-end processing from acquisition to early warning is completed, and pre-warning is moved forward. This design significantly improves the response speed of flight safety monitoring, reduces the risk of potential unsafe behavior, supports pilots to adjust operations in simulation training or actual flights, and improves the overall level of aviation safety.

[0036] 2. The present application overcomes the problem of insufficient multi-modal data correlation in the prior art through multi-modal feature fusion and dynamic algorithm selection. Specifically, in the Sp3 feature optimization step, two-stage architecture is used to expand and select time domain, frequency domain and physiological features, realizing effective fusion of expression data (such as valence, arousal) and physiological indicators (such as heart rate variability); in Sp5 model training, an algorithm pool is constructed and a suitable model is dynamically selected, further strengthening the correlation between behavior precursors and flight patterns. This innovation improves the accuracy and interpretability of recognition, enabling the system to reliably convert physiological signals into behavior warning basis, optimizing the risk management and decision support of airlines.

[0037] 3. The present application solves the problem of limited imbalance data processing capability in the prior art through differential category balance enhancement and four-stage intelligent model selection. Specifically, in Sp4, an intelligent balance or extreme enhancer is selected to process different imbalance levels (such as 8 times enhancement for super-minority class), combined with physiological constraints to ensure sample authenticity; in Sp6, a four-stage strategy (such as abnormal detection to exclude over-prediction) optimizes model performance. This mechanism significantly improves the recall rate and precision of minority class patterns (such as rarely seen unsafe behavior), reduces the risk of false positives, improves the robustness and applicability of the system in real flight scenarios, and has broad market prospects and social benefits. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 is the system framework diagram of the present application;

[0039] Figure 2 is the system core flowchart of the present application;

[0040] Figure 3 is the model training module flowchart of the present application;

[0041] Figure 4 is the Enhanced_GRU model structure diagram of the present application;

[0042] Figure 5This is a flowchart of the intelligent model selection process of the present invention. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Specific Implementation Example 1:

[0045] like Figures 1 to 5 As shown, this embodiment provides a comprehensive and detailed explanation of the "method and system for identifying unsafe behaviors based on pilot facial expressions and physiological indicators". This method achieves accurate identification and early warning of unsafe behavior patterns of pilots through the coordinated execution of six core steps, Sp1 to Sp6. The specific implementation process, technical parameters and operation details of each step are as follows.

[0046] I. Sp1: Detailed Implementation of Multidimensional Data Acquisition

[0047] The core objective of this step is to achieve the synchronous collection and aggregation of pilot facial expression data, physiological index data, and flight parameter data, providing a complete and time-consistent raw dataset for subsequent processing. Specific implementation details are as follows:

[0048] 1. Data Acquisition Equipment and Tool Configuration: A high-definition industrial camera (1920×1080 resolution, 30fps) is installed directly in front of the cockpit to ensure complete capture of the pilot's facial area; the facial expression calculation software uses an authoritative certified facial expression analysis system (FaceReader), which supports the recognition of seven basic emotions (happiness, sadness, anger, surprise, fear, disgust, and neutrality) with an average accuracy of over 96%, and can output time-series data of core emotional quantitative indicators such as valence and arousal; the wearable device is a smartwatch supporting PPG technology, with a sampling frequency set to 100Hz, and has the function of real-time acquisition of data such as heart rate time interval, average heart rate, body temperature (accuracy 0.1℃), and blood oxygen value (accuracy 1%); flight parameter data is collected through the aircraft's built-in Flight Quality Assurance (FOQA) system, which can record more than 20 flight parameters such as control stick actions, flight path, altitude, and speed in real time, and has a built-in pattern recognition submodule for generating behavioral pattern labels.

[0049] 2. Data collection process: Before the pilot performs the flight task, complete the camera calibration, smart watch binding (associated with the pilot's unique number) and FOQA system data interface configuration, and ensure that the three are time synchronized (based on the FOQA system time, the time error of the camera and the smart watch is controlled within ±50ms). During the flight, the camera captures a frame of facial image every 1 second and transmits it to the expression calculation software, which analyzes and outputs the valence and arousal quantization values (range 0-100) in real time, forming time series data; The smart watch collects the mean heart rate every 5 seconds, the body temperature every 10 seconds, and the blood oxygen value every minute, while continuously collecting and storing heart rate interval data; FOQA system collects flight parameter data at a frequency of 10Hz, and generates behavior pattern labels (including 6 patterns such as high slope and gradual change slope, coded as 0-1 integer) based on flight parameters (such as descent rate, airspeed, pitch angle, etc.). During the collection process, the camera, expression calculation software and wearable device interact with the FOQA system through Ethernet, and synchronize the timestamp every 30 seconds to ensure the time sequence consistency of the three types of data.

[0050] 3. Construction of raw data set: The three types of data collected are associated and matched according to "pilot unique number + flight number + timestamp", and are summarized to form a raw data set. The data set contains five categories of information: one is the index identification information (pilot number, flight number, time window identification); two is the facial expression data (time series of valence and arousal); three is the physiological index data (heart rate variability time domain features, heart rate statistical features, blood oxygen saturation features, body temperature features); four is the flight metadata (takeoff time, landing time, landing phase marker, operator marker); five is the behavior pattern label (0-1 coding of 6 flight patterns). The raw data set is stored in CSV format, and a single record contains 32 fields, covering all the information dimensions mentioned above.

[0051] II. Sp2: Details of the original data preprocessing

[0052] This step screens, quality screens, fills missing values, standardizes and divides the raw data set to ensure that the data quality meets the requirements of subsequent model training. The specific implementation details are as follows:

[0053] 1. Data screening: Two layers of screening logic are used to perform screening operations. The first layer is flight phase screening. Since the invention focuses on the identification of unsafe behaviors in the landing phase, the "landing phase marker" field is screened to only retain data records marked as "landing phase"; The second layer is operator screening, based on the "operator marker" field, the pilot data actually performing the landing operation is screened out (excluding monitoring personnel data) to avoid non-target data interference. After screening, the data volume is about 35%-40% of the original data (due to differences in landing time of different flights).

[0054] 2. Quality screening and reorganization: First, perform data column reorganization, delete columns that do not contribute to model training (such as departure time, aircraft number, etc.), and retain identification columns, core expression and physiological feature columns (a total of 23 columns), flight mode label columns; Then perform data quality evaluation, evaluation criteria are: ① Each feature has at least 3 valid values (to ensure statistical reliability); ② At least 50% of the features are valid (to avoid too sparse data). The pilot data that passes the evaluation is retained, and the data that does not pass the evaluation (such as a pilot's blood oxygen data missing rate of 60%) is directly removed, and the final data quality qualified rate is about 85%.

[0055] 3. Multi-level intelligent filling strategy implementation: For the missing values in the screened data, based on individual baseline and group baseline, and combined with the physiological meaning of the features, a differentiated filling scheme is adopted: ① Individual difference index (such as heart rate variability): preferentially use the historical median of the pilot's index to fill (at least 3 valid values are required), and use the group median if insufficient; ② Time domain indicators with mathematical correlation (such as heart rate mean, median): use correlation between indicators to fill, for example, when the heart rate mean is missing, use the median x 0.98 to estimate, when the median is missing, use the mean x 1.02 to estimate, the maximum value is estimated by mean + 2σ, and the minimum value is estimated by mean - 2σ, if the correlation filling fails, fall back to individual mean or physiological default value; ③ Oxygen saturation: based on the physiological characteristics of the pilot's oxygen level stability, randomly select a value between the 95th and 99th percentiles from the group statistics to fill; ④ Body temperature: fill with individual mean, seasonal adjustment, diurnal adjustment and random noise, subtract 0.2°C in winter (December-February), add 0.1°C in summer (June-August), add 0.1°C in daytime (6-14 o'clock), subtract 0.2°C at night (22-6 o'clock), and finally add normal distribution noise with mean 0 and standard deviation 0.1. The missing data rate after filling is controlled within 5%.

[0056] 4. Feature standardization: Due to the large difference in the dimensions of different physiological features (such as heart rate variability unit in milliseconds, body temperature unit in °C), Z-score standardization method is used. Stack all samples vertically to form a large matrix, calculate the mean and standard deviation for each feature independently, and standardize according to the formula (x-mean) / standard deviation to ensure that all feature values are mapped to the same order of magnitude (mean 0, standard deviation 1), avoiding the model being overly sensitive to large numerical values.

[0057] 5. Variable-length time series dataset construction and division: Use "pilot unique number + flight number" as the grouping key to aggregate all time window data for the same landing, arrange in chronological order to form continuous time series, and use variable-length sequence storage (retain original time length, do not fill or truncate); use five-fold multi-label stratified K-fold algorithm to divide training set and test set, first calculate the label combination of each sample, then stratified sampling to ensure that the proportion of each label combination in each fold is consistent with the overall dataset, avoid data leakage, finally the training set accounts for 70%, and the test set accounts for 30%.

[0058] III. Sp3: Detailed implementation of feature optimization

[0059] This step uses a two-stage progressive optimization architecture to achieve feature expansion and intelligent selection, and obtains the optimal feature subset. The specific implementation details are as follows:

[0060] 1. First stage: candidate feature pool expansion (core feature extraction): Integrate frequency domain, statistical, and physiological feature extraction techniques to expand the original 23-dimensional features to generate a 251-dimensional candidate feature pool: ① Frequency domain feature extraction: For heart rate variability signals, use Welch power spectral density estimation method, divide into three frequency bands: very low frequency band (0.0033-0.04Hz), low frequency band (0.04-0.15Hz), and high frequency band (0.15-0.4Hz), extract power values and derived indicators (such as low / high frequency ratio) in each frequency band, actively exclude total power frequency band to reduce redundancy, and generate 114-dimensional frequency domain features; ② Statistical feature extraction: Use adaptive sliding window strategy (window size is the minimum of sequence length 1 / 5 and 10, not less than 3), extract mean stability, standard deviation variability, trend strength, and volatility rate, a total of 76-dimensional statistical features; ③ Physiological feature extraction: Predefine normal ranges for each physiological indicator (heart rate 40-200 times / min, blood oxygen 85%-100%, etc.), calculate the proportion of outliers (out-of-range time steps) and range utilization (signal span / normal range span) for each feature, a total of 38-dimensional physiological features.

[0061] 2. Second stage: intelligent feature selection (optimal feature subset acquisition): based on a multi-dimensional feature importance evaluation system to screen the optimal feature subset, the evaluation system contains five dimensions: ①RandomForest importance evaluation: build a RandomForest classifier of 100 decision trees (maximum depth 10, minimum leaf node sample number 5), get the Gini importance score of each feature; ②Correlation analysis: calculate the Pearson correlation coefficient matrix between features, set the threshold to 0.95, and mark the highly correlated redundant features; ③Variance analysis: calculate the variance and coefficient of variation of each feature, and remove low-variance features whose variance is lower than the first percentile of the variance of all features; ④Statistical significance test: use F test (P<0.05 is significant) and mutual information method to evaluate the correlation strength of features and labels; ⑤Physiological prior knowledge: assign weights to different types of features (original physiological features 3.0, heart rate variability features 2.5, etc.). After normalizing the scores of the five dimensions, sum them up according to the weights (RandomForest 40%, variance 20%, mutual information 20%, physiological prior 20%), and select the top 100 features according to the scores as the optimal feature subset, with a feature reduction ratio of 40%.

[0062] Four, Sp4: detailed implementation of differential class balance enhancement

[0063] This step is based on intelligent selection of enhancer according to the degree of class imbalance to solve the problem of class imbalance in training data. The specific implementation details are as follows:

[0064] 1. Class imbalance degree classification: by counting the proportion of positive samples of each flight mode label, the imbalance degree is divided into four levels: super minority class (<5%), high minority class (5%-10%), moderate minority class (10%-25%), and super extreme minority class (<5% and sample number <100). According to statistics, the positive sample proportion of the "ground micro-stable" mode in the target application scenario is only 3.2% (super extreme minority class), the positive sample proportion of the "high slope" mode is 7.4% (high minority class), and the positive sample proportion of the "progressive change slope" mode is 20.3% (moderate minority class).

[0065] 2. Balanced class enhancer implementation (mainly used): Adopt a mild balancing strategy, configure different parameters for different imbalance levels: ① Super-minority class / highly minority class: adopt 5-8 times positive sample enhancement multiple, target positive sample ratio 25%, enable time distortion (scaling factor 0.8-1.2), amplitude distortion (scaling factor 0.7-1.3), and Gaussian noise injection (standard deviation 5 times) three techniques; ② Moderate minority class: adopt 2.5 times enhancement multiple, target ratio 35%, only enable time distortion and amplitude distortion techniques. Physiological constraints are applied during enhancement to ensure that the feature values after enhancement are within a reasonable range (such as heart rate variability change amplitude ≤30%, blood oxygen change amplitude ≤7%). Implement multiple rounds of iterative enhancement (up to 5 rounds), randomly select basic samples to generate new samples in each round until the target sample quantity is reached.

[0066] 3. Extreme class enhancer implementation (auxiliary): For super-extreme minority classes, adopt an aggressive balancing strategy, configure 20 times positive sample enhancement multiple, 0.8 times negative sample downsampling multiple, target positive sample ratio 40%, enable time distortion, amplitude distortion, Gaussian noise, frequency domain offset (frequency offset -0.1 to 0.1), and synthetic sample generation five techniques. Synthetic sample generation is achieved by mixing two real samples of the same class: after uniform sample length, randomly select a mixing coefficient of 0.3-0.7 to weight and combine, then inject different noise according to feature groups (heart rate variability noise standard deviation 8%, body temperature 2%), to ensure sample authenticity.

[0067] 4. Enhancement effect verification: After enhancement, the "ground micro-stable" mode positive sample ratio is increased to 38.5%, the "high slope" mode is increased to 24.5%, and the "progressive change slope" mode is increased to 33.3%, significantly improving class imbalance problems; after physiological reasonableness inspection, the feature values of the enhanced samples are within the normal range, and no false data is generated.

[0068] Five, Sp5: Detailed implementation of model training

[0069] This step builds an algorithm pool to realize multi-algorithm parallel competition training, and the specific implementation details are as follows:

[0070] 1. Algorithm pool construction: contains two categories of traditional machine learning algorithms and deep learning algorithms, a total of 10 algorithms: ① Traditional machine learning algorithms: RandomForest (integrates 100 decision trees), XGBoost (learning rate 0.1, maximum depth 6), LightGBM (leaf number 31, learning rate 0.05), SVM (Gaussian kernel function), LogisticRegression (regularization strength 0.1), GradientBoosting (learning rate 0.1); ② Deep learning algorithms: Enhanced_GRU (core innovative model), GRU (hidden layer dimension 64), LSTM_Attention (attention head number 4), CNN-LSTM (convolution kernel size 3x3, LSTM hidden layer dimension 64). Establish a feature profile for each algorithm, record its class imbalance handling ability, training speed, memory usage, etc.

[0071] 2. Dynamic algorithm selection: generate the distribution profile of the enhanced data, including three dimensions of balance state (positive sample ratio), data complexity (sample number, feature number), and enhancement multiple; based on the profile and algorithm feature profile matching, divide the core algorithm (3-4 optimal adaptation algorithms), exploration algorithm (1-4 potential algorithms), and standby algorithm (emergency algorithm) into three levels. For example, the core algorithm for super-extreme minority class data (poor balance state, high complexity) is Enhanced_GRU, XGBoost, and LSTM_Attention, and the exploration algorithm is LightGBM and CNN-LSTM; the core algorithm for moderate minority class data is RandomForest, XGBoost, and LightGBM. Calculate the priority through a multi-dimensional scoring function (imbalance handling ability 40%, training speed 30%, memory usage 30%) to determine the training order.

[0072] 3. Parallel competitive training mechanism: adopt a hybrid strategy of "traditional machine learning parallel + deep learning sequential": ① Traditional machine learning algorithms: concurrent training through thread pool (size ≤ 4), each thread independently creates algorithm instances, executes training, and calculates evaluation indicators to avoid resource competition; ② Deep learning algorithms: train in priority order (avoid GPU memory overflow), enable GPU acceleration (support CUDA), integrate early stop mechanism (patience value 10) and checkpoint saving function. Monitor loss, accuracy, recall rate, etc. during training, and record training history.

[0073] 4. Enhanced GRU model training: As the core innovative model, its training process is as follows: ① Input layer: receiving enhanced variable-length time series data (shape: batch x time x 100), processing variable-length sequences through mask mechanism; ② Network backbone: 1-2 layers of GRU stacked layers (hidden layer dimension 64) to extract time sequence dependence, global average pooling layer to obtain fixed-length vector (to reduce overfitting), fully connected classification head (activation function Sigmoid) to map to pattern probability space; ③ Loss function: FocalLoss (γ=2, α=0.25) is used to focus on difficult classification samples and alleviate class imbalance; ④ Optimizer: AdamW (learning rate 0.001, weight decay 0.01) is used, combined with Dropout (probability 0.2) regularization, training for 50 rounds, and early stopping mechanism is used to monitor the F1 score of the validation set.

[0074] Six, Sp6: Detailed implementation of intelligent model selection

[0075] This step adopts a four-stage selection strategy to determine the optimal model and realize unsafe behavior pattern recognition and early warning. The specific implementation details are as follows:

[0076] 1. First stage: quality screening stage: set performance thresholds suitable for imbalanced data: precision ≥ 0.25, recall ≥ 0.3, F1 score ≥ 0.3, precision-recall gap ≤ 0.5. All models that have completed training are subjected to threshold testing, and only models that pass all thresholds enter the next stage. For example, a LogisticRegression model with a recall rate of 0.28 (lower than the threshold of 0.3) is directly eliminated; six models such as Enhanced GRU and XGBoost pass the screening.

[0077] 2. Second stage: anomaly detection stage: focus on identifying over-prediction patterns, divided into two levels: ① Severe over-prediction (precision ≤ 0.5 and recall ≥ 0.95), such models completely lose their discrimination ability and are directly excluded; ② Moderate over-prediction (precision < 0.6, recall > 0.85, and gap > 0.3), also excluded. After detection, a LSTM_Attention model with a precision of 0.48 and a recall of 0.96 (severe over-prediction) is determined as an abnormal model and excluded.

[0078] 3. The third stage: business scoring stage: calculate the total score of the business based on the flight safety scene weight configuration, the weight is: accuracy 40%, recall rate 35%, F1 score 15%, training efficiency 10%. Scoring calculation method: ① Accuracy score = accuracy x 40%; ② Recall rate score = recall rate x 35%; ③ F1 score = F1 x 15%; ④ Training efficiency score = (1-current model training time / maximum training time) x 10%. For example, the Enhanced_GRU model has an accuracy of 0.72, a recall rate of 0.68, an F1 score of 0.70, and a training time of 80 minutes (maximum training time of 120 minutes), and its business total score is: 0.72 x 0.4 + 0.68 x 0.35 + 0.70 x 0.15 + (1-80 / 120) x 0.1 = 0.288 + 0.238 + 0.105 + 0.033 = 0.664. According to the business total score, generate a model adaptation degree ranking list.

[0079] 4. The fourth stage: final selection stage: under normal circumstances, select the model with the highest business total score as the optimal model, and generate detailed selection reasoning (including score details, performance indicators, selection advantages, etc.). For example, the Enhanced_GRU model has a business total score of 0.664 (ranked first), and the selection advantage is "high accuracy adaptation to flight safety scene, support for unbalanced data processing, strong time series modeling capability". If the candidate model pool is empty (all models do not pass the previous two stages), start the alternative mechanism, select the model with the highest F1 score as the emergency solution, and mark warning information (suggesting manual review). After the optimal model is output, it is used for real-time identification and early warning of pilot unsafe behavior patterns, with an identification delay of ≤2 seconds. Specific embodiment two:

[0081] As shown in Figures 1 to 5 , this embodiment comprehensively and in detail develops a "pilot unsafe behavior pattern recognition system based on facial expressions and physiological indicators", which includes six core modules, each module independently implements specific functions and works cooperatively to ensure efficient execution of the recognition method.

[0082] The data acquisition module is responsible for multi-source synchronous acquisition and preliminary aggregation, including a camera sub-module, a smartwatch sub-module, and a flight quality monitoring sub-module. The module uses an embedded processor (such as ARM Cortex-A53) as the core control unit to ensure low latency and high reliability in the acquisition process. The camera sub-module is equipped with a high-definition industrial camera (resolution at least 1920x1080, frame rate above 30fps) installed in the front of the cockpit to cover the entire area of the pilot's face, which captures real-time video and inputs into the expression analysis software. The software calculates the probability distribution of seven basic emotions (happy, sad, angry, surprised, scared, disgusted, and neutral) based on a deep learning model (average accuracy above 96%), and outputs the valence (emotional intensity, range 0-100), arousal (activation level, range 0-100), and time series data of emotional attitude indicators (such as interest, confusion, and boredom) updated once per second. The smartwatch sub-module uses a commercial-grade device that supports PPG technology, with a sampling frequency of 100Hz or higher, capable of continuously collecting heart rate time intervals (used to calculate variability indicators such as SDNN and RMSSD), heart rate mean values every 5 seconds (range 40-200 beats / min), body temperature every 10 seconds (accuracy 0.1°C, normal range 36.0-38.0°C), and blood oxygen saturation every minute (accuracy 1%, normal range 85%-100%), and data is uploaded to the module buffer through Bluetooth or Wi-Fi. The flight quality monitoring sub-module is integrated with the FOQA system of the aircraft, collecting more than 20 flight parameters (such as height change rate, speed, attitude, and ground load), recording once every 10Hz, and using the built-in pattern recognition algorithm to generate six unsafe mode labels (high slope descent, gradual change slope, specific height range anomaly, multi-level height change, insufficient stability after landing, and micro-stable anomaly on the ground). The labels are output in the form of 0-1 integer encoding. The three sub-modules interact through a timestamp synchronization mechanism (based on FOQA time, error controlled within ±50ms) to ensure multi-modal data alignment, such as camera data and physiological data calibrated every 30 seconds to avoid time sequence drift. The aggregated raw data set is stored in a structured format (such as JSON or CSV), including pilot identification, flight metadata, expression sequence, physiological sequence, and label information, which is directly transmitted to the input queue of the data preprocessing module, realizing zero-delay connection.

[0083] The data preprocessing module, as the system data entry, has multiple layers of intelligent filling units and standardization units. It uses multi-core CPUs (such as Intel i7 series) to process high-concurrency data streams, ensuring processing speeds of over 1000 records per second. The module first performs data screening and quality screening: the screening logic includes flight phase filtering (only retaining data marked in the landing phase, accounting for about 35-40% of the original data) and operator screening (extracting actual pilot records based on markers, excluding monitor data); quality screening deletes irrelevant columns (such as takeoff time and aircraft number), and evaluates each pilot data set, requiring at least 3 valid values for each feature and at least 50% of the features to be valid, with unqualified data automatically excluded (the pass rate is about 85%). The multi-layer intelligent filling unit constructs a filling strategy based on individual baselines (calculating the mean, median, standard deviation, and quartiles for each pilot) and group baselines (summarizing global statistics for all pilots), and selects associated derivation, multivariate sampling, or constraint fallback for different physiological characteristics: for example, heart rate variability is filled with individual median values first, and group median values are used when insufficient; heart rate statistics use mathematical association derivation (such as estimating the maximum value as the mean plus 2 times the standard deviation); blood oxygen saturation uses group 95-99% quantile random sampling; body temperature adds seasonal adjustment (subtract 0.2°C in winter, add 0.1°C in summer, subtract 0.1°C in autumn) and day-night adjustment (add 0.1°C during the day, subtract 0.2°C at night) based on individual mean values, and adds normal noise with mean 0 and standard deviation 0.1 to simulate natural fluctuations, with a missing rate of less than 5% after filling. The standardization unit uses the Z-score global method to stack all sample time steps into a matrix, then independently calculates the mean and standard deviation for each feature to normalize, ensuring dimensional consistency. The module also includes a time series reorganization unit that aggregates data by pilot-flight combination to form variable-length sequence samples (keeping the original length, without truncation or padding), with labels taken from the end-of-sequence state; the division unit uses a five-fold multi-label hierarchical K-fold algorithm to ensure consistent label proportions for each fold and avoid data leakage. After processing, the clean dataset is transmitted to the feature optimization engineering module buffer in NumPy array format, supporting batch parallel input.

[0084] The feature optimization engineering module contains an intelligent feature engineer and an intelligent feature selector, which uses GPU acceleration (such as NVIDIA RTX series) to handle high-dimensional calculations, ensuring that the expansion and screening time does not exceed 30 seconds. The input of this module is the pre-processed sequence dataset, and the intelligent feature engineer integrates time domain, frequency domain and nonlinear extraction techniques to expand the initial 23-dimensional original features to generate a pool of 251-dimensional candidate features: the frequency domain part uses Welch power spectral density estimation for heart rate variability, divides it into three segments of very low frequency (0.0033-0.04 Hz), low frequency (0.04-0.15 Hz) and high frequency (0.15-0.4 Hz), calculates the power value and derived indicators (such as low / high frequency ratio, normalized power), actively excludes total power to reduce redundancy, and a total of 114 dimensions; the statistical part uses an adaptive sliding window (size is the minimum of 1 / 5 and 10 of the sequence length, not less than 3), extracts mean stability (rolling mean standard deviation), standard deviation variability (rolling standard deviation standard deviation), trend intensity (Pearson correlation absolute value of sequence and time index) and volatility (first-order difference standard deviation), a total of 76 dimensions; the physiological part calculates the abnormal value proportion (time step ratio exceeding the normal range) and range utilization rate (signal span relative to the normal range ratio), and the normal range includes heart rate 40-200 beats / min, blood oxygen 85%-100%, body temperature 36.0-38.0°C, heart rate variability 20-200 ms, a total of 38 dimensions. The intelligent feature selector selects the optimal subset based on a multi-dimensional evaluation system: fusion of RandomForest Gini importance (construct 100 trees, depth 10, minimum leaf 5), correlation analysis (threshold 0.95 to mark redundant pairs, identify intra-group duplicates through hierarchical clustering), analysis of variance (coefficient of variation calculation, remove low-information features below the 1% percentile), statistical significance (F-test P<0.05 is significant, mutual information quantifies dependence), and physiological prior weight (original physiology 3.0, heart rate variability 2.5, frequency domain 1.8, etc.); normalized weighted sum (RandomForest 40%, variance 20%, mutual information 20%, physiological prior 20%), select the top 100-dimensional features according to the score, reduce the proportion by 40%, and verify the physiological range (such as heart rate feature 40-200 beats / min). The output feature subset is transmitted to the data augmentation module through memory mapping, realizing efficient connection.

[0085] The data augmentation module includes a balanced class enhancer and an extreme class enhancer, supports CUDA accelerated calculation, and the enhancement speed reaches 100,000 samples per minute. The input of the module is the optimal feature subset. First, the class imbalance degree is evaluated (positive sample ratio classification: super minority <5%, high minority 5%-10%, moderate minority 10%-25%, super extreme <5% and sample extremely few), and the enhancer is intelligently switched according to the severity. The balanced class enhancer is used as the main one, which is responsible for the mild strategy: integrated time distortion (scaling factor 0.8-1.2, adjusting sequence length by linear interpolation), amplitude distortion (0.7-1.3 times random scaling, simulating individual differences) and Gaussian noise injection (standard deviation 5% times, dynamically adjusting noise intensity); stratified configuration parameters, such as 8 times enhancement of super minority class to 25% target ratio, 5 times of high minority class to 25%, 2.5 times of moderate minority class to 35%, through a maximum of 5 iterations (randomly extracting basic samples in each round, applying 1-2 kinds of technology to generate new samples, and conserving coefficient 0.9 to control the number), and applying physiological constraints (heart rate variability change <30%, blood oxygen <7%, body temperature <3%) to ensure authenticity. The extreme class enhancer is used as the auxiliary one, which adopts aggressive strategy for super extreme scene: 20 times positive sample enhancement + 0.8 times negative sample downsampling to 40% target, integrating ADASYN adaptive sampling, boundary sample emphasis enhancement, frequency domain shift (shift parameter -0.1-0.1, sine modulation depth <5%) and synthetic sample generation (randomly selecting two samples of the same class, length uniformity, weighted mixing with 0.3-0.7 coefficient, and then grouping noise injection: heart rate variability 80% probability 8% standard deviation, heart rate 70% probability 5%); a maximum of 10 rounds of multi-generation iteration (large multiple generation in early stage, fine adjustment in later stage). The enhanced training set is transmitted to the model training module in tensor format, and the event trigger mechanism is used to monitor the enhancement completion state between modules.

[0086] The model training module integrates an algorithm pool, a dynamic algorithm selector and an algorithm factory, and a consumer-grade graphics card (such as RTX5090) can meet the calculation amount demand, ensuring that the training period does not exceed 2 hours. The input of the module is the enhanced training set. The dynamic algorithm selector first generates a data distribution portrait (balanced state four levels, complexity three levels, enhancement multiple), loads an algorithm characteristic file (5-dimensional features of 10 algorithms: type, imbalance processing rating, speed rating, memory rating, applicable state), calculates the matching score and sorts through multi-dimensional scoring (imbalance 40%, speed 30%, memory 30%) to form a core (3-4), exploration (1-4) and standby algorithm group.

[0087] The algorithm factory provides a unified interface: the registry maintains algorithm metadata (implementation class, type, default parameters), intelligently creates instances (merges user-defined parameters, supports GPU acceleration flags), exposes training (input features / labels, returns history dictionary), prediction (input test features, returns class), and probability prediction methods; traditional algorithms are wrapped in scikit-learn interfaces, and deep learning is wrapped in PyTorch logic. Training uses a hybrid strategy: traditional machine learning is multi-threaded parallel (thread pool size min (candidate number, 4)), deep learning is sequentially executed to avoid resource conflicts, and integrated performance monitoring, early stopping (patience 10), and checkpoint saving are used. The core Enhanced_GRU model implements SMOTE pre-processing (generates new samples to balance batches), GRU stacking (1-2 layers, 64-dimensional hidden extraction of temporal dependencies), global average pooling (fixed-length vector to reduce overfitting), fully connected head (Sigmoid output probability), Focal Loss optimization (focus on difficult samples), and AdamW (learning rate 0.001, weight decay 0.01, Dropout 0.2). The trained model and index are serialized and transmitted to the intelligent selection module.

[0088] The intelligent selection module includes quality screening, anomaly detection, business scoring, and final selection, uses a decision tree engine to process multi-model evaluation, and ensures that the selection time is <10 seconds. The module input is a pool of trained models, which first sets thresholds (precision ≥0.25, recall ≥0.3, F1 ≥0.3, gap ≤0.5) for quality screening to filter out unqualified models; anomaly detection identifies over-prediction (severe: precision ≤0.5 and recall ≥0.95; moderate: precision <0.6 and recall >0.85 and gap >0.3) to exclude false-positive risk models; business scoring calculates the total score based on flight safety weights (precision 40%, recall 35%, F1 15%, efficiency 10%, efficiency is one minus relative training time), sorts the ranking list according to the score, and selects the highest score model for final selection to generate an inference report (method identifier, detailed dictionary, performance dictionary, advantage list, health / normal identifier). When the pool is empty, the F1 highest is selected with a warning. The module simultaneously feeds back instructions to the data acquisition module to dynamically adjust parameters (such as increasing the camera frame rate to 60fps in high-risk stages or reducing low-priority physiological sampling to optimize power consumption) according to the identification results, realizes system adaptive closed-loop, and improves overall robustness. Embodiment Three

[0090] As Figures 1 to 5As shown, the embodiments of the present application are described in detail for the input and output relationship between the modules of the system and the data transmission path, and the core algorithm is described in detail, including its input data, output results and calculation process, and specific application in the system. The system forms a complete data processing link through the close cooperation between the modules, avoids functional isolation, and ensures the end-to-end efficient operation from raw collection to final warning.

[0091] The data transmission path of the system is mainly linear flow, supplemented by feedback mechanism to realize closed loop optimization. The output of the data acquisition module is the original data set, including facial expression sequence, physiological index sequence and flight parameter label, which is directly transmitted as input to the data preprocessing module, transmitted in real time through the internal communication interface to ensure no data loss. After receiving the original data set, the data preprocessing module performs screening, filling and standardization processing, and its output is a clean variable-length sequence training set and test set, which is transmitted to the feature optimization engineering module in array form to avoid intermediate storage overhead. The input of the feature optimization engineering module is the preprocessed sequence data set, and the output is a 100-dimensional optimal feature subset, which is directly injected into the input buffer of the data enhancement module through the memory sharing mechanism. The data enhancement module receives the optimal feature subset as input to generate an enhanced balanced training set, which is transmitted to the algorithm pool entrance of the model training module in tensor format. The output of the model training module is a plurality of trained model instances and performance indicators, which are transmitted to the candidate pool of the intelligent selection module through serialization. The final output of the intelligent selection module is the optimal model and its inference explanation, which is deployed to the system warning engine, and at the same time the recognition result is fed back to the data acquisition module through the feedback channel for dynamic adjustment of the acquisition parameters, such as increasing the sampling rate in high-risk mode detection frequency, or reducing the physiological index sampling accuracy in low imbalance scene to save resources. The path ensures that the delay of data from acquisition to warning is not more than 5 seconds, supporting real-time flight monitoring.

[0092] In terms of algorithms, the core algorithms of the system include dynamic algorithm selection algorithm, Enhanced GRU model, RandomForest algorithm, XGBoost algorithm, LightGBM algorithm, SVM algorithm, Logistic Regression algorithm, Gradient Boosting algorithm, GRU algorithm, LSTM Attention algorithm and CNN-LSTM algorithm, which are distributed in the algorithm pool of the model training module for processing unsafe behavior pattern recognition under different data distribution. The input data, output results, calculation process and specific application of each algorithm in the system are described in detail as follows.

[0093] The input of the dynamic algorithm selection algorithm is the distribution portrait of the enhanced data, including the class balance state (determined by calculating the number of positive samples divided by the total number of samples, and divided into four levels of balance, mild imbalance, moderate imbalance and severe imbalance), data complexity (evaluated by the sample size multiplied by the feature dimension, and divided into three levels of low, medium and high), and the enhancement multiple (obtained by dividing the number of positive samples after enhancement by the original number of positive samples). The calculation process first loads the algorithm characteristic profile matrix, which records the type, imbalance handling ability rating (excellent, good, general), training speed rating (very fast, fast, medium, slow, very slow), memory occupation rating (low, medium, high, very high) and applicable balance state of each algorithm. Then, the matching score of each algorithm is calculated, including the balance matching score (if the algorithm's applicable state covers the input balance state, add points, otherwise subtract points) and the complexity matching score (deep learning algorithm is preferred for high complexity data, with weighted adjustment); subsequently, a multi-dimensional scoring function is applied, and the imbalance handling ability, training speed and memory occupation are respectively weighted and summed to obtain the comprehensive priority score of each algorithm; finally, the scores are sorted in descending order to form a candidate list, including the core algorithm group (selecting the top 3-4 highest score algorithms), the exploration algorithm group (selecting the middle 1-4 potential algorithms) and the standby algorithm group (selecting the remaining algorithms with strong explanatory power). The output is the sorted candidate algorithm list, which is applied in the system to optimize the execution order of the model training module, for example, in severe imbalance data, Enhanced_GRU is placed at the top to ensure that the algorithm that adapts to imbalance is trained first, reducing invalid calculations and improving training efficiency by more than 20%.

[0094] The input of the Enhanced GRU model is the enhanced variable-length sequence data. First, the imbalance is handled by pre-SMOTE processing, the nearest neighbors between minority class samples are calculated, and new samples are generated to balance the class distribution while ensuring reasonable feature values under physiological constraints. The calculation process is divided into network forward propagation and optimization: the input sequence enters the GRU stacking layer, which controls the degree of new information integration through the update gate, decides to forget the historical state through the reset gate, and fuses the current input and historical information through the candidate hidden state, finally generates a hidden state sequence; the global average pooling layer averages all hidden states along the time dimension to obtain a fixed-length vector to reduce the risk of overfitting; the fully connected classification head maps the vector to the mode probability space through linear transformation and activation function. The loss function focuses on difficult-to-classify samples through weighting, reduces the influence of easy-to-classify samples, and alleviates the class imbalance problem. The optimization process uses the AdamW optimizer to gradually adjust the model parameters, combines Dropout to randomly discard part of the units in the hidden layer to prevent overfitting, and monitors the performance of the validation set to achieve early stopping. The output is the unsafe mode probability of each input sequence. In the system, it is used as the core component of the early warning engine, which receives real-time flight data, calculates the probability threshold to trigger an alarm, such as activating a voice prompt when the probability exceeds 0.5, to help pilots correct their behavior in a timely manner.

[0095] The input of the RandomForest algorithm is tabular feature data, including the optimal feature subset and behavior mode label. The calculation process first builds multiple decision trees, each tree randomly samples a subset of data and a subset of features, recursively splits nodes until purity maximization, and selects the best split point through the Gini impurity index; then, the results of all trees are aggregated by majority voting to generate the final classification decision. The output is the mode prediction probability and feature importance ranking. In the system, it is used as the preferred algorithm for balancing data scenarios, which is used to quickly identify low-complexity patterns, such as evaluating feature contributions in the training module to help optimize subsequent iterations.

[0096] The input of the XGBoost algorithm is the feature matrix and label vector. The calculation process iteratively builds decision trees through the gradient boosting framework, calculates the residual as the target of the new tree in each round, applies a learning rate to scale the tree contribution, controls the tree complexity with built-in regularization terms, and supports built-in imbalance handling such as weight adjustment. The output is the boosted prediction score. In the system, it is used to handle moderately imbalanced data and provide high-accuracy predictions, such as quickly detecting grounding stability anomalies in conjunction with physiological indicators in the early warning system.

[0097] The input of the LightGBM algorithm is similar to XGBoost, but the calculation process uses histogram approximation to accelerate split finding, leaf growth strategy to reduce computational overhead, and supports class weight adjustment for imbalance. The output is an efficient prediction result. In the system, it is used as a backup in memory-constrained scenarios, providing fast training for real-time feedback adjustments.

[0098] The input of SVM algorithm is high-dimensional feature space. The calculation process maps data to high dimension through kernel function, finds the maximum interval hyperplane, optimizes support vector to minimize classification error, and supports multiple kernels such as Gaussian kernel to handle nonlinearity. The output is the classification boundary decision, and the application in the system is the classification of high-dimensional physiological data, which ensures stable prediction such as blood oxygen abnormality detection.

[0099] The input of Logistic Regression algorithm is linear feature combination. The calculation process maps linear combination to probability through logistic function, optimizes coefficients through gradient descent, and adds regularization to avoid overfitting. The output is probability estimation, and the application in the system is an explanatory benchmark model used to verify the performance of other algorithms in simple patterns.

[0100] The input of Gradient Boosting algorithm is feature and residual sequence. The calculation process is similar to XGBoost, but pays more attention to robustness by sequentially adding weak learners to correct the previous round error. The output is cumulative prediction, and the application in the system is progressive optimization of complex time series patterns, such as multi-level height change recognition.

[0101] The input of GRU algorithm is time series sequence. The calculation process updates the hidden state through the gating mechanism, simplifies LSTM to improve efficiency. The output is time series encoding, and the application in the system is basic time series modeling, which assists Enhanced_GRU in pre-training on light imbalance data.

[0102] The input of LSTM_Attention algorithm is long sequence data. The calculation process combines LSTM memory unit and attention mechanism to weight key time steps. The output is focus point encoding, and the application in the system is to capture long dependency patterns, such as progressive change slope detection.

[0103] The input of CNN-LSTM algorithm is spatio-temporal sequence. The calculation process first extracts spatial features with convolution, and then processes time dependence with LSTM. The output is hybrid mode prediction, and the application in the system is to process physiological flight data that integrates "space + time", and to improve the accuracy of post-landing stability recognition. Specific embodiment four:

[0105] As Figures 1 to 5 shown, the following provides specific use cases:

[0106] 1. Application in pilot training simulator:

[0107] The system is deployed in the simulator cabin of an aviation training center and interfaces with the simulator software. During training, a new pilot performs a landing simulation task, and the system captures facial expressions through cockpit cameras, collects heart rate variability, body temperature, and blood oxygen data through a smartwatch, and the simulator provides virtual flight parameters such as altitude changes and speed. The data collection module aggregates the raw data set in real time, the preprocessing module filters the landing phase data and fills in the missing heart rate values (using the individual baseline median), and the feature optimization module extracts 100-dimensional features such as frequency domain power and abnormality proportion. The model training module dynamically selects XGBoost and Enhanced_GRU as the core algorithm for parallel prediction; the intelligent selection module selects the Enhanced_GRU model (total score 0.68) through four-stage evaluation. In the simulation, the system identifies an abnormal increase in pilot facial arousal combined with heart rate variability fluctuations, predicts a "high slope descent" probability of 0.75, and immediately triggers the simulator warning light and voice prompt "adjust the descent rate." After training, the system generates a report showing a recall rate of 65%, helping instructors analyze pilot behavior pattern deviations under stress and optimize training plans, increasing the pass rate of new pilots by 15%.

[0108] 2. Real-time monitoring in long-haul commercial flights:

[0109] The system is integrated into the cockpit electronics of a passenger aircraft and is networked with the onboard computer. After takeoff on a transatlantic flight, the system starts continuous monitoring of the facial expressions and physiological indicators of the two pilots, while simultaneously collecting parameters from the flight quality monitoring system such as attitude and ground stability. The data collection module ensures synchronization of the three types of data, and the preprocessing module reorganizes them into variable-length sequences and standardizes the features (such as global Z-score processing of blood oxygen saturation). The feature optimization module expands statistical features to detect trend intensity anomalies. The model training module prioritizes LightGBM and LSTM_Attention based on data portraits, and after prediction, the intelligent selection module excludes over-predicting models and selects LSTM_Attention (precision weight dominant, total score 0.71). As the flight approaches landing, the system analyzes a decrease in valence and an increase in heart rate standard deviation for one of the pilots, combined with flight parameters to predict a "ground micro-stability anomaly" probability of 0.62, and sends a real-time alert to the head-mounted device, prompting "check the ground attitude." The pilot adjusts the operation accordingly, avoiding potential risks of ground overload, with a delay of only 1.5 seconds. Post-event logs show that the system reduces false alarms by 5%, improves crew alertness, and ensures safe landing of the flight.

[0110] 3. High-intensity early warning in military flight missions:

[0111] The system is installed in the cockpit of a fighter plane, adapting to high G forces and complex environments. During a mission, the system collects facial data through an impact-resistant camera embedded in the helmet, monitors physiological indicators with a military smartwatch, and provides real-time flight parameters such as multi-level altitude changes from a mission computer. The data collection module aggregates the dataset, and the preprocessing module applies physiological constraints to fill in missing temperature values (with an overlay for day-night adjustments). The feature optimization module focuses on extracting high-frequency power features to capture stress responses. The model training module intelligently selects Enhanced_GRU and CNN-LSTM as the core, performing sequential prediction to manage resources; the intelligent selection module excludes models with high recall rates through anomaly detection, selecting CNN-LSTM (F1 score dominant, total score 0.65). During the mission, the system detects the pilot's angry expression combined with a decrease in blood oxygen, predicting a "specific altitude range anomaly" probability of 0.81, activating a vibration alarm and displaying "stable altitude" on the HUD. The pilot corrects in time, avoiding potential collisions. This application improves mission success rates and reduces losses caused by human errors.

[0112] It should be noted that the relational terms herein, such as first and second, and the like, are used solely to distinguish one from another entity or action without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0113] While embodiments of the present application have been shown and described with reference to particular embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application. The scope of the application is defined by the appended claims and their equivalents.

Claims

1. A method for identifying unsafe behavior based on pilot facial expression and physiological indicators, characterized in that, Comprise the following steps: Sp1: multi-dimensional data acquisition, through the camera and expression calculation software to collect pilot facial expression data, through wearable devices to collect pilot physiological index data, through flight quality monitoring system to collect flight parameter data and generate pilot behavior mode label, the pilot facial expression data, pilot physiological index data and flight parameter data three kinds of data are summarized to form the original data set; the facial expression data includes time series data of valence and arousal, the physiological index data includes time domain feature data of heart rate variability, heart rate, blood oxygen saturation and body temperature; the camera and expression calculation software, wearable devices are interactive with flight quality monitoring system data, ensure the time synchronization of the collected data; Sp2: original data preprocessing, screening, quality screening and reorganization, using multi-level intelligent filling strategy to process missing values, standardizing the features, reorganizing the data into variable length time series sample data set, and dividing the training set and test set by hierarchical cross-validation; the multi-level intelligent filling strategy is based on individual baseline and group baseline, combined with the physiological meaning of the features, using correlation calculation, multivariate derivation, baseline quantile sampling and physiological constraint bottom-up scheme to realize missing feature filling; the Sp2 receives the original data set summarized in Sp1, and the processed data is transmitted to Sp3; Sp3: feature optimization, using a two-stage progressive optimization architecture, the first stage expands the candidate feature pool through frequency domain feature extraction, statistical feature extraction and physiological feature extraction, the second stage selects intelligent features based on a multi-dimensional feature importance evaluation system to obtain the optimal feature subset; the multi-dimensional feature importance evaluation system integrates RandomForest importance evaluation, correlation analysis, variance analysis, statistical significance test and physiological prior knowledge dimension; Sp3 interacts with Sp2 data, receives the preprocessed data and outputs the optimal feature subset to Sp4; Sp4: differential category balance enhancement, based on the category imbalance degree of each flight mode label, intelligent selection of balance category enhancer or extreme category balancer for data enhancement of training set; the balance category enhancer adopts a mild balance strategy, integrating time distortion, amplitude distortion and Gaussian noise injection technology and applying physiological constraints; the extreme category enhancer adopts an aggressive balance strategy, integrating synthetic sample generation and frequency domain shift technology, with negative sample downsampling mechanism; Sp4 receives the optimal feature subset output by Sp3, and the enhanced training set is transmitted to Sp5; Sp5: model training, construct an algorithm pool containing traditional machine learning algorithms and deep learning algorithms, dynamically select candidate algorithm combinations based on the distribution characteristics of the enhanced data, and complete model training using parallel competitive training mechanism; the deep learning algorithm includes Enhanced_GRU model, which integrates FocalLoss and SMOTE technology and supports variable length time series processing; Sp5 interacts with Sp4 data, receives the enhanced training set, and transmits the trained model to Sp6; Sp6: intelligent model selection, adopting a four-stage selection strategy to determine the optimal model, realizing the identification and early warning of the unsafe behavior pattern of the pilot; the four-stage selection strategy is quality screening stage, anomaly detection stage, business scoring stage and final selection stage in turn, the business scoring stage adopts preset weight configuration to calculate the comprehensive business score; Sp6 receives the trained model output by Sp5, and outputs the optimal model after selection for unsafe behavior pattern identification and early warning.

2. The pilot expression and physiological index based unsafe behavior recognition method according to claim 1, characterized in that, The screening of the original data set in Sp2 includes flight phase screening and operator screening, and data records marked as landing phase and actually executed by the pilot are retained; the evaluation standard of the quality screening is that each feature has at least 3 valid values, and at least 50% of the features are in the valid state. 3.The pilot expression and physiological index based unsafe behavior recognition method according to claim 1, characterized in that, The process of reorganizing the data into a variable-length time sequence sample data set in Sp2 adopts a variable-length sequence storage method to maintain the original time length of each flight sample; the stratified cross-validation adopts a five-fold multi-label stratified K-fold algorithm to ensure that the proportion of various label combinations in each fold is consistent with the overall data set.

4. The pilot expression and physiological index based unsafe behavior recognition method according to claim 1, characterized in that, The first-stage frequency domain feature extraction in Sp3 adopts Welch power spectral density estimation method for heart rate variability signal to extract power values and derived indicators in very low frequency band, low frequency band and high frequency band, and actively excludes the total power frequency band to reduce redundancy; the first-stage physiological feature extraction calculates the abnormal value proportion and range utilization of the physiological feature relative to the normal range.

5. The pilot expression and physiological index based unsafe behavior recognition method according to claim 1, characterized in that, The category imbalance degree division in Sp4 is divided into four levels, namely super-minority class, high-minority class, moderate-minority class and super-extreme-minority class; the balanced class enhancer configures different enhancement multiples and technical combinations for super-minority class, high-minority class and moderate-minority class; the extreme class enhancer configures aggressive enhancement parameters for super-extreme-minority class, and the synthesized sample generation is realized by mixing two real samples of the same class and injecting noise.

6. The pilot expression and physiological index based unsafe behavior recognition method according to claim 1, characterized in that, The process of dynamically selecting candidate algorithm combination based on the distribution characteristics of the enhanced data in Sp5 is to generate data distribution portraits, construct algorithm characteristic archives, implement stratified algorithm selection strategy based on data balance state, data complexity and enhancement multiple, and calculate algorithm comprehensive priority through multi-dimensional scoring function.

7. The pilot expression and physiological index based unsafe behavior recognition method according to claim 1, characterized in that, In Sp6, the quality screening stage of the four-stage selection strategy sets a performance threshold suitable for unbalanced data; the anomaly detection stage of the four-stage selection strategy identifies the over-prediction mode of the algorithm; the weight configuration of the business scoring stage of the four-stage selection strategy is that the accuracy weight is 40%, the recall rate weight is 35%, the F1 score weight is 15%, and the efficiency weight is 10%; the final selection stage of the four-stage selection strategy selects the model with the highest comprehensive score, and the standby selection mechanism is started when the candidate model pool is empty.

8. A system corresponding to the pilot expression and physiological index based unsafe behavior recognition method according to any one of claims 1-7, characterized in that, The system comprises a data acquisition module, a data preprocessing module, a feature optimization engineering module, a data enhancement module, a model training module and an intelligent selection module, and is used for executing the method. The data acquisition module is used for collecting facial expression data, physiological index data and flight parameter data of the pilot and generating behavior mode labels, and the original data set is formed by summarizing; the data acquisition module includes a camera and expression calculation software, a wearable device and a flight quality monitoring submodule; the camera and expression calculation software, the wearable device and the flight quality monitoring submodule are in data interaction, ensuring the time synchronization of the collected data; the data acquisition module is in communication connection with the data preprocessing module, and the original data set is transmitted to the data preprocessing module; The data preprocessing module is used for screening, quality screening and reorganization of the original data set, processing missing values, standardizing the features, reorganizing the data and dividing the training set and the test set; the data preprocessing module is built-in multi-level intelligent filling unit, which constructs individual baseline and group baseline, and selects corresponding filling scheme combined with the physiological meaning of the features; The data preprocessing module and the feature optimization engineering module are in data interaction, and the processed data is transmitted to the feature optimization engineering module; The feature optimization engineering module is used for realizing feature expansion and intelligent feature selection, and obtaining an optimal feature subset; the feature optimization engineering module includes an intelligent feature engineer and an intelligent feature selector, the intelligent feature engineer integrates frequency domain, statistical and physiological feature extraction technologies, and the intelligent feature selector adopts a multi-dimensional feature importance evaluation system; The feature optimization engineering module and the data enhancement module are in communication connection, and the optimal feature subset is transmitted to the data enhancement module; The data enhancement module is used for solving the class imbalance problem of the training data, and includes a balanced class enhancer and an extreme class balancer; the balanced class enhancer adopts a mild balancing strategy and applies physiological constraints, and the extreme class balancer adopts an aggressive balancing strategy and supports synthetic sample generation and negative sample downsampling; the data enhancement module and the model training module are in data interaction, and the enhanced training set is transmitted to the model training module; The model training module is used for realizing multi-algorithm parallel competition training, and includes an algorithm pool, a dynamic algorithm selector and an algorithm factory; the algorithm pool includes traditional machine learning algorithms and deep learning algorithms, the dynamic algorithm selector selects candidate algorithm combinations based on data distribution characteristics, and the algorithm factory provides a unified algorithm instance creation and calling interface; the model training module and the intelligent selection module are in communication connection, and the trained model is transmitted to the intelligent selection module; The intelligent selection module is used for determining an optimal model through a four-stage selection strategy, and realizing identification and early warning of the pilot's unsafe behavior mode; the intelligent selection module includes a quality screening unit, an anomaly detection unit, a business scoring unit and a final selection unit; the intelligent selection module and the data acquisition module keep instruction transmission, and the data acquisition parameters can be adjusted according to the identification result.

9. The system corresponding to the pilot expression and physiological index based unsafe behavior identification method according to claim 8, characterized in that, The wearable device included in the data acquisition module is a smart watch, which supports collecting PPG-based heart rate time interval, heart rate mean value, body temperature and blood oxygen value data; the expression calculation software included in the data acquisition module is a facial expression analysis system, which can analyze seven basic emotions, valence and arousal degree indexes.

10. The system corresponding to the pilot expression and physiological index based unsafe behavior identification method according to claim 8, characterized in that, The deep learning algorithm in the algorithm pool contained in the model training module includes Enhanced_GRU, GRU, LSTM_Attention and CNN-LSTM, and the traditional machine learning algorithm includes RandomForest, XGBoost, LightGBM, SVM, LogisticRegression and GradientBoosting; the Enhanced_GRU model integrates FocalLoss and SMOTE technology, and includes a GRU stacking layer, a global average pooling layer and a fully connected classification head.

Citation Information

Patent Citations

  • Pilot abnormal state monitoring and early warning system and method

    CN117854229A

  • Pilot airworthiness state sensing method based on multi-task rotation pairing learning

    CN119832593A

  • Pilot fatigue detection and alert technology

    US20250313346A1