Machine learning based aircraft ice accretion prediction model training and evaluation test method

CN122451472BActive Publication Date: 2026-09-25NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610930898.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25
Estimated Expiration
2046-06-26

AI Technical Summary

Technical Problem

然而,此类指数存在明显局限性:一方面,经验阈值在不同季节、不同地域、不同天气系统下的适用性差异大,泛化能力不足;另一方面,这类方法无法充分利用数值模式输出的丰富物理量信息,对积冰等级的区分能力较弱,尤其在严重积冰的判定上准确率偏低,难以满足现代航空精细化预警的需求

Benefits of technology

[0043](1)本发明将遗传算法引入飞机积冰预测的集成学习框架,以类别平衡验证集上的宏平均F1分数为适应度准则,自动搜索最优异构基模型组合,避免了人工选择基模型的盲目性和主观性,同时保证了所选模型子集在全部积冰类别上的综合识别能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451472B_ABST
    Figure CN122451472B_ABST
Patent Text Reader

Abstract

The application discloses a method for training and evaluating an aircraft icing prediction model based on machine learning, in the training method, first, original data is acquired and physical derived features are constructed, the optimal feature subset is screened out by using a random forest recursive feature elimination method, then a heterogeneous base model pool is constructed, and a verification set is divided from the optimal feature subset. Subsequently, a heuristic search is used to automatically optimize the base model combination, and the macro average F1 score obtained by the candidate subset on the verification set is used as the fitness function for optimization; the optimal machine model subset obtained by optimization is used to generate a meta-feature matrix through stratified K-fold cross-validation on the verification set; and the meta-feature matrix and the corresponding real icing grade label are used to train a meta-classifier, so that intelligent fusion of base model output is realized. The application effectively overcomes the class imbalance problem of icing samples, fully utilizes the complementarity between heterogeneous models, and significantly improves the recognition accuracy and recall rate of moderate and severe icing grades.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of aviation meteorology and artificial intelligence, specifically relating to a method for training and evaluating an aircraft icing prediction model based on machine learning. Background Technology

[0002] Aircraft icing is a major weather phenomenon that threatens flight safety. When an aircraft passes through clouds containing supercooled water droplets, ice may form on its fuselage surface, leading to decreased lift, increased drag, and deteriorated handling performance. Therefore, accurately predicting the level of aircraft icing is crucial for ensuring flight safety and optimizing flight path planning.

[0003] Current aircraft icing prediction technologies primarily rely on traditional meteorological indices, such as the IC icing index and its modified forms. These methods typically depend on only a limited number of meteorological factors, such as temperature and humidity, and use empirical thresholds to determine the likelihood and intensity of icing. However, these indices have significant limitations: firstly, the applicability of empirical thresholds varies greatly across different seasons, regions, and weather systems, resulting in insufficient generalization ability; secondly, these methods cannot fully utilize the rich physical quantity information output by numerical models, exhibiting weak ability to differentiate icing levels, particularly in determining severe icing, where accuracy is low and fails to meet the demands of modern, refined aviation early warning systems.

[0004] In recent years, with the development of machine learning technology, some studies have attempted to apply single classifiers (such as decision trees and support vector machines) to aircraft icing prediction. However, such methods face the following prominent problems: First, actual icing samples naturally exhibit class imbalance. Severe icing events are rare but extremely harmful, while non-icing samples constitute the majority, resulting in imbalance in the data for each of the four categories. Conventional classifiers tend to favor the majority class, leading to a very high false negative rate for severe icing levels. Second, single algorithm models show significant differences in performance across different icing levels, making it difficult to achieve ideal results simultaneously across all categories. Third, existing ensemble methods often use simple voting or fixed-weight average probabilities for fusion, failing to fully consider the reliability differences of each base model under different meteorological conditions. The fusion strategy lacks data-driven impetus, limiting further improvement in overall prediction performance.

[0005] In summary, there is an urgent need for an intelligent prediction method for aircraft icing levels using small-sample, multi-class imbalanced data. This method should be able to overcome class imbalance, flexibly integrate the complementary advantages of multiple heterogeneous machine learning models, and learn the optimal fusion strategy in a data-driven manner, thereby improving the prediction accuracy for each category, especially severe icing. Summary of the Invention

[0006] To address the aforementioned problems, this invention discloses a machine learning-based method for predicting and evaluating aircraft icing, which overcomes the shortcomings of traditional methods in terms of class imbalance and model fusion, and improves the prediction accuracy for various categories, especially severe icing.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A training method for an aircraft icing level prediction model based on a machine learning ensemble model includes the following steps:

[0009] Step 1: Obtain a historical dataset containing multiple meteorological elements and aircraft icing level labels, construct physical derived features, and use the random forest recursive feature elimination method to iteratively select the optimal feature subset from all features;

[0010] Step 2: Construct a base model pool containing multiple machine learning algorithms based on different principles, and extract equal amounts of samples from the optimal feature subset according to the ice accumulation level category to form a fixed test set and isolate it for subsequent evaluation of the model's generalization performance. The remaining samples constitute the training and validation set.

[0011] Step 3: On the training and validation set, a heuristic search algorithm is used to search for a subset of base models to be optimized from the base model pool; wherein, the heuristic search algorithm uses the macro-average F1 score obtained by the subset of base models to be optimized on the training and validation set as the fitness function for selection, and outputs the optimal subset of base models obtained by the search.

[0012] Step 4: Based on the optimal machine model subset, input the training and validation set into the hierarchical K-fold cross-validation module, generate corresponding prediction results using each base model, and summarize the prediction results of all base models to construct a meta-feature matrix; train a meta-classifier using the meta-feature matrix and its corresponding true ice accumulation level labels.

[0013] The optimal base model subset and the meta-classifier together constitute the aircraft icing level prediction model.

[0014] Furthermore, in step 1, the physical derived features include one or more of the following: the total wind speed synthesized from zonal and meridional winds, the humidity difference and temperature difference at different altitude levels, the specific humidity ratio and gradient, and the ratio of liquid water path to precipitable water.

[0015] Furthermore, the optimal feature subset includes: cloud-water mixing ratio at the icing point (Qc), rainwater mixing ratio at the icing point (Qr), 700 hPa temperature (T700), overall specific humidity (Q), temperature-dew point difference (T-Td), relative humidity at the icing point (RH), dew point temperature at the icing point (Td), icing point temperature (T), zero-degree layer height (Height_0°), 0-3 km vertical wind shear (SH03), ratio of precipitable water to liquid water path (PWAT / LWP), overall liquid water path (LWP), 500 hPa temperature (T500), air pressure (P), total wind speed (Wind), and ratio of specific humidity at 600 hPa to 500 hPa (Q600 / 500).

[0016] Furthermore, in step 2, the base model pool includes at least three of the following: logistic regression, support vector machine, decision tree, random forest, XGBoost, LightGBM, and K-nearest neighbor algorithm.

[0017] Furthermore, in step 3, the heuristic search algorithm is a genetic algorithm.

[0018] Furthermore, in step 3, the fitness function—the macro-average F1 score—is... The calculation formula is:

[0019]

[0020] in, The number of folds for cross-validation. The total number of categories for ice accumulation levels. and Let be the precision and recall of class c on the i-th validation set, respectively.

[0021] Furthermore, in step 3, when calculating the fitness, and in each fold cross-validation of generating the meta-feature matrix in step 4, the EasyEnsemble undersampling strategy is used on the training data to balance the number of samples in each class.

[0022] Furthermore, in step 4, the classifier is a logistic regression model, and the class weight parameter is set to a balanced mode during training, with the weight of each class being inversely proportional to the number of samples of that class in the training data.

[0023] This invention also discloses a machine learning-based aircraft icing prediction model evaluation and verification method, which uses the above-mentioned aircraft icing level prediction model for prediction; during model performance evaluation, multiple independent evaluation loops are executed, and the above-mentioned training method is executed independently in each loop to obtain a prediction model for this loop.

[0024] The prediction model used in this iteration is used to predict the fixed test set, and the evaluation index is calculated.

[0025] The average value of each evaluation index obtained from multiple iterations is used as the final performance verification basis for the aircraft icing level prediction model.

[0026] Furthermore, the evaluation metrics include at least one of accuracy, precision for each category, recall, F1 score, and critical success index (CSI).

[0027] Furthermore, accuracy for:

[0028]

[0029] In the formula, TP is the total number of samples correctly predicted for all classes, and TN is the total number of negative samples correctly predicted for all classes;

[0030] Recall rates for various categories c for:

[0031]

[0032] In the formula, The number of samples in category c that were correctly predicted. This represents the number of samples that were missed in category c. The macro average of the recall rates for each category is... C represents the total number of ice accumulation categories;

[0033] Precision of various categories c for:

[0034]

[0035] In the formula, The number of samples in category c that were correctly predicted. This represents the number of samples that were misclassified as category c from other categories.

[0036] F1 scores by category c for:

[0037]

[0038] In the formula, For category c, Let C be the recall rate for category c;

[0039] Critical Success Indices for Various Categories for:

[0040]

[0041] In the formula, The number of samples in category c that were correctly predicted. This represents the number of samples from other categories that were misclassified as category c. This represents the number of samples that were missed in category c.

[0042] Beneficial effects:

[0043] (1) This invention introduces genetic algorithm into the integrated learning framework for aircraft icing prediction. It uses the macro-average F1 score on the class balance validation set as the fitness criterion to automatically search for the best combination of base models, avoiding the blindness and subjectivity of manual selection of base models, while ensuring the comprehensive recognition ability of the selected model subset on all icing categories.

[0044] (2) This invention adopts a meta-model stacking and fusion strategy based on hierarchical K-fold cross-validation “out-of-fold” prediction. It learns the mapping relationship between the prediction probability of each base model and the final icing level in a data-driven manner. Compared with the traditional hard voting or fixed weight average probability method, it can more finely characterize the reliability differences of each base model under different meteorological conditions and significantly improve the accuracy of fusion prediction. At the same time, the whole process integrates the EasyEnsemble undersampling mechanism, which effectively overcomes the class imbalance problem of scarce severe icing samples.

[0045] (3) The present invention combines the evaluation and testing method of dividing a fixed test set with multiple different random seeds and taking the average, which eliminates the impact of the randomness of a single data division on the performance evaluation, and the evaluation results are more objective and robust. According to the actual icing case data, the method of the present invention has significantly improved the accuracy of icing level classification, macro average F1 score and critical success index compared with the traditional IC index and single machine learning model. It solves the problems of high false negative rate and insufficient prediction accuracy of existing methods for severe icing levels, and can provide more reliable decision support for aviation flight safety early warning. Attached Figure Description

[0046] Figure 1 This is a flowchart of an embodiment of the present invention;

[0047] Figure 2 This is a graph showing the recursive feature elimination curve of a random forest according to an embodiment of the present invention.

[0048] Figure 3 This is a feature importance ranking diagram according to an embodiment of the present invention;

[0049] Figure 4 This is a bar chart comparing the performance of the method of this invention with four traditional physical index methods in the task of determining the degree of ice accumulation.

[0050] Figures 5(a) to 5(d) are heatmaps comparing the performance evaluation indicators of the method of the present invention and the traditional physical index method on the four-class ice accumulation level identification task; where Figures 5(a) to 5(d) correspond to the evaluation results of precision, recall, F1 score and critical success index (CSI) respectively.

[0051] Figures 6(a) to 6(e) are comparison diagrams of the confusion matrices of the method of the present invention and the traditional physical index method on the four-category ice accumulation level discrimination task; wherein Figures 6(a) to 6(d) correspond to the confusion matrices of four traditional physical index methods, namely IC index, improved IC index, TF method and LWC method, respectively, and Figure 6(e) corresponds to the confusion matrix of the GA-Stacking Ensemble method of the present invention. Detailed Implementation

[0052] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0053] Example 1

[0054] like Figures 1-4 As shown in the figure, this embodiment discloses a training method for an aircraft icing level prediction model based on a machine learning ensemble model, including the following steps:

[0055] Step 1: Data preprocessing and feature engineering based on random forest iteration.

[0056] Step 101: Collect historical aircraft icing and non-icing datasets. In this embodiment, aircraft icing observation reports and corresponding spatiotemporal meteorological field data are collected. The aircraft icing observation reports include the actual time, location, and intensity level of icing encountered by the aircraft. The meteorological field data is dynamically downscaled using an ERA5-driven WRF model to obtain high-resolution meteorological fields (such as temperature, humidity, wind, and air pressure). The downscaled meteorological variables are then matched with the aircraft icing observation reports in both time and spatial grid dimensions to construct a meteorological element dataset corresponding to each icing sample. Each sample represents an icing (or non-icing) event and its corresponding meteorological environment. In this embodiment, icing levels are divided into four categories: no icing (category 0), light icing (category 1), moderate icing (category 2), and severe icing (category 3). The sample counts for each category are: 771 for no icing, 40 for light icing, 129 for moderate icing, and 83 for severe icing, totaling 1023 samples. The data exhibits a significant imbalance.

[0057] Step 102: Construct meteorological and physical derived features and perform missing value processing and standardization.

[0058] In this embodiment, not only are the original variables directly output by the numerical weather prediction model (such as temperature T, relative humidity RH, zonal wind U / meridian wind V, air pressure P, and specific humidity Q) used, but new composite features (i.e. meteorological physical derived features) are also constructed based on the principles of cloud microphysics and atmospheric thermodynamics.

[0059] Meteorological and physical derivative characteristics include:

[0060] Dynamic characteristics: Total wind speed composed of zonal and meridional winds, and vertical wind shear of 0-3 km;

[0061] Thermal and humidity characteristics: humidity and temperature differences between different altitude layers, specific humidity ratio of each layer, and vertical gradient;

[0062] Cloud microphysical characteristics: the ratio of liquid water path to precipitable water, etc.

[0063] The meteorological and physical derivative features constructed in this embodiment can more directly describe the specific conditions required for ice formation (such as the presence of supercooled water, uplift motion, water vapor supply, etc.) than single variables, providing the model with more physically meaningful and discriminative information.

[0064] Step 103, data preprocessing, including missing value handling and standardization. Missing values ​​in the data are imputed; in this embodiment, the mean is used for imputation. All features are standardized using Z-scores, which involves subtracting the mean and dividing by the standard deviation. This brings features of different dimensions and magnitudes to the same numerical scale.

[0065] Step 104: Train a random forest model based on the preprocessed feature set (including original variables and derived features). The random forest model outputs a feature importance score, which is based on the contribution of each feature to reducing model impurity when used to split nodes across all decision trees in the forest. All features are ranked according to their importance scores. In each iteration, the feature with the lowest importance ranking (or a few features) is removed. Then, a new random forest model is retrained using the remaining feature subset, and its performance is evaluated on a fixed validation set or through cross-validation. In this embodiment, the macro-average F1 score or CSI is used as the evaluation metric. The above cycle of removal, retraining, and evaluation is repeated until either of the following conditions is met:

[0066] Performance degradation means that the model performance (macro-average F1 score) on the new feature subset no longer improves, or even begins to decline; quantity target is met means that the number of features is reduced to a preset lower limit.

[0067] Stop iterating over the feature subset from the previous round, in order to select the optimal feature subset, that is, the most discriminative feature combination selected from all original features and derived features.

[0068] like Figures 2-3 As shown, the 16 optimal features retained after random forest recursive feature elimination and their relative importance ranking are as follows: ice point cloud-water mixing ratio (Qc), ice point rainwater mixing ratio (Qr), 700hPa temperature (T700), overall specific humidity (Q), temperature-dew point difference (T-Td), ice point relative humidity (RH), ice point dew point temperature (Td), ice point temperature (T), zero-degree layer height (Height_0°), 0-3km vertical wind shear (SH03), ratio of precipitable water to liquid water path (PWAT / LWP), overall liquid water path (LWP), 500hPa temperature (T500), air pressure (P), total wind speed (Wind), and ratio of 600hPa to 500hPa specific humidity (Q600 / 500). The ranking results indicate that the mixing ratio of cloud water and rainwater near the icing point, as well as the temperature and humidity-related factors in the middle and lower atmosphere, have a dominant contribution to the determination of icing level, which is highly consistent with the physical mechanisms such as the supercooled water accumulation environment and thermal uplift conditions on which icing formation depends.

[0069] Step 2: Construct a heterogeneous integrated learning framework and partition the dataset.

[0070] Considering the nonlinearity and heteroscedasticity of icing data, this embodiment constructs a heterogeneous base model pool containing seven algorithms: linear model (logistic regression), kernel method model (support vector machine), tree model (decision tree, random forest, XGBoost, LightGBM), and distance model (K-nearest neighbors). Each model has the ability to output class probabilities (i.e., to output the probability vector of a sample belonging to each icing level) to ensure the diversity of learning algorithms in subsequent fusion and reduce the bias of a single model.

[0071] In terms of dataset partitioning, 10 samples were randomly selected from each of the four ice accumulation level categories in the original data to form a fixed test set, which was then isolated. This test set did not participate in any training or validation processes. All remaining samples served as the training and validation set for subsequent base model selection and meta-model training.

[0072] Step 3: Optimize the base model based on a combination of genetic algorithm and EasyEnsemble undersampling.

[0073] This step aims to automatically and intelligently select the optimal combination (especially for a few categories) from a pre-built set of heterogeneous base models, thereby avoiding the subjectivity and blindness of manual trial and error.

[0074] Step 301: Find an optimal subset of models from the seven base models (logistic regression, SVM, decision tree, random forest, XGBoost, LightGBM, and K-nearest neighbors) constructed in step 2.

[0075] In this embodiment, a genetic algorithm is used for the search. Considering that the heterogeneous base model pool allows for repeated selection of the same algorithm to further enhance ensemble diversity, this step performs a combined search with a maximum selection of 10 models as a constraint.

[0076] In terms of encoding, each candidate model combination scheme is encoded as a "chromosome," with the integers at the gene positions corresponding to the algorithm indices in the base model pool. In this embodiment, integer encoding is used, where each gene position on the chromosome represents which model in the base model pool was selected (for example, number 1 represents logistic regression, 2 represents SVM, and so on). A chromosome might be [2, 3, 5, 7], representing the selection of the combination of SVM, decision tree, XGBoost, and LightGBM.

[0077] Step 302: Calculate the fitness function.

[0078] Genetic algorithm parameters are set as follows: population size 30, maximum number of generations 50, crossover probability 0.8, mutation probability 0.1. Fitness calculation is the core step in model selection. The specific process is as follows: for each candidate base model combination, hierarchical 3-fold cross-validation is performed on the training and validation sets described in step 2. In each fold cross-validation, EasyEnsemble undersampling is first performed on the majority class samples in the training part, that is, a subset of samples equal in size to the minority class samples is randomly drawn from the majority class samples (e.g., samples without ice accumulation), so that the number of samples in each class is balanced on the training set of this fold; then, all base models in this combination are trained using the balanced training set. The macro-average F1 score is calculated on the validation set after the same undersampling balancing process.

[0079] Macro average F1 score The calculation formula is:

[0080]

[0081] in, The number of folds for cross-validation. This represents the total number of categories for ice accumulation levels. and Let be the precision and recall of class c on the i-th validation set, respectively.

[0082] Based on the above formula, the F1 score for each fold and each category is first calculated, then the average score (macro-average) is taken over all categories, and finally the average score for the 3-fold results is taken again. This average score is used as the fitness value of the candidate model combination (chromosome). The higher the fitness, the stronger the comprehensive recognition ability of the combination across all categories. If a combination is empty, the fitness is directly set to 0.

[0083] Step 303, Genetic Evolution and Search. A certain number of model combinations are randomly generated to form the initial population. The genetic algorithm continuously evolves the population through selection, crossover, and mutation operations. After 50 generations of iterative search, it finally converges to obtain the optimal base model combination that maximizes the macro-average F1 score on the balanced validation set. This combination includes five base models: decision tree, random forest, support vector machine, XGBoost, and LightGBM, covering three heterogeneous paradigms: tree models, kernel method models, and ensemble tree models. In the specific aircraft icing dataset used in this embodiment, logistic regression and K-nearest neighbors algorithms provide relatively little help in accurately identifying categories like "severe icing" with a small number of samples, and therefore were not included in the optimal subset. This result demonstrates the effectiveness and objectivity of the genetic algorithm's automatic selection.

[0084] Step 4, Metamodel stacking and fusion.

[0085] This step aims to automatically learn, in a data-driven manner, how to merge the prediction results of multiple optimal base models selected in step 3 into a more powerful final prediction model.

[0086] Based on the optimal base model combination obtained in step 3 and the training and validation sets partitioned in step 2, a meta-classifier is trained to learn how to make the best final decision based on the probabilities output by each base model. Specifically, the training and validation sets are further divided into K folds (K=5 in this embodiment). This hierarchical partitioning ensures that the proportion of samples of each ice accumulation level in each fold is consistent with the overall distribution. K rounds of training and prediction are performed on each base model in the optimal base model subset. In the i-th round, the i-th fold is used as the validation set, and the remaining K-1 folds are used as the training set. During training, the EasyEnsemble undersampling strategy is also used to maintain class balance. The trained model is used to predict the i-th fold validation set that has not participated in this round of training, outputting the probability vector of each sample in that fold belonging to each class. For example, for a 4-class classification problem, each sample receives a 4-dimensional vector: [probability of belonging to class 0, probability of belonging to class 1, probability of belonging to class 2, probability of belonging to class 3]. After K rounds, all samples in the training and validation sets receive a set of probability outputs generated by unseen samples. This process is repeated once for each base model in the optimal subset. The 4-dimensional probability vectors output by the five base models for each sample are concatenated vertically to form a meta-feature matrix. For example, with five base models, each outputting a 4-dimensional probability vector, the concatenation results in a 20-dimensional vector. Each row of the matrix corresponds to a sample, and each column corresponds to the predicted probability of a base model for a specific class. This matrix serves as the input feature for training the meta-classifier.

[0087] Using the aforementioned meta-feature matrix as the new feature X, and the true icing level labels corresponding to the training and validation set samples as the target y, the meta-classifier in this embodiment adopts logistic regression. When training the original logistic regression classifier, the class weight parameters are set to a balanced mode. That is, when calculating the loss, logistic regression automatically assigns higher penalty weights to the prediction errors of a few classes (such as severe icing). In the fusion stage, the influence of class imbalance on the decision boundary is further suppressed, and the optimal fusion mapping relationship from the probability output of each base model to the final icing level is learned in a data-driven manner.

[0088] Finally, the optimal subset of base models obtained in step 3 and the meta-classifier trained in step 4 together constitute a complete aircraft icing level prediction model that can be used for final prediction.

[0089] Unlike fixed voting or averaging methods, stacking allows a meta-model to automatically learn "when to trust which base model" from the data. Different base models perform differently on different types of data subsets or categories. Stacking ensemble integrates the strengths of these heterogeneous models through the meta-model. Furthermore, by reintroducing class weight balancing at the meta-classifier stage, it further strengthens the overall model's tendency to identify minority classes (moderate, severe icing), consolidating the effectiveness of addressing the false negative problem from the fusion strategy level.

[0090] It should be noted that the sample size distribution of different icing levels (no, light, moderate, and severe) in aircraft icing observation data is extremely uneven, especially the proportion of moderate and severe icing samples is very low. If conventional machine learning methods are used for training directly, the model is prone to bias towards the majority class (no icing), resulting in a severe deficiency in the ability to identify the critical minority class (severe icing). To solve the above problem, this embodiment adopts a dual class imbalance processing mechanism.

[0091] The first stage of processing, in the base model training and genetic algorithm fitness evaluation phase (step 3), employs the EasyEnsemble undersampling strategy to balance the training data. This strategy generates multiple balanced subsets through multiple random undersampling operations, and trains the base model on each subset separately. This ensures that the base model can fully learn the discriminative features of the minority class while retaining information from the majority class samples.

[0092] The second layer of processing occurs during the training phase of the meta-classifier (logistic regression) (step 4). The class weight parameters are set to "balanced" mode, meaning the weight coefficients in the loss function are automatically adjusted based on the inverse ratio of the number of samples in each class. This causes the meta-classifier to impose a higher penalty on prediction errors of minority class samples when fusing the base model output, thereby further correcting the decision boundary shift caused by class imbalance at the final decision level.

[0093] The dual-class imbalance handling mechanism works synergistically at both the base model learning and meta-model fusion levels, significantly improving the recognition accuracy and recall rate for moderate and severe icing levels.

[0094] Example 2

[0095] This embodiment discloses a method for evaluating and verifying aircraft icing levels, used to evaluate and report the generalization performance of the final model in Embodiment 1.

[0096] First, determine the number of times N will be repeated for the evaluation.

[0097] The following steps are required for a single independent evaluation cycle:

[0098] A. Use a new random seed for partitioning: At the start of each loop, use a different random seed than the previous one.

[0099] B. Repeat steps 2 to 4 of Example 1, performing the entire process of data partitioning, model selection, and training prediction. Each time, calculate the accuracy, macro-average F1 score, F1 score for each category, critical success index (CSI), and macro-average AUC on an isolated, fixed test set. Use the average value of each indicator across multiple independent runs as the final performance evaluation basis to eliminate the influence of randomness in a single data partitioning on the evaluation conclusion.

[0100] accuracy for:

[0101]

[0102] In the formula, TP is the total number of samples correctly predicted for all classes, and TN is the total number of negative samples correctly predicted for all classes;

[0103] Recall rates for various categories c for:

[0104]

[0105] In the formula, The number of samples in category c that were correctly predicted. This represents the number of samples that were missed in category c. The macro average of the recall rates for each category is... C represents the total number of ice accumulation categories;

[0106] Precision of various categories c for:

[0107]

[0108] In the formula, The number of samples in category c that were correctly predicted. This represents the number of samples that were misclassified as category c from other categories.

[0109] F1 scores by category c for:

[0110]

[0111] In the formula, For category c, Let be the recall rate for category c;

[0112] Critical Success Indices for Various Categories for:

[0113]

[0114] In the formula, The number of samples in category c that were correctly predicted. This represents the number of samples from other categories that were misclassified as category c. This represents the number of samples that were missed in category c.

[0115] To verify the effectiveness of this invention, four internationally recognized traditional methods for predicting the physical indices of aviation icing were selected as benchmarks. The definitions and calculation methods of each method are as follows:

[0116] (1) Icing Index (IC):

[0117]

[0118] In the formula: RH is relative humidity (unit: %), and T is ambient temperature (unit: ℃). IC index < 0 indicates no ice accumulation, 0 ≤ IC index < 40 indicates light ice accumulation, 40 ≤ IC index < 70 indicates moderate ice accumulation, and 70 ≤ IC index indicates heavy ice accumulation.

[0119] (2) Improve the IC index (Improved Icing Index):

[0120] In the formula: RH is relative humidity (unit: %), and T is ambient temperature (unit: ℃). Below, IIC index < 0 indicates no ice accumulation, 0 ≤ IIC index < 4 indicates light ice accumulation, 4 ≤ IIC index < 7 indicates moderate ice accumulation, and 7 ≤ IIC index indicates severe ice accumulation.

[0121] (3) TF method (Temperature-Fog method):

[0122]

[0123] In the formula: V is the false frost point temperature, V is the relative air speed (hereinafter referred to as airspeed, unit: 100 km·h-1), T is the temperature (unit: ℃), and Td is the dew point temperature (unit: ℃). No ice buildup There is ice accumulation. Moderate or higher level of icing

[0124] (4) LWC method (Liquid Water Content method):

[0125]

[0126] In the formula Liquid water content (unit: g / m³); 1≤ <6 indicates light icing, 6≤ <8 indicates moderate icing, 8≤ The ice accumulation is moderate.

[0127] like Figure 4The figure shows a comprehensive comparison of the accuracy, macro-average precision, macro-average recall, and macro-average critical success index of five methods (IC index, improved IC index, TF method, LWC method, and this invention) in the four-category icing classification. Among them, the IC index and improved IC index are based on empirical thresholds for temperature and humidity; the TF method combines temperature and dew point temperature; and the LWC method is based on a single threshold for liquid water content. Figure 4 As can be seen, the performance of the four traditional physical index methods is at a low level: the IC index has an accuracy of 0.400, a macro-average precision of 0.425, a macro-average recall of 0.400, and a macro-average CSI of 0.270, indicating limited discriminative ability; the improved IC index has an accuracy of 0.325, a macro-average precision of 0.280, a macro-average recall of 0.325, and a macro-average CSI of 0.171, showing a further decline in overall performance; the TF method has an accuracy of 0.300, a macro-average precision of 0.316, a macro-average recall of 0.300, and a macro-average CSI of 0.116, performing relatively weakly among traditional methods; the LWC method has an accuracy of 0.343, a macro-average precision that drops sharply to 0.183, a macro-average recall of 0.400, and a macro-average CSI of 0.165, with acceptable recall but severely low precision. Overall, the macro-average CSI of the four traditional physical index methods is below 0.27, indicating that the discrimination strategy based on empirical thresholds has an inherent limitation in its ability to identify the overall ice accumulation level. In particular, the problem of missed reports is prominent in a few types of samples such as severe ice accumulation, making it difficult to meet the actual early warning needs.

[0128] In comparison, the GA-Stacking Ensemble method of this application achieves significant breakthroughs in all four core metrics: accuracy reaches 0.675, macro-average precision is 0.681, macro-average recall is 0.675, and macro-average CSI is 0.510. Compared with the IC index, which has the best overall performance among traditional methods, accuracy is improved by 0.275, macro-average precision by 0.256, macro-average recall by 0.275, and macro-average CSI by 0.24; compared with the LWC method, which has a higher recall, accuracy is improved by 0.332, and macro-average CSI by 0.345. The above quantitative comparison results fully demonstrate that this invention, through genetic algorithm-driven adaptive combinatorial search of heterogeneous base models, supplemented by the EasyEnsemble undersampling mechanism to effectively suppress class imbalance, and the data-driven optimal mapping of the output probabilities of each base model by the meta-model stacking and fusion strategy, improves the inherent defects of traditional empirical threshold methods in failing to characterize the physical conditions of ice formation and the ability to identify scarce sample categories such as severe ice accumulation. Furthermore, to further analyze the specific classification error patterns of each method in the four-class ice accumulation level discrimination, Figures 5(a)-5(d) show the confusion matrices of the five methods. Rows correspond to the true labels (0 to 3 for no ice accumulation, light ice accumulation, moderate ice accumulation, and severe ice accumulation, respectively), columns correspond to the predicted labels, diagonal elements represent the proportion of each category correctly classified, and off-diagonal elements reflect the distribution of misclassification and omission.

[0129] Figures 5(a)-5(d) are heatmaps comparing the performance evaluation metrics of this application and traditional physical index methods on a four-class icing level recognition task. They show the precision, recall, F1 score, and critical success index (CSI) evaluation results of the IC index, improved IC index, TF method, LWC method, and the GA-Stacking Ensemble method of this invention at different icing levels. The horizontal axis corresponds to different icing level categories, and the vertical axis corresponds to different recognition methods. The color intensity and numerical value reflect the performance level of each method in the corresponding category. Figures 5(a)-5(d) Evaluation metrics of different methods at various icing levels reveal significant differences in the aircraft icing identification capabilities between traditional empirical diagnostic methods and machine learning methods. Among these, the GA-Stacking Ensemble multivariate heterogeneous ensemble learning method proposed in this application demonstrates superior overall performance across multiple evaluation metrics, including precision, recall, F1 score, and Critical Success Index (CSI). This indicates that the method can more effectively characterize the nonlinear relationship of icing formation under complex weather conditions and exhibits strong stability and generalization ability across different icing levels.

[0130] From the accuracy results, the IC index method has high accuracy in the "no icing" category, but its performance drops significantly in the light, moderate, and severe icing levels, indicating that the traditional IC index is more inclined to identify large-scale stable environments and has limited ability to distinguish complex icing processes. The improved IC index enhances the identification ability of moderate and severe icing to some extent, but the overall accuracy remains low, especially in the light icing level where it almost fails, indicating that a single empirical threshold method cannot accurately describe the multi-factor coupling effect in the icing formation process. The TF method performs well in the no-icing category, but has almost no ability to identify severe icing, exhibiting a significant class bias. The LWC method has high accuracy in the light icing category, but its adaptability to other categories, especially moderate and severe icing, reflects that a single liquid water content index cannot comprehensively reflect the thermodynamic and dynamic conditions of icing. In contrast, the GA-Stacking Ensemble method has a more balanced accuracy distribution across all icing levels. The precision for the no-ice category reached 0.889, while the precision for moderate and severe icing reached 0.600, indicating that the model not only effectively reduces the false alarm rate but also maintains high classification reliability in complex icing environments. This demonstrates that multi-model heterogeneous ensemble can fully integrate the complementary advantages of different algorithms in feature representation, thereby improving the model's ability to identify multiple categories of icing states. From the recall results, traditional methods generally suffer from category omission. For example, while the TF method can identify moderate icing well, its recall for severe icing is 0%; the LWC method achieves a recall of 0.800 for light icing but also fails to effectively identify moderate and severe icing. The IC index and the improved IC index also have relatively low recall rates for severe icing, indicating that traditional methods have a high risk of underreporting under severe icing conditions. In contrast, the GA-Stacking Ensemble method achieved high recall rates in the categories of no icing, light icing, and moderate icing, with a recall rate of 0.900 for moderate icing. This indicates that the method can more effectively capture key feature information during the icing process, reduce missed detections, and improve the completeness of icing identification. From the F1 score, the GA-Stacking Ensemble method significantly outperforms other methods across all categories. Since the F1 score considers both precision and recall, it better reflects the overall classification ability of the model. Traditional methods show significant fluctuations between different categories. For example, the TF method has some advantages in moderate icing but poor overall stability; the LWC method only performs well in light icing.In comparison, the GA-Stacking Ensemble achieved higher F1 scores in all three icing levels: no icing, light icing, and moderate icing. Specifically, it achieved 0.842 for no icing and 0.720 for moderate icing, indicating that the method maintains high recall while ensuring accuracy, resulting in more stable overall classification performance. Furthermore, the Critical Success Index (CSI) of the GA-Stacking Ensemble method is higher than traditional methods in most icing levels. Specifically, it achieves 0.727 for no icing and 0.562 for moderate icing, significantly outperforming the IC index, the improved IC index, and the TF and LWC methods. CSI considers hits, misses, and false alarms simultaneously, making it more suitable for meteorological disaster identification and assessment. These results demonstrate that the proposed multivariate heterogeneous ensemble learning method can effectively reduce false alarms and misses in icing identification, improving the overall reliability of icing event identification.

[0131] Figures 6(a)-6(e) show a comparison of the confusion matrices of the present application and traditional physical index methods in the four-class icing level discrimination task. Figures 6(a)-6(d) correspond to the confusion matrices of four traditional physical index methods: IC index, improved IC index, TF method, and LWC method, respectively. Figure 6(e) corresponds to the confusion matrix of the GA-Stacking Ensemble method of the present invention. Rows correspond to the true icing level labels, columns correspond to the predicted level labels, diagonal elements represent the correct recognition ratio of each category, and off-diagonal elements reflect the distribution of misclassification and omission. The figures visually demonstrate that the confusion matrices of different methods in aircraft icing level recognition show significant differences. Traditional physical index methods still exhibit strong class bias and structural misclassification characteristics, while the GA-Stacking Ensemble method proposed in this application shows more reasonable class discrimination ability and more stable multi-level recognition performance. Among them, although the IC index method has a certain recognition ability in the no-icing category (class 0), with a correct recognition rate of 0.60, its classification stability for other icing levels is poor. In the light icing samples (Class 1), only 0.10 were correctly identified, while a high percentage (0.60) were misclassified as moderate icing (Class 2). In the moderate icing samples (Class 2), only 0.40 were correctly identified, while the remaining 0.60 were misclassified as severe icing (Class 3). In the severe icing samples (Class 3), although 0.50 were correctly identified, 0.30 were still misclassified as light icing. Overall, the IC index exhibits a significant misclassification phenomenon in class transitions, meaning the model tends to shift icing levels to higher or adjacent categories, reflecting the limited ability of traditional empirical indices to characterize changes in icing intensity. Compared to the original IC index, the improved IC index method shows improved recognition ability in the non-icing category, achieving a correct recognition rate of 0.70, but its recognition of icing categories still has significant shortcomings. In the light icing category, 0.70 samples were misclassified as having no icing, and only 0.30 were identified as having moderate icing, failing to achieve effective identification of this category. In the moderate icing category, only 0.40 samples were correctly classified, with the remaining 0.40 misclassified as having no icing. In the severe icing category, only 0.20 samples were correctly identified, with the rest mainly misclassified as having no icing or light icing. This indicates that while the improved IC index improves the background field discrimination ability under some weather conditions, it still struggles to effectively establish nonlinear boundary relationships between icing levels, especially lacking stable identification ability for a few categories and high-level icing. The confusion matrix of the TF method exhibits extreme category concentration. In the no-icing category, 0.80 samples were directly misclassified as having moderate icing, with only 0.20 correctly identified. In the light, moderate, and severe icing categories, almost all samples were predicted as having moderate icing (2 categories), with the correct identification rate for both light and severe icing being 0%.This indicates that while the TF method can respond strongly to certain icing characteristic signals, the model lacks fine-grained discrimination ability for different icing levels, leading to a severe collapse of classification results into a single category and exhibiting a significant over-concentration problem. The LWC method performs relatively well in the no-ice and light-ice categories, with a correct recognition rate of 0.80 for both, but its ability to identify moderate and severe icing is significantly insufficient. In the moderate-ice category, 0.80 was misclassified as no ice, and only 0.20 was misclassified as light ice, failing to achieve correct classification at all; in the severe-ice category, 0.60 was misclassified as no ice, and 0.40 was misclassified as light ice, also failing to achieve effective identification. This shows that while a single liquid water content index can reflect some thermal conditions in low-level icing environments, it lacks sufficient characterization of the dynamic structure and phase transition characteristics in complex icing processes, thus making it difficult to apply to the identification of high-level icing. In contrast, the GA-Stacking Ensemble method proposed in this application exhibits significantly better diagonal clustering characteristics in the confusion matrix. The correct identification rate for the no-ice-accumulation category reached 0.80%, with only 0.20 samples misclassified as light ice accumulation. In the light ice accumulation category, 0.70 samples were correctly identified, with only a small number of samples misclassified as no ice accumulation or severe ice accumulation. The correct identification rate for the moderate ice accumulation category was as high as 0.90%, with only 0.10 samples misclassified as light ice accumulation, indicating that the model has a very strong ability to distinguish moderate ice accumulation. In the severe ice accumulation category, although only 0.30 samples were correctly identified, the main misclassification flow was towards moderate ice accumulation (0.60), with only 0.10 samples misclassified as light ice accumulation, and no extreme false negatives of direct misclassification as no ice accumulation occurred.

[0132] In summary, this embodiment, based on actual aircraft icing report data, uses random forest recursive feature elimination to select 16 optimal features with clear physical meaning, effectively extracting meteorological factors highly correlated with the icing formation mechanism. An EasyEnsemble undersampling strategy is introduced in each fold cross-validation of the genetic algorithm fitness evaluation to balance the class distribution. A flexible search mechanism allowing repeated model selection and setting a maximum of 10 models successfully obtains the optimal base model combination including decision trees, random forests, support vector machines, XGBoost, and LightGBM. "Out-of-fold" predictions are generated on a strictly isolated meta-model validation set to train the meta-classifier, learning the optimal fusion mapping in a data-driven manner. After multiple random partitioning and averaging evaluations, the method of this invention significantly outperforms traditional physical index methods in terms of class recognition ability and overall stability. The confusion matrix and various indicators verify its effective solution to the problem of missed reports of severe icing samples, providing reliable decision support for aviation flight safety early warning.

[0133] The embodiments described above are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.

Claims

1. A training method for an aircraft icing prediction model based on machine learning, characterized in that, Includes the following steps: Step 1: Obtain a historical dataset containing several meteorological elements and aircraft icing level labels, construct physical derived features, and use the random forest recursive feature elimination method to iteratively select the optimal feature subset from all features; Step 2: Construct a base model pool containing multiple machine learning algorithms based on different principles, and extract equal amounts of samples from the optimal feature subset according to the ice accumulation level category to form a fixed test set and isolate it for subsequent evaluation of the model's generalization performance. The remaining samples constitute the training and validation set. Step 3: On the training and validation set, a heuristic search algorithm is used to search for a subset of base models to be optimized from the base model pool; wherein, the heuristic search algorithm uses the macro-average F1 score obtained by the subset of base models to be optimized on the training and validation set as the fitness function for selection, and outputs the optimal subset of base models obtained by the search. Step 4: Based on the optimal machine model subset, input the training and validation set into the hierarchical K-fold cross-validation module, generate corresponding prediction results using each base model, and summarize the prediction results of all base models to construct a meta-feature matrix; train a meta-classifier using the meta-feature matrix and its corresponding true ice accumulation level labels. The optimal base model subset and the meta-classifier together constitute the aircraft icing level prediction model.

2. The training method according to claim 1, characterized in that, In step 1, the physical derived characteristics include one or more of the following: the total wind speed synthesized from zonal and meridional winds, the humidity and temperature differences at different altitudes, the specific humidity ratio and gradient, and the ratio of liquid water path to precipitable water.

3. The training method according to claim 2, characterized in that, The optimal feature subset includes: cloud-water mixing ratio at the icing point, rainwater mixing ratio at the icing point, 700hPa temperature, overall specific humidity, temperature-dew point difference, relative humidity at the icing point, dew point temperature at the icing point, icing point temperature, zero-degree layer height, 0-3km vertical wind shear, ratio of precipitable water to liquid water path, overall liquid water path, 500hPa temperature, air pressure, total wind speed, and ratio of specific humidity at 600hPa to that at 500hPa.

4. The training method according to claim 1, characterized in that, In step 2, the base model pool includes at least three of the following: logistic regression, support vector machine, decision tree, random forest, XGBoost, LightGBM, and K-nearest neighbor algorithm.

5. The training method according to claim 1, characterized in that, In step 3, the heuristic search algorithm is a genetic algorithm.

6. The training method according to any one of claims 1-5, characterized in that, In step 3, the fitness function—the macro-average F1 score—is... The calculation formula is: , in, The number of folds for cross-validation. This represents the total number of categories for ice accumulation levels. and Let be the precision and recall of class c on the i-th validation set, respectively.

7. The training method according to claim 1, characterized in that, In step 3, when calculating fitness, and in each fold cross-validation step 4 when generating the meta-feature matrix, the EasyEnsemble undersampling strategy is used on the training data to balance the number of samples in each class.

8. The training method according to claim 1, characterized in that, In step 4, the meta-classifier is a logistic regression model, and the class weight parameters are set to a balanced mode during training, with the weight of each class being inversely proportional to the number of samples of that class in the training data.

9. A method for evaluating and verifying an aircraft icing prediction model based on machine learning, characterized in that, The aircraft icing level prediction model obtained by the training method described in any one of claims 1-8 is used for prediction; During model performance evaluation, multiple independent evaluation loops are executed, and the training method is executed independently in each loop to obtain a prediction model for this loop. The prediction model used in this iteration is used to predict the fixed test set, and the evaluation index is calculated. The average value of each evaluation index obtained from multiple iterations is used as the final performance verification basis for the aircraft icing level prediction model.

10. The evaluation and testing method according to claim 9, characterized in that, The evaluation metrics include at least one of accuracy, precision for each category, recall, F1 score, and critical success index (CSI).

Citation Information

Patent Citations

  • Fan blade icing prediction method based on feature selection and XGBoost

    CN109026563A

  • Aircraft ice accretion prediction method based on ensemble forecasting thought

    CN118468685A