A Recognition Method and System for Flight Return and Diversion Based on Machine Learning

Through machine learning-based identification methods, the feature samples and gradient improvement decision tree algorithm of flight data are used to quickly determine the flight return and landing reserve and extract key data, which solves the problems of low efficiency and insufficient accuracy of flight return return reserve and landing identification in the existing technology, and improves the accuracy and automation level of flight operation safety assessment.

CN119249155BActive Publication Date: 2025-07-01QINGDAO CIVIL AVIATION KAIYA SYST INTEGRATION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411764222.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-07-01
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

In the comprehensive consideration of flight data, the existing technology cannot quickly determine whether the flight returns to the air and land on reserve, and it is difficult to effectively extract the key quantitative data of the flight returns to the air and land on reserve, resulting in low accuracy in flight operation safety assessment.

Method used

Using machine learning-based identification method, through data analysis of flight data time, heading, track, air pressure altitude, radio altitude, etc., a flight learning feature sample is established, and the machine learning algorithm of gradient enhancement decision tree GBDT is used for training and learning, to quickly determine whether the flight returns to the ready for landing, and extract the key quantized data for the ready for landing.

Benefits of technology

It realizes rapid identification and critical data extraction of flight return and landing, improves the accuracy of flight operation safety assessment, provides airlines with automated auxiliary analysis support, and reduces the cost of manual identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119249155B_ABST
    Figure CN119249155B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of civil aviation operation safety assessment, and discloses a method and system for identifying flight return and diversion based on machine learning. The method comprehensively considers data such as time, heading, track, barometric altitude, and radio altitude of flight data of flights to establish flight learning feature samples; conducts a large number of trainings and learnings with the flight data actually generated during the flight of the flight; based on the machine learning algorithm of gradient boosting decision tree, quickly determines whether it is a flight for return and diversion from the massive flight data after training and learning, extracts the key quantified data of flight return and diversion, and conducts an operation safety assessment of the flight. The present invention provides automated auxiliary analysis support for the operation safety analysis of airlines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of civil aviation operation safety assessment, and particularly relates to a method and system for identifying flight return and diversion based on machine learning. Background Art

[0002] Currently, the identification of flight return and diversion by existing airlines mostly involves manually recording and marking key information such as flight numbers, finding the decoded flight data in the information system for flight return and diversion marking, and then conducting safety analysis such as information analysis of return and diversion risks. The actual identification of flight return and diversion requires comprehensive consideration and quantification of data in aspects such as time (UTC), heading (magnetic heading), track (latitude and longitude coordinates), barometric altitude, and radio altitude for evaluation.

[0003] Through the above analysis, the problems and defects of the existing technology are as follows: In the comprehensive consideration of flight data by the existing technology, it cannot quickly determine whether a flight is a return or diversion flight, cannot effectively extract the key quantified data of flight return and diversion, and has a low accuracy in evaluating the operation safety of flights. Summary of the Invention

[0004] To overcome the problems in the related art, the disclosed embodiments of the present invention provide a method and system for identifying flight return and diversion based on machine learning.

[0005] The technical solution is as follows: The method for identifying flight return and diversion based on machine learning includes:

[0006] S1. By analyzing the data of time UTC, heading, track, barometric altitude, and radio altitude of flight data, establish a flight learning feature sample;

[0007] S2. Use the established flight learning feature sample and the actual flight data generated during the flight process for training and learning. Use the training label dataset to train the model. Through the supervised training dataset containing input labels and corresponding output labels, use the validation set to monitor the performance of the model;

[0008] S3. Based on the machine learning algorithm of Gradient Boosting Decision Tree (GBDT), quickly determine whether a flight is a return or diversion flight from the trained and learned flight data, extract the key quantified data of flight return and diversion, and evaluate the operation safety of the flight.

[0009] In step S1, establishing a flight learning feature sample includes: establishing the features of flight return and diversion, known a set composed of samples and corresponding labels, and the expression is:

[0010] ;

[0011] In the formula, the sample is described by attributes, and the value of the label is 1 or -1;

[0012] Using the supervised learning algorithm, the dependence between the input and the output is determined from the data set, and the prediction function is determined to obtain the predicted value ;

[0013] In flight risk identification, the feature represents the terminal factor vector, and each component takes an integer between 1 and 10, represents the serial number of the flight ID, indicates abnormal, and abnormal includes flight return or alternate landing, indicates that the flight is normal.

[0014] In step S2, training and learning are performed, including:

[0015] Using the cost-sensitive method or sampling-based technology, the class imbalance problem training and learning are performed on the established flight learning feature samples and the actual flight data generated during the flight process, so as to identify the flight risk according to the number of training examples of different classes in the classification task.

[0016] In step S3, the machine learning algorithm of the gradient boosting decision tree GBDT adopts the method of oversampling the training data set with the SMOTE function and combining data replication to solve the imbalance problem of the number of two types of samples.

[0017] In step S3, the flight operation safety assessment is carried out, including:

[0018] Step 1. Use the SMOTE technology for oversampling to obtain a new data set newtrain;

[0019] Step 2. Combine the abnormal data replication in newtrain and the abnormal data replication in ORItrain to form the tmpabnormal data set, and compare the sizes of the normal data set and the tmpabnormal data set to obtain keynum;

[0020] Step 3. Randomly sample the tmpabnormal and normal data sets with keynum as the sampling number respectively to form SYNtrain, and output SYNtrain as the training data of XGBOOST;

[0021] Step 4. Use the training data and xgboost to obtain a model; output the confusion matrix of the training data and the misclassified flight information;

[0022] Step 5. Call the model to predict the test data; output the confusion matrix of the test data and record the misclassified flight numbers.

[0023] In Step 1, before obtaining the new dataset newtrain, it is necessary to: combine the data received by the gradient boosting decision tree model, namely flightnormal, flightabnormal0, and flightabnormal1, to form the training dataset train1.

[0024] In Step 1, set the parameters of the data received by the gradient boosting decision tree model for training: the sampling volume is 1000%, and use 5-nearest neighbors.

[0025] In Step 2, for the abnormal data in newtrain and the abnormal data in ORItrain, check = 0;

[0026] The abnormal data in newtrain and the abnormal data in ORItrain are both copied 20 times;

[0027] Among the keynum obtained by comparing the normal dataset and the tmpabnormal dataset, take the one with the smaller number of data in the two datasets.

[0028] In Step 3, randomly sample to form SYNtrain, and the ratio of the number of 0s and 1s in SYNtrain is 1:1.

[0029] Another object of the present invention is to provide an identification system for flight return and alternate landing based on machine learning. This system implements the identification method for flight return and alternate landing based on machine learning, and this system includes:

[0030] Flight learning feature sample establishment module, used to establish flight learning feature samples by analyzing the data of time UTC, heading, track, barometric altitude, and radio altitude of flight data;

[0031] Class imbalance problem training and learning module, used to train and learn using the established flight learning feature samples and the actual flight data generated during the flight of the flight, train the model using the training label dataset, and use the validation set to monitor the performance of the model;

[0032] Gradient Boosting Decision Tree Learning Module, which is used to quickly determine whether a flight is a return-to-base alternate flight based on the machine learning algorithm of Gradient Boosting Decision Tree (GBDT), extract key quantified data for flight return-to-base alternate, and conduct an operational safety assessment of the flight.

[0033] Combining all the above technical solutions, the beneficial effects of the present invention are as follows: Based on aspects such as time (UTC), heading (magnetic heading), track (latitude and longitude coordinates), barometric altitude, and radio altitude as automated flight return-to-base alternate identification features, the present invention can quickly determine whether a flight is a return-to-base alternate flight, extract key quantified data for return-to-base alternate, conduct an operational safety assessment of the flight, and provide automated auxiliary analysis support for the airline's operational safety analysis. Traditional return-to-base alternate identification relies on manual data analysis to determine the compliance of return-to-base alternate data. The present invention can quickly identify return-to-base alternate flights, improve the reliability and availability of data, and reduce the cost of manual identification. Brief Description of the Drawings

[0034] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure;

[0035] Figure 1 is a flowchart of a method for identifying flight return-to-base alternate based on machine learning provided by an embodiment of the present invention;

[0036] Figure 2 is a diagram of the input-output relationship of the i-th sample provided by an embodiment of the present invention;

[0037] Figure 3 is a schematic diagram of the operation principle of the Boosting model provided by an embodiment of the present invention; among them, (a) is a diagram of the Boosting model (M = 1), (b) is a diagram of the Boosting model (M = 2), (c) is a diagram of the Boosting model (M = 3), (d) is a diagram of the Boosting model (M = 6), (e) is a diagram of the Boosting model (M = 10), and (f) is a diagram of the Boosting model (M = 150);

[0038] Figure 4 is an effect diagram of the Boosting process provided by an embodiment of the present invention;

[0039] Figure 5 is a schematic diagram of a system for identifying flight return-to-base alternate based on machine learning provided by an embodiment of the present invention;

[0040] In the figure: 1. Flight Learning Feature Sample Establishment Module; 2. Class Imbalance Problem Training and Learning Module; 3. Gradient Boosting Decision Tree Learning Module. Detailed Embodiments

[0041] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0042] The innovation of the present invention lies in: based on machine learning algorithms, the present invention automatically identifies whether a flight returns or diverts by analyzing flight data, solving the problem of low efficiency in traditional manual identification and authentication.

[0043] Example 1, as Figure 1 shown, the method for identifying flight return and diversion based on machine learning provided by the embodiment of the present invention includes:

[0044] S1, by analyzing the data of UTC time, heading, track, barometric altitude, and radio altitude of flight data, establish a flight learning feature sample;

[0045] S2, use the established flight learning feature sample to train and learn with the actual flight data generated during the flight of the flight, use the training label data set to train the model, and use the validation set to monitor the performance of the model;

[0046] S3, based on the machine learning algorithm of Gradient Boosting Decision Tree (GBDT), quickly determine whether a flight is a return or diversion flight from the trained and learned flight data, extract the key quantified data of the flight's return and diversion, conduct a safety assessment of the flight operation, and provide automated auxiliary analysis support for the airline's flight operation safety analysis.

[0047] Example 2, as another detailed embodiment of the present invention, the method for identifying flight return and diversion based on machine learning provided by the embodiment of the present invention includes:

[0048] The first step, feature recognition of flight return and diversion:

[0049] Given a set composed of

[0050] ;

[0051] In the formula, the example is described by attributes, and the value of the label is 1 or -1;

[0052] Determine the dependency between the input and output from the dataset using a supervised learning algorithm, and determine the prediction function and obtain the predicted value ; ;

[0053] In flight risk identification, the feature represents the terminal factor vector, and each component takes an integer between 1 and 10, represents the serial number of the flight ID, indicates abnormality, and the abnormality includes the flight returning or making an alternate landing, indicates that the flight is normal. The input-output relationship of the th sample is shown in Figure 2 .

[0054] It can be understood that according to formula (1), the prediction model is obtained from the dataset, so that , for the feature vector of any given flight, the predicted output is predicted for this flight, where indicates that the flight is abnormal (returning or making an alternate landing); otherwise, it indicates that the flight is normal.

[0055] Among them, in the data of the flight risk identification problem, the number of "abnormal" flights is very small, which belongs to the class imbalance problem.

[0056] Second step, class imbalance problem:

[0057] Common classification learning methods generally have a common basic assumption, that is, the number of training examples of different classes is quite the same. If the number of training examples of different classes is slightly different, it usually has little impact, but if the difference is large, it will cause trouble to the learning process. For example, there are 998 negative examples , but there are only 2 positive examples. Then the learning method only needs to return a learner that always predicts new samples as negative examples to achieve an accuracy of 99.8%; however, such a learner often has no value because it cannot predict any positive examples.

[0058] The class imbalance problem refers to the situation where the number of training examples of different classes in the classification task varies greatly. Without loss of generality, the present invention assumes that the number of positive class examples is small and the number of negative class examples is large. For the flight risk identification problem in the present invention, the number of examples of flights returning or making an alternate landing is small, and the number of examples of normal flights is large, which is a typical class imbalance problem.

[0059] ​There are two basic methods to solve the class imbalance problem. One is the cost-sensitive method, such as weighted support vector machines; the other is sampling-based techniques. Existing techniques generally have three approaches: The first is to directly perform "undersampling" on the negative samples in the training set, that is, removing some negative examples to make the number of positive and negative examples close, and then learning; the second is to perform "oversampling" on the positive examples in the training set, that is, adding some positive examples to make the number of positive and negative examples close, and then learning; the third is directly based on the combination of the two, that is, performing undersampling on negative examples and oversampling on positive examples. The representative algorithm of oversampling, SMOTE, generates additional positive examples by interpolating the positive examples in the training set. On the other hand, if the undersampling method randomly discards negative examples, some important information may be lost; the representative algorithm of the undersampling method, EasyEnsemble, uses the ensemble learning mechanism to divide the negative examples into several sets for different learners to use. In this way, for each learner, undersampling is performed, but globally, important information will not be lost.

[0060] The third step, the gradient boosting decision tree machine learning algorithm:

[0061] Model combinations, such as Gradient Boosting Decision Tree (GBDT) or Random Forest (RF), combine simple models, and the prediction effect is better than that of a single more complex model. There are many combination methods, and randomization (such as RF) and Boosting (such as GBDT) are typical methods among them. The present invention selects the Gradient Boosting method. An initial model is established as the iteration basis. In each iteration, it is necessary to calculate the residual between the predicted value of the current model and the actual value as the target variable, train a new weak classifier, and add the prediction result of the newly trained weak classifier to the predicted value of the current model to update the model. Repeat calculating the residual, training a new weak classifier, and updating the model until the preset number of iterations is reached or the model performance reaches the preset accuracy rate.

[0062] As Figure 3 shown, the idea of Boosting is quite simple. Generally speaking, for a data set, M models (such as classification) are established. Generally, these models are relatively simple and are called weak classifiers. Each time of classification, the weight of the data misclassified in the previous time is increased a little and then classified. In this way, the final classifier can achieve relatively good results on both the test data and the training data. Figure 3In Figures (a)-(f), the green line represents the currently obtained model (the model is obtained by merging the models obtained in the previous M times), the dashed line represents the current model, and the red and blue dots are the data. The larger the dot, the higher the weight.

[0063] Figure 4 As shown, it is a Boosting process. The solid line represents the currently obtained model (the model is obtained by merging the models obtained in the previous M times), and the dashed line represents the current model. Each time of classification, more attention will be paid to the misclassified data. Figure 4 Among them, the dots shown by the circles and dots are the data. The larger the dot, the higher the weight. When M = 150, the obtained model can almost distinguish between the red and blue dots.

[0064] Step 4, the learning method and steps of the gradient boosting decision tree.

[0065] Regarding the problem of the imbalance in the number of two types of samples in the experiment (seed = 300 and seed = 321), the method of oversampling the training data set with the SMOTE function and combining data replication is adopted to solve it. The detailed steps are as follows:

[0066] Step 1. The training data received by the gradient boosting decision tree model, that is, flightnormal, flightabnormal0, and flightabnormal1, are used to form the training data set train1. Then, the SMOTE technology is used for oversampling. K nearest neighbor samples are selected from the minority class samples, and the difference between the two samples is calculated to randomly generate new samples. For each difference, a random number is multiplied, and then the result is added to the original sample to obtain a synthesized new sample and a new data set newtrain. The parameter settings for this experiment are: sampling volume 1000%, using 5-nearest neighbors.

[0067] Step 2. Next, the abnormal data (check = 0) in newtrain are copied 20 times and the abnormal data (check = 0) in ORItrain are copied 20 times to form the tmpabnormal data set. Then, the sizes of the normal data set and the tmpabnormal data set are compared to obtain keynum (take the smaller number of data in the two data sets).

[0068] Step 3. Random sampling is performed on the tmpabnormal and normal data sets with keynum as the sampling number to form SYNtrain (so that the ratio of the number of 0s and 1s in SYNtrain is 1:1), and SYNtrain is output as the training data for XGBOOST.

[0069] Step 4. Use the training data and xgboost to obtain a gradient boosting decision tree model. Output the confusion matrix of the training data and the misclassified flight information. The confusion matrix is a matrix sample table of the gradient boosting decision tree for two types of data samples.

[0070] Step 5. Call the model to predict the test data. Output the confusion matrix of the test data and record the misclassified flight numbers.

[0071] Examples of gradient boosting decision trees. Table 1 records the confusion matrices of two groups of experiments with seeds = 300 and seeds = 321 on 4 data sets. Table 2 gives the information of the misclassified flights in the corresponding experiments.

[0072] Observing the results of the model on 8 test sets in Table 1 and Table 2, corresponding to the confusion matrices in the second and fourth columns of Table 1 and the misclassified flight information in the second and fourth columns of Table 2 respectively, the following phenomena can be summarized:

[0073] (1) Using abnormal1 to identify return and diversion flights does not bring any new information, that is, the number of return and diversion flights identified by the model (the number of true samples) is not any different. Combining with Table 2, it is found that the misclassified return and diversion flights abnormal0 in the test set are exactly the same in the two cases of not using abnormal1 data (the second column) and using abnormal1 data (the fourth column). After using abnormal1, one new sample is added to each of the test sets before and after takeoff, and this sample is misjudged.

[0074] (2) Using abnormal1 increases the number of false negative samples, that is, the number of misclassified normal flights in the fourth column of Table 2 is more than the number of misclassified normal flights in the second column.

[0075] A possible explanation for the above phenomena is that when the data abnormal1 is put in, as seen from the third and fifth columns of Table 2, indeed a small number of abnormal1 are identified in the training, and the recognition rate of normal flights in the training set is also improved. However, the price paid here is overfitting, resulting in some normal flights in the test data being misjudged as abnormal flights.

[0076] Confusion matrices of two groups of experiments (seed = 300) and (seed = 321) for four times of learning in Table 1

[0077]

[0078] Flight numbers predicted incorrectly corresponding to Table 1 in Table 2

[0079]

[0080] Example 2, as Figure 5As shown in the figure, the recognition system for flight return and alternate landing based on machine learning provided by the embodiments of the present invention includes:

[0081] A flight learning feature sample establishment module 1, configured to establish flight learning feature samples by analyzing data such as UTC time, heading, track, barometric altitude, and radio altitude of flight data.

[0082] A class imbalance problem training and learning module 2, configured to use the established flight learning feature samples and the actual flight data generated during the flight of the flight for training and learning, train the model using the training label dataset, and monitor the performance of the model through the supervised training dataset including input labels and corresponding output labels, and use the validation set.

[0083] A gradient boosting decision tree learning module 3, configured to quickly determine whether it is a return and alternate landing flight from the trained and learned flight data based on the machine learning algorithm of gradient boosting decision tree GBDT, extract the key quantified data of the flight for alternate landing, and perform a safety assessment of the flight operation.

[0084] Exemplarily, the flight learning feature sample establishment module 1 includes:

[0085] A feature recognition module for flight return and alternate landing, configured to determine the dependency relationship between the input and output from a set composed of known samples and their corresponding labels, and determine the prediction function , and obtain the predicted value .

[0086] The class imbalance problem training and learning module 2 includes: a class imbalance problem training and learning module, configured to identify the number of training examples of different classes in the classification task by a cost-sensitive method or a sampling-based technique.

[0087] The gradient boosting decision tree learning module 3 includes:

[0088] A sub-module for obtaining a new dataset newtrain, configured to use the data received by the model as training data, that is, flightnormal, flightabnormal0, and flightabnormal1 to form a training dataset train1, and then perform oversampling using the SMOTE technique to obtain a new dataset newtrain.

[0089] The Keynu acquisition sub-module is used to copy the abnormal data in newtrain 20 times and the abnormal data in ORItrain 20 times to form the tmpabnormal data set, and then compare the sizes of the normal data set and the tmpabnormal data set to obtain keynum.

[0090] The SYNtrain composition sub-module is used to randomly sample the tmpabnormal and normal data sets with keynum as the number of samples to form SYNtrain and output SYNtrain as the training data for XGBOOST.

[0091] The confusion matrix and misclassified flight information output sub-module is used to obtain a model using the training data and xgboost. Output the confusion matrix of the training data and the misclassified flight information.

[0092] The prediction sub-module is used to call the model to predict the test data. Output the confusion matrix of the test data and record the misclassified flight numbers.

[0093] Currently, airlines still analyze the return and diversion of flight operation data through manual interpretation to verify whether the flight data meets the conditions for return and diversion, with very low efficiency. However, the present invention can automatically identify the characteristics of flight return and diversion based on data such as the time (UTC), course (magnetic course), track (latitude and longitude coordinates), barometric altitude, and radio altitude of the flight data, quickly determine whether a flight returns or diverts, extract the key quantified data for return and diversion, conduct a safety assessment of flight operations, and provide automated auxiliary analysis support for airlines' flight safety analysis.

[0094] The above is only a relatively preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made by those skilled in the art within the technical scope disclosed by the present invention, as long as they are made within the spirit and principle of the present invention, shall be covered by the protection scope of the present invention.

Claims

1. A method for identifying flight return and alternate landing based on machine learning, characterized in that: The method includes: S1, establish flight learning feature samples by analyzing the flight data of time UTC, heading, track, pressure altitude, and radio altitude; S2, using the established flight learning feature samples and the flight data actually generated during the flight process to train and learn, using the training label data set to train the model, and using the supervised training data set to contain input labels and corresponding output labels, and using the validation set to monitor the performance of the model; S3, a machine learning algorithm based on the gradient boosting decision tree (GBDT), quickly determines whether a flight is a return flight from the flight data after training and learning, and extracts the key quantitative data of the return flight to conduct an operational safety assessment of the flight; In step S2, training and learning are performed, including: Using cost-sensitive methods or sampling-based techniques, the established flight learning feature samples and the flight data actually generated during the flight process are trained and learned for the class imbalance problem, which is used to identify flight risks based on the number of training samples of different categories in the classification task; In step S3, the machine learning algorithm of the gradient boosting decision tree GBDT adopts the method of oversampling the training data set with the SMOTE function and combining data replication to solve the imbalance problem of the number of samples of the two categories; In step S3, an operational safety assessment is performed on the flight, including: Step 1. Use SMOTE technology to perform oversampling to obtain a new data set newtrain; Step 2. Copy the abnormal data in newtrain and ORItrain to form the tmpabnormal data set, and compare the sizes of the normal data set and the tmpabnormal data set to get keynum; Step 3. Take keynum as the sampling number and randomly sample the tmpabnormal and normal data sets to form SYNtrain and output SYNtrain as the training data of XGBOOST; Step 4. Use the training data and xgboost to get the model; output the confusion matrix of the training data and the misclassified flight information; Step 5. Call the model to predict the test data; output the confusion matrix of the test data and record the misclassified flight numbers; In step 1, the parameters of the data training data received by the gradient boosting decision tree model are set as follows: sampling volume 1000%, using 5-nearest neighbors; In step 2, check=0 for the abnormal data in newtrain and the abnormal data in ORItrain; The abnormal data in newtrain and the abnormal data in ORItrain are replicated 20 times; Compare the keynums obtained from the normal dataset and the tmpabnormal dataset, and take the one with the smaller number of data in the two datasets; In step 3, a random sample is formed into SYNtrain, and the ratio of the number of 0s and 1s in SYNtrain is 1:

1.

2. The method for identifying flight return and alternate landing based on machine learning according to claim 1, characterized in that: In step S1, a flight learning feature sample is established, including: establishing the features of flight return and alternate landing, a set of N samples and corresponding labels is known, and the expression is: D={(x1, y1), (x2, y2)…(xN, yN)} In the formula, the sample xi is described by P attributes, and the value of the label yi is 1 or -1; Use the supervised learning algorithm to determine the dependency between input x and output y from the data set, determine the prediction function f(x), and obtain the predicted value Y=f(x); In flight risk identification, feature xi represents the terminal factor vector, each component takes an integer between 1 and 10, i represents the serial number of the flight ID, yi=1 indicates abnormality, which includes flight return or diversion, and yi=-1 indicates normal flight.

3. The method for identifying flight return and alternate landing based on machine learning according to claim 1, characterized in that: In step 1, before obtaining the new data set newtrain, it is necessary to: combine the data training data received by the gradient boosting decision tree model, namely flightnormal, flightabnormal0 and flightabnormal1, into a training data set train1.

4. A flight return and alternate landing identification system based on machine learning, characterized in that: The system implements the flight return and alternate landing identification method based on machine learning as claimed in any one of claims 1 to 3, and the system includes: A flight learning feature sample establishment module (1) is used to establish a flight learning feature sample by analyzing the flight data including the time UTC, heading, track, pressure altitude, and radio altitude; The class imbalance problem training and learning module (2) is used to train and learn using the established flight learning feature samples and the flight data actually generated during the flight process, using the training label data set to train the model, and using the supervised training data set to include input labels and corresponding output labels, and using the validation set to monitor the performance of the model; The gradient boosting decision tree learning module (3) is used for a machine learning algorithm based on the gradient boosting decision tree GBDT to quickly determine whether a flight is a return or diversion flight from the flight data after training and learning, and to extract the key quantified data of the return or diversion flight to conduct an operational safety assessment of the flight.

Citation Information

Patent Citations

  • Flight behavior detection method and device

    CN107808552A

  • Aero-engine time series abnormality detection method based on CNN feature extraction

    CN109035488A

  • Flight safety risk assessment method and device, electronic equipment and storage medium

    CN117972336A