A flight test safety risk quantitative evaluation system and method based on XGBoost-SHAP

CN122840665APending Publication Date: 2026-09-29BEIJING XINMINGSHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610994784.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了一种基于XGBoost-SHAP的试飞安全风险量化评估系统及方法,解决了现有层次分析法、单一机器学习模型、LIME解释算法的缺陷

Benefits of technology

(1)、本发明彻底解决传统层次分析法风险淹没问题:基于XGBoost非线性拟合能力自动挖掘多层级指标耦合风险,不会埋没高风险小权重指标;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840665A_ABST
    Figure CN122840665A_ABST
Patent Text Reader

Abstract

This invention discloses a flight test safety risk quantification assessment system and method based on XGBoost-SHAP, belonging to the field of aviation flight test safety risk assessment technology. Addressing the shortcomings of traditional analytic hierarchy process (AHP), such as risk inundation, inability to automatically iterate and optimize, low utilization of historical flight test data, and lack of interpretability of risk assessment results, this invention constructs a hierarchical, multi-level flight test risk indicator system. Data preprocessing is completed through data quality verification, multi-mode indicator standardization, and small-sample data augmentation. Hyperparameter optimization is performed using grid search. The performance of XGBoost and ResDNN residual deep network models is compared, and XGBoost is selected as the core risk prediction model to achieve high-precision risk quantification. The SHAP algorithm is introduced to replace the traditional LIME algorithm for risk contribution analysis, and the robustness of the interpretation algorithm is verified based on Spearman correlation coefficient, quantifying the contribution of each level of indicator to the total risk. Finally, the invention provides visualized outputs including risk prediction comparison charts, indicator Sankey diagrams, and multi-dimensional risk contribution waterfall charts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aviation flight test safety risk assessment technology, and in particular to a flight test safety risk quantification assessment system and method based on XGBoost-SHAP. Background Technology

[0002] Aircraft flight testing is a core component of type certification. Flight test scenarios include various high-risk scenarios such as high temperature, stall, strong crosswinds, and avionics verification. Multiple factors are coupled, including the crew, aircraft condition, weather and airspace conditions, ground support, and the complexity of the flight test mission, resulting in highly concealed risks. The current mainstream method for flight test risk assessment in the industry is the Analytic Hierarchy Process (AHP), which has the following core shortcomings: Risk inundation problem: In the process of multi-level indicator weighting calculation, key low-weight high-risk indicators are easily covered by the overall value, making it impossible to identify hidden major risk points; No automatic iterative optimization capability: The weights rely entirely on manual settings by experts, and the evaluation rules cannot be automatically corrected using historical test flight data. The evaluation accuracy cannot be improved with the accumulation of test flight samples. Insufficient data utilization: Unable to automatically mine nonlinear risk correlations in historical 62 test flights and multi-subject expert scoring data; The results lack interpretability: only a comprehensive risk value is output, which cannot quantify the contribution of a single indicator to the overall risk and is difficult to support risk rectification and source tracing.

[0003] Existing machine learning risk assessment solutions suffer from two main shortcomings: Firstly, using XGBoost and neural networks alone only outputs risk predictions, making them black-box models that cannot explain the causes of risk. Secondly, existing model interpretation algorithms, including LIME and SHAP, have not undergone robust comparison and selection for the multi-dimensional, hierarchical indicator system of flight tests. LIME is sensitive to non-linear risk data disturbances in flight tests, resulting in poor stability of interpretation results. Furthermore, flight test scenarios are typically characterized by small sample sizes, making direct model training prone to overfitting. A complete data standardization and data augmentation process tailored to the characteristics of flight test indicators is lacking.

[0004] In summary, there is a need for a comprehensive quantitative assessment system and method for flight test safety risks that fully covers indicator preprocessing, high-precision risk prediction, and stable and explainable risk source tracing. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a flight test safety risk quantification assessment system and method based on XGBoost-SHAP, which overcomes the deficiencies of existing analytic hierarchy process, single machine learning model, and LIME interpretation algorithm.

[0006] Technical Solution: To solve the above-mentioned technical problems, according to one aspect of the present invention, more specifically, a flight test safety risk quantification assessment system based on XGBoost-SHAP, comprising 7 functional sub-modules, wherein data flows unidirectionally between sub-modules and collaboratively completes the full-process quantitative assessment of flight test risks, as detailed below: Flight Test Risk Indicator Collection Submodule Two types of basic data sources were collected: 30 expert scoring questionnaires on the importance and relevance of valid indicators; and data from 62 complete test flights, covering 15 test flight subjects. A five-level hierarchical indicator library was constructed, with a total of 202 original indicators collected, divided into five primary risk dimensions: "human" crew risk factors, "aircraft" aircraft status risk factors, "environment" operating environment factors, ground support risk factors, and test flight method risk factors. After expert screening, 132 valid indicators were retained, including 90 input terminal risk indicators and 32 output risk indicators.

[0007] Indicator Data Quality Assessment Submodule Configure automatic verification rules based on two dimensions: distinguishability and independence. Discrimination verification: The standard deviation of the indicator is ≥0.5; the range of the maximum and minimum values ​​of the indicator matches the numerical span of the actual test flight scenario; indicators with a standard deviation of less than 0.5, such as friction effect, airport elevation, and complex terrain, are deemed qualified if they conform to the actual test flight distribution after expert verification. Independence check: Only the input metrics are constrained, and the correlation coefficient between any two input metrics is <1, to avoid multicollinearity interference in model training; the output metrics have a natural hierarchical relationship, so no independence check is performed.

[0008] Multi-mode indicator standardization submodule Upon receiving quality verification indicators, four standardization methods are matched based on the indicator's numerical characteristics and business attributes. The standardized output value range is uniformly set to [0, 10] risk score. Indicators without valid collected data are uniformly assigned a value of 0 to reduce computational interference. Classification method: Adapt to discrete type indicators, such as the aircraft development stage and whether there is ground monitoring, and map fixed risk values ​​according to risk trends; Scaling method: Adapts to continuous numerical indicators and preserves numerical details and risk trends through linear transformation, such as the number of times the unit operates the machine, the age of the machine, and the crosswind speed; Hybrid approach: Standardization of the same indicator is achieved by combining classification and scaling in segments; Subjective method: Subjective indicators without objective quantitative values ​​are filled in by experts with risk scores in the range of 0 to 10, such as unit fatigue index and operational difficulty.

[0009] Data Preprocessing and Enhancement Submodule Includes data segmentation units and data augmentation units: The data segmentation unit is configured with two sets of segmentation logic: ① Hyperparameter search and model pre-training: The dataset is divided into training set and test set in an 8:2 ratio; ② Horizontal performance comparison of the model: 62-fold K-fold cross-validation is used, with 61 flights taken as the training set and 1 flight taken as the test set each time; the logic of "segment first, then augment" is strictly followed to avoid data pollution in the test set; Data augmentation unit: Adds uniform random noise in the range [-0.1, 0.1] to the standardized training set, expanding the original 62 sets of data to 6200 sets (100 times the original data), reducing the risk of overfitting of the small sample model.

[0010] Dual-model alternative risk assessment submodule It incorporates the XGBoost main model and the ResDNN residual deep network backup model, and performs hyperparameter optimization through grid search, using RMSE and MAE as model performance evaluation metrics. Optimal hyperparameters for XGBoost: objective=reg:squarederror, eval_metric=rmse, learning_rate=0.05, max_depth=3, subsample=0.5, colsample_bytree=0.9; The optimal hyperparameters for ResDNN are: residual block depth of 4, intermediate hidden dimension of 128, loss function MSELoss, Adam optimizer learning rate of 0.001, and built-in Dropout layer to suppress overfitting. Model selection logic: After 62-fold cross-validation, XGBoost has an average RMSE of 0.529 and an average MAE of 0.521; ResDNN has an average RMSE of 0.843 and an average MAE of 0.840. The system automatically selects XGBoost to complete the risk prediction, and a reserved interface is provided to support switching to ResDNN after the amount of data increases.

[0011] SHAP Risk Contribution Analysis Submodule Built-in LIME and SHAP dual interpretation algorithms, automatically select SHAP through robustness assessment to complete the quantification of indicator risk contribution; Robustness assessment process: Add a small perturbation of [-0.1, 0.1] to the standardized indicators of a single flight, and calculate the risk contribution value of each indicator before and after the perturbation using LIME and SHAP respectively. Use Spearman correlation coefficient to measure the stability of the interpretation results; the correlation coefficients of all dimensions of SHAP indicators are higher than those of LIME, indicating better robustness. SHAP calculation logic: Based on the game theory additive interpretation model, it quantifies the positive / negative contribution of 90 input indicators to the comprehensive risk value, supports custom low-risk benchmark values, and outputs risk contribution results independently in five dimensions.

[0012] Visualization output submodule By integrating XGBoost risk prediction results and SHAP risk contribution results, three types of visualization charts are generated: a bar chart comparing the actual value and model prediction value of a single flight indicator; a Sankey flow chart of secondary risk indicators; and a multi-dimensional SHAP waterfall chart, which intuitively displays the contribution size and positive and negative impact of individual indicators.

[0013] According to another aspect of the present invention, more specifically, a method for quantitatively assessing flight test safety risks based on XGBoost-SHAP, using the above-described system, includes the following steps: S1: Collect basic flight test data and construct a five-level hierarchical flight test risk indicator system. We collected 30 valid expert indicator scoring questionnaires and 62 full-process test flight data, covering 15 types of test flight subjects. The original 202 indicators were screened by experts to obtain 132 valid indicators, which were divided into 90 input indicators and 32 output indicators, and divided into five primary risk dimensions: human, machine, environment, ground, and test flight method.

[0014] S2: Dual-dimensional verification of indicator data quality Perform discrimination verification and input indicator independence verification on all indicators, remove abnormal indicators that do not meet the verification rules and do not conform to the real flight test scenario, and output a set of qualified indicators.

[0015] S3: Multi-mode standardized processing indicators Based on the four standardization methods of matching indicator business attributes—classification method, scaling method, hybrid method, and subjective method—a standardized risk score in the range of 0 to 10 is uniformly output; indicators without collected data are assigned a value of 0.

[0016] S4: Dataset Splitting and Small Sample Data Augmentation Choose either 8:2 simple splitting or 62-fold K-fold cross-splitting based on the usage scenario; after splitting, add uniform random noise of [-0.1, 0.1] to the training set to augment the data to 100 times the original data, thus completing the data augmentation.

[0017] S5: Dual-model grid hyperparameter search and optimal model selection Grid search optimization was performed on XGBoost and ResDNN respectively to determine the optimal hyperparameters of the two models. The average RMSE and MAE of the two models were calculated based on 62-fold K-fold cross-validation. The XGBoost with better performance was automatically selected as the risk prediction model, and the model training and storage were completed.

[0018] S6: Quantitative Prediction of Risks for Test Flights to be Evaluated The standardized input indicators for a single flight are input into the trained XGBoost model, which outputs the predicted value of the comprehensive flight test safety risk for that flight.

[0019] S7: Explaining the robustness comparison of algorithms, using SHAP to calculate the risk contribution indicator. A small perturbation is added to the current flight index, and the Spearman correlation coefficient of risk contribution before and after the LIME and SHAP perturbation is calculated. The more robust SHAP algorithm is selected to quantify the contribution of each input index to the overall risk and distinguish between positive risk-increasing indicators and negative risk-reducing indicators.

[0020] S8: Layered Visualization Output of Evaluation Results Generate indicator prediction comparison charts, secondary indicator Sankey diagrams, and waterfall charts showing the independent SHAP contributions of five dimensions, completing the entire process of flight test risk quantification and risk tracing.

[0021] S9: Model Iterative Update After adding new test flight data, repeat steps S2 to S5 to re-execute hyperparameter search and model training to achieve continuous optimization of the evaluation model.

[0022] The beneficial effects of the XGBoost-SHAP-based flight test safety risk quantification assessment system and method of this invention are as follows: (1) This invention completely solves the risk inundation problem of traditional analytic hierarchy process: Based on the nonlinear fitting capability of XGBoost, it automatically mines the risk of multi-level index coupling, and will not bury high-risk small-weight indicators. Adapting to small-sample test flight scenarios: The sample size is expanded by 100 times through random noise data augmentation in the range of [-0.1, 0.1]. Combined with Dropout and regularization terms to suppress model overfitting, the data from 62 historical test flights and 30 expert questionnaires are fully reused. High model accuracy and lightweight: XGBoost has an average RMSE of only 0.529, which is far superior to ResDNN. It has fast training and inference speeds and also reserves a ResDNN switching interface to adapt to future scenarios with large-scale flight test data. Stable and aligned with flight test operations: Compared to LIME, SHAP is more robust, allows for customizable risk benchmarks, and provides layered output of risk contributions across five dimensions: human, machine, environment, ground, and flight test methodology. It accurately identifies risk sources and supports flight test rectification. Full-process standardization and automation: Built-in indicator quality verification and four types of standardization rules eliminate the need for manual repetition of indicator weight setting. New test flight data can automatically iterate the model, reducing manual operation and maintenance costs. Covering all types of flight test subjects: The model training data includes 15 typical flight test scenarios such as high temperature, stall, strong crosswind, and avionics verification, which are suitable for the flight test safety assessment needs of various civil and military aircraft models. Attached Figure Description

[0023] The present invention will now be described in further detail with reference to the accompanying drawings and specific implementation methods.

[0024] Figure 1 This is a flowchart illustrating the overall selection process for the flight test risk assessment algorithm of this invention. Figure 2 This is a flowchart of the flight test index data preprocessing and enhancement process of the present invention; Figure 3 This is a flowchart of the grid search hyperparameter optimization process of the present invention; Figure 4 This is a flowchart illustrating the performance comparison of the 62-fold K-fold model of the present invention. Figure 5 This is a diagram of the XGBoost gradient boosting tree risk prediction model architecture of the present invention; Figure 6 This is a diagram of the alternative model architecture for the ResDNN residual deep neural network of the present invention; Figure 7 This is a graph showing the performance results of the ResDNN hyperparameter grid search of this invention; Figure 8 This is a heatmap comparing the XGBoost learning rate and tree depth hyperparameters of the present invention. Figure 9 This is a thermal comparison diagram of the hyperparameters of the XGBoost sampling parameters of the present invention; Figure 10 This is a bar chart comparing the average RMSE of the XGBoost and ResDNN models of this invention. Figure 11 This is a bar chart comparing the average MAE of the XGBoost and ResDNN models of this invention. Figure 12 This is a schematic diagram illustrating the principle of the LIME locally interpretable algorithm of the present invention; Figure 13 This is a schematic diagram of the SHAP game theory risk contribution analysis algorithm of the present invention; Figure 14 This is a comparison chart of the Spearman coefficients of the robustness of SHAP and LIME in this invention; Figure 15 This is a comparison chart of the actual value and the model prediction value of the indicator in Case 1 of the present invention; Figure 16 This is a comparison chart of the actual values ​​and model predictions of the indicators in Case 2 of this invention; Figure 17 This is a comparison chart of the actual values ​​and model predictions of the indicators in Case 3 of this invention; Figure 18 This is a Sankey flow diagram, representing the first and second level risk indicators of this invention. Figure 19This is the Sankey flow diagram, a secondary risk indicator in Case Two of the present invention. Figure 20 This is the Sankey flow diagram, a secondary risk indicator in Case 3 of the present invention. Figure 21 Example 1 of this invention is a SHAP waterfall diagram of the risk factors of the "human" crew. Figure 22 Example 1 of this invention: SHAP waterfall diagram of aircraft status risk factors; Figure 23 This is a SHAP waterfall diagram of the "ring" operating environment elements in Example 1 of the present invention; Figure 24 This is a SHAP waterfall diagram illustrating the risk factors of the flight test method in Case 1 of the present invention. Figure 25 This is the second example of the present invention: a SHAP waterfall diagram of the risk factors for the "human" crew. Figure 26 Example 2 of this invention: SHAP waterfall diagram of aircraft status risk factors; Figure 27 This is the SHAP waterfall diagram of the "ring" operating environment elements in Case 2 of the present invention; Figure 28 This is a SHAP waterfall diagram illustrating the risk factors of the flight test method in Case Study 2 of this invention. Figure 29 This is the SHAP waterfall diagram of the risk factors for the "human" crew in Case 3 of the present invention; Figure 30 Example 3 of this invention is the SHAP waterfall diagram of aircraft status risk factors. Figure 31 This is the SHAP waterfall diagram of the "ring" operating environment elements in Case 3 of the present invention; Figure 32 This is a SHAP waterfall diagram illustrating the risk factors of the test flight method in Case 3 of this invention. Detailed Implementation

[0025] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.

[0026] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] Example 1: Complete Training Process for Data Model from 62 Flight Tests Data collection: 30 valid expert indicator scoring questionnaires were collected, along with complete test flight data from 62 sorties, covering 15 test flight subjects; 202 original indicators were collected, and 132 valid indicators were obtained after expert screening, including 90 input indicators and 32 output indicators. Data quality verification: The discrimination and independence were verified, and the standard deviation and correlation of all indicators were verified. The friction effect and the low standard deviation of airport elevation were verified by experts and retained. The correlation coefficient of all input indicators was <1, and the data was qualified. Multi-mode standardization processing: Standardization is completed using classification, scaling, hybrid, and subjective methods according to the business type of the indicator, and a risk score of 0-10 is output uniformly; indicators with no collected data are assigned a value of 0; Dataset Segmentation and Data Augmentation: A 62-fold K-fold split was used, strictly followed by augmentation; uniform random noise of [-0.1, 0.1] was added to the training set to expand the original 62 sets of samples to 6200 sets, corresponding to the... Figure 2 ; Dual-model mesh hyperparameter search (corresponding appendix) Figure 3 , Figure 7 , Figure 8 , Figure 9 ) (1) XGBoost model and core algorithm formula XGBoost is a gradient boosting ensemble tree model with the following overall objective function:

[0028] In the formula, To predict the loss term, This is a tree complexity regularization term used to suppress overfitting in small test flight samples; Loss term expression:

[0029] Regularization term expression:

[0030] Parameter description: For L1 regularization terms, The L2 regularization coefficient; The number of tree nodes in a single decision tree. The weights are those of the leaf nodes; The predicted flight test risk values ​​are generated using a forward step-by-step addition iterative method.

[0031]

[0032]

[0033]

[0034] : No. After the first iteration Predicted risk values ​​for each test flight; No. Gradient boosting decision trees; XGBoost architecture is attached. Figure 5 .

[0035] XGBoost hyperparameter search range: learning_rate∈[0,0.05,0.1,0.15,0.2], max_depth∈[3,4,5,6,7], subsample∈[0.5,0.6,0.7,0.8,0.9], colsample_bytree∈[0.5,0.6,0.7,0.8,0.9]; Optimal parameters found through grid search: learning_rate=0.05, max_depth=3, subsample=0.5, colsample_bytree=0.9; Search results correspond to... Figure 8 Appendix Figure 9 .

[0036] (2) ResDNN residual deep network and core algorithm formula ResDNN is a backup deep learning model that incorporates residual connections and Dropout layers to address the issue of multi-level indicator feature loss during flight testing. The architecture is shown in the appendix. Figure 6 ; Basic residual calculation unit:

[0037] : 90 standardized flight test input risk indicators; Network weights and biases; Dropout layers are used to prevent overfitting in small-sample test flights; Formula for stacking multi-layer residual blocks:

[0038] Final risk output mapping:

[0039] ResDNN hyperparameter search range: block depth [3,4,5,6,7], hidden dimension [16,32,64,128]; optimal parameters: block depth 4, hidden dimension 128. See attached search results. Figure 7 .

[0040] A comparative analysis of the performance of 62-fold K-fold models (with appendix) Figure 4 , Figure 10 , Figure 11 ) Using RMSE and MAE as evaluation metrics, 62-fold cross-validation was used to compare model performance. XGBoost had an average RMSE of 0.529 and an average MAE of 0.521, while ResDNN had an average RMSE of 0.843 and an average MAE of 0.840. XGBoost's prediction accuracy was superior to ResDNN across the board, and XGBoost was selected as the main model for the system.

[0041] Explanation of algorithm robustness comparison (with appendix) Figure 12 , Figure 13 , Figure 14 ) The algorithm incorporates both LIME and SHAP interpretation algorithms, adds a small perturbation of [-0.1, 0.1] to the standardized indicators, and uses the Spearman correlation coefficient to measure the stability of risk contribution before and after the perturbation. The Spearman coefficients of each dimension of SHAP are higher than those of LIME, indicating better robustness. Therefore, SHAP was selected as the algorithm for calculating risk contribution.

[0042] SHAP core additive decomposition formula:

[0043] In the formula: : The overall risk prediction value for each flight output by XGBoost; Customize the low-risk baseline value for test flights; : Total number of standardized input risk indicators; : No. The SHAP risk contribution value of each indicator. This indicates an increased risk during flight testing. The representative indicator reduces the risk of flight testing; the principle of the SHAP algorithm is attached. Figure 13 The principles of LIME are detailed in the appendix. Figure 12 The robustness comparison results are attached. Figure 14 .

[0044] 4.5.2 Example 2: Quantitative Evaluation and Verification of Three Typical Flight Test Cases Three independent flight test samples (Case 1, Case 2, and Case 3) from the original documents were selected, and the complete evaluation method of this invention was executed: Input three sets of cases, standardize 90 input indicators, and feed them into the trained XGBoost model. Output a comprehensive risk prediction value; output a comparison chart of the actual and predicted values ​​of the indicators (attached). Figure 15 , 16 17); The SHAP algorithm is used to calculate the risk contribution of individual indicators across four dimensions: crew, aircraft, environment, and flight test methods, generating multi-dimensional SHAP waterfall charts (see attached). Figures 21-32 ); Generate a secondary indicator Sankey flow diagram to illustrate the risk transmission relationships at each level (attached). Figure 18 , 19 20); Verification results: The model's predicted values ​​and the actual values ​​of expert scores show a high degree of consistency, and the error is controllable; the SHAP waterfall chart can accurately identify the core risk indicators of each case group (such as lateral wind and captain fatigue in case 1, and incomplete configuration deviation in case 3). The risk tracing results are fully consistent with the experience of test flight operations, proving the practicality and accuracy of the present invention.

[0045] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A flight test safety risk quantification assessment system based on XGBoost-SHAP, characterized in that, It includes a flight test risk indicator collection submodule, an indicator data quality assessment submodule, a multi-mode indicator standardization submodule, a data preprocessing and enhancement submodule, a dual-model alternative risk assessment submodule, a SHAP risk contribution analysis submodule, and a visualization output submodule, which are connected sequentially. The flight test risk indicator collection submodule is used to collect expert questionnaire data and actual test data from 62 flight test sorties, covering 15 types of flight test subjects. A total of 202 indicators were originally collected, and 132 valid indicators were retained after expert screening. These include 90 input terminal risk indicators and 32 output risk indicators, which are divided into five primary risk dimensions: "human" crew risk factors, "aircraft" aircraft status risk factors, "environment" operating environment factors, ground support risk factors, and flight test method risk factors. The expert questionnaire data consists of 30 valid questionnaires scoring the importance and relevance of indicators. The indicator data quality assessment submodule is configured with dual rules for discrimination and independence verification. The discrimination verification requires the standard deviation of the indicator to be ≥0.5 and the extreme value range to match the actual flight test span. The independence verification only constrains the input indicators, and the correlation coefficient between the input indicators is <1. The multi-mode indicator standardization submodule is configured with four standardization methods: classification, scaling, hybrid, and subjective. The standardized output value range is uniformly [0, 10], and indicators without valid data collection are uniformly assigned a value of 0. The four standardization methods are applicable to the following scenarios: classification is used for discrete indicators; scaling is used for continuous numerical indicators; hybrid is used for segmented combination classification and scaling of the same indicator; and subjective is used for subjective risk indicators without objective quantitative values, where experts fill in a risk score in the range of 0 to 10. The data preprocessing and augmentation submodule includes a data segmentation unit and a data augmentation unit. The data segmentation unit supports two modes: 8:2 training / test set splitting and 62-fold K-fold cross-validation. The data augmentation unit adds uniform random noise of [-0.1, 0.1] to the training set, expanding the original data by 100 times. Data augmentation is performed after data segmentation to avoid contamination of the test set data. The dual-model alternative risk assessment submodule incorporates the XGBoost main model and the ResDNN residual deep network backup model, using RMSE and MAE as performance evaluation metrics. After grid hyperparameter search and 62-fold cross-validation, XGBoost achieved an average RMSE of 0.529 and an average MAE of 0.521, while ResDNN achieved an average RMSE of 0.843 and an average MAE of 0.

840. The system automatically selects XGBoost to complete the risk prediction. The optimal hyperparameter configuration for XGBoost is: objective=reg:squarederror, eval_metric=rmse, learning_rate=0.05, max_depth=3, subsample=0.5, colsample_bytree=0.9; The optimal hyperparameter configuration for the ResDNN backup model is as follows: residual block depth of 4, intermediate hidden dimension of 128, loss function of MSELoss, optimizer of Adam with learning rate of 0.001, and built-in Dropout layer to suppress overfitting. The SHAP risk contribution analysis submodule incorporates both LIME and SHAP dual interpretation algorithms. By adding a small perturbation to the indicators, the robustness comparison is completed using the Spearman correlation coefficient, and the more robust SHAP algorithm is selected to quantify the positive / negative risk contribution of each indicator. It supports custom low-risk benchmark values ​​and outputs risk contribution results independently across five primary risk dimensions. The visualization output submodule is used to generate a comparison chart of the actual and predicted values ​​of indicators, a Sankey diagram of secondary risk indicators, and a multi-dimensional SHAP risk contribution waterfall chart.

2. The flight test safety risk quantification assessment system based on XGBoost-SHAP according to claim 1, characterized in that, The XGBoost model uses an objective function. Complete risk fitting and overfit suppression, where the loss term Regularization term ; Using forward step-by-step iterative formula Iteratively generate flight test risk prediction values.

3. The flight test safety risk quantification assessment system based on XGBoost-SHAP according to claim 1, characterized in that, The basic unit of the ResDNN residual network is Multiple layers stacked as Ultimate risk output .

4. The flight test safety risk quantification assessment system based on XGBoost-SHAP according to claim 1, characterized in that, The SHAP algorithm uses the additive decomposition formula. Quantify the risk contribution of individual indicators. This is a comprehensive risk forecast value. As a low-risk benchmark, To input the total number of indicators, The contribution value for the i-th indicator.

5. The flight test safety risk quantification assessment system based on XGBoost-SHAP according to claim 1, characterized in that, The original test flight samples totaled 62 groups, which were expanded to 6200 groups after data augmentation for model training.

6. The flight test safety risk quantification assessment system based on XGBoost-SHAP according to claim 1, characterized in that, The 15 categories of flight test subjects include high-temperature flight test, delivery flight test, engine takeoff without air, rear center of gravity expansion, production flight test, RVSM flight test, night landing light flight test, crosswind flight test, stall flight test, engine operating characteristics, route demonstration flight, VMCA, takeoff performance, new avionics version verification, and climb performance flight test.

7. The flight test safety risk quantification assessment system based on XGBoost-SHAP according to claim 1, characterized in that, In the discrimination test, friction effects, airport properties, complex terrain, border terrain, airport elevation, and complex weather indicators with a standard deviation of less than 0.5 are deemed qualified if they conform to the actual test flight distribution as verified by experts.

8. A method for quantitatively assessing flight test safety risks based on XGBoost-SHAP, characterized in that, Using the XGBoost-SHAP-based flight test safety risk quantification assessment system as described in any one of claims 1-7, the following steps are included: S1: Collect basic flight test data and construct a five-level hierarchical flight test risk indicator system; S2: Dual-dimensional verification of indicator data quality, outputting a set of qualified indicators; S3: Use the matching standardization method to standardize the indicators and output a risk score in the range of 0 to 10. S4: Split the dataset and add random noise of [-0.1, 0.1] to the training set to complete data augmentation; S5: Grid search is used to obtain the optimal hyperparameters of XGBoost and ResDNN, and 62-fold K-fold cross-validation is used to select XGBoost as the optimal risk prediction model and train and store it. S6: Input the standardized indicators of the flights to be evaluated into the XGBoost model and output the comprehensive flight test safety risk prediction value; S7: Add a small perturbation to the indicators, compare the Spearman correlation coefficients of LIME and SHAP, and use the SHAP algorithm to quantify the risk contribution of each input indicator; S8: Layered visualization output of indicator prediction comparison chart, Sankey diagram, and multi-dimensional SHAP waterfall chart; S9: After adding new test flight data, repeat S2~S5 to iteratively update the risk prediction model.

9. The method for quantitative assessment of flight test safety risks based on XGBoost-SHAP according to claim 8, characterized in that, In step S4, data augmentation expands the original 62 test flight samples to 6200 samples.

10. The method for quantitative assessment of flight test safety risks based on XGBoost-SHAP according to claim 8, characterized in that, In step S7, SHAP distinguishes between indicators that contribute positively to risk increase and those that contribute negatively to risk decrease, thus identifying the core sources of risk in flight testing.