Hierarchical prediction method for abnormal value processing and interpretability analysis of mine disaster data
The outliers and SHAP models were screened through the Marshall distance discrimination method for interpretability analysis, and combined with the deep forest model of the fusion algorithm, the problems of mining disaster data processing and prediction accuracy were solved, achieving higher prediction accuracy and model interpretability.
Patent Information
- Application Number
- CN202510073038.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
Due to the small sample and high-dimensional characteristics of mine disaster data, it is difficult for the existing technology to effectively deal with outliers, resulting in a decrease in the prediction accuracy of machine learning algorithms and a lack of interpretability analysis.
The Mahayana distance discrimination method was used to filter the mine disaster data, delete outliers, and interpretability analysis was performed through the SHAP model. Fusion of CatBoost, LightGBM, and XGBoost algorithms into the deep forest algorithm to form CGCF, LGCF, and XGCF algorithms to improve prediction accuracy and general applicability.
In the case of a small data set, more data information is retained, the accuracy of the prediction algorithm is improved, and the interpretability of the model is enhanced through the SHAP model.
Smart Images

Figure CN120013277A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graded prediction of mine disasters, and in particular relates to a graded prediction method for outlier processing and interpretability analysis of mine disaster data. Background Art
[0002] Traditional outlier screening, such as the isolation forest algorithm, is essentially an unsupervised anomaly detection method based on a tree model. Its core idea is to construct multiple binary trees by randomly selecting features and split points to gradually isolate data points. Since outliers are usually located in sparse areas of data distribution, their average isolation paths in the tree are short. Therefore, by evaluating the average path length of multiple trees, the isolation forest can effectively distinguish outliers from normal points. However, this modeling algorithm often screens the entire data. When the data is determined to be an outlier, the data will be deleted as a whole. The premise of the machine learning algorithm is to have sufficient high-quality data. The particularity of mine disasters means that its data has the characteristics of small samples and multiple dimensions. Therefore, mine disaster-related data is particularly precious. In addition, due to the harsh data mining environment in mines, the stability of the equipment for collecting data is easily affected, and the collected data is prone to have many outliers. The existence of outliers will interfere with the prediction of the machine learning algorithm. Therefore, it is necessary to reasonably screen and process the data for outliers to maximize the retention of the information stored in the data to improve the accuracy of the machine learning algorithm.
[0003] From a data perspective, most of the current mine disaster data are still small sample high-dimensional data, so the problem of graded prediction of mine disasters still requires high-performance algorithms. Therefore, this patent uses the machine learning method of integrated learning to improve the deep forest algorithm.
[0004] In addition, the current artificial intelligence algorithms still have a serious "black box" problem, that is, most scholars are limited to using artificial intelligence algorithms to predict or classify disasters, but are unable to well match physical parameters or disaster mechanisms with prediction or classification results. Summary of the invention
[0005] The purpose of the present invention is to provide a hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data, which is suitable for situations where the data set is small. Compared with the conventional isolation forest algorithm, it can retain more information retained in the data and improve the accuracy of the prediction algorithm.
[0006] In order to solve the above technical problems, the technical solution of the present invention is: a hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data, comprising the following steps:
[0007] The mine disaster data are screened for a preset round by using the Mahalanobis distance discrimination method to obtain the Mahalanobis distance of each mine disaster data and the abnormal threshold of each mine disaster data is calculated according to the distribution of its Mahalanobis distance. The data of each mine disaster data whose Mahalanobis distance is greater than the abnormal threshold is identified as an abnormal quantity. The mine disaster data whose abnormal quantity is less than or equal to one-half of its corresponding data quantity are deleted, and the deleted abnormal quantity is replaced by taking the median of the remaining data; the mine disaster data whose abnormal quantity is greater than one-half of its corresponding data quantity are deleted;
[0008] The interpretability analysis of each mine disaster data after deleting the abnormal amount was carried out through the SHAP model to obtain the average absolute value of the SHAP value of each mine disaster data, and the relative importance of each mine disaster data was obtained based on its average absolute value.
[0009] The method also includes the following steps: inputting the mine disaster data after deleting the abnormal amount into a preset disaster prediction algorithm model, training the model, and performing algorithm performance analysis on the data set obtained by processing the trained model through the Mahalanobis distance discriminant method.
[0010] The workflow of the disaster prediction algorithm model is:
[0011] Perform a preset preprocessing operation on the input data through multi-granularity scanning to obtain a feature vector;
[0012] The feature vector is trained through the cascade forest module;
[0013] Specifically, the cascade forest module is as follows: the deep forest algorithm adopts the idea of layer-by-layer processing to form a cascade structure, and each level is composed of multiple random forests; the random forest learns the input feature vector information at each level and passes it to the next level after processing; in order to improve the generalization ability, each level uses a completely random forest and a random forest as the base learner; the cascade forest module uses cross-validation to adaptively adjust the number of cascade layers.
[0014] The deep forest algorithm selects the extreme gradient boosting algorithm XGBoost, the classification boosting algorithm CatBoost and the light gradient boosting machine algorithm LightGBM as the basic classifier integration, and integrates and optimizes to form the extreme gradient boosting and deep forest algorithm XGCF, the classification boosting and deep forest algorithm CGCF and the light gradient boosting and deep forest algorithm LGCF.
[0015] A hierarchical prediction system for outlier processing and interpretability analysis of mine disaster data is also provided, including:
[0016] The mine disaster data abnormal value processing module is used to screen the mine disaster data for preset rounds through the Mahalanobis distance discrimination method, obtain the Mahalanobis distance of each mine disaster data and calculate the abnormal threshold of each mine disaster data according to the distribution of its Mahalanobis distance, identify the data of each mine disaster data whose Mahalanobis distance is greater than the abnormal threshold as abnormal quantity, delete the abnormal quantity of the mine disaster data whose abnormal quantity is less than or equal to one-half of its corresponding data quantity, and replace the deleted abnormal quantity with the median of its remaining data; delete the mine disaster data whose abnormal quantity is greater than one-half of its corresponding data quantity;
[0017] The mine disaster index interpretability analysis module is used to perform interpretability analysis on each mine disaster data after deleting the abnormal amount through the SHAP model, obtain the average absolute value of the SHAP value of each mine disaster data, and obtain the relative importance of each mine disaster data based on its average absolute value.
[0018] It also includes: a disaster prediction algorithm module, which is used to input the mine disaster data after deleting the abnormal amount into a preset disaster prediction algorithm model, train the model, and perform algorithm performance analysis on the data set obtained by processing the trained model through the Mahalanobis distance discriminant method.
[0019] The workflow of the disaster prediction algorithm model is:
[0020] Perform a preset preprocessing operation on the input data through multi-granularity scanning to obtain a feature vector;
[0021] The feature vector is trained through the cascade forest module;
[0022] Specifically, the cascade forest module is as follows: the deep forest algorithm adopts the idea of layer-by-layer processing to form a cascade structure, and each level is composed of multiple random forests; the random forest learns the input feature vector information at each level and passes it to the next level after processing; in order to improve the generalization ability, each level uses a completely random forest and a random forest as the base learner; the cascade forest module uses cross-validation to adaptively adjust the number of cascade layers.
[0023] The deep forest algorithm selects the extreme gradient boosting algorithm XGBoost, the classification boosting algorithm CatBoost and the light gradient boosting machine algorithm LightGBM as the basic classifier integration, and integrates and optimizes to form the extreme gradient boosting and deep forest algorithm XGCF, the classification boosting and deep forest algorithm CGCF and the light gradient boosting and deep forest algorithm LGCF.
[0024] A computer device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the hierarchical prediction method described in any one of the above items are implemented.
[0025] A computer-readable storage medium is also provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the hierarchical prediction method as described in any one of the above items are implemented.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] The present invention screens mine disaster data through the Mahalanobis distance discriminant method. When the data set is small, compared with the conventional isolation forest algorithm, it can retain more information retained in the data and improve the accuracy of the prediction algorithm; CatBoost, LightGBM, and XGBoost algorithms are integrated into the deep forest algorithm to form CGCF, LGCF, and XGCF algorithms, which improves the versatility of the present invention in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of the workflow of each module in the hierarchical prediction system for outlier processing and interpretability analysis of mine disaster data in an embodiment of the present invention;
[0029] FIG2( a ) is a schematic diagram of calculating the X1 feature threshold by the Mahalanobis distance discriminant method in an embodiment of the present invention;
[0030] FIG2( b ) is a schematic diagram of calculating the X2 feature threshold by the Mahalanobis distance discriminant method in an embodiment of the present invention;
[0031] FIG2( c ) is a schematic diagram of calculating the X3 feature threshold by the Mahalanobis distance discrimination method in an embodiment of the present invention;
[0032] FIG2(d) is a schematic diagram of calculating the X4 feature threshold by the Mahalanobis distance discrimination method in an embodiment of the present invention;
[0033] FIG2(e) is a schematic diagram of calculating the X5 feature threshold by the Mahalanobis distance discrimination method in an embodiment of the present invention;
[0034] FIG2( f ) is a schematic diagram of calculating the X6 feature threshold by the Mahalanobis distance discrimination method in an embodiment of the present invention;
[0035] Figure 3 Schematic diagram of the principle of the deep forest algorithm in an embodiment of the present invention;
[0036] FIG4( a ) is a schematic diagram of the ROC curve of the XGBoost model in an embodiment of the present invention;
[0037] FIG4( b ) is a schematic diagram of the ROC curve of the XGCF model in an embodiment of the present invention;
[0038] FIG4( c ) is a schematic diagram of the ROC curve of the CatBoost model in an embodiment of the present invention;
[0039] FIG4( d ) is a schematic diagram of the ROC curve of the CGCF model in an embodiment of the present invention;
[0040] Figure 5 Schematic diagram of the principle of the SHAP model in an embodiment of the present invention;
[0041] FIG6( a ) is a SHAP model analysis diagram of slope sample 1 in an embodiment of the present invention;
[0042] FIG6( b ) is a SHAP model analysis diagram of slope sample 2 in an embodiment of the present invention;
[0043] FIG6( c ) is a SHAP model analysis diagram of slope sample 3 in an embodiment of the present invention;
[0044] FIG6( d ) is a SHAP model analysis diagram of slope sample 4 according to an embodiment of the present invention;
[0045] Figure 7 It is a SHAP feature analysis diagram after summarizing all sample features in the embodiment of the present invention;
[0046] Figure 8 This is a schematic diagram of SHAP average absolute value sorting after the features of all samples are summarized in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0048] The technical solution of the present invention is: a hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data, comprising the following steps:
[0049] The mine disaster data are screened for a preset number of rounds (3 in this embodiment) by using the Mahalanobis distance discrimination method to obtain the Mahalanobis distance of each mine disaster data and calculate the abnormal threshold of each mine disaster data according to the distribution of its Mahalanobis distance. The data of each mine disaster data whose Mahalanobis distance is greater than the abnormal threshold is identified as an abnormal amount. The mine disaster data whose abnormal amount is less than or equal to one-half of its corresponding data amount is deleted, and the deleted abnormal amount is replaced by taking the median of its remaining data; the mine disaster data whose abnormal amount is greater than one-half of its corresponding data amount is deleted;
[0050] The interpretability analysis of each mine disaster data after deleting the abnormal amount was carried out through the SHAP model to obtain the average absolute value of the SHAP value of each mine disaster data, and the relative importance of each mine disaster data was obtained based on its average absolute value.
[0051] The method also includes the following steps: inputting the mine disaster data after deleting the abnormal amount into a preset disaster prediction algorithm model, training the model, and performing algorithm performance analysis on the data set obtained by processing the trained model through the Mahalanobis distance discriminant method.
[0052] The workflow of the disaster prediction algorithm model is:
[0053] Perform a preset preprocessing operation on the input data through multi-granularity scanning to obtain a feature vector;
[0054] The feature vector is trained through the cascade forest module;
[0055] Specifically, the cascade forest module is as follows: the deep forest algorithm adopts the idea of layer-by-layer processing to form a cascade structure, and each level is composed of multiple random forests; the random forest learns the input feature vector information at each level and passes it to the next level after processing; in order to improve the generalization ability, each level uses a completely random forest and a random forest as the base learner; the cascade forest module uses cross-validation to adaptively adjust the number of cascade layers.
[0056] The deep forest algorithm selects the extreme gradient boosting algorithm XGBoost, the classification boosting algorithm CatBoost and the light gradient boosting machine algorithm LightGBM as the basic classifier integration, and integrates and optimizes to form the extreme gradient boosting and deep forest algorithm XGCF, the classification boosting and deep forest algorithm CGCF and the light gradient boosting and deep forest algorithm LGCF.
[0057] It also provides a hierarchical prediction system for outlier processing and interpretability analysis of mine disaster data, such as Figure 1 As shown, including:
[0058] Mine disaster data outlier processing module
[0059] In order to properly handle the outliers in mine disaster data, the present invention uses the Mahalanobis distance method to screen the data. The Mahalanobis distance (MD) represents the covariance distance of the data. It is an effective method for calculating the similarity of two unknown sample sets and is suitable for anomaly detection. Unlike the Euclidean distance, it takes into account the connection between various characteristics and is scale-independent, that is, independent of the measurement scale. In summary, the Mahalanobis distance can be independent of the measurement unit and comprehensively reflect the correlation between multiple variables.
[0060]
[0061] For a multivariate vector with a mean of μ and a covariance matrix of Σ, the Mahalanobis distance in n-dimensional space is calculated as shown in Formula 1, where d is the Mahalanobis distance between each individual sample Y and the average sample μ, and T represents the transposition. Using the Mahalanobis distance method, a regression relationship can be established between a single indicator and the remaining indicators, and the influence of each indicator on the indicator can be comprehensively considered to determine whether there are outliers in the indicator.
[0062] Taking the slope data of the mine as an example, for the six features X1 to X6, their overall Mahalanobis distances are calculated respectively. The frequency distribution of the Mahalanobis distance of each feature is as follows: Figure 2(a)-Figure 2(f) As shown, the confidence interval is set to 85% (the closer the confidence interval is to 100%, the more outliers there are in the data set). The abnormal thresholds of the X1-X6 features are calculated to be 1.73, 0.92, 1.58, 1.80, 1.05, and 1.24, that is, feature data greater than these values will be determined as abnormal, and then the abnormal feature data will be replaced with the median of the feature data in which the abnormal feature is found. Unlike the traditional method of directly deleting the entire set of abnormal data, the method used in the present invention only targets some of the internal features of the data, replaces or deletes the abnormal features of the data, retains some of the information of the data, and makes it easier for the algorithm to understand and learn the data set.
[0063] The slope database of the present invention includes six features. When using Mahalanobis distance for discrimination, for data with an abnormal feature number less than or equal to 2, median replacement is performed after screening the abnormal features, and for features with an abnormal feature number greater than 2, that is, the abnormal amount in the data is greater than 1 / 2, it is deleted from the process.
[0064] The accuracy data of different algorithms in each round of screening records are shown in Table 1. IF1-IF5 refers to the data set after 1-5 rounds of isolation forest algorithm screening, and MD1 represents the remaining data set after the screening of the data set using the above-mentioned Mahalanobis distance screening. Taking the CGCF algorithm as an example, it performs best on the data after three rounds of isolation forest screening, with an accuracy rate of up to 0.93. However, in the fourth and fifth rounds of screening, the accuracy rate dropped to 0.91 and 0.91 respectively. This shows that the first and second rounds of screening cannot completely screen out the outliers in the data, and in the fourth and fifth rounds of screening, the isolation forest algorithm may "misjudge" some normal data and remove them. After adopting the Mahalanobis distance discriminant method, the accuracy of the CGCF algorithm can reach 0.94. In addition, the screening method of outliers in the isolation forest is more "rough" than the Mahalanobis distance discriminant method used in the present invention. Almost every round of screening requires the deletion of at least 15 groups of data. After the most ideal third round of screening, only 316 groups of data are left. After the third round of screening, there are 316 data sets left in the data set, which is 15 less than the 331 data sets left by the Mahalanobis distance discriminant method. Therefore, under the same effect, the Mahalanobis distance discriminant method can retain more data information.
[0065] Table 1
[0066] LightGBM XGBoost CatBoost GCF XGD LGCF CGCF RF SVM Numberofstoreddata IF1 0.87 0.87 0.86 0.87 0.88 0.88 0.88 0.84 0.76 351 IF2 0.9 0.89 0.89 0.9 0.9 0.9 0.91 0.86 0.76 333 IF3 0.87 0.88 0.93 0.91 0.93 0.93 0.93 0.88 0.78 316 IF4 0.92 0.93 0.91 0.9 0.91 0.91 0.91 0.86 0.76 300 IF5 0.91 0.91 0.91 0.87 0.91 0.9 0.91 0.86 0.76 285 MD1 0.91 0.94 0.94 0.91 0.94 0.92 0.94 0.89 0.78 331
[0067] Disaster prediction algorithm module
[0068] Deep Forest is a new tree-based model that is comparable to deep neural networks and was proposed by Professor Zhou Zhihua and Dr. Feng Ji in their paper Deep Forest: Towards An Alternative to Deep Neural Networks published on February 28, 2017. Figure 3 As shown in the figure. This algorithm incorporates the three major advantages of traditional deep learning: layer-by-layer processing, feature conversion within the model, and sufficient model complexity. At the same time, the complexity of the deep forest algorithm will automatically adjust with the size of the data set, and it is not sensitive to the adjustment of hyperparameters, effectively avoiding the shortcomings of deep models requiring too many adjustment parameters and overly complex models.
[0069] (1) Cascade Forest
[0070] The Deep Forest algorithm adopts the idea of layer-by-layer processing to form a cascade structure, and each level consists of multiple random forests. The random forest learns the input feature vector information at each level and passes it to the next layer after processing. To improve the generalization ability, each level uses a completely random forest and a random forest as the base learner. At the same time, the cascade forest module uses cross-validation to adaptively adjust the number of cascade layers.
[0071] (2) Multi-granularity scanning
[0072] By setting the sliding window width to s, this study reduces the M-dimensional original feature vector to (Ms)+1. Subsequently, the features are further transformed through the random forest and completely random forest algorithms to extract the class distribution vector. These vectors are concatenated to construct the ultimate enhanced feature vector for subsequent data analysis and pattern recognition.
[0073] Therefore, the overall process of the deep forest algorithm is as follows Figure 3 As shown in the figure, it is divided into the following two steps: (1) Use multi-granularity scanning to preprocess the input features; (2) Send the obtained feature vector to the cascade forest for training.
[0074] In order to enhance the deep forest model, extreme gradient boosting (XGBoost), classification boosting (CatBoost) and light gradient boosting machine (LightGBM) are integrated as basic classifiers to form extreme gradient boosting and deep forest (XGCF), classification boosting and deep forest (CGCF), and light gradient boosting and deep forest (LGCF) algorithms. The accuracy, precision, and F1 value evaluation indicators are used to verify the model, as shown in Table 2. The accuracy of XGBoost, CatBoost, XGCF, and CGCF algorithms is 0.94,
[0075] In order to further characterize the four algorithms with the highest accuracy, a five-fold cross-validation was performed on the mine slope data, and the area under the curve (AUC) value was calculated. It can be seen that the CGCF algorithm AUC is 0.96+0.02, which is the most stable among the above four algorithms and is the optimal algorithm.
[0076] Table 2
[0077]
[0078] Mine disaster indicator interpretability analysis module
[0079] SHAP is an additive explanatory model inspired by cooperative game theory. Its core is to calculate the Shap value of each feature to reflect the contribution of the feature to the predictive ability of the entire model. The principle diagram is as follows: Figure 5. Feature Importance in traditional machine learning can intuitively reflect the importance of features, but its implementation process is more like a "black box" and it is impossible to judge the relationship between features and the final prediction results. The SHAP value can reflect the impact of each feature on the final prediction value and can show the positive and negative impact, which increases the interpretability of the model. Shap interprets the model's prediction value as the sum of the attribution values (Shap Values) of each input feature, that is:
[0080]
[0081] In the formula, is the model prediction value, f i is the attribute value corresponding to each feature, and f0 is the predicted mean of all training samples.
[0082] In this embodiment, a slope is taken as an example to analyze the case in detail, and the Shap model is used to calculate the Shap value in the data set. In order to clearly illustrate the working mechanism of the Shap model, some slope cases in the test set are selected for separate display, such as Figure 6(a)-Figure 6(d) , where base value = 1.524 is the basic standard value of the data set. Values above this value will be judged as stable, and values below this standard value will be judged as unstable. Therefore, the samples corresponding to Figure 6(a), Figure 6(b), and Figure 6(c) are judged as unstable slopes, while the sample corresponding to Figure 6(d) has a value of 2.01 and is judged as stable.
[0083] The six characteristics that affect slope stability and their corresponding Shap values are plotted in the form of a scatter plot as shown below: Figure 7 , the color in the figure represents the size of the corresponding eigenvalue. The closer the color is to red, the larger the corresponding eigenvalue is, and the closer it is to blue, the smaller the eigenvalue is. The ordinate of the figure represents the corresponding feature name, and the abscissa represents the Shap value corresponding to the feature point. A positive Shap value is equivalent to a positive contribution to the increase in the prediction result. In the present invention, a positive Shap value means a positive contribution to the slope stability. Figure 7 It can be seen that the larger the corresponding values of X1 bulk density and X3 internal friction angle, the more stable the slope is, while the larger the corresponding values of X4 slope angle and X6 pore pressure ratio, the easier the slope is to become unstable. This conclusion is consistent with the traditional slope theory and enhances the interpretability and credibility of the model.
[0084] Then, the mean absolute value of the Shap value for each feature is calculated and plotted as follows Figure 8 .according to Figure 8 The results show that the importance of the six characteristics on slope stability is ranked as follows: X1 bulk density, X5 slope height, X2 cohesion, X6 pore pressure ratio, X4 slope angle, and X3 internal friction angle.
[0085] (1) This patent invents a method for processing mine disaster outliers based on Mahalanobis distance. After comparative studies, this method is suitable for situations where the data set is small. Compared with the conventional isolation forest algorithm, it can retain more information in the data and improve the accuracy of the prediction algorithm.
[0086] (2) This study integrates CatBoost, LightGBM, and XGBoost algorithms into the deep forest algorithm to form CGCF, LGCF, and XGCF algorithms. The CGCF algorithm successfully predicts slope stability with an identification accuracy of 0.94 and an AUC value of 0.96±0.02. This shows that the CGCF method is very suitable for slope stability prediction and has high practical value in engineering applications.
[0087] (3) The SHAP model was used to analyze the six key features that affect slope stability and determine the impact of each feature on the final prediction. The study found that the greater the unit deadweight and internal friction angle, the better the slope stability, while the greater the slope angle and pore pressure ratio, the greater the possibility of slope failure. These findings are consistent with rock mechanics theory and enhance the interpretability and credibility of the machine learning model. In addition, the SHAP model analysis ranked the importance of the six features that affect slope stability as follows: X1 (unit weight), X5 (slope height), X2 (cohesion), X6 (pore pressure ratio), X4 (slope angle), and X3 (internal friction angle). This ranking can provide a reference for formulating relevant safety measures in actual engineering practice.
[0088] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data, characterized in that: The following steps are involved: The mine disaster data are screened for a preset round by using the Mahalanobis distance discrimination method to obtain the Mahalanobis distance of each mine disaster data and the abnormal threshold of each mine disaster data is calculated according to the distribution of its Mahalanobis distance. The data of each mine disaster data whose Mahalanobis distance is greater than the abnormal threshold is identified as an abnormal quantity. The mine disaster data whose abnormal quantity is less than or equal to one-half of its corresponding data quantity are deleted, and the deleted abnormal quantity is replaced by taking the median of the remaining data; the mine disaster data whose abnormal quantity is greater than one-half of its corresponding data quantity are deleted; The interpretability analysis of each mine disaster data after deleting the abnormal amount was carried out through the SHAP model to obtain the average absolute value of the SHAP value of each mine disaster data, and the relative importance of each mine disaster data was obtained based on its average absolute value.
2. The hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data according to claim 1 is characterized in that: The following steps are also included: The mine disaster data after deleting the abnormal data are input into the preset disaster prediction algorithm model, the model is trained, and the algorithm performance analysis is performed on the data set obtained by processing the trained model using the Mahalanobis distance discriminant method.
3. The hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data according to claim 2 is characterized in that: The workflow of the disaster prediction algorithm model is: Perform a preset preprocessing operation on the input data through multi-granularity scanning to obtain a feature vector; The feature vector is trained through the cascade forest module; Specifically, the cascade forest module is as follows: the deep forest algorithm adopts the idea of layer-by-layer processing to form a cascade structure, and each level is composed of multiple random forests; the random forest learns the input feature vector information at each level and passes it to the next level after processing; in order to improve the generalization ability, each level uses a completely random forest and a random forest as the base learner; the cascade forest module uses cross-validation to adaptively adjust the number of cascade layers.
4. The hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data according to claim 3 is characterized in that: The deep forest algorithm selects the extreme gradient boosting algorithm XGBoost, the classification boosting algorithm CatBoost and the light gradient boosting machine algorithm LightGBM as the basic classifier integration, and integrates and optimizes to form the extreme gradient boosting and deep forest algorithm XGCF, the classification boosting and deep forest algorithm CGCF and the light gradient boosting and deep forest algorithm LGCF.
5. A hierarchical prediction system using the hierarchical prediction method for outlier processing and interpretability analysis of mine disaster data as claimed in claim 1, characterized in that: include: The mine disaster data abnormal value processing module is used to screen the mine disaster data for preset rounds through the Mahalanobis distance discrimination method, obtain the Mahalanobis distance of each mine disaster data and calculate the abnormal threshold of each mine disaster data according to the distribution of its Mahalanobis distance, identify the data of each mine disaster data whose Mahalanobis distance is greater than the abnormal threshold as abnormal quantity, delete the abnormal quantity of the mine disaster data whose abnormal quantity is less than or equal to one-half of its corresponding data quantity, and replace the deleted abnormal quantity with the median of its remaining data; delete the mine disaster data whose abnormal quantity is greater than one-half of its corresponding data quantity; The mine disaster index interpretability analysis module is used to perform interpretability analysis on each mine disaster data after deleting the abnormal amount through the SHAP model, obtain the average absolute value of the SHAP value of each mine disaster data, and obtain the relative importance of each mine disaster data based on its average absolute value.
6. The classification prediction system according to claim 5, characterized in that: Also includes: The disaster prediction algorithm module is used to input the mine disaster data after deleting the abnormal amount into the preset disaster prediction algorithm model, train the model, and perform algorithm performance analysis on the data set obtained by processing the trained model through the Mahalanobis distance discriminant method.
7. The classification prediction system according to claim 6, characterized in that: The workflow of the disaster prediction algorithm model is: Perform a preset preprocessing operation on the input data through multi-granularity scanning to obtain a feature vector; The feature vector is trained through the cascade forest module; Specifically, the cascade forest module is as follows: the deep forest algorithm adopts the idea of layer-by-layer processing to form a cascade structure, and each level is composed of multiple random forests; the random forest learns the input feature vector information at each level and passes it to the next level after processing; in order to improve the generalization ability, each level uses a completely random forest and a random forest as the base learner; the cascade forest module uses cross-validation to adaptively adjust the number of cascade layers.
8. The classification prediction system according to claim 7, characterized in that: The deep forest algorithm selects the extreme gradient boosting algorithm XGBoost, the classification boosting algorithm CatBoost and the light gradient boosting machine algorithm LightGBM as the basic classifier integration, and integrates and optimizes to form the extreme gradient boosting and deep forest algorithm XGCF, the classification boosting and deep forest algorithm CGCF and the light gradient boosting and deep forest algorithm LGCF.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the hierarchical prediction method according to any one of claims 1 to 4 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the hierarchical prediction method according to any one of claims 1 to 4 are implemented.