Multi-disaster grading prediction method based on artificial intelligence

Through the multi-disaster hierarchical prediction method based on artificial intelligence, the mining disaster data is screened and dimensionalized, and the problems of redundant and unbalanced multi-disaster prediction systems in the existing technology are solved, significantly improving the accuracy and sensitivity of prediction.

CN119988932AActive Publication Date: 2025-05-13WUHAN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510014722.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Existing mine disaster prediction methods are difficult to effectively deal with multiple disaster situations, resulting in complex systems, low response sensitivity, uneven data and low quality problems.

Method used

Using a multi-disaster hierarchical prediction method based on artificial intelligence, we use filtering and dimensionality reduction of mine disaster data, build a balanced data set, and use machine learning and deep learning algorithms to perform hierarchical prediction. Specific steps include data collection, indicator filtering, data set reconstruction, algorithm selection and model verification.

Benefits of technology

It significantly improves the imbalance of the mine disaster data set, improves the accuracy of prediction, and achieves the sensitivity and accuracy of multi-hazard hierarchical prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988932A_ABST
    Figure CN119988932A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-disaster grading prediction method based on artificial intelligence, and the method comprises the following steps: collecting mine disaster data, and forming a mine disaster database; constructing a mine disaster prediction index system; constructing a mine multi-disaster prediction database; evaluation characteristics are selected according to indexes in the mine disaster prediction index system, grading prediction is carried out on data in the mine multi-disaster prediction database according to all algorithms in a preset mine disaster prediction algorithm library, and prediction results of grading prediction of all the algorithms are summarized and compared under different preset evaluation indexes. Obtaining an algorithm with the optimal score of the comprehensive evaluation index of the prediction result, and taking the algorithm as an algorithm of a mine disaster prediction optimal model; and inputting the mine disaster data into the mine disaster prediction optimal model, and performing multi-disaster grading prediction. According to the method, the imbalance of a mine disaster data set is remarkably improved, data dimension reduction is realized, and the prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mine disaster classification prediction, and in particular relates to a multi-disaster classification prediction method based on artificial intelligence. Background Art

[0002] The geological conditions of many mines in my country are complex. During the mining process, they are inevitably threatened by disasters such as water disasters, fire, dust disasters, roof disasters, and slope instability. Deep mining areas are also plagued by various disasters such as rock bursts, pillar instability, and goaf instability. These disasters not only delay construction schedules and cause huge economic losses, but also threaten the lives and safety of construction workers, bringing severe challenges to the operation and safety of underground projects.

[0003] Mine disaster prediction is a difficult problem that needs to be solved urgently. However, the mechanisms of many mine disasters are complex and have not yet been unified. Its prediction is a complex nonlinear problem. It is the result of the combined action of multiple factors. It has randomness, fuzziness and uncertainty. No mathematical or mechanical method can fully and accurately describe it. Therefore, it is necessary to explore how to improve the existing prediction methods and improve the accuracy of mine disaster prediction, so as to provide a scientific basis for the safety protection and reasonable construction of mine mining in my country. In addition, due to the particularity of mine disasters and their relatively harsh data collection environment, they will have problems such as low data quality and difficulty in selecting evaluation indicators. Therefore, it is meaningful to study how to improve the quality of mine disaster data sets and the optimization of indicators.

[0004] There is not just a single disaster in a mine. If one disaster corresponds to a set of algorithms, it will lead to the complexity of the smart mine system, reduce its response sensitivity, greatly increase the cost of software and hardware, and cause waste of resources and difficulties in use and management. Summary of the invention

[0005] The purpose of the present invention is to provide a multi-disaster graded prediction method based on artificial intelligence, which significantly improves the imbalance of mine disaster data sets, realizes data dimensionality reduction, and improves the accuracy of prediction.

[0006] In order to solve the above technical problems, the technical solution of the present invention is: a multi-disaster graded prediction method based on artificial intelligence, comprising the following steps:

[0007] S1. Collect mining disaster data to form a mining disaster database; mining disaster data at least includes hanging wall stability, slope stability, rock burst, pillar stability and goaf stability;

[0008] S2. Constructing a mine disaster prediction index system: Screening the data in the mine disaster database, eliminating the indexes with high feature overlap and low feature importance, and forming a mine disaster prediction index system after dimensionality reduction based on the screened data;

[0009] S3. Constructing a mine disaster prediction database: reconstructing the data set in the mine disaster database, processing the abnormal values ​​and missing values ​​therein, obtaining a balanced data set and using it as the data in the mine disaster prediction database;

[0010] S4. Select evaluation features according to the indicators in the mine disaster prediction index system, perform graded prediction on the data in the mine multi-disaster prediction database according to the algorithms in the preset mine disaster prediction algorithm library, summarize and compare the prediction results of the graded prediction of each algorithm under different preset evaluation indicators, obtain the algorithm with the best score of the comprehensive evaluation index of the prediction result, and use the algorithm as the algorithm of the optimal model for mine disaster prediction;

[0011] S5. Input the mine disaster data into the optimal mine disaster prediction model to conduct graded prediction of multiple disaster types.

[0012] In S2, data indicators are screened according to the Pearson correlation coefficient and feature importance. The specific steps are as follows:

[0013] Calculate the feature importance of any type of disaster in the mine disaster data, remove the features with low feature importance, and calculate the Pearson correlation coefficient of the remaining features;

[0014] Weights are assigned according to the absolute values ​​of the Pearson correlation coefficients between the remaining features. The larger the absolute value of the Pearson correlation coefficient between the features, the greater the weight. The Pearson correlation coefficient is used to indicate the strength of the correlation between the features. The larger the absolute value of the Pearson correlation coefficient between two features, the stronger the correlation.

[0015] According to different preset feature importance algorithms, the features corresponding to the absolute values ​​of the weighted Pearson correlation coefficients are ranked in importance, and the filtered data is obtained based on the importance ranking results and the correlation between the features.

[0016] The Pearson correlation coefficient is a statistical indicator used to measure the strength and direction of the linear relationship between two continuous variables. The value range is between -1 and 1. The larger the absolute value, the stronger the correlation. If the value is positive, it is a positive correlation, and if the value is negative, it is a negative correlation. The calculation formula is expressed as:

[0017]

[0018] Among them, ρ X,Y is the Pearson correlation coefficient, cov(X,Y) represents the covariance between the two variables X and Y, σ X , σ Y are the standard deviations of variables X and Y respectively.

[0019] The feature importance is output by the random forest algorithm and measured by the gini coefficient. The principle is as follows: suppose there are J features X1, X2, X3, ..., XJ, a total of I decision trees, and C categories; for each feature Xj, calculate its contribution to each tree in the random forest, and then take the average value to compare the importance of the features;

[0020] When calculating the Gini index score, the impurity change of each node of each tree is considered, expressed as the Gini coefficient of the node; among them, the Gini coefficient of node q The calculation formula is expressed as:

[0021]

[0022] Where C represents the number of categories, represents the proportion of category c in node q;

[0023] The importance of feature Xj in the i-th tree node q, that is, the change in the gini coefficient before and after the node q branches for:

[0024]

[0025] in, and Respectively represent the gini coefficients of the two new nodes after the branch.

[0026] The specific steps for reconstructing the data set in the mine disaster database in S3 are:

[0027] The Smote algorithm is improved by Tomek Link algorithm to reduce the outliers generated in the reconstruction of the data set;

[0028] The steps of reconstructing the data set using the improved SmoteTomek algorithm are expressed as follows:

[0029] (1) The mine disaster database is used as a sample set, and the data in it is used as samples. After the minority class and the majority class are determined according to the sample set, the Euclidean distance between each minority class sample and all minority class samples is calculated to obtain its k nearest neighbor samples; then, the same operation is performed on other minority class samples in turn;

[0030] (2) According to the sample ratio of the majority class samples to the minority class samples, determine the number of samples to be generated. For each minority class sample k, randomly select a corresponding number of samples from its nearest neighbor set. Assume that the selected nearest neighbor is x.

[0031] (3) For each selected neighbor sample x, construct a new sample x according to the following formula new :

[0032]

[0033] (4) Construct Tomek Link pairs with the new samples and their original samples, delete the paired synthetic samples, and clean up each minority class sample in turn;

[0034] (5) Repeat steps (1) to (4) until the sample categories of the reconstructed data are balanced.

[0035] The specific method for processing data outliers in the reconstructed data set in S3 is:

[0036] The Mahalanobis distance discriminant method is used to screen the data, and its calculation method is expressed as:

[0037]

[0038] The above formula is the Mahalanobis distance in n-dimensional space for a multivariate vector with mean μ and covariance matrix Σ;

[0039] Among them, D M (x) is the Mahalanobis distance between each individual sample Y and the average sample μ; T represents the transpose of the matrix.

[0040] The method also includes the following steps: performing algorithm analysis on the mine multi-disaster prediction database through several cross-validations to verify the reliability of the sample feature importance.

[0041] The mine disaster prediction algorithm library includes a machine learning algorithm library and a deep learning algorithm library. The algorithms in the mine disaster prediction algorithm library include at least: the support vector machine algorithm, the random forest algorithm and the extreme random tree algorithm in the machine learning algorithm library, and the BP neural network algorithm, the CNN algorithm and the LSTM algorithm in the deep learning algorithm library.

[0042] A computer device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the computer program.

[0043] A computer-readable storage medium is also provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] (1) The present invention constructs and reconstructs a database of multiple categories of mine disasters. In order to address the data imbalance and outlier problems caused by the characteristics of mine disasters themselves, the Mahalanobis distance discriminant method is used to screen outliers in the mine disaster sample data, and Tomek Link is used to optimize the SMOTE algorithm, which significantly improves the imbalance of the mine disaster data set and improves the overall prediction accuracy.

[0046] (2) The present invention aims at solving the problems that there are too many indicators for the classification and prediction of mine disasters, some indicators are highly correlated and interfere with each other, and there is no relatively unified prediction indicator system. Starting from the mine disaster sample data itself, the importance of the indicators and the correlation coefficient are calculated. The mine disaster indicators are screened twice, and a mine disaster classification prediction indicator system based on data analysis is established to achieve data dimensionality reduction and enhance the prediction accuracy.

[0047] (3) The present invention uses machine learning and deep learning algorithms to perform hierarchical prediction on the constructed mining disaster data set, and uses parameters such as accuracy, precision, F1 value, etc. as prediction evaluation indicators to analyze the hierarchical prediction results of different algorithms, thereby further improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of a flow chart of an embodiment of the present invention;

[0049] FIG2( a ) is a feature importance ranking diagram of the RF algorithm used to calculate feature importance in an embodiment of the present invention;

[0050] FIG2( b ) is a feature importance ranking diagram of the XGB algorithm used to calculate feature importance in an embodiment of the present invention;

[0051] FIG2( c ) is a feature importance ranking diagram of the ET algorithm used to calculate feature importance in an embodiment of the present invention;

[0052] FIG2( d ) is a feature importance ranking diagram of the ADA algorithm used to calculate feature importance in an embodiment of the present invention;

[0053] FIG2( e ) is a feature importance ranking diagram of the GBDT algorithm used to calculate feature importance in an embodiment of the present invention;

[0054] FIG2( f ) is a feature importance ranking diagram of the DT algorithm used to calculate feature importance in an embodiment of the present invention;

[0055] Figure 3 Schematic diagram of goaf stability evaluation index in an embodiment of the present invention;

[0056] Figure 4 Schematic diagram of the SMOTE algorithm improved by Tomek Link in an embodiment of the present invention;

[0057] Figure 5 Schematic diagram of ten-fold cross validation in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0059] The technical solution of the present invention is:

[0060] A multi-disaster classification prediction method based on artificial intelligence includes the following steps:

[0061] S1. Collect mining disaster data to form a mining disaster database; mining disaster data at least includes hanging wall stability, slope stability, rock burst, pillar stability and goaf stability;

[0062] S2. Constructing a mine disaster prediction index system: Screening the data in the mine disaster database, eliminating the indexes with high feature overlap and low feature importance, and forming a mine disaster prediction index system after dimensionality reduction based on the screened data; wherein, the method for eliminating the indexes with high feature overlap is: calculating the Pearson correlation coefficient, and then using the weighting method in Table 1 to determine the feature overlap, and deleting the features with high overlap, see Table 1 below;

[0063] S3. Constructing a mine disaster prediction database: reconstructing the data set in the mine disaster database, processing the abnormal values ​​and missing values ​​therein, obtaining a balanced data set and using it as the data in the mine disaster prediction database; wherein the missing values ​​are processed by using the median interpolation method;

[0064] S4. Select evaluation features according to the indicators in the mine disaster prediction index system, perform graded prediction on the data in the mine multi-disaster prediction database according to the algorithms in the preset mine disaster prediction algorithm library, summarize and compare the prediction results of the graded prediction of each algorithm under different preset evaluation indicators, obtain the algorithm with the best score of the comprehensive evaluation index of the prediction result, and use the algorithm as the algorithm of the optimal model for mine disaster prediction;

[0065] S5. Input the mine disaster data into the optimal mine disaster prediction model to conduct graded prediction of multiple disaster types.

[0066] There is a strong correlation between the data indicators of mine disasters. In order to reduce the dimension of disaster data, the Pearson correlation coefficient and feature importance are used to simplify the data indicators. The specific principles are as follows:

[0067] (1) Pearson correlation coefficient principle

[0068] The Pearson correlation coefficient, also known as Pearson correlation, is a statistical indicator used to measure the strength and direction of the linear relationship between two continuous variables. The value range is between -1 and 1. The larger the absolute value, the stronger the correlation. If the value is positive, it is a positive correlation, and if the value is negative, it is a negative correlation. The calculation formula is as follows:

[0069]

[0070] Where cov(X,Y) represents the covariance between the two variables X and Y, σ X σ Y is the standard deviation of the variable.

[0071] (2) Principles of feature importance assessment

[0072] Taking random forest as an example, the variable importance measures (VIM) are measured using the gini coefficient. Assume there are J features X1, X2, X3, ..., XJ, a total of I decision trees, and C categories. For each feature Xj, its contribution to each tree in the random forest needs to be calculated, and then the average is taken to compare the importance of the features.

[0073] When calculating the Gini index score, the impurity change of each node of each tree is considered, which is expressed as the Gini coefficient of the node. The Gini coefficient (GI) calculation formula of node q is:

[0074]

[0075] Where C represents the number of categories, Represents the proportion of category c in node q.

[0076] The importance of feature Xj in the i-th tree node q, that is, the change in the Gini coefficient before and after the node q branch is:

[0077]

[0078] in, and Respectively represent the Gini coefficients of the two new nodes after the branch.

[0079] First, the feature importance of a single disaster type is calculated based on the above feature importance, and the lower features are eliminated according to the prediction system requirements. Then the Pearson correlation coefficient of the remaining features is calculated.

[0080] Taking the goaf data as an example, the Pearson correlation coefficient is calculated, where the absolute value of the value represents the strength of the correlation between the features. A negative sign represents a negative correlation between the features, and vice versa for a positive correlation. The absolute value of the correlation coefficient below 0.2 is defined as no correlation, 0.2-0.3 is defined as a general correlation, 0.3-0.5 is defined as a strong correlation, and 0.5 or above is defined as a significant correlation. * indicates a general correlation, ** indicates a strong correlation, and *** indicates a significant correlation.

[0081] Table 1

[0082] X1 X2 X3 X4 X5 X6 X7 X8 X9 X10 X11 *** 8 10 6 6 3 5 6 5 6 7 6 ** 2 0 0 3 2 2 2 5 2 1 1 * 0 0 2 1 3 2 0 0 0 2 2 Scores 2.8 3 2 2.5 1.6 2.1 2.2 2.5 2.2 2.5 2.2

[0083] The correlation between the two feature variables is counted by the number of times, and the weights of ***, **, and * are set to 0.3, 0.2, and 0.1, respectively, so that the correlation between each feature and the remaining features is shown in Table 1. The correlation between the 11 indicators affecting the stability of the goaf is strong, so it is necessary to reduce the features. As shown in Table 1, X3, X5, and X6 score 2, 1.6, and 2.1, respectively, and the correlation with other features is relatively low, so they can be regarded as independent features for selection.

[0084] The following six algorithms with the ability to calculate feature importance are selected: Random Forest (RF), Xgboost (XGB), ExtraTrees (ET), Adaboost (ADA), Gradient Boosting Decision Trees (GBDT), and DecisionTrees (DT) to rank the importance of the features that affect the stability of the goaf. The ranking results are as follows: Figure 2(a)-Figure 2(f) The gray shaded area in the figure represents the threshold of 0.10 importance (the sum of all indicator importances is 1). When the importance of a feature in the corresponding algorithm exceeds 0.1, a √ mark is made at the corresponding position in Table 2.

[0085] It can be seen from Table 2 that the importance of X4, X5, X6, X7, X9, and X10 is relatively high, while the importance of X1, X2, X8, and X11 is relatively low.

[0086] Table 2

[0087] X1 X2 X3 X4 X5 X6 X7 X8 X9 X10 X11 RF √ √ √ XGB √ √ √ √ ET √ √ √ √ ADA √ √ √ √ √ GBDT √ √ √ √ √ √ DT √ √ √ √

[0088] like Figure 3 As shown, these are the indicators represented by X1-X9.

[0089] Based on the above analysis of feature correlation and importance, the goaf indicators X3, X4, X5, X6, X7, X9, and X10 are selected.

[0090] The Tomek Link pair is used to improve the SMOTE algorithm to reduce the abnormal samples generated in the dataset reconstruction. The Tomek Link algorithm helps to identify and remove the overlapping samples in adjacent classes, thereby reducing the generation rate of abnormal samples.

[0091] The main idea of the Tomek Links algorithm is that if X and Y in the samples come from different classes respectively, and there is no other sample Z such that d(X,Z) < d(X,Y) or d(Y,Z) < d(X,Y) holds, then the above samples X and Y are called Tomek Link pairs. Here, d represents the Euclidean distance between two samples. In Figure 4 the figure, the part under the red shadow is the Tomek Link pair, and the data pairs under the red shadow are deleted to achieve the effect of cleaning the data at the classification boundary.

[0092] The Tomek Link method can be used to undersample the synthesized samples, thereby improving the quality of the training samples. The steps for reconstructing the Smote Tomek dataset are as follows:

[0093] (1) After determining the minority class according to the sample set, for each minority class sample x, calculate the Euclidean distance between it and all minority class samples to obtain the sample set of its k nearest neighbors. Subsequently, the same operation is performed on other minority class samples in turn.

[0094] (2) According to the sample ratio of the majority class samples to the minority class samples, determine the number of samples to be generated. For each minority class sample k, randomly select the corresponding number of samples from its nearest neighbor set. Suppose the selected nearest neighbor is x.

[0095] (3) For each selected nearest neighbor sample x, construct a new sample according to the following formula:

[0096]

[0097] (4) According to Figure X, construct Tomek Link pairs for the newly generated samples and the original samples, and delete the paired synthesized samples, and clean each minority class sample in turn.

[0098] (5) Repeat steps (1) to (4) until the sample categories of the reconstructed data are basically balanced.

[0099] The main problems of the mine disaster dataset mainly include the existence of outliers, unbalanced dataset, large difference in data dimensions, etc.

[0100] In order to properly handle the outliers in the mine disaster data, the Mahalanobis distance method is used to screen the data. Unlike the Euclidean distance, it takes into account the relationship between various characteristics and is scale-independent, that is, independent of the measurement scale. Therefore, the Mahalanobis distance can be independent of the measurement unit and comprehensively reflect the correlation between multiple variables. The specific calculation rules are as follows:

[0101]

[0102] For a multivariate vector with a mean of μ and a covariance matrix of Σ, the Mahalanobis distance in n-dimensional space is calculated as follows, where d is the Mahalanobis distance between each individual sample Y and the average sample μ. Using the Mahalanobis distance method, a regression relationship can be established between a single indicator and the other indicators, and the influence of each indicator on the indicator can be comprehensively considered to determine whether there are outliers in the indicator.

[0103] The dimensions of the indicator data in the mine disaster database are inconsistent, and the range of values ​​is large, which is not conducive to the algorithm's feature learning. In order not to affect the distribution of the samples, consider using formula 6 to normalize the rockburst sample data, and use the shuffle function in python to disrupt the order of the data set. In order to ensure the reliability of the sample importance obtained by the algorithm library analysis, a 10-fold cross-validation of all samples was used for algorithm analysis. The specific operation is to divide the database sample data into 10 parts, and each part of the data is called a fold. Then, each time cross-validation is performed, one of the parts is selected as the test set, and the remaining 9 parts are used as training sets. Next, train the model on the training set and make predictions on the test set. Repeat this process 10 times, selecting a different test set each time. Finally, calculate the accuracy of these 10 cross-validations, and take the average as the final cross-validation accuracy. The schematic diagram is shown in the figure. Figure 5 This method can avoid the randomness caused by data set division and ensure the reliability of the obtained sample importance.

[0104]

[0105] The support vector machine, random forest, extreme random tree in the machine learning algorithm library and the BP neural network, CNN, and LSTM algorithms in the deep learning algorithm library are used to classify and predict the five types of mine disasters, and the PSO algorithm is used to optimize the hyperparameters in machine learning. Different evaluation indicators such as accuracy, precision, and recall rate are used to summarize and compare the prediction results of different algorithms for the five types of mine disasters, and the optimal model for mine classification prediction is obtained through analysis.

[0106] Using the above-constructed mining disaster dataset, the algorithms in the machine learning algorithm library and the deep learning algorithm library are used to perform graded predictions on five types of mining disasters. In order to achieve the best effect of the algorithm, the PSO algorithm is used to optimize the hyperparameters of the algorithm.

[0107] With the continuous development of artificial intelligence, the iteration speed of algorithms is gradually accelerating. In order to ensure the effectiveness of the prediction system, it is necessary to continuously iterate the algorithms in the algorithm library, and transfer the newly developed algorithms into the mine disaster prediction algorithm library. By comparing different evaluation indicators such as accuracy, precision, and recall rate with the optimal algorithm in the algorithm library, the optimal algorithm can be obtained in a timely manner.

[0108] The iteration of the database is that every time new data is predicted, the data is included in the database, and then the new database is used as the mine multi-disaster prediction database when the system is used next time, thus completing the iteration of the database.

[0109] (1) Construction and reconstruction of a database of multiple categories of mine disasters. In order to solve the problem of data imbalance and outliers caused by the characteristics of mine disasters, the Mahalanobis distance discriminant method is used to screen outliers in the mine disaster sample data, and the SMOTE algorithm is optimized using Tomek Link, which significantly improves the imbalance of the mine disaster data set and improves the accuracy of the overall prediction.

[0110] (2) In view of the problems that there are too many indicators for the classification and prediction of mine disasters, some indicators are highly correlated and interfere with each other, and there is no relatively unified prediction indicator system, the importance and correlation coefficient of indicators are calculated based on the mine disaster sample data itself. The mine disaster indicators are screened twice, and a mine disaster classification prediction indicator system based on data analysis is established to achieve data dimensionality reduction and enhance prediction accuracy.

[0111] (3) Machine learning and deep learning algorithms are used to perform graded predictions on the constructed mine disaster dataset, and parameters such as accuracy, precision, and F1 value are used as prediction evaluation indicators to analyze the graded prediction results of different algorithms. In the current mine disaster prediction database, the optimal machine learning algorithm is the extreme random tree algorithm, and the optimal algorithm is the deep forest algorithm. The accuracy of the deep forest algorithm in predicting five levels of mine disasters is as follows: goaf risk level prediction is 92.31%; slope stability prediction is 96.77%; rock burst level prediction is 92.50%; pillar stability prediction is 91.67%; and upper wall stability prediction is 95.00%.

[0112] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-disaster classification prediction method based on artificial intelligence, characterized in that: The following steps are involved: S1. Collect mining disaster data to form a mining disaster database; mining disaster data at least includes hanging wall stability, slope stability, rock burst, pillar stability and goaf stability; S2. Constructing a mine disaster prediction index system: Screening the data in the mine disaster database, eliminating the indexes with high feature overlap and low feature importance, and forming a mine disaster prediction index system after dimensionality reduction based on the screened data; S3. Constructing a mine disaster prediction database: reconstructing the data set in the mine disaster database, processing the abnormal values ​​and missing values ​​therein, obtaining a balanced data set and using it as the data in the mine disaster prediction database; S4. Select evaluation features according to the indicators in the mine disaster prediction index system, perform graded prediction on the data in the mine multi-disaster prediction database according to the algorithms in the preset mine disaster prediction algorithm library, summarize and compare the prediction results of the graded prediction of each algorithm under different preset evaluation indicators, obtain the algorithm with the best score of the comprehensive evaluation index of the prediction result, and use the algorithm as the algorithm of the optimal model for mine disaster prediction; S5. Input the mine disaster data into the optimal mine disaster prediction model to conduct graded prediction of multiple disaster types.

2. According to the artificial intelligence-based multi-disaster classification prediction method of claim 1, it is characterized in that: In S2, data indicators are screened according to the Pearson correlation coefficient and feature importance. The specific steps are as follows: Calculate the characteristic importance of any disaster type in the mine disaster data, and eliminate the characteristic importance The low features are selected and the Pearson correlation coefficient is calculated for the remaining features; Weights are assigned according to the absolute values ​​of the Pearson correlation coefficients between the remaining features. The larger the absolute value of the Pearson correlation coefficient between the features, the greater the weight. The Pearson correlation coefficient is used to indicate the strength of the correlation between the features. The larger the absolute value of the Pearson correlation coefficient between two features, the stronger the correlation. According to different preset feature importance algorithms, the features corresponding to the absolute values ​​of the weighted Pearson correlation coefficients are ranked in importance, and the filtered data is obtained based on the importance ranking results and the correlation between the features.

3. The multi-disaster classification prediction method based on artificial intelligence according to claim 2 is characterized in that: The Pearson correlation coefficient is a statistical indicator used to measure the strength and direction of the linear relationship between two continuous variables. The value range is between -1 and 1. The larger the absolute value, the stronger the correlation. If the value is positive, it is a positive correlation, and if the value is negative, it is a negative correlation. The calculation formula is expressed as: Among them, ρ X,Y is the Pearson correlation coefficient, cov(X,Y) represents the covariance between the two variables X and Y, ρ X , Y are the standard deviations of variables X and Y respectively.

4. The method for multi-disaster classification prediction based on artificial intelligence according to claim 2 is characterized in that: The feature importance is output by the random forest algorithm and measured by the gini coefficient. The principle is as follows: suppose there are J features X1, X2, X3, ..., XJ, a total of I decision trees, and C categories; for each feature Xj, calculate its contribution to each tree in the random forest, and then take the average value to compare the importance of the features; When calculating the Gini index score, the impurity change of each node of each tree is considered, expressed as the Gini coefficient of the node; among them, the Gini coefficient of node q The calculation formula is expressed as: Where C represents the number of categories, represents the proportion of category c in node q; The importance of feature Xj in the i-th tree node q, that is, the change in the gini coefficient before and after the node q branches for: in, and Respectively represent the gini coefficients of the two new nodes after the branch.

5. The method for multi-disaster classification prediction based on artificial intelligence according to claim 4 is characterized in that: The specific steps for reconstructing the data set in the mine disaster database in S3 are: The Smote algorithm is improved by Tomek Link algorithm to reduce the outliers generated in the reconstruction of the data set; The steps of reconstructing the data set using the improved SmoteTomek algorithm are expressed as follows: (1) The mine disaster database is used as a sample set, and the data in it is used as samples. After the minority class and the majority class are determined according to the sample set, the Euclidean distance between each minority class sample and all minority class samples is calculated to obtain its k nearest neighbor samples; then, the same operation is performed on other minority class samples in turn; (2) According to the sample ratio of the majority class samples to the minority class samples, determine the number of samples to be generated. For each minority class sample k, randomly select a corresponding number of samples from its nearest neighbor set. Assume that the selected nearest neighbor is x. (3) For each selected neighbor sample x, construct a new sample x according to the following formula new : (4) Construct Tomek Link pairs with the new samples and their original samples, delete the paired synthetic samples, and clean up each minority class sample in turn; (5) Repeat steps (1) to (4) until the sample categories of the reconstructed data are balanced.

6. The method for multi-disaster classification prediction based on artificial intelligence according to claim 5 is characterized in that: The specific method for processing data outliers in the reconstructed data set in S3 is: The Mahalanobis distance discriminant method is used to screen the data, and its calculation method is expressed as: The above formula is the Mahalanobis distance in n-dimensional space for a multivariate vector with mean μ and covariance matrix Σ; Among them, D M (x) is the Mahalanobis distance between each individual sample Y and the average sample μ; T represents the transpose of the matrix.

7. The method for multi-disaster classification prediction based on artificial intelligence according to claim 1 is characterized in that: The following steps are also included: The algorithm analysis of the mine multi-disaster prediction database was carried out through several cross-validations to verify the reliability of the sample feature importance.

8. The method for multi-disaster classification prediction based on artificial intelligence according to claim 1 is characterized in that: The mine disaster prediction algorithm library includes a machine learning algorithm library and a deep learning algorithm library. The algorithms in the mine disaster prediction algorithm library include at least: the support vector machine algorithm, the random forest algorithm and the extreme random tree algorithm in the machine learning algorithm library, and the BP neural network algorithm, the CNN algorithm and the LSTM algorithm in the deep learning algorithm library.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Chinese elderly cognitive impairment prediction model

    CN114420300A

  • Drought disaster weather prediction method based on semi-supervised ensemble learning

    CN114841064A

  • Multi-sensor information fusion fire prediction algorithm and system, electronic equipment and medium

    CN115619013A

  • Method and system for optimizing turbine of thermal power unit on basis of sparse big data mining

    WO2022193569A1