A high-resilience safety assessment method for complex mining projects based on AdaBoost algorithm
By applying a high-strength security assessment method based on the AdaBoost algorithm in complex mining engineering, abnormal data is identified and processed, data inaccuracy caused by sensor degradation in harsh environments is solved, and the accuracy and reliability of security assessment is improved.
Patent Information
- Application Number
- CN202310882745.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-07-18
AI Technical Summary
The harsh environment in complex mining projects leads to accelerated degradation of sensors, resulting in inaccurate data collection and affecting the accuracy of safety assessment.
A high-resilience security evaluation method based on AdaBoost algorithm is adopted to identify normal and abnormal data, update data weights, and build a highly accurate security evaluation model.
Effectively resist the interference of harsh external conditions on sensor data, improve the accuracy of safety assessment, reduce false alarm rates and missed reports, and support the decision-making of accident warnings and prevention measures.
Smart Images

Figure CN117151915B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a high-toughness safety assessment method for complex mining engineering based on an AdaBoost algorithm, and belongs to the field of safety assessment of coal mining engineering. Background Art
[0002] Coal is my country's basic energy and important industrial raw material. my country's current energy demand is still on the rise, so how to ensure the safe mining of coal has become a key issue. Coal mining is divided into two methods: open-pit mining and underground mining. At present, the proportion of underground mining in my country is as high as 85%, making it the country with the highest proportion of underground coal mining in the world. In the process of underground coal mining, in order to meet the needs of coal mining and transportation, a large number of working faces, tunnels and chambers need to be excavated, collectively referred to as mining projects. In the process of mining and excavation engineering operations, complex geological conditions such as groundwater, geological structures, and high stress are often encountered. Once an accident occurs, it will directly threaten the lives of underground workers. Therefore, it is very important to ensure the safety of mining and excavation projects.
[0003] In order to monitor the actual situation in the mine in real time, each coal mine has a large number of sensors arranged at multiple key points to collect various data in real time to judge and analyze whether the current mining project is safe, including tunnel cover thickness, cover span ratio, soil internal friction angle, soil compression modulus, soil cohesion, soil bin pressure, grouting volume, etc. These data are called mining data. Since the working conditions of mining projects are underground high temperature and high humidity environments, the working conditions of multiple sensors are far lower than the ideal working conditions, so the accelerated degradation of multiple sensors is inevitable, which will cause serious distortion of the measured data and easily affect the authenticity and validity of the data. At present, most mining projects do not process mining data, but directly conduct safety assessment based on mining data. For example, the threshold-based mining project safety assessment method: when the shape variable of any point in the mining working face exceeds the safety threshold, it must be stopped for inspection, and the safety indicators exceeding the threshold are adjusted (such as reducing the excavation speed, reducing the torque, etc.) to make the safety indicators fall back to the safe range. This threshold-based mining engineering safety assessment method is easy to understand, convenient to practice, and has good operability. However, the phenomenon of reduced credibility of mining data caused by the degradation of multi-sensor performance objectively exists in this process, so it is very likely to increase the false alarm rate and missed alarm rate, causing serious economic losses. Roof accidents in coal mine mining projects, whether in terms of the number of occurrences or the number of deaths, have seriously endangered the safety of life and property of the people. The root cause of roof accidents lies in the fracture and fall of the roof rock layer of the working face, which is the safety indicator of "mine settlement value" concerned by the present invention. Most accidents occur because the monthly cumulative value of mine settlement value exceeds the safety range during the mining process, which causes the fracture and collapse of the rock layer in the long run, and then evolves into roof accidents. Therefore, the output indicator "mine settlement value" in the mining data is of great importance in the safety assessment of mining projects. Summary of the invention
[0004] The purpose of the present invention is to solve the problem of accelerated degradation of multiple sensors caused by harsh environments in complex mining projects, which in turn leads to inaccurate collected data. A high-toughness safety assessment method for complex mining projects based on the AdaBoost algorithm is proposed. The high toughness mentioned in the present invention refers to the ability to resist the interference of harsh external construction conditions on sensor collected data. That is to say, in the case where the accuracy of some data in the known on-site mining data has decreased, this method can still support the improvement of the accuracy of safety assessment, provide early warning for the occurrence of accidents, and provide decision-making support for the formulation of necessary preventive measures.
[0005] Technical solution of the present invention
[0006] A high-toughness safety assessment method for complex mining projects based on the AdaBoost algorithm. In the present invention, a safety assessment model is constructed based on the data collected by sensors at the construction site to identify potential dangers and risk sources, and an early warning is provided for the occurrence of accidents based on the output value of the model, providing decision support for the next step of control measures. Therefore, the closer the output value of the model is to the true value, the greater the false alarm rate and missed alarm rate will be, and the project can run safely and smoothly.
[0007] The present invention first divides the data collected at the construction site into a training data set and a test data set, where each sample contains 11 input feature parameters of three types of influencing factors, namely tunnel design parameters, geological conditions and construction parameters, and uses the mine settlement value as the output parameter to build a safety assessment model. Then, the neural network is used as the benchmark model to divide the training data set into normal data and abnormal data and update the weight according to the data type. Finally, a high-toughness safety assessment method for complex mining projects based on the AdaBoost algorithm is proposed. The research results show that the AdaBoost high-toughness safety assessment method constructed based on normal data and abnormal data has a higher accuracy than the results obtained by directly using the AdaBoost algorithm. This advantage is also reflected in the experimental results of adding additional noise to the training set and the test set, which fully verifies the high-toughness characteristics of the proposed method.
[0008] The method of the present invention comprises the following steps:
[0009] (1) Data collection and data set division. Tunnel design parameters, geological condition parameters and construction parameters are used as input parameters of the model. The data of the input parameters are collected from the time series data collected from the actual coal mining construction site to construct a complete data set D. In the complete data set D, the first 80% is divided into training set D in order. T , the remaining 20% is divided into the test set D V .
[0010] (2) From the training set D T In D, S sub-datasets are randomly selected to build S sub-models. T Randomly select S sub-datasets in the same proportion To build S sub-models, these sub-datasets It should account for the majority of the data in the complete training set and the number of sub-models should be at least 30, that is, S ≥ 30, in order to reduce the random error in the modeling process. In this step, the extraction of the training sub-dataset can be repeated, that is, the intersection of any two training subsets can be non-empty, and it is only necessary to ensure that the sizes of the subsets are the same.
[0011] (3) Computational sub-model The errors and weights of the S sub-models constructed in step (2) are used to calculate the weight of each sub-model, and then the S sub-models are calculated based on these weights. The average error of each group of data.
[0012] (4) Weight the data and build a prediction model. The average error of each group of data is arranged in descending order, and the first 10% of the data is defined as abnormal data in the present invention, and the remaining 90% of the data is defined as normal data. Normal data helps to establish an accurate safety assessment model, and the default weight is 1; while abnormal data will interfere with the safety assessment modeling process and should be given a smaller weight, which is 0.01 in the present invention by default. T After weighting each set of data in, a new weighted data set is returned To build the final prediction model.
[0013] (5) Verify on the test set. As the data weights are updated, the entire weighted training set is used to construct the final prediction model, and the prediction model finally constructed in step (4) is verified on the test set.
[0014] Advantages and beneficial effects of the present invention:
[0015] 1. In traditional mining engineering practice, it is generally assumed that all mining data are completely accurate, but the accelerated degradation of sensors caused by the harsh underground mining engineering environment may reduce the accuracy of mining data. Therefore, the present invention starts from a new perspective of the consistency relationship between mining data input and output, defines normal / abnormal data in mining data, and thus proposes a security assessment method for a high-toughness AdaBoost algorithm.
[0016] Second, the high-toughness AdaBoost method can effectively deal with the interference of external harsh construction conditions on sensor data collection, and can effectively improve the accuracy of safety assessment. Compared with directly using the complete training data set without identifying various types of data, the high-toughness AdaBoost method has achieved superior performance; further verification shows that under different levels of external noise, the high-toughness AdaBoost method is better than the direct use of the AdaBoost method, which shows that the method proposed in this invention has good toughness. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall framework diagram of the method of the present invention.
[0018] Figure 2 This is the partition result diagram of the original data, where (a) is sorted by error size and (b) is sorted by the original order of errors.
[0019] Figure 3 This is the experimental result diagram on the original test set without adding noise, where (a) the output values of the two methods are compared with the true values, and (b) the mean absolute error values of the two methods.
[0020] Figure 4 1 is a graph of the experimental results of the present invention in an actual case with noise added to the training set, where (a) is the mean absolute error and (b) is the improvement rate (%).
[0021] Figure 5 1 is a graph of the experimental results of the present invention in an actual case with a test set plus noise, where (a) is the mean absolute error and (b) is the improvement rate (%).
[0022] Figure 6 The following are specific results of adding 15% noise to the test set in an actual case, where (a) the output values of the two methods are compared with the true values, and (b) the mean absolute error values of the two methods. DETAILED DESCRIPTION
[0023] The present invention relates to a high-toughness safety assessment method for complex mining projects based on the AdaBoost algorithm. First, the data collected at the construction site is divided into a training data set and a test data set. Then, a neural network is used as a benchmark model to divide the training data set into normal data and abnormal data, and the weight is updated according to the data type. Finally, a high-toughness safety assessment method for complex projects based on AdaBoost is proposed. The research results show that the AdaBoost high-toughness safety assessment method constructed based on normal data and abnormal data has a higher accuracy than the results obtained by directly using the AdaBoost algorithm. This advantage is also reflected in the experimental results of adding additional noise to the training set and the test set.
[0024] The following describes in detail the embodiments of the method of the present invention in conjunction with the accompanying drawings.
[0025] The high-toughness safety assessment method for complex mining engineering based on AdaBoost algorithm proposed in this paper includes five steps. The overall framework of the method is as follows: Figure 1 shown.
[0026] (1) Data collection and data set division. The present invention uses tunnel design parameters, geological condition parameters and construction parameters as input parameters of the model, collects data of the input parameters from the time series data collected from the actual coal mining construction site, and constructs a complete data set D. In the complete data set D, the first 80% is sequentially divided into a training set D T , the remaining 20% is divided into the test set D V .
[0027] (2) From the training set D T30 sub-datasets are randomly selected from D T 30 sub-datasets are randomly selected in the same proportion to build 30 sub-models. The specific method of extracting sub-datasets to build sub-models is: each time from the complete training set D T 90% of the data are randomly selected to form a sub-dataset This step is repeated 30 times to obtain 30 sub-datasets. Each sub-dataset All D T A subset of S=30.
[0028] (3) Calculate the errors and weights of the sub-models. The 30 output errors generated by the 30 sub-models are used to calculate the weight of each sub-model, and then the sub-models are calculated based on these weights. The average error of each group of data.
[0029] (3.1) Calculate the sub-model error.
[0030] First, for each sub-model Calculate the output error for each set of data in As shown below:
[0031]
[0032] Among them, p represents the sub-model The size of the data in , p∈[1,360] and p is an integer, y p Represents the true output of the pth group of data; Represents the output of the p-th group of data in the sub-model s. If the p-th group of data does not appear in the sub-model In
[0033] Then, the BP neural network model is used to calculate the mean absolute error percentage (MAPE) of the sth sub-model. s ), as follows:
[0034]
[0035] Wherein, P represents P groups of data for constructing the s-th sub-model, in the present invention, P is 360, p represents the size of the data in the s-th sub-model, p∈[1,360] and p is an integer.
[0036] (3.2) Calculate sub-model weights
[0037] Calculate the MAPE of the sub-model s to represent their ability to identify the corresponding sub-training datasets, and MAPE s ∈[0,100%]. Where MAPEs = 0 or MAPE s =100% represents the highest and lowest modeling capabilities of the model. Therefore, the weight of the sth sub-model can be obtained according to the following formula:
[0038] w s =1-MAPE s (3)
[0039] Among them, w s ∈[0,1].
[0040] (3.3) Calculate the average error of all P groups of data.
[0041] The average error of all P groups of data is calculated by the output error of each group of data and the weight of the sub-model, as shown below:
[0042]
[0043] in, represents the output error of the pth group of data in the sth sub-model; w s represents the weight of the s-th sub-model; s * Indicates that this submodel is not constructed using the p-th group of data. * sub-models, then e p =0. The above formula can also be converted into the following formula:
[0044]
[0045] (4) Data weighting and prediction model building. Reweight each set of data according to the average error of each set of data, and use the weighted data set to Build the final prediction model.
[0046] First, according to the average error of the pth group of data calculated The pth group of data is weighted as shown in the following formula:
[0047]
[0048] Among them, w p,0 represents the initial weight of the pth group of data (the default is 1); μ abnormal Represents the update weight of abnormal data; μ normal represents the update weight of normal data. For abnormal data, a smaller update weight should be given. The present invention defines μ abnormal =0.01; for normal data, the present invention keeps the weight unchanged by default, i.e. μ normal =1.
[0049] Then, after updating the weights for each set of data, a new weighted training set is returned. It is used to train AdaBoost and generate the final high-resilience AdaBoost model, where the objective function of the model is as follows:
[0050]
[0051] in, It has been given in (3.1); y p represents the true output of the pth group of data; w p It has been given in the above formula.
[0052] (5) Verification on the test set. As the data weights are updated, the entire weighted training set is used to build the final prediction model and verified on the test set.
[0053] The performance indicator verified on the test set in step (5) is MAE, which is an indicator used to evaluate the accuracy of the model. It measures the average difference between the predicted value and the true value. The smaller the indicator, the better the model prediction ability. The calculation formula is as follows:
[0054]
[0055] Where Q represents the Q group of data in the test set. In the present invention, Q is 100. q represents the amount of data in the test set. q∈[1,100] and q is an integer. q Represents the true output of the qth group of data in the test set; Represents the output of the qth group of data in the test set in the prediction model.
[0056] Example 1
[0057] The following is an introduction to the experimental results of the high-resilience security assessment method based on AdaBoost without adding noise:
[0058] Since all data are time series data collected from actual coal mining construction sites, each sample contains 11 characteristic parameters of three types of influencing factors, namely tunnel design parameters, geological conditions and construction parameters. Tunnel design parameters include characteristic parameters x1 (tunnel cover thickness), x2 (cover span ratio); geological conditions include characteristic parameters x3 (soil internal friction angle), x4 (soil compression modulus), x5 (soil cohesion); construction parameters include characteristic parameters x6 (excavation speed), x7 (cutter head torque), x8 (shield thrust), x9 (cutter head speed), x10 (shield thrust), x11 (cutter head speed), x12 (cutter head speed), x13 (cutter head speed), x14 (cutter head speed), x15 (cutter head speed), x16 (cutter head speed), x17 (cutter head speed), x18 (cutter head speed), x19 (cutter head speed), x20 (cutter head speed), x21 (cutter head speed), x22 (cutter head speed), x23 (cutter head speed), x24 (cutter head speed), x25 (cutter head speed), x26 (cutter head speed), x27 (cutter head speed), x28 (cutter head speed), x29 (cutter head speed), x30 (cutter head speed), x31 (cutter head speed), x32 (cutter head speed), x33 (cutter head speed), x34 (cutter head speed), x35 (cutter head speed), x36 (cutter head speed), x37 (cutter head speed), x38 (cutter head speed), x39 (cutter head speed), x40 (cutter head speed), x41 (cutter head speed), x42 (cutter head speed), x43 (cutter head speed), x44 (cutter head speed), x45 (cutter head speed), x46 (cutter head speed), x47 (cutter head speed), x48 (cutter head speed), x49 (cutter head speed), x50 (cutter head speed). 10 (soil bin pressure), x 11(grouting volume), the upper limit (ub) and lower limit (lb) of the 11 characteristic parameters are given in Table 1. The next step is to build a safety assessment model with these 11 characteristic parameters as input and the mine settlement value as output.
[0059] Table 1 Upper and lower limits of input characteristic parameters
[0060]
[0061] With the above 11 characteristic parameters as input and the mine settlement value as output, 500 sets of mining engineering field data were collected, of which the first 400 sets of data were used as training data sets to build the model, and the last 100 sets of data were used as test data sets. The parameters of the benchmark model are set as follows:
[0062] [1] The BP neural network was used as the benchmark model and implemented in Python using the MLPRegressor function: the number of hidden layers in the network was 4, the number of nodes in each hidden layer was set to 10, the function was implemented as the activation function of the network, using relu, the optimization tolerance was set to 1e-4, and the maximum number of iterations was set to 10000;
[0063] [2] According to step (2), set up 30 random samplings from the original data set, each time taking 90% (400*90%=360) of the training data set as the sub-data set;
[0064] [3] After completing [2], take the top 10% (400*10%=40) groups of the training data set with the highest average error from the largest to the smallest as abnormal data, and the remaining 360 groups as normal data;
[0065] [4] Assign update weight μ to normal data normal =1, give update weight μ to abnormal data abnormal = 0.01 to construct a new training set. The final prediction model is implemented in Python using the function AdaBoostRegressor, where the number of iterations of the baseline model is set to 50 and the iterative learning rate is set to 0.5;
[0066] [5] Since the process of extracting training subsets is random, steps [2] and [3] are repeated 30 times to obtain a comprehensive result.
[0067] Figure 2 The original data division results are given in: The original training set is divided into 40 groups of abnormal data and 360 groups of normal data. The threshold between abnormal data and normal data is about 0.032, as shown in Figure 2 (b) Based on the normal data / abnormal data identification results, the effectiveness of the high-resilience AdaBoost method is verified on 100 test data sets. Figure 3The comparison results of the original data, AdaBoost and high-toughness AdaBoost methods are given. For 100 test data sets, the average error of AdaBoost is 0.0124 (see Figure 3 (b)), while the average prediction error of the method proposed in the present invention is 0.0038 (an increase of 69.35%). Figure 3 The table in (b) shows that the method directly using AdaBoost has 89 data in the final output falling within the error range of 0-0.03; 7 data falling within the error range of 0.03-0.06; and 4 data falling within the error range of 0.06-0.1. In the method proposed in the present invention, all data fall within the smaller error range of 0-0.03. This further proves the stability of the method proposed in the present invention.
[0068] Example 2
[0069] The following introduces the experimental results of the high resilience security assessment method based on AdaBoost when the training set is noised:
[0070] In order to further test the robustness of the proposed method to additional noise, different amounts of noise are added to the training dataset and the performance on the test set is verified. The specific parameter settings are given below:
[0071] [1] Taking the above 11 characteristic parameters as input and the mine settlement value as output, 500 sets of mining engineering field data were collected, and 400 sets of data were randomly selected from the complete data set as the training data set. 5%-50% additional noise was randomly added to these data, and the noise amount ranged from Δ=5% to Δ=50%; the method of adding noise was to randomly add [-Δ, +Δ];
[0072] [2] The remaining 100 sets of original data are used as test data sets;
[0073] [3] In the training data set, the top 10% (40 sets of data) with the highest to lowest average errors are selected as abnormal data, and the rest are normal data;
[0074] [4] The benchmark model is implemented using the MLPRegressor function. The number of hidden layers in the network is 4, the number of nodes in each hidden layer is 2, the activation function of the network is logistic, the optimization tolerance is 1e-4, and the maximum number of iterations is 10000.
[0075] [5] The prediction model is implemented using the AdaBoostRegressor function; the baseline model is iterated 10 times and the learning rate is set to 0.1.
[0076] The results are as follows Figure 4As shown in (a), based on Figure 4 The results in (a) can lead to the following conclusions:
[0077] 1. Compared with direct AdaBoost, the high-toughness AdaBoost method produces results with much smaller MAE. Figure 4 As shown in (b), with the increase of disturbance, the advantage of the proposed method always remains in the range of 35% to 50%. Specifically, when 20% of the noise data is added, the superiority of the proposed method is 36.1%; when 35% of the noise data is added, the superiority of the proposed method is 47.4%.
[0078] 2. Compared with the direct AdaBoost method, the advantage of the high-toughness AdaBoost method remains above 35%, and then further increases when the noise increases to 20%-40%. When the noise data is within 20%, the advantage of the proposed method shows a downward trend; the advantage drops from 41.7% to 36.7%. However, when the noise level increases to 20%-40%, the advantage shows an upward trend, rising again from 36.7% to 47.2%.
[0079] 3. When the noise ratio is further increased, the high-toughness AdaBoost method can produce better results. When the noise data increases to 45%, the MAE drops from 0.172 (identifying 10% of abnormal data) to 0.146 (identifying 40% of abnormal data). Compared with direct AdaBoost, the improvement rate of the proposed method increases from 43.8% to 52.4%. When 50% of noise data is added to the training dataset, the MAE decreases from 0.164 (identifying 10% of abnormal data) to 0.148 (identifying 40% of abnormal data). Compared with direct AdaBoost, the improvement of the proposed method increases from 46.1% to 51.3%.
[0080] Example 3
[0081] The following introduces the experimental results of the high resilience security assessment method based on AdaBoost when the test set is noised:
[0082] Since the proposed method has been tested and verified for its resilience to noise in the training dataset, the present invention will further test the proposed method's resilience to noise added to the test dataset. The specific parameter settings are given below.
[0083] [1] Noise was added only to the test dataset (no noise was added to the training dataset, and the original training dataset was used to build the benchmark model). The perturbation range was Δ=5% to Δ=50%. The perturbation was added by randomly adding [-Δ, +Δ].
[0084] [2] The number of hidden layers of the benchmark model is set to 4, with 10 neurons in each layer; the optimization tolerance tol is set to 0; the activation function activation = 'relu'; the maximum number of iterations max_iter is set to 10000;
[0085] [3] The prediction model is implemented using the AdaBoostRegressor function. The baseline model is iterated 20 times, and the function is implemented as n_estimators = 20. The learning rate is set to 1, and the function is implemented as learning_rate = 1.
[0086] The comprehensive results are as follows Figure 5 As shown in the figure, according to the results, it can be found that the proposed method is better than direct AdaBoost under all these conditions. More specific conclusions are summarized as follows:
[0087] 1. When the noise in the test data set is small, i.e., increasing from 1% to 3%, it can be observed that the improvement rate of the proposed method increases steadily, which proves the superiority of the method of the present invention in low interference conditions.
[0088] 2. On the contrary, when the noise in the test dataset is larger (i.e., from 5% to 15%), the advantage of the proposed method decreases, but it still remains above 10%. This is because the noise level is too high, destroying the original data distribution in the test dataset. When the noise increases from 5% to 15%, the improvement rate of the proposed method decreases from 20.4% to 14.4%.
[0089] Figure 6 The specific results when 15% noise is added to the test set are shown. Based on the results in the figure, the following conclusions can be drawn:
[0090] (1) Based on Figure 6 (a) It can be seen that the output value of the AdaBoost method directly cannot track the real output well, while the present invention (high-toughness AdaBoost) can better track the real output. This shows that the method proposed by the present invention has better toughness than the direct AdaBoost method.
[0091] (2) Based on Figure 6 (b) It can be seen that for 100 sets of test data, the average error of AdaBoost is 0.9479, while the average prediction error of the present invention is 0.4789 (an increase of 49.48%). Figure 6The table in (b) shows that the method directly using AdaBoost has 53 data in the final output falling within the range of 0 to 1, and 47 data falling within the range of 1 to 2. However, in the method proposed in the present invention, 95 data fall within the range of 0 to 1, and 5 data fall within the range of 1 to 2. This further proves the stability of the present invention.
[0092] The case study results of the present invention show that the direct use of the AdaBoost method has a large output error and unstable results. However, the high-resilience AdaBoost method that re-weights the abnormal data in the data set has a small output error and is relatively stable. The specific results are shown in Table 2 below.
[0093] Table 2 Average absolute errors of the two methods in three cases
[0094]
[0095] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. A high-resilience safety assessment method for complex mining projects based on the AdaBoost algorithm, characterized by The method comprises the following steps: (1) Data collection and data set division: Tunnel design parameters, geological condition parameters and construction parameters are used as input parameters of the model, and the data of the input parameters are collected from the time series data collected from the actual coal mining construction site to construct a complete data set D; the first 80% of the complete data set D is divided into a training set D in sequence. T , the remaining 20% is divided into the test set D V ; (2) From the training set D T Randomly select S sub-datasets from D to build S sub-models; T Randomly select S sub-datasets in the same proportion To build S sub-models, these sub-datasets The majority of data in the complete training set and the number of sub-models is at least 30, that is, S ≥ 30, in order to reduce random errors in the modeling process; the extraction of training sub-data sets in this step allows repeated extraction, that is, the intersection of any two training subsets is not allowed to be empty, and it is only necessary to ensure that the sizes of the subsets are the same; (3) Computational sub-model The error and weight of (3.1) Calculate sub-model error; First, for each sub-model Calculate the output error for each set of data in As shown below: Among them, p represents the sub-model The size of the data in , p∈[1,360] and p is an integer, y p Represents the true output of the pth group of data; Represents the output of the p-th group of data in the sub-model s; if the p-th group of data does not appear in the sub-model In Then, the BP neural network model is used to calculate the mean absolute error percentage (MAPE) of the sth sub-model. s ), as follows: Where P represents the P groups of data used to construct the s-th sub-model, P is 360, p represents the amount of data in the s-th sub-model, p∈[1,360] and p is an integer; (3.2) Calculate sub-model weights Calculate the MAPE of the sub-model s to represent their ability to identify the corresponding sub-training datasets, and MAPE s ∈[0,100%]; where MAPE s =0 indicates the highest modeling ability of the model, MAPE s =100% represents the minimum modeling capability of the model; therefore, the weight of the sth sub-model is obtained according to the following formula: In s =1-MAPE s (3) Among them, w s ∈[0,1]; (3.3) Calculate the average error of all P groups of data; The average error of all P groups of data is calculated by the output error of each group of data and the weight of the sub-model, as shown below: in, represents the output error of the pth group of data in the sth sub-model; w s represents the weight of the s-th sub-model; s * Indicates that this submodel is not constructed using the p-th group of data; if the p-th group of data is not used to construct the s-th * sub-models, then e p =0; the above formula can also be converted into the following formula: (4) Weight the data and build a prediction model; use the sub-model calculated in step (3) The average error of each group of data is arranged in descending order, the first 10% of the data is defined as abnormal data, and the remaining 90% of the data is defined as normal data; normal data helps to establish an accurate safety assessment model, and its weight coefficient is set to 1; while abnormal data will interfere with the safety assessment modeling process, and should be given a smaller weight, and its weight coefficient is set to 0.01; after obtaining the updated weights of all data, the weights are normalized to avoid the sum of the weights being greater than 1, and the training set D T After weighting each set of data in, a new weighted data set is returned To build the final prediction model; (5) Verify on the test set; as the data weights are updated, the entire weighted training set is used to construct the final prediction model, and the prediction model finally constructed in step (4) is verified on the test set.
2. The high-toughness safety assessment method for complex mining engineering based on the AdaBoost algorithm according to claim 1 is characterized by: The input parameters of the model in step (1) include tunnel design parameters: tunnel cover thickness x1, cover span ratio x2; geological condition parameters: soil internal friction angle x3, soil compression modulus x4, soil cohesion x5; construction parameters: excavation speed x6, cutterhead torque x7, shield thrust x8, cutterhead speed x9, soil bin pressure x 10 , grouting volume x 11 .
3. The high-toughness safety assessment method for complex mining engineering based on AdaBoost algorithm according to claim 1 is characterized by: The specific method of extracting sub-datasets and constructing sub-models in step (2) is as follows: each time from the complete training set D T 90% of the data are randomly selected to form a sub-dataset This step is repeated 30 times to obtain 30 sub-datasets. Each sub-dataset All D T A subset of 4. The high-toughness safety assessment method for complex mining engineering based on AdaBoost algorithm according to claim 1 is characterized by: The specific steps of data weighting and building a prediction model in step (4) are as follows: First, according to the average error of the pth group of data calculated The pth group of data is weighted as shown in the following formula: Among them, w p represents the update weight of the pth group of data, w p,0 Represents the initial weight of the pth group of data, the default is 1; μ counterintuitive Represents the weight coefficient of abnormal data; μ intuitive Represents the weight coefficient of normal data; for abnormal data, define μ counterintuitive =0.01; for normal data, the default weight is kept unchanged, that is, μ intuitive =1; after obtaining the update weights of all data, normalize the update weights, that is, After updating the weights for each set of data, return a new weighted training set It is used to train AdaBoost and generate the final high-resilience AdaBoost model, where the objective function of the model is as follows: in, It has been given in (3.1); y p Represents the true output of the pth group of data; It has been given in the above formula.
5. The high-toughness safety assessment method for complex mining engineering based on AdaBoost algorithm according to claim 1 is characterized by: The performance indicator verified on the test set in step (5) is MAE, which is an indicator used to evaluate the accuracy of the model. It measures the average difference between the predicted value and the true value. The smaller the indicator, the better the model prediction ability. The calculation formula is as follows: Among them, Q represents the Q group of data in the test set, Q is 100, q represents the size of the data in the test set, q∈[1,100] and q is an integer, y q Represents the true output of the qth group of data in the test set; Represents the output of the qth group of data in the test set in the prediction model.
Citation Information
Patent Citations
Method for preprocessing data in mine based on Bagging improvement
CN114862612A