Method for predicting CO concentration and coal temperature in spontaneous combustion latency period of coal
Through the improved POA algorithm and LSTM neural network, combined with feature selection and model optimization, the problem of insufficient accuracy in the prediction of the latent period of coal spontaneous combustion was solved, high-precision CO concentration and coal temperature prediction was achieved, and timely coal spontaneous combustion warning was provided.
Patent Information
- Application Number
- CN202510840930.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Existing technologies lack accuracy in predicting the latent period of coal spontaneous combustion. Sensor calibration errors and data missing lead to prediction difficulties, and traditional methods cannot achieve timely and effective early warning.
An improved POA algorithm combined with LSTM neural network is used to optimize model parameters through dynamic convergence coefficient, memory pool and hybrid search strategy. The Spearman correlation coefficient and SHAP algorithm are used for feature selection to construct a CO concentration and coal temperature prediction model to improve the prediction accuracy.
The accuracy and stability of the prediction of CO concentration and coal temperature during the latent period of coal spontaneous combustion have been improved, which can identify the risk of coal spontaneous combustion in advance and enhance the accuracy and reliability of the prediction.
Smart Images

Figure CN120724045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of coal spontaneous combustion monitoring and early warning, and in particular to a method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion. Background Art
[0002] As one of the world's most important energy sources, coal plays an irreplaceable role in industrial production, transportation, and daily life. As a key energy source, coal mining, transportation, storage, and utilization are crucial to the stability of the energy supply. However, coal mining and use also face numerous safety and environmental challenges. In particular, spontaneous combustion of coal is a common occurrence in mine operations, leading to serious safety incidents such as gas and dust explosions, posing a threat to mine safety, environmental protection, and personnel safety. To effectively prevent the frequent occurrence of gas and dust explosions caused by coal spontaneous combustion, numerous monitoring and early warning methods have been proposed, with gas and temperature testing being the most widely used. Gas testing uses abnormalities in the volume fraction of a single or combined gas to warn of coal spontaneous combustion risks, while temperature testing uses the heat release characteristics generated by the heat release during the spontaneous combustion process to predict hazards. However, the incubation period of coal spontaneous combustion is typically long, and there are no obvious visible symptoms in the early stages. Therefore, relying on traditional monitoring methods cannot provide timely and effective early warnings. In particular, in the absence of obvious gas anomalies, potential fire risks in coal mines can be easily overlooked.
[0003] In the early stages of coal spontaneous combustion, changes in CO concentration and coal temperature are the most sensitive indicators, reflecting the initial signs of coal spontaneous combustion. In particular, CO, a key gas in the coal spontaneous combustion process, increases in its concentration are closely related to the pyrolysis and oxidation reactions of the coal. Changes in coal temperature, on the other hand, are directly linked to the heat release process of coal spontaneous combustion. Therefore, real-time monitoring and prediction of CO concentration and coal temperature can help identify coal spontaneous combustion risks early, provide timely warnings to coal mines, and avoid catastrophic safety accidents.
[0004] With the rapid development of big data and machine learning technologies, their application to the prediction and early warning of coal spontaneous combustion has become a growing trend. Machine learning methods can deeply exploit nonlinear relationships in monitoring data such as gas concentration and temperature, providing more accurate predictions than traditional methods. This is particularly effective in situations where the incubation period for coal spontaneous combustion is long and early warning is difficult. Combined with traditional monitoring methods, machine learning provides strong technical support for the accurate prediction and risk warning of coal spontaneous combustion disasters.
[0005] Although the theoretical research on coal spontaneous combustion prediction and forecasting is now very complete, the research on using machine learning and big data technology to predict the gas and coal temperature during coal spontaneous combustion is still in its infancy. For example, the monitoring error is large: due to the influence of the environment during long-term use, sensors are prone to calibration errors and drift, and even data loss. Existing sensors are often not sensitive enough in the case of low gas concentrations or small temperature changes, affecting the accuracy and reliability of the data. Summary of the Invention
[0006] In order to make up for the shortcomings of the existing technology, the present invention proposes a method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion, aiming to solve the problems of insufficient accuracy and prediction difficulty in the prediction of the latent period of coal spontaneous combustion. By introducing a dynamic convergence coefficient, a memory pool, a mechanism and a hybrid search strategy through an improved POA algorithm, the convergence process of the model is effectively accelerated, the inference time is reduced, and the robustness of the model is enhanced. In addition, by combining with the LSTM neural network, the present invention optimizes the hyperparameter combination of the model, improves the prediction accuracy, and can provide a scientific basis for the accurate prediction of coal spontaneous combustion disasters.
[0007] The technical solution adopted by the present invention to solve the technical problem is: a method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion, comprising the following steps:
[0008] S1: Import historical data of coal spontaneous combustion in the mine goaf into the prediction system to construct a time series dataset of CO concentration, coal temperature and preset indicators;
[0009] S2: Perform data preprocessing on the dataset, including missing value filling and data normalization, to improve model accuracy;
[0010] S3: Use the Spearman correlation coefficient method to determine the correlation between the preset indicators, and perform feature selection based on the size of the correlation coefficient;
[0011] S4: Perform feature importance analysis on the feature selection indicators of the model and use the SHAP algorithm to calculate the Shapley value of the features to obtain different indicators;
[0012] S5: After data preprocessing, the dataset is divided into training set, validation set, and test set in a ratio of 8:1:1. The LSTM neural network model is constructed and the model parameters are initialized.
[0013] S6: Use the improved POA algorithm to optimize the LSTM model parameters. When the improved POA-LSTM optimization model meets the preset conditions, select the hyperparameter configuration;
[0014] S7: Build an LSTM model for predicting CO concentration and coal temperature based on the output optimal parameter combination, and retrain and test the dataset;
[0015] S8: Train and establish the LSTM, WOA-LSTM, PSO-LSTM, POA-LSTM, and IPOA-LSTM algorithms in sequence, and compare the prediction results and prediction accuracy by comparing the evaluation indicators to verify the superiority of the IPOA-LSTM model.
[0016] Specifically, step S2 includes:
[0017] S21: Missing value filling, where the data filling method uses linear interpolation. In data preprocessing, linear interpolation assumes that the missing data points change linearly between the known data points before and after them. There are two known points (x1, y1) and (x2, y2). The value y at a point x between these two points is estimated. The formula for linear interpolation is:
[0018]
[0019] Where: x is the point to be interpolated, y1 and y2 are the values of the known data points, and x1 and x2 are the independent variable values of the known data points;
[0020] S22: Outlier processing, wherein an IQR-based outlier processing method detects and processes outliers in the data preprocessing stage; this includes calculating the quartiles of the data and using the IQR to determine whether the data contains outliers. The calculation steps are as follows:
[0021] Calculate the first quartile (Q1, 25% quantile) and the third quartile (Q3, 75% quantile) of the data;
[0022] Calculate IQR = Q3-Q1, where IQR represents the distribution range of the middle 50% of the data set;
[0023] The lower and upper thresholds of outliers are set as follows: Lower Bound = Q1 - 1.5 × IQR, Lower Bound = Q3 + 1.5 × IQR. Values below the lower threshold or above the upper threshold are considered outliers.
[0024] S23: Data normalization, used to reduce the range of model parameter updates during training, including the following calculation formula:
[0025]
[0026] Among them, x is the original data, x min and x max are the minimum and maximum values of the data respectively, and x′ is the normalized data.
[0027] Specifically, step S3 includes:
[0028] S31: Collect multi-dimensional data related to coal spontaneous combustion, including CO concentration, coal temperature, carbon dioxide concentration, oxygen concentration, methane concentration, and olefin concentration;
[0029] S32: Set a data matrix, where each row represents a sample, each column represents an indicator or target variable, and the target variable is the coal spontaneous combustion risk or spontaneous combustion index;
[0030] S33: The Spearman correlation coefficient is used to calculate the correlation between each indicator and the risk of coal spontaneous combustion, and the linear correlation between the relevant indicators is quantitatively evaluated. The Spearman correlation coefficient between each indicator is calculated using the following formula:
[0031]
[0032] Where: d i is the difference between the rankings of each pair of variables, n is the number of samples, and ρ is the Spearman correlation coefficient;
[0033] S34: Generate heatmap: Use the heatmap function in MATLAB to visualize the indicator weights, and based on the normalized results, obtain the Spearman correlation coefficient between each indicator.
[0034] Specifically, step S4 includes:
[0035] S41: Quantify the contribution of each feature in a single prediction;
[0036] S42: Given a prediction, use the Shapley value to calculate the marginal contribution of each feature, including calculations based on all possible feature subsets. Assuming there are m features, the calculation formula for the Shapley value is as follows:
[0037]
[0038] Where M represents the set of all features, P represents the subset of M that does not include feature j, |.| represents the number of elements in the set, and F x (P∪{j}) represents the predicted value of the sample point x when it contains the feature subset P and feature j, F x (P) represents the predicted value of the sample point x when the feature subset P is included, F x (P∪{j})-F x (P) measures the influence of feature j on the predicted value when controlling feature subset P. The SHAP value of feature j of sample point x is the weighted sum of the influence of feature j under all possible feature subsets P;
[0039] S43: Obtain the importance of different predictors in the prediction model by calculating feature importance, and calculate the average absolute SHAP value of feature j across all sample points
[0040]
[0041] Among them, X represents the sample set, express The absolute value of , len(X) represents the total number of samples, That is, the overall feature importance of feature j in the prediction model, The larger the value, the more important the feature j is.
[0042] Specifically, step S5 includes:
[0043] S51: input layer dimension;
[0044] S52: output layer dimension;
[0045] S53: hidden layer dimension;
[0046] S54: activation function,
[0047] S55: the number of neurons in the hidden layer,
[0048] S56: input layer time steps,
[0049] S57: learning rate;
[0050] S58: Number of iterations.
[0051] Specifically, step S6 includes:
[0052] S61: Run the POA algorithm, calculate the objective function value and randomly generate prey. In the search phase, update the position. In the hunting phase, update the position and save the optimal solution until the number of iterations is reached.
[0053] Among them, it includes population initialization, which is used to generate the initial population. The random value rand ensures that the initial positions of the pelicans are evenly distributed between the upper and lower bounds of the solution, covering the entire solution space;
[0054]
[0055] Among them, X i,j is the position of the i-th pelican in the j-th dimension; rand is a random number in the range [0,1]; u j Indicates the upper bound of the problem in the jth dimension, l j is the lower bound; N is the number of pelicans; m is the dimension of the problem;
[0056] S62: The pelican population in POA is represented by the matrix X, which is used to record the current solution of all pelicans for subsequent iterative updates and optimizations:
[0057]
[0058] Among them, x ij is the position of the i-th pelican, X is the matrix representation of the pelican population;
[0059] S63: The objective function value vector is in the form of , which is used to measure the quality of each pelican's solution. Each pelican individual represents a set of LSTM parameters, and the LSTM parameters are optimized to find the optimal solution.
[0060]
[0061] Among them, F i is the objective function value of the i-th pelican, and F is the vector of objective function values for the entire pelican population;
[0062] S64: The search phase, the way the pelican approaches its prey, is accurately modeled.
[0063] S65: hunting stage;
[0064] S66: Improve the POA algorithm by introducing dynamic convergence coefficient, memory pool mechanism and hybrid search strategy into the Pelican optimization algorithm.
[0065] The beneficial effects of the present invention are as follows:
[0066] The present invention proposes a method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion. This method is based on multi-factor indicators to jointly predict CO concentration and coal temperature to provide early warning for the progress of coal spontaneous combustion. It has high accuracy and stability in the CO concentration and coal temperature prediction tasks, can identify the risk of coal spontaneous combustion in advance, and improve the accuracy of the prediction of the latent period of coal spontaneous combustion. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0068] Figure 1 To improve the prediction flow chart of POA algorithm and LSTM neural network model;
[0069] Figure 2 This is the architecture diagram of the coal mine bundle pipe monitoring system;
[0070] Figure 3 This is the original data graph in the CO concentration data visualization;
[0071] Figure 4 Data graph for processing missing and outliers in CO concentration data visualization;
[0072] Figure 5 This is the normalized data graph for CO concentration data visualization;
[0073] Figure 6 It is the correlation coefficient graph of the preset indicators;
[0074] Figure 7 This is the CO concentration prediction map in the SHAP global interpretability feature summary;
[0075] Figure 8 This is the temperature prediction graph in the SHAP global interpretability feature summary;
[0076] Figure 9 It is the average impact ranking diagram of input variables SHAP;
[0077] Figure 10 This is the LSTM network structure diagram;
[0078] Figure 11 This is the IPOA-LSTM model training curve (staircase state);
[0079] Figure 12 This is the CO concentration prediction graph for the model prediction curve comparison;
[0080] Figure 13 Temperature prediction graph for comparison of model prediction curves;
[0081] Figure 14 The CO concentration evaluation chart predicted by different models in the radar chart is shown;
[0082] Figure 15 The coal temperature evaluation chart predicted by different models in the radar chart for different model prediction evaluation is shown. DETAILED DESCRIPTION
[0083] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following, in combination with specific implementation methods, further elaborates and illustrates the implementation and application process of a method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion proposed by the present invention.
[0084] like Figures 1 to 15 As shown, the process of implementation of the present invention is:
[0085] S1: Import historical data on spontaneous combustion in mine goafs into the prediction system to construct a time series dataset of CO concentration, coal temperature, and preset indicators (oxygen concentration, olefin concentration, etc.). The dataset is derived from 507 sets of publicly available coal mine experimental data, including CO concentration, coal temperature, oxygen concentration, C2H4 concentration, and C2H4 / C2H6. Using the MATLAB platform, a corresponding program module is compiled to analyze and process the data, providing reliable data support for the prediction and prevention of coal spontaneous combustion risks.
[0086] S2: Perform data preprocessing on the dataset, including missing value filling and data normalization, to improve model accuracy;
[0087] The specific implementation process of step S2 of the present invention includes:
[0088] S21: Missing value filling, where the data filling method uses linear interpolation. In data preprocessing, linear interpolation assumes that the missing data points change linearly between the known data points before and after them. There are two known points (x1, y1) and (x2, y2). The value y at a point x between these two points is estimated. The formula for linear interpolation is:
[0089]
[0090] Where: x is the point to be interpolated, y1 and y2 are the values of the known data point, and x1 and x2 are the independent variable values of the known data point; S22: outlier processing, wherein the outlier processing method based on IQR detects and processes outliers in the data preprocessing stage; including calculating the quartiles of the data and using IQR to determine whether the data contains outliers, wherein the calculation steps are as follows:
[0091] Calculate the first quartile (Q1, 25% quantile) and the third quartile (Q3, 75% quantile) of the data;
[0092] Calculate IQR = Q3-Q1, where IQR represents the distribution range of the middle 50% of the data set;
[0093] The lower and upper thresholds of outliers are set as follows: Lower Bound = Q1 - 1.5 × IQR, Lower Bound = Q3 + 1.5 × IQR. Values below the lower threshold or above the upper threshold are considered outliers.
[0094] Outliers can be caused by data collection errors, equipment failures, or environmental anomalies, which can severely impact data analysis and prediction results. Therefore, outliers must be detected and processed during the data preprocessing phase. To this end, this paper introduces an outlier processing method based on the Interquartile Range (IQR). This method calculates the quartiles of the data and uses the IQR to determine whether the data contains outliers.
[0095] S23: Data normalization, used to reduce the range of model parameter updates during training, including the following calculation formula:
[0096]
[0097] Among them, x is the original data, x min and x max are the minimum and maximum values of the data, respectively, and x′ is the normalized data. Machine learning algorithms are very sensitive to the scale of the input data. If the scale difference of the input data is too large, the model may experience gradient vanishing or gradient explosion during training, thereby affecting the learning process and causing the model to converge more slowly or even fail to converge. Through normalization, the range of model parameter updates during training is reduced, making the training process smoother. Usually, scaling the values of all feature data to a uniform range of [0,1] can effectively avoid this situation. In addition, in the inference stage, the normalized data allows the model to make predictions in a shorter time, supporting real-time monitoring and early warning.
[0098] Taking CO concentration data as an example, linear interpolation is used to fill missing values. The processed data has the same trend as the original data, which greatly reduces data noise. Figure 3-5 shown.
[0099] S3: The Spearman correlation coefficient method is used to determine the correlation between the preset indicators, and feature selection is performed based on the size of the correlation coefficient. Because the coal spontaneous combustion time series data set contains many indicators with weak correlation with coal spontaneous combustion, the Spearman correlation coefficient method is used to determine the correlation between the preset indicators, and feature selection is performed based on the size of the correlation coefficient. Indicators with weak correlation are eliminated, so that only the most important features are included in the model training, thereby optimizing the model structure, being more conducive to extracting the rules of time series data and improving the accuracy and stability of prediction.
[0100] The specific implementation process of step S3 of the present invention includes:
[0101] S31: Collect multi-dimensional data related to coal spontaneous combustion, including but not limited to CO concentration, coal temperature, carbon dioxide concentration, oxygen concentration, methane concentration, olefin concentration, etc.;
[0102] S32: Set a data matrix, where each row represents a sample and each column represents an indicator or target variable. The target variable is usually the coal spontaneous combustion risk or spontaneous combustion index.
[0103] S33: The Spearman correlation coefficient is used to calculate the correlation between each indicator and the risk of coal spontaneous combustion, and the linear correlation between the relevant indicators is quantitatively evaluated. The Spearman correlation coefficient between each indicator is calculated using the following formula:
[0104]
[0105] Where: d i is the difference between the rankings of each pair of variables, n is the number of samples, and ρ is the Spearman correlation coefficient;
[0106] S34: Generate heatmap: Use the heatmap function in MATLAB to visualize the indicator weights, and based on the normalized results, obtain the Spearman correlation coefficient between each indicator.
[0107] The Spearman correlation coefficient method was used to compare the correlation coefficients between CO concentration, experimental temperature and other preset indicators. The correlation coefficient calculation results are as follows: Figure 6 As shown,
[0108] CO, CO2, CO / O2, CO2 / O2, C2H4 / C2H6 are positively correlated with the experimental temperature, while O2 is negatively correlated with the experimental temperature. The reason for this may be that as the experimental temperature increases, the redox reaction is intensified, a large amount of oxygen is consumed, and a large amount of CO is produced. At the same time, high temperature promotes the oxidation of organic matter in coal to produce more carbon oxides.
[0109] Experimental temperature, CO2, C2H4, CO / O2, CO2 / O2, CH4, C2H4, and C2H6 are positively correlated with CO concentration, while O2 is negatively correlated with CO concentration. Temperature and pyrolysis gas have a greater impact on CO. The fundamental reason is that the higher the temperature, the faster the coal oxidation rate and the greater the CO concentration. At the same time, a large amount of pyrolysis gas as a by-product of CO is also produced. O2 acts as an inhibitory factor because incomplete combustion is more likely to occur in a low-oxygen environment, producing a large amount of CO, which leads to an increase in CO concentration.
[0110] S4: Perform feature importance analysis on the feature selection indicators of the model and use the SHAP (SHapley AdditiveexPlanations) algorithm to calculate the Shapley value of the features to obtain different indicators;
[0111] The importance of the feature in the prediction model, and then the weight of the feature input in the machine learning model through the average absolute SHAP value of the feature;
[0112] The specific implementation process of step S4 of the present invention includes:
[0113] S41: Quantify the contribution of each feature to a single prediction. SHAP is a model interpretation method based on game theory that is used to quantify the contribution of each feature to a single prediction. It draws on the idea of Shapley values, which are mainly used to allocate the contributions of each participant in a cooperative game. The core idea of the SHAP method is to calculate the marginal contribution of each feature to the prediction result, thereby obtaining the degree of importance of different indicators in the prediction model.
[0114] S42: Given a prediction model, use the Shapley value to calculate the marginal contribution of each feature, including calculations based on all possible feature subsets. Assuming there are m features, the calculation formula for the Shapley value is as follows:
[0115]
[0116] Where M represents the set of all features, P represents the subset of M that does not include feature j, |.| represents the number of elements in the set, and F x (P∪{j}) represents the predicted value of the sample point x when it contains the feature subset P and feature j, F x (P) represents the predicted value of the sample point x when the feature subset P is included, F x (P∪{j})-F x (P) measures the influence of feature j on the predicted value when controlling feature subset P. The SHAP value of feature j of sample point x is the weighted sum of the influence of feature j under all possible feature subsets P;
[0117] like Figure 7-8 As shown in the global interpretability feature summary diagram, each point represents the impact of the characteristic value of a sample on the prediction, and the color represents the size of the characteristic value, red represents high and blue represents low. The size represents the impact of each characteristic variable on the prediction of CO concentration and coal temperature;
[0118] In terms of the degree of influence on CO concentration prediction, the coal temperature is at the top, and the SHAP value distribution range is the widest, indicating that it has the most significant impact on CO concentration prediction; in addition, the SHAP values of O2 concentration, CO2 concentration, C2H4, CH4 and C2H6 values are also relatively high, indicating that carbon oxides and pyrolysis gases have a greater influence on the prediction of CO concentration; finally, the CO / O2, CO2 / O2 and C2H4 / C2H6 values are relatively small, indicating that the Graham coefficient and the pyrolysis gas ratio have almost no effect on CO concentration prediction.
[0119] In the prediction of coal temperature, the SHAP value of CO concentration is the largest, indicating that it has the most significant impact on the coal temperature prediction; followed by O2, CO2, CO / O2 and CO2 / O2, whose SHAP values are also large, indicating that carbon oxides and Graham coefficient have a greater impact on the model output; and there is a significant negative correlation between O2 concentration and SHAP value, indicating that an increase in O2 concentration will inhibit the increase in coal temperature and also affect the prediction of coal temperature; finally, the values of C2H6, C2H4, C2H4 / C2H6 and CH4 are relatively small, indicating that pyrolysis gas has a smaller impact on the prediction of CO concentration.
[0120] S43: Obtain the importance of different predictors in the prediction model by calculating feature importance, and calculate the average absolute SHAP value of feature j across all sample points
[0121]
[0122] Among them, X represents the sample set, express The absolute value of , len(X) represents the total number of samples, That is, the overall feature importance of feature j in the prediction model, The larger it is, the more important feature j is;
[0123] The significance analysis of indicators in the prediction model can help determine the degree of influence of each predictor on the result. Therefore, it can be calculated through the above steps. By using the SHAP algorithm, the contribution of each input variable to the model prediction can be quantified, the internal correlation logic of the model can be revealed, and the interpretability of the model can be improved. Figure 12-13 Clearly show the contribution of each feature variable to the prediction result. The average absolute SHAP value is shown in the figure. The average absolute SHAP value The weights of the features can be used as inputs to the prediction model. By optimizing the weight ratio of these indicators, the problems of category imbalance and noise data can be solved, the model can focus more on key information, the recognition accuracy of minority categories and the model performance can be improved, and the prediction reliability of CO concentration and coal temperature can be increased.
[0124] S5: After completing data preprocessing, the dataset is divided into training set, validation set, and test set in a ratio of 8:1:1. LSTM neural network model is constructed and model parameters are initialized. The core structure of LSTM consists of three main parts: forget gate, input gate, and output gate. These gates control the flow and retention of information, thereby effectively retaining and updating the memory in the network. The LSTM network structure is as follows: Figure 10 As shown;
[0125] Among them, first, the cell state is the key part of LSTM, which can be understood as the "long-term memory" of the network. It flows throughout the entire sequence and carries information. LSTM uses these gates to decide which information to keep and which to discard;
[0126] Second, LSTM uses a gating mechanism to decide which information to discard or retain. The output of each gate is a value between 0 and 1, indicating the proportion of information that should be retained or discarded. The input of each gate comes from the output of the previous layer and the input of the current time step.
[0127] Specifically, the forget gate vector of the model is expressed as:
[0128]
[0129] Among them, ft is the output of the forget gate, σ is the Sigmoid activation function, Wf is the weight matrix of the forget gate, [h t -1,X t ] is the concatenation of the previous hidden state and the current input, b f is the bias of the forget gate;
[0130] The input gate vector of the model is represented as:
[0131]
[0132] Among them, i t represents the output of the input gate, W i is the weight matrix of the input gate, b i is the input gate bias;
[0133] The candidate cell states for the model are:
[0134]
[0135] in, represents the new candidate cell state, tanh is the tanh activation function, W c is the weight matrix of the cell state, b C is the cell state bias;
[0136] The output gate vector of the model is represented as:
[0137]
[0138] Among them, O t represents the output of the output gate, h t-1 is the hidden state of the previous time step, X t is the input of the current time step, W o is the weight matrix of the output gate, b ois the bias of the output gate;
[0139] Third: Update of cell state, in LSTM, C t The update of the cell state is determined by the forget gate, input gate and candidate cell state:
[0140]
[0141] Among them, f t Control discarded information, i t The table controls the added information. is a candidate cell state;
[0142] Fourth, the final output h of the hidden state t is calculated through the output gate:
[0143]
[0144] Among them, h t is the implicit state at the current moment, O t Represents the output of the output gate, C t is the current cell state.
[0145] The specific implementation process of step S5 of the present invention includes:
[0146] S51: Input layer dimensions; for this model, multi-dimensional data related to coal spontaneous combustion will be input, including temperature, CO2, O2, C2H4, etc., that is, the input layer determines the specific input values according to the number of input indicators;
[0147] S52: Output layer dimension; this model is to predict the progress of coal spontaneous combustion, which is determined by CO concentration and coal temperature, so the dimension is 2;
[0148] S53: Hidden layer dimension; since the problem solved by the present invention is a multi-input single-output problem, which is relatively simple, the hidden layer is set to 1;
[0149] S54: Activation function. For nonlinear data, the network effect of using nonlinear activation functions such as ReLU is better.
[0150] S55: The number of hidden layer neurons determines the capacity and complexity of the LSTM model and represents the number of features the model can learn and remember. For simple tasks, it is usually recommended to start with a small value (such as 64) and gradually increase it to 128, 256, and so on, to observe the training speed and performance of the model.
[0151] S56: Input layer time steps. Time steps refer to the number of consecutive time points observed by the LSTM model. For short-term prediction tasks, a smaller time step of 5 can be used. The step size constraint is usually set to 5-10.
[0152] S57: Learning rate; The learning rate determines the step size of the parameter adjustment during each gradient update of the model. If the learning rate is too high, the model may oscillate during training and fail to converge. If the learning rate is too low, it may lead to very slow convergence or fall into a local optimal solution. Generally, a smaller initial learning rate is used, i.e. 0.001-0.01.
[0153] S58: Iterations. Iterations refer to the number of times the model completes training on the entire training set. At the end of each epoch, all data is passed to the model for forward and backward propagation. This is typically done starting with 50 to 100 epochs and combined with early stopping. When performance on the validation set no longer improves, training can be stopped to avoid wasting computing resources.
[0154] S6: Use the improved POA algorithm to optimize the LSTM model parameters. When the improved POA-LSTM optimization model meets the preset conditions, select the optimal hyperparameter configuration;
[0155] The specific implementation process of step S6 of the present invention includes:
[0156] S61: Run the POA algorithm, calculate the objective function value and randomly generate prey. In the search phase, the position is updated. In the hunting phase, the position is updated and the optimal solution is saved, and the maximum number of iterations is finally reached.
[0157] Among them, it includes population initialization, which is used to generate the initial population. The random value rand ensures that the initial positions of the pelicans are evenly distributed between the upper and lower bounds of the solution, thus covering the entire solution space.
[0158]
[0159] Among them, X i,j is the position of the i-th pelican in the j-th dimension; rand is a random number in the range [0,1]; u j Indicates the upper bound of the problem in the jth dimension, l j is the lower bound; N is the number of pelicans; m is the dimension of the problem;
[0160] S62: The pelican population in POA is represented by the matrix X, which is used to record the current solution of all pelicans for subsequent iterative updates and optimizations:
[0161]
[0162] Among them, xij is the position of the i-th pelican, X is the matrix representation of the pelican population;
[0163] S63: The objective function value vector is in the form of , which is used to measure the quality of each pelican's solution. Each pelican individual represents a set of LSTM parameters, and the LSTM parameters are optimized to find the optimal solution.
[0164]
[0165] Among them, F i is the objective function value of the i-th pelican, and F is the vector of objective function values for the entire pelican population;
[0166] S64: The search phase, or how the pelican approaches its prey, is precisely modeled so that the POA algorithm can effectively scan the search space and fully utilize its ability to explore different areas. At the same time, the prey's position in the search space is randomly set, which further improves the POA algorithm's exploration efficiency when dealing with precise search problems. This feature constitutes a key component of the POA algorithm. The following is a mathematical approximation of the above principles and hunting strategies:
[0167]
[0168] in, is the position of the i-th pelican in the j-th dimension after the update in the first phase (exploration phase), where I is an integer value in the range [1, 2]; P j is the position of the prey in the jth dimension; F p is the objective function value of the prey;
[0169] If the objective function value decreases at the new position, the POA algorithm adopts the new position of the pelican. This update is considered a valid update and can prevent the pelican from moving to a non-optimal position. The process can be described by the following formula:
[0170]
[0171] Where, is the new position of the i-th pelican, for The objective function value of
[0172] S65: During the hunting phase, this flat-surface flight strategy helps pelicans expand their attack range during the development phase. By simulating this behavior, individual pelicans can be moved to more advantageous locations within the prey capture area, thereby enhancing their local search capabilities. The program then checks whether the pelican's position changes during the hunting process to move to a more suitable location. The mathematical model of the pelican's behavior is described below.
[0173]
[0174] in, is the position of the i-th pelican in the j-th dimension after the second phase update; R is a random integer between [0,2]; R.(1-t / T) is The radius of the domain; t is the current number of iterations; T is the maximum number of iterations;
[0175] The mathematical model for updating the pelican's position during the capture phase is as follows:
[0176]
[0177] Where, is the new position of the i-th pelican; for The objective function value of
[0178] S66: Improved POA algorithm: This algorithm simultaneously introduces a dynamic convergence coefficient, a memory pool mechanism, and a hybrid search strategy into the Pelican Optimization Algorithm (POA). These three optimization strategies work together to find the best balance between global and local search, accelerating the search process, avoiding premature convergence, and improving the quality of the optimal solution.
[0179] Among them, in the traditional POA algorithm, the search process often lacks a balance between global search and local development, especially in the early and late stages of optimization, the step size of the algorithm remains unchanged. This will result in insufficient exploration ability of the individual in the solution space in the early stages of the search, making it difficult to cover a wide area and easily falling into local optimality. In the later stages of optimization, if the step size is large, the individual is likely to skip high-quality solutions and cannot effectively carry out fine local development. This paper proposes a dynamic convergence coefficient mechanism to dynamically adjust the search range and step size of the individual according to the different stages of the optimization process. Through the introduction of the dynamic convergence coefficient, the algorithm can achieve a smooth transition from the initial global extensive search to the later fine local development. Definition of dynamic convergence coefficient:
[0180]
[0181] Where t is the current iteration number; T is the total number of iterations; k is a parameter that controls the convergence speed, and its value range is [0.05, 0.2]. The larger the value, the faster the convergence.
[0182] Furthermore, in the traditional POA algorithm, individual position updates only rely on the information of the current population, ignoring the use of high-quality solutions found in the historical iteration process, which often results in the inability of high-quality solutions to continue to play a role in the optimization process, and then fall into local optimality. To this end, the present invention introduces a memory pool mechanism to save the historical high-quality solutions generated during the optimization process to prevent the algorithm from losing valuable information. The solutions in the memory pool are not only used for population updates, but also can guide individuals to search in a more optimal direction, enhancing the synergistic effect of global and local searches. Through the memory pool mechanism, historical information can be fully utilized, search efficiency can be improved, and the overall performance of the algorithm can be improved. The update rules of the memory pool are as follows:
[0183]
[0184] Where x memory is an individual randomly selected from the memory pool; r is the perturbation coefficient, which controls the introduction intensity of the memory pool solution;
[0185] Furthermore, in the traditional POA algorithm, individual updates rely on searching for solutions within the population, but this approach often fails to effectively utilize the global information of high-quality solutions, resulting in a single individual search direction, insufficient solution diversity, and an increased probability of falling into local optimality. In order to further improve the diversity and optimization capabilities of the population, the present invention introduces the idea of differential evolution (DE) into the individual update process. The hybrid search strategy combines the current optimal solution, the memory pool solution, and the differential information between individuals through differential evolution, prompting individuals to form diverse exploration paths in the search space. Through the hybrid search strategy, the algorithm can more efficiently balance global search and local development, significantly improving the search capability during the optimization process. The differential evolution formula is as follows:
[0186]
[0187] Where F is the differential scaling factor, ranging from [0.5, 1]; xbest1 and xbest2 are two different individuals randomly selected from the memory pool;
[0188] Furthermore, during the search phase, the algorithm needs to balance global exploration and local development capabilities. By dynamically adjusting the individual search step size and introducing a memory pool mechanism, individuals can leverage historical high-quality solutions for more efficient searches, ensuring a balance between global search and local development, and improving the algorithm's convergence accuracy and efficiency. The improved formula for the search phase is as follows:
[0189]
[0190] Where M j: A randomly selected solution in the memory pool; w: Contribution weight of the memory pool, which controls the strength of the memory pool, with a value range of [0.1, 0.5];
[0191] Furthermore, the hunting phase is mainly responsible for in-depth development in local areas while maintaining a certain level of global exploration capabilities. By comparing high-quality solutions and mining differential information, the search precision and diversity are improved. The individual update formula in the hunting phase is as follows:
[0192]
[0193] Where, and are two randomly selected solutions from the memory pool.
[0194] S7: Based on the output optimal parameter combination, an LSTM model for predicting CO concentration and coal temperature is constructed, and the dataset is retrained and tested. The model is trained and validated using the constructed dataset. The IPOA algorithm outputs the optimal LSTM parameter combination. Table 1 below shows the initial range and optimal values of the optimized parameters.
[0195] Table 1 shows the model hyperparameter values.
[0196] Serial number Hyperparameters Random initialization range Optimal parameters 1 Input layer time steps [5-10] 8 2 Number of neurons in the hidden layer [50-200] 135 3 Learning rate [0.001-0.01] 0.006 4 Number of iterations [100-300] 153
[0197] In order to intuitively compare the performance of the four optimization algorithms WOA-LSTM, PSO-LSTM, POA-LSTM and IPOA-LSTM, as shown in the figure: Figure 11 As shown in the figure, the fitness curves of four different models in the process of finding the optimal parameter combination are plotted. MSE represents the average value of the square of the error between the model's predicted value and the true value. The smaller the value, the smaller the prediction error and the better the model performance. Analysis shows that in the hyperparameter optimization with MSE as the objective function, the convergence curves of all algorithms go through three stages: a sharp decline, fluctuations, and steady convergence. Specifically, in the early stages of training, the algorithm globally explores the search space and the MSE drops sharply; as the number of iterations increases, the algorithm approaches the optimal solution area, and the MSE fluctuates due to the complexity of the objective function and the local sensitivity of the parameters; eventually, under the adjustment of the dynamic convergence factor and the search strategy of the optimization algorithm, the hyperparameters tend to stabilize and the MSE tends to stabilize, marking the end of the optimization process. The local graph shows that the MSE size of the four algorithms is ranked as follows:
[0198] WOA-LSTM>PSO-LSTM>POA-LSTM>IPOA-LSTM,
[0199] Among them, IPOA-LSTM achieved the smallest MSE and required the fewest iterations, reaching convergence after only 136 iterations. The other three algorithms all converged after 175 iterations. WOA-LSTM and PSO-LSTM were prone to falling into local optima, resulting in large and fluctuating MSEs. While POA-LSTM achieved some optimization results, it converged slowly. In contrast, IPOA-LSTM, through its rational parameter settings, dynamic convergence coefficient, and hybrid strategy, not only quickly found parameter combinations close to the global optimal solution, reducing training time, but also improving the model's prediction accuracy and stability.
[0200] S8: Train and establish LSTM, WOA-LSTM, PSO-LSTM, POA-LSTM and IPOA-LSTM algorithms in sequence, and compare the evaluation indicators (MAE, RMSE, R 2 ), judge the prediction results and compare the prediction accuracy to verify the superiority of the IPOA-LSTM model;
[0201] In order to evaluate the feasibility and accuracy of the prediction of this model, the test set results of the trained model are further solved by some evaluation indicators. For regression models, common evaluation indicators include mean absolute error MAE (Mean Absolute Error), root mean square error RMSE (Root Mean Square Error), determination coefficient R 2 (Rsquared) etc.
[0202] MAE represents the mean of the absolute errors between the true value and the predicted value, which can better reflect the error of the predicted value. The calculation formula is as follows:
[0203]
[0204] RMSE represents the deviation between the model's predicted value and the true value. The closer the value is to 0, the better the model's generalization performance and the higher the accuracy. Its calculation formula is as follows:
[0205]
[0206] R 2 The coefficient of determination reflects the degree of correlation between the dependent variable and the independent variable. The closer its value is to 1, the better the correlation between the predicted value and the true value. The calculation formula is as follows:
[0207]
[0208] In order to verify the performance of the IPOA-LSTM model, this paper established a prediction model of CO concentration and coal temperature during the latent period of coal spontaneous combustion based on LSTM, WOA-LSTM, PSO-LSTM, POA-LSTM and IPOA-LSTM algorithms according to the optimal parameter combination obtained by the algorithm in step S7. The prediction effects of the training set and test set after optimizing the LSTM model with different optimization algorithms are shown in the following figure. Figure 9 As shown in the figure, the unoptimized LSTM model has very poor prediction results in both the training and test sets. The prediction trends of the four improved LSTM models are basically consistent with the change trends of the original data, indicating that the optimization algorithm has certain advantages. From the local magnification diagram, we can see that the prediction results of WOA-LSTM, PSO-LSTM, and POA-LSTM differ greatly from the original values when the CO concentration fluctuates greatly, while the prediction results of IPOA-LSTM are always roughly consistent with the original data, with the smallest error.
[0209] In order to more intuitively observe the prediction results of the four models and quantitatively evaluate the performance of the models, a radar chart showing the model prediction evaluation was drawn based on the prediction results of the above training set and test set. Figure 14-15 It can be found that IPOA-LSTM has a higher R 2 The values of the IPOA-LSTM model are 0.95 and 0.97 respectively, which are closest to 1 compared with the other models. The MAE is 8 and 12, and the RMSE is 10 and 6. Compared with the other four models, the closer the value is to 0, the better the generalization performance of the model and the higher the accuracy. The above indicators fully demonstrate the effectiveness of the IPOA-LSTM model in improving prediction accuracy. It has high accuracy and stability in the CO concentration and coal temperature prediction tasks, can identify the risk of coal spontaneous combustion in advance, and improve the accuracy of the prediction of the coal spontaneous combustion latent period.
[0210] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion, characterized by: The steps include: S1: Import historical data of coal spontaneous combustion in the mine goaf into the prediction system to construct a time series dataset of CO concentration, coal temperature and preset indicators; S2: Perform data preprocessing on the dataset, including missing value filling and data normalization, to improve model accuracy; S3: Use the Spearman correlation coefficient method to determine the correlation between the preset indicators, and perform feature selection based on the size of the correlation coefficient; S4: Perform feature importance analysis on the feature selection indicators of the model and use the SHAP algorithm to calculate the Shapley value of the features to obtain different indicators; S5: After data preprocessing, the dataset is divided into training set, validation set, and test set in a ratio of 8:1:
1. The LSTM neural network model is constructed and the model parameters are initialized. S6: Use the improved POA algorithm to optimize the LSTM model parameters. When the improved POA-LSTM optimization model meets the preset conditions, select the hyperparameter configuration; S7: Build an LSTM model for predicting CO concentration and coal temperature based on the output optimal parameter combination, and retrain and test the dataset; S8: Train and establish the LSTM, WOA-LSTM, PSO-LSTM, POA-LSTM, and IPOA-LSTM algorithms in sequence, and compare the prediction results and prediction accuracy by comparing the evaluation indicators to verify the superiority of the IPOA-LSTM model.
2. The method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion according to claim 1, wherein: Step S2 includes: S21: Missing value filling, where the data filling method uses linear interpolation. In data preprocessing, linear interpolation assumes that the missing data points change linearly between the known data points before and after them. There are two known points (x1, y1) and (x2, y2). The value y at a point x between these two points is estimated. The formula for linear interpolation is: Where: x is the point to be interpolated, y1 and y2 are the values of the known data points, and x1 and x2 are the independent variable values of the known data points; S22: Outlier processing, wherein an IQR-based outlier processing method detects and processes outliers in the data preprocessing stage; this includes calculating the quartiles of the data and using the IQR to determine whether the data contains outliers. The calculation steps are as follows: Calculate the first quartile (Q1, 25% quantile) and the third quartile (Q3, 75% quantile) of the data; Calculate IQR = Q3-Q1, where IQR represents the distribution range of the middle 50% of the data set; The lower and upper thresholds of outliers are set as follows: Lower Bound = Q1 - 1.5 × IQR, Lower Bound = Q3 + 1.5 × IQR. Values below the lower threshold or above the upper threshold are considered outliers. S23: Data normalization, used to reduce the range of model parameter updates during training, including the following calculation formula: Among them, x is the original data, x min and x max are the minimum and maximum values of the data respectively, and x′ is the normalized data.
3. The method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion according to claim 1, wherein: Step S3 includes: S31: Collect multi-dimensional data related to coal spontaneous combustion, including CO concentration, coal temperature, carbon dioxide concentration, oxygen concentration, methane concentration, and olefin concentration; S32: Set a data matrix, where each row represents a sample, each column represents an indicator or target variable, and the target variable is the coal spontaneous combustion risk or spontaneous combustion index; S33: The Spearman correlation coefficient is used to calculate the correlation between each indicator and the risk of coal spontaneous combustion, and the linear correlation between the relevant indicators is quantitatively evaluated. The Spearman correlation coefficient between each indicator is calculated using the following formula: Where: d i is the difference between the rankings of each pair of variables, n is the number of samples, and ρ is the Spearman correlation coefficient; S34: Generate heatmap: Use the heatmap function in MATLAB to visualize the indicator weights, and based on the normalized results, obtain the Spearman correlation coefficient between each indicator.
4. The method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion according to claim 1, wherein: Step S4 includes: S41: Quantify the contribution of each feature in a single prediction; S42: Given a prediction, use the Shapley value to calculate the marginal contribution of each feature, including calculations based on all possible feature subsets. Assuming there are m features, the calculation formula for the Shapley value is as follows: Where M represents the set of all features, P represents the subset of M that does not include feature j, |.| represents the number of elements in the set, and F x (P∪{j}) represents the predicted value of the sample point x when it contains the feature subset P and feature j, F x (P) represents the predicted value of the sample point x when the feature subset P is included, F x (P∪{j})-F x (P) measures the influence of feature j on the predicted value when controlling feature subset P. The SHAP value of feature j of sample point x is the weighted sum of the influence of feature j under all possible feature subsets P; S43: Obtain the importance of different predictors in the prediction model by calculating feature importance, and calculate the average absolute SHAP value of feature j across all sample points Among them, X represents the sample set, express The absolute value of , len(X) represents the total number of samples, That is, the overall feature importance of feature j in the prediction model, The larger the value, the more important the feature j is.
5. The method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion according to claim 1, wherein: Step S5 includes: S51: input layer dimension; S52: output layer dimension; S53: hidden layer dimension; S54: activation function, S55: the number of neurons in the hidden layer, S56: input layer time steps, S57: learning rate; S58: Number of iterations.
6. The method for predicting CO concentration and coal temperature during the latent period of coal spontaneous combustion according to claim 1, wherein: Step S6 includes: S61: Run the POA algorithm, calculate the objective function value and randomly generate prey. In the search phase, update the position. In the hunting phase, update the position and save the optimal solution until the number of iterations is reached. Among them, it includes population initialization, which is used to generate the initial population. The random value rand ensures that the initial positions of the pelicans are evenly distributed between the upper and lower bounds of the solution, covering the entire solution space; Among them, Xi , j is the position of the i-th pelican in the j-th dimension; rand is a random number in the range [0,1]; uj represents the upper boundary of the problem in the j-th dimension, and lj is the lower boundary; N is the number of pelicans in the population; m is the dimension of the problem; S62: The pelican population in POA is represented by the matrix X, which is used to record the current solution of all pelicans for subsequent iterative updates and optimizations: Where xij is the position of the i-th pelican, and X is the matrix representation of the pelican population; S63: The objective function value vector is in the form of , which is used to measure the quality of each pelican's solution. Each pelican individual represents a set of LSTM parameters, and the LSTM parameters are optimized to find the optimal solution. Among them, F i is the objective function value of the i-th pelican, and F is the vector of objective function values for the entire pelican population; S64: The search phase, the way the pelican approaches its prey, is accurately modeled. S65: hunting stage; S66: Improve the POA algorithm by introducing dynamic convergence coefficient, memory pool mechanism and hybrid search strategy into the Pelican optimization algorithm.
Citation Information
Patent Citations
Coal mine driving face gas emission quantity prediction method based on KPCA-POA-LSTM model
CN115470887A
Modeling method for optimizing thermal error of main shaft of long short-term memory network by pelican algorithm
CN116861793A
Goaf CO concentration prediction method for coal spontaneous combustion risk prediction
CN118191233A
VMD-PCA-LSTM ultra-short-term wind power prediction method based on pelicans optimization algorithm
CN119272598A
Propylene mass fraction prediction method, system and medium
CN119940411A