Coal mine goaf carbon storage suitability evaluation and prediction method based on machine learning
By constructing a feature dictionary and training an integrated machine learning model, the consistency and repeatability issues of traditional methods in goaf evaluation are solved, realizing intelligent and accurate evaluation of the suitability of carbon sequestration in goaf, improving evaluation accuracy and reliability, and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNIV OF MINING & TECH
- Filing Date
- 2026-03-03
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional methods for evaluating carbon dioxide geological storage sites suffer from poor consistency and repeatability in goaf environments, difficulty in describing complex nonlinear relationships and simulating dynamic changes, and an inability to accurately assess potential risks during the storage process.
A machine learning-based approach is adopted to construct a feature dictionary, collect multi-source data, and train an ensemble machine learning model, including random forest, support vector machine, extreme gradient boosting model and logistic regression model, to evaluate the suitability of carbon sequestration in goaf areas, thereby achieving data-driven intelligent and objective evaluation.
It improves the accuracy and reliability of the evaluation, ensures the consistency and repeatability of the results, can accurately capture the complex nonlinear relationships between multi-dimensional evaluation indicators, reduce evaluation costs, and realize a systematic survey of the carbon sequestration potential of large-scale goaf resources.
Smart Images

Figure CN122134196A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of suitability assessment and prediction technology, and particularly relates to a machine learning-based method for assessing and predicting the suitability of carbon sequestration in coal mine goaf areas. Background Technology
[0002] In action plans to address global climate change, carbon dioxide capture, utilization, and storage (CCUS) is a key and fundamental technology, with carbon dioxide geological storage (CGS) being a crucial component of CCUS technology. With increasingly stringent carbon emission controls, finding safe, efficient, and economical carbon storage sites has become a top priority. Meanwhile, the transformation, utilization, and stability assessment of mined-out areas—a potential underground space resource—remain a significant technical challenge.
[0003] Currently, the evaluation methods for carbon dioxide geological storage sites mainly draw on traditional geological evaluation methods from oil and gas reservoir exploration and groundwater hydrology. In the site selection process, the weights of elements at each level are obtained and ranked using the analytic hierarchy process (AHP) and multi-factor spatial overlay method, thereby evaluating and classifying the grid of the study area. Finally, a weighted calculation is used to determine the suitability of carbon dioxide geological storage sites in the study area. While these methods have shown some practicality in preliminary regional screening, they reveal many limitations when dealing with complex, data-sparse goaf environments.
[0004] Traditional analytic hierarchy process (AHP) relies heavily on expert experience, leading to significant subjectivity in the allocation of indicator weights. In complex and highly uncertain geological environments like goaf areas, different experts may have significantly different judgments on the importance of the same indicator, resulting in poor consistency and low repeatability of evaluation results. Furthermore, the suitability of carbon sequestration in goaf areas is influenced by a variety of interacting factors, including geological structure, fracture development, rock mass mechanical properties, and hydrogeological conditions. These factors exhibit complex nonlinear relationships, and traditional multi-factor superposition methods, based on linear weighting assumptions, struggle to accurately describe these complex interactions, leading to discrepancies between evaluation results and actual suitability. Moreover, carbon dioxide injection into goaf areas triggers a series of physicochemical changes, including fluid pressure variations, rock-fluid interactions, and fracture propagation. These dynamic processes directly impact the long-term safety and stability of carbon sequestration. Therefore, traditional evaluation methods cannot simulate and predict these temporal changes, making it difficult to assess potential risks during the sequestration process.
[0005] In recent years, machine learning algorithms have been widely used in the field of carbon sequestration, demonstrating unique advantages in processing complex geological systems and multi-source data. Summary of the Invention
[0006] The purpose of this invention is to provide a machine learning-based method for evaluating and predicting the suitability of carbon sequestration in coal mine goaf areas, aiming to solve the problems mentioned in the background art.
[0007] The present invention is implemented as follows: a machine learning-based method for assessing and predicting the suitability of carbon sequestration in coal mine goaf areas includes the following steps:
[0008] S1. Constructing a feature dictionary: Construct a feature dictionary containing 27 evaluation factors from five dimensions, namely, mine safety, formation stability, formation sealing, resource storage, and economic suitability, to characterize the suitability of carbon sequestration in goaf areas.
[0009] S2. Data Acquisition and Preprocessing: Collect historical exploration data, real-time monitoring data, and regional socio-economic data of the target goaf area, and process them in a unified manner to obtain a dataset.
[0010] S3. Construct and train an ensemble machine learning model: Use random forest, support vector machine, extreme gradient boosting, and logistic regression models as training models for the dataset, and train the model using a dataset with labeled suitability levels to obtain the optimal model.
[0011] S4. Conduct suitability level evaluation: Based on the standardized feature data of the goaf to be evaluated, the trained optimal model is used to process the data to obtain the probability that the goaf to be evaluated belongs to each suitability level, and the final evaluation result is output.
[0012] This invention employs a data-driven integrated machine learning model to achieve objectivity and intelligence in the evaluation process of carbon sequestration suitability in goaf areas, effectively overcoming the excessive reliance on expert experience in traditional methods and ensuring the consistency and repeatability of evaluation results.
[0013] The embodiments of the present invention can accurately capture the complex nonlinear relationships between multi-dimensional evaluation indicators, significantly improving the evaluation accuracy and reliability. At the same time, through systematic data preprocessing and feature engineering, it realizes the efficient integration of multi-source heterogeneous data such as geological exploration, mining history, real-time monitoring and socio-economic data, expanding the evaluation dimensions from a single geological factor to a comprehensive consideration of mining area, strata and economy.
[0014] The method provided by the embodiments of the present invention significantly improves evaluation efficiency and reduces evaluation costs, making it possible to conduct a systematic survey and selection of carbon sequestration potential for resources in large-scale goaf areas. Attached Figure Description
[0015] Figure 1 The flowchart illustrates a machine learning-based method for evaluating and predicting the suitability of carbon sequestration in coal mine goaf areas, as provided in this embodiment of the invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.
[0018] Example 1, such as Figure 1 The flowchart shown is a machine learning-based method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs, according to an embodiment of the present invention. The method includes the following steps:
[0019] (1) Constructing a feature dictionary: Based on five dimensions of mining area safety, formation stability, formation sealing, resource storage and economic suitability, a feature dictionary containing multi-source evaluation factors is constructed to characterize the suitability of carbon sequestration in mining areas. The features of the mining area safety dimension include: mining subsidence level, maximum surface subsidence rate, absolute gas emission, water inflow level, rockburst risk level and earthquake impact index.
[0020] The characteristics of the geological stability dimension include: distance from adjacent mines, coal seam thickness, mining depth ratio, goaf volume, and surface load intensity.
[0021] The characteristics of the stratigraphic sealing dimension include: fault distance, structural complexity index, permeability grade, roof lithology type, roof thickness, and roof compressive strength;
[0022] The characteristics of resource restorability include: remaining coal reserves, volume of water accumulation in goaf areas, CO2 adsorption capacity, and coalbed methane content.
[0023] The characteristics of the economic suitability dimension include: distance to the nearest carbon source, distance to major transportation routes, storage costs, ecological and environmental impact, social benefits, and economic benefits.
[0024] (2) Data preprocessing and dataset construction: missing value processing, outlier removal, numerical normalization, category feature encoding, and class imbalance processing are performed on the collected historical exploration data, real-time monitoring data and regional socio-economic data. The training set and validation set are divided according to the preset ratio to form a standardized dataset.
[0025] (3) Construct and train machine learning models: Use the model with the highest accuracy among RandomForest, Support Vector Machine (SVM), Extreme Gradient Boosting (XGBoost) and Logistic Regression to train the model. Use 5-fold hierarchical cross-validation and use ROC-AUC, F1 score and overall accuracy as the model selection index to evaluate the model performance. Select the model with the highest prediction accuracy as the optimal model.
[0026] (4) Model optimization: Optimize the parameters of the optimal model based on grid search, Bayesian optimization or other hyperparameter tuning methods to improve the model's generalization ability;
[0027] (5) Feature data check: The subsidence level, water inflow level, rock burst risk level, structural complexity, and permeability level of the goaf are used as feature data and are checked before model prediction. Specifically, this includes:
[0028] Feature completeness check;
[0029] Check the reasonableness of the feature range;
[0030] Feature physical consistency check;
[0031] If the features meet the requirements, proceed to step (6); if they do not meet the requirements, give a direct suggestion that the evaluation is not appropriate.
[0032] (6) Goaf suitability prediction analysis: The standardized characteristic data of any goaf are input into the optimal model to obtain the probability distribution of the goaf belonging to each suitability level. The probability distribution output by the model is a four-dimensional vector:
[0033] ;
[0034] in, These represent the probabilities of being suitable for sealing, relatively suitable for sealing, generally suitable for sealing, and unsuitable for sealing, respectively.
[0035] In the case of incomplete feature data, the missing values are estimated and the compensation prediction results are obtained through the missing feature inference module of the optimal model.
[0036] (7) Output carbon sequestration suitability level: The final carbon sequestration suitability level of the goaf is determined based on the maximum probability, wherein the level includes: suitable for sequestration, relatively suitable for sequestration, generally suitable for sequestration and unsuitable for sequestration.
[0037] Example 2: A machine learning-based suitability evaluation system for carbon sequestration in goaf areas, used to implement the above method, comprising:
[0038] Feature construction module: used to construct a dictionary of features suitable for carbon sequestration in goaf areas;
[0039] Data processing module: used to perform data cleaning, normalization, encoding, and sample balancing.
[0040] Machine learning training module: used for training and selecting the optimal model based on a multi-model framework;
[0041] Model optimization module: used to perform hyperparameter tuning on the optimal model;
[0042] Predictive analysis module: used to predict the suitability probability of the target goaf area;
[0043] Feature checking module: Used to check the completeness and rationality of input features;
[0044] Evaluation output module: Used to output the suitability level of carbon sequestration in goaf areas.
[0045] Example 3: Carbon sequestration suitability evaluation based on simulated goaf data:
[0046] To verify the feasibility and effectiveness of the method in this embodiment of the invention, a goaf area in a deep coal mine in North China was selected as the research object. A typical goaf area formed in the mining area was chosen as the evaluation unit. Based on historical exploration data, production monitoring data, and regional socio-economic conditions, a goaf carbon sequestration suitability evaluation dataset was constructed, and predictive analysis was performed according to the method proposed in this embodiment of the invention.
[0047] Basic conditions of the goaf:
[0048] The goaf is located at a depth of approximately 720m, with an average coal seam thickness of 4.2m and a goaf volume of approximately 1.15 × 10⁻⁶ m. 6 m 3 The lithology of the top plate is mainly sandstone-mudstone interbedded, with an average thickness of 18m and a uniaxial compressive strength of approximately 42MPa.
[0049] Feature data construction:
[0050] Following step S1 in Example 1, a feature dictionary was constructed, selecting 27 evaluation factors from five dimensions: mining area safety, formation stability, formation sealing, resource reservableness, and economic suitability. Some feature data are shown in Table 1.
[0051] Table 1
[0052]
[0053] The remaining feature data are generated using historical statistical data or simulations based on parameters from similar mining areas;
[0054] Data preprocessing:
[0055] The data is processed according to step S2 in Example 1, including:
[0056] Continuous features are normalized; categorical features such as "lithological type" are encoded using one-hot encoding; missing data are compensated using mean imputation; a dataset containing 120 samples is constructed, with the training set and validation set divided in an 8:2 ratio.
[0057] Model training and optimization:
[0058] Following step S3 in Example 1, the random forest model, support vector machine model, extreme gradient boosting model and logistic regression model were trained respectively, and the performance was evaluated using 5-fold hierarchical cross-validation.
[0059] Experimental results show that the random forest model has an overall accuracy of 87.6% on the validation set, an ROC-AUC of 0.91, and an F1 score of 0.88. Its overall performance is better than other models, so the random forest model is selected as the optimal prediction model.
[0060] Suitability prediction results:
[0061] The standardized characteristic data of the goaf are input into the optimal model, and the model outputs the probability distribution of its belonging to each suitability level as follows:
[0062] Suitable for sealing: 0.62
[0063] Suitable for sealing: 0.24
[0064] Generally suitable for sealing: 0.10,
[0065] Unsuitable for sealing: 0.04;
[0066] Based on the principle of maximum probability, the goaf was ultimately classified as "suitable for sealing".
[0067] Results analysis:
[0068] The prediction results show that the goaf has good carbon sequestration conditions in terms of formation sealing, structural stability and resource storage, and its economic suitability is within an acceptable range. This verifies the effectiveness and engineering applicability of the method proposed in the embodiments of the present invention in the evaluation of the suitability of carbon sequestration in goaf.
[0069] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A machine learning-based method for evaluating and predicting the suitability of carbon sequestration in coal mine goaf areas, characterized in that, Includes the following steps: S1. Constructing a feature dictionary: Construct a feature dictionary containing 27 evaluation factors from five dimensions, namely, mine safety, formation stability, formation sealing, resource storage, and economic suitability, to characterize the suitability of carbon sequestration in goaf areas. S2. Data Acquisition and Preprocessing: Collect historical exploration data, real-time monitoring data, and regional socio-economic data of the target goaf area, and process them in a unified manner to obtain a dataset. S3. Construct and train an ensemble machine learning model: Use random forest, support vector machine, extreme gradient boosting, and logistic regression models as training models for the dataset, and train the model using a dataset with labeled suitability levels to obtain the optimal model. S4. Conduct suitability level evaluation: Based on the standardized feature data of the goaf to be evaluated, the trained optimal model is used to process the data to obtain the probability that the goaf to be evaluated belongs to each suitability level, and the final evaluation result is output.
2. The method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs based on machine learning according to claim 1, characterized in that, In step S1, the five dimensions and 27 evaluation factors are as follows: The characteristics of mine safety dimensions include: goaf subsidence level, maximum surface subsidence rate, absolute gas emission, water inflow level, rockburst risk level, and earthquake impact index. The characteristics of the geological stability dimension include: distance from adjacent mines, coal seam thickness, mining depth ratio, goaf volume, and surface load intensity. The characteristics of the stratigraphic sealing dimension include: fault distance, structural complexity index, permeability grade, roof lithology type, roof thickness, and roof compressive strength; The characteristics of resource restorability include: remaining coal reserves, volume of water accumulation in goaf areas, CO2 adsorption capacity, and coalbed methane content. The characteristics of the economic suitability dimension include: distance to the nearest carbon source, distance to major transportation routes, storage costs, ecological and environmental impact, social benefits, and economic benefits.
3. The method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs based on machine learning according to claim 1, characterized in that, In step S2, the unified processing to obtain the dataset specifically includes: Data cleaning: Cleaning data, handling missing values, and removing outliers; Data standardization: Normalizing continuous features; Encoding transformation: One-hot encoding of categorical features; Sample balancing: To avoid model bias caused by uneven distribution of suitability levels, a standardized feature dataset is formed.
4. The machine learning-based method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs according to claim 1, characterized in that, In step S3, the suitability level labels in the dataset are divided into four levels: "suitable for sealing", "relatively suitable for sealing", "generally suitable for sealing" and "unsuitable for sealing".
5. The method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs based on machine learning according to claim 1, characterized in that, In step S3, the model is trained using a 5-fold stratified cross-validation method, which divides the dataset into training and validation sets in an 8:2 ratio, and selects the model with the highest accuracy as the final prediction model.
6. The method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs based on machine learning according to claim 1, characterized in that, Between steps S3 and S4, the following steps are also included: Model optimization: Optimize the parameters of the optimal model using hyperparameter tuning methods based on grid search or Bayesian optimization; Feature data verification: The subsidence level, water inflow level, rock burst risk level, structural complexity index, and permeability level of the goaf are used as feature data and verified before model prediction. If any of the feature data is abnormal, the goaf is determined to be unsuitable for carbon sequestration, and the quality of other data is no longer considered.
7. The method for evaluating and predicting the suitability of carbon sequestration in coal mine goafs based on machine learning according to claim 6, characterized in that, The specific checks performed before model prediction include: feature completeness check, feature range reasonableness check, and feature physical consistency check.