CO2 leakage risk analysis method
By constructing a CO2 leakage risk analysis method, using sample sets and simulator optimization models, and combining the SHAP method, the high cost and environmental dependence problems of traditional methods are solved, and efficient and accurate CO2 leakage risk assessment and management are achieved.
Patent Information
- Application Number
- CN202510927997.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing CO2 leakage risk assessment methods have problems such as high cost, environmental dependence and high computing resource requirements in sensor monitoring, numerical simulation and geochemical imaging technology, making it difficult to achieve large-scale or long-term monitoring.
The CO2 leakage risk analysis method is adopted to obtain a sample set and randomly divide it to construct an initial prediction model. The simulator and data preprocessing are combined with the SHAP method to perform model verification and key factor analysis, optimize hyperparameter selection, and improve the generalization ability and prediction accuracy of the model.
It significantly improves the prediction accuracy of CO2 leakage risk, identifies key influencing factors, optimizes model parameters, adapts to different geological conditions, improves risk management efficiency, and reduces environmental risks and economic losses.
Smart Images

Figure CN120634271A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of CO2 leakage risk analysis in geological carbon sequestration, and in particular relates to a method for analyzing CO2 leakage risk. Background Art
[0002] Potential CO2 storage sites include deep saline aquifers, deep-sea formations, and depleted oil and gas reservoirs. Deep saline aquifers are considered the most promising due to their enormous storage capacity. However, CO2 storage in saline aquifers presents numerous challenges. One of these challenges is the density difference between the formation water and the injected CO2. The low-density CO2 rises within the reservoir and leaks into the atmosphere through the caprock. Therefore, before selecting a CO2 storage site, rapid and accurate assessment of CO2 leakage risk is essential. Current CO2 leakage assessment methods primarily rely on sensor monitoring, numerical simulations, and geochemical imaging. However, in practice, sensor monitoring is significantly affected by environmental factors and has high deployment and maintenance costs, making it unsuitable for large-scale or long-term monitoring. Numerical simulations, on the other hand, require extensive geological data, significant iterative computation time, and significant computational resources. Geochemical imaging technology is expensive and requires high site conditions and operator expertise. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes a CO2 leakage risk analysis method, which is suitable for large-scale or long-term monitoring and reduces iterative calculation time and computing resources.
[0004] To achieve the above objectives, the present invention provides a method for analyzing CO2 leakage risk, comprising:
[0005] Obtain a sample set of CO2 leakage;
[0006] Randomly dividing the sample set to obtain a first sample set and a second sample set;
[0007] Constructing an initial CO2 leakage risk prediction model based on the first sample set;
[0008] Using the second sample set to test and verify the initial CO2 leakage risk prediction model to obtain an optimal CO2 leakage risk prediction model, wherein the CO2 leakage risk prediction model is used to output a CO2 leakage distance and make a prediction based on the leakage distance;
[0009] The SHAP method is used to analyze the prediction results and obtain the key factors affecting the CO2 leakage risk.
[0010] Optionally, the sample set for obtaining CO2 leakage includes:
[0011] Obtain variables affecting CO2 leakage;
[0012] The influencing factors are processed by using Latin hypercube sampling method to obtain the sample set.
[0013] Optionally, the influencing factor variables are processed using a Latin hypercube sampling method to obtain the sample set, including:
[0014] S1. Divide the influencing factor variable into equal probability intervals to obtain several sub-intervals;
[0015] S2. Randomly select a variable from each subinterval to form a sample point;
[0016] S3. Repeat S2 until a preset number of sample points are obtained to form the sample set.
[0017] Optionally, before constructing the initial CO2 leakage risk prediction model based on the sample set, the following steps are included:
[0018] simulating the sample set using a simulator to obtain a simulated CO2 data set, wherein the simulator accurately simulates the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase change in porous media;
[0019] Anomaly detection and data standardization are performed on the simulated CO2 data set to obtain a processed sample set.
[0020] Optionally, constructing an initial CO2 leakage risk prediction model using the first sample set, where obtaining the initial CO2 leakage risk prediction model includes:
[0021] Determine the hyperparameters and corresponding value ranges of the initial CO2 leakage risk prediction model:
[0022] Combining the hyperparameters with corresponding hyperparameters to obtain a hyperparameter combination solution;
[0023] Use cross-validation to verify the hyperparameter combination scheme and obtain the optimal parameter combination as the final hyperparameter setting;
[0024] Based on the final hyperparameter setting, the initial CO2 leakage risk prediction model is obtained.
[0025] Optionally, using the second sample set to test and verify the initial CO2 leakage risk prediction model to obtain the optimal CO2 leakage risk prediction model includes:
[0026] By calculating the performance indicators of the test verification results, the parameters of the initial CO2 leakage risk prediction model are iteratively optimized until the prediction error of the model on the second sample set no longer decreases, thereby determining the optimal CO2 leakage risk prediction model.
[0027] Optionally, the CO2 leakage risk prediction model is:
[0028]
[0029] Among them, y(x) is the prediction result, α i is the Lagrange multiplier, K(x,x i ) is the kernel function, b is the bias term, i is the index used to traverse the sample points, N is the total number of sample points, x is the input variable to be predicted, x i is the input value of the i-th training sample.
[0030] Optionally, the SHAP method is used to analyze the prediction results:
[0031]
[0032] in, is the Shapley value of feature m, which reflects the average marginal contribution of feature m to the model output. F represents the set of all features, S is the feature subset that does not contain feature m, |S| represents the number of features in subset S, |F| is the number of all features, u(S) is the output value when the model only uses feature subset S for prediction, and u(S∪{m}) is the predicted output value of the model after adding feature m to subset S.
[0033] Compared with the prior art, the present invention has the following advantages and technical effects:
[0034] 1. Improve prediction accuracy:
[0035] This method obtains a sample set of CO2 leaks and randomly divides it into a first sample set and a second sample set. Using the first sample set to construct an initial prediction model, and then testing and verifying it using the second sample set, this method effectively improves the model's generalization and prediction accuracy. In particular, by simulating the sample set with a simulator and performing preprocessing on the simulated data (such as anomaly detection and data normalization), the quality of the sample data is further optimized, providing more accurate input for model training and significantly improving the accuracy of CO2 leak risk prediction.
[0036] 2. Efficiently identify key influencing factors:
[0037] This paper uses the SHAP (SHapley Additive exPlanations) method to analyze prediction results, systematically identifying key factors influencing CO2 leakage risk and their contributions. Based on Shapley values, the SHAP method quantifies the average marginal contribution of each feature to the model output, providing a deeper explanation and basis for risk assessment. This method not only helps users understand the inherent logic of model predictions but also guides the control and optimization of key factors in actual projects to reduce CO2 leakage risks.
[0038] 3. Optimize model parameter selection:
[0039] This method uses the first sample set to perform supervised learning on the initial prediction model. By determining hyperparameters and their value ranges and verifying the hyperparameter combination scheme using cross-validation, it can quickly find the optimal parameter combination. This method avoids the blindness and inefficiency of traditional parameter selection methods, significantly improves the efficiency and effectiveness of model training, and ensures that the resulting CO2 leakage risk prediction model has higher performance and reliability.
[0040] 4. Strong adaptability:
[0041] This method, based on a variety of advanced techniques and algorithms (such as Latin hypercube sampling, simulator simulation, supervised learning, and the SHAP method), can adapt to the needs of CO2 leakage risk analysis in diverse geological conditions, engineering scenarios, and data characteristics. By flexibly adjusting the sample acquisition method, model construction and optimization process, and key factor analysis methods, the method can be widely applied to leakage risk assessment of different types of CO2 geological storage, demonstrating its high versatility and adaptability.
[0042] 5. Improve risk management efficiency:
[0043] By effectively predicting CO2 leakage risks and identifying key factors, this method can provide a scientific basis for the planning, design, operation, and supervision of CO2-related projects. For example, in CO2 geological storage projects, potential leakage risk areas can be identified in advance, storage plans can be optimized, and environmental risks and economic losses can be reduced.
[0044] In summary, the CO2 leakage risk analysis method of the present invention has significant technical effects in improving prediction accuracy, identifying key factors, optimizing model parameter selection, strong adaptability, and improving risk management efficiency, and can provide strong technical support for risk assessment and management in CO2-related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0046] Figure 1 This is a flow chart of a method for analyzing CO2 leakage risk according to an embodiment of the present invention;
[0047] Figure 2 is a data variable distribution diagram of an embodiment of the present invention;
[0048] Figure 3 is a schematic diagram of an outlier detection result according to an embodiment of the present invention;
[0049] Figure 4 Schematic diagrams of comparative prediction results of the LSSVM model of an embodiment of the present invention and other machine learning models in a training set, (a) is a schematic diagram of the prediction results of the LSSVM model in the training set, (b) is a schematic diagram of the prediction results of the GRNN model in the training set, (c) is a schematic diagram of the prediction results of the MLPNN model in the training set, and (d) is a schematic diagram of the prediction results of the ELM model in the training set;
[0050] Figure 5 Schematic diagrams of comparative prediction results of the LSSVM model of an embodiment of the present invention and other machine learning models on a test set, (a) is a schematic diagram of the prediction results of the LSSVM model on the test set, (b) is a schematic diagram of the prediction results of the GRNN model on the test set, (c) is a schematic diagram of the prediction results of the MLPNN model on the test set, and (d) is a schematic diagram of the prediction results of the ELM model on the test set;
[0051] Figure 6 4 is a schematic diagram of the SHAP analysis results of an embodiment of the present invention. DETAILED DESCRIPTION
[0052] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0053] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0054] This embodiment proposes a method for analyzing the risk of CO2 leakage, such as Figure 1 As shown, the specific steps include:
[0055] Obtain a sample set of CO2 leakage;
[0056] Randomly divide the sample set to obtain a first sample set and a second sample set;
[0057] Constructing an initial CO2 leakage risk prediction model based on the first sample set;
[0058] The initial CO2 leakage risk prediction model is tested and verified using the second sample set to obtain the optimal CO2 leakage risk prediction model (LSSVM model). The CO2 leakage risk prediction model is used to output the CO2 leakage distance and make predictions based on the leakage distance.
[0059] The SHAP method is used to analyze the prediction results and obtain the key factors affecting the CO2 leakage risk.
[0060] Furthermore, the sample set for obtaining CO2 leakage includes:
[0061] Obtain variables affecting CO2 leakage;
[0062] The Latin hypercube sampling method was used to process the influencing factors and obtain the sample set.
[0063] Furthermore, the Latin hypercube sampling method is used to process the influencing factor variables, and the sample set obtained includes:
[0064] S1. Divide the influencing factor variables into equal probability intervals to obtain several sub-intervals;
[0065] S2. Randomly select a variable from each subinterval to form a sample point;
[0066] S3. Repeat S2 until a preset number of sample points are obtained to form a sample set.
[0067] Specifically, the Latin hypercube sampling method is:
[0068] Determine the variables and sample size: First, identify the input variables and determine the number of samples to be generated;
[0069] Divide variable intervals: For each variable, divide it into equal probability intervals according to its value range;
[0070] Random sampling and combined sampling: randomly select a value from each variable's subinterval, and then combine these random values from different variable subintervals to form a sample point. Repeat this process until the required number of sample points are generated. For variable X, if its value range is [a, b], divide the interval into n subintervals, and the length of each subinterval is Randomly draw a value c from each subinterval [a+(j-1)Δc,a+jΔc] (j=1,2,…,n) j, these c j The sample values that make up the variable X.
[0071] Furthermore, before constructing the initial CO2 leakage risk prediction model based on the sample set, the following steps are included:
[0072] Simulating the sample set using a simulator to obtain a simulated CO2 data set, wherein the simulator accurately simulates the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase changes in porous media;
[0073] Anomaly detection and data normalization are performed on the simulated CO2 dataset to obtain the processed sample set.
[0074] Specifically, the sample data was simulated using the CMG-GEM module, which accurately predicts the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase changes in porous media.
[0075] The process of CO2 dissolving in water is realized using Henry's law:
[0076]
[0077] Where H(T) is the Henry's constant at temperature T; H0 is the Henry's constant at the reference temperature; C1 and C2 are empirical fitting constants; and T is the absolute temperature (K).
[0078] The capillary effect is expressed using the Leverett J function:
[0079]
[0080] Among them, J Sg Enter the dimensionless J function in the SGT or SLT table, SRFTNG is the air-water surface tension, is the matrix porosity, K is the matrix permeability, and CF is the conversion factor (CF = 31.8316 when the pressure unit is kP).
[0081] Furthermore, preprocessing the simulated CO2 dataset to obtain a processed first sample set includes:
[0082] Anomaly detection and data normalization are performed on the simulated CO2 dataset to obtain the first processed sample set.
[0083] Specifically, the outlier detection method is:
[0084]
[0085] Where: x is the sample point; u is the mean vector of the data set; -1 is the inverse of the covariance matrix of the data set; D M(x) is the Mahalanobis distance from the sample point x to the mean.
[0086] The dataset normalization method is:
[0087]
[0088] Where: x' is the standardized sample point, min(x) is the minimum feature value; max(x) is the maximum feature value.
[0089] Furthermore, an initial CO2 leakage risk prediction model is constructed using the first sample set. Obtaining the initial CO2 leakage risk prediction model includes:
[0090] Determine the hyperparameters and corresponding value ranges of the initial CO2 leakage risk prediction model:
[0091] Based on the combination of hyperparameters and corresponding hyperparameters, a hyperparameter combination solution is obtained;
[0092] Use cross-validation to verify the hyperparameter combination scheme and obtain the best parameter combination as the final hyperparameter setting;
[0093] Based on the final hyperparameter settings, the initial CO2 leakage risk prediction model is obtained.
[0094] Specifically, model evaluation: use the second sample set to predict the trained model, compare the predicted results with the actual values, and use a variety of error evaluation indicators for evaluation:
[0095]
[0096]
[0097] Among them, y i is the true value, y' i is the model prediction value, y mean is the average value of the data set, and p is the number of training samples per round.
[0098] Furthermore, the CO2 leakage risk prediction model is:
[0099]
[0100] Among them, y(x) is the prediction result, α i is the Lagrange multiplier, K(x,x i ) is the kernel function, b is the bias term, i is the index used to traverse the sample points, N is the total number of sample points, x is the input variable to be predicted, x i is the input value of the i-th training sample.
[0101] Furthermore, the method of using the SHAP method to analyze the prediction results is as follows:
[0102]
[0103] in, is the Shapley value of feature m, which reflects the average marginal contribution of feature m to the model output. F represents the set of all features, S is the feature subset that does not contain feature m, |S| represents the number of features in subset S, |F| is the number of all features, u(S) is the output value when the model only uses feature subset S for prediction, and u(S∪{m}) is the predicted output value of the model after adding feature m to subset S.
[0104] The present embodiment will be described below with reference to the accompanying drawings:
[0105] (1) Variable range selection:
[0106] To establish a broadly applicable prediction model, 12 sets of input parameters were selected, including caprock properties, reservoir properties, and formation water properties. These variables have different effects on the CO2 breakthrough distance, as shown in Table 1.
[0107] Table 1
[0108]
[0109] (2) Sample data generation:
[0110] Latin hypercube sampling combined with numerical simulation was used to generate a data set containing 500 samples (due to the limitation of computing resources, the sum of injection time and diffusion time does not exceed 100 years). The characteristic variable distribution of the data set is as follows Figure 2 shown.
[0111] (3) Data processing:
[0112] In order to ensure the accuracy and rationality of the dataset, the Mahalanobis distance method is used to detect outliers in the dataset. The threshold is determined by calculating the quantile of the Mahalanobis distance, and the 5% quantile is selected as the lower threshold and the 95% quantile is selected as the upper threshold. Values exceeding this threshold are considered outliers. The detection results are shown in Figure 2. Figure 3 As shown in the figure, a total of 50 outliers were detected and deleted.
[0113] (4) Model parameter optimization:
[0114] Grid search technology is used to adjust the control parameters and network structure of the model. The results are as follows:
[0115] LSSVM: The kernel function used is RBF, the regularization parameter is 200, the kernel parameter is 0.01, and the order of the polynomial kernel is 2.
[0116] (5) Comparison with other prediction models:
[0117] Three machine learning prediction methods (GRNN, MLPNN, ELM) were selected to verify the effectiveness of the model of the present invention, and the training and testing results were obtained as follows: Figure 4 (a)-(d) and Figure 5 As shown in (a)-(d).
[0118] In order to more intuitively reflect the model differences, the MAE, MSE, RMSE, and R2 of the model of the present invention and three machine learning models were calculated respectively. The results are shown in Table 2.
[0119] Table 2
[0120]
[0121] (6) Model explanation:
[0122] In order to reveal the impact of each input feature on the prediction results of the LSSVM model, the SHAP method is used to determine the importance of each input feature to the output of the LSSVM prediction model. Figure 6 The cap rock permeability has the most significant and positive effect on the CO2 breakthrough distance, followed by temperature and pressure, which also have a significant impact on the CO2 breakthrough distance.
[0123] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for analyzing CO2 leakage risk, characterized in that: include: Obtain a sample set of CO2 leakage; Randomly dividing the sample set to obtain a first sample set and a second sample set; Constructing an initial CO2 leakage risk prediction model based on the first sample set; Using the second sample set to test and verify the initial CO2 leakage risk prediction model to obtain an optimal CO2 leakage risk prediction model, wherein the CO2 leakage risk prediction model is used to output a CO2 leakage distance and make a prediction based on the leakage distance; The SHAP method is used to analyze the prediction results and obtain the key factors affecting the CO2 leakage risk.
2. A method for analyzing CO2 leakage risk according to claim 1, characterized in that: The sample set for obtaining CO2 leakage includes: Obtain variables affecting CO2 leakage; The influencing factors are processed by using Latin hypercube sampling method to obtain the sample set.
3. A method for analyzing CO2 leakage risk according to claim 2, characterized in that: The Latin hypercube sampling method is used to process the influencing factor variables to obtain the sample set, including: S1. Divide the influencing factor variable into equal probability intervals to obtain several sub-intervals; S2. Randomly select a variable from each subinterval to form a sample point; S3. Repeat S2 until a preset number of sample points are obtained to form the sample set.
4. A method for analyzing CO2 leakage risk according to claim 1, characterized in that: Before constructing the initial CO2 leakage risk prediction model based on the first sample set, the following steps are included: simulating the sample set using a simulator to obtain a simulated CO2 data set, wherein the simulator accurately simulates the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase change in porous media; Anomaly detection and data standardization are performed on the simulated CO2 data set to obtain a processed sample set.
5. A method for analyzing CO2 leakage risk according to claim 4, characterized in that: Constructing an initial CO2 leakage risk prediction model using the first sample set, wherein obtaining the initial CO2 leakage risk prediction model includes: Determine the hyperparameters and corresponding value ranges of the initial CO2 leakage risk prediction model: Combining the hyperparameters with corresponding hyperparameters to obtain a hyperparameter combination solution; Use cross-validation to verify the hyperparameter combination scheme and obtain the optimal parameter combination as the final hyperparameter setting; Based on the final hyperparameter setting, the initial CO2 leakage risk prediction model is obtained.
6. A method for analyzing CO2 leakage risk according to claim 1, characterized in that: Using the second sample set to test and verify the initial CO2 leakage risk prediction model to obtain the optimal CO2 leakage risk prediction model includes: By calculating the performance indicators of the test verification results, the parameters of the initial CO2 leakage risk prediction model are iteratively optimized until the prediction error of the model on the second sample set no longer decreases, thereby determining the optimal CO2 leakage risk prediction model.
7. A method for analyzing CO2 leakage risk according to claim 1, characterized in that: The CO2 leakage risk prediction model is: Among them, y(x) is the prediction result, α i is the Lagrange multiplier, K(x,x i ) is the kernel function, b is the bias term, i is the index used to traverse the sample points, N is the total number of sample points, x is the input variable to be predicted, x i is the input value of the i-th training sample.
8. A method for analyzing CO2 leakage risk according to claim 1, characterized in that: The method of using SHAP method to analyze the prediction results is as follows: in, is the Shapley value of feature m, which reflects the average marginal contribution of feature m to the model output. F represents the set of all features, S is the feature subset that does not contain feature m, |S| represents the number of features in subset S, |F| is the number of all features, u(S) is the output value when the model only uses feature subset S for prediction, and u(S∪{m}) is the predicted output value of the model after adding feature m to subset S.
Citation Information
Patent Citations
Method for extracting leakage risk in carbon dioxide geological sequestration body
CN112800592A
Carbon dioxide geological sequestration leakage assessment method, device, equipment and medium
CN116205487A
Intelligent monitoring method for CO2 leakage sampling
CN119223547A
Method, device and equipment for determining safety of carbon dioxide burying and storage medium
CN119831314A
Method of analyzing influence factor for predicting carbon dioxide concentration of any spatiotemporal position
US20230186173A1