A method for analyzing a risk of CO2 leakage
By acquiring a CO2 leak sample set and constructing an initial prediction model, and combining the SHAP method and Latin hypercube sampling method to optimize parameters, the high cost and low efficiency of CO2 leak risk assessment in traditional methods are solved, achieving high-precision risk prediction and key factor identification, and improving risk management efficiency.
Patent Information
- Application Number
- CN202510927997.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing technologies for assessing the risk of CO2 leakage at CO2 burial sites are hampered by high costs for sensor monitoring, high computational resources for numerical simulation, and expensive geochemical imaging equipment, making it difficult to achieve large-scale or long-term monitoring, and lacking effective risk prediction methods.
By acquiring a CO2 leakage sample set, randomly dividing it into first and second sample sets, constructing an initial prediction model and validating it using the second sample set, and finally using the SHAP method to analyze the prediction results, combined with Latin hypercube sampling and simulator data processing, optimizing model parameters, and identifying key influencing factors.
It improves the accuracy of CO2 leak risk prediction and the generalization ability of the model, identifies key influencing factors, optimizes model parameter selection, adapts to different geological conditions, enhances risk management efficiency, and reduces environmental and economic losses.
Smart Images

Figure CN120634271B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of CO2 leakage risk analysis in geological carbon sequestration, and particularly relates to a CO2 leakage risk analysis method. BACKGROUND
[0002] Current potential carbon dioxide sequestration sites include deep saline aquifers, deep sea formations, depleted oil and gas reservoirs, etc. Among them, deep saline aquifers are evaluated as the best site due to their huge storage capacity. However, there are still many challenges in CO2 storage in saline aquifers, one of which is the density difference between the formation water in the saline aquifer and the injected carbon dioxide, and the low-density carbon dioxide floats in the reservoir and leaks into the atmosphere through the cap rock. Therefore, before selecting the CO2 storage address, it is necessary to quickly and accurately evaluate the CO2 leakage risk. Current CO2 leakage evaluation methods mainly include sensor monitoring, numerical simulation research and geochemical imaging technology, but in practical application, the sensor monitoring technology is greatly affected by the environment, and the deployment and maintenance cost is high, which is not suitable for large-scale or long-term monitoring. The numerical simulation method requires a large amount of geological data, and also requires a large amount of iterative calculation time and computing resources. The geochemical imaging technology is expensive and has high requirements for site conditions and professional ability of operating personnel. SUMMARY
[0003] To solve the above technical problems, the application provides a CO2 leakage risk analysis method, which is suitable for large-scale or long-term monitoring and reduces the iterative calculation time and computing resources.
[0004] To achieve the above purpose, the application provides a CO2 leakage risk analysis method, which comprises:
[0005] obtaining a sample set of CO2 leakage;
[0006] randomly dividing the sample set to obtain a first sample set and a second sample set;
[0007] constructing an initial CO2 leakage risk prediction model according to the first sample set;
[0008] testing and verifying the initial CO2 leakage risk prediction model using the second sample set to obtain an optimal CO2 leakage risk prediction model, wherein the CO2 leakage risk prediction model is used to output a CO2 leakage distance, and the leakage distance is used for prediction;
[0009] using the SHAP method to analyze the prediction result to obtain key factors affecting the CO2 leakage risk.
[0010] Optionally, obtaining the sample set of CO2 leakage comprises:
[0011] An influencing factor variable of CO2 leakage is acquired;
[0012] The influencing factor is processed by using a Latin hypercube sampling method to acquire the sample set.
[0013] Optionally, processing the influencing factor variable by using the Latin hypercube sampling method to acquire the sample set comprises:
[0014] S1, the influencing factor variable is divided into equal probability intervals to acquire a plurality of sub-intervals;
[0015] S2, a variable is randomly selected from each sub-interval to form a sample point;
[0016] S3, repeating S2 until a preset number of sample points are acquired to form the sample set.
[0017] Optionally, before constructing an initial CO2 leakage risk prediction model according to the sample set, the method comprises:
[0018] The sample set is simulated by using a simulator to acquire a simulated CO2 data set, wherein the simulator accurately simulates the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation and phase change in a porous medium;
[0019] The simulated CO2 data set is subjected to anomaly detection and data standardization processing to acquire a processed sample set.
[0020] Optionally, the initial CO2 leakage risk prediction model is constructed by using the first sample set to acquire an initial CO2 leakage risk prediction model, and the method comprises:
[0021] The hyperparameters of the initial CO2 leakage risk prediction model and corresponding value ranges are determined:
[0022] Based on the hyperparameters and corresponding hyperparameters, a hyperparameter combination scheme is acquired;
[0023] The hyperparameter combination scheme is verified by using cross-validation to acquire the best parameter combination as the final hyperparameter setting;
[0024] Based on the final hyperparameter setting, the initial CO2 leakage risk prediction model is acquired.
[0025] Optionally, the initial CO2 leakage risk prediction model is tested and verified by using the second sample set to acquire an optimal CO2 leakage risk prediction model, and the method comprises:
[0026] The performance indicators of the test verification results are calculated, and the parameters of the initial CO2 leakage risk prediction model are iteratively optimized until the prediction error of the model on the second sample set is no longer reduced, so as to determine the optimal CO2 leakage risk prediction model.
[0027] Optionally, the CO2 leakage risk prediction model is:
[0028]
[0029] Wherein, y(x) is a prediction result, alpha i is a Lagrange multiplier, K(x,x i ) is a kernel function, b is a bias term, i is an index, used to traverse sample points, N is the total number of sample points, x is an input variable to be predicted, x i is the input value of the i-th training sample.
[0030] Optionally, the method for analyzing the prediction result by using the SHAP method is:
[0031]
[0032] Wherein, is the Shapley value of feature m, which reflects the average marginal contribution of feature m to the model output, F represents the set of all features, S is a feature subset not containing feature m, |S| represents the number of features in the subset S, |F| is the number of all features, u(S) is the output value when the model only uses the feature subset S for prediction, and u(S∪{m}) is the prediction output value of the model after adding feature m to the subset S.
[0033] Compared with the prior art, the present application has the following advantages and technical effects:
[0034] 1. Improve the prediction accuracy:
[0035] The present application can effectively improve the generalization ability and prediction accuracy of the model by obtaining a sample set of CO2 leakage, randomly dividing it into a first sample set and a second sample set, using the first sample set to construct an initial prediction model, and then using the second sample set for test verification. In particular, the sample set is simulated by a simulator, and the simulated data is preprocessed (such as anomaly detection and data standardization), which further optimizes the sample data quality and provides more accurate input for model training, thereby significantly improving the accuracy of CO2 leakage risk prediction.
[0036] 2. Efficiently identify key influencing factors:
[0037] The application adopts the SHAP (SHapley Additive exPlanations) method to analyze the prediction results, which can systematically identify the key factors affecting the CO2 leakage risk and their contribution. The SHAP method is based on Shapley value, which can quantify the average marginal contribution of each feature to the model output, thus providing a deeper explanation and basis for risk assessment. This method not only helps users understand the internal logic of model prediction, but also guides the control and optimization of key factors in actual engineering to reduce the risk of CO2 leakage.
[0038] 3. Optimization of model parameter selection:
[0039] In the process of supervised learning of the initial prediction model using the first sample set, the application determines the hyperparameters and their value range, and verifies the hyperparameter combination scheme using cross-validation, which can quickly find the best parameter combination. This method avoids the blindness and inefficiency of traditional parameter selection methods, significantly improves the efficiency and effectiveness of model training, and ensures that the final CO2 leakage risk prediction model has higher performance and reliability.
[0040] 4. Strong adaptability:
[0041] Based on a variety of advanced technologies and algorithms (such as Latin hypercube sampling method, simulator simulation, supervised learning, SHAP method, etc.), the application can adapt to the CO2 leakage risk analysis needs of different geological conditions, engineering scenarios and data characteristics. By flexibly adjusting the sample set acquisition method, model construction and optimization process, and key factor analysis method, this method can be widely applied to different types of CO2 geological storage leakage risk assessment, and has strong universality and adaptability.
[0042] 5. Improve risk management efficiency:
[0043] Through effective prediction of CO2 leakage risk and identification of key factors, the application can provide scientific basis for the planning, design, operation and supervision of CO2 related projects. For example, in CO2 geological storage projects, potential leakage risk areas can be identified in advance, storage schemes can be optimized, and environmental risks and economic losses can be reduced.
[0044] In summary, the CO2 leakage risk analysis method of the application has significant technical effects in improving prediction accuracy, identifying key factors, optimizing model parameter selection, strong adaptability and improving risk management efficiency, which can provide strong technical support for risk assessment and management in the field of CO2 related fields. BRIEF DESCRIPTION OF DRAWINGS
[0045] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The illustrations, together with their description, serve to explain the application without unduly
[0046] Figure 1 is a flow chart of a CO2 leakage risk analysis method according to an embodiment of the present application;
[0047] Figure 2 is a data variable distribution diagram according to an embodiment of the present application;
[0048] Figure 3 is an abnormal value detection result diagram according to an embodiment of the present application;
[0049] Figure 4 is a comparative prediction result diagram of an LSSVM model and other machine learning models in a training set according to an embodiment of the present application, (a) is a prediction result diagram of the LSSVM model in the training set, (b) is a prediction result diagram of the GRNN model in the training set, (c) is a prediction result diagram of the MLPNN model in the training set, and (d) is a prediction result diagram of the ELM model in the training set;
[0050] Figure 5 is a comparative prediction result diagram of an LSSVM model and other machine learning models in a test set according to an embodiment of the present application, (a) is a prediction result diagram of the LSSVM model in the test set, (b) is a prediction result diagram of the GRNN model in the test set, (c) is a prediction result diagram of the MLPNN model in the test set, and (d) is a prediction result diagram of the ELM model in the test set;
[0051] Figure 6 is a SHAP analysis result diagram according to an embodiment of the present application. DETAILED DESCRIPTION
[0052] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0053] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0054] The present embodiment proposes a CO2 leakage risk analysis method, as shown in Figure 1 specifically comprising the following steps:
[0055] obtaining a sample set of CO2 leakage;
[0056] randomly divide the sample set to obtain a first sample set and a second sample set;
[0057] construct an initial CO2 leakage risk prediction model according to the first sample set;
[0058] test and verify the initial CO2 leakage risk prediction model by using the second sample set to obtain an optimal CO2 leakage risk prediction model (LSSVM model), wherein the CO2 leakage risk prediction model is used to output a CO2 leakage distance, and the leakage distance is used for prediction;
[0059] adopt the SHAP method to analyze the prediction result to obtain key factors affecting the CO2 leakage risk.
[0060] Further, obtaining the sample set of CO2 leakage includes:
[0061] obtaining the influence factor variable of CO2 leakage;
[0062] adopting the Latin hypercube sampling method to process the influence factor to obtain the sample set.
[0063] Further, adopting the Latin hypercube sampling method to process the influence factor variable to obtain the sample set includes:
[0064] S1, equally-probability interval division is performed on the influence factor variable to obtain a plurality of subintervals;
[0065] S2, a variable is randomly extracted from each subinterval to form a sample point;
[0066] S3, S2 is repeated until a preset number of sample points are obtained to form the sample set.
[0067] Specifically, the Latin hypercube sampling method is:
[0068] determining the variable and the sample number: first, the input variable is determined, and the number of samples to be generated is determined;
[0069] dividing the variable interval: for each variable, equally-probability interval division is performed according to the value range thereof;
[0070] random sampling and combination of samples: a value is randomly extracted from each variable subinterval, and then these random values from different variable subintervals are combined to form a sample point. This process is repeated until the required number of sample points is generated. For the variable X, if the value range thereof is [a, b], the interval is equally divided into n subintervals, and the length of each subinterval is a value c is randomly extracted from each subinterval [a + (j-1)Δc, a + jΔc] (j = 1, 2,..., n) jThese c j The sample values constituting the variable X.
[0071] Further, before constructing the initial CO2 leakage risk prediction model according to the sample set, the following steps are included:
[0072] The sample set is simulated by using a simulator to obtain a simulated CO2 data set, wherein the simulator accurately simulates the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase change in a porous medium.
[0073] The simulated CO2 data set is subjected to anomaly detection and data standardization processing to obtain a processed sample set.
[0074] Specifically, the sample data is simulated by using a CMG-GEM module, which accurately predicts the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase change in a porous medium.
[0075] The process of CO2 dissolving in water is realized by using Henry's law:
[0076]
[0077] where H(T) is the Henry constant at temperature T; H0 is the Henry constant at the reference temperature; C1 and C2 are empirical fitting constants; and T is the absolute temperature (K).
[0078] The capillary effect is represented by the Leverett J function:
[0079]
[0080] where J Sg SGT or SLT table input dimensionless J function, SRFTNG is the gas-water surface tension, is the matrix porosity, K is the matrix permeability, and CF is the conversion factor (CF = 31.8316 when the pressure unit is kP).
[0081] Further, the simulated CO2 data set is preprocessed to obtain a processed first sample set, including:
[0082] The simulated CO2 data set is subjected to anomaly detection and data standardization processing to obtain a processed first sample set.
[0083] Specifically, the anomaly value detection method is:
[0084]
[0085] where x is the sample point; u is the mean vector of the data set; -1 is the inverse of the covariance matrix of the data set; D M(x) is the Mahalanobis distance from sample point x to the mean.
[0086] The data set standardization method is:
[0087]
[0088] Wherein: x' is the normalized sample point, min(x) is the minimum value of the feature; max(x) is the maximum value of the feature.
[0089] Further, an initial CO2 leakage risk prediction model is constructed using the first sample set, and the initial CO2 leakage risk prediction model comprises:
[0090] The hyperparameters of the initial CO2 leakage risk prediction model and the corresponding value range are determined:
[0091] Based on the hyperparameters and the corresponding hyperparameters, a hyperparameter combination scheme is obtained;
[0092] The hyperparameter combination scheme is verified using cross-validation, and the best parameter combination is obtained as the final hyperparameter setting;
[0093] Based on the final hyperparameter setting, the initial CO2 leakage risk prediction model is obtained.
[0094] Specifically, model evaluation: using the second sample set to predict the trained model, comparing the prediction results with the actual values, and using multiple error evaluation indexes for evaluation:
[0095]
[0096]
[0097] Wherein, y i is the true value, y' i is the model prediction value, y mean is the average value of the data set, and p is the number of samples per round of training.
[0098] Further, the CO2 leakage risk prediction model is:
[0099]
[0100] Wherein, y(x) is the prediction result, a i is the Lagrange multiplier, K(x, x i ) is the kernel function, b is the bias term, i is the index, used to traverse the sample points, N is the total number of sample points, x is the input variable to be predicted, and x i is the input value of the i-th training sample.
[0101] Further, the method for analyzing the prediction result by using the SHAP method is:
[0102]
[0103] wherein, is the Shapley value of the feature m, which reflects the average marginal contribution of the feature m to the model output, F represents the set of all features, S is a feature subset not containing the feature m, |S| represents the number of features in the subset S, |F| is the number of all features, u(S) is the output value when the model only uses the feature subset S for prediction, and u(S∪{m}) is the prediction output value of the model after adding the feature m to the subset S.
[0104] The present embodiment will be described below in conjunction with the accompanying drawings:
[0105] (1) Variable range selection:
[0106] In order to establish a prediction model with wide applicability, 12 groups of input parameters are selected, including caprock properties, reservoir properties, and formation water properties. These variables have different effects on the breakthrough distance of CO2, as shown in Table 1.
[0107] Table 1
[0108]
[0109] (2) Sample data generation:
[0110] Latin hypercube sampling is applied in combination with numerical simulation to generate a data set containing 500 samples (due to the limitation of computing resources, the total of injection time and diffusion time is not more than 100 years). The distribution of feature variables of the data set is shown in Figure 2 .
[0111] (3) Data processing:
[0112] In order to ensure the accuracy and reasonableness of the data set, the Mahalanobis distance method is used for outlier detection. The threshold is determined by calculating the quantile of Mahalanobis distance, and the 5th percentile is selected as the lower threshold and the 95th percentile as the upper threshold. Values exceeding the threshold are considered outliers. The detection results are shown in Figure 3 , and a total of 50 outliers are detected and deleted.
[0113] (4) Model parameter optimization:
[0114] The grid search technique is used to adjust the control parameters and network structure of the model, and the results are as follows:
[0115] LSSVM: the kernel function adopts RBF, the regularization parameter is 200, the kernel parameter is 0.01, and the order of the polynomial kernel is 2.
[0116] (5) Compared with other prediction models:
[0117] Three machine learning prediction methods (GRNN, MLPNN, and ELM) were selected to verify the effectiveness of the model of the application, and the training and test results are shown in Figure 4 (a)-(d) and Figure 5 (a)-(d) shown.
[0118] In order to more intuitively reflect the differences between the models, the MAE, MSE, RMSE and R2 of the model of the application and the three machine learning models were calculated respectively, and the results are shown in Table 2.
[0119] Table 2
[0120]
[0121] (6) Model explanation:
[0122] In order to reveal the influence of each input feature on the prediction result of the LSSVM model, the SHAP method was used to determine the importance of each input feature on the output of the LSSVM prediction model. The results are shown in Figure 6 The capillary permeability has the most significant influence on the CO2 breakthrough distance and has a positive effect, followed by temperature and pressure, which also have a greater influence on the CO2 breakthrough distance.
[0123] The above is only the preferred specific embodiment of the application, but the protection scope of the application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the application, which should be covered within the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims.
Claims
1. A method for analyzing CO2 leakage risk, characterized in that, include: Obtain a sample set of CO2 leak samples; The sample set is randomly divided to obtain a first sample set and a second sample set; An initial CO2 leakage risk prediction model is constructed based on the first sample set; The initial CO2 leakage risk prediction model is tested and verified using the second sample set to obtain the optimal CO2 leakage risk prediction model, wherein the CO2 leakage risk prediction model is used to output the CO2 leakage distance and make predictions based on the leakage distance; The CO2 leakage risk prediction model is as follows: in, For the predicted results, For Lagrange multipliers, For kernel function, This is the bias term, where i is the index used to traverse the sample points, N is the total number of sample points, and x is the input variable to be predicted. i Let be the input value of the i-th training sample; The SHAP method was used to analyze the prediction results and identify the key factors affecting the risk of CO2 leakage. The method for analyzing prediction results using the SHAP method is as follows: in, is the Shapley value of feature m, which reflects the average marginal contribution of feature m to the model output. F represents the set of all features, and S is the subset of features that does not include feature m. This represents the number of features in subset S. It is the number of all features. This is the output value when the model uses only a subset of features S for prediction. This refers to the predicted output value of the model after adding feature m to the subset S.
2. The method for analyzing CO2 leakage risk according to claim 1, characterized in that, The sample set of CO2 leaks obtained includes: Identify the variables influencing CO2 leakage; The influencing factors were processed using the Latin hypercube sampling method to obtain the sample set.
3. The method for analyzing CO2 leakage risk according to claim 2, characterized in that, The Latin hypercube sampling method was used to process the influencing factor variables to obtain the sample set, which includes: S1. Divide the influencing factor variables into equal probability intervals to obtain several sub-intervals; S2. Randomly select one variable from each sub-interval to form a sample point; S3. Repeat S2 until a preset number of sample points are obtained to form the sample set.
4. The method for analyzing CO2 leakage risk according to claim 1, characterized in that, Before constructing the initial CO2 leakage risk prediction model based on the first sample set, the following steps are included: The sample set is simulated using a simulator to obtain a simulated CO2 dataset. The simulator accurately simulates the behavior of CO2 based on thermodynamic equilibrium, mass conservation, energy conservation, and phase changes in porous media. Anomaly detection and data standardization are performed on the simulated CO2 dataset to obtain a processed sample set.
5. The method for analyzing CO2 leakage risk according to claim 4, characterized in that, Using the first sample set, an initial CO2 leakage risk prediction model is constructed. The initial CO2 leakage risk prediction model includes: Determine the hyperparameters and corresponding value ranges of the initial CO2 leakage risk prediction model: Based on the hyperparameters and their corresponding hyperparameters, a hyperparameter combination scheme is obtained. Cross-validation was used to verify the hyperparameter combination scheme, and the optimal parameter combination was obtained as the final hyperparameter setting. Based on the final hyperparameter settings, the initial CO2 leakage risk prediction model is obtained.
6. The method for analyzing CO2 leakage risk according to claim 1, characterized in that, The initial CO2 leakage risk prediction model was tested and verified using the second sample set to obtain the optimal CO2 leakage risk prediction model, including: By calculating the performance indicators of the test verification results, the parameters of the initial CO2 leakage risk prediction model are iteratively optimized until the prediction error of the model on the second sample set no longer decreases, thereby determining the optimal CO2 leakage risk prediction model.
Citation Information
Patent Citations
Method for extracting leakage risk in carbon dioxide geological sequestration body
CN112800592A
Carbon dioxide geological sequestration leakage assessment method, device, equipment and medium
CN116205487A