An interpretable data-driven prediction and optimization method for CO2 flooding and storage
By constructing the LightUFTrans multimodal proxy model, the challenges of multimodal data fusion and interpretable modeling in CCUS-EOR technology were solved, enabling accurate prediction and optimization of CO2 oil recovery and storage, improving computational efficiency and engineering reliability, and obtaining the Pareto optimal injection and production scheme.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YANGTZE UNIVERSITY
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-03
AI Technical Summary
Existing CCUS-EOR technology has significant shortcomings in multimodal data fusion, interpretable modeling, and multi-objective collaborative optimization, making it difficult to achieve efficient and robust underground CO2 storage and energy utilization, and lacking a unified evaluation system to guide intelligent decision-making.
A LightUFTrans multimodal proxy model was constructed, and multimodal data was obtained through the CMG-GEM component simulator. A hybrid model was built that integrates a lightweight gradient booster, a time-series fusion Transformer, and a U-Net enhanced Fourier neural operator. Combined with the LA-MONSGA-II optimization framework, multi-objective optimization and multi-dimensional interpretability analysis were performed to achieve accurate prediction and optimization of CO2 oil displacement and storage.
It significantly improves computational efficiency and engineering reliability, achieves accurate characterization of the spatiotemporal dynamics of underground CO2/oil, obtains Pareto optimal injection and production schemes, improves the prediction accuracy of oil production, CO2 sequestration and net present value, and overcomes the limitations of black box models.
Smart Images

Figure CN122334080A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of carbon dioxide capture-enhanced oil recovery and storage technology, and in particular to an interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage. Background Technology
[0002] Carbon dioxide capture, enhanced oil recovery (EOR) utilization and storage (CCUS-EOR) technology is a key pathway to mitigate global warming, achieve carbon emission reduction, and efficiently develop underground energy resources. It can integrate CO2 capture, utilization, and storage, reducing CO2 emissions in the atmosphere while enhancing oil recovery through CO2-driven oil recovery, thus balancing environmental and economic benefits. It is currently a research hotspot in the energy and environment fields.
[0003] However, CCUS-EOR technology faces many challenges in actual engineering deployment. The complex engineering uncertainties make it difficult to design a robust operating strategy that can effectively balance the effects of CO2 geological sequestration, geoenergy utilization efficiency, and economic benefits. This results in insufficient reliability design and limits the large-scale promotion and application of the technology.
[0004] Traditional numerical simulation methods are the main means of modeling and optimizing CCUS-EOR processes. However, these methods are computationally expensive and have limited scalability, making it difficult to efficiently solve geoenergy management problems such as the synergistic optimization of geological storage and energy utilization efficiency. In recent years, data-driven methods have gradually become a research hotspot. Related studies have attempted to use models such as artificial neural networks and long short-term memory networks (LSTM) to combine static engineering parameters or dynamic production data for prediction and optimization. However, most existing data-driven methods rely only on a single type of data and fail to fully integrate multimodal information such as tabular data (static), time series data (dynamic), and spatially distributed data (field scale). This limits the ability to characterize reservoir dynamic behavior and the system optimization effect. Currently, there is still a lack of a unified modeling framework that can integrate multi-source heterogeneous data.
[0005] Meanwhile, existing modeling methods generally suffer from the black box problem of insufficient interpretability. Even if some studies use posterior interpretation methods or attention mechanisms, it is difficult to fully reveal the complex multi-scale physical mechanisms such as multiphase flow, mass transfer, and fluid migration in the CCUS-EOR system. Moreover, single interpretation methods lack systematicity and rigor, cannot provide comprehensive mechanistic insights, and are unable to bridge the gap between data-driven prediction and physical mechanism cognition, thus weakening the engineering credibility of the model.
[0006] Furthermore, there is an inherent trade-off between CO2 sequestration, energy utilization, and economic benefits, and optimizing the CO2 injection scheme is key to achieving a balance among these three factors. Currently, surrogate models and multi-objective optimization algorithms have been applied in this field, but for the CCUS-EOR scenario, there is a lack of systematic comparison of the performance of various surrogate models and optimization algorithms, resulting in significant uncertainty in model selection. Moreover, the suitability of existing optimization algorithms for specific trade-off problems in CO2-EOR still requires in-depth research, and a unified evaluation system is lacking to guide intelligent decision-making for underground carbon-energy systems.
[0007] In summary, the current CCUS-EOR technology has significant shortcomings in multimodal data fusion, interpretable modeling, and multi-objective collaborative optimization. There is an urgent need for an advanced technical framework to solve the above-mentioned technical bottlenecks, improve the modeling accuracy, optimization effect, and engineering credibility of the CCUS-EOR process, and promote the industrial application of this technology. Summary of the Invention
[0008] The purpose of this invention is to provide an interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage, constructing the LightUFTrans multimodal proxy model to accurately characterize the spatiotemporal dynamic evolution of underground systems, significantly improving computational efficiency and greatly reducing engineering design and decision-making costs; through multidimensional interpretability analysis, key engineering control parameters are identified, their temporal dynamic characteristics and spatial migration mechanisms are analyzed, overcoming the limitations of black-box models and improving engineering credibility.
[0009] To achieve the above objectives, this invention provides an interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage, comprising the following steps: Step S1, Multimodal Data Acquisition: Using the CMG-GEM component simulator, numerical simulation is carried out through Monte Carlo sampling to acquire a multimodal data pipeline that integrates static tabular data, dynamic time series data, and spatial distribution data; Step S2, Hybrid Model Construction: Build the LightUFTrans hybrid model, which integrates a lightweight gradient booster, a temporal fusion Transformer, and a U-Net enhanced Fourier neural operator to process static, temporal, and spatial data respectively, and achieve three-dimensional prediction of CO2-EOR and sealing performance. Step S3, Multi-objective optimization: LightUFTrans is embedded as a surrogate model into the LA-MONSGA-II optimization framework, with the cumulative crude oil production, CO2 storage, and net present value as optimization objectives, to obtain the Pareto optimal solution; Step S4, Multidimensional Interpretability Analysis: Combining SHAP, attention mechanism, and spatiotemporal validation, the model decision-making mechanism is analyzed from the dimensions of global features, temporal dependence, and spatial distribution.
[0010] Preferably, in step S1, a CMG-GEM component simulator is used to simulate the multiphase and multicomponent flow behavior involved in CO2 injection and oil recovery. The specific process is as follows: Based on the Peng–Robinson equation of state, by introducing the temperature-dependent attraction term κ and fluid characteristic parameters a and b, the gas-liquid phase equilibrium behavior under reservoir conditions is predicted, as shown below: (1); in, Indicates pressure; It is the gas constant; Indicates temperature; Indicates molar volume; and These are fluid characteristic parameters; For temperature-related attraction terms; In a multiphase, multicomponent system with CO2 injection, the mass conservation equation is used to describe the spatiotemporal evolution of each component in all petroleum phases, as shown below: (2); in, Indicates the porosity of porous media; Indicates phase saturation; Indicates phase mass density; Indicates that component i is in phase The mole fraction in; Indicates phase Darcy velocity is used to describe the volumetric flow behavior under the influence of pressure gradient and gravity. Represents the source and sink terms of component i, used to characterize the injection or production process; For each fluid phase Darcy's velocity is described by the multiphase Darcy's law, as shown below: (3); in, Indicates absolute penetration rate; Indicates phase Relative penetration rate; Indicates phase viscosity; Indicates phase The pressure; Represents gravitational acceleration; Indicates depth; The CO2 sequestration mechanism is comprehensively characterized by considering three forms: structural sequestration, dissolution sequestration, and bound sequestration, as detailed below: The mass of CO2 to be sealed is calculated directly from the in-situ gas phase mass, as shown below: (4); in, Indicates the quality of CO2 structure sequestration; Represents a grid cell; Indicates porosity; Indicates the volume of a grid cell; Indicates CO2 gas phase saturation; Indicates the density of the gas phase; In the CMG-GEM simulation, the solubility of CO2 in brine is calculated using Henry's Law for dissolution and sequestration, as shown below: (5); in, This indicates the concentration of CO2 in the aqueous phase; This represents the CO2 fugacity calculated using the equation of state. is the Henry's constant, whose value depends on temperature T, pressure P, and ionic strength I; CMG-GEM simulates bound CO2 sequestration by introducing a relative permeability and capillary pressure model with hysteresis characteristics, thereby characterizing the transition process of the fluid from a kinetic to a kinetic state after WAG displacement; the total mass of bound CO2 is shown below: (6); in, This indicates the total mass of CO2 bound and sealed. This indicates the residual gas saturation.
[0011] Preferably, the Monte Carlo sampling method is used for numerical simulation to generate static tabular data, historical time series data and spatial field data. The three types of data are then input into three different machine learning model components to predict the CO2-EOR and geological sequestration capacity, dynamic processes and spatial distribution characteristics, thereby achieving cross-modal constraints. The static table data includes: gas injection time, water injection time, gas injection rate, water injection rate, bottom-hole flowing pressure of injection well, bottom-hole flowing pressure of production well, and surface production rate of production well. The static table data is input into the lightweight gradient lift model, which outputs the cumulative crude oil production, CO2 storage, and net present value. Historical time series data includes: historical crude oil production rate and historical CO2 sequestration rate over time; historical time series data and static tabular data are input into the time series fusion Transformer model to output crude oil production rate and CO2 sequestration rate over time. The spatial field data includes: permeability distribution map and initial water saturation distribution map; the spatial field data and static tabular data are input into the U-Net enhanced Fourier neural operator model to output crude oil saturation distribution map and gas saturation distribution map.
[0012] Preferably, in step S2, the process by which the lightweight gradient booster achieves predictive performance on static tabular data in the LightUFTrans hybrid model is as follows: First, gradient-based one-sided sampling prioritizes samples with large gradients, thereby significantly reducing computational costs while ensuring model accuracy. Secondly, the histogram-based algorithm discretizes continuous features into multiple intervals, effectively reducing memory usage and accelerating computation. Furthermore, the mutual exclusion feature bundling technique is used to combine non-overlapping features, reducing the number of effective features and improving computational efficiency. Finally, the leaf node-based growth strategy prioritizes splitting the leaf nodes with the highest gain, thereby reducing errors and improving the overall model performance.
[0013] Preferably, in terms of time series forecasting, a time-series fusion Transformer is used to model the dynamics of CO2 sequestration and the oil production process, and a dedicated encoding module for static parameters is introduced to enhance the model's sensitivity to static features. The specific process is as follows: First, a deep fusion of static and dynamic data is achieved through a static covariate encoder and a dedicated encoding module for static parameters. Second, a gated residual network is introduced to enhance the model's autoregressive modeling capability and enable the processing of complex time series data. Third, through a feature selection network and an interpretable multi-head attention mechanism, the temporal fusion Transformer identifies key influencing factors and generates attention weights, improving the model's interpretability. Fourth, the temporal fusion Transformer supports conditional prediction capabilities, combining known future information for prediction and achieving engineering optimization. Finally, the temporal fusion Transformer has the ability to process multi-scale temporal data, meeting the requirements of different time resolutions in CCUS-EOR. The static covariate encoder generates four distinct context vectors through independent GRN modules, used for temporal feature selection, local temporal feature processing, and enhancement and fusion of temporal features with static information, respectively. The GRN efficiently models the input through nonlinear transformations, as shown below: (7); (8); (9); (10); in, Represents the gated residual network; This represents the learnable network parameters; This represents the main input feature vector; Represents a static eigenvector; Presentation layer normalization operation; Represents a gated linear unit; Represents the activation function of the exponential linear unit; This represents the Sigmoid activation function; Indicates intermediate implicit variables; express The function's output variable; express The input variables of the function; , , , , Indicates weight; , , , Indicates bias; Represents the Hadamard product; In neural networks, the attention mechanism enables the model to focus on key parts of the input sequence during the prediction process, as shown below: (11); in, Represents the query matrix; Represents the key matrix; Represents a value matrix; Indicates the dimension of the key vector; The function is used to convert the scaled dot product result into a probability distribution; Represents single attention; Multi-head attention mechanisms introduce multiple attention heads, enabling the model to simultaneously focus on different parts of the input sequence, as shown below: (12); (13); in, , and Represent the i-th attention head respectively The weight matrix of queries, keys, and values; This represents the linear mapping weight matrix after multi-head attention output; Indicates the number of attention heads; This indicates a high level of attention from multiple parties. This indicates a concatenation operation, used to merge the outputs of multiple attention heads into a unified matrix; To enhance the interpretability of the model, an interpretable multi-head attention mechanism with shared-value weights is employed to highlight the importance of specific features, as shown below: (14); (15); (16); (17); in, This indicates that multi-headed attention can be explained; Represents the final mapping matrix. This represents the aggregated result of all attention head outputs; This represents the normalized attention weights. This represents a value mapping matrix shared among all attention heads; and Let represent the projection weight matrices of the query and key for the h-th attention head, respectively; This indicates the number of attention heads.
[0014] Preferably, the LightUFTrans hybrid model introduces the U-Net enhanced Fourier neural operator to achieve spatial prediction of CO2 and oil saturation distribution, and is used to visualize and verify the optimization results; The U-Net enhanced Fourier neural operator introduces a multi-scale U-Net structural path to characterize the spatial hierarchical features in the data, enabling the model to simultaneously capture large-scale oscillation modes and fine local structures, and predict the spatial distribution dynamics of oil recovery and CO2 sequestration.
[0015] Preferably, in step S3, LightUFTrans is embedded as a surrogate model into the LA-MONSGA-II optimization framework, and the termination condition for multi-objective optimization is set to a fixed maximum number of iterations, which ensures that the Pareto front fully converges and explores while taking into account computational feasibility. During the training of the surrogate model, by combining an elite retention strategy, a selection mechanism based on crowding distance, and a surrogate model evaluation with an early shutdown strategy, LA-MONSGA-II obtains the Pareto optimal solution with the optimization objectives of cumulative crude oil production, CO2 storage, and net present value.
[0016] Preferably, an economic evaluation model is constructed to estimate the net present value (NPV) by comprehensively considering oil revenue, CO2 processing costs, carbon trading revenue, and tax incentives, as shown below: (18); (19); in, This represents the revenue generated from hydrocarbon production in year n. Represents the operating costs for the same year; The incentive price is set as the price for CO2 geological sequestration using EOR technology in year n. Indicates the discount rate; The hydrocarbon production in year n; Indicates the price of hydrocarbons; Indicates the cumulative water injection volume; Indicates the cumulative gas injection volume; This represents the water injection cost corresponding to the cumulative water injection volume; This represents the CO2 procurement cost corresponding to the cumulative gas injection volume; This represents the cumulative gas production in year n. Represents the cost of CO2 recycling; Tax credits for carbon sequestration; To verify the model performance, several commonly used evaluation metrics were employed, including root mean square error (RMSE) and coefficient of determination (R²). 2 Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and Structural Similarity Index (SSIM) are used to measure the error between predicted and observed values.
[0017] Preferably, in step S4, the model decision-making mechanism is analyzed from multiple dimensions, including global features, temporal dependencies, and spatial distribution, by combining SHAP, attention mechanism, and spatiotemporal verification. The specific process is as follows: (1) Quantitative analysis of the importance of key engineering parameters based on SHAP; The SHAP method was used to analyze the feature importance in the LightUFTrans model. The model predictions and the SHAP feature importance ranking were highly consistent. The parameters that play a key controlling role in cumulative oil production and net present value (NPV) are the surface production rate of the production well, the gas injection rate / injection time, and the bottom hole flowing pressure, as well as the secondary parameter, the water injection rate / injection time. For CO2 sequestration, the key parameters are the gas injection rate, gas injection time, and water injection time, which is consistent with the traditional understanding of oil and gas reservoir mechanisms and verifies the reliability of the model. (2) Temporal pattern interpretation based on attention mechanism: dynamic analysis of CO2 sequestration and oil production; In oil production forecasting, the fluctuation of attention weights is highly consistent with the key trends in historical data. A conservative strategy is adopted, resulting in a smooth output fluctuation and a small overall forecast bias. In CO2 sequestration prediction, the model effectively identifies extreme points and inflection points, captures subtle changes in time series, and has a good tracking effect on the fluctuation trend of the optimization scheme, demonstrating excellent out-of-distribution generalization ability. (3) Interpretability analysis of CO2 and oil saturation distribution based on spatiotemporal verification; A spatiotemporal validation method was introduced to systematically analyze the spatial distribution and evolution of CO2 and oil saturation predicted by the LightUFTrans model, verifying the reliability and strong generalization ability of the LightUFTrans model in predicting the spatial distribution of CO2 and oil saturation. By integrating SHAP analysis, attention mechanisms, and spatiotemporal verification methods, a systematic interpretability analysis of the LightUFTrans model was conducted from three dimensions: global, temporal, and spatial.
[0018] Therefore, the present invention employs the above-mentioned interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage, with the following beneficial effects: (1) Significantly improved prediction accuracy: The LightUFTrans multimodal proxy model innovatively integrates quantitative evaluation, dynamic trend prediction and spatial distribution prediction, accurately depicts the spatiotemporal dynamic evolution of underground CO2 / oil, significantly improves computational efficiency, greatly reduces engineering design and decision-making costs, fills the key cognitive gap in underground system modeling, and achieves a comprehensive characterization of the system.
[0019] (2) The multi-objective synergistic optimization effect is outstanding: the Pareto optimal injection and production scheme is obtained, with oil production, CO2 storage and net present value reaching 806,000 tons, 1.47 million tons and US$443 million respectively, which are 5.06%, 15.98% and 3.53% higher than traditional numerical simulation.
[0020] (3) The interpretability of the model is significantly enhanced: Through multidimensional interpretability analysis, the key engineering control parameters are identified, their time dynamic characteristics and spatial migration mechanism are analyzed, the limitations of the black box model are overcome, and the credibility of the project is improved.
[0021] (4) Forming a general intelligent decision-making paradigm: By further integrating multi-source data, optimizing multimodal learning methods, and expanding data scale and key parameters, its applicability in the comprehensive geological system can be further improved, adapting to the needs of multiple fields such as geothermal development, CO2 geological storage, and intelligent optimization of underground resources.
[0022] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the framework of an interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to the present invention. Figure 2 This is a schematic diagram of reservoir numerical sub-model extraction in an embodiment of the present invention; Figure 3This is a schematic diagram of the overall architecture of the LightUFTrans hybrid model of the present invention; wherein, a is a TFT structure; b is a LightGBM structure; and c is a U-FNO structure; Figure 4 This invention is based on static tabular data; where a is the model's prediction result for the cumulative crude oil production in CCUS-EOR operations; b is the model's prediction result for the CO2 sequestration volume in CCUS-EOR operations; and c is the model's prediction result for the net present value (NPV) in CCUS-EOR operations. Figure 5 This invention presents the prediction results of the proposed model on crude oil production rate using historical time steps in two randomly selected CCUS-EOR operation cases; where a is Case 5 and b is Case 320. Figure 6 This invention is in Figure 5 In two randomly selected cases, the proposed model uses historical time steps to predict the time-varying CO2 sequestration amount; where a is case 5 and b is case 320. Figure 7 This invention performs K-means clustering on 1069 simulations after PCA dimensionality reduction; Figure 8 The table shows the observed (first row, from CMG-GEM simulation) and predicted (second row, from the model proposed in this invention) results of crude oil saturation distribution at three selected time steps (from left to right: steps 1, 7, and 11, each step representing 1.67 years) in different layers of this invention; where a represents layer 3; b represents layer 7; and the injection well location is indicated by... Marked as a reference point for CO2 saturation behavior; Figure 9 The first row shows the observational results (from CMG-GEM simulation) and the prediction results (from the model proposed in this invention) of CO2 saturation distribution at three selected time steps (from left to right: steps 1, 7, and 11, each step representing 1.67 years) at different layers of this invention; where a is layer 3 and b is layer 7. Figure 10 This presents the CO2-EOR and storage optimization results of this invention; where ab represents the superiority of LA-MONSGA-II (combined with LightUFTrans) over other methods in terms of Pareto front and supervolume indicators; c represents the comparison of the optimization results with the Monte Carlo baseline; and d represents the cumulative crude oil production, CO2 storage capacity, and net present value (NPV) displayed in three dimensions. Figure 11 This invention demonstrates the performance advantages of its optimized scheme compared to a 20-year CO2-WAG Monte Carlo simulation; where a represents cumulative crude oil production (HP) and b represents CO2 storage capacity. Figure 12This invention is based on the eigenvalue importance of SHAP; where ac is the eigenvalue importance of cumulative crude oil production, df is the eigenvalue importance of CO2 sequestration, and gi is the eigenvalue importance of net present value. Figure 13 These are the prediction results under the optimized scheme of this invention; where a is the prediction result of crude oil production rate and b is the prediction result of time-varying CO2 sequestration. Figure 14 The first row shows the observed results (from CMG-GEM simulation) and the predicted results (from the model proposed in this invention) of crude oil saturation distribution at three selected time steps (from left to right: steps 1, 7, and 11, each step representing 1.67 years) at different layers under the optimization scheme of this invention; where a is layer 3 and b is layer 7. Figure 15 The first row shows the observation results (from CMG-GEM simulation) and the prediction results (from the model proposed in this invention) of CO2 saturation distribution at three selected time steps (from left to right: steps 1, 7, and 11, each step representing 1.67 years) in different layers under the optimization scheme of this invention; where a is layer 3 and b is layer 7. Detailed Implementation
[0024] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0025] like Figure 1 As shown, this invention presents an interpretable data-driven framework for predicting and optimizing CO2-EOR and storage performance under a CO2-WAG (CO2-water-gas alternating injection) strategy. It fully utilizes multimodal input data from large-scale Monte Carlo reservoir numerical simulations, including static tabular data (such as injection schemes and well control parameters), time-series data (such as historical crude oil production dynamics and CO2 storage rate variations), and spatial data (such as permeability distribution and fluid saturation field). A hybrid surrogate model, LightUFTrans, is constructed by fusing Lightweight Gradient Booster (LightGBM), Temporal Fusion Transformer (TFT), and U-Net Enhanced Fourier Neural Operator (U-FNO), allowing each module to model different data modalities. This architecture enables high-precision predictions at three levels: (i) static prediction of cumulative recovery, CO2 storage, and net present value (NPV); (ii) dynamic evolution prediction of the production and storage process; and (iii) prediction of the saturation distribution of oil and CO2 in the underground space.
[0026] Furthermore, LightUFTrans is embedded as a surrogate model within a multi-objective optimization framework to explore the trade-offs between recovery efficiency, storage capacity, and economic benefits. This unified and physically constrained framework provides a high-resolution, interpretable modeling method for intelligent decision-making in CO2 utilization-related underground energy systems.
[0027] Example 1 Step S1: Multimodal data acquisition: Using the CMG-GEM component simulator, numerical simulation is carried out through Monte Carlo sampling to acquire a multimodal data pipeline that integrates static tabular data, dynamic time series data, and spatial distribution data.
[0028] Based on field engineering and production data from the Turpan-Hami Basin in Northwest China, a component reservoir numerical model was constructed. To reduce computational costs while preserving reservoir heterogeneity and flow characteristics, a representative sub-model centered on a specific well group was extracted from the full-field-scale geological model. Figure 2 As shown in the figure, this sub-model retains the spatial variability and key physical characteristics of the original system and serves as the computational basis for simulating the geological sequestration process under CO2 enhanced oil recovery (EOR) and water-gas alternating injection (WAG) conditions.
[0029] The reservoir numerical sub-model used includes 4 injection wells and 15 production wells, distributed in a structurally heterogeneous reservoir. The average well spacing between injection and production wells is approximately 417m, consistent with conventional well spacing in actual oilfield development. The effective vertical depth of the wellbore penetrating the reservoir is approximately 300m. Since this model is derived from an actual oilfield, the depth and spacing of different wells vary; therefore, this invention only reports the average values.
[0030] The reservoir numerical sub-model was discretized into 20×18×25 grid cells in the x, y, and z directions, corresponding to a depth range of 2203–2647 m (total thickness approximately 444 m). Each grid cell has a size of approximately 51.9 m in the x and y directions and approximately 17.8 m in the z direction. Rock physical properties (such as porosity, permeability, and initial hydrocarbon saturation) were determined based on field measurement data. Except for the top boundary, all outer boundaries were set as flow-free boundaries; the top boundary was set to constant pressure conditions to simulate fluid communication with the overlying aquifer. The basic parameters of the reservoir numerical sub-model are shown in Table 1.
[0031] Table 1. Basic parameters of the reservoir numerical sub-model
[0032] The component simulator CMG-GEM, developed by Computer Modelling Group Ltd. (CMG), was used to simulate the multiphase, multicomponent flow behavior involved in CO2 injection and oil recovery. This simulator is based on a series of fundamental physical principles, including the Peng–Robinson equation of state (EOS), the mass conservation equation, Darcy's law, and Henry's law. The calculations of oil production and CO2 sequestration are both determined by these governing equations.
[0033] The Peng–Robinson equation of state is widely used to describe the phase behavior of non-ideal fluids such as CO2 and oil. By introducing a temperature-dependent attraction term κ and fluid characteristic parameters a and b, this equation can more accurately predict the gas-liquid phase equilibrium behavior under reservoir conditions, as shown below: (1); in, Indicates pressure; It is the gas constant; Indicates temperature; Indicates molar volume; and These are fluid characteristic parameters; This is an attraction term related to temperature.
[0034] In multiphase, multicomponent systems involving CO2 injection, the mass conservation equation is used to describe the spatiotemporal evolution of each component in all petroleum phases, as shown below: (2); in, Indicates the porosity of porous media; Indicates phase saturation; Indicates phase mass density; Indicates that component i is in phase The mole fraction in; Indicates phase Darcy velocity is used to describe the volumetric flow behavior under the influence of pressure gradient and gravity. This represents the source and sink terms of component i, used to characterize the injection or production process.
[0035] For each fluid phase Its Darcy velocity is described by the multiphase Darcy law, as shown below: (3); in, Indicates absolute penetration rate; Indicates phase Relative penetration rate; Indicates phase viscosity; Indicates phase The pressure; Represents gravitational acceleration; Indicates depth.
[0036] The CO2 sequestration mechanism is comprehensively characterized by considering three forms: tectonic sequestration, dissolution sequestration, and bound sequestration. These processes are not only controlled by multiphase flow mechanisms but also constrained by thermodynamic principles, including equation of state (EOS) modeling and dissolution equilibrium relationships. Specifically, the Peng–Robinson equation of state is used to describe the non-ideal phase behavior of the CO2-petroleum system, while the dissolution equilibrium of CO2 in brine is governed by Henry's law.
[0037] Structural sequestration refers to the physical retention of CO2 in its free gas phase or supercritical state at high geological structures (such as the top of anticlines or beneath impermeable faults and caprocks). Its simulation depends on the integrity of the caprock and the reservoir's structural morphology, as CO2 tends to accumulate at structural high points. The mass of structurally sequestered CO2 can be directly calculated from the in-situ gas phase mass, as shown below: (4); in, Indicates the quality of CO2 structure sequestration; Represents a grid cell; Indicates porosity; Indicates the volume of a grid cell; Indicates CO2 gas phase saturation; This represents the density of the gas phase.
[0038] Dissolution sequestration refers to the process where injected CO2 dissolves in formation water, forming bicarbonate ions (HCO3-). - ) and carbonate ions (CO3) 2- The process involves calculating the solubility of CO2 in salt water using Henry's Law in CMG-GEM simulations, as shown below: (5); in, This indicates the concentration of CO2 in the aqueous phase; This represents the CO2 fugacity calculated using a state equation (such as the Peng–Robinson equation). is the Henry's constant, whose value depends on temperature T, pressure P, and ionic strength I.
[0039] Binding and sequestration refers to the retention of CO2 in pore spaces as discontinuous ganglia due to capillary forces, thus losing its fluidity. CMG-GEM simulates this process by introducing a relative permeability and capillary pressure model with hysteresis characteristics, thereby characterizing the transition of the fluid from a mobile to a non-mobile state after WAG displacement. The total mass of bound and sequestered CO2 is shown below: (6); in, This indicates the total mass of CO2 bound and sealed. This indicates the residual gas saturation. This form of storage provides short- to medium-term storage security.
[0040] According to the development plan, it is assumed that all existing water injection wells will be uniformly converted into WAG injection wells, and coordinated water-gas alternating injection (WAG) operations and joint control strategies for injection and production wells will be implemented. To fully cover actual operating conditions, this invention employs the Monte Carlo sampling method to explore various WAG scheme designs and well control parameter combinations. The Monte Carlo method is a statistical method based on random sampling, which characterizes the uncertainty and variability of the system by repeatedly randomly sampling input parameters within a preset range. All static operating parameters listed in Table 2 are independently and randomly sampled within a preset range according to a uniform distribution, conforming to the standard Monte Carlo simulation procedure.
[0041] Table 2. Basic parameters and ranges of the Monte Carlo experiment
[0042] This invention conducted 1069 Monte Carlo numerical simulations, with inputs including static parameters (such as WAG scheme design and well control parameters, as shown in Table 2) and historical time-varying data (such as oil production and CO2 sequestration rate). All simulations were performed using the component simulator CMG-GEM, generating observational data on cumulative oil production and CO2 sequestration, while also recording the time-series changes in oil production and CO2 sequestration. Each simulation generated 3150 time-series data records over a twenty-year CO2-WAG injection process.
[0043] The number of WAG cycles in each simulation depends on the duration of the gas injection and water injection phases. Gas and water injection times were sampled independently between 100 and 360 days (see Table 2), resulting in a WAG cycle count varying from approximately 10 to over 30 over a 20-year injection cycle. This setup allows the model to cover diverse CCUS-EOR operating strategies. The selected CO2-WAG strategy is consistent with the field development plan, aiming to improve carbon sequestration efficiency while optimizing economic benefits.
[0044] In each Monte Carlo simulation, CO2 injection was conducted under surface liquid production (STL) constraints, with STL serving as the upper limit of production. To ensure miscibility (minimum miscibility pressure MMP is approximately 21 MPa), a lower limit of 22 MPa was set for the bottom hole pressure (BHP) of all production wells. In each WAG cycle, the injection rate of the injection well and the well control parameters (i.e., STL and BHP) of the production well were simultaneously set. It is important to note that the bottom hole pressure of the injection well must not exceed the formation fracture pressure (approximately 60 MPa). This setting allows the simulator to dynamically adjust the injection intensity according to reservoir conditions while maintaining favorable phase conditions for miscible displacement. Furthermore, this method enhances the realism of injection strategy exploration, simulating the flexibility of switching between flow-limited and pressure-limited well control modes in field operations. Therefore, this simulation framework can generate rich and engineering-practical CO2-EOR and storage optimization design spaces.
[0045] Furthermore, this invention combines the permeability field and initial water saturation distribution map with various static parameters to predict the spatial distribution of oil production and CO2 sequestration in an oil reservoir. Permeability and initial water saturation are represented as a 20×18 two-dimensional matrix to depict the planar distribution of the reservoir. The static parameters, originally represented as scalars, are extended to 20×18 matrices to match the reservoir's spatial grid. These matrices together constitute an input tensor with a dimension of 20×18×8. The reservoir model contains 25 vertical layers, and time-dimensional data is generated by introducing 12 uniform time sampling points, resulting in a final input tensor with a dimension of 20×18×25×12×8. The model output is the distribution of oil saturation and CO2 saturation in the reservoir, with a tensor dimension of 20×18×25×12×2.
[0046] The aforementioned static tabular data, historical time-series data, and spatial field data were input into three different machine learning model components to predict CO2-EOR and the capacity, dynamic processes, and spatial distribution characteristics of geological sequestration. Table 3 summarizes the input and output parameters used by each model component.
[0047] Table 3 shows the inputs and outputs of the proposed model.
[0048] It is worth noting that the same set of static table parameters used by LightGBM are also introduced as auxiliary inputs into the TFT and U-FNO models, thereby achieving cross-modal condition constraints and improving the prediction accuracy of temporal and spatial targets.
[0049] Step S2, Hybrid Model Construction: Build the LightUFTrans hybrid model, which integrates LightGBM, Temporal Fusion Transformer (TFT), and U-Net Enhanced Fourier Neural Operator (U-FNO) to process static, temporal, and spatial data respectively, and achieve three-dimensional prediction of CO2-EOR and sealing performance.
[0050] LightUFTrans is a deep learning model that integrates Lightweight Gradient Boosting Machine (LightGBM), U-Net Enhanced Fourier Neural Operator (U-FNO), and Temporal Fusion Transformer (TFT). The TFT structure is used to predict future crude oil production and CO2 sequestration dynamics; the LightGBM structure is used to predict cumulative crude oil production, CO2 sequestration, and net present value (NPV); and the U-FNO structure is used to model the spatial distribution of crude oil and gas saturation. Therefore, LightUFTrans is specifically designed for predicting CO2-EOR and sequestration performance in oil reservoirs, and can efficiently process large-scale static tabular data, dynamic time-series data, and spatially distributed data in CCUS-EOR systems, such as… Figure 3 As shown.
[0051] In the LightUFTrans hybrid model, LightGBM achieves superior prediction performance through the following four core technologies, leveraging its efficient processing capabilities for static tabular data: First, gradient-based one-sided sampling (GOSS) prioritizes samples with larger gradients, thereby significantly reducing computational costs while ensuring model accuracy. Secondly, the histogram-based algorithm discretizes continuous features into multiple intervals (i.e., input parameter binning), which effectively reduces memory usage and accelerates computation. In addition, the mutually exclusive feature bundling (EFB) technique combines non-overlapping features, thereby reducing the number of effective features and improving computational efficiency. Finally, the leaf-wise growth strategy prioritizes splitting the leaf nodes with the highest gain, thereby more effectively reducing errors and improving the overall model performance.
[0052] In this invention, LightGBM has good adaptability to small sample data and can maintain high prediction accuracy even when the data scale is small (such as thousands of samples). LightGBM has a faster training speed, is especially suitable for small and medium-sized datasets, and is relatively easy to use. Compared with the complex and time-consuming hyperparameter tuning process of MLNN, LightGBM requires less parameter tuning and is more practical for engineering.
[0053] In time series forecasting, the Temporal Fusion Transformer (TFT) is employed to model CO2 sequestration dynamics and oil production processes. The Transformer model has achieved significant success in Natural Language Processing (NLP) and Computer Vision (CV), particularly excelling in handling complex dependencies and deep feature extraction. In recent years, its application has expanded to time series forecasting, outperforming traditional Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) in various tasks. In the CCUS-EOR scenario, the Transformer model demonstrates promising application potential in CO2 sequestration forecasting and reservoir production dynamics modeling.
[0054] However, traditional Transformers still have certain limitations in time series tasks. First, their ability to model autoregressive properties is insufficient, making it difficult to effectively capture sequence dependencies in time series. Second, they are inadequate in integrating static and dynamic data, as the original structure fails to naturally support unified modeling of time-invariant and time-varying data. Furthermore, Transformer models are often considered black boxes, with weak interpretability, making it difficult to reveal the decision-making mechanisms behind the prediction results.
[0055] To overcome the aforementioned problems, TFT has made several improvements to the Transformer, introducing a dedicated encoding module for static parameters to enhance the model's sensitivity to static features. This gives it a significant advantage in dynamic prediction of CO2 sequestration and oil production under different engineering parameter conditions. TFT has demonstrated good application results in the energy and engineering fields and shows great promise in time-dependent prediction of the CCUS-EOR system.
[0056] The core design improvements of TFT are mainly reflected in the following aspects: First, by using a static covariate encoder and a dedicated encoding module for static parameters, deep fusion of static and dynamic data is achieved, thereby significantly improving prediction accuracy and model generalization ability. Second, the introduction of a gated residual network (GRN) enhances the model's autoregressive modeling ability, enabling it to handle complex time series more accurately. Furthermore, through a feature selection network and an interpretable multi-head attention mechanism, TFT can identify key influencing factors and generate attention weights, thereby improving model interpretability and providing decision support for CO2 sequestration and oil recovery strategies. Simultaneously, TFT supports conditional prediction capabilities, allowing predictions to be made in conjunction with known future information (such as changes in injection strategies), which is helpful for engineering optimization. Finally, TFT has the ability to process multi-scale time data (from daily to annual scales), meeting the needs of different time resolutions in the CCUS-EOR project.
[0057] Specifically, the static covariate encoder generates four distinct context vectors through independent GRN modules. These vectors are used for temporal feature selection, local temporal feature processing, and the enhancement and fusion of temporal features with static information, respectively. The GRN efficiently models the input through nonlinear transformations, as shown below: (7); (8); (9); (10); in, Represents the gated residual network; This represents the learnable network parameters; This represents the main input feature vector; Represents a static eigenvector; Presentation layer normalization operation; Represents a gated linear unit; Represents the activation function of the exponential linear unit; This represents the Sigmoid activation function; Indicates intermediate implicit variables; express The function's output variable; express The input variables of the function; , , , , Indicates weight; , , , Indicates bias; This represents element-wise multiplication (Hadamard product).
[0058] In neural networks, the attention mechanism enables the model to focus on key parts of the input sequence during the prediction process, as shown below: (11); in, Represents the query matrix (Query); Represents the key matrix; Represents a value matrix; Indicates the dimension of the key vector; Represents single attention; The function is used to convert the scaled dot product result into a probability distribution.
[0059] Multi-head attention mechanisms introduce multiple attention heads, enabling the model to simultaneously focus on different parts of the input sequence, as shown below: (12); (13); in, , and Represent the i-th attention head respectively The weight matrix of queries, keys, and values; This represents the linear mapping weight matrix after multi-head attention output; Indicates the number of attention heads; This indicates a high level of attention from multiple parties. This indicates a concatenation operation, used to merge the outputs of multiple attention heads into a unified matrix.
[0060] Furthermore, to enhance the interpretability of the model, an interpretable multi-head attention mechanism with shared-value weights is employed to highlight the importance of specific features, as shown below: (14); (15); (16); (17); in, This indicates that multi-headed attention can be explained; Represents the final mapping matrix. This represents the aggregated result of all attention head outputs; This represents the normalized attention weights. This represents a value mapping matrix shared among all attention heads; and Let represent the projection weight matrices of the query and key for the h-th attention head, respectively; This indicates the number of attention heads.
[0061] Furthermore, LightUFTrans introduces the U-Net Enhanced Fourier Neural Operator (U-FNO) to achieve high-precision spatial prediction of CO2 and oil saturation distributions, and to visualize and verify the optimization results. U-FNO is a neural network architecture specifically designed for solving complex partial differential equations (PDEs). Its core idea is to map data to the frequency domain through Fourier transform, process complex spatial features in the frequency domain, and then transform it back to physical space, thereby balancing global feature extraction with local detail preservation.
[0062] As an extension of the Fourier Neural Operator (FNO), U-FNO introduces a multi-scale U-Net structural path to characterize the spatial hierarchical features in the data, enabling the model to simultaneously capture large-scale oscillation modes and fine local structures. Therefore, U-FNO is particularly suitable for complex problems requiring both global continuity and local accuracy. Through accurate modeling of complex physical and chemical processes, U-FNO can reliably predict the spatial distribution dynamics of oil recovery and CO2 sequestration, demonstrating significant potential in CCUS-EOR applications.
[0063] In summary, thanks to its modular design and powerful multimodal modeling capabilities, LightUFTrans demonstrates significant application potential in oil production and CO2 sequestration prediction and optimization tasks in this invention.
[0064] Step S3, Multi-objective optimization: LightUFTrans is embedded as a surrogate model into the LA-MONSGA-II optimization framework, with cumulative crude oil production, CO2 sequestration, and net present value (NPV) as optimization objectives to obtain the Pareto optimal solution.
[0065] Surrogate-Assisted Multi-Objective Evolutionary Algorithm (SA-MOEA) is particularly suitable for computationally expensive multi-objective optimization problems, such as those relying on numerical simulation software or physical experiments for data acquisition, due to its advantages of low time cost, high computational efficiency, and strong adaptability. NSGA-II, as a classic and efficient SA-MOEA method, hierarchically stratifies solutions in the population through non-dominated sorting. In the non-dominated sorting process, individuals are ordered according to Pareto dominance, with individuals in the first layer (Pareto front) not dominated by any other solution, thus representing the current optimal solution set. This sorting mechanism enables NSGA-II to effectively search for Pareto optimal solutions in multi-objective optimization problems.
[0066] To maintain population diversity, NSGA-II measures the distribution of solutions in the target space by calculating the crowding distance between individuals. A larger crowding distance indicates that the individual is more sparse in the target space and has greater exploration potential. During population updates, individuals with larger crowding distances are preferentially retained to avoid over-clustering and maintain a uniform distribution of the Pareto front solution set. Furthermore, NSGA-II employs an elitist retention strategy, preserving the non-dominated optimal solution from the previous generation in each iteration to prevent the loss of the global optimum due to random crossover or mutation. With its efficient search capabilities and elitist strategy, NSGA-II has been widely applied to complex multi-objective optimization problems, such as reservoir production optimization and energy system optimization.
[0067] This invention proposes a multi-objective non-dominated sorting genetic algorithm II (LA-MONSGA-II) based on the LightUFTrans surrogate model. By introducing the LightUFTrans surrogate model, LA-MONSGA-II can achieve efficient optimization while significantly reducing the number of evaluations of high-cost functions, thereby improving computational efficiency while ensuring a wide and uniform distribution of the Pareto front.
[0068] The termination condition for multi-objective optimization is set to a fixed maximum number of iterations, i.e., 1000 generations. This setting ensures sufficient convergence and exploration of the Pareto front while also considering computational feasibility. During the surrogate model training process, an early stopping strategy is introduced to prevent overfitting. Specifically, training stops when the validation set loss no longer decreases over several consecutive training iterations, and the model parameters at which the validation performance is optimal are retained, thereby ensuring good generalization ability of the surrogate model during the optimization process.
[0069] By combining an elite retention strategy, a selection mechanism based on congestion distance, and a proxy model with an early shutdown strategy, LA-MONSGA-II can achieve a balanced trade-off between cumulative oil production, CO2 sequestration, and net present value (NPV), while maintaining the diversity of the Pareto frontier.
[0070] Taking into account key factors such as oil revenue, CO2 processing costs, carbon trading revenue, and tax incentives, this invention constructs an economic evaluation model to estimate the net present value (NPV), as shown below: (18); (19); in, This represents the revenue generated from hydrocarbon production in year n. Represents the operating costs for the same year; The incentive price is set as the price for CO2 geological sequestration using EOR technology in year n. Indicates the discount rate; The hydrocarbon production in year n; Indicates the price of hydrocarbons; Indicates the cumulative water injection volume; Indicates the cumulative gas injection volume; This represents the water injection cost corresponding to the cumulative water injection volume; This represents the CO2 procurement cost corresponding to the cumulative gas injection volume; This represents the cumulative gas production in year n. Represents the cost of CO2 recycling; Tax credits for carbon sequestration. In this example, the carbon tax is set at $35 per tonne.
[0071] This economic model is primarily used to evaluate the effectiveness of the proposed optimization framework. To simplify the calculation process, this invention sets the discount rate to 0, i.e., ignoring the time value of money, thereby avoiding complex economic discounting analysis and allowing the evaluation to focus more on the overall performance of the framework. Specifically, the oil price is $60 / STB, the CO2 purchase price is $25.2 / ton, the gas reinjection cost is $0.6 / MSCF, and the water treatment cost is $0.85 / STB.
[0072] To verify the model performance, several commonly used evaluation metrics were employed, including root mean square error (RMSE) and coefficient of determination (R²). 2 The mean absolute error (MAE), mean absolute percentage error (MAPE), and structural similarity index (SSIM) are used to measure the error between predicted and observed values.
[0073] (1) Root mean square error (RMSE) is used to measure the deviation between the predicted value and the observed value. It is a commonly used indicator to evaluate the performance of a model. The smaller the RMSE value, the higher the accuracy of the model prediction.
[0074] (20); in, Let be the predicted value for the i-th sample. These are the corresponding observations; This represents the number of samples.
[0075] (2) Coefficient of determination (R) 2 The value of is used to measure the ability of a regression model to explain the total variation of the dependent variable. Its value ranges from 0 to 1, and the closer it is to 1, the better the model fit.
[0076] (twenty one); (3) Mean Absolute Error (MAE) represents the average absolute value of the prediction error, which can avoid the cancellation of positive and negative errors, thus reflecting the magnitude of the error more accurately.
[0077] (twenty two); (4) Mean Absolute Percentage Error (MAPE) represents the average relative error between the predicted and observed values. Compared to RMSE, MAPE normalizes the error, thereby reducing the impact of extreme values.
[0078] (twenty three); (5) Structural Similarity Index (SSIM) is used to evaluate the quality of spatial prediction results. It is comprehensively evaluated by comparing the similarity between the predicted and observed values in terms of brightness, contrast, and structure. The closer the SSIM value is to 1, the more similar the structures are and the stronger the model's spatial prediction ability.
[0079] (twenty four); in, and These are the predicted value and the observed value, respectively. and Indicates average brightness; and For variance; For covariance; , is a constant used for stability calculations.
[0080] Step S4: Combining SHAP, attention mechanism, and spatiotemporal verification, analyze the model decision mechanism from multiple dimensions, including global features, temporal dependencies, and spatial distribution.
[0081] (1) Quantitative analysis of the importance of key engineering parameters based on SHAP; The SHAP method was used to analyze the feature importance in the LightUFTrans model. The model predictions and the SHAP feature importance ranking were highly consistent. The parameters that play a key controlling role in cumulative oil production and net present value (NPV) are production well surface production rate (PRO STL), gas injection rate / injection time (GASI), and bottom hole flowing pressure (BHP), as well as the secondary parameter water injection rate / injection time (WATI). For CO2 sequestration, the key parameters are gas injection rate (GASI), gas injection time (GASTIME), and water injection time (WATERTIME), which is consistent with the traditional understanding of oil and gas reservoir mechanisms and verifies the reliability of the model.
[0082] (2) Temporal pattern interpretation based on attention mechanism: dynamic analysis of CO2 sequestration and oil production; In oil production forecasting, the fluctuation of attention weights is highly consistent with the key trends in historical data. A conservative strategy is adopted, resulting in a smooth output fluctuation and a small overall forecast bias. In CO2 sequestration prediction, the model effectively identifies extreme points and inflection points, captures subtle changes in time series, and has a good tracking effect on the fluctuation trend of the optimization scheme, demonstrating excellent out-of-distribution generalization ability.
[0083] (3) Interpretability analysis of CO2 and oil saturation distribution based on spatiotemporal verification; To compensate for the shortcomings in spatial interpretability, a spatiotemporal validation method was introduced to systematically analyze the spatial distribution and evolution of CO2 and oil saturation predicted by the LightUFTrans model. This not only verified the reliability and strong generalization ability of the LightUFTrans model in predicting the spatial distribution of CO2 and oil saturation, but also provided a new perspective for evaluating the accuracy and robustness of the prediction and optimization results.
[0084] In summary, by integrating SHAP analysis, attention mechanisms, and spatiotemporal validation methods, a systematic interpretability analysis of the LightUFTrans model was conducted from three dimensions: global, temporal, and spatial. SHAP analysis was used to quantify the global importance of input parameters, the attention mechanism was used to identify key time steps in the time series, and the spatiotemporal validation method further revealed the spatial distribution variation characteristics. This multidimensional interpretability analysis deepened the understanding of the model's prediction mechanism, constructed a clear explanation path, and improved the transparency and reliability of the optimization process.
[0085] Example 2 This embodiment verifies the prediction accuracy of the proposed framework for static and dynamic data in the CCUS-EOR system, as well as its spatiotemporal prediction capability. It evaluates the optimization performance of the framework in practical engineering scenarios and conducts an in-depth interpretability analysis of the model based on SHAP, attention mechanism, and spatiotemporal verification method.
[0086] Before conducting prediction and optimization, the LightGBM model (combining fair segmentation tree and synthetic minority class oversampling technique, i.e., FCT-SMOTE-LightGBM) was used to conduct a preliminary assessment of the CO2-EOR and storage potential of a typical reservoir block. The model has been validated in various reservoir types, and the assessment results are good, indicating that this low-permeability reservoir is suitable for CO2-EOR development.
[0087] I. Performance evaluation of LightUFTrans in CO2-EOR and static and dynamic prediction of storage.
[0088] A hybrid hyperparameter optimization strategy is employed. First, Bayesian optimization and stochastic search are used to efficiently explore the global hyperparameter space and locate potentially superior regions. Then, manual fine-tuning is performed within a limited range around the optimal configuration obtained through automatic search to further optimize model performance. This step aims to compensate for model-specific sensitivities and empirical characteristics that automatic optimization methods struggle to capture. Manual tuning is used only as a supplementary method and is strictly limited to a small range to avoid redundancy with automatic optimization.
[0089] The dataset was divided into training and test sets in a ratio of 0.85:0.15, ensuring sufficient model training while retaining a separate test set for performance evaluation. Table 4 summarizes the final hyperparameter settings of the LightUFTrans model.
[0090] Table 4 shows the range of hyperparameters of the proposed model and their optimal values for prediction tasks using static tabular data.
[0091] Figure 4 This demonstrates its performance in static prediction tasks. The predicted points on both the training and test sets are closely distributed around the 45° diagonal, indicating a good model fit. Therefore, LightUFTrans exhibits excellent performance in predicting static indicators such as cumulative crude oil production, CO2 sequestration, and net present value (NPV). Specific metrics are shown in Table 5: On the test set, R... 2 The R² values were 0.9618, 0.9631, and 0.9625, respectively; the RMSE values were 0.1088, 0.5677, and 0.3705, respectively; and the MAPE values were 0.0079, 0.0247, and 0.0070, respectively. This indicates that the model has high prediction accuracy on different static targets. The high R² values... 2 The values reflect the model's strong explanatory power for data variability, while the lower RMSE and MAPE further validate the model's reliability and accuracy.
[0092] Table 5 shows the prediction performance metrics of the proposed model for three different static targets on the training and test sets.
[0093] Table 6 also compares the performance of the proposed model with several mainstream baseline models. All models were optimized using a unified strategy of random search combined with manual tuning to ensure fairness in the comparison. Baseline models included XGBoost, Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Support Vector Regression (SVR), Multilayer Neural Network (MLNN), and Radial Basis Function Network (RBF). The results show that tree-based models (XGBoost, RF, GBDT) generally outperform shallow models (SVR, RBF). Among them, LightGBM (R 2 =0.9625) and XGBoost (R 2 =0.9567) performed the best, demonstrating the advantages of tree models in handling complex nonlinear relationships and feature interactions. LightGBM showed particularly excellent generalization ability on high-dimensional complex data, while XGBoost followed closely behind. In contrast, deep neural networks (MLNN, R 2The performance of MLNN (0.9449) is superior to shallow models (SVR: 0.7830; RBF: 0.9231), but it is still slightly inferior to ensemble tree models on structured data. Although MLNN has advantages in nonlinear modeling, its performance is still limited by the data size. As the data size increases, its performance is expected to improve further.
[0094] Table 6 compares the predictions of the proposed model with other baseline models on the test set for static tabular data.
[0095] For dynamic prediction, the tree-structured Parzen estimation (TPE) method based on Bayesian optimization was adopted, and hyperparameter optimization was performed in combination with the Optuna framework to obtain near-optimal parameter configurations, as shown in Table 7. The prediction results are shown in Table 8.
[0096] Table 7 shows the range of hyperparameters of the proposed model and their optimal values for predicting dynamic time series data.
[0097] The result is the average error of 1069 sets of tests: For oil production prediction, R 2 The R² value reached 0.9283, MAPE was 0.0334, and RMSE was 0.1479; for dynamic prediction of CO2 sequestration, R²... 2 The R-value reached 0.9473, MAPE was 0.0309, and RMSE was 0.0448. The overall average R-value was... 2 The MAPE and RMSE values were 0.9378, 0.0322, and 0.0964, respectively, indicating that the model can accurately capture the dynamic changes of the CO2-EOR process under different engineering parameter conditions.
[0098] Table 8 shows the dynamic prediction performance of the proposed model on time series data.
[0099] From a visualization perspective, Figure 5 and Figure 6 The model demonstrates predictions for the next 300 steps based on 600 steps of historical data. It accurately captures the trends in oil production and CO2 sequestration, successfully identifying inflection points, although slight deviations exist at these points. For oil production, the P25–P75 prediction range performs best, covering all true values under relatively lenient conditions. For CO2 sequestration dynamics, the P10–P90 range is more suitable, also covering all true values. Although the P2 and P98 ranges have high accuracy, their wide range limits their practical application value.
[0100] Table 9 further compares the performance of this model with several mainstream time series models, including Historical Average (HA), ARIMA, LSTM, and GRU. The results show that HA and ARIMA have higher R-values. 2 The low values (all below 0.3559) indicate that it is difficult to effectively capture data fluctuations, and the prediction results tend to be average, failing to reflect the true dynamic trend. In contrast, LSTM and GRU significantly improve performance (R0). 2 The R values were 0.8944 and 0.9127 respectively, but still lower than the model proposed in this invention. The proposed model performs best across all metrics (R²). 2 =0.9378, MAPE=0.0322, RMSE=0.0964), compared to LSTM and GRU, R 2 The improvement of 2.68%–4.63%, the reduction of MAPE by 32.35%–60.45%, and the reduction of RMSE by 26.76%–45.67% fully demonstrate its superiority in dynamic prediction tasks.
[0101] Table 9 Comparison of Dynamic Prediction Models
[0102] All experiments were conducted on a computing platform equipped with an Intel i5-13600KF CPU and an NVIDIA RTX 4070Ti (12GB VRAM). Under these hardware conditions, LightUFTrans's training time for static prediction was approximately 1.52 seconds, and for dynamic prediction, it was approximately 6,719.46 seconds. After training, the inference time for static prediction on the entire test set was only about 0.006 seconds, demonstrating extremely high computational efficiency, especially suitable for multi-objective optimization processes. For the dynamic prediction task, the overall inference time was approximately 28 seconds, and it supports up to 300 future prediction steps, meeting the near real-time monitoring and analysis needs of practical engineering applications.
[0103] II. Spatiotemporal modeling of CO2 and oil saturation based on LightUFTrans.
[0104] To more intuitively understand the spatiotemporal dynamics of oil recovery and CO2 sequestration, this invention extends the prediction task from tabular data to spatial data. However, spatial prediction for all 1069 sets of experiments would incur high computational costs. To improve computational efficiency, K-means clustering was used to divide these 1069 sets of experiments into 40 clusters. Within each cluster, a distance-based uniform sampling strategy was used to select 4 representative experiments, resulting in 160 representative samples. The clustering results were visualized using principal component analysis (PCA), as shown below. Figure 7 As shown.
[0105] This method significantly reduces the computational load required for spatial prediction while ensuring sample representativeness and uniform distribution, thus substantially improving computational efficiency. To further optimize efficiency while maintaining prediction accuracy, the top 10 layers with the highest oil distribution density were selected for modeling and analysis. Finally, the dataset was divided into training and testing sets in an 8:2 ratio, and model training was conducted based on a three-dimensional field-scale reservoir model.
[0106] Figure 8 and Figure 9 This paper presents the predicted distributions of oil saturation and CO2 saturation in layers 3 and 7 at three different time steps in a randomly selected test case (Case 823), including the actual distribution, predicted distribution, and spatial error map. Due to space limitations, only these two layers are shown from the complete set of 10 layers. To mitigate boundary effects and ensure the representativeness of the internal layer predictions, layers 3 and 7 are selected to represent shallow and deep layers, respectively, to highlight the model's performance under different depth conditions.
[0107] In Case 823, the model successfully captured a key characteristic: shallow oil is easier to extract, while deep oil is more difficult to recover. For example... Figure 8 As shown, the oil saturation in layer 7 is generally lower (lighter in color compared to layer 3), but its displacement resistance is stronger, especially in the lower right region. Furthermore, the model demonstrates high accuracy in predicting oil displacement in the well perimeter area and effectively characterizes the dynamic changes of the displacement front. At different time steps, the model's predictions for high, medium, and low saturation regions (corresponding to dark red, light red, and white regions, respectively) are highly consistent with the numerical simulation results. However, as time progresses, the model's prediction accuracy in the displacement front region slightly decreases; for example, in layer 3, R... 2 The SSIM decreased from 0.9981 to 0.991, and the overall prediction accuracy decreased from 0.9968 to 0.9906, but the overall prediction accuracy remained high.
[0108] Compared to oil saturation prediction, the model's performance in predicting CO2 saturation distribution has improved over time. Figure 9 As shown. This is mainly attributed to the increased CO2 accumulation with prolonged injection time, making it easier for the model to capture its spatiotemporal evolution characteristics. Similar to oil distribution, CO2 saturation prediction accuracy is higher in shallow layers. However, in deeper layers, the model's prediction accuracy for CO2 saturation is lower than that for oil, possibly due to the upward migration of CO2 during displacement, resulting in limited retention in deeper layers. This phenomenon indicates that CO2 tends to accumulate in the upper strata after injection, thus forming tectonic sequestration. Despite some errors in the boundary region, the model can still accurately predict the overall trend of CO2 migration. For example, at time step 11 of layer 7, R... 2The value is 0.8966, and SSIM is 0.9195; while in the 3rd layer at the same time step, R... 2 The SSIM values were 0.9735 and 0.9762, respectively.
[0109] Table 10 presents the model's average prediction performance on the test set for the distribution of oil and CO2 saturation, demonstrating that the model has high prediction accuracy. For oil saturation, the model's R²... 2 The value is 0.9424, SSIM is 0.9335, and RMSE is 0.0430; for CO2 saturation, R... 2 The RSI is 0.9048, SSIM is 0.9392, and RMSE is 0.0309. Overall, the model's average performance is R... 2 =0.9236, SSIM=0.9364, RMSE=0.0370. This indicates that, under engineering uncertainties, this model can effectively predict the dynamic processes of oil recovery and CO2 sequestration in low-permeability reservoirs, and has good potential for engineering applications.
[0110] Table 10 shows the spatial prediction performance of the proposed model on the test set.
[0111] Furthermore, the quantitative comparisons in Table 11 further demonstrate that the proposed LightUFTrans model exhibits significant advantages over commonly used convolutional neural networks (CNNs) and operator learning methods (FNOs) on the test set. LightUFTrans in R... 2 Higher than SSIM and lower RMSE indicate improvements in both numerical accuracy and structural consistency. Compared to widely used CNN models, this performance improvement highlights the limitations of pure convolutional structures in capturing long-range spatial dependencies and multi-scale contextual information. Compared to FNO models, which focus on global operator learning, LightUFTrans still shows a stable improvement, indicating its superior ability to fuse local and global spatial features. This advantage is mainly attributed to its U-Net encoder-decoder structure and skip connection mechanism, which enables multi-scale feature fusion while preserving fine-grained spatial information, thus achieving higher accuracy in saturation field prediction.
[0112] Table 11 Performance comparison of different models on the test set
[0113] III. Multi-objective optimization of CO2-EOR and storage performance based on LA-MONSGA-II: Pareto front and hypervolume comparison analysis.
[0114] First, the hybrid model LightUFTrans was trained using experimental data generated from 1069 sets of different engineering parameters to predict cumulative oil production, CO2 sequestration, and net present value (NPV). Then, based on the prediction results of the LightUFTrans model, the Pareto front was calculated using the LA-MONSGA-II algorithm, and the optimization results were further validated through reservoir numerical simulation. Specific algorithm parameter settings are shown in Table 12.
[0115] Table 12 Hyperparameters in the LA-MONSGA-II algorithm
[0116] Based on this, the optimal combination of engineering parameters was selected, and the LightUFTrans model was used to visualize and analyze the time series changes of oil production rate and CO2 sequestration rate, as well as the spatial distribution of CO2 and oil saturation.
[0117] The results show that the model proposed in this invention can effectively identify ideal optimization solutions. Furthermore, this optimization method exhibits significant advantages in computational efficiency. Pareto front calculations based on the surrogate model typically take only a few seconds, achieving a speedup of approximately 2–3 orders of magnitude (approximately 10^6 seconds) compared to traditional numerical simulations (approximately 30–50 minutes per run). 2 -10 3 (Times). Considering the 1069 initial simulations required during the surrogate model training phase, the overall computational efficiency improvement of the surrogate-assisted optimization framework proposed in this study is approximately 10 times. 5 -10 6 It has a significant advantage over optimization methods based on direct numerical simulation.
[0118] Furthermore, in practical applications, the optimization process typically requires multiple independent runs (e.g., with different random seeds, hyperparameter settings, or number of iterations), and each additional run yields a speedup of the same magnitude. Therefore, as the number of optimization tasks increases, the cumulative computational savings brought by this method grow linearly, thus providing an efficient and scalable alternative to traditional numerical simulation-based optimization methods.
[0119] To ensure a fair comparison of the performance of different optimization algorithms, five independent runs of SMPSO, MOPSO, IBEA, NSGA-II, and NSGA-III were performed under the same parameter settings, and the results were averaged to evaluate the Pareto front quality and hypervolume performance. The LightUFTrans model was used as a surrogate model in all experiments.
[0120] Figure 10Figure 'a' shows the Pareto front comparison results of five algorithms, with NSGA-II generally outperforming the others, especially in the trade-off region between cumulative oil production and CO2 sequestration. The solution set obtained by NSGA-II (orange squares) is significantly better than SMPSO (green dots) and MOPSO (blue diamonds) in the upper, middle, and lower regions, indicating that it can obtain superior solutions. This advantage is mainly attributed to NSGA-II's superiority in non-dominated sorting and population diversity maintenance. Compared to Particle Swarm Optimization (PSO) algorithms, NSGA-II has a greater advantage in Pareto front exploration and distribution preservation.
[0121] Figure 10 Figure b presents the hypervolume comparison results for each algorithm. NSGA-II obtained the largest hypervolume value, further indicating its optimal overall performance in multi-objective optimization. Generally, a larger hypervolume indicates a wider coverage of the solution set in the objective space, and it is closer to the optimal front, reflecting the diversity and superiority of the solutions. The reason why NSGA-II outperforms NSGA-III may lie in the characteristics of three-objective optimization problems: although NSGA-III performs well in high-dimensional optimization through reference directions, its complex mechanism reduces efficiency in low-dimensional objective spaces, while NSGA-II is more flexible and efficient. In addition, compared with IBEA, NSGA-II achieves a better balance between computational complexity and convergence performance, making it more suitable for the multi-objective optimization problem in this invention. Therefore, this invention combines NSGA-II with a surrogate model to propose the LA-MONSGA-II algorithm for three-objective optimization.
[0122] Figure 10 Figure c shows the Pareto front comparison between the LA-MONSGA-II optimization results and the original data. The results indicate that the optimized Pareto optimal solution set significantly improves cumulative oil production, CO2 sequestration, and net present value (NPV): cumulative crude oil production increased from 6.36 × 10⁻⁶ to 6.36 × 10⁻⁶. 5 The tonnage increased to 7.85 × 10 5 tons, CO2 storage capacity increased from 1.11 × 10 6 The tonnage increased to 1.24 × 10 6 tons, NPV from 3.68×10 8 The US dollar rose to 4.37 x 10 8 Dollar.
[0123] Furthermore, the Pareto optimal solution generated by LA-MONSGA-II was input into a numerical simulator for verification. Figure 11 The results were compared with those of the Monte Carlo experiment: the optimal result obtained by the Monte Carlo method was a cumulative crude oil production of 7.67 × 10⁻⁶. 5 Tons, CO2 storage capacity 1.27 × 106 tons, NPV is 4.28×10 8 US dollars; while the optimized results of this invention reached 8.06 × 10⁻⁶. 5 Tons, 1.47 × 10 6 tons and 4.43 × 10 8 US dollars. Compared to traditional methods, the method of this invention increases cumulative crude oil production by at least 3.88 × 10⁻⁶. 4 The amount of CO2 stored increased by 2.03 × 10⁻⁶ tons. 5 tons, NPV increased by 1.51×10 7 The three core indicators, namely US dollars, increased by 5.06%, 15.98%, and 3.53%, respectively. Comparative verification based on Monte Carlo results further demonstrates the effectiveness of the proposed optimization algorithm in improving the overall benefits of the CCUS-EOR project.
[0124] In summary, this embodiment verifies the superior performance of the LA-MONSGA-II algorithm in multi-objective optimization, providing an efficient and reliable optimization decision-making method for the CCUS-EOR project.
[0125] IV. Multidimensional interpretability analysis of the framework proposed in this invention.
[0126] In deep learning frameworks, interpretability analysis is crucial for revealing the predictive mechanisms of models. This invention conducts a systematic interpretability analysis of the model from multiple dimensions, specifically including: (1) Quantitative analysis of the importance of key engineering parameters based on SHAP.
[0127] The feature importance in the LightUFTrans model was analyzed using the Shapley Additive Explanations (SHAP) method. SHAP is a widely used posterior interpretability method used to quantify the contribution of each input parameter to the model's prediction results. Through SHAP analysis, the engineering parameters with key impacts on cumulative oil production, CO2 sequestration, and net present value (NPV) under different engineering scenarios were identified.
[0128] The model in this invention and SHAP analysis show a high degree of consistency in the ranking of the importance of features in cumulative oil production and net present value (NPV), both identifying key parameters (such as PRO STL, GASI, and BHP) and minor parameters (such as WATI). Surface fluid production (PRO STL) has a significant impact on cumulative oil production and NPV, as the upper limit of fluid production usually directly constrains oil and gas production capacity, thus further affecting economic returns. Gas injection rate (GASI) is also a key factor affecting recovery efficiency, with its importance almost twice that of other parameters. This is because a higher injection rate can rapidly expand the displacement range, especially under miscible displacement conditions, where the interaction between gas and oil can reduce fluid viscosity, thereby significantly improving recovery efficiency. Furthermore, bottom hole pressure (BHP) plays an important role in displacement efficiency and ultimate recovery by influencing miscibility conditions and pressure gradients.
[0129] In CO2 sequestration prediction, the model results are highly consistent with the SHAP ranking. Key parameters (such as GASI, GASTIME, and WATER TIME) are identified as the most important factors, while secondary parameters (such as WATI and PRO BHP) and moderately important parameters (such as PRO STL) exhibit similar ranking patterns. Furthermore, the contribution proportions in the histograms are almost identical. GASI has a particularly significant impact on CO2 sequestration, with its importance nearly twice that of other parameters, as it directly determines the amount of CO2 injected and the scale of storage. Although GASTIME and WATER TIME have a significant impact on CO2 sequestration, their impact on oil production and NPV is relatively weak, reflecting the inherent trade-off between recovery and sequestration. Specifically, these two parameters affect sequestration uniformity by regulating CO2 sweep efficiency, but their direct impact on recovery is limited by dominant factors such as injection rate and displacement pressure. Meanwhile, parameters that significantly affect recovery, such as PRO STL and PRO BHP, have a smaller impact on CO2 sequestration, further highlighting the differences between the two objectives. Overall, the ranking of the importance of this feature is highly consistent with traditional mechanistic understanding.
[0130] like Figure 12 As shown in the bar chart of average SHAP values ( Figure 12 (a, d, g) and bee colony diagram ( Figure 12 The contribution plots of b, e, and h in the model ranking (and model ranking) Figure 12 The parameters c, f, and i in the model are visualized. Local interpretation analysis reveals the differences in harvesting and storage performance among different samples. Figure 12 Figures b and h indicate that high PRO STL and high GASI (red) have a positive promoting effect on oil production and NPV, while lower BHP (blue) helps to improve displacement efficiency by increasing pressure differential. Figure 12 As shown in Figure 'e', high GASI and long GAS TIME (red), as well as lower WATER TIME and PRO STL, are beneficial for enhancing CO2 sequestration capabilities. Furthermore, the figure also presents the distribution characteristics of parameter effects. Figure 12 The left-tailed distribution of PRO STL, GASI, and GASTIME in b and h indicates that these parameters have a stronger negative impact on model predictions at lower values. Similarly, although GASTIME ranks low in global importance, its SHAP value exhibits a significant negative long tail, suggesting that it may still play a crucial role in specific scenarios.
[0131] Overall, the proposed model's feature importance ranking in oil recovery and CO2 sequestration is highly consistent with the SHAP analysis results and aligns with domain knowledge. This consistency validates the model's interpretability and enhances its decision-making reliability in practical CCUS-EOR applications.
[0132] (2) Temporal pattern interpretation based on attention mechanism: dynamic analysis of CO2 sequestration and oil production.
[0133] To address the shortcomings of SHAP in explaining time dependencies, this paper further utilizes the attention mechanism in the LightUFTrans model to analyze time series patterns. Attention weights reflect the importance of historical time steps to the current prediction, thus revealing the model's temporal decision-making basis.
[0134] like Figure 13 As shown, the historical and future values, attention weights, and confidence intervals over time steps are visualized. Figure 13 The 'a' in the figure indicates that the significant fluctuations in attention weights are highly consistent with key trends in historical data. However, in the optimized scheme, due to the higher production level, the time series fluctuations are more drastic, making it difficult for the model to accurately capture inflection points, thus affecting its out-of-distribution generalization ability. Therefore, the model tends to adopt a relatively conservative strategy when predicting oil production, resulting in lower prediction fluctuations. Nevertheless, the overall prediction bias of the model is small.
[0135] In comparison, Figure 13 The 'b' in the model indicates that it can effectively identify extreme points (inflection points) and accurately capture subtle changes in CO2 sequestration prediction. Prediction results show that the model can track the fluctuation trends in the optimized scheme well, with only slight deviations at a few locations. Therefore, compared to oil production prediction, the model of this invention exhibits stronger out-of-distribution generalization ability in dynamic CO2 sequestration prediction.
[0136] (3) Interpretability analysis of CO2 and oil saturation distribution based on spatiotemporal verification.
[0137] To address the shortcomings of previous research in spatial interpretability, spatiotemporal verification methods are further introduced. For example... Figure 14 As shown, the prediction results for layers 3 and 7 indicate that oil saturation gradually decreases with increasing depth. Although there are some deviations in local details, especially at time step 11 of layer 3, the model is still able to effectively capture the overall boundary evolution trend. This difference may stem from the uneven distribution of data in the optimization scheme, which typically corresponds to higher oil production. This can be seen by comparing... Figure 14 The optimization results and Figure 8 The original distribution of (first row) is further verified.
[0138] At time step 11, significant oil displacement occurred around the injection well in the upper right region, such as... Figure 14 The error distribution plot (third row, third column) in section a is shown. This result verifies the reliability of the optimization scheme, indicating that oil recovery in the vicinity of the injection well has been significantly enhanced. Compared to traditional studies that typically focus only on total production indicators, the method of this invention can deeply analyze spatial distribution patterns, thus providing a higher level of interpretability.
[0139] Figure 15 The CO2 saturation predictions for layers 3 and 7 also show a decreasing trend with increasing depth, successfully capturing the overall boundary changes. Figure 9 Compared to b in the middle, Figure 15 The overall color of b in the image is darker, further validating the effectiveness of the optimization scheme. From an interpretability perspective, the improvement in storage efficiency mainly stems from enhanced deep CO2 retention, rather than shallow CO2 retention. This can be demonstrated through... Figure 15 a in Figure 9 The small color difference between 'a' values in the model provides further evidence. This suggests that shallow CO2 sequestration may be nearing its saturation limit, and that optimization strategies should focus on enhancing deep sequestration capabilities.
[0140] Further analysis revealed that in the optimized scheme, the predicted CO2 saturation R at the three time steps of the 7th layer... 2 They reached 0.9907, 0.9151, and 0.9113 respectively, such as Figure 15 b in the value is significantly higher than Figure 9 The values of 0.8258, 0.8867, and 0.8966 corresponding to 'b' in the model demonstrate its strong generalization ability in out-of-distribution predictions. This advantage may stem from the fact that higher CO2 saturation makes the model more readily able to identify its spatial distribution characteristics. Furthermore, although at the 11th time step of the 3rd layer, such as... Figure 15 In the third column of column a, the model prediction accuracy decreased slightly, but it can still accurately reflect the overall distribution of the shallow CO2 storage range.
[0141] These results demonstrate that the proposed model exhibits good generalization ability in out-of-distribution prediction of CO2 saturation distribution. Meanwhile, Figure 15 The error distribution in a (third row, third column) also reveals the trend of CO2 migrating to deeper layers over time, further clarifying the potential migration path and distribution range of CO2 under efficient storage strategies.
[0142] Therefore, spatiotemporal prediction effectively compensates for the limitations of tabular data-based prediction models, provides a new perspective for evaluating the accuracy and robustness of prediction and optimization results, and serves as an indirect means of verification and interpretation, thereby significantly enhancing the engineering application value and credibility of the model.
[0143] By integrating SHAP analysis, attention mechanisms, and spatiotemporal validation methods, a systematic interpretability analysis of the LightUFTrans model was conducted from three dimensions: global, temporal, and spatial. SHAP analysis quantifies the global importance of input parameters, the attention mechanism identifies key time steps in the time series, and the spatiotemporal validation method further reveals spatial distribution variations. This multidimensional interpretability analysis significantly deepens the understanding of the model's prediction mechanism, constructs a clear explanation path, and thus improves the transparency and reliability of the optimization process.
[0144] Therefore, this invention employs the aforementioned interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage, constructing the LightUFTrans multimodal proxy model. This innovatively integrates quantitative evaluation, dynamic trend prediction, and spatial distribution prediction to accurately characterize the spatiotemporal dynamic evolution of underground CO2 / oil, significantly improving computational efficiency, greatly reducing engineering design and decision-making costs, and filling key cognitive gaps in underground system modeling, achieving a comprehensive characterization of the system. Through multidimensional interpretability analysis, key engineering control parameters are clarified, their temporal dynamic characteristics and spatial migration mechanisms are analyzed, overcoming the limitations of black-box models and enhancing engineering credibility. Furthermore, by further integrating multi-source data, optimizing multimodal learning methods, and expanding data scale and key parameters, this invention can further enhance its applicability in integrated geological systems, adapting to the needs of multiple fields such as geothermal development, CO2 geological storage, and intelligent optimization of underground resources.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage, characterized in that, Includes the following steps: Step S1, Multimodal Data Acquisition: Using the CMG-GEM component simulator, numerical simulation is carried out through Monte Carlo sampling to acquire a multimodal data pipeline that integrates static tabular data, dynamic time series data, and spatial distribution data; Step S2, Hybrid Model Construction: Build the LightUFTrans hybrid model, which integrates a lightweight gradient booster, a temporal fusion Transformer, and a U-Net enhanced Fourier neural operator to process static, temporal, and spatial data respectively, and achieve three-dimensional prediction of CO2-EOR and sealing performance. Step S3, Multi-objective optimization: LightUFTrans is embedded as a surrogate model into the LA-MONSGA-II optimization framework, with the cumulative crude oil production, CO2 storage, and net present value as optimization objectives, to obtain the Pareto optimal solution; Step S4, Multidimensional Interpretability Analysis: Combining SHAP, attention mechanism, and spatiotemporal validation, the model decision-making mechanism is analyzed from the dimensions of global features, temporal dependence, and spatial distribution.
2. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 1, characterized in that, In step S1, the CMG-GEM component simulator is used to simulate the multiphase and multicomponent flow behavior involved in CO2 injection and oil recovery. The specific process is as follows: Based on the Peng–Robinson equation of state, by introducing the temperature-dependent attraction term κ and fluid characteristic parameters a and b, the gas-liquid phase equilibrium behavior under reservoir conditions is predicted, as shown below: (1); in, Indicates pressure; It is the gas constant; Indicates temperature; Indicates molar volume; and These are fluid characteristic parameters; For temperature-related attraction terms; In a multiphase, multicomponent system with CO2 injection, the mass conservation equation is used to describe the spatiotemporal evolution of each component in all petroleum phases, as shown below: (2); in, Indicates the porosity of porous media; Indicates phase saturation; Indicates phase mass density; Indicates that component i is in phase The mole fraction in; Indicates phase Darcy velocity is used to describe the volumetric flow behavior under the influence of pressure gradient and gravity. Represents the source and sink terms of component i, used to characterize the injection or production process; For each fluid phase Darcy's velocity is described by the multiphase Darcy's law, as shown below: (3); in, Indicates absolute penetration rate; Indicates phase Relative penetration rate; Indicates phase viscosity; Indicates phase The pressure; Represents gravitational acceleration; Indicates depth; The CO2 sequestration mechanism is comprehensively characterized by considering three forms: structural sequestration, dissolution sequestration, and bound sequestration, as detailed below: The mass of CO2 to be sealed is calculated directly from the in-situ gas phase mass, as shown below: (4); in, Indicates the quality of CO2 structure sequestration; Represents a grid cell; Indicates porosity; Indicates the volume of a grid cell; Indicates CO2 gas phase saturation; Indicates the density of the gas phase; In the CMG-GEM simulation, the solubility of CO2 in brine is calculated using Henry's Law for dissolution and sequestration, as shown below: (5); in, This indicates the concentration of CO2 in the aqueous phase; This represents the CO2 fugacity calculated using the equation of state. is the Henry's constant, whose value depends on temperature T, pressure P, and ionic strength I; CMG-GEM simulates bound CO2 sequestration by introducing a relative permeability and capillary pressure model with hysteresis characteristics, thereby characterizing the transition process of the fluid from a kinetic to a kinetic state after WAG displacement; the total mass of bound CO2 is shown below: (6); in, This indicates the total mass of CO2 bound and sealed. This indicates the residual gas saturation.
3. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 2, characterized in that, The Monte Carlo sampling method was used for numerical simulation to generate static tabular data, historical time series data and spatial field data. The three types of data were then input into three different machine learning model components to predict the CO2-EOR and geological sequestration capacity, dynamic processes and spatial distribution characteristics, thereby achieving cross-modal constraints. The static table data includes: gas injection time, water injection time, gas injection rate, water injection rate, bottom-hole flowing pressure of injection well, bottom-hole flowing pressure of production well, and surface production rate of production well. The static table data is input into the lightweight gradient lift model, which outputs the cumulative crude oil production, CO2 storage, and net present value. Historical time series data includes: historical crude oil production rate and historical CO2 sequestration rate over time; historical time series data and static tabular data are input into the time series fusion Transformer model to output crude oil production rate and CO2 sequestration rate over time. The spatial field data includes: permeability distribution map and initial water saturation distribution map; the spatial field data and static tabular data are input into the U-Net enhanced Fourier neural operator model to output crude oil saturation distribution map and gas saturation distribution map.
4. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 3, characterized in that, In step S2, the process by which the lightweight gradient booster achieves predictive performance on static tabular data in the LightUFTrans hybrid model is as follows: First, gradient-based one-sided sampling prioritizes samples with large gradients, thereby significantly reducing computational costs while ensuring model accuracy. Secondly, the histogram-based algorithm discretizes continuous features into multiple intervals, effectively reducing memory usage and accelerating computation. Furthermore, the mutual exclusion feature bundling technique is used to combine non-overlapping features, reducing the number of effective features and improving computational efficiency. Finally, the leaf node-based growth strategy prioritizes splitting the leaf nodes with the highest gain, thereby reducing errors and improving the overall model performance.
5. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 4, characterized in that, In terms of time series forecasting, a time-series fusion Transformer is used to model the dynamics of CO2 sequestration and the oil production process. A dedicated encoding module for static parameters is introduced to enhance the model's sensitivity to static features. The specific process is as follows: First, a deep fusion of static and dynamic data is achieved through a static covariate encoder and a dedicated encoding module for static parameters. Second, a gated residual network is introduced to enhance the model's autoregressive modeling capability and enable the processing of complex time series data. Third, through a feature selection network and an interpretable multi-head attention mechanism, the temporal fusion Transformer identifies key influencing factors and generates attention weights, improving the model's interpretability. Fourth, the temporal fusion Transformer supports conditional prediction capabilities, combining known future information for prediction and achieving engineering optimization. Finally, the temporal fusion Transformer has the ability to process multi-scale temporal data, meeting the requirements of different time resolutions in CCUS-EOR. The static covariate encoder generates four distinct context vectors through independent GRN modules, used for temporal feature selection, local temporal feature processing, and enhancement and fusion of temporal features with static information, respectively. The GRN efficiently models the input through nonlinear transformations, as shown below: (7); (8); (9); (10); in, Represents the gated residual network; This represents the learnable network parameters; This represents the main input feature vector; Represents a static eigenvector; Presentation layer normalization operation; Represents a gated linear unit; Represents the activation function of the exponential linear unit; This represents the Sigmoid activation function; Indicates intermediate implicit variables; express The function's output variable; express The input variables of the function; , , , , Indicates weight; , , , Indicates bias; Represents the Hadamard product; In neural networks, the attention mechanism enables the model to focus on key parts of the input sequence during the prediction process, as shown below: (11); in, Represents the query matrix; Represents the key matrix; Represents a value matrix; Indicates the dimension of the key vector; The function is used to convert the scaled dot product result into a probability distribution; Represents single attention; Multi-head attention mechanisms introduce multiple attention heads, enabling the model to simultaneously focus on different parts of the input sequence, as shown below: (12); (13); in, , and Represent the i-th attention head respectively The weight matrix of queries, keys, and values; This represents the linear mapping weight matrix after multi-head attention output; Indicates the number of attention heads; This indicates a high level of attention from multiple parties. This indicates a concatenation operation, used to merge the outputs of multiple attention heads into a unified matrix; To enhance the interpretability of the model, an interpretable multi-head attention mechanism with shared-value weights is employed to highlight the importance of specific features, as shown below: (14); (15); (16); (17); in, This indicates that multi-headed attention can be explained; Represents the final mapping matrix. This represents the aggregated result of all attention head outputs; This represents the normalized attention weights. This represents a value mapping matrix shared among all attention heads; and Let represent the projection weight matrices of the query and key for the h-th attention head, respectively; This indicates the number of attention heads.
6. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 5, characterized in that, The LightUFTrans hybrid model introduces the U-Net enhanced Fourier neural operator to achieve spatial prediction of CO2 and oil saturation distribution, and is used to visualize and verify the optimization results. The U-Net enhanced Fourier neural operator introduces a multi-scale U-Net structural path to characterize the spatial hierarchical features in the data, enabling the model to simultaneously capture large-scale oscillation modes and fine local structures, and predict the spatial distribution dynamics of oil recovery and CO2 sequestration.
7. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 6, characterized in that, In step S3, LightUFTrans is embedded as a surrogate model into the LA-MONSGA-II optimization framework. The termination condition for multi-objective optimization is set to a fixed maximum number of iterations, which ensures that the Pareto front fully converges and explores while taking into account computational feasibility. During the training of the surrogate model, by combining an elite retention strategy, a selection mechanism based on crowding distance, and a surrogate model evaluation with an early shutdown strategy, LA-MONSGA-II obtains the Pareto optimal solution with the optimization objectives of cumulative crude oil production, CO2 storage, and net present value.
8. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 7, characterized in that, Taking into account oil revenues, CO2 processing costs, carbon trading revenues, and tax incentives, an economic evaluation model is constructed to estimate the net present value (NPV), as shown below: (18); (19); in, This represents the revenue generated from hydrocarbon production in year n. Represents the operating costs for the same year; The incentive price is set as the price for CO2 geological sequestration using EOR technology in year n. Indicates the discount rate; The hydrocarbon production in year n; Indicates the price of hydrocarbons; Indicates the cumulative water injection volume; Indicates the cumulative gas injection volume; This represents the water injection cost corresponding to the cumulative water injection volume; This represents the CO2 procurement cost corresponding to the cumulative gas injection volume; This represents the cumulative gas production in year n. Represents the cost of CO2 recycling; Tax credits for carbon sequestration; To verify the model performance, several commonly used evaluation metrics were employed, including root mean square error (RMSE) and coefficient of determination (R²). 2 Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and Structural Similarity Index (SSIM) are used to measure the error between predicted and observed values.
9. The interpretable data-driven prediction and optimization method for CO2 enhanced oil recovery and storage according to claim 8, characterized in that, In step S4, combining SHAP, attention mechanism, and spatiotemporal validation, the model decision mechanism is analyzed from multiple dimensions, including global features, temporal dependencies, and spatial distribution. The specific process is as follows: (1) Quantitative analysis of the importance of key engineering parameters based on SHAP; The SHAP method was used to analyze the feature importance in the LightUFTrans model. The model predictions and the SHAP feature importance ranking were highly consistent. The parameters that play a key controlling role in cumulative oil production and net present value (NPV) are the surface production rate of the production well, the gas injection rate / injection time, and the bottom hole flowing pressure, as well as the secondary parameter, the water injection rate / injection time. For CO2 sequestration, the key parameters are the gas injection rate, gas injection time, and water injection time, which is consistent with the traditional understanding of oil and gas reservoir mechanisms and verifies the reliability of the model. (2) Temporal pattern interpretation based on attention mechanism: dynamic analysis of CO2 sequestration and oil production; In oil production forecasting, the fluctuation of attention weights is highly consistent with the key trends in historical data. A conservative strategy is adopted, resulting in a smooth output fluctuation and a small overall forecast bias. In CO2 sequestration prediction, the model effectively identifies extreme points and inflection points, captures subtle changes in time series, and has a good tracking effect on the fluctuation trend of the optimization scheme, demonstrating excellent out-of-distribution generalization ability. (3) Interpretability analysis of CO2 and oil saturation distribution based on spatiotemporal verification; A spatiotemporal validation method was introduced to systematically analyze the spatial distribution and evolution of CO2 and oil saturation predicted by the LightUFTrans model, verifying the reliability and strong generalization ability of the LightUFTrans model in predicting the spatial distribution of CO2 and oil saturation. By integrating SHAP analysis, attention mechanisms, and spatiotemporal verification methods, a systematic interpretability analysis of the LightUFTrans model was conducted from three dimensions: global, temporal, and spatial.