A method for constructing a physical-driven reservoir proxy model based on multi-source data

By using a deep learning network architecture driven by multi-source data and physical constraints, the problems of data utilization efficiency and physical constraints of reservoir surrogate models in complex reservoirs are solved, achieving efficient and accurate prediction of reservoir dynamic response and improving the applicability and reliability of the model.

CN121787285BActive Publication Date: 2026-05-19CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNIV OF PETROLEUM (EAST CHINA)
Filing Date
2026-03-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing reservoir proxy models suffer from insufficient data utilization efficiency and complex physical constraint introduction methods during the construction process, making them difficult to adapt to complex reservoir models. This results in prediction results that do not conform to physical mechanisms and high computational costs.

Method used

A deep learning network architecture driven by multi-source data is adopted, combined with the physical constraints of the reservoir numerical simulator, and a deep learning surrogate model is trained by a hybrid loss function. By fusing temporal and spatial data and introducing a physical driving mechanism, the prediction accuracy and consistency are improved.

Benefits of technology

It enables efficient and accurate dynamic response prediction in complex reservoir models, reduces computational costs, and enhances the physical consistency and engineering reliability of the models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787285B_ABST
    Figure CN121787285B_ABST
Patent Text Reader

Abstract

The application discloses a kind of physical driving reservoir agent model construction method based on multi-source data, it is related to intelligent oil and gas field development technical field, including the following steps: using corner point grid to establish reservoir simulation numerical model;Multi-source data is constructed and coded injection-production time sequence;Disturb injection-production parameters to generate random time sequence injection-production sample;Get training data and construct simulation data and virtual data mixed multi-source data set;Build multi-source data driven deep learning agent model;Set mixed loss function based on virtual data and simulation data calculation;Using mixed data strategy completes model training update.The method improves the prediction accuracy of the agent model by upgrading the deep learning network architecture and considering the time series and spatial reservoir data. Furthermore, the method uses simulated and virtual mixed multi-source data during the agent model training process and introduces a physical driving mechanism, which enables efficient prediction of reservoir dynamic response while enhancing the physical consistency and engineering reliability of the model results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent oil and gas field development technology, and in particular to a method for constructing a physical-driven reservoir proxy model based on multi-source data. Background Technology

[0002] Oil reservoir production is influenced by multiple factors, including geological structure, fluid properties, and injection-production conditions, resulting in dynamic behavior characterized by significant nonlinearity and spatiotemporal coupling. To characterize reservoir response under different development regimes, numerical simulation methods are commonly used in engineering to predict and analyze multiphase flow processes, thereby optimizing production and adjusting strategies. While these methods theoretically offer high accuracy, their computational efficiency relies on complex discrete grids and iterative solutions. In applications with large-scale models or those requiring repeated iterations, the computational efficiency falls short of practical needs. As oil and gas field development moves towards complex structural reservoirs and refined management, the number of faults and heterogeneity in reservoir models are increasing, often necessitating the use of corner grids to describe complex geological structures. Under these conditions, the computational and storage requirements of numerical simulations increase significantly, limiting their application in production optimization, parameter search, and multi-scheme evaluation. Therefore, there is an urgent need for more efficient prediction methods as supplements or alternatives.

[0003] In recent years, data-driven reservoir surrogate models have gradually attracted attention. These methods learn the inherent mapping relationships in historical simulation or production data to achieve rapid prediction of reservoir dynamic responses, thus significantly reducing computational costs. Existing surrogate models are mostly built using deep learning frameworks, which improves prediction efficiency to some extent. However, their performance is highly dependent on the scale of training data, and the models themselves lack constraints on reservoir physics, easily producing prediction results that do not conform to physical mechanisms under unseen conditions, affecting their engineering applicability. To improve these issues, some studies have attempted to introduce physical information into the surrogate model construction process to enhance the rationality of the model's prediction results. However, existing methods usually require explicit construction of seepage control equations or numerical approximation of physical constraints, leading to complex model training processes, high implementation costs, and difficulty in flexibly adapting to different types of reservoir models. Especially when facing actual reservoir systems with complex fault structures and irregular grids, the introduction of physical constraints is often limited by both computational complexity and implementation difficulty. Furthermore, existing technologies still have shortcomings in data utilization methods. Most surrogate models rely on a single type of data for training, failing to fully explore the synergistic relationship between multiple sources of data, such as injection and production parameters, production data, and reservoir state variables. This results in room for improvement in data utilization efficiency and generalization ability.

[0004] Therefore, how to effectively integrate multi-source information and introduce physical driving mechanisms to construct an efficient surrogate model suitable for complex reservoir models while reducing data and computing costs remains a technical problem that needs to be solved in the field of reservoir engineering. Summary of the Invention

[0005] To address the common problems in the construction of existing reservoir surrogate models, such as insufficient data utilization efficiency, complex introduction of physical constraints, and difficulty in adapting to complex reservoir models, this invention discloses a physical-driven reservoir surrogate model construction method based on multi-source data. This method improves the prediction accuracy of the surrogate model by upgrading the deep learning network architecture to comprehensively consider temporal and spatial reservoir data. Furthermore, it introduces a physical-driven mechanism during the training process of the surrogate model to achieve efficient prediction of reservoir dynamic response, while enhancing the physical consistency and engineering reliability of the model results.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for constructing a physical-driven reservoir proxy model based on multi-source data includes the following steps:

[0008] S1. Use corner mesh to establish a numerical model for reservoir simulation;

[0009] S2. Construct multi-source data and encode the injection and acquisition time series;

[0010] S3. Perturb the injection sampling parameters to generate random time-series injection samples;

[0011] S4. Acquire training data and construct a multi-source dataset that mixes simulated and virtual data;

[0012] S5. Build a deep learning proxy model driven by multi-source data;

[0013] S6. Define a hybrid loss function based on virtual and simulated data;

[0014] S7. Use a hybrid data strategy to complete model training and updates.

[0015] Optionally, in step S1, a three-dimensional oil-water two-phase seepage physical model of the reservoir is constructed using a reservoir numerical simulator. The reservoir is discretized using corner grids. The model includes several injection wells and production wells, and faults and well locations are set.

[0016] The governing equations for the physical model of oil-water two-phase flow are:

[0017] ;

[0018] in, Represents the Laplace operator. Indicates phase state, , oil phase, It is an aqueous phase. For density, For source and sink items, Porosity It is saturation and satisfies , For Darcy's velocity, the expression is as follows:

[0019] ;

[0020] in, For the absolute permeability tensor, This refers to relative penetration rate. For viscosity, For depth, It is the acceleration due to gravity. Considering the phase pressure and capillary pressure, the following conditions are met: ;

[0021] The finite volume method is used to spatially discretize the governing equations, and a time implicit scheme is used for time discretization to obtain the discrete residual expression of each grid element at each time step. The solution is then obtained by nonlinear iterative solving using a numerical simulator.

[0022] Optionally, in step S2, to meet the requirements of building the proxy model, the input information is divided into multi-source data and uniformly encoded, specifically including:

[0023] First source: Spatial static geological data, including permeability fields, well location indication information, and injection-production sparse matrix;

[0024] Second source: Time-series injection and production data, setting the control quantities of injection / production wells at each time step as sequence vectors. Use Transformer to process time series injection data separately;

[0025] The third source: spatial dynamic state data, including water saturation and pressure distribution at each time step, is used for supervised and constrained training;

[0026] The fourth source: time-series production data, including the time-series oil production rate and water production rate of each production well at each time step, used for supervisory and constraint training.

[0027] Optionally, in step S4, the injection-production sequence and static spatial data generated in step S3 are input into the reservoir numerical simulator to obtain the corresponding reservoir state and production response. For each sample, the following is extracted:

[0028] Spatial static geological data: static distribution of geological parameters Well location information and injection-production sparse matrix;

[0029] Spatial dynamic state data: pressure field at each time step With water saturation field And form a state vector ;

[0030] Time-series injection and production data: Well operating system control data ;

[0031] Time-series production data: Oil production rate of each production well at each time step. With water production rate and form ;

[0032] Based on this, two types of datasets are constructed:

[0033] Simulated dataset: Supervision labels are provided by the output of the numerical simulator;

[0034] Virtual dataset: It only includes model input and does not include , Labels are used to calculate the physics-driven loss.

[0035] Optionally, in step S5, based on the data organization format of step S2, an MTU-Net deep learning proxy model capable of simultaneously processing spatial multi-source information and time series injection is established, including:

[0036] Encoding module: Performs convolutional encoding on static spatial input and sampled sparse field, downsamples to extract multi-scale spatial features and forms latent variable representation F;

[0037] The temporal module models the temporal dependencies of the feature vector sequence at each time step, employing a Transformer encoder to capture long-range correlations across time steps, and outputting a feature sequence with temporal context. ;

[0038] Decoding module: It fuses temporal features with spatial multi-scale features and upsamples them layer by layer to output the predicted state field and production response prediction values ​​at each time step;

[0039] The model output is: , ,in, and These are the prediction results of the proxy model for state variables and production data, respectively.

[0040] Optionally, step S6 specifically includes:

[0041] S61. Model Initialization and Batch Data Input

[0042] Load the surrogate model from step S5, input the mixed dataset from step S4, and set the training hyperparameters;

[0043] S62. Constructing a hybrid loss training objective

[0044] The hybrid loss consists of both physical-driven loss and data-supervised loss. ;

[0045] in, For physical driving loss, For state variable monitoring loss, To prevent losses in production data monitoring, , These are the weighting coefficients;

[0046] S63. Calculate the physical drive loss

[0047] The predicted state output by the surrogate model at each time step Corresponding injection Input the reservoir numerical simulator and calculate the mass-conserving residual vector within the framework of discrete control equations. And define the physical driving loss as the normalization of the residuals. norm ,in, For the number of grid cells, For time steps;

[0048] To achieve effective backpropagation of the physics-driven loss to the network parameters, the Jacobian matrix generated by the numerical simulator in the current state is preserved. And based on the chain rule, the physical driving loss on network parameters is obtained. gradient: ,in, It is obtained by automatic differentiation from a deep learning agent model, and It is generated naturally or directly output by the numerical simulator when solving the control equations;

[0049] S64. Calculate the monitoring loss and

[0050] For the simulated dataset Normalization between predicted and simulated data is used. Norm constructs the supervised loss, and the calculation method is as follows: , ;

[0051] in, The number of production wells; the gradient of the supervised loss is automatically generated by the deep learning proxy model.

[0052] Optionally, in step S7, the training process uses a virtual dataset. The primary batch is a cyclical batch, while also randomly matching simulated data batches. Stable convergence and improved prediction accuracy for each training iteration:

[0053] (1) Forward calculation yields , ;

[0054] (2) Calculate the residuals using the reservoir numerical simulator based on the virtual dataset. Jacobian matrix ,form and its gradient information;

[0055] (3) Calculate based on simulation data , And obtain the corresponding gradient;

[0056] (4) Combined loss And backpropagate to update network parameters Complete the training of the proxy model.

[0057] The beneficial effects of this invention are that, in the process of constructing existing reservoir surrogate models, there are problems such as insufficient data utilization efficiency, complex introduction of physical constraints, and difficulty in adapting to complex reservoir models. The purpose of this invention is to propose a physical-driven method for constructing complex reservoir surrogate models based on multi-source data. This method improves the prediction accuracy of the surrogate model by upgrading the deep learning network architecture to comprehensively consider temporal and spatial reservoir data. Furthermore, a physical-driven mechanism is introduced during the training process of the surrogate model to achieve efficient prediction of reservoir dynamic response, while enhancing the physical consistency and engineering reliability of the model results.

[0058] This invention fully leverages the mature advantages of reservoir numerical simulators in physical computation, using the physical constraint information generated during the simulation as a crucial basis for model training. This information is then integrated with multi-source information such as injection-production data and production data. This allows the surrogate model to learn mapping relationships that conform to reservoir physical laws without explicitly solving complex governing equations. This approach reduces the surrogate model's dependence on large-scale labeled data and avoids the problems of high implementation difficulty and computational cost associated with traditional physical constraint methods under complex reservoir conditions.

[0059] The method provided by this invention can effectively balance prediction accuracy and computational efficiency, improve the applicability of surrogate models in reservoirs with complex geological structures and irregular grids, and provide a more efficient, robust and easily promoted technical approach for reservoir production optimization and intelligent numerical simulation. Attached Figure Description

[0060] Figure 1 This is a flowchart of a physical-driven reservoir proxy model construction method based on multi-source data according to the present invention.

[0061] Figure 2 Map showing the permeability, well locations, and fault distribution of the SAIGUP reservoir;

[0062] Figure 3 This is a schematic diagram of the multi-source MTU-Net network structure;

[0063] Figure 4A To enable the trained MTU-Net to predict water saturation under random injection and sampling parameters;

[0064] Figure 4B To predict stress under random injection parameters using the trained MTU-Net;

[0065] Figure 5A To predict oil production rates under random injection and production parameters using the trained MTU-Net;

[0066] Figure 5B To predict water production rates under random injection and extraction parameters using the trained MTU-Net;

[0067] Figure 6 Comparison of the spatial distribution of water saturation and pressure prediction errors of PCDD(50) and DD(50) at the final time step;

[0068] Figure 7 Histograms of errors in water saturation, pressure, oil production rate, and water production rate for 50 test samples using different surrogate models;

[0069] Figure 8 Comparison of quality residual distributions for different surrogate models at the final time step;

[0070] Figure 9 This represents the statistical distribution of the residuals due to mass conservation. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0072] A method for constructing a physical-driven reservoir proxy model based on multi-source data, such as Figure 1 As shown, it includes the following steps:

[0073] S1. Use corner mesh to establish a numerical model for reservoir simulation.

[0074] A three-dimensional two-phase oil-water flow physical model of a complex reservoir is constructed using a reservoir numerical simulator. The reservoir is discretized using corner grids, which can represent irregular boundaries and fault structures. In this embodiment, the reservoir grid size can be 40×120×20 (or other actual engineering scale), and interface transfer capabilities (conduction / closure) are set at fault locations to reflect the impact of faults on fluid exchange. Figure 2 The model sequentially displays permeability, well location distribution, and permeability distribution near the fault, as well as enlarged views of the local fault and the sealed fault plane. The model includes several injection wells and production wells. The injection and production methods can be injection rate control for injection wells, bottom hole pressure control for production wells (or both types of wells can use bottom hole pressure control), and includes settings for fault and well location distribution, total simulation duration, and injection / production update cycle.

[0075] The governing equations for the physical model of oil-water two-phase flow are:

[0076] ;

[0077] in, Represents the Laplace operator. Indicates phase state, , oil phase, It is an aqueous phase. For density, For source and sink items, Porosity It is saturation and satisfies , For Darcy's velocity, the expression is as follows:

[0078] ;

[0079] in, For the absolute permeability tensor, This refers to relative penetration rate. For viscosity, For depth, It is the acceleration due to gravity. Considering the phase pressure and capillary pressure, the following conditions are met: ;

[0080] The governing equations are spatially discretized using the finite volume method, and temporally discretized using a temporal implicit scheme to obtain the element mesh. The discrete mass conservation equation is:

[0081] ;

[0082] in, for unit Phase quality residuals for Cell grid volume, For a unit discrete time step, For discrete time steps, For unit Index of adjacent cells, The total number of all adjacent units. The out-of-plane normal vector of the element. for unit Phase production / injection volume;

[0083] After obtaining the discrete residual representation of each grid cell at each time step, a nonlinear iterative solution is performed using a numerical simulator. Represents the reservoir state variables, starting from the initial state. Start, iteration step The solution process is as follows:

[0084] ;

[0085] in, express The Jacobian matrix of the iteration step, express The state variables of the iteration step Indicates reservoir condition The corresponding set of quality residuals;

[0086] The reservoir state variable update process is as follows:

[0087]

[0088] in, for Damping coefficient of the iteration step;

[0089] Numerical simulation iterations, the norm of the quality residual converges to the set minimum iteration error. , represented as:

[0090] ;

[0091] This completes the entire discrete difference iterative solution process.

[0092] S2. Construct multi-source data and encode the injection and acquisition time series.

[0093] To address the requirements for building a proxy model, the input information is divided into multi-source data and uniformly encoded, specifically including:

[0094] First source: Spatial static geological data, including permeability fields, well location indication information, and injection-production sparse matrix;

[0095] Second source: Time-series injection and production data, setting the control quantities of injection / production wells at each time step as sequence vectors. Use Transformer to process time series injection data separately;

[0096] The third source: spatial dynamic state data, including water saturation and pressure distribution at each time step, is used for supervised and constrained training;

[0097] The fourth source: time-series production data, including the time-series oil production rate and water production rate of each production well at each time step, used for supervisory and constraint training.

[0098] S3. Perturb the injection sampling parameters to generate random time-series injection samples.

[0099] Taking a production optimization scenario as an example, the injection and sampling sequence is randomly perturbed within a set control period to obtain different combinations of injection and sampling parameters. In the example embodiment, the total simulation period is 5 years, the control update period is 1 year, and the state prediction is output at a six-month interval, thus the prediction time steps are... =10, which can be set according to the specific problem.

[0100] The disturbance range can be set according to engineering constraints, and the injection rate range of the injection well can be set to [100m]. 3 / day,300m 3 / day], the bottom hole pressure range of production wells [23Mpa, 28Mpa] or other corresponding upper and lower limits of reservoir pressure levels). A sample set is constructed through random sampling or Latin hypercube sampling to form training scenarios covering different control strategies.

[0101] S4. Obtain training data and construct a multi-source dataset that mixes simulated and virtual data.

[0102] Input the injection-production sequence and static spatial data generated in step S3 into the reservoir numerical simulator, run it to obtain the corresponding reservoir state and production response, and extract the following for each sample:

[0103] Spatial static geological data: static distribution of geological parameters Well location information and injection-production sparse matrix;

[0104] Spatial dynamic state data: pressure field at each time step With water saturation field And form a state vector ;

[0105] Time-series injection and production data: Well operating system control data ;

[0106] Time-series production data: Oil production rate of each production well at each time step. With water production rate and form ;

[0107] Based on this, two types of datasets are constructed:

[0108] Simulated dataset: Supervision labels are provided by the output of the numerical simulator;

[0109] Virtual dataset: It only includes model input and does not include , Labels are used to calculate the physics-driven loss.

[0110] S5. Build a deep learning proxy model driven by multi-source data.

[0111] Based on the data organization format in step S2, a Transformer-U-Net (MTU-Net) deep learning proxy model capable of simultaneously processing spatial multi-source information and time-series injection is established, such as... Figure 3 As shown, it includes:

[0112] Encoding module: Performs convolutional encoding on static spatial inputs (permeability field, well location / fault information, etc.) and injection-production sparse field, downsamples to extract multi-scale spatial features and forms a latent variable representation F;

[0113] The temporal module models the temporal dependencies of the feature vector sequence at each time step, employing a Transformer encoder to capture long-range correlations across time steps, and outputting a feature sequence with temporal context. ;

[0114] Decoding module: It fuses temporal features with spatial multi-scale features and upsamples them layer by layer to output the predicted state field and production response prediction values ​​at each time step;

[0115] The model output is: , ,in, and These are the prediction results of the proxy model for state variables and production data, respectively.

[0116] S6. Define a hybrid loss function based on virtual and simulated data; specifically including:

[0117] S61. Model Initialization and Batch Data Input

[0118] Load the surrogate model from step S5, input the mixed dataset from step S4, and set the training hyperparameters (learning rate, batch size, number of iterations, and loss weights, etc.).

[0119] S62. Constructing a hybrid loss training objective

[0120] The training loss consists of both the physics-driven loss and the data-supervised loss. ;

[0121] in, For physical driving loss, For state variable monitoring loss, To prevent losses in production data monitoring, , These are the weighting coefficients;

[0122] S63. Calculate the physical drive loss

[0123] The predicted state output by the surrogate model at each time step With corresponding injection and extraction parameters Input the reservoir numerical simulator and calculate the mass-conserving residual vector within the framework of discrete control equations. And define the physical driving loss as the normalization of the residuals. norm ,in, For the number of grid cells, For time steps;

[0124] To achieve effective backpropagation of the physics-driven loss to the network parameters, the Jacobian matrix generated by the numerical simulator in the current state is preserved. And based on the chain rule, the physical driving loss on network parameters is obtained. gradient: ,in, It is obtained by automatic differentiation from a deep learning agent model, and The process is generated naturally or directly by the numerical simulator when solving the control equations, thus it has good engineering feasibility and simulator portability.

[0125] S64. Calculate the monitoring loss and

[0126] For the simulated dataset Normalization between predicted and simulated data is used. The norm is used to construct the supervised loss, and the calculation process is as follows: , ;

[0127] in, The number of production wells; the gradient of the supervised loss is automatically generated by the deep learning proxy model.

[0128] S7. Use a hybrid data strategy to complete model training and updates.

[0129] The training process uses a virtual dataset The primary batch is a cyclical batch, while also randomly matching simulated data batches. Stable convergence and improved prediction accuracy for each training iteration:

[0130] (1) Forward calculation yields , ;

[0131] (2) Calculate the residuals using the reservoir numerical simulator based on the virtual dataset. Jacobian matrix ,form and its gradient information;

[0132] (3) Calculate based on simulation data , And obtain the corresponding gradient;

[0133] (4) Combined loss And backpropagate to update network parameters Complete the training of the proxy model.

[0134] Figure 4A , Figure 4B , Figure 5A , Figure 5B The training results of MTU-Net on water saturation, pressure, and yield data are shown. The training process in this embodiment uses the proposed method to train a surrogate model PCDD(50) with only 50 sets of data and simulated data under constraints. As shown in the figure, MTU-Net can accurately predict the distribution of water saturation, pressure, and yield curves under random injection-production parameters. The distributions of pressure and saturation errors are well controlled, and the yield curve is almost perfectly fitted, making it an accurate surrogate model for numerical simulation.

[0135] Figure 6 The figure shows the predictions of the surrogate model PCDD(50), trained using 50 sets of virtual and simulated data, and the surrogate model DD(50), trained directly using 50 sets of simulated data, for saturation distribution and pressure at the last time step. As can be seen from the figure, for the prediction error of water saturation, DD(50) is greater than 0.04, while PCDD(50) is less than 0.02; for the prediction error of pressure, DD(50) is greater than 4 MPa, while PCDD(50) is less than 2 MPa. Therefore, compared to DD(50), PCDD(50) exhibits significantly superior prediction performance.

[0136] Figure 7 The figure shows the prediction statistics of PCDD(50), DD(50), and MTU-Net (built using only 300 sets of simulated data) for water saturation, pressure, oil production rate, and water production rate of 50 test samples. As can be seen from the figure, DD(300) and PCDD(50) significantly outperform DD(50) in all four prediction tasks, indicating that both increasing the scale of simulated data and introducing physical constraints on virtual data can effectively improve model accuracy. More importantly, PCDD(50) has achieved a prediction level close to that of DD(300) using only 50 sets of simulated data. Taking water saturation as an example, the average errors of DD(300) and PCDD(50) are 0.0015 and 0.0017, respectively, representing a nearly 50% reduction compared to the 0.0030 error of DD(50). Pressure and production forecasts also showed consistent results: compared with DD(50), the average errors of pressure, oil production rate and water production rate of PCDD(50) decreased by 42.2%, 40.3% and 48.0% respectively, with an overall average error reduction of 45.1%.

[0137] Figure 8The distribution of mass residuals calculated based on the predicted pressure and saturation fields at the final time step of the same test sample is shown. It can be observed that the residuals of PCDD(50) are significantly smaller overall: most regions have residuals below 0.01, with only a few fault neighborhoods showing local residuals of approximately 0.03. In contrast, both DD(300) and DD(50) show regions with large area residuals exceeding 0.04, indicating that even though DD(300) is slightly better than PCDD(50) in terms of data error, its prediction results may still exhibit more significant inconsistencies in the sense of mass conservation.

[0138] Figure 9 The statistical distribution of the quality residuals of the test samples is shown in the figure. As can be seen from the figure, PCDD(50) exhibits a clear advantage in terms of quality residuals: the average residual of PCDD(50) is 0.0020, which is about 23% lower than that of DD(300) (0.0026) and about 78% lower than that of DD(50) (0.0091). This indicates that physical constraints can continuously correct the model output during the training process, making it more strictly satisfy the control equations followed by the reservoir simulation process, thereby significantly improving the physical consistency and interpretability of the prediction.

[0139] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for constructing a physical-driven reservoir proxy model based on multi-source data, characterized in that, Includes the following steps: S1. Use corner mesh to establish a numerical model for reservoir simulation; S2. Construct multi-source data and encode the injection and acquisition time series; S3. Perturb the injection sampling parameters to generate random time-series injection samples; S4. Obtain training data and construct a multi-source dataset that mixes simulated and virtual data; S5. Build a deep learning proxy model driven by multi-source data; S6. Define a hybrid loss function based on virtual and simulated data; S7. Employ a hybrid data strategy to complete model training and updates; In step S2, to meet the requirements of the agent model construction, the input and output information are divided into multi-source data and uniformly encoded, specifically including: First source: Spatial static geological data, including permeability fields, well location indication information, and injection-production sparse matrix; Second source: Time-series injection and production data, setting the control quantities of the injection / production wells at each time step as injection and production parameters. Use Transformer to process time series injection data separately; The third source: spatial dynamic state data, including water saturation and pressure distribution at each time step, is used for supervised and constrained training; The fourth source: time-series production data, including the time-series oil production rate and water production rate of each production well at each time step, used for supervision and constraint training; In step S4, the injection-production sequence and static spatial data generated in step S3 are input into the reservoir numerical simulator to obtain the corresponding reservoir state and production response. For each sample, the following is extracted: Spatial static geological data: Absolute permeability tensor Well location information and injection-production sparse matrix; Spatial dynamic state data: pressure field at each time step With water saturation field And form a state vector ; Time-series injection and acquisition data: injection and acquisition parameters ; Time-series production data: Oil production rate of each production well at each time step. With water production rate and form ; Based on this, two types of datasets are constructed: Simulated dataset: Supervision labels are provided by the output of the numerical simulator; Virtual dataset: It only includes model input and does not include , Labels are used to calculate the physics-driven loss; Step S6 specifically includes: S61. Model initialization and batch data input; Load the surrogate model from step S5, input the mixed dataset from step S4, and set the training hyperparameters; S62. Construct a hybrid loss training objective; The hybrid loss consists of both physical-driven loss and data-supervised loss. ; in, For physical driving loss, For state variable monitoring loss, To prevent losses in production data monitoring, , These are the weighting coefficients; S63. Calculate the physical drive loss ; The predicted state output by the surrogate model at each time step With corresponding injection and extraction parameters Input the reservoir numerical simulator and calculate the mass-conserving residual vector within the framework of discrete control equations. And define the physical driving loss as the normalization of the residuals. norm ,in, For the number of grid cells, For time steps; To achieve effective backpropagation of the physics-driven loss to the network parameters, the Jacobian matrix generated by the numerical simulator in the current state is preserved. And based on the chain rule, the physical driving loss on network parameters is obtained. gradient: ,in, It is obtained by automatic differentiation from a deep learning agent model, and It is generated naturally or directly output by the numerical simulator when solving the control equations; S64. Calculate the monitoring loss and ; For the simulated dataset Normalization between predicted and simulated data is used. Norm constructs the supervised loss, and the calculation method is as follows: , ; in, The number of production wells; the gradient of the supervised loss is automatically generated by the deep learning proxy model; In step S7, the training process uses a virtual dataset. The primary batch is a cyclical batch, while also randomly matching simulated data batches. Stable convergence and improved prediction accuracy for each training iteration: (1) Forward calculation yields , ; (2) Calculate the residuals using the reservoir numerical simulator based on the virtual dataset. Jacobian matrix ,form and its gradient information; (3) Calculate based on simulation data , And obtain the corresponding gradient; (4) Combined loss And backpropagate to update network parameters Complete the training of the proxy model.

2. The method for constructing a physical-driven reservoir proxy model based on multi-source data as described in claim 1, characterized in that, In step S1, a three-dimensional oil-water two-phase seepage physical model of the reservoir is constructed using a reservoir numerical simulator. The reservoir is discretized using corner grids. The model includes several injection wells and production wells, and faults and well locations are set. The governing equations for the physical model of oil-water two-phase flow are: ; in, Represents the Laplace operator. Indicates phase state, , oil phase, It is an aqueous phase. For density, For source and sink items, Porosity It is saturation and satisfies , For Darcy's velocity, the expression is as follows: ; in, For the absolute permeability tensor, This refers to relative penetration rate. For viscosity, For depth, It is the acceleration due to gravity. Considering the phase pressure and capillary pressure, the following conditions are met: ; The finite volume method is used to spatially discretize the governing equations, and a time implicit scheme is used for time discretization to obtain the discrete residual expression of each grid element at each time step. The solution is then obtained by nonlinear iterative solving using a numerical simulator.

3. The method for constructing a physical-driven reservoir proxy model based on multi-source data as described in claim 1, characterized in that, In step S5, based on the data organization format of step S2, an MTU-Net deep learning proxy model capable of simultaneously processing spatial and temporal series multi-source information is established, including: Encoding module: Performs convolutional encoding on static spatial input and sampled sparse field, downsamples to extract multi-scale spatial features and forms latent variable representation F; The temporal module models the temporal dependencies of the feature vector sequence at each time step, employing a Transformer encoder to capture long-range correlations across time steps, and outputting a feature sequence with temporal context. ; Decoding module: It fuses temporal features with spatial multi-scale features and upsamples them layer by layer to output the predicted state field and production response prediction values ​​at each time step; The model output is: , ,in, and These are the prediction results of the proxy model for state variables and production data, respectively.