A Smart Optimization Method for Multi-Type Well Fracture Joint Control and Fine Injection-Production Mode
By combining well-fracture joint control with precision injection and production mode, and integrating deep learning and reinforcement learning, we have optimized injection and production schemes for various types of well-fracture joint control. This has solved the problem that traditional methods have difficulty in directionally spreading injected fluids in heterogeneous reservoirs, and has enabled efficient optimization and real-time adjustment of injection and production schemes, thereby improving reservoir development efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA UNIV OF PETROLEUM (EAST CHINA)
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional macro-injection-production well networks struggle to directionally spread injected fluids in heterogeneous reservoirs, resulting in untapped residual oil between well sections and fractures and poor water sweep efficiency. Existing injection-production scheme optimization methods are computationally time-consuming and inefficient, making it difficult to support rapid optimization and decision-making in scenarios involving multiple types of well-fracture joint control.
A fine-grained injection-production model with well-fracture joint control was adopted. Combined with reservoir physical parameters, a reservoir numerical simulation model was established. Sample data was obtained using Latin hypercube sampling. A deep learning surrogate model and a reinforcement learning dynamic decision model were established. The optimal injection-production development scheme was generated through rapid optimization using LSTM network and particle swarm optimization algorithm.
It enables rapid optimization and decision support under a new multi-type well-fracture joint control model, improves the efficiency of injection and production schemes, can cope with real-time adjustments under complex working conditions, ensures that the injection and production system is optimal globally on a macro level and has the ability to adjust locally, and enhances reservoir development results.
Smart Images

Figure CN121675830B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas field development technology, specifically to an intelligent optimization method for a multi-type well-fracture joint control precision injection and production mode. Background Technology
[0002] Unconventional oil and gas resources are abundant and have great development potential, serving as an important replacement force for ensuring oil and gas energy security. However, they are limited by poor reservoir properties and low energy, requiring large-scale fracturing to improve oil and gas seepage conditions and achieve effective development of unconventional oil and gas.
[0003] Considering the significantly enhanced heterogeneity of the reservoir matrix and fractures after compression, the traditional macro-injection-production well network has poor global adaptability, making it difficult for injected fluids to achieve directional sweeping. This results in a large amount of residual oil between wells and fracture sections that is difficult to utilize, and poor sweeping efficiency of injected water, which greatly limits the post-compression production enhancement effect of the reservoir. The well-fracture joint control method achieves directional guidance and fine control of the displacement flow field in heterogeneous reservoirs by independently controlling the injection and production properties of each fracture section in horizontal wells. This allows for precise utilization of residual oil between wells and fracture sections, improving sweeping efficiency.
[0004] Asynchronous injection and production schemes within the same well section and staggered injection and production schemes between different well sections are two new modes of fine injection and production with well-fracture joint control. They can effectively expand the sweep range of injected fluid. However, their actual application effect depends on the design of a reasonable injection and production scheme. However, the traditional numerical simulation method for optimizing the injection and production scheme has limitations such as long computation time and low efficiency. The computational cost of screening the optimal injection and production scheme through massive simulations is huge. Moreover, the current optimization method for injection and production schemes in well-fracture joint control scenarios is not yet mature.
[0005] Therefore, there is an urgent need to propose an efficient, accurate, and adaptive intelligent optimization method for multi-type well-fracture joint control fine injection and production modes, so as to achieve rapid optimization and decision support for injection and production schemes under the new multi-type well-fracture joint control mode. Summary of the Invention
[0006] This invention aims to solve the above-mentioned problems and proposes an intelligent optimization method for a multi-type well-fracture joint control fine injection and production mode. By designing a well-fracture joint control fine injection and production mode for the reservoir, an intelligent algorithm coupled with single-objective pre-search and reinforcement learning is established to refine and dynamically control the benchmark injection and production system of the reservoir, obtain the optimal injection and production development scheme of the reservoir, realize the rapid optimization and decision support of reservoir injection and production scheme under the new multi-type well-fracture joint control mode, and provide technical support for the development strategy of actual reservoirs.
[0007] The present invention adopts the following technical solution:
[0008] A method for intelligent optimization of multi-type well fracture joint control precision injection and production mode includes the following steps:
[0009] Step 1: Set up a well-fracture joint control fine injection and production mode. Combine reservoir physical property parameters and horizontal well design parameters to establish a reservoir numerical simulation model in the reservoir numerical simulation software to simulate the injection and production process of the fractured horizontal well.
[0010] Step 2: Set the core sampling parameters and their ranges. Based on the Latin hypercube sampling method, obtain multiple reservoir injection and production schemes. Use the reservoir numerical simulation model to simulate each reservoir injection and production scheme to obtain the state, action and reward of the reservoir numerical simulation model at each time step when each reservoir injection and production scheme is adopted. Generate multiple sample data and establish a sample database.
[0011] Step 3: Build a deep learning agent model based on the LSTM network structure, and train and validate the deep learning agent model using a sample database to obtain the validated deep learning agent model.
[0012] Step 4: Replace the reservoir numerical simulation model with the validated deep learning surrogate model, and use the particle swarm optimization algorithm for single-objective pre-search global optimization. Determine the optimal benchmark strategy through fast search to obtain the benchmark injection-production scheme.
[0013] Step 5: Establish a reinforcement learning dynamic decision-making model based on the Actor-Critic dual network structure of policy gradient, train the reinforcement learning dynamic decision-making model based on the PPO proximal policy optimization algorithm, and obtain a reinforcement learning agent for fine-tuning the benchmark injection and extraction system scheme.
[0014] Step 6: Obtain the initial formation state, total development time, and time step of the actual reservoir. Input the benchmark injection-production regime scheme obtained in Step 4 into the reinforcement learning agent. Use the reinforcement learning agent to obtain the injection flow rate and production pressure of all fractures in the actual reservoir at all time steps within the total development time. Generate an intelligent injection-production regime table containing the injection flow rate sequence and production pressure sequence to obtain the optimal injection-production development scheme for the actual reservoir.
[0015] Preferably, in step 1, the well-fracture joint control precision injection and production mode includes an asynchronous injection and production scheme between the same well section and an interleaved injection and production scheme between different well sections. The asynchronous injection and production scheme between the same well section is designed for a single fractured horizontal well, and the injection and production development of the fractured horizontal well is carried out by alternately arranging production fractures and injection fractures in the horizontal well section of the single fractured horizontal well. The interleaved injection and production scheme between different well sections is designed for multiple fractured horizontal wells, and the injection and production development of the fractured horizontal well is carried out by alternately setting fractures between the same well sections of two adjacent fractured horizontal wells.
[0016] Preferably, step 2 includes the following sub-steps:
[0017] Step 2.1: Based on the actual working conditions of the reservoir, set the total development time and initial formation pressure of the reservoir numerical simulation model, and set the core sampling parameters and their value ranges. The core sampling parameters include the mean and fluctuation range of the injected fracture flow rate, the mean and fluctuation range of the produced fracture pressure, and the adjustment frequency.
[0018] Step 2.2: Based on the Latin hypercube sampling method, hierarchical sampling is performed in the parameter space of each core sampling parameter to obtain multiple sets of non-overlapping core sampling parameters. According to the preset threshold judgment rule, the collected core sampling parameters are converted into injection and sampling parameters to obtain multiple injection and sampling schemes.
[0019] Step 2.3: Based on each injection-production scheme and the preset time step, simulate the reservoir using the reservoir numerical simulation model in the reservoir numerical simulator to obtain the reservoir state, actions, and rewards obtained from the reservoir numerical simulation model at each time step, and then... The state of the reservoir at each time step ,action and rewards and the The state of the reservoir at each time step As sample data, the state is set as the average pressure and water saturation around the fracture, the action is set as the average injection volume and average production pressure of the fracture, and the reward is set as the step-by-step net present value.
[0020] Step 2.4: Preprocess the sample data obtained from the simulation. First, remove the sample data that failed the simulation and contained outliers. Then, use the Z-Score standardization method to normalize the sample data. Randomly allocate the normalized sample data to the training set and the validation set to construct the sample database.
[0021] Preferably, step 3 includes the following sub-steps:
[0022] Step 3.1: Establish a deep learning agent model based on the LSTM network structure, and set the maximum number of iterations and loss function of the deep learning agent model;
[0023] Step 3.2: Randomly select multiple sample data from the training set and input them into the deep learning proxy model. Train the deep learning proxy model to predict the production index of the reservoir numerical simulation model at the current moment and the state of the reservoir numerical simulation model at the next moment based on the state and action of the reservoir numerical simulation model in the sample data. Calculate the loss function value of the deep learning proxy model during the training process.
[0024] Step 3.3: If the current iteration count has reached the maximum number of iterations, stop training the deep learning proxy model and proceed to step 3.4. Otherwise, compare the calculated loss function value with the preset loss function value. If the calculated loss function value exceeds the preset loss function value, adjust the hyperparameters of the deep learning proxy model and return to step 3.2 to continue training the deep learning proxy model. If the calculated loss function value does not exceed the preset loss function value, stop training the deep learning proxy model and proceed to step 3.4.
[0025] Step 3.4: Randomly select multiple sample data from the test set and input them into the trained deep learning agent model. Evaluate the performance of the trained deep learning agent model based on the coefficient of determination and the mean absolute percentage error. If the coefficient of determination of the trained deep learning agent model is greater than the preset coefficient of determination value and the mean absolute percentage error is less than the preset error value, proceed to step 3.5. Otherwise, adjust the hyperparameters of the deep learning agent model and return to step 3.2 to continue training the deep learning agent model.
[0026] Step 3.5: Output the validated deep learning agent model.
[0027] Preferably, the deep learning proxy model is built based on an LSTM network structure, including an input layer, a hidden layer, a fully connected layer, and an output layer. The input layer is used to input the state and actions of the reservoir numerical simulation model at the current time step. The hidden layer uses multiple stacked LSTM units to learn the pressure propagation characteristics and saturation front advancement characteristics during reservoir seepage through a gating mechanism. The fully connected layer is set between the hidden layer and the output layer to map the high-dimensional feature data output by the hidden layer to the physical space. The output layer is used to output the predicted state of the reservoir numerical simulation model at the next time step and the production indicators at the current time step.
[0028] The mapping function of the deep learning agent model is:
[0029] ;
[0030] In the formula, For the first Predicted values of reservoir state at each time step; , , The first Predicted values, status, and actions of reservoir production indicators at each time step; LSTM network weight parameters The proxy model mapping function; The hidden state memory passed to the LSTM network;
[0031] The loss function of the deep learning agent model is:
[0032] ;
[0033] In the formula, For the loss function of the deep learning agent model; These are the weight parameters of the LSTM network; The sample data sequence number; The total number of sample data points used in training; , These are all weighting coefficients of the loss function; , The first The reservoir status and production indicators at each time step; For the first Predicted values of reservoir production indicators at each time step; It is an L2 norm.
[0034] Preferably, step 4 includes the following sub-steps:
[0035] Step 4.1 sets maximizing the cumulative net present value over the entire cycle as the objective function of the pre-search process, expressed as:
[0036] ;
[0037] in,
[0038] ;
[0039] In the formula, It is a function for maximizing the value; The objective function is... This refers to the time step number; This represents the total number of time steps. For the first The stepwise net present value of the reservoir at each time step; , , These are crude oil price, water injection cost, and produced water treatment cost, respectively. , , The first The cumulative oil production, cumulative water injection, and water production of the reservoir at each time step; The annual discount rate; For time step;
[0040] Using production regime parameters, including injection volume and production pressure, as optimization variables, the deep learning surrogate model is set with the following parameters: The crack in the first The actions at each time step are:
[0041] ;
[0042] In the formula, The crack number; For the first The crack in the first Injected traffic at each time step; For the first The crack in the first The extraction pressure at each time step; For the first The average injection flow rate of the crack; For the first Average bottom hole flowing pressure of the fracture; For the first The amplitude of the crack fluctuation; It is a sine function;
[0043] Step 4.2: Set the initial parameters of the particle swarm optimization algorithm, including the total number of particles, the maximum number of iterations, the inertia weight, and the learning factor;
[0044] Step 4.3: Using the particle swarm to represent the injection and collection regime scheme, the position and velocity of each particle in the particle swarm are randomly set. The velocity of the particle is the rate of change of the particle position, which is used to represent the magnitude of the injection and collection parameter adjustment. The position of the particle is the injection and collection regime parameter.
[0045] For each particle in the particle swarm, the injection-extraction scheme represented by the particle is input into the validated deep learning proxy model for prediction, so as to obtain the production indicators and cumulative net present value corresponding to the injection-extraction scheme, and use them as the fitness value of the particle.
[0046] Step 4.4: Update the optimal value of each particle, making the particle with the highest fitness the global optimal particle, and update the velocity and position of each particle in the particle swarm; Step 4.5: Determine whether the current iteration count has reached the preset maximum iteration count. If it has, proceed to Step 4.6; otherwise, update the iteration count and return to Step 4.4 to continue particle swarm optimization.
[0047] Step 4.6: Output the injection and extraction regime scheme corresponding to the globally optimal particle as the preferred benchmark strategy to obtain the benchmark injection and extraction regime scheme.
[0048] Preferably, step 5 includes the following sub-steps:
[0049] Step 5.1: Establish a reinforcement learning dynamic decision-making model based on the Actor-Critic dual network structure of policy gradient, and connect the reinforcement learning dynamic decision-making model and the validated deep learning agent model based on the Markov decision process. Use the reinforcement learning dynamic decision-making model to predict the cumulative net present value of the reservoir within a specified time in the future based on the current injection and production scheme.
[0050] Step 5.2: Train the reinforcement learning dynamic decision-making model using the PPO (Proximity Policy Optimization) algorithm, and update the network parameters of the reinforcement learning dynamic decision-making model.
[0051] Step 5.3: Save the network parameters of the reinforcement learning dynamic decision-making model to obtain the reinforcement learning agent.
[0052] Preferably, the reinforcement learning dynamic decision-making model is constructed based on an Actor-Critic dual-network structure with policy gradients, including a policy network and a value network. The policy network is used to adjust the injection-production regime, including an input layer, a fully connected hidden layer, and an output layer connected in sequence. It is used to determine the normal distribution function of the adjustment coefficient based on the current state of the reservoir, and the adjustment coefficient is determined by random sampling. The value network is used to evaluate the merits of the injection-production regime adjustment, including an input layer, a fully connected hidden layer, and an output layer connected in sequence. It is used to determine the state value function based on the current state of the reservoir, and the value function is calculated to guide the policy network update.
[0053] The state-value function is used to predict the cumulative net present value of the reservoir over a specified future period based on the current reservoir state and the injection-production scheme. The expression is:
[0054] ;
[0055] In the formula, The state value function; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step; Expected based on experience; Discount factor; For the first The reward for each time step of the reservoir;
[0056] The Markov Decision Process (MDP) is used for parameter passing in reinforcement learning dynamic decision-making models and deep learning agent models, by establishing a continuous state-action MDP quintuple. ,in, This represents the state space, which serves as the input to the reinforcement learning dynamic decision-making model. This is the action space, used to represent adjustments to the injection and extraction system; This is a state transition model, responsible for receiving the current state of the reservoir and the adjusted actions, and calculating the state of the reservoir in the next time step. For the reward function; This is the discount factor.
[0057] Preferably, step 5.2 includes the following sub-steps:
[0058] Step 5.2.1: Initialize the PPO proximal policy optimization algorithm parameters, setting the learning rate, truncation parameter, discount factor, generalized advantage estimation parameter, maximum number of training epochs, and composite total loss function;
[0059] Step 5.2.2: During iterative training, the policy network of the reinforcement learning dynamic decision-making model is first used to identify the current state of the reservoir and output adjustment coefficients. The adjusted injection-production scheme is then input into the validated deep learning agent model. The validated deep learning agent model is used to obtain the reward for the current time step and the reservoir state for the next time step, acquiring a reservoir database composed of reservoir data from all time steps. The data in the reservoir database are then input into the reinforcement learning dynamic decision-making model. The value network of the reinforcement learning dynamic decision-making model outputs the state value function for each time step, determining the reward and advantage function of the reservoir at each time step. The calculation formula is as follows:
[0060] ;
[0061] ;
[0062] In the formula, For the first The return on the reservoir at each time step; For the first The reward for each time step of the reservoir; Discount factor; For the first The return on the reservoir at each time step; For the first The dominant function of a reservoir at each time step; For summation index; This represents the total number of time steps. This refers to the time step number; For parameters to estimate the generalized dominance; For the first The reward for each time step of the reservoir; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step;
[0063] Step 5.2.3: Based on the gradient optimization algorithm, calculate the gradient of the composite total loss function value relative to the policy network weight parameters and the value network weight parameters, and update the network parameters of the reinforcement learning dynamic decision model by combining the preset learning rate.
[0064] Step 5.2.4: Continue iterative calculations until the current number of iterations reaches the preset maximum number of training rounds or the preset reward convergence condition is met, then end the training of the reinforcement learning dynamic decision model.
[0065] Preferably, the composite total loss function is set as follows:
[0066] ;
[0067] in,
[0068] ;
[0069] ;
[0070] ;
[0071] In the formula, This is the composite total loss function; To maximize the strategy's reward function; The policy entropy function; , All are weighting coefficients of the composite total loss function; These are the policy network weight parameters; These are the weight parameters of the value network; The current injection / collection strategy is determined by the strategy network weight parameters; The predicted value is the expected value based on experience; It is a minimum value function; The new and old injection and extraction system schemes were proposed in the 1st year. Select an action within a time step The probability ratio; The dominant function; This is a truncation function; To truncate parameters; The value network loss function; This is an estimate of the state value; For the first The state of the reservoir at each time step; For the first The return on the reservoir at each time step; The crack number; To optimize the total number of cracks; It is a natural constant; For the first The standard deviation of the crack action; It is a logarithmic function with base 10.
[0072] The present invention has the following beneficial effects:
[0073] (1) This invention proposes an intelligent optimization method for fine injection and production mode with multi-type well fracture joint control. Based on the Latin hypercube sampling and physical constraint mapping method, high-quality sample data is obtained, a sample database is constructed, and the sample data in the sample database is used to train and verify a deep learning proxy model based on LSTM network structure. The deep learning proxy model is used to replace the reservoir numerical simulation model to obtain simulation results, which effectively solves the problem that traditional numerical simulation methods are time-consuming and difficult to support high-frequency iterative optimization.
[0074] Meanwhile, the method of this invention introduces physical boundary constraints to ensure the generalization ability of the data-driven model to complex working conditions. It uses the unique gating mechanism of the LSTM network structure to capture the temporal characteristics of reservoir pressure transmission and material balance, reducing the time cost of simulating a single injection and production scheme. This provides an efficient and high-fidelity environment for realizing large-scale, long-cycle intelligent injection and production optimization, and significantly improves the optimization efficiency of injection and production schemes.
[0075] (2) This invention proposes an intelligent optimization method for a multi-type fracture joint control fine injection-production mode, and establishes an intelligent optimization algorithm framework coupled with single-objective pre-search and reinforcement learning. Addressing the strong heterogeneity and dynamic interference problems existing in multiple fractures and between multiple wells, it solves the difficulties of single heuristic algorithms in handling complex time-varying reservoir conditions, and the challenges of slow startup and training convergence when directly applying reinforcement learning for reservoir injection-production system optimization. This method rapidly searches for the optimal benchmark strategy globally through single-objective pre-search, providing high-quality initial values for the reinforcement learning dynamic decision model. By introducing a residual control mechanism into the reinforcement learning of the dynamic decision model, the optimal benchmark strategy is dynamically fine-tuned based on the real-time formation pressure and water saturation state, resulting in the optimal injection-production development scheme for the reservoir.
[0076] The method of this invention not only ensures the global optimization of the injection and production system from a macroscopic perspective, but also has the ability to adjust in real time to deal with sudden reservoir conditions such as local water channeling and pressure exhaustion. It enables rapid optimization and decision support of injection and production schemes under the new model of multi-type well and fracture joint control, effectively improving the optimization efficiency of reservoir injection and production development schemes. Attached Figure Description
[0077] Figure 1 This is a flowchart of an intelligent optimization method for a multi-type well fracture joint control fine injection and production mode according to the present invention.
[0078] Figure 2 This is a schematic diagram of the asynchronous injection and production scheme between the same well sections according to the present invention.
[0079] Figure 3 This is a schematic diagram of the staggered injection and production scheme between different well sections according to the present invention.
[0080] Figure 4 This is the relative permeability curve of oil and water.
[0081] Figure 5 This is a table showing the dynamic injection and production work schedule for the asynchronous injection and production scheme between the same well sections in this invention.
[0082] Figure 6 The figure shows the reservoir parameter variation curves based on the asynchronous injection and production scheme between the same well sections. In the figure, (a) is the curve of cumulative oil production of the reservoir over time, (b) is the curve of cumulative water production of the reservoir over time, (c) is the curve of water cut of the reservoir over time, and (d) is the curve of net present value of the reservoir over time.
[0083] Figure 7 The table shows the dynamic injection and production work schedule for the staggered injection and production scheme between different well sections of the present invention. In the figure, (a) is the dynamic injection and production work schedule for the first horizontal well in the staggered injection and production scheme between different well sections, and (b) is the dynamic injection and production work schedule for the second horizontal well in the staggered injection and production scheme between different well sections.
[0084] Figure 8 The figure shows the reservoir parameter variation curves based on the staggered injection and production scheme between different well sections. In the figure, (a) is the curve of cumulative oil production of the reservoir over time, (b) is the curve of cumulative water production of the reservoir over time, (c) is the curve of water cut of the reservoir over time, and (d) is the curve of net present value of the reservoir over time. Detailed Implementation
[0085] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0086] This invention proposes an intelligent optimization method for multi-type well fracture joint control and fine injection-production mode, such as... Figure 1 As shown, this method is mainly used for intelligent design of injection and production schemes in fractured horizontal wells. By coupling single-objective pre-search with reinforcement learning, the injection and production operation system is dynamically controlled and optimized. Specifically, it includes the following steps:
[0087] Step 1: Set up a well-fracture joint control fine injection and production mode. Combine reservoir physical parameters and horizontal well design parameters to establish a reservoir numerical simulation model in the reservoir numerical simulation software to simulate the injection and production process of the fractured horizontal well.
[0088] Specifically, to improve the sweep efficiency of injected fluids after reservoir fracturing and enhance the production properties of traditional single fractures, this embodiment sets up two well-fracture joint control precision injection and production modes: an asynchronous injection and production scheme between the same well section and a staggered injection and production scheme between different well sections. The asynchronous injection and production scheme between the same well section is designed for a single fracturing horizontal well, such as... Figure 2 As shown, fracturing horizontal well injection-production development is carried out by alternately arranging production fractures and injection fractures in the horizontal section of a single fracturing horizontal well; the staggered injection-production scheme between different well sections is designed for multiple fracturing horizontal wells, such as... Figure 3 As shown, horizontal well injection and production development is carried out by staggering fractures between the same well sections of two adjacent fractured horizontal wells.
[0089] In this embodiment, a reservoir numerical simulation model is established in the reservoir numerical simulation software based on the asynchronous injection-production scheme between the same well sections. The size of the reservoir numerical simulation model is set to... , set the grid cell size to The length of the horizontal well model in the reservoir numerical simulation model is set to The crack spacing is set to Crack half length set The initial pressure in the basic physical properties of the formation is set to Porosity set to Permeability set to The initial water saturation is set to 0.2. The relative permeability data used in the fluid model of the reservoir numerical simulation model are as follows: Figure 4 As shown.
[0090] A reservoir numerical simulation model is established in the reservoir numerical simulation software based on the staggered injection and production scheme between different well sections. The size of the reservoir numerical simulation model is set as follows: , set the grid cell size to Two horizontal well models were established in the reservoir numerical simulation model, namely the first horizontal well model and the second horizontal well model. The length of the horizontal well models was set to... The crack spacing is set to The spacing between fractures in the horizontal direction of the horizontal well model is set as follows: The fracture spacing at the fracture tip along the vertical direction of the horizontal well model is set to... Crack half length set The initial pressure in the basic physical properties of the formation is set to Porosity set to Permeability set to The initial water saturation is set to 0.2. The relative permeability data used in the fluid model of the reservoir numerical simulation model are as follows: Figure 4 As shown.
[0091] Step 2: Set the core sampling parameters and their value ranges. Based on the Latin hypercube sampling method, obtain multiple reservoir injection-production schemes. Use a reservoir numerical simulation model to simulate each scheme, obtaining the state, actions, and rewards of the reservoir numerical simulation model at each time step when using each scheme. Generate multiple sample data points and establish a representative sample database that fully covers the parameter space. This includes the following sub-steps:
[0092] Step 2.1: Based on the actual working conditions of the reservoir, set the total development time for the reservoir numerical simulation model. and initial formation pressure Set the core sampling parameters and their value range, wherein the core sampling parameters include the average value of the injected fracture flow rate. With fluctuation range Mean value of extracted fracture pressure With fluctuation range Adjustment frequency The adjustment frequency The average injection fracture flow rate is set as the reciprocal of the total number of days in the injection-production regime. Fluctuation range of injected fracture flow Average pressure of produced fractures Range of pressure fluctuations in the produced fracture and adjustment frequency .
[0093] Step 2.2: Based on the Latin hypercube sampling method, hierarchical sampling is performed in the parameter space of each core sampling parameter to obtain multiple sets of non-overlapping core sampling parameters. The collected produced fracture pressure, injected fracture flow rate and time are judged according to the preset threshold judgment rules. The collected core sampling parameters are transformed into injection and production parameters that meet the actual requirements of the field, and multiple injection and production schemes are obtained.
[0094] In this embodiment, the threshold determination rule is set as follows:
[0095] ;
[0096] In the formula, The collected pressure from the extracted fracture; The collected injection fracture flow rate; For time; Total development time.
[0097] In the threshold judgment rule, when the extracted fracture pressure and injected fracture flow rate have no negative values, the values of the extracted fracture pressure and injected fracture flow rate are set to 0. In actual engineering, the corresponding fracture is in a closed state and no injection or extraction is performed. At the same time, the extracted fracture pressure is limited to not exceeding the initial formation pressure, and the total duration of the injection and production scheme constructed by sampling does not exceed the total development time.
[0098] Step 2.3: Based on each injection and extraction plan, and in conjunction with the preset time step... The reservoir numerical simulation model in the reservoir numerical simulator is used to simulate the reservoir state, actions, and rewards obtained from the reservoir numerical simulation model at each time step, and the results are then used to simulate the reservoir state, actions, and rewards at each time step. The state of the reservoir at each time step ,action and rewards and the The state of the reservoir at each time step As sample data, the state is set as the average pressure and water saturation around the fracture, the action is set as the average injection volume and average production pressure of the fracture, and the reward is set as the step-by-step net present value.
[0099] Specifically, the formula for calculating the step-by-step net present value is as follows:
[0100] ;
[0101] In the formula, For the first The stepwise net present value of the reservoir at each time step; , , These are crude oil price, water injection cost, and produced water treatment cost, respectively. The crude oil price... Water injection cost Produced water treatment costs All units are ; , , The first The cumulative oil production, cumulative water injection, and water production of the reservoir at each time step, the first... Cumulative oil production of reservoir at each time step Cumulative water injection volume and water production All units are ; The annual discount rate; For time step.
[0102] In this embodiment, reservoir numerical simulation models are established in reservoir numerical simulation software based on asynchronous injection-production schemes within the same well section and staggered injection-production schemes between different well sections, respectively. The total development time of the reservoir numerical simulation model is set to 1080 days, the average injection fracture flow rate is set to 10~30 MPa, the fluctuation range of the injection fracture flow rate is set to 5~15 MPa, the probability of variation ranges from 0.01 to 0.1, indicating that the working system adjustment will occur within 10~100 days, the time step is set to 30 days, the annual discount rate is set to 0.1, and the crude oil price is set to 3600. The cost of water injection is set at 50. The cost of produced water treatment is set at 80. .
[0103] Step 2.4: Preprocess the sample data obtained from the simulation. First, remove the sample data that failed the simulation and contained outliers. Then, use the Z-Score standardization method to normalize the sample data. After normalization, randomly distribute the sample data to the training set and the validation set in an 8:2 ratio to construct the sample database.
[0104] Step 3: To achieve rapid simulation of reservoir dynamics, a deep learning surrogate model is established based on the LSTM (Long Short-Term Memory Network) network structure. The sample database established in Step 2 is used to train and validate the deep learning surrogate model, resulting in a validated deep learning surrogate model. This validated surrogate model is then used to replace the reservoir numerical simulation model to obtain simulation results, significantly improving the efficiency of subsequent pre-search and reinforcement learning optimization. Specifically, this includes the following sub-steps:
[0105] Step 3.1: Establish a deep learning agent model based on the LSTM network structure, and set the maximum number of iterations and loss function of the deep learning agent model.
[0106] In this embodiment, the deep learning agent model is built based on an LSTM network structure, including an input layer, a hidden layer, a fully connected layer, and an output layer. The input layer is used to input the state and actions of the reservoir numerical simulation model at the current time step. The hidden layer uses multiple stacked LSTM units to learn the pressure propagation characteristics and saturation front advancement characteristics during reservoir seepage through a gating mechanism. The fully connected layer is set between the hidden layer and the output layer. The fully connected layer is used to map the high-dimensional feature data output by the hidden layer to the physical space. The output layer is used to output the predicted state of the reservoir numerical simulation model at the next time step and the production indicators at the current time step.
[0107] The mapping function of the deep learning agent model is:
[0108] ;
[0109] In the formula, For the first Predicted values of reservoir state at each time step; , , The first Predicted values, status, and actions of reservoir production indicators at each time step; LSTM network weight parameters The proxy model mapping function; The hidden state memory passed to the LSTM network.
[0110] In this embodiment, when training the deep learning agent model, the gradient-based Adam optimization algorithm is used to update the network parameters of the deep learning agent model, and Dropout units are introduced into the LSTM network to prevent overfitting. To simultaneously measure the prediction accuracy of state and output, the mean squared error is used to construct the loss function of the deep learning agent model, resulting in:
[0111] ;
[0112] In the formula, For the loss function of the deep learning agent model; These are the weight parameters of the LSTM network; The sample data sequence number; The total number of sample data points used in training; , These are all weighting coefficients of the loss function, used to determine the degree of fit between the equilibrium state and the output. , The first The reservoir status and production indicators at each time step; For the first Predicted values of reservoir production indicators at each time step; As an L2 norm, combined with the outer square sign, this term represents the mean square error as a whole.
[0113] Step 3.2: Randomly select multiple sample data from the training set and input them into the deep learning proxy model. Train the deep learning proxy model to predict the production indicators of the reservoir numerical simulation model at the current moment and the state of the reservoir numerical simulation model at the next moment based on the state and actions of the reservoir numerical simulation model in the sample data. Calculate the loss function value of the deep learning proxy model during the training process.
[0114] Step 3.3: If the current iteration count has reached the maximum iteration count, stop training the deep learning proxy model and proceed to step 3.4. Otherwise, compare the calculated loss function value with the preset loss function value. If the calculated loss function value exceeds the preset loss function value, adjust the hyperparameters of the deep learning proxy model and return to step 3.2 to continue training the deep learning proxy model. If the calculated loss function value does not exceed the preset loss function value, stop training the deep learning proxy model and proceed to step 3.4.
[0115] Step 3.4: Randomly select multiple sample data from the test set and input them into the trained deep learning agent model. Evaluate the performance of the trained deep learning agent model based on the coefficient of determination and mean absolute percentage error. If the coefficient of determination of the trained deep learning agent model is greater than 0.9 and the mean absolute percentage error is less than 10%, proceed to step 3.5. Otherwise, adjust the hyperparameters of the deep learning agent model and return to step 3.2 to continue training the deep learning agent model.
[0116] Specifically, the formula for calculating the coefficient of determination is as follows:
[0117] ;
[0118] In the formula, The coefficient of determination; The total number of sample data used in the validation; The true value of the sample data; The predicted value output by the deep learning agent model; This is the arithmetic mean of the sample data used for validation.
[0119] The formula for calculating the mean absolute percentage error is:
[0120] ;
[0121] In the formula, The mean absolute percentage error.
[0122] Step 3.5: Output the validated deep learning agent model.
[0123] Step 4: Replace the reservoir numerical simulation model with the validated deep learning surrogate model, and use the particle swarm optimization algorithm for single-objective pre-search global optimization. Determine the optimal benchmark strategy through fast search to obtain the benchmark injection-production scheme. This includes the following sub-steps:
[0124] Step 4.1 sets maximizing the cumulative net present value over the entire cycle as the objective function of the pre-search process, expressed as:
[0125] ;
[0126] In the formula, It is a function for maximizing the value; The objective function is... This refers to the time step number; This represents the total number of time steps. For the first The stepwise net present value of the reservoir at each time step.
[0127] Production regime parameters, including injection volume and production pressure, are used as optimization variables. To reduce dimensionality, a combination of mean and fluctuation range is used to construct the injection-production regime scheme. The third parameter in the deep learning proxy model is set. The crack in the first The actions at each time step are:
[0128] ;
[0129] In the formula, The crack number; For the first The crack in the first Injected traffic at each time step; For the first The crack in the first The extraction pressure at each time step; For the first The average injection flow rate of the crack; For the first Average bottom hole flowing pressure of the fracture; For the first The amplitude of the crack fluctuation; It is a sine function.
[0130] In this embodiment, the total number of time steps of the deep learning proxy model is set. The value is 36, and the average injected flow rate is set to 36. The average bottom hole flowing pressure was set to 10~30MPa, and the fluctuation range was set to 10%~30%.
[0131] Step 4.2: Set the initial parameters of the particle swarm optimization algorithm, including the total number of particles, the maximum number of iterations, the inertia weight, and the learning factor. In this embodiment, the total number of particles is set to 80, the maximum number of iterations is set to 100, the inertia weight adopts a linear decreasing strategy, decreasing from 0.9 to 0.4 with the number of iterations, and the learning factor is set to 1.5.
[0132] Step 4.3: Using particles from the particle swarm to represent the injection-progression regime, randomly set the position and velocity of each particle in the particle swarm. The particle velocity is the rate of change of the particle position, used to represent the magnitude of the injection-progression parameter adjustment. The particle position is the injection-progression regime parameter, expressed as:
[0133] ;
[0134] In the formula, For the first The position of each particle; This represents the total number of time steps. For the first The fluctuation range of the crack.
[0135] For each particle in the particle swarm, the injection-procurement scheme represented by the particle is input into the validated deep learning proxy model for prediction, so as to obtain the production indicators and cumulative net present value corresponding to the injection-procurement scheme, which are used as the fitness value of the particle.
[0136] Step 4.4: Update the optimal value of each particle, and select the particle with the highest fitness as the globally optimal particle, resulting in:
[0137] ;
[0138] In the formula, For the first During the nth iteration calculation The optimal fitness value of each particle; For the first During the nth iteration calculation The optimal fitness value of each particle; For the first During the nth iteration calculation The position of each particle; This is the fitness value.
[0139] The formula for updating the velocity and position of each particle in the particle swarm is as follows:
[0140] ;
[0141] ;
[0142] In the formula, For the first During the nth iteration calculation The velocity of each particle; Inertial weight; , All are learning factors; , All values are random numbers, ranging from [0,1]. By introducing random perturbations into the particle velocity update, population diversity is increased, preventing premature entry into local optima. , The first During the nth iteration calculation The velocity and position of each particle; This is the globally optimal value.
[0143] Step 4.5: Determine whether the current iteration count has reached the preset maximum iteration count. If it has, proceed to step 4.6; otherwise, update the iteration count and return to step 4.4 to continue particle swarm optimization.
[0144] Step 4.6: Output the injection and extraction regime scheme corresponding to the globally optimal particle as the preferred benchmark strategy to obtain the benchmark injection and extraction regime scheme.
[0145] Step 5: Establish a reinforcement learning dynamic decision-making model based on the Actor-Critic dual network structure with policy gradient. Train the reinforcement learning dynamic decision-making model using the PPO proximal policy optimization algorithm to obtain a reinforcement learning agent, which is used for fine-tuning the benchmark injection and collection scheme. This step includes the following sub-steps:
[0146] Step 5.1: Establish a reinforcement learning dynamic decision-making model based on the Actor-Critic dual network structure of policy gradient, and connect the reinforcement learning dynamic decision-making model and the validated deep learning agent model based on the Markov decision process. Use the reinforcement learning dynamic decision-making model to predict the cumulative net present value of the reservoir within a specified time in the future according to the current injection and production scheme.
[0147] Specifically, the Markov decision process is used for parameter passing in reinforcement learning dynamic decision models and deep learning agent models by establishing a continuous state-action MDP quintuple. ,in, The state space is specifically a state vector composed of the local formation pressure and water saturation around each fracture at all time steps, which serves as the input to the reinforcement learning dynamic decision-making model. This is the action space, used to represent adjustments to the injection-production regime, by introducing an adjustment coefficient relative to the benchmark injection-production regime. The adjusted injection and collection system is as follows:
[0148] ;
[0149] In the formula, For the adjusted number The actions of the reservoir at each time step; For the first The actions of the reservoir at each time step; To adjust the coefficient, the adjustment coefficient in this embodiment is... The value range is set to [-0.2, 0.2].
[0150] The state transition model, namely the deep learning agent model verified in step 3, is responsible for receiving the current state of the reservoir and the adjusted actions, and calculating the state of the reservoir in the next time step. The reward function is the same as the step-by-step net present value calculation formula in step 2; This is a discount factor used to balance current returns with long-term returns.
[0151] In this embodiment, the reinforcement learning dynamic decision-making model is constructed based on an Actor-Critic dual-network structure with policy gradients, including a policy network and a value network. The policy network, used to adjust the injection-production regime, includes an input layer, a fully connected hidden layer, and an output layer connected sequentially, used to adjust the current reservoir state based on the input. Determine the adjustment coefficient The normal distribution function is used to determine the adjustment coefficient through random sampling. The value network, used to evaluate the merits of injection-production regime adjustments, comprises an input layer, a fully connected hidden layer, and an output layer connected in sequence, used to evaluate the current reservoir state based on the input data. Determine the state value function and calculate the advantage function. Guidance strategy network update.
[0152] The state value function is used to determine the current state of the reservoir. The cumulative net present value of a reservoir over a specified future time period is predicted by the injection-production scheme, expressed as:
[0153] ;
[0154] In the formula, The state value function; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step; Expected based on experience; Discount factor; For the first The reward for each time step of the reservoir.
[0155] Step 5.2: Train the reinforcement learning dynamic decision-making model using the PPO (Proximity Policy Optimization) algorithm, and update the network parameters of the reinforcement learning dynamic decision-making model. The specific process is as follows:
[0156] Step 5.2.1: Initialize the PPO proximal policy optimization algorithm parameters, setting the learning rate, truncation parameter, discount factor, generalized advantage estimation parameter, maximum number of training epochs, and composite total loss function.
[0157] Specifically, in this embodiment, the learning rate is set to 0.0003, the truncation parameter is set to 0.2, the discount factor is set to 0.99, the generalized advantage estimation parameter is set to 0.95, the maximum number of training epochs is set to 1000, and the composite total loss function is set as follows:
[0158] ;
[0159] in,
[0160] ;
[0161] ;
[0162] ;
[0163] In the formula, This is the composite total loss function; To maximize the strategy's reward function; The policy entropy function; , These are all weighting coefficients of the composite total loss function. In this embodiment, the weighting coefficients of the composite total loss function are... The value is set to 0.5, and the weighting coefficient of the composite total loss function is... The value is set to 0.01; These are the policy network weight parameters; These are the weight parameters of the value network; The current injection / collection strategy is determined by the strategy network weight parameters; The predicted value is the expected value based on experience; It is a minimum value function; The new and old injection and extraction system schemes were proposed in the 1st year. Select an action within a time step The probability ratio; The dominant function; This is a truncation function, and its value range is... This is used to prevent excessively large parameter updates; To truncate parameters; The value network loss function; This is an estimate of the state value; For the first The state of the reservoir at each time step; For the first The return on the reservoir at each time step; The crack number; To optimize the total number of cracks; It is a natural constant; For the first The standard deviation of the crack action; It is a logarithmic function with base 10.
[0164] Step 5.2.2: During the iterative training process, the policy network of the reinforcement learning dynamic decision-making model is first used to identify the current state of the reservoir. And output adjustment coefficient The adjusted injection and collection scheme is then input into the validated deep learning agent model, which is used to obtain the reward for the current time step. and the state of the reservoir at the next time step. A reservoir database consisting of reservoir data from all time steps is obtained. The data from this database are then input into a reinforcement learning dynamic decision model. The value network of the dynamic decision model outputs the state-value function for each time step, determining the reward and advantage function for each reservoir at each time step. The calculation formula is as follows:
[0165] ;
[0166] ;
[0167] In the formula, For the first The return on the reservoir at each time step; For the first The reward for each time step of the reservoir; Discount factor; For the first The return on the reservoir at each time step; For the first The dominant function of a reservoir at each time step; This is a summation index used to indicate the number of future delay steps; This represents the total number of time steps. This refers to the time step number; For parameters to estimate the generalized dominance; For the first The reward for each time step of the reservoir; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step.
[0168] Step 5.2.3: Based on the gradient optimization algorithm and the preset learning rate, calculate the composite total loss function value relative to the policy network weight parameters. and value network weight parameters The gradient is used to update the network parameters of the reinforcement learning dynamic decision model.
[0169] Step 5.2.4: Continue iterative calculations until the current number of iterations reaches the preset maximum number of training rounds or the preset reward convergence condition is met, then end the training of the reinforcement learning dynamic decision model.
[0170] In this embodiment, the reward convergence condition is used to evaluate the steady-state characteristics of the cumulative net present value over the entire cycle. By introducing a window size based on the moving average method to smooth the training curve of the reinforcement learning dynamic decision-making model, the relative growth rate of the average step-by-step net present value in the current window relative to the previous window is calculated. When the relative growth rate is less than the preset growth rate threshold of 0.005, it is determined that the reinforcement learning dynamic decision-making model has converged to the optimal strategy, and the iterative calculation is stopped at this time.
[0171] Step 5.3: Save the network parameters of the reinforcement learning dynamic decision-making model to obtain the reinforcement learning agent.
[0172] Step 6: Obtain the initial formation state, total development time, and time step of the actual reservoir. Input the benchmark injection-production regime scheme obtained in Step 4 into the reinforcement learning agent obtained in Step 5. Use the reinforcement learning agent to generate the optimal dynamic injection-production control scheme for the entire life cycle of the actual reservoir. Obtain the injection flow rate and production pressure of all fractures in the actual reservoir at all time steps within the total development time. Generate an intelligent injection-production regime table containing the injection flow rate sequence and production pressure sequence to obtain the optimal injection-production development scheme for the actual reservoir.
[0173] In this embodiment, a dynamic injection-production work schedule table for asynchronous injection-production schemes between the same well section is obtained based on the optimization of the method of the present invention, as follows: Figure 5 As shown, Figure 5 The negative values on the vertical axis represent the injection flow rate of the injection joint, while the positive values represent the pressure of the production joint. Figure 6 To obtain the reservoir parameter variation curves based on the optimized dynamic injection-production scheme after asynchronous injection-production schemes between the same well sections, the following steps were taken: Figure 6 The curves of cumulative oil production, cumulative water production, water cut, and net present value of the reservoir over time were obtained when an asynchronous injection and production scheme was implemented between the same well sections.
[0174] In this embodiment, a dynamic injection-production work schedule table for the inter-well section staggered injection-production scheme, optimized based on the method of the present invention, is obtained, as follows: Figure 7 As shown, Figure 7 The negative values on the vertical axis represent the injection flow rate of the injection joint, while the positive values represent the pressure of the production joint. Figure 8 To obtain the reservoir parameter variation curves based on the optimized dynamic injection-production scheme of inter-well sections, the following steps were taken: Figure 8The curves of cumulative oil production, cumulative water production, water cut, and net present value of the reservoir over time were obtained when the injection and production development was carried out based on the staggered injection and production scheme between different well sections.
[0175] In summary, the method of this invention addresses the strong heterogeneity and dynamic disturbances present in multi-fracture and multi-well environments by combining single-objective pre-search and reinforcement learning dynamic decision-making. By introducing a residual control mechanism into the reinforcement learning dynamic decision-making model, the optimal baseline strategy is dynamically fine-tuned based on real-time formation pressure and water saturation. This yields the optimal injection-production development scheme for reservoirs under different well-fracture joint control fine-tuning modes. This effectively solves the limitations of traditional numerical simulation methods for optimizing injection-production schemes, which suffer from long computation times and low efficiency. It reduces the computational cost of massive simulations to screen for the optimal injection-production scheme, enabling rapid optimization and decision support for reservoir injection-production schemes under new multi-type well-fracture joint control modes, and providing technical support for actual reservoir development strategies.
[0176] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for intelligent optimization of multi-type well fracture joint control precision injection and production mode, characterized in that, Includes the following steps: Step 1: Set up a well-fracture joint control fine injection and production mode. Combine reservoir physical property parameters and horizontal well design parameters to establish a reservoir numerical simulation model in the reservoir numerical simulation software to simulate the injection and production process of the fractured horizontal well. Step 2: Set the core sampling parameters and their ranges. Based on the Latin hypercube sampling method, obtain multiple reservoir injection and production schemes. Use the reservoir numerical simulation model to simulate each reservoir injection and production scheme to obtain the state, action and reward of the reservoir numerical simulation model at each time step when each reservoir injection and production scheme is adopted. Generate multiple sample data and establish a sample database. Step 3: Build a deep learning agent model based on the LSTM network structure, and train and validate the deep learning agent model using a sample database to obtain the validated deep learning agent model. Step 4: Replace the reservoir numerical simulation model with the validated deep learning surrogate model, and use the particle swarm optimization algorithm for single-objective pre-search global optimization. Determine the optimal benchmark strategy through fast search to obtain the benchmark injection-production scheme. Step 5: Establish a reinforcement learning dynamic decision-making model based on the Actor-Critic dual network structure of policy gradient, train the reinforcement learning dynamic decision-making model based on the PPO proximal policy optimization algorithm, and obtain a reinforcement learning agent for fine-tuning the benchmark injection and extraction system scheme. Step 6: Obtain the initial formation state, total development time, and time step of the actual reservoir. Input the benchmark injection-production regime scheme obtained in Step 4 into the reinforcement learning agent. Use the reinforcement learning agent to obtain the injection flow rate and production pressure of all fractures in the actual reservoir at all time steps within the total development time. Generate an intelligent injection-production regime table containing the injection flow rate sequence and production pressure sequence to obtain the optimal injection-production development scheme for the actual reservoir.
2. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 1, characterized in that, In step 1, the well-fracture joint control precision injection and production mode includes an asynchronous injection and production scheme between the same well section and an interleaved injection and production scheme between different well sections. The asynchronous injection and production scheme between the same well section is designed for a single fractured horizontal well, and the injection and production development of the fractured horizontal well is carried out by alternately arranging production fractures and injection fractures in the horizontal well section of the single fractured horizontal well. The interleaved injection and production scheme between different well sections is designed for multiple fractured horizontal wells, and the injection and production development of the fractured horizontal well is carried out by alternately setting fractures in the same well section between two adjacent fractured horizontal wells.
3. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 1, characterized in that, Step 2 includes the following sub-steps: Step 2.1: Based on the actual working conditions of the reservoir, set the total development time and initial formation pressure of the reservoir numerical simulation model, and set the core sampling parameters and their value ranges. The core sampling parameters include the mean and fluctuation range of the injected fracture flow rate, the mean and fluctuation range of the produced fracture pressure, and the adjustment frequency. Step 2.2: Based on the Latin hypercube sampling method, hierarchical sampling is performed in the parameter space of each core sampling parameter to obtain multiple sets of non-overlapping core sampling parameters. According to the preset threshold judgment rule, the collected core sampling parameters are converted into injection and sampling parameters to obtain multiple injection and sampling schemes. Step 2.3: Based on each injection-production scheme and the preset time step, simulate the reservoir using the reservoir numerical simulation model in the reservoir numerical simulator to obtain the reservoir state, actions, and rewards obtained from the reservoir numerical simulation model at each time step, and then... The state of the reservoir at each time step ,action and rewards and the The state of the reservoir at each time step As sample data, the state is set as the average pressure and water saturation around the fracture, the action is set as the average injection volume and average production pressure of the fracture, and the reward is set as the step-by-step net present value. Step 2.4: Preprocess the sample data obtained from the simulation. First, remove the sample data that failed the simulation and contained outliers. Then, use the Z-Score standardization method to normalize the sample data. Randomly allocate the normalized sample data to the training set and the validation set to construct the sample database.
4. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 1, characterized in that, Step 3 includes the following sub-steps: Step 3.1: Establish a deep learning agent model based on the LSTM network structure, and set the maximum number of iterations and loss function of the deep learning agent model; Step 3.2: Randomly select multiple sample data from the training set and input them into the deep learning proxy model. Train the deep learning proxy model to predict the production index of the reservoir numerical simulation model at the current moment and the state of the reservoir numerical simulation model at the next moment based on the state and action of the reservoir numerical simulation model in the sample data. Calculate the loss function value of the deep learning proxy model during the training process. Step 3.3: If the current iteration count has reached the maximum number of iterations, stop training the deep learning proxy model and proceed to step 3.
4. Otherwise, compare the calculated loss function value with the preset loss function value. If the calculated loss function value exceeds the preset loss function value, adjust the hyperparameters of the deep learning proxy model and return to step 3.2 to continue training the deep learning proxy model. If the calculated loss function value does not exceed the preset loss function value, stop training the deep learning proxy model and proceed to step 3.
4. Step 3.4: Randomly select multiple sample data from the test set and input them into the trained deep learning agent model. Evaluate the performance of the trained deep learning agent model based on the coefficient of determination and the mean absolute percentage error. If the coefficient of determination of the trained deep learning agent model is greater than the preset coefficient of determination value and the mean absolute percentage error is less than the preset error value, proceed to step 3.
5. Otherwise, adjust the hyperparameters of the deep learning agent model and return to step 3.2 to continue training the deep learning agent model. Step 3.5: Output the validated deep learning agent model.
5. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 4, characterized in that, The deep learning proxy model is based on an LSTM network structure, including an input layer, a hidden layer, a fully connected layer, and an output layer. The input layer is used to input the state and actions of the reservoir numerical simulation model at the current time step. The hidden layer uses multiple stacked LSTM units to learn the pressure propagation characteristics and saturation front advancement characteristics during reservoir seepage through a gating mechanism. The fully connected layer is set between the hidden layer and the output layer to map the high-dimensional feature data output by the hidden layer to the physical space. The output layer is used to output the predicted state of the reservoir numerical simulation model at the next time step and the production indicators at the current time step. The mapping function of the deep learning agent model is: ; In the formula, For the first Predicted values of reservoir state at each time step; , , The first Predicted values, status, and actions of reservoir production indicators at each time step; LSTM network weight parameters The proxy model mapping function; The hidden state memory passed to the LSTM network; The loss function of the deep learning agent model is: ; In the formula, For the loss function of the deep learning agent model; These are the weight parameters of the LSTM network; The sample data sequence number; The total number of sample data points used in training; , These are all weighting coefficients of the loss function; , The first The reservoir status and production indicators at each time step; For the first Predicted values of reservoir production indicators at each time step; It is an L2 norm.
6. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 1, characterized in that, Step 4 includes the following sub-steps: Step 4.1 sets maximizing the cumulative net present value over the entire cycle as the objective function of the pre-search process, expressed as: ; in, ; In the formula, It is a function for maximizing the value; The objective function is... This refers to the time step number; This represents the total number of time steps. For the first The stepwise net present value of the reservoir at each time step; , , These are crude oil price, water injection cost, and produced water treatment cost, respectively. , , The first The cumulative oil production, cumulative water injection, and water production of the reservoir at each time step; The annual discount rate; For time step; Using production regime parameters, including injection volume and production pressure, as optimization variables, the deep learning surrogate model is set with the following parameters: The crack in the first The actions at each time step are: ; In the formula, The crack number; For the first The crack in the first Injected traffic at each time step; For the first The crack in the first The extraction pressure at each time step; For the first The average injection flow rate of the crack; For the first Average bottom hole flowing pressure of the fracture; For the first The amplitude of the crack fluctuation; It is a sine function; Step 4.2: Set the initial parameters of the particle swarm optimization algorithm, including the total number of particles, the maximum number of iterations, the inertia weight, and the learning factor; Step 4.3: Using the particle swarm to represent the injection and collection regime scheme, the position and velocity of each particle in the particle swarm are randomly set. The velocity of the particle is the rate of change of the particle position, which is used to represent the magnitude of the injection and collection parameter adjustment. The position of the particle is the injection and collection regime parameter. For each particle in the particle swarm, the injection-extraction scheme represented by the particle is input into the validated deep learning proxy model for prediction, so as to obtain the production indicators and cumulative net present value corresponding to the injection-extraction scheme, and use them as the fitness value of the particle. Step 4.4: Update the optimal value of each particle, making the particle with the highest fitness the global optimal particle, and update the velocity and position of each particle in the particle swarm; Step 4.5: Determine whether the current iteration count has reached the preset maximum iteration count. If it has, proceed to Step 4.6; otherwise, update the iteration count and return to Step 4.4 to continue particle swarm optimization. Step 4.6: Output the injection and extraction regime scheme corresponding to the globally optimal particle as the preferred benchmark strategy to obtain the benchmark injection and extraction regime scheme.
7. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 1, characterized in that, Step 5 includes the following sub-steps: Step 5.1: Establish a reinforcement learning dynamic decision-making model based on the Actor-Critic dual network structure of policy gradient, and connect the reinforcement learning dynamic decision-making model and the validated deep learning agent model based on the Markov decision process. Use the reinforcement learning dynamic decision-making model to predict the cumulative net present value of the reservoir within a specified time in the future based on the current injection and production scheme. Step 5.2: Train the reinforcement learning dynamic decision-making model using the PPO (Proximity Policy Optimization) algorithm, and update the network parameters of the reinforcement learning dynamic decision-making model. Step 5.3: Save the network parameters of the reinforcement learning dynamic decision-making model to obtain the reinforcement learning agent.
8. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 7, characterized in that, The reinforcement learning dynamic decision-making model is constructed based on an Actor-Critic dual-network structure with policy gradients, comprising a policy network and a value network. The policy network is used to adjust the injection-production regime, including an input layer, a fully connected hidden layer, and an output layer connected in sequence. It is used to determine the normal distribution function of the adjustment coefficients based on the current state of the reservoir, and the adjustment coefficients are determined through random sampling. The value network is used to evaluate the merits of the injection-production regime adjustment, including an input layer, a fully connected hidden layer, and an output layer connected in sequence. It is used to determine the state value function based on the current state of the reservoir, and the value function is used to guide the policy network update by calculating the advantage function. The state-value function is used to predict the cumulative net present value of the reservoir over a specified future period based on the current reservoir state and the injection-production scheme. The expression is: ; In the formula, The state value function; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step; Expected based on experience; Discount factor; For the first The reward for each time step of the reservoir; The Markov Decision Process (MDP) is used for parameter passing in reinforcement learning dynamic decision-making models and deep learning agent models, by establishing a continuous state-action MDP quintuple. ,in, This represents the state space, which serves as the input to the reinforcement learning dynamic decision-making model. This is the action space, used to represent adjustments to the injection and extraction system; This is a state transition model, responsible for receiving the current state of the reservoir and the adjusted actions, and calculating the state of the reservoir in the next time step. For the reward function; This is the discount factor.
9. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 8, characterized in that, Step 5.2 includes the following sub-steps: Step 5.2.1: Initialize the PPO proximal policy optimization algorithm parameters, setting the learning rate, truncation parameter, discount factor, generalized advantage estimation parameter, maximum number of training epochs, and composite total loss function; Step 5.2.2: During iterative training, the policy network of the reinforcement learning dynamic decision-making model is first used to identify the current state of the reservoir and output adjustment coefficients. The adjusted injection-production scheme is then input into the validated deep learning agent model. The validated deep learning agent model is used to obtain the reward for the current time step and the reservoir state for the next time step, acquiring a reservoir database composed of reservoir data from all time steps. The data in the reservoir database are then input into the reinforcement learning dynamic decision-making model. The value network of the reinforcement learning dynamic decision-making model outputs the state value function for each time step, determining the reward and advantage function of the reservoir at each time step. The calculation formula is as follows: ; ; In the formula, For the first The return on the reservoir at each time step; For the first The reward for each time step of the reservoir; Discount factor; For the first The return on the reservoir at each time step; For the first The dominant function of a reservoir at each time step; For summation index; This represents the total number of time steps. This refers to the time step number; For parameters to estimate the generalized dominance; For the first The reward for each time step of the reservoir; For the first The state of the reservoir at each time step; For the first The state of the reservoir at each time step; Step 5.2.3: Based on the gradient optimization algorithm, calculate the gradient of the composite total loss function value relative to the policy network weight parameters and the value network weight parameters, and update the network parameters of the reinforcement learning dynamic decision model by combining the preset learning rate. Step 5.2.4: Continue iterative calculations until the current number of iterations reaches the preset maximum number of training rounds or the preset reward convergence condition is met, then end the training of the reinforcement learning dynamic decision model.
10. The intelligent optimization method for multi-type well fracture joint control fine injection and production mode according to claim 9, characterized in that, The composite total loss function is set as follows: ; in, ; ; ; In the formula, This is the composite total loss function; To maximize the strategy's reward function; The policy entropy function; , All are weighting coefficients of the composite total loss function; These are the policy network weight parameters; These are the weight parameters of the value network; The current injection / collection strategy is determined by the strategy network weight parameters; The predicted value is the expected value based on experience; It is a minimum value function; The new and old injection and extraction system schemes were proposed in the 1st year. Select an action within a time step The probability ratio; The dominant function; This is a truncation function; To truncate parameters; The value network loss function; This is an estimate of the state value; For the first The state of the reservoir at each time step; For the first The return on the reservoir at each time step; The crack number; To optimize the total number of cracks; It is a natural constant; For the first The standard deviation of the crack action; It is a logarithmic function with base 10.
Citation Information
Patent Citations
Layered water injection optimization method based on long short-term memory neural network and particle swarm optimization algorithm
CN114357852A
Oil reservoir injection-production optimization method based on deep reinforcement learning
CN114444402A