A multi-scene oil reservoir generalization injection-production optimization method based on economic state perception
By introducing an economic state perception mechanism and a multi-view value network, the problems of fixed cycles and insufficient safety constraints in reinforcement learning methods in reservoir optimization are solved, and adaptive optimization and generalized application in reservoirs under multiple scenarios are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA UNIV OF PETROLEUM (EAST CHINA)
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-26
AI Technical Summary
Existing reinforcement learning methods for reservoir production optimization suffer from problems such as fixed optimization cycles, insufficient generalization ability across multiple geological scenarios, and lack of engineering safety constraints. As a result, they are difficult to achieve effective generalization in reservoir scenarios with different permeability, water cut stages, and development paces.
An adaptive weighting mechanism based on economic status perception is introduced to construct a multi-perspective value network. By dynamically adjusting and optimizing the weights of the cycle and time perspectives through real-time monitoring of economic indicators, and combining this with an engineering safety constraint module, parallel evaluation and fusion of values at different time scales can be achieved.
It achieves improved economic rationality and computational efficiency throughout the entire lifecycle, enables adaptive decision-making at different penetration rates and development stages, has cross-scenario generalization capabilities, avoids ineffective or negative-return production, and ensures engineering safety.
Smart Images

Figure CN121684558B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of petroleum engineering technology, specifically relating to a multi-scenario reservoir generalized injection and production optimization method based on economic status perception. Background Technology
[0002] Reservoir production optimization is an important research direction in the field of petroleum development. Its core objective is to maximize cumulative oil production or net present value (NPV) by controlling variables such as water injection rate and production volume, while meeting engineering constraints. Traditional optimization methods usually rely on reservoir numerical simulation and heuristic optimization algorithms (such as genetic algorithms and particle swarm optimization). However, these methods are highly dependent on geological models, have high computational complexity, and are difficult to achieve real-time optimal control under dynamically changing reservoir conditions.
[0003] In recent years, Deep Reinforcement Learning (DRL) has been increasingly applied to solve complex reservoir injection-production optimization problems due to its powerful decision-making capabilities. Its basic idea is to automatically learn control strategies that maximize long-term rewards through the interaction between an agent and the reservoir numerical simulation environment. However, existing reinforcement learning optimization frameworks face significant technical bottlenecks when generalizing to reservoir development across multiple scenarios and the entire lifecycle, mainly in the following two aspects.
[0004] First, the fixed optimization cycle and single time perspective lead to insufficient generalization ability of the model: Existing reservoir injection and production optimization methods usually adopt a fixed optimization cycle, such as formulating and implementing optimization schemes in monthly, quarterly or annual units. Although this approach is convenient for engineering management, it also has obvious limitations: First, the dynamic response of reservoirs is affected by geological heterogeneity, fluid changes and well network interference. Its change rhythm does not strictly follow the artificially divided time period. The fixed cycle may cause the optimization action to lag behind the actual dynamics and miss the best control opportunity; Second, in the case of oil price fluctuations, equipment status or development stage changes, the fixed cycle is difficult to flexibly adapt to rapidly changing economic and technical conditions, which may lead to economic loss; In addition, the long-term fixed optimization window may not be able to fully capture short-term fluctuations or sudden anomalies, affecting the real-time performance and adaptability of the optimization scheme. Therefore, the fixed optimization time length restricts the agility and accuracy of the injection and production system in dealing with complex dynamics to a certain extent. Specifically, in reinforcement learning applications, this "one-size-fits-all" setting seriously violates the physical and economic laws of reservoir development: (1) Differences in physical laws: Reservoirs with different geological properties have huge differences in fluid transport and pressure transmission speed. High-permeability reservoirs develop quickly and may reach their development limit in a short time step; while low-permeability reservoirs take longer to show results and require a longer water injection and pressure-holding period. (2) Differences in economic cycles: The effective economic life cycle of a reservoir depends on the input-output ratio. When the output benefits are lower than the water injection and treatment costs, further optimization is not only meaningless but also "negative returns". In addition, existing reinforcement learning algorithms generally use a single and fixed discount factor to define the "farsightedness" of the agent. This pre-set time preference makes the model very easy to overfit to the specific reservoir environment during training. For example, a "shortsighted" strategy trained on a short-cycle high-permeability reservoir will lead to a significant drop in recovery rate once it is applied to a long-cycle low-permeability reservoir due to a lack of patience and neglect of formation energy recovery; and vice versa. Due to the lack of a mechanism that can perceive the current economic state and dynamically adjust the optimization time and time vision weight accordingly, existing models are difficult to achieve effective generalization and transfer between reservoir scenarios with different permeabilities and different water-cut stages.
[0005] Second, there is a lack of deterministic engineering safety guarantees: Reinforcement learning is essentially a "black box" decision-making process based on probability distributions, and the directly output actions are highly susceptible to violating physical engineering constraints. Existing technologies mostly employ "soft constraint" mechanisms, that is, imposing penalties in the reward function, but this cannot guarantee 100% safety compliance. In actual production, any probabilistic violation can lead to a production stoppage, therefore, there is an urgent need to introduce hard constraint mechanisms with physical meaning.
[0006] In summary, while existing technologies can achieve a certain degree of injection-production optimization under specific reservoir conditions, their optimization cycles typically rely on pre-set parameters and have a fixed time horizon, making it difficult to dynamically adjust based on the reservoir's real-time economic performance. This often necessitates redesigning and retraining models for reservoirs with different permeability characteristics, water-cut stages, and development paces, severely limiting the generalization ability of reinforcement learning methods across the entire reservoir lifecycle and various scenarios. Furthermore, the lack of a unified mechanism to explicitly incorporate reservoir economic status into the decision-making process and drive optimization termination and adaptive adjustment of time preferences makes it difficult for existing methods to achieve accurate assessment and efficient utilization of reservoir development value while ensuring engineering safety. Summary of the Invention
[0007] To address the limitations of existing reinforcement learning methods in reservoir production optimization, such as fixed optimization cycles, insufficient generalization ability across multiple geological scenarios, and lack of engineering safety constraints, this invention proposes a multi-scenario reservoir generalized injection-production optimization method based on economic state awareness. This method introduces an economic state-driven adaptive weight mechanism within the reinforcement learning framework to construct a multi-view value network, enabling parallel evaluation and fusion of value expectations at different time scales. The economic state vector plays a crucial dual-driving role: firstly, it serves as a dynamic truncation criterion, automatically determining the optimal optimization termination time for the current reservoir by monitoring whether economic indicators fall below a preset threshold in real time, thus dynamically adapting to the differentiated development cycles of reservoirs with varying geological conditions; secondly, it dynamically allocates weights for different discount factors, allowing the agent to automatically adjust its focus on short-term returns and long-term value based on the current economic development stage. Combined with an engineering safety constraint module, this invention effectively utilizes multi-discount factor advantage estimation technology to achieve injection-production optimization decisions that combine physical safety, time preference adaptability, and efficient economic returns across multiple reservoir lifecycle scenarios.
[0008] The technical solution of the present invention is as follows:
[0009] A multi-scenario reservoir generalized injection-production optimization method based on economic status perception includes the following steps:
[0010] Step 1: Based on the physical constraints, a reinforcement learning interactive environment is used to run the reinforcement learning interactive loop in the reservoir numerical simulation environment.
[0011] Step 2: Update the economic state vector and use the economic state vector to perform real-time economic evaluation;
[0012] Step 3: Construct an experience replay mechanism that incorporates both physical and economic information;
[0013] Step 4: Generate multi-view target Q-values through the multi-view value network, and update the parameters of the multi-view value network based on the loss function;
[0014] Step 5: Construct a dynamic weighted architecture driven by economic status, realize the weighted fusion of multi-view value assessment results based on the adaptive weighted network, construct the optimization objective function of the policy network based on the weighted fusion results, and update the policy network parameters.
[0015] Step 6: Repeat steps 1 to 5 until the training termination condition is met, and output the optimal strategy network; deploy the optimal strategy network on the target reservoir to be optimized, and directly generate the corresponding optimal injection and production control action sequence by collecting the original physical state vector and economic state vector of the target reservoir in real time and inputting them into the optimal strategy network.
[0016] Furthermore, in step 1, the reinforcement learning interaction loop performs the following operations at each time step:
[0017] Step 1.1: Obtain the real-time raw physical state vector of the reservoir and input it into the strategy network to generate candidate control actions. ;definition For the current time step, This is the original physical state vector at the current time step; the physical state vector is used as the input to the state representation in normalized form within the policy network, defined as follows. This is the normalized physical state vector for the current time step.
[0018] Step 1.2: Recover the physical space characteristics using the state decoding module;
[0019] Step 1.3: Combine the engineering safety constraint module to perform safety checks and corrections on the candidate control actions to obtain the final safe control actions to be executed;
[0020] Step 1.4: Input the final safety control action to be executed in the current time step into the reservoir simulator, and output the physical state vector of the next time step returned by the reservoir numerical simulation environment and the instantaneous reward of the current time step.
[0021] Furthermore, in step 1.2, the specific working process of the state decoding module is as follows:
[0022] Step 1.2.1: First, perform inverse normalization using the pre-stored mean. and standard deviation Normalize the physical state vector at the current time step Restored to the corresponding actual physical parameter scale:
[0023] (1);
[0024] in, Represents element-wise product; This is the physical state vector at the current time step;
[0025] Step 1.2.2: Next, perform spatial reconstruction processing to reconstruct the physical state vector into a form with spatial dimensions. 3D reservoir physical property mesh:
[0026] (2);
[0027] in, This is the 3D reservoir physical property mesh for the current time step; , , These represent the number of grid cells in the reservoir model along the x, y, and z directions, respectively. This is a dimension reconstruction function.
[0028] Furthermore, in step 1.3, the specific working process of the engineering safety constraint module is as follows:
[0029] Step 1.3.1: Determine the safety space; based on the three-dimensional reservoir physical property mesh and a set of preset engineering rules, determine the real-time safety action space boundary. :
[0030] (3);
[0031] in, This is the injection control action vector; For the first The constraint function of the project; This represents the total number of preset engineering conditions;
[0032] Step 1.3.2: Perform security checks on candidate control actions and generate the final safe control action to be executed at the current moment; the security check process is as follows: judge the candidate control actions generated by the policy network. Does it exceed the boundary of the safe movement space? ;
[0033] When candidate control action Exceeding the boundaries of the safe operating space At that time, candidate control actions are corrected; specifically, the components of the current candidate control actions that exceed the boundary of the safety action space are projected or truncated to the nearest safety compliance value defined by the boundary of the safety action space, so as to generate the safety control action to be finally executed at the current time step. :
[0034] (4);
[0035] in, For the first The final safety control action executed corresponding to each control variable; It is an interval cutoff function; The corresponding first candidate control action generated by the policy network Each action component; , These represent the upper and lower safety limits of the action components calculated based on the three-dimensional reservoir physical property grid;
[0036] When candidate control action Not exceeding the safe operating space boundary When the current candidate control action is selected, it is directly determined as the final safety control action to be executed in the current time step. .
[0037] Furthermore, in step 1.4, the immediate reward combines the economic benefits and costs of reservoir development:
[0038] (5);
[0039] in, The immediate reward for the current time step; This indicates the price per unit of crude oil; This indicates the increase in oil production within the current time step; This indicates the unit cost of water injection; This indicates the increase in water injection volume within the current time step; This indicates the unit cost of wastewater treatment. This indicates the increase in water production within the current time step.
[0040] Furthermore, the specific process of step 2 is as follows:
[0041] Step 2.1: Update the economic state vector at each time step based on immediate rewards; the economic state vector specifically includes long-term return indicators and short-term return trend indicators:
[0042] (6);
[0043] in, This is the economic state vector at the current time step; This is a long-term return indicator for the current time step. This is a short-term return trend indicator for the current time step.
[0044] The calculation formula is as follows:
[0045] (7);
[0046] in, This represents the final value of positive cash flow up to the current time step; This represents the present value of negative cash flows up to the current time step; This represents the cumulative number of running days at the current time step.
[0047] , The calculation formulas are as follows:
[0048] (8);
[0049] (9);
[0050] in, For the first The instantaneous economic benefit value corresponding to each time step; The discount rate; For initial investment; For the first The cumulative number of operating days for each time step;
[0051] The calculation formula is:
[0052] (10);
[0053] in, , These represent the short-term and long-term exponential moving averages of the instantaneous reward at the current time step, respectively. To prevent division by zero of small constants;
[0054] , The calculation formulas are as follows:
[0055] (11);
[0056] (12);
[0057] in, The smoothing coefficient for a short-term exponential moving average. The smoothing coefficient of an exponential moving average over a long period; For the previous time step; , These represent the short-term and long-term exponential moving averages of the instantaneous reward at the previous time step, respectively.
[0058] Step 2.2: Perform real-time economic evaluation using the updated economic state vector at each time step to assess the reservoir's development potential and economic life stage in real time. First, determine whether the economic state vector at the current time step meets the preset economic limit cutoff condition. If it does, it is determined that the current reservoir development process has entered an economically unsustainable stage and no longer has development potential, and the current round is terminated. If it does not meet the condition, the economic state vector at the current time step is retained. The economic limit cutoff condition is: when the long-term return index in the economic state vector is lower than the preset economic threshold, and the short-term return trend index shows a continuous deterioration trend within a preset number of consecutive time steps, the economic limit cutoff condition is met.
[0059] Furthermore, in step 3, the experience replay mechanism involves: generating a physical state vector containing the current time step. The control actions generated and executed by the policy network at the current time step. Instant rewards for the current time step The physical state vector of the next time step and the economic state vector at the current time step Transfer data Store in the experience replay pool.
[0060] Furthermore, in step 4, the multi-view value network includes two sets of target value networks, and the calculation process is as follows:
[0061] (13);
[0062] in, Indicates the training of the first The target Q value of the visual action value network; For the first Discount factor corresponding to each field of view; This marks the completion of the task; Indicates the sequence number of the two sets of target value networks; This means that, under the corresponding viewpoint, the minimum value is selected from the target Q values output by the two target value networks as the target Q value estimate under that viewpoint; This indicates that under the corresponding field of view, by the first The action value function obtained by applying the target value network; For the first Set target value network Network parameters; This is the normalized physical state vector for the next time step; In order to be in The following is from the current policy network The actions obtained from sampling; For parameters Representational policy network; This is the entropy regularization coefficient;
[0063] loss function Specifically, it is set to minimize the mean square error between the estimated Q value and the target Q value based on the sampled data, as shown in the formula:
[0064] (14);
[0065] in, For the first The parameters of the target value network corresponding to each viewpoint; For mathematical expectation operations; For experience replay pool; For the first Action value function under a single perspective; For the first Action value function under a single viewpoint Parameters; For input and Then through the action value function The estimated Q value was obtained;
[0066] The loss function is minimized using gradient descent to update the parameters of the multi-view value network.
[0067] Furthermore, the specific process of step 5 is as follows:
[0068] Step 5.1: Sample the physical state vector Input the multi-view value network to obtain the current estimated Q-value evaluation results. The estimated Q-value evaluation result is a set of estimated Q-values corresponding to multiple fields of view, where each component in the set corresponds to an estimated Q-value under a single field of view.
[0069] Step 5.2: Sample the economic state vector Input the adaptive weight network to generate dynamic weights corresponding to each discount factor:
[0070] (15);
[0071] in, For the first Dynamic weights corresponding to each field of view; It is a natural constant; , The adaptive weight network is for the first The, the The raw log odds of each field of view output; The total number of multi-view fields is set;
[0072] Step 5.3: Use dynamic weights to weight and combine the Q-value evaluation results of the multi-view estimation to construct the optimization objective function of the policy network. :
[0073] (16);
[0074] in, For policy networks Parameters;
[0075] By using an optimization objective function to update the policy network parameters, adaptive time preference adjustment based on economic state awareness is achieved.
[0076] Furthermore, in step 6, the training termination condition is: when the fluctuation range of the net present value of the economy for a consecutive preset number of training rounds is less than a set threshold, or when the preset maximum number of training iterations is reached, it is determined that the training termination condition is met.
[0077] The beneficial technical effects brought about by this invention are as follows.
[0078] (1) By introducing an economic state perception mechanism, this invention breaks the limitation of relying on a preset fixed optimization cycle in traditional reinforcement learning injection and production optimization, enabling the agent to dynamically determine the optimal optimization termination time based on the real-time economic performance of the reservoir, effectively avoiding the continued execution of invalid simulations and negative-return production in the later stages of development or inefficient stages, and improving the economic rationality and computational efficiency of the whole life cycle optimization.
[0079] (2) This invention constructs a multi-view value assessment network driven by economic status and an adaptive weight fusion mechanism, which enables the reinforcement learning strategy to automatically adjust time preference between short-term gains and long-term value. This overcomes the problem in existing methods where the strategy is difficult to adapt to different development stages due to the single and fixed discount factor. It realizes adaptive decision-making of the same model in different scenarios such as high-permeability short-cycle reservoirs and low-permeability long-cycle reservoirs.
[0080] (3) This invention integrates reservoir physical and economic information under a unified reinforcement learning framework, so that the trained strategy model can be directly transferred and deployed to multiple reservoir models without redesigning the algorithm structure or retraining the optimization model for different reservoir scenarios. This improves the generalization ability and engineering practical value of reinforcement learning injection and production optimization method in multi-scenario reservoir development. Attached Figure Description
[0081] Figure 1 This is a flowchart illustrating the design of the multi-scenario reservoir generalized injection-production optimization method based on economic status perception, as proposed in this invention.
[0082] Figure 2This is a schematic diagram of the economic adaptive value assessment and update framework in the multi-scenario reservoir generalized injection and production optimization method based on economic state perception of the present invention.
[0083] Figure 3 This is a schematic diagram of the convergence curve of the economic net present value iterative optimization using the method of the present invention in a typical reservoir test case in an embodiment of the present invention.
[0084] Figure 4 This is a schematic diagram illustrating the adaptive evolution of the time-vision weights as economic conditions change when the discount factor is 0.90 during the training process, using the method of this invention.
[0085] Figure 5 This is a schematic diagram illustrating the adaptive evolution of the time-vision weights as economic conditions change when the discount factor is 0.99 during the training process, using the method of this invention.
[0086] Figure 6 This is a schematic diagram illustrating the convergence process of the economic net present value as training rounds when the differential evolution algorithm optimizes injection and production parameters on a target reservoir model that has not participated in training. Detailed Implementation
[0087] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0088] like Figure 1 and Figure 2 As shown, this invention proposes a multi-scenario reservoir generalized injection-production optimization method based on economic state perception, aiming to solve the problems of insufficient generalization ability, lack of safety constraints, and lack of decision-making flexibility caused by fixed time perspective in existing reinforcement learning algorithms in multi-scenario reservoir development. The method specifically includes the following steps:
[0089] Step 1: Based on a physical constraint-based reinforcement learning interactive environment, run a reinforcement learning interactive loop in the reservoir numerical simulation environment; the loop performs the following operations at each time step:
[0090] Step 1.1: Obtain the real-time raw physical state vector of the reservoir and input it into the strategy network to generate candidate control actions. ;definition For the current time step, This is the original physical state vector at the current time step. To meet the numerical stability requirements of the reinforcement learning model, the physical state vector is normalized within the policy network. As the input for state representation, the normalized form is the representation of the original physical state vector after numerical scaling. The normalized physical state vector for the current time step is used for internal calculations within the reinforcement learning model. The policy network is a reinforcement learning model, specifically implemented using a feedforward fully connected neural network structure to ensure computational stability and deployability in real-time reservoir optimization applications. This network is used to map the reservoir physical state vector to injection-production control actions, with its input being the original physical state vector for the current time step and its output being the corresponding candidate control actions.
[0091] Step 1.2: Recover the physical space features through the state decoding module; the specific working process of the state decoding module is as follows:
[0092] Step 1.2.1: First, perform inverse normalization using pre-stored statistical parameters (including the mean). and standard deviation The physical state vector normalized to the current time step. When restored to the corresponding actual physical parameter scale, its calculation satisfies the following relationship:
[0093] (1);
[0094] in, Represents element-wise product; This is the physical state vector at the current time step, which has had its physical dimensions restored through inverse normalization.
[0095] Step 1.2.2: Next, perform spatial reconstruction processing (also known as anti-flattening processing) to reconstruct the physical state vector, which has recovered its physical dimensions, into a vector with spatial dimensions. The calculation of the three-dimensional reservoir physical property mesh satisfies the following relationship:
[0096] (2);
[0097] in, This is the 3D reservoir physical property mesh for the current time step; , , These represent the number of grid cells in the reservoir model along the x, y, and z directions, respectively. For dimension reconstruction functions;
[0098] This invention provides a three-dimensional reservoir physical property grid with physical spatial distribution characteristics for subsequent engineering safety constraint modules through inverse normalization and inverse flattening processing.
[0099] Step 1.3: Combine the engineering safety constraint module to perform safety checks and corrections on the candidate control actions, and obtain the final safe control actions to be executed. The main function of the engineering safety constraint module is to perform constraint correction, and the specific working process is as follows:
[0100] Step 1.3.1: Determine the safety space; based on the three-dimensional reservoir physical property mesh and a set of preset engineering rules (including maximum injection rate, maximum bottom hole pressure, etc.), determine the real-time safety action space boundary. Its mathematical expression is:
[0101] (3);
[0102] in, This is the injection and extraction control action vector, representing the values of control variables such as water injection rate and production control parameters at the current time step; For the first The constraint function of the project; This represents the total number of preset engineering conditions;
[0103] Step 1.3.2: Perform security checks on candidate control actions and generate the final safe control action to be executed at the current moment; the security check process is as follows: judge the candidate control actions generated by the policy network. Does it exceed the boundary of the safe action space? ;
[0104] When candidate control action Beyond the boundary of the safety action space At that time, candidate control actions are corrected; specifically, the components of the current candidate control actions that violate the rules (i.e., exceed the boundary of the safety action space) are projected or truncated to the nearest safety compliance value defined by the boundary of the safety action space, so as to generate the safety control action to be finally executed at the current time step. Its calculation satisfies the following relationship:
[0105] (4);
[0106] in, For the first The final safety control action executed corresponding to each control variable; This is a range cutoff function used to restrict the input value to a given safe upper and lower limit range. When the input value is greater than the upper limit, the upper limit value is taken; when the input value is less than the lower limit, the lower limit value is taken; otherwise, the value remains unchanged. The corresponding first candidate control action generated by the policy network Each action component; , These represent the upper and lower safety limits of the action components calculated based on the three-dimensional reservoir physical property grid;
[0107] When candidate control action Not exceeding the boundary of the safe action space When a candidate control action is approved, that is, the current candidate control action is directly determined as the final safety control action to be executed in the current time step. .
[0108] Step 1.4: Execute the final safety control action at the current time step. Input the reservoir simulator, output the physical state vector of the next time step returned by the reservoir numerical simulation environment. And the instant reward for the current time step; Indicates the next time step;
[0109] The immediate reward combines the economic benefits and costs of reservoir development, and its calculation satisfies the following relationship:
[0110] (5);
[0111] in, The immediate reward for the current time step; This indicates the price per unit of crude oil; This indicates the increase in oil production within the current time step; This indicates the unit cost of water injection; This indicates the increase in water injection volume within the current time step; This indicates the unit cost of wastewater treatment. This indicates the increase in water production within the current time step.
[0112] By explicitly incorporating the economic benefits and costs of reservoir development into the immediate reward calculation, the optimization objectives of reinforcement learning are aligned with actual economic evaluation indicators, thereby avoiding economic biases caused by decision-making driven solely by production or a single technical indicator.
[0113] Step 2: Update the economic state vector and use it to perform real-time economic evaluation; the specific process is as follows:
[0114] Step 2.1: Update the economic state vector once at each time step based on immediate rewards;
[0115] The economic state vector serves two purposes: firstly, it characterizes the long-term economic feasibility and short-term return trends of the reservoir at the current development stage, and acts as a dynamic termination criterion to determine whether the optimization process should continue; secondly, it serves as the core input feature driving the adaptive weight network, guiding the multi-view value network evaluation to achieve differentiated fusion across different time scales, enabling the reinforcement learning strategy to adaptively adjust time preferences at different reservoir development stages. Specifically, the economic state vector includes long-term return indicators and short-term return trend indicators; the long-term return indicators characterize the overall economic return level of the project from its inception to the current time step, while the short-term return trend indicators characterize the recent trend of economic returns; the economic state vector is specifically defined as follows:
[0116] (6);
[0117] in, This is the economic state vector at the current time step; This is a long-term return indicator for the current time step. This is a short-term return trend indicator for the current time step.
[0118] The calculation formula is as follows:
[0119] (7);
[0120] in, This represents the final value of positive cash flow up to the current time step; This represents the present value of negative cash flows up to the current time step; This represents the cumulative number of running days at the current time step.
[0121] , The calculation formulas are as follows:
[0122] (8);
[0123] (9);
[0124] in, For the first The immediate economic benefit value corresponding to each time step is used to represent the cash flow generated by reservoir development at that time step. Indicates positive cash flow. Indicates cost expenditure; The discount rate; For initial investment; For the first The cumulative number of operating days for each time step;
[0125] The calculation formula is:
[0126] (10);
[0127] in, This represents the short-term exponential moving average of the immediate reward at the current time step; This represents the long-term exponential moving average of the immediate reward at the current time step; To prevent division by zero of small constants;
[0128] , The calculation formulas are as follows:
[0129] (11);
[0130] (12);
[0131] in, This represents the smoothing coefficient of the EMA (Exponential Moving Average) for short periods. Indicates the window length for short periods; The EMA smoothing coefficient represents the long-period smoothing factor. Indicates the length of the long-period window; For the previous time step; , These represent the short-term exponential moving average and the long-term exponential moving average of the immediate reward at the previous time step, respectively.
[0132] Step 2.2: Perform real-time economic evaluation using the updated economic state vector at each time step to assess the reservoir's development potential and economic life stage in real time. First, determine whether the economic state vector at the current time step meets the preset economic limit truncation condition. If it does, it is determined that the current reservoir development process has entered an economically unsustainable stage and no longer has good development potential. The current round is terminated, and an optimization cycle suitable for the current reservoir characteristics is determined. If it does not meet the condition, the economic state vector at the current time step is retained and passed to subsequent steps as input features to drive the adaptive weight network for multi-view value fusion. The economic limit truncation condition is constructed based on the joint judgment of long-term return indicators and short-term return trend indicators. Specifically, the economic limit truncation condition is met when the long-term return indicator in the economic state vector is lower than the preset economic threshold, and the short-term return trend indicator shows a continuous deterioration trend within a preset number of consecutive time steps.
[0133] Step 3: Construct an experience replay mechanism that incorporates both physical and economic information, namely, a physical state vector containing the current time step. The control actions generated and executed by the policy network at the current time step. Instant rewards for the current time step The physical state vector of the next time step and the economic state vector at the current time step Transfer data Store in the experience replay pool.
[0134] Step 4: Sample data in batches from the experience replay pool, use a set of different preset discount factors, estimate the value function at different time scales in parallel through the multi-view value network, generate the multi-view target Q value, and update the parameters of the multi-view value network based on the corresponding loss function to achieve value consistency learning at different time views.
[0135] The multi-view value network is a dual-value network structure containing two sets of target value networks, each consisting of a target network and a main network. The main principle of the multi-view value network is as follows: upon input of the physical state vector and actions, the main network in each target value network calculates an estimated Q-value, and the target network calculates a target Q-value. The two estimated Q-values and the target Q-value are then updated using a loss function. When calculating the target Q-value, the network with the smallest Q-value between the two target value networks is selected. The calculation process of the multi-view value network is as follows:
[0136] (13);
[0137] in, Indicates the training of the first The target Q value of the visual action value network; For the first Discount factor corresponding to each field of view; This marks the completion of the task; Indicates the sequence number of the two sets of target value networks; This means that, under the corresponding viewpoint, the minimum value is selected from the target Q values output by the two target value networks as the target Q value estimate under that viewpoint; This indicates that under the corresponding field of view, by the first The action value function obtained by applying the target value network; For the first Set target value network Network parameters; This is the normalized physical state vector for the next time step; In order to be in The following is from the current policy network The actions obtained from sampling; For parameters The policy network is represented to describe the probability distribution of control actions in a given state; This is the entropy regularization coefficient;
[0138] loss function Specifically, it is set to minimize the mean square error between the estimated Q value and the target Q value based on the sampled data, as shown in the formula:
[0139] (14);
[0140] in, For the first The parameters of the target value network corresponding to each viewpoint; For mathematical expectation operations; The control action generated and executed by the policy network at the current time step; This serves as an experience replay pool, used to store state, action, reward, and state transition samples collected during reinforcement learning. For the first Each field of view (corresponding discount factor) Action value function under ( ); For the first Action value function under a single viewpoint Parameters; For input and Then through the action value function The estimated Q value was obtained;
[0141] The loss function is minimized using gradient descent to update the parameters of the multi-view value network.
[0142] Step 5: Construct a dynamic weighted architecture driven by economic status, achieve weighted fusion of multi-view value assessment results based on an adaptive weighted network, and construct an optimization objective function for the policy network based on the weighted fusion results, updating the policy network parameters; the specific process is as follows:
[0143] Step 5.1: Sample the physical state vector Input the multi-view value network to obtain the current estimated Q-value evaluation results. The estimated Q-value evaluation result is a set of estimated Q-values corresponding to multiple fields of view, that is, each component in the set corresponds to an estimated Q-value under a field of view.
[0144] Step 5.2: Sample the economic state vector Input the adaptive weight network to generate dynamic weights corresponding to each discount factor;
[0145] The adaptive weighted network uses Softmax normalization to generate dynamic weights based on the economic state vector, satisfying the following constraints:
[0146] (15);
[0147] in, For the first Dynamic weights corresponding to each field of view; It is a natural constant; , The adaptive weight network is for the first The, the The raw log odds of each field of view output; To set the total number of multi-views, ensure that the sum of all weights is 1 and non-negative.
[0148] Step 5.3: Utilize dynamic weights Weighted combination of Q-value evaluation results from multi-view estimation The optimization objective function of the policy network is constructed to achieve adaptive time preference adjustment based on economic state; the optimization objective function The following relationship must be satisfied:
[0149] (16);
[0150] in, For policy networks Parameters;
[0151] By updating the policy network parameters using an optimization objective function, adaptive time preference adjustment based on economic state awareness can be achieved.
[0152] Step 6: Repeat steps 1 to 5 until the training termination condition is met and the optimal strategy network is output. Specifically, when the economic net present value fluctuation of a preset number of training rounds is less than a set threshold, or when the preset maximum number of training iterations is reached, the training termination condition is met. After training terminates, the model is saved, specifically the parameters of the optimal strategy network, the parameters of the multi-view value network, and the parameters of the adaptive weight network, forming a complete optimization model. Finally, the optimal model is deployed to the target reservoir to be optimized. By collecting the original physical state vector and economic state vector of the target reservoir in real time, the corresponding optimal injection and production control action sequence is directly generated by inputting it into the optimal strategy network. This can guide the on-site production of the target reservoir throughout its entire life cycle without retraining.
[0153] Compared with existing reinforcement learning injection and acquisition optimization methods, this invention does not only improve policy performance under a given optimization cycle and fixed field of view, but also enables the optimization process to have three key capabilities by introducing an economic state awareness mechanism:
[0154] Development cycle adaptive capability: It can dynamically determine the optimal timing for termination of optimization based on the real-time economic status of the reservoir, avoiding the continued execution of ineffective or negative-return production decisions during economically infeasible phases;
[0155] Time preference adaptive capability: It can automatically adjust the weight allocation between short-term benefits and long-term value according to the development stage of the reservoir, without the need to manually select a fixed discount factor;
[0156] Cross-scenario generalization capability: Enables the same trained strategy model to adapt to reservoir scenarios with different permeability structures, different water-cutting evolution characteristics and different development paces, reducing dependence on specific reservoir models.
[0157] Through the synergistic effect of the above capabilities, this invention realizes the transformation of reinforcement learning injection and production optimization from "parameter adjustment optimization for a single reservoir" to "general intelligent decision-making for reservoirs in multiple scenarios".
[0158] The effectiveness and superiority of the method of this invention are further demonstrated through numerical simulation experiments using the Egg Model, a standard benchmark model commonly used in the international petroleum engineering field:
[0159] The reservoir model based on this invention is a typical three-dimensional heterogeneous channel sandstone reservoir model, which can fully represent the complex nonlinear characteristics of underground fluid flow. The model has a grid dimension of 60×60×7, with a total of 25,200 grids, simulating the geological heterogeneity of alternating high-permeability channels and low-permeability backgrounds, thus fully representing the complex nonlinear characteristics of underground fluid flow. The model well network includes 8 injection wells and 4 production wells. To verify the generalization ability of the method of this invention, 10 geological scenario cases with different permeability distributions, different initial water saturation, and different formation energy characteristics were constructed experimentally. In this invention, the overall decision-making unit, including the policy network, multi-view value network, and their update mechanism used to perform reinforcement learning decision-making and environmental interaction, is collectively referred to as the agent; at each time step, the agent needs to simultaneously optimize the bottomhole flowing pressure of the 4 production wells and the water injection rate of the 8 injection wells, with a single-step action dimension of 12 dimensions.
[0160] Unlike traditional optimization methods that use a fixed number of decision steps, this embodiment employs a dynamic termination mechanism driven by economic state. A maximum allowed number of decision steps is set as a safety limit, but in actual operation, the optimization step size for each round is autonomously determined by the agent based on the economic state vector. To construct the economic state vector, this embodiment sets the following economic parameters: crude oil price at $20 / barrel, water injection cost at $3 / barrel, and water treatment cost at $1 / barrel. When the agent perceives that the current geological model's output benefit is lower than the preset economic threshold, a dynamic truncation decision is triggered, automatically terminating the current round of production simulation. This means that for models with high permeability and rapid water emergence, the agent may terminate optimization within a shorter time step to lock in profits; while for low permeability models, the agent may execute a longer time step to achieve full utilization, realizing adaptive optimization time for different reservoir models.
[0161] This example uses the commercial numerical simulator Eclipse as the reservoir numerical simulation environment and interacts with the reinforcement learning agent through a Python interface. The experiment is set to a total of 1500 training rounds, covering the 10 different reservoir geological cases mentioned above. During the training process, this invention utilizes a multi-view-economic state framework, enabling the agent to automatically adjust the temporal view weights of value assessment when facing different geological models in different rounds, achieving cross-scenario policy generalization learning under a unified training framework.
[0162] Based on the experimental setup of the Egg model described above, the specific implementation steps of the multi-scenario reservoir generalized injection-production optimization method based on economic state perception described in this invention are as follows:
[0163] (1) A numerical simulation model was established based on Egg reservoir geological data. The production control variables were set as the bottomhole flowing pressure of 4 production wells and the water injection rate of 8 injection wells, totaling 12 control actions. The maximum number of training rounds was set to 1500. A reinforcement learning algorithm framework was constructed, and the network parameters were initialized. A set of multi-view discount factors was defined. These correspond to four time preferences: extremely short-sighted, short-sighted, medium-sighted, and far-sighted, covering different optimization needs from prioritizing immediate gains to prioritizing long-term cumulative returns.
[0164] (2) At each time step, read the current three-dimensional physical field data of the reservoir from the reservoir numerical simulator, including the pressure field and water saturation field, normalize and reconstruct the dimensions, and convert the three-dimensional physical field into a one-dimensional physical state vector that can be input into the strategy network.
[0165] (3) Perform parallel generation and security constraints. The input is fed into the policy network to output 12-dimensional candidate control actions. Subsequently, the engineering safety constraint module... Perform component-wise verification and correction. Map the normalized action output of the neural network to the physical control range [300, 400] bar and [0, 80] m. 3 / d performs a safety boundary check. The final output conforms to all engineering red lines in terms of safety control actions. .
[0166] (4) Conduct environmental interaction and economic status updates. The simulation is executed using a reservoir numerical simulator with a time step of 360 days. The immediate reward and the physical state vector for the next time step are obtained. The economic state vector is updated based on the immediate reward. Specifically, long-term return indicators and short-term return trend indicators are calculated; when calculating the long-term return indicator, a discount rate of 10% is set, and the calculation is based on... =0 to the current time step, the adjusted internal rate of return; when calculating the short-term return trend indicator, calculate the short-term EMA smoothing coefficient and the long-term EMA smoothing coefficient of the immediate reward, and find their normalized difference.
[0167] Determine if the economic state vector meets the termination condition. If the long-term return indicator is below the preset economic threshold, or the short-term return trend indicator shows a continuous deterioration, then the current reservoir model is determined to have reached its economic exploitation life limit, and the current round is immediately terminated. The optimal optimization step size suitable for the current reservoir type is dynamically determined. If not triggered, then... Store in the experience replay pool.
[0168] (5) Training the multi-view value network. Randomly sample batch data from the experience replay pool, with a batch size of 256. The multi-view value network contains four parallel value evaluation heads, each corresponding to one of the four discount factors: For each field of view, the corresponding discount factor... The target Q value is calculated using formula (13); and the mean square error between the predicted Q value and the target Q value is minimized according to formula (14). The parameters of the four value assessment heads are updated in parallel so that they can learn to assess the expected returns at different time scales.
[0169] (6) Training the policy network. The policy network parameters are updated using a decoupling framework: the sampled physical state vector is input into the multi-view value network to obtain four sets of Q-value scores. The sampled economic state vector is input into the adaptive weight network, and four dynamic weights are output via Softmax. The weighted objective is calculated. This is then used as the optimization objective of the policy network for backpropagation updates. For the first The target Q-value evaluation result corresponds to each field of view. This process enables the agent to learn to adaptively adjust its strategy. For example, when When the short-term returns show a downward trend (such as during the high water content period or the late stage of development), the adaptive weight network will automatically increase the weight corresponding to the low discount factor, forcing the strategy to prioritize maximizing immediate cash-in returns and preventing ineffective inflation caused by excessive pursuit of long-term inflated returns; conversely, in the early stage of development, the weight corresponding to the high discount factor will be increased to focus on energy preservation.
[0170] (7) Iterative Loop and Model Deployment. Determine whether the current round has ended or whether the maximum number of training rounds has been reached. If not, return to step (2) to continue the decision for the next time step; if training is complete, save the optimal strategy network parameters. Then proceed to the generalization deployment stage: deploy the optimal strategy network model to the new reservoir to be optimized, input the real-time physical and economic states of the target reservoir, and use the strategy network to directly generate the optimal injection and production control sequence end-to-end, which can guide on-site production without retraining.
[0171] Figure 3 This diagram illustrates the convergence curve of the economic net present value (NPV) iterative optimization of the multi-scenario reservoir generalized injection-production optimization method based on economic state perception during training on a typical reservoir test case, as described in this embodiment of the invention. The horizontal axis represents the number of training rounds, and the vertical axis represents the NPV corresponding to the current optimal injection-production strategy. The curve shows a significant "step-like rise" and rapid convergence during training: in the early stages, thanks to the timely truncation of ineffective optimization steps by the economic state vector and the adaptive guidance of multi-view value weights, the agent quickly abandons inefficient strategies, leading to a significant leap in NPV. Subsequently, the curve reaches its peak around 100 rounds and remains highly stable, indicating that the model has successfully found a globally optimal strategy that balances long-term and short-term benefits under specific geological conditions. This demonstrates that the method of this invention possesses extremely high learning efficiency and convergence stability when facing specific reservoir models, and can rapidly evolve high-value production control schemes from scratch.
[0172] Figure 4 and Figure 5 This is a schematic diagram illustrating the adaptive change of the time-view weights corresponding to different discount factors in a single round of training for the multi-scenario reservoir generalized injection-production optimization method based on economic state awareness, as described in this embodiment of the invention. Figure 4 Corresponding discount factor = 0.90, Figure 5 Corresponding discount factor = 0.99.
[0173] by Figure 4 For example, in the 0–300 round phase, the discount factor The weight corresponding to 0.90 remained relatively stable across all time steps, maintaining a low and stable level. In the 300–800 round phase, this weight gradually increased in the later time steps, indicating that the agent began to focus more on short-term returns in the later stages of development. In the 800–1500 round phase, the discount factor… The weight of 0.90 remains low at the beginning of the round and increases significantly as the time step progresses, reflecting that the agent has learned to dynamically adjust its time vision preference according to the development progress within a single round.
[0174] Figure 5 Discount factor shown = 0.99 weight change and Figure 4 They exhibit a complementary relationship. In different rounds, the weight corresponding to this discount factor is higher in the initial time steps and gradually decreases in the later stages of the round, thus working with the short-sighted discount factor to achieve adaptive adjustment of time preference as the development stage changes. The time horizons corresponding to the other discount factors change less significantly during training and mainly play a supporting transitional role, therefore they are not shown separately in the figure.
[0175] Figure 6 The diagram shows the convergence process of the net present value (NPV) as the differential evolution algorithm optimizes injection and production parameters on a target reservoir model that has not been trained. It can be seen that this algorithm requires repeated numerical simulations and iterative searches to gradually approach the optimal solution. In contrast, deploying the policy network trained by the method of this invention directly onto the target reservoir model, without any retraining or parameter adjustments, generates the injection and production control strategy through only one forward inference, achieving an NPV of 6.34 × 10⁻⁶. 6 USD is superior to USD under the conditions of this embodiment. Figure 6 The final convergence level of the differential evolution algorithm shown indicates that the method of this invention can directly obtain better economic optimization results on reservoir models that have not participated in training, and has good cross-scenario generalization ability. To ensure the fairness of the above comparison results, in this comparative experiment, the population size of the differential evolution algorithm was set to 50, the crossover probability CR was set to 0.5, the scaling factor F was set to 0.5, and the maximum number of function evaluations was set to 1500. This setting is comparable to the number of environmental interactions experienced by the method of this invention during training.
[0176] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A multi-scenario reservoir generalized injection-production optimization method based on economic status perception, characterized in that, Includes the following steps: Step 1: Based on the physical constraints, a reinforcement learning interactive environment is used to run the reinforcement learning interactive loop in the reservoir numerical simulation environment. Step 2: Update the economic state vector and use the economic state vector to perform real-time economic evaluation; Step 3: Construct an experience replay mechanism that incorporates both physical and economic information; the content of the experience replay mechanism is: to generate a physical state vector containing the current time step. The control actions generated and executed by the policy network at the current time step. Instant rewards for the current time step The physical state vector of the next time step and the economic state vector at the current time step Transfer data Store in the experience replay pool; Step 4: Generate multi-view target Q-values through the multi-view value network, and update the parameters of the multi-view value network based on the loss function; The multi-view value network comprises two sets of target value networks, and the calculation process is as follows: (13); in, Indicates the use of training the first The target Q value of the visual action value network; For the first Discount factor corresponding to each field of view; This marks the completion of the task; Indicates the sequence number of the two sets of target value networks; This means that, under the corresponding viewpoint, the minimum value is selected from the target Q values output by the two target value networks as the target Q value estimate under that viewpoint; This indicates that under the corresponding field of view, by the first The action value function obtained by applying the target value network; For the first Set target value network Network parameters; This is the normalized physical state vector for the next time step; In order to be in The following is from the current policy network The actions obtained from sampling; For parameters Representational policy network; This is the entropy regularization coefficient; loss function Specifically, it is set to minimize the mean square error between the estimated Q value and the target Q value based on the sampled data, as shown in the formula: (14); in, For the first The parameters of the target value network corresponding to each viewpoint; For mathematical expectation calculation; For experience replay pool; For the first Action value function under a single perspective; For the first Action value function under a single viewpoint Parameters; For input and Then through the action value function The estimated Q value was obtained; Step 5: Minimize the loss function using gradient descent to update the parameters of the multi-view value network; Step 6: Construct an economic state-driven dynamic weight architecture, implement weighted fusion of multi-view value assessment results based on the adaptive weight network, and construct the optimization objective function of the policy network based on the weighted fusion results to update the policy network parameters. Step 6: Repeat steps 1 to 5 until the training termination condition is met, and output the optimal strategy network; deploy the optimal strategy network on the target reservoir to be optimized, and directly generate the corresponding optimal injection and production control action sequence by collecting the original physical state vector and economic state vector of the target reservoir in real time and inputting them into the optimal strategy network.
2. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception as described in claim 1, characterized in that, In step 1, the reinforcement learning interaction loop performs the following operations at each time step: Step 1.1: Obtain the real-time raw physical state vector of the reservoir and input it into the strategy network to generate candidate control actions. ;definition For the current time step, This is the original physical state vector at the current time step; the physical state vector is used as the input to the state representation in normalized form within the policy network, defined as follows. This is the normalized physical state vector for the current time step. Step 1.2: Recover the physical space characteristics using the state decoding module; Step 1.3: Combine the engineering safety constraint module to perform safety checks and corrections on the candidate control actions to obtain the final safe control actions to be executed; Step 1.4: Input the final safety control action to be executed in the current time step into the reservoir simulator, and output the physical state vector of the next time step returned by the reservoir numerical simulation environment and the instantaneous reward of the current time step.
3. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception as described in claim 2, characterized in that, In step 1.2, the specific working process of the state decoding module is as follows: Step 1.2.1: First, perform inverse normalization using the pre-stored mean. and standard deviation Normalize the physical state vector at the current time step Restored to the corresponding actual physical parameter scale: (1); in, Represents element-wise product; This is the physical state vector at the current time step; Step 1.2.2: Next, perform spatial reconstruction processing to reconstruct the physical state vector into a form with spatial dimensions. 3D reservoir physical property mesh: (2); in, This is the 3D reservoir physical property mesh for the current time step; , , These represent the number of grid cells in the reservoir model along the x, y, and z directions, respectively. This is a dimension reconstruction function.
4. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception as described in claim 3, characterized in that, In step 1.3, the specific working process of the engineering safety constraint module is as follows: Step 1.3.1: Determine the safety space; based on the three-dimensional reservoir physical property mesh and a set of preset engineering rules, determine the real-time safety action space boundary. : (3); in, This is the injection control action vector; For the first The constraint function of the project; This represents the total number of preset engineering conditions; Step 1.3.2: Perform security checks on candidate control actions and generate the final safe control action to be executed at the current moment; the security check process is as follows: judge the candidate control actions generated by the policy network. Does it exceed the boundary of the safe movement space? ; When candidate control action Exceeding the boundaries of the safe operating space At that time, candidate control actions are corrected; specifically, the components of the current candidate control actions that exceed the boundary of the safety action space are projected or truncated to the nearest safety compliance value defined by the boundary of the safety action space, so as to generate the safety control action to be finally executed at the current time step. : (4); in, For the first The final safety control action executed corresponding to each control variable; It is an interval cutoff function; The corresponding first candidate control action generated by the policy network Each action component; , These represent the upper and lower safety limits of the action components calculated based on the three-dimensional reservoir physical property grid; When candidate control action Not exceeding the safe operating space boundary When the current candidate control action is selected, it is directly determined as the final safety control action to be executed in the current time step. .
5. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception according to claim 4, characterized in that, In step 1.4, the immediate reward combines the economic benefits and costs of reservoir development: (5); in, The immediate reward for the current time step; This indicates the price per unit of crude oil; This indicates the increase in oil production within the current time step; This indicates the unit cost of water injection; This indicates the increase in water injection volume within the current time step; This indicates the unit cost of wastewater treatment. This indicates the increase in water production within the current time step.
6. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception as described in claim 5, characterized in that, The specific process of step 2 is as follows: Step 2.1: Update the economic state vector at each time step based on immediate rewards; the economic state vector specifically includes long-term return indicators and short-term return trend indicators: (6); in, This is the economic state vector at the current time step; This is a long-term return indicator for the current time step. This is a short-term return trend indicator for the current time step. The calculation formula is as follows: (7); in, This represents the final value of positive cash flow up to the current time step; This represents the present value of negative cash flows up to the current time step; This represents the cumulative number of running days at the current time step. , The calculation formulas are as follows: (8); (9); in, For the first The instantaneous economic benefit value corresponding to each time step; The discount rate; For initial investment; For the first The cumulative number of operating days for each time step; The calculation formula is: (10); in, , These represent the short-term and long-term exponential moving averages of the instantaneous reward at the current time step, respectively. To prevent division by zero of small constants; , The calculation formulas are as follows: (11); (12); in, The smoothing coefficient for a short-term exponential moving average. The smoothing coefficient of an exponential moving average over a long period; For the previous time step; , These represent the short-term and long-term exponential moving averages of the instantaneous reward at the previous time step, respectively. Step 2.2: Perform real-time economic evaluation using the updated economic state vector at each time step to assess the reservoir's development potential and economic life stage in real time. First, determine whether the economic state vector at the current time step meets the preset economic limit cutoff condition. If it does, it is determined that the current reservoir development process has entered an economically unsustainable stage and no longer has development potential, and the current round is terminated. If it does not meet the condition, the economic state vector at the current time step is retained. The economic limit cutoff condition is: when the long-term return index in the economic state vector is lower than the preset economic threshold, and the short-term return trend index shows a continuous deterioration trend within a preset number of consecutive time steps, the economic limit cutoff condition is met.
7. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception as described in claim 6, characterized in that, The specific process of step 5 is as follows: Step 5.1: Sample the physical state vector Input the multi-view value network to obtain the current estimated Q-value evaluation results. The estimated Q-value evaluation result is a set of estimated Q-values corresponding to multiple fields of view, where each component in the set corresponds to an estimated Q-value under a single field of view. Step 5.2: Sample the economic state vector Input the adaptive weight network to generate dynamic weights corresponding to each discount factor: (15); in, For the first Dynamic weights corresponding to each field of view; It is a natural constant; , The adaptive weight network is for the first The, the The raw log odds of each field of view output; The total number of multi-view fields is set; Step 5.3: Use dynamic weights to weight and combine the Q-value evaluation results of the multi-view estimation to construct the optimization objective function of the policy network. : (16); in, For policy networks Parameters; By using an optimization objective function to update the policy network parameters, adaptive time preference adjustment based on economic state awareness is achieved.
8. The multi-scenario reservoir generalized injection-production optimization method based on economic state perception according to claim 7, characterized in that, In step 6, the training termination condition is: when the fluctuation range of the net present value of the economy for a consecutive preset number of training rounds is less than a set threshold, or when the preset maximum number of training iterations is reached, it is determined that the training termination condition is met.