Horizontal well regulation and control reinforcement learning method based on evolutionary algorithm assistance
By applying hybrid deep reinforcement learning intelligent body model assisted by evolutionary algorithms in oil well regulation, the problems of low efficiency and difficulty in responding to reservoir changes in traditional regulation methods are solved, and the maximum recovery of oil and gas resources and the improvement of oil field mining efficiency are achieved.
Patent Information
- Application Number
- CN202510615970.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Traditional oil well regulation methods rely on manual experience and manual regulation, are inefficient and difficult to quickly respond to dynamic changes in reservoirs. Especially under complex and unpredictable reservoir conditions, it is difficult to achieve optimal production regulation.
A hybrid deep reinforcement learning intelligent model assisted by evolutionary algorithm is adopted. Through evolutionary algorithms, the training and exploration process of reinforcement learning is optimized, and the hybrid action space is built, and the valve opening of each well section and the production system of each well are adjusted in real time.
It accelerates the learning process, improves the global search capability of the model, realizes the maximization of oil and gas resources, can quickly adapt to the dynamic changes of the reservoir, and improves the mining efficiency and economic benefits of the oil field.
Smart Images

Figure CN120124501A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of reservoir production optimization, and particularly relates to a horizontal well control reinforcement learning method assisted by an evolutionary algorithm. Background Technique
[0002] During the development of oil and gas fields, horizontal well technology is widely used in oil and gas production under complex geological conditions. However, due to the heterogeneity and multi-layer structure of the reservoir, there are significant differences in the liquid production and fluid flow characteristics of each well section during the production process. In order to improve the oil and gas recovery efficiency, it is necessary to carry out refined segmented management of oil wells and accurately control the valve opening of each well section, so as to ensure the balanced production of oil and gas resources. With the dynamic changes of reservoir conditions, the production system of each well also needs to be optimized and adjusted in real time. Traditional oil well control methods mostly rely on manual experience and manual adjustment. This method not only has low efficiency, but also is difficult to quickly respond to the dynamic changes of the reservoir. In the face of complex and unpredictable reservoir conditions, it is often difficult to achieve optimal production control.
[0003] As a powerful optimization algorithm, reinforcement learning can achieve adaptive adjustment of the production system under the condition of dynamic changes in the reservoir by continuously learning and adjusting strategies. However, although reinforcement learning can optimize production control to a certain extent, its efficiency and effect are limited by the quality of training data and the convergence speed of the algorithm. In practical applications, the training of reinforcement learning models often requires a large number of samples and computing resources, and the exploration ability during the model training process is limited, making it difficult to effectively avoid local optimal solutions. In order to improve the efficiency and global search ability of the reinforcement learning model, an auxiliary optimization scheme combined with an evolutionary algorithm has become an effective choice. The evolutionary algorithm simulates the process of natural selection and searches in a large range through genetic and selection operations, thereby optimizing the decision-making process. The diversity and randomness of the evolutionary algorithm enable it to search globally and avoid the dilemma of reinforcement learning falling into local optimality. In addition, the population characteristics of the evolutionary algorithm can promote the comprehensive exploration of the search space through the evolution and adaptive update of the population, further improving the performance of the reinforcement learning model. Through this method, the valve opening and production system of each well section can be optimized in real time under the dynamic changes of the reservoir, so as to achieve the maximum recovery of oil and gas resources, and has broad application prospects. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a horizontal well control reinforcement learning method assisted by an evolutionary algorithm. This method combines a hybrid deep reinforcement learning agent model with an evolutionary algorithm to optimize the training and exploration process of reinforcement learning through the evolutionary algorithm, thereby accelerating the learning process and improving the global search ability of the model. Through this method, the horizontal well section division can be determined according to the reservoir state, and the valve opening of each section and the production regime of each well can be adjusted in real time to ensure the maximization of oil and gas recovery rate and quickly adapt to different production requirements under dynamically changing reservoir conditions. Finally, the optimal control scheme for each well can be obtained, thus greatly improving the oilfield exploitation efficiency and economic benefits.
[0005] The technical solution of the present invention is as follows: A horizontal well control reinforcement learning method assisted by an evolutionary algorithm, comprising the following steps: Step 1: Determine the positions of the segmentation points, valve openings, and injection-production variables to be optimized, and construct a hybrid action space combining discrete and continuous control measures; Step 2: Construct a hybrid deep reinforcement learning agent model and a prioritized experience replay pool for recent experiences; Step 3: Interact the hybrid deep reinforcement learning agent model with a numerical simulator, and train the agent through the generated experience samples; Step 4: Perform replacement, selection, crossover, and mutation operations of the evolutionary algorithm to optimize the hybrid deep reinforcement learning agent model; Step 5: Repeat Steps 3 to 4 until the maximum number of episodes is reached.
[0006] Further, the specific process of Step 1 is as follows: Step 1.1: Determine the horizontal well segmentation points, valves, and the number of oil and water wells to be controlled according to the target reservoir. The variables to be optimized include the positions of the segmentation points, valve openings, and injection-production variables. The oil and water wells include production wells and injection wells, and the injection-production variables include production rates and injection rates; specifically expressed as: ; In the formula, is the control time step is the variable to be optimized at time step is the position of the th segmentation point at the current time step; is the opening of the th valve at the current time step; is the production rate of the th production well at the current time step; is the injection rate of the th injection well at the current time step; Step 1.2: Construct a discrete action space is: ; Step 1.3. Construct a continuous action space is: ; Step 1.4. The hybrid action space combining discrete and continuous control measures is represented as: ; In the formula, is the hybrid action space.
[0007] Furthermore, in the said Step 2, the hybrid deep reinforcement learning agent model includes a policy network, an action-value network, and a target action-value network; among them, the policy network consists of two parts: a discrete policy network and a continuous policy network; the discrete policy network consists of a convolutional neural network and a fully connected neural network, the input is the reservoir saturation field and pressure field, the data first extracts the reservoir data features through the convolutional neural network, and then outputs the discrete action probability distribution through the fully connected neural network; the continuous policy network consists of a convolutional neural network and a fully connected neural network, the input is the reservoir saturation field, pressure field, and discrete action probability distribution, the reservoir saturation field and pressure field first extract the reservoir data features through the convolutional neural network, and after the extracted reservoir data features are concatenated with the discrete action probability distribution, they are input into the fully connected neural network to output the Gaussian distribution of the continuous action; The network structures of the action-value network and the target action-value network are the same, both consisting of a convolutional neural network and a fully connected neural network; the input of the action-value network is the reservoir saturation field, pressure field, and continuous action at a given time step, and the input of the target action-value network is the reservoir saturation field, pressure field, and continuous action at the next time step of the given time step; the reservoir saturation field and pressure field first extract the reservoir data features through the convolutional neural network, and after the extracted reservoir data features are concatenated with the continuous action, they are input into the fully connected neural network, the action-value network outputs the Q value of the action at the given time step, and the target action-value network outputs the Q value of the action at the next time step of the given time step; Construct a recent experience prioritized replay pool The sampling ratio factor and the actual sampling range are as follows: ; ; In the formula, is the sampling ratio factor; is the current update count; is the total update count; is the hyperparameter for controlling sampling bias; is the actual sampling range; is the number of experiences in the current experience pool.
[0008] Further, in step 3, the agent is an individual for reinforcement learning; the specific process of step 3 is as follows: Step 3.1: Set the reward value at the current time step as the economic net present value at the current time step: ; In the formula, is the current time step; is the reward value at the current time step; is the economic net present value at the current time step; is the number of production wells; is the number of injection wells; is the unit price of crude oil sales; is the treatment cost of production well water; is the injection cost of injection wells; , are respectively the daily oil production and daily water production of the th production well at the current time step; is the th injection rate of the injection well; Step 3.2: Use the saturation field and pressure field data of the reservoir model at the current time step as the state and input it into the policy network. The policy network outputs the discrete action and the continuous action at the current time step; calculate the reward value at the current time step and the state at the next time step through a numerical simulator. If the current time step is the last time step, the end flag is 1, otherwise it is 0; save the trajectory of the experience sample in the recent experience prioritized replay pool ; Step 3.3: Use the experience samples in the recent experience prioritized replay pool to train the action value loss, specifically as follows: ; ; ; In the formula, is the Q value of the target action; is the discount factor; represents the discrete action probability distribution; is to find the minimum value; is the th target action value network; Represents the continuous action for the next time step; Represents the temperature coefficient of the continuous action; Represents the probability distribution of the continuous action; Represents the discrete action for the next time step; Represents the temperature coefficient of the discrete action; Represents the loss function of the th action value network, and the serial number of the action value network corresponds to that of the target action value network; Is the action value network; Represents the total loss function of the action value network; Step 3.4: Use the experience samples in the recent experience prioritized replay pool to train the policy network parameters, specifically as follows: ; ; ; Among them, is the loss function of the discrete policy network; is for calculating the expectation; is the loss function of the continuous policy network; is the total loss function of the policy network; Step 3.5: Design the loss functions for the temperature coefficients of the continuous action and the discrete action, and update and specifically as follows: ; ; Among them, is the loss function of the temperature coefficient of the continuous action; is the loss function of the temperature coefficient of the discrete action; and respectively represent the entropies of the continuous policy network and the discrete policy network; and respectively represent the target entropies of the continuous policy network and the discrete policy network.
[0009] Furthermore, the specific process of step 4 is as follows: Step 4.1: The evolutionary algorithm has a total of Individuals, each individual is a policy network; the sum of all rewards within one round of the policy network is recorded as the initial fitness value of the individual; a penalty term is set. If the continuous actions output by the policy network are equal to the boundary values, the fitness value of the policy network is deducted accordingly; calculate the fitness values of all individuals, and the formula is: ; In the formula, is the fitness value of individual ; is the initial fitness value of individual ; is the penalty term of individual . For each boundary value taken by the individual, the penalty term is incremented by 1; Calculate the minimum tolerance of all individuals: ; In the formula, is the minimum tolerance of individual ; is the fitness value of the previous round of individual ; is the tolerance threshold; If the fitness value of the current individual is the lowest or lower than the minimum tolerance of the individual, then replace the current individual with a reinforcement learning individual; Step 4.2: The individual with the highest fitness value directly enters the winning pool, and then the roulette wheel algorithm is used to select several individuals to enter the winning pool. The probability of selecting an individual with a high fitness value is high, and the probability calculation formula is as follows: ; ; In the formula, is the normalized fitness value of individual ; is the fitness value; is a constant; is the probability of individual being selected; Step 4.3: All individuals who do not enter the winning pool enter the remaining evolution pool, and the three-source crossover operation is performed on all individuals in the remaining evolution pool; this operation performs crossover on the convolutional layer, fully connected layer, and bias term, where the convolution kernel, row, and element are used as the basic units of crossover respectively; during the crossover process of each basic unit, the probability of the target individual retaining itself is , the probability of selecting an individual from the winning pool is , and the probability of selecting an individual from the remaining evolution pool is , satisfying the following formula: ; Step 4.4. Perform mutation operations on all individuals in the remaining evolution pool. Mutation means adjusting the values of the parameters of an individual by adding noise based on a standard normal distribution to adjust its value, for the individual the value of the parameter participating in mutation, is the variance; each individual triggers a mutation event according to the mutation probability When a mutation is triggered, for the multi-dimensional parameter vector of the individual, according to the ratio of the parameters of the preset normal mutation and the ratio of the parameters of the super-strong mutation select parameters to perform normal mutation and parameters to perform super-strong mutation respectively by sampling without replacement; where is the number of parameters of the individual , is rounding up; define the variance parameter of the normal mutation is less than the variance parameter of the super-strong mutation , and satisfy the following formula: .
[0010] Furthermore, the specific process of Step 5 is as follows: Completing one round of Step 3 and Step 4 is called a round. In each round, Step 3 is executed times, and Step 4 is executed once; after Step 4 ends, the individual with the highest fitness replaces the reinforcement learning individual in Step 3, and the tolerance threshold and the mutation probability change with the number of rounds. The formula is: ; ; In the formula, is the minimum value of the tolerance threshold; is the maximum value of the tolerance threshold; is the minimum value of the mutation probability; is the maximum value of the mutation probability; is the current number of rounds; is the maximum number of rounds; Repeat Step 3.1 to Step 4.4, and execute a total of rounds to finally obtain the optimal horizontal well section point position, valve opening and optimal injection-production system plan.
[0011] Beneficial technical effects brought by the present invention: The present invention proposes a reinforcement learning method assisted by an evolutionary algorithm for the refined regulation and production system optimization of horizontal wells. By combining the evolutionary algorithm with a hybrid deep reinforcement learning agent model, the present invention can effectively optimize the well section division and valve opening of horizontal wells, adjust the production system of each well in real time, thereby improving the oil and gas recovery rate and quickly adapting to the dynamic changes of the reservoir. This method can accelerate the learning process, enhance the global search ability, and improve the economic benefits of oilfield exploitation through automated regulation, with significant technical advantages and wide application value for popularization. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the design flow chart of the horizontal well regulation reinforcement learning method assisted by an evolutionary algorithm according to the present invention.
[0013] Figure 2 is the comparison diagram of the economic net present value convergence curves between the horizontal well regulation reinforcement learning method assisted by an evolutionary algorithm according to the present invention and the theoretical maximum economic net present value of the horizontal well without sectioning but only adjusting the production system.
[0014] Figure 3 is the comparison diagram of the cumulative oil production between the horizontal well regulation reinforcement learning method assisted by an evolutionary algorithm according to the present invention and the theoretical maximum economic net present value of the horizontal well without sectioning but only adjusting the production system.
[0015] Figure 4 is the comparison diagram of the cumulative water production between the horizontal well regulation reinforcement learning method assisted by an evolutionary algorithm according to the present invention and the theoretical maximum economic net present value of the horizontal well without sectioning but only adjusting the production system.
[0016] Figure 5 is the oil saturation and well location map of the reservoir after optimization according to the present invention.
[0017] Figure 6 is the schematic diagram of the optimized sectional point position of production well P1 according to the present invention.
[0018] Figure 7 is the schematic diagram of the optimized sectional point position of production well P2 according to the present invention.
[0019] Figure 8 is the schematic diagram of the optimized sectional point position of production well P3 according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments: Taking a certain reservoir model as an example, the size of this model is 25*25*3, the porosity is 0.2, the initial pressure is 6000 psi, and the initial water saturation is 0.2. There are a total of three production wells (numbered P1 - P3) and four injection wells (numbered I1 - I4), and all of the production wells are horizontal wells. The possible number of segment points for production wells P1, P2, and P3 are 9, 10, and 9 respectively, and the maximum number of segments for each well are 3, 4, and 3 respectively. The upper limit of the valve opening is 100%, and the lower limit is 0%. The upper limit of the production rate of the production wells is 178.61 m 3 / d, and the lower limit is 0 m 3 / d. The injection rate of the injection wells is 178.61 m 3 / d.
[0021] As Figure 1 shown in the horizontal well control reinforcement learning method assisted by evolutionary algorithm, it specifically includes the following steps: Step 1. Determine the segment point positions, valve openings, and injection-production variables to be optimized, and construct a hybrid action space that combines discrete and continuous control measures; the specific process is as follows: Step 1.1. According to the target reservoir, determine the horizontal well segment points, valves, and the number of oil and water wells (including production wells and injection wells) to be controlled. The variables to be optimized include segment point positions, valve openings, and injection-production variables (including production rates and injection rates), which can be specifically expressed as: ; In the formula, is the variable to be optimized at the control time step ; the number of all possible segment points is 28; the number of valves is 10; the number of optimized production wells is 3; the number of optimized injection wells is 0; is the control time step, and each control time step is an optimization period, with a total of 5 control time steps; is the position of the th segment point at the current time step, taking values 0 or 1, which is only valid for the first time step; is the opening of the th valve at the current time step, %; is the production rate of the th production well at the current time step, m 3 / d; is the injection rate of the th injection well at the current time step, 178.61 m 3 / d.
[0022] Step 1.2. Construct a discrete action space is: ; Set to represent the dimension of the discrete action space, with a total of dimensions; Step 1.3, construct the continuous action space is: ; Set to represent the dimension of the continuous action space, with a total of dimensions; Step 1.4, the hybrid action space combining discrete and continuous control measures can be expressed as: ; In the formula, is the hybrid action space.
[0023] Step 2, construct the hybrid deep reinforcement learning agent model and the prioritized experience replay pool for recent experiences, which specifically include the following steps: The hybrid deep reinforcement learning agent model includes a policy network, an action-value network, and a target action-value network; among them, the policy network consists of two parts: a discrete policy network and a continuous policy network. The discrete policy network consists of a convolutional neural network and a fully connected neural network. The input is the reservoir saturation field and pressure field. The data first extracts the reservoir data features through the convolutional neural network, and then outputs the discrete action probability distribution through the fully connected neural network; the continuous policy network consists of a convolutional neural network and a fully connected neural network. The input is the reservoir saturation field, pressure field, and discrete action probability distribution. The reservoir saturation field and pressure field first extract the reservoir data features through the convolutional neural network, and after the extracted reservoir data features are concatenated with the discrete action probability distribution, they are input into the fully connected neural network to output the Gaussian distribution of the continuous action; The network structures of the action-value network and the target action-value network are the same, and both consist of a convolutional neural network and a fully connected neural network. The input of the action-value network is the reservoir saturation field, pressure field, and continuous action at a given time step. The input of the target action-value network is the reservoir saturation field, pressure field, and continuous action at the next time step of a given time step. The reservoir saturation field and pressure field first extract the reservoir data features through the convolutional neural network, and after the extracted reservoir data features are concatenated with the continuous action, they are input into the fully connected neural network. The action-value network outputs the Q value of the action at a given time step, and the target action-value network outputs the Q value of the action at the next time step of a given time step; Construct the prioritized experience replay pool for recent experiences , which is used to store the experience samples generated during the agent's exploration; the sampling ratio factor and the actual sampling range of the prioritized experience replay pool for recent experiences are as follows: ; ; In the formula, is the sampling scale factor; is the current number of updates; is the total number of updates, which is 50,000 times; is the hyperparameter for controlling the sampling bias, taking 0.4; is the actual sampling range; is the number of experiences in the current experience pool. The sampling range in the subsequent step 3 is the latest pieces of data.
[0024] Step 3: Interact the hybrid deep reinforcement learning agent model with the numerical simulator, and train the agent through the generated experience samples. In the present invention, the agent is a reinforcement learning individual. Specifically, it includes the following steps: Step 3.1: Set the reward value at the current time step as the economic net present value at the current time step: ; In the formula, is the current time step; is the reward value at the current time step; is the economic net present value at the current time step; the number of production wells is 3; the number of injection wells is 4; is the unit price of crude oil sales, 3522 yuan / m 3 ; is the treatment cost of production well water, 220 yuan / m 3 ; is the injection cost of injection wells, 132 yuan / m 3 ; , are respectively the daily oil production and daily water production of the th production well at the current time step, m 3 / d; is the injection rate of the th injection well, 178.61 m 3 / d.
[0025] Step 3.2: Input the saturation field and pressure field data of the reservoir model at the current time step as the state into the policy network. The policy network outputs the discrete action at the current time step and the continuous action at the current time step; calculate the reward value at the current time step and the state at the next time step through the numerical simulator, if the current time step is the last time step, the end flag is 1, otherwise 0. Save the trajectory of the experience sample in the recent experience prioritized replay pool .
[0026] Step 3.3: Use the experience samples in the recent experience prioritized replay pool to train the action value loss, which is specifically as follows: ; ; ; In the formula, is the Q value of the target action; is the discount factor, indicating the relative importance of the current reward and future rewards, with a value of 0.99; represents the discrete action probability distribution; represents the state at the next time step; is to find the minimum value; is the th target action value network. In the present invention, two target action value networks and two action value networks are set. The first target action value network corresponds to the first action value network, and the second target action value network corresponds to the second action value network. Therefore, here ; represents the continuous action at the next time step; represents the temperature coefficient of the continuous action; represents the continuous action probability distribution; represents the discrete action at the next time step; represents the temperature coefficient of the discrete action; represents the th loss function of the action value network; represents calculating the expectation; is the action value network; represents the total loss function of the action value network; Step 3.4: Use the experience samples in the recent experience prioritized replay pool to train the policy network parameters, which is specifically as follows: ; ; ; Among them, is the loss function of the discrete policy network; is to calculate expectation; is the loss function of the continuous policy network; is the total loss function of the policy network; Step 3.5: Design the loss functions for the temperature coefficients of continuous actions and discrete actions, and update the temperature coefficients of continuous actions and discrete actions and as follows: ; ; where is the loss function of the temperature coefficient of continuous actions; is the loss function of the temperature coefficient of discrete actions; and represent the entropies of the continuous policy network and the discrete policy network respectively; and represent the target entropies of the continuous policy network and the discrete policy network respectively, with values of -13 and 0.416 respectively.
[0027] Step 4: Execute the replacement, selection, crossover, and mutation operations of the evolutionary algorithm to optimize the hybrid deep reinforcement learning agent model. Specifically, it includes the following steps: Step 4.1: The evolutionary algorithm has individuals, which is 19, and each individual is a policy network. The sum of all rewards within one episode of the policy network is recorded as the initial fitness value of the individual. Set a penalty term. If the continuous action output by the policy network is exactly equal to the boundary value, it is considered that the network has a risk of damaged network structure, and the fitness value of the network is deducted accordingly. Calculate the fitness values of all individuals, and the calculation formula is: ; In the formula, is the fitness value of individual ; is the initial fitness value of individual ; is the penalty term of individual . For each boundary value taken by the individual, the penalty term is incremented by 1, and the initial value is 0.
[0028] Calculate the minimum tolerance of all individuals: ; In the formula, is the minimum tolerance of individual ; is the previous round's fitness value of individual ; is the tolerance threshold, and the value range is ; If the fitness value of the current individual is the lowest or lower than the individual's minimum tolerance, it is replaced by a reinforcement learning individual; Step 4.2: The individual with the highest fitness value directly enters the winning pool, and then the roulette wheel algorithm is used to select several individuals to enter the winning pool. Individuals with higher fitness values have a greater probability of being selected. The probability calculation formula is as follows: ; ; In the formula, is the fitness value of individual after normalization; is the fitness value; is the minimum value among the fitness values of all individuals; is a constant with a very small value to avoid the normalized fitness value being zero, and it is taken as 1e-8. is the probability that individual is selected; is the sum of the normalized fitness values of all individuals.
[0029] Step 4.3: All individuals that do not enter the winning pool enter the remaining evolution pool, and the three-source crossover operation is performed on all individuals in the remaining evolution pool. This operation performs crossover on parameters such as convolutional layers, fully connected layers, and bias terms, where the convolution kernel, row, and element are used as the basic units of crossover respectively. During the crossover process of each basic unit, the probability that the target individual retains itself is , the probability of selecting an individual from the winning pool is , and the probability of selecting an individual from the remaining evolution pool is , satisfying the following formula: ; Among them, is 0.5, is 0.3, is 0.2.
[0030] Step 4.4: The mutation operation is performed on all individuals in the remaining evolution pool. Mutation means adjusting the value of an individual's parameters by adding noise based on the standard normal distribution, is the value of the parameter of individual participating in the mutation, is the variance, representing the average squared distance of the data from the mean. Each individual triggers a mutation event according to the mutation probability . When a mutation is triggered, for the multi-dimensional parameter vector of the individual, according to the preset ratio of the parameters of the conventional mutation and the ratio of the parameters of the super-strong mutation , they are respectively selected through sampling without replacement Execute conventional mutation on parameters and execute super-strong mutation on parameters. Among them, is the number of parameters of individual , is the variance parameter of conventional mutation, , and satisfy the following formula: ; where is 0.2, is 10, is 0.1, is 0.02.
[0031] Step 5. Repeat Step 3 to Step 4 until the maximum number of rounds is reached. Specifically, it includes the following steps: Completely executing one round of Step 3 and Step 4 is called one round. In each round, Step 3 is executed times, and Step 4 is executed once. is 200. After Step 4 ends, the individual with the highest fitness replaces the reinforcement learning individual in Step 3. The tolerance threshold and the mutation probability change with the number of rounds. The formula is: ; ; In the formula, is the minimum value of the tolerance threshold, which is 0.6; is the maximum value of the tolerance threshold, which is 0.9; is the minimum value of the mutation probability, which is 0.1; is the maximum value of the mutation probability, which is 0.8; is the current number of rounds; is the maximum number of rounds, which is 250.
[0032] Repeat Step 3.1 to Step 4.4, and execute rounds in total, and finally obtain the optimal horizontal well section point position, valve opening degree and injection-production system optimal plan.
[0033] The comparative analysis of the final optimization results is as follows: Figure 2It is a comparison chart of the economic net present value convergence curves of the theoretical maximum economic net present value of the horizontal well control enhanced learning method assisted by the evolutionary algorithm (abbreviated as evolutionary-assisted reinforcement learning) and the horizontal well non-segmented only control system in the embodiments of the present invention. It can be seen from the curve convergence that the method proposed in the present invention can obtain a higher economic net present value within the same control time step.
[0034] Figure 3 It is a comparison chart of the cumulative oil production of the theoretical maximum economic net present value of the horizontal well control enhanced learning method assisted by the evolutionary algorithm (abbreviated as evolutionary-assisted reinforcement learning) and the horizontal well non-segmented only control system in the embodiments of the present invention. It can be seen that the method proposed in the present invention has a larger cumulative oil production, effectively realizing oil production increase.
[0035] Figure 4 It is a comparison chart of the cumulative water production of the theoretical maximum economic net present value of the horizontal well control enhanced learning method assisted by the evolutionary algorithm (abbreviated as evolutionary-assisted reinforcement learning) and the horizontal well non-segmented only control system in the embodiments of the present invention. It can be seen that the method proposed in the present invention has less cumulative water production, effectively realizing water control.
[0036] Figure 5 It is the oil saturation and well location map of the optimized reservoir, showing the specific locations of three production wells (numbered P1 - P3) and four injection wells (numbered I1 - I4).
[0037] Figure 6 、 Figure 7 、 Figure 8 They are respectively schematic diagrams of the optimized sectional point positions of production well P1, production well P2, and production well P3. The coordinates of each grid are shown in the figure. For example, (4, 4, 1) is the coordinate of a grid. It can be seen from Figures 6 - 8 that each well has been effectively segmented and all meet the constraint of the maximum number of segments per well.
[0038] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A horizontal well control reinforcement learning method based on evolutionary algorithm assistance, characterized in that: The steps include: Step 1: Determine the segment point position, valve opening and injection and production variables that need to be optimized, and construct a hybrid action space that combines discrete and continuous control measures; Step 2: Build a hybrid deep reinforcement learning agent model and a recent experience priority replay pool; Step 3: Interact the hybrid deep reinforcement learning agent model with the numerical simulator and train the agent through the generated experience samples; Step 4: Execute the replacement, selection, crossover, and mutation operations of the evolutionary algorithm to optimize the hybrid deep reinforcement learning agent model; Step 5: Repeat steps 3 to 4 until the maximum number of rounds is reached.
2. The horizontal well control reinforcement learning method based on evolutionary algorithm assistance according to claim 1 is characterized in that: The specific process of step 1 is as follows: Step 1.1: Determine the horizontal well segmentation points, valves, and the number of oil and water wells that need to be regulated according to the target reservoir. The variables that need to be optimized include segmentation point positions, valve openings, and injection and production variables. Oil and water wells include production wells and water injection wells. The injection and production variables include production rate and water injection rate. The specific expression is: ; In the formula, To control the time step The variables that need to be optimized are: is the number of The location of the segment points; is the number of The opening of a valve; is the number of The production rate of each production well; is the number of The injection rate of each injection well; Step 1.2: Construct discrete action space for: ; Step 1.3: Construct a continuous action space for: ; Step 1.4: The hybrid action space combining discrete and continuous control measures is expressed as: ; In the formula, It is a mixed action space.
3. The horizontal well control reinforcement learning method based on evolutionary algorithm assistance according to claim 2 is characterized in that: In step 2, the hybrid deep reinforcement learning agent model includes a strategy network, an action value network and a target action value network; The strategy network consists of two parts: a discrete strategy network and a continuous strategy network. The discrete strategy network consists of a convolutional neural network and a fully connected neural network. The input is the reservoir saturation field and pressure field. The data is first extracted from the reservoir data through a convolutional neural network, and then the discrete action probability distribution is output through a fully connected neural network. The continuous strategy network consists of a convolutional neural network and a fully connected neural network. The input is the reservoir saturation field, pressure field and discrete action probability distribution. The reservoir saturation field and pressure field are first extracted from the reservoir data through a convolutional neural network. The extracted reservoir data features are spliced with the discrete action probability distribution and then input into the fully connected neural network, which outputs the Gaussian distribution of continuous actions. The network structures of the action value network and the target action value network are consistent, both consisting of a convolutional neural network and a fully connected neural network. The input of the action value network is the reservoir saturation field, pressure field and continuous action of a given time step, and the input of the target action value network is the reservoir saturation field, pressure field and continuous action of the next time step of a given time step. The reservoir saturation field and pressure field are first extracted from reservoir data features through a convolutional neural network, and the extracted reservoir data features are concatenated with the continuous action and input into a fully connected neural network. The action value network outputs the Q value of the action of a given time step, and the target action value network outputs the Q value of the action of the next time step of a given time step. Build a recent experience priority replay pool The sampling scale factor and actual sampling range are as follows: ; ; In the formula, is the sampling scale factor; is the current update number; is the total number of updates; A hyperparameter to control sampling bias; is the actual sampling range; Is the amount of experience in the current experience pool.
4. The horizontal well control reinforcement learning method based on evolutionary algorithm assistance according to claim 3 is characterized in that: In step 3, the agent is a reinforcement learning individual; the specific process of step 3 is: Step 3.1: Assume the reward value of the current time step is the economic net present value of the current time step: ; In the formula, is the current time step; is the reward value at the current time step; is the economic net present value at the current time step; is the number of producing wells; is the number of injection wells; is the unit price of crude oil sales; Treatment costs for production well water; the cost of filling injection wells with water; , They are respectively Daily oil and water production of production wells; For the The injection rate of the injection wells; Step 3.2: Take the saturation field and pressure field data of the reservoir model at the current time step as the state Input into the policy network, the policy network outputs the discrete action of the current time step and the continuous action at the current time step ; Calculate the reward value at the current time step through the numerical simulator and the state of the next time step , if the current time step is the last time step, then the end mark is 1, otherwise it is 0; the trajectory of the empirical sample Save in the recent experience priority replay pool middle; Step 3.3: Use recent experience to prioritize the replay pool The experience sample training action value loss in is as follows: ; ; ; In the formula, is the Q value of the target action; is the discount factor; represents the probability distribution of discrete actions; To find the minimum value; For the A target action value network; represents the continuous action at the next time step; Indicates the temperature coefficient of continuous action; represents the probability distribution of continuous actions; represents the discrete action at the next time step; represents the temperature coefficient of discrete action; Indicates The loss function of the action value network, the sequence number of the action value network corresponds to the sequence number of the target action value network; represents the computational expectation; is the action value network; represents the total loss function of the action-value network; Step 3.4: Use recent experience to prioritize the replay pool The network parameters of the experience sample training strategy in are as follows: ; ; ; in, is the loss function of the discrete policy network; For calculation expectations; is the loss function of the continuous policy network; is the total loss function of the policy network; Step 3.5: Design the loss function of the temperature coefficient of continuous action and discrete action. and Update as follows: ; ; in, is the loss function of the continuous action temperature coefficient; is the loss function of the discrete action temperature coefficient; and Represent the entropy of continuous policy network and discrete policy network respectively; and denote the target entropy of the continuous policy network and the discrete policy network respectively.
5. The horizontal well control reinforcement learning method based on evolutionary algorithm assistance according to claim 4 is characterized in that: The specific process of step 4 is as follows: Step 4.1: Evolutionary Algorithm individuals, each of which is a policy network; the sum of all rewards of the policy network in one round is recorded as the initial fitness value of the individual; a penalty term is set, and if the continuous action output by the policy network is equal to the boundary value, the fitness value of the policy network is deducted accordingly; the fitness values of all individuals are calculated, and the formula is: ; In the formula, For individuals The fitness value of For individuals The initial fitness value of For individuals The penalty term is increased by 1 for each boundary value taken by the individual; Calculate the minimum tolerance of all individuals: ; In the formula, For individuals minimum tolerance level; For individuals The fitness value of the previous round; is the tolerance threshold; If the fitness value of the current individual is the lowest or lower than the individual's minimum tolerance, the reinforcement learning individual will be used to replace the current individual; Step 4.2: The individual with the highest fitness value directly enters the winning pool, and then the roulette algorithm is used to select several individuals to enter the winning pool. The individual with a high fitness value has a high probability of being selected. The probability calculation formula is as follows: ; ; In the formula, Is an individual Normalized fitness value; is the fitness value; is a constant; For individuals The probability of being selected; Step 4.3: All individuals that did not enter the winning pool enter the remaining evolution pool, and a three-source crossover operation is performed on all individuals in the remaining evolution pool. This operation crosses the convolution layer, the fully connected layer, and the bias term, where the convolution kernel, the row, and the element are used as the basic units of crossover, respectively. In the crossover process of each basic unit, the probability of the target individual retaining itself is , the probability of selecting an individual from the winning pool is , the probability of selecting an individual from the remaining evolution pool is , satisfying the following formula: ; Step 4.4: Perform mutation operation on all individuals in the remaining evolution pool. Mutation refers to adding noise based on standard normal distribution to the parameters of individuals. To adjust its value, For individuals The values of the parameters involved in the mutation, is the variance; each individual has a probability of mutation Trigger a mutation event. When a mutation is triggered, the multidimensional parameter vector of the individual is adjusted according to the ratio of the parameters of the preset conventional mutation. and the ratio of parameters with strong variation Selected by sampling without replacement parameters to perform regular mutation and parameters to perform super strong mutation; For individuals The number of parameters, is rounded up; defines the variance parameter of the normal variation Smaller than the variance parameter of the super strong mutation , and Satisfies the following formula: 。 6. The horizontal well control reinforcement learning method based on evolutionary algorithm assistance according to claim 5 is characterized in that: The specific process of step 5 is: a complete execution of steps 3 and 4 is called a round, and each round of step 3 is executed times, and step 4 is executed once; after step 4, the individual with the highest fitness replaces the reinforcement learning individual in step 3, and the tolerance threshold in step 4 and mutation probability It changes with the number of rounds, and the formula is: ; ; In the formula, is the minimum tolerance threshold; is the maximum value of the tolerance threshold; is the minimum mutation probability; is the maximum probability of mutation; is the current round number; is the maximum number of rounds; Repeat steps 3.1 to 4.4 for a total of After several rounds, the optimal horizontal well segmentation point position, valve opening and injection-production system are finally obtained.
Citation Information
Patent Citations
Optimizing well management policy
CA2766437A1
Injection-production joint debugging intelligent decision-making method and system based on shaft parameters
CN112836349A
Recent experience-guided oil reservoir multi-measure flow field regulation and control reinforcement learning method
CN118095667A
Digital twin-driven power quality monitoring and intelligent optimization method and system
CN118842191A
Horizontal subdivision capacity optimization method based on adaptive entropy deep reinforcement learning
CN119558468A
Cited By
Inventory optimization method and system based on evolution-assisted multi-agent reinforcement learning
CN120931208A