Gas well group production pressure differential control method based on reinforcement learning and safety constraint projection
By using a method based on reinforcement learning and safety constraint projection, the problems of reliance on experience and low computational efficiency in the production pressure differential control of gas well clusters are solved, realizing intelligent and executable control of the production pressure differential of gas well clusters, and improving the scientificity and safety of production scheduling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-07
AI Technical Summary
Existing gas well cluster differential pressure control methods rely heavily on experience, have low computational efficiency, and fail to adequately characterize the well-network coupling effect, making it difficult to achieve efficient and stable production control. In particular, under conditions of dense well networks, heterogeneous reservoirs, or large supply and demand fluctuations, they suffer from response lag and operational instability.
By employing a method based on reinforcement learning and safety constraint projection, executable valve action commands are generated through the construction of training samples, reinforcement learning models, and safety constraint projections, thereby achieving intelligent control of the production pressure differential of gas well clusters.
It has achieved efficient, stable and executable control of the production pressure differential of the gas well group, improved the level of intelligent production scheduling, and enhanced the scientific nature and safety of production.
Smart Images

Figure CN121277006B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas field production optimization and intelligent control technology, and more specifically, to a method for production differential pressure control of gas well clusters based on reinforcement learning and safety constraint projection. Background Technology
[0002] In the process of gas field development, the production pressure differential is a key parameter that determines the release of gas well production capacity, the maintenance of reservoir energy, and the stability of gas supply. Reasonable control of the production pressure differential of a group of wells not only helps to delay reservoir pressure depletion, but also ensures stable pipeline pressure supply and improves overall development efficiency. Traditional gas well group control methods mainly rely on manual experience, single well start-up and shutdown rules, or fixed opening strategies, lacking a systematic pressure differential control framework. They also have difficulty taking into account the coupling effect between wells and pipeline constraints. Especially in the case of dense well networks, heterogeneous reservoirs, or large fluctuations in supply and demand, the coordinated control of well groups faces problems of delayed response and unstable operation, which restricts the level of intelligent production scheduling.
[0003] While existing numerical simulation or optimization methods can be used for differential pressure prediction and control scheme evaluation, they are highly dependent on geological models, initial conditions, and manually set boundary constraints. The calculation process is complex and time-consuming, making it difficult to support real-time production control needs. At the same time, these methods often ignore the safety window of differential pressure in a single well and the pressure limit of pipeline nodes, resulting in output schemes that lack executability and cannot directly guide field operations.
[0004] Therefore, existing gas well group differential pressure control methods generally suffer from problems such as strong reliance on experience, low computational efficiency, insufficient characterization of well-network coupling effects, and inadequate safety assurance. There is an urgent need for an intelligent control method that balances optimization and safety to achieve efficient, stable, and executable control of the gas well group differential pressure. Summary of the Invention
[0005] The purpose of this invention is to provide a method for differential pressure control in gas well cluster production based on reinforcement learning and safety constraint projection. It aims to construct an intelligent and executable control scheme for differential pressure regulation and valve parameter optimization tasks in the collaborative production of gas field clusters.
[0006] The above-mentioned technical objective of the present invention is achieved through the following technical solution:
[0007] In the first aspect, this application provides a method for controlling the production differential pressure of a gas well cluster based on reinforcement learning and safety constraint projection, including the following specific steps:
[0008] Obtain historical production data of each gas well in the group, and construct training samples using the preprocessed historical production data;
[0009] Construct an initial model based on reinforcement learning and train the initial model using training samples until the initial model reaches the training termination condition. The initial model that reaches the training termination condition is determined as the control model.
[0010] The real-time production data of the gas well to be regulated is obtained and input into the regulation model for processing to obtain the valve action range of the gas well to be regulated.
[0011] Based on the preset safety constraints, the valve action range is projected onto the constraint feasible region of the safety constraints, and the executable actions of the valve action range are obtained by solving.
[0012] The valves of the gas well to be controlled are regulated based on executable actions, and the control model is dynamically adjusted based on the production data after the executable actions are executed.
[0013] Based on the above technical solution, the present invention can be further improved as follows.
[0014] Furthermore, the above training termination condition is: the reward function of the initial model is not lower than the threshold.
[0015] Furthermore, the reward function described above is as follows:
[0016] ;
[0017] In the formula, To pass The value of the reward function calculated from the data generated at each moment. This indicates the total output of the well group. gas well exist Pressure difference at any moment Indicates gas well The upper limit of pressure differential safety, Indicates a constraint or penalty item. This indicates the total number of gas wells in the well group. This represents the maximum total output of the well group. These are the weight coefficients for the corresponding items.
[0018] Furthermore, the total production of the aforementioned well group is as follows:
[0019] ;
[0020] In the formula, The total production of the well group Indicates gas well Daily gas production This indicates the total number of gas wells in the well group.
[0021] Furthermore, the aforementioned production data includes at least daily gas production, bottom hole pressure, wellhead pressure, valve opening, and pipeline node pressure.
[0022] Furthermore, the training samples mentioned above include the daily gas production, bottom hole pressure, wellhead pressure, pipeline node pressure, and gas well pressure differential of each gas well at different times; among which, the gas well pressure differential is specifically:
[0023] ;
[0024] In the formula, For gas well pressure differential, This represents the formation pressure value. This represents the bottom hole pressure value for the corresponding gas well.
[0025] Furthermore, the aforementioned executable actions are as follows:
[0026] ;
[0027] in:
[0028] ;
[0029] In the formula, For executable actions, Let be the optimization variable for the quadratic programming problem, representing the valve opening adjustment amount; The valve action amplitude is obtained through the scheduling model. This is the constraint coefficient matrix obtained through the safety constraints. Let be the constraint upper bound margin vector obtained through the safety constraints, and let st represent the constraints used for the optimization problem.
[0030] Secondly, this application provides a gas well cluster production differential pressure control system based on reinforcement learning and safety constraint projection, applied to the gas well cluster production differential pressure control method based on reinforcement learning and safety constraint projection in any of the first aspects, including:
[0031] The sample construction module is used to obtain historical production data of each gas well in the group and construct training samples through preprocessed historical production data.
[0032] The model training module is used to build an initial model based on reinforcement learning and train the initial model using training samples until the initial model reaches the training termination condition. The initial model that reaches the training termination condition is determined as the control model.
[0033] The data processing module is used to acquire real-time production data of the gas well to be regulated, input the real-time production data into the regulation model for processing, and obtain the valve action range of the gas well to be regulated.
[0034] The action determination module is used to project the valve action range onto the constraint feasible region of the safety constraint conditions based on preset safety constraints, and solve for the executable actions of the valve action range.
[0035] The cyclic control module is used to control the valves of the gas well to be controlled based on executable actions, and to dynamically adjust the control model based on the production data after the executable actions are executed.
[0036] Thirdly, this application provides an electronic device, including: at least one processor, at least one memory, and a data bus;
[0037] In this system, the processor and memory communicate with each other via a data bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method as described in any of the first aspects.
[0038] Fourthly, this application provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to perform any of the methods in the first aspect.
[0039] Compared with the prior art, the present invention has at least the following beneficial effects:
[0040] Using gas well production pressure differential as the core control indicator, this method generates valve opening increment predictions through a reinforcement learning model. Combined with a safety constraint projection mechanism, it overcomes the bottlenecks of slow response and insufficient executability of traditional experience-based control and numerical simulation methods in complex well network scenarios. The reinforcement learning component learns the dynamic relationship between well production and pressure differential, enabling adaptive control strategy generation. The safety constraint projection layer corrects the model output, ensuring that control actions meet engineering conditions such as single-well pressure differential windows, pipeline node pressures, and the upper limit of well flow rates. Under this closed-loop architecture, safe and feasible valve control commands for the well network can be output in real time. It possesses advantages such as strong adaptive optimization capabilities, clear engineering constraint expression, and on-site implementability, making it a general auxiliary tool for gas well network production operation management, effectively improving the scientific, safe, and economical aspects of gas field production pressure differential control. Attached Figure Description
[0041] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:
[0042] Figure 1 This is a flowchart of the control method in an embodiment of the present invention;
[0043] Figure 2 This is a connection diagram of the control system in an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram of the connection of an electronic device in an embodiment of the present invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0046] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0047] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0048] In the description of the embodiments of the present invention, "multiple" means at least two.
[0049] S1: Obtain historical production data of each gas well in the group, and construct training samples using the preprocessed historical production data.
[0050] The production data includes, but is not limited to, daily gas production, bottom hole pressure, wellhead pressure, valve opening, and pipeline node pressure. Specifically, the preprocessing includes missing value imputation, outlier removal, and standardized resampling. Linear interpolation is used for missing values, and Z-Score (threshold=3) is used to remove outliers (such as sudden jump points). Finally, the sampling period is unified and normalized to complete the data preprocessing.
[0051] Optionally, the training samples include the daily gas production, bottom hole pressure, wellhead pressure, pipeline node pressure, and gas well pressure differential of each gas well at different times; wherein, the gas well pressure differential is specifically:
[0052] ;
[0053] In the formula, For gas well pressure differential, This represents the formation pressure value. This represents the bottom hole pressure value for the corresponding gas well.
[0054] Specifically, the training samples formed above can be represented as: ,in, gas well Daily gas production gas well exist The pressure at the bottom of the well at any given moment. Indicates gas well exist Wellhead pressure at any given time express The pressure at each pipeline node at any given time. gas well exist The pressure difference of the gas well at any given moment.
[0055] S2, Construct an initial model based on reinforcement learning, and train the initial model using training samples until the initial model reaches the training termination condition. The initial model that reaches the training termination condition is determined as the control model. The training termination condition is: the reward function of the initial model is not lower than a threshold.
[0056] The initial reinforcement learning model, which takes the valve opening increment as output, is constructed by taking the well group state as input, learning the dynamic relationship between production, differential pressure, and energy consumption, and generating the original valve opening control command. The initial model can be an LSTM neural network model; specifically, the input to the initial model is as described above. The output is the valve actuation range. It should be noted that the valve's operating range... The value can be positive or negative. A positive value indicates an increase in valve opening, while a negative value indicates a decrease in valve opening.
[0057] Optionally, the reward function described above is as follows:
[0058] ;
[0059] In the formula, To pass The value of the reward function calculated from the data generated at each moment. This indicates the total output of the well group. gas well exist Pressure difference at any moment Indicates gas well The upper limit of pressure differential safety, Indicates a constraint or penalty item. This indicates the total number of gas wells in the well group. This represents the maximum total output of the well group. These are the weight coefficients for the corresponding items.
[0060] Optionally, in the reward function, the total output of the well swarm can be expressed as:
[0061] ;
[0062] In the formula, The total production of the well group Indicates gas well Daily gas production This indicates the total number of gas wells in the well group.
[0063] S3: Obtain real-time production data of the gas well to be regulated, input the real-time production data into the regulation model for processing, and obtain the valve action range of the gas well to be regulated.
[0064] S4. Based on the preset safety constraints, the valve action range is projected onto the constraint feasible region of the safety constraints, and the executable actions of the valve action range are solved.
[0065] The safety constraints may include:
[0066] Single-well pressure differential constraint: ; and These are the corresponding minimum and maximum values;
[0067] Node pressure constraints: ; This represents the corresponding maximum value;
[0068] Total production constraints for the well group: ; This corresponds to the maximum output value;
[0069] Valve actuation range constraints: , This corresponds to the maximum allowable valve movement range.
[0070] Optional, the specific actions that can be performed are:
[0071] ;
[0072] in:
[0073] ;
[0074] In the formula, For executable actions, Let be the optimization variable for the quadratic programming problem, representing the valve opening adjustment amount; The valve action amplitude is obtained through the scheduling model. This is the constraint coefficient matrix obtained through the safety constraints. Let be the constraint upper bound margin vector obtained through the safety constraints, and let st represent the constraints used for the optimization problem.
[0075] in, This is the constraint coefficient matrix obtained through the safety constraints. This is the constraint upper bound margin vector obtained through the safety constraints. The rows represent the coupling relationship between various engineering constraints and valve actions (including differential pressure, node pressure, output, and valve amplitude constraints). This represents the remaining safety margin of each constraint under the current state, and both are constructed from the aforementioned safety constraints.
[0076] S5 regulates the valves of the gas well to be regulated based on executable actions, and dynamically adjusts the regulation model based on the production data after executing the executable actions.
[0077] Among them, the executable actions The data is sent to the field control system for execution, and data such as gas production, pressure and energy consumption are collected after execution. The training model and constraint parameters are updated periodically to form a closed-loop control process of "prediction-optimization-projection-execution-feedback".
[0078] The control method provided in this embodiment will be further illustrated by the following examples.
[0079] Example 1: Application of intelligent control of differential pressure in well clusters in old well areas.
[0080] In a typical old well area, there are 20 producing gas wells and 3 pipeline nodes. The coupling relationship between the wells is complex, the production of some gas wells is continuously declining, and the manual valve control method is lagging behind, resulting in large fluctuations in gas supply from the group of wells. In order to improve the accuracy of differential pressure control and the overall stability of gas supply, the "Gas Well Group Production Differential Pressure Control Method Based on Reinforcement Learning and Security Constraint Projection" in this embodiment is applied for dynamic control, and the specific implementation is as follows:
[0081] Step 1: Extract nearly 90 days of operational data from the regional production database, including daily gas production, bottom hole pressure, wellhead pressure, valve opening and node pressure; the data is then processed by linear interpolation to fill gaps, Z-Score (threshold=3) to remove outliers, and normalized, with a sampling interval of 10 minutes to form a standard time series input.
[0082] Step 2: Construct a reinforcement learning control model. The state inputs are the production rate of the entire well group, wellhead / bottomhole pressure, differential pressure, and node pressure; the output is the valve opening increment for each well. The DDPG algorithm is used for training, with a sliding window length of 12, a discount factor of 0.99, and a learning rate of 3×10⁻⁻⁶. 4 Batch size 256, buffer size 10 6 The target network soft update coefficient τ = 5 × 10⁻³.
[0083] Step 3: Define safety constraints: Single-well pressure differential range is set to 0.5–5.0 MPa, node pressure does not exceed 6.5 MPa, and the total production limit of the group of wells is 8.0 × 10⁻⁶ MPa. 5 m³ / d, single-step valve actuation amplitude not exceeding 5%, and valve opening degree maintained within the allowable range.
[0084] Step 4: Perform secondary planning projection on the action output by DDPG to obtain the safety valve adjustment amount that meets the engineering constraints, and update the valve opening accordingly.
[0085] Step five involves sending safety control commands to the field control system, collecting execution feedback data, and periodically updating the model and constraint parameters to form a closed-loop control system. Continuous operation for 30 days showed that the gas supply pressure fluctuation range decreased by approximately 32%, the total production of the well group increased by approximately 4.8%, and boundary violations were significantly reduced.
[0086] Example 2: Application of dynamic control in high-pressure gathering and transmission system in new well areas
[0087] In a typical new well area, 30 gas wells were put into production simultaneously, resulting in high node pressures and significant production fluctuations. Manual regulation was insufficient to balance the safety of individual wells with the stability of the pipeline network. To ensure gas supply stability and pressure differential safety, the "Gas Well Cluster Production Pressure Differential Control Method Based on Reinforcement Learning and Security Constraint Projection" in this embodiment was applied for dynamic regulation, as detailed below:
[0088] Step 1: Collect nearly 60 days of operational data at a 5-minute interval. Fields include daily gas production, bottom hole pressure, wellhead pressure, valve opening, and node pressure. The data is processed by interpolation, outlier removal, and normalization to form a high-frequency time-series input.
[0089] Step 2: Construct a reinforcement learning control model. The state vector includes the production rate, differential pressure, and node pressure of each well, and the output is the valve opening increment. The SAC algorithm is used for training, with a sliding window length of 12, and the remaining parameter configurations are the same as in Example 1.
[0090] Step 3, set safety constraints: single-well pressure differential should be between 1.0 and 6.0 MPa, node pressure should not exceed 10 MPa, and the total production capacity of the well group should be capped at 1.5 × 10⁻⁶ MPa. 6 m³ / d, single-step valve actuation amplitude not exceeding 3%, and valve opening degree must be within the physically permissible range.
[0091] Step 4: Perform constraint projection on the SAC model output, construct and solve the quadratic programming problem to obtain safe and executable actions, and update the valve opening on site.
[0092] Step 5 involves issuing and executing safety actions, collecting feedback data, and periodically updating the model and constraint parameters. Actual operational results show that the node pressure stabilized within the range of 9.2 ± 0.3 MPa, gas supply stability improved by approximately 27%, pressure differential exceedance events decreased by approximately 90%, and unit energy consumption decreased by 2–4%.
[0093] Example 2: This application provides a gas well cluster production pressure differential control system based on reinforcement learning and safety constraint projection, applied to the gas well cluster production pressure differential control method based on reinforcement learning and safety constraint projection in Example 1, such as... Figure 2 As shown, it includes:
[0094] The sample construction module is used to obtain historical production data of each gas well in the group and construct training samples through preprocessed historical production data.
[0095] The model training module is used to build an initial model based on reinforcement learning and train the initial model using training samples until the initial model reaches the training termination condition. The initial model that reaches the training termination condition is determined as the control model.
[0096] The data processing module is used to acquire real-time production data of the gas well to be regulated, input the real-time production data into the regulation model for processing, and obtain the valve action range of the gas well to be regulated.
[0097] The action determination module is used to project the valve action range onto the constraint feasible region of the safety constraint conditions based on preset safety constraints, and solve for the executable actions of the valve action range.
[0098] The cyclic control module is used to control the valves of the gas well to be controlled based on executable actions, and to dynamically adjust the control model based on the production data after the executable actions are executed.
[0099] Example 3: This application provides an electronic device, such as... Figure 3 As shown, it includes: at least one processor, at least one memory, and a data bus;
[0100] In this embodiment, the processor and the memory communicate with each other through a data bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method as described in Embodiment 1.
[0101] Example 4: This application provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to execute the method of Example 1.
[0102] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0106] Those skilled in the art will understand that all or part of the steps in the above facts and methods can be implemented by a program instructing related hardware. The program or the program involved can be stored in a computer-readable storage medium. When the program is executed, it includes the following steps: at this time, the corresponding method steps are introduced. The storage medium can be ROM / RAM, magnetic disk, optical disk, etc.
[0107] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for controlling production differential pressure in gas well clusters based on reinforcement learning and safety constraint projection, characterized in that, The specific steps include the following: Obtain historical production data of each gas well in the group, and construct training samples using the preprocessed historical production data; Construct an initial model based on reinforcement learning, and train the initial model using the training samples until the initial model reaches the training termination condition. The initial model that reaches the training termination condition is determined as the control model. The real-time production data of the gas well to be regulated is obtained, and the real-time production data is input into the regulation model for processing to obtain the valve action range of the gas well to be regulated. Based on preset safety constraints, the valve actuation range is projected onto the constraint feasible region of the safety constraints, and the executable actions of the valve actuation range are obtained by solving. The valves of the gas well to be regulated are controlled based on the executable actions, and the regulation model is dynamically adjusted based on the production data after the executable actions are executed. The training termination condition is: the reward function of the initial model is not lower than a threshold; the reward function is specifically: ; In the formula, To pass The value of the reward function calculated from the data generated at each moment. This indicates the total output of the well group. gas well exist Pressure difference at any moment Indicates gas well The upper limit of pressure differential safety, Indicates a constraint or penalty item. This indicates the total number of gas wells in the well group. This represents the maximum total output of the well group. These are the weight coefficients for the corresponding items.
2. The gas well cluster production pressure differential control method based on reinforcement learning and safety constraint projection according to claim 1, characterized in that, The total production of the well group is as follows: ; In the formula, The total production of the well group Indicates gas well Daily gas production This indicates the total number of gas wells in the well group.
3. The gas well cluster production pressure differential control method based on reinforcement learning and safety constraint projection according to claim 1, characterized in that, The production data includes at least daily gas production, bottom hole pressure, wellhead pressure, valve opening, and pipeline node pressure.
4. The gas well cluster production pressure differential control method based on reinforcement learning and security constraint projection according to claim 3, characterized in that, The training samples include the daily gas production, bottom hole pressure, wellhead pressure, pipeline node pressure, and gas well pressure differential of each gas well at different times; wherein, the gas well pressure differential specifically refers to: ; In the formula, For gas well pressure differential, This represents the formation pressure value. This represents the bottom hole pressure value for the corresponding gas well.
5. The gas well cluster production pressure differential control method based on reinforcement learning and safety constraint projection according to claim 1, characterized in that, The specific executable action is as follows: ; in: ; In the formula, For executable actions, Let be the optimization variable for the quadratic programming problem, representing the valve opening adjustment amount; The valve action amplitude is obtained through the scheduling model. This is the constraint coefficient matrix obtained through the safety constraints. Let be the constraint upper bound margin vector obtained through the safety constraints, and let st represent the constraints used for the optimization problem.
Citation Information
Patent Citations
Training method, determining method and device for natural gas well and control system
CN118350486A