Well control methods, devices, apparatus, storage media, and program products
Patent Information
- Application Number
- CN202510233690.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-08-28
AI Technical Summary
[0002]传统的井控方案通常固定且反应迟缓,不足以应对复杂和多变的油藏条件
[0009] The technical solution of this application embodiment acquires real-time production data of each production well in an oil and gas field or gas storage facility. The production data includes production volume, water injection volume, and well pressure. A pre-trained reinforcement learning model, combined with the real-time production data, determines the well control strategy for each production well, thereby optimizing the production of the oil and gas field or gas storage facility based on the well control strategy. The reinforcement learning model is obtained through iterative training using a pre-built environmental model. The well control strategy includes a water injection volume adjustment strategy, a well pressure adjustment strategy, and a production well on/off state adjustment strategy. This embodiment of the application, by using a reinforcement learning model and real-time production data to determine the well control strategy for each production well in real time, can optimize the production of each production well in an oil and gas field or gas storage facility in real time, thereby improving the overall production of the oil and gas field or gas storage facility.
Smart Images

Figure CN122655486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of oil and gas engineering technology, and in particular to a well control method, apparatus, equipment, storage medium and program product. Background Technology
[0002] Traditional well control schemes are typically fixed and slow to react, making them insufficient to cope with complex and variable reservoir conditions. Traditional methods rely on experience and simplified physical models to formulate water injection strategies and well pressure controls. These strategies often fail to reflect the actual production status of the oilfield or gas storage facility in real time, resulting in low resource utilization and low production efficiency. Furthermore, due to a lack of flexibility and adaptability, traditional well control schemes often cannot adjust in time to sudden changes in the production process of the oilfield or gas storage facility, thus missing opportunities to optimize production. Summary of the Invention
[0003] This application provides a well control method, apparatus, equipment, storage medium, and program product that can optimize the production of oil and gas fields or gas storage facilities in real time, thereby increasing the overall output of oil and gas fields or gas storage facilities.
[0004] In a first aspect, embodiments of this application provide a well control method, comprising: acquiring real-time production data of each production well in an oil and gas field or gas storage facility; wherein the production data includes production rate, water injection rate, and well pressure; determining a well control strategy for each production well by combining the real-time production data with a pre-trained reinforcement learning model, so as to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy; wherein the reinforcement learning model is obtained by iterative training using a pre-constructed environmental model; the well control strategy includes a water injection rate adjustment strategy, a well pressure adjustment strategy, and a production well on / off state adjustment strategy.
[0005] Secondly, embodiments of this application also provide a well control device, comprising: a production data acquisition module, used to acquire real-time production data of each production well in an oil and gas field or gas storage facility; wherein the production data includes production rate, water injection rate, and well pressure; and a well control strategy determination module, used to determine the well control strategy for each production well by combining the real-time production data with a pre-trained reinforcement learning model, so as to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy; wherein the reinforcement learning model is obtained by iterative training using a pre-built environmental model; and the well control strategy includes adjusted water injection rate, adjusted well pressure, and the on / off state of the production well.
[0006] Thirdly, embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the well control method as described in embodiments of this application.
[0007] Fourthly, embodiments of this application also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the well control method as described in embodiments of this application.
[0008] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the well control method as described in embodiments of this application.
[0009] The technical solution of this application embodiment acquires real-time production data of each production well in an oil and gas field or gas storage facility. The production data includes production volume, water injection volume, and well pressure. A pre-trained reinforcement learning model, combined with the real-time production data, determines the well control strategy for each production well, thereby optimizing the production of the oil and gas field or gas storage facility based on the well control strategy. The reinforcement learning model is obtained through iterative training using a pre-built environmental model. The well control strategy includes a water injection volume adjustment strategy, a well pressure adjustment strategy, and a production well on / off state adjustment strategy. This embodiment of the application, by using a reinforcement learning model and real-time production data to determine the well control strategy for each production well in real time, can optimize the production of each production well in an oil and gas field or gas storage facility in real time, thereby improving the overall production of the oil and gas field or gas storage facility. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of a well control method provided in an embodiment of this application;
[0012] Figure 2 This is a schematic diagram of a well control device provided in an embodiment of this application;
[0013] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". It should be noted that the concepts of "first," "second," etc., mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications "a" and "a plurality" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless explicitly indicated otherwise in the context, they should be understood as "one or more". It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of applicable laws, regulations, and relevant provisions.
[0016] Figure 1 This is a schematic flowchart of a well control method provided in an embodiment of this application. The method can be executed by a well control device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method includes:
[0017] S110. Obtain real-time production data for each production well in an oil and gas field or gas storage facility.
[0018] The production data includes output, water injection volume, and well pressure.
[0019] Here, output refers to daily production. In oil and gas fields or gas storage facilities, output can be the production of crude oil or natural gas.
[0020] S120. By using a pre-trained reinforcement learning model and combining it with the real-time production data, a well control strategy for each production well is determined, so as to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy.
[0021] The reinforcement learning model is obtained through iterative training using a pre-built environment model; the well control strategy includes a water injection adjustment strategy, a well pressure adjustment strategy, and a production well on / off state adjustment strategy. In this embodiment, well control parameters, namely water injection volume, well pressure, and production well on / off state parameters, can be adjusted according to the well control strategy.
[0022] In this embodiment, the specific structure of the environmental model is not limited and can be obtained using deep learning technology. The environmental model is used to simulate the production of oil and gas fields or gas storage facilities. This model uses physical simulation technology combined with deep learning methods to accurately reflect various physical phenomena and operational impacts during the production process of production wells.
[0023] The environment model includes an input layer, an output layer, and multiple hidden layers; the hidden layers use convolutional neural networks and recurrent neural networks.
[0024] In this embodiment, a multi-layer (e.g., 3-layer) convolutional neural network can be used to simulate oil-water two-phase flow (flow conditions in an oil reservoir). A multi-layer (e.g., 2-layer) recurrent neural network can be used to predict time-varying production.
[0025] Optionally, the training process of the environmental model is as follows: acquiring the first production training data of each production well in the oil and gas field or the gas storage facility; preprocessing the first production training data of each production well; and training the basic environmental model based on the preprocessed first production training data of each production well to obtain the environmental model.
[0026] In this embodiment, no specific limitations are placed on the first production training data. For example, it may include historical production data (production rate, water injection rate, well pressure) and well control strategies (production rate, water injection rate, on / off status of production wells); it may also include environmental data for each production well at the corresponding historical time; and it may also include physical characteristic data of the wellbore and reservoir. In this embodiment, no limitations are placed on the environmental data, which may include geological data, temperature data, pressure data, etc.
[0027] In this embodiment, the preprocessing method is not limited. For example, cleaning and standardization can be performed to ensure the quality and consistency of the environmental model training data.
[0028] Among them, the basic environment model is the environment model that has not been trained, while the trained basic environment model can be directly referred to as the environment model.
[0029] In this embodiment, the reinforcement learning model can use the PPO (Proximal Policy Optimization) algorithm to conduct a large number of simulation experiments through the environment model. Through these experiments, experience is accumulated, and the optimal well control strategy is learned under different production environments (or conditions).
[0030] Optionally, the training method of the reinforcement learning model is as follows: each round of iterative training process is as follows: based on the second production training data corresponding to the first time step of each production well, determine the state training data corresponding to the second time step of the multi-agent and the reward value corresponding to the first time step; wherein, the state training data of the multi-agent is the output of the environment model; the second production training data corresponding to the first time step includes the production training data corresponding to the first time step and the action training data corresponding to the first time step; wherein, the action training data corresponding to the first time step is obtained by the multi-agent based on the state training data corresponding to the first time step; based on the state training data corresponding to the second time step of the multi-agent and combined with the reward value corresponding to the first time step, determine the action training data corresponding to the second time step of the multi-agent; wherein, the action training data represents the well control strategy of each production well during the training phase; when the set iterative training termination condition is met, training is stopped, and the trained reinforcement learning model is obtained.
[0031] The reinforcement learning model includes multiple agents. In this embodiment, a multi-agent reinforcement learning framework is used, which allows multiple agents to operate and optimize different production wells simultaneously. Through cooperation and competition among the agents, the overall production efficiency of oil and gas fields or gas storage facilities can be optimized.
[0032] In this embodiment, a reinforcement learning model can be used to handle high-dimensional decision-making problems related to production optimization in oil and gas fields or gas storage facilities.
[0033] In this embodiment, there is no limit to the number of iterations for training the reinforcement learning model; for example, 100,000 PPO training iterations can be performed to obtain the optimal well control strategy. The training process for each round (or iteration) is similar; the following explanation uses one round of iteration training as an example:
[0034] The second production training data corresponding to the first time step for each production well is input into the environment model. The state training data corresponding to the second time step and the reward value corresponding to the first time step are output and used as input for each agent in the multi-agent system. The second time step is greater than the first time step. Both the second and first time steps are historical times. The state training data of the multi-agent system is the output of the environment model; the second production training data corresponding to the first time step includes the production training data and action training data corresponding to the first time step. The production training data may include production rate, water injection rate, and well pressure. The action training data corresponding to the first time step is obtained by the multi-agent system based on the state training data corresponding to the first time step.
[0035] The multi-agent system determines the action training data for the second time step based on the state training data at the second time step and the reward value at the first time step. The action training data represents the well control strategy for each production well during the training phase, and may include production rate, water injection rate, and the on / off status of the production well. It should be noted that the action training data for the second time step output by the multi-agent policy network is calculated using a uniformly weighted average of the action probability distributions of each agent.
[0036] When the set termination condition for iterative training is met, training stops, and the trained reinforcement learning model is obtained.
[0037] In this embodiment, there are no restrictions on setting the termination condition for iterative training. For example, it can be that the number of training iterations is greater than a set threshold for the number of training iterations.
[0038] It should be noted that during training, an adaptive learning rate adjustment mechanism can be used to update the parameters of the policy network in the multi-agent system, thereby optimizing the training process and accelerating the convergence speed.
[0039] It should be noted that during the training process, a multi-objective optimization algorithm is applied to minimize environmental impact and operating costs, i.e., total production costs, while increasing the total output (or total output revenue) of oil and gas fields or gas storage facilities, thus achieving a balance between economic and ecological objectives.
[0040] Optionally, the reward value corresponding to the first time point is obtained based on the total production revenue and total production cost generated by all production wells at the first time point; wherein, the total production revenue and total production cost are obtained based on the second production training data corresponding to the first time point.
[0041] The reward value is obtained based on a reward function, which is a global reward function to maximize long-term cumulative rewards.
[0042] The reward value can be understood as the difference between the total production revenue generated by all production wells (i.e., the economic benefits corresponding to the total production) and the total production cost.
[0043] It should be noted that the reward function corresponding to the first moment can consider not only the total production revenue and total production cost generated by all production wells at that first moment, but also the total production revenue and total production cost generated by all production wells in the long term (starting from the first moment). This is to consider the long-term health and sustainable production capacity of the oil and gas field or gas storage facility, optimize the balance between short-term and long-term interests, and promote the effective utilization and protection of resources. The reward value corresponding to the first moment can also be understood as the discounted total reward starting from the first time point. The global reward function corresponding to the first moment can be composed of the time rewards (i.e., reward values) received at multiple time points (starting from the first moment), the discount factor, and the number of steps starting from the first moment.
[0044] In this embodiment, by using an environment model to train or simulate a large number of times the reinforcement learning model can be trained, it can learn how to make the optimal well control strategy under various conditions.
[0045] Optionally, after determining the well control strategy for each production well using a pre-trained reinforcement learning model and the real-time production data to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy, the method further includes: acquiring first actual production data for each production well in the oil and gas field or gas storage facility after optimization according to the well control strategy; acquiring first predicted production data obtained based on the real-time production data and the environmental model; determining an evaluation result based on the first actual production data and the first predicted production data; and adjusting the reinforcement learning model according to the evaluation result; wherein the adjustment method includes using an adaptive learning rate adjustment mechanism.
[0046] In this embodiment, a trained reinforcement learning model can be applied to the actual production process of an oil and gas field or gas storage facility. Real-time, up-to-date well control strategies can be obtained based on real-time production data, and well control parameters, such as water injection volume, well pressure, and the on / off status of production wells, can be adjusted according to these strategies to optimize production in the oil and gas field or gas storage facility. The first actual production data (i.e., optimized production data) for each production well in the optimized oil and gas field or gas storage facility is then acquired. The real-time production data is input into the environmental model to obtain the first predicted production data (i.e., predicted production data obtained based on the real-time production data). The first actual production data and the first predicted production data are compared and analyzed to obtain an evaluation result, which represents the effectiveness of the well control strategy implementation. The environmental model and the PPO algorithm are then adjusted based on the evaluation result to improve the prediction accuracy and efficiency of the environmental model and the reinforcement learning model, and to reduce resource waste. Methods for adjusting the environmental model include, but are not limited to, retraining the model and modifying hyperparameters.
[0047] In this embodiment, if the reinforcement model is adjusted by retraining the model, an adaptive learning rate adjustment mechanism or a modified learning rate strategy can be used.
[0048] Optionally, after determining the well control strategy for each production well using a pre-trained reinforcement learning model in conjunction with the real-time production data, and optimizing the production of the oil and gas field or the gas storage facility according to the well control strategy, the method further includes: acquiring environmental data and second actual production data for each production well in real time; and updating the environmental model based on the environmental data and the second actual production data.
[0049] In this embodiment, the second actual production data may be the same as or different from the first actual production data.
[0050] In this embodiment, environmental data and second actual production data of each production well can be acquired in real time, and the parameters of the environmental model can be updated according to the environmental data and the second actual production data to dynamically adjust the environmental model to adapt to changes in input data. This can improve the responsiveness and accuracy of the environmental model to changes in actual oil and gas fields or gas storage facilities, thereby improving the accuracy and real-time performance of well control strategies.
[0051] The environmental data can include geological data, temperature data, pressure data, etc. This environmental data can be obtained by receiving data from sensors in the oil and gas field or the gas storage facility.
[0052] In this embodiment, the data predicted by the reinforcement learning model and the environmental model can be integrated into a unified database. Data analysis and visualization tools can be used to compare and analyze the data in the database with the actual production data after the application of the well control strategy, identify potential problems and anomalies, so as to achieve dynamic optimization of the production process, respond to changes in oil and gas fields or gas storage facilities in a timely manner, and improve production efficiency and resource utilization.
[0053] Optionally, after determining the well control strategy for each production well using a pre-trained reinforcement learning model in conjunction with the real-time production data, and optimizing the production of the oil and gas field or the gas storage facility based on the well control strategy, the method further includes: acquiring monitoring data during the implementation of the well control strategy; determining risk results based on the monitoring data; and adjusting the well control strategy based on the risk results.
[0054] The monitoring data can be sensor data and / or camera data for each production well in an oil and gas field or gas storage facility at the current moment. In this embodiment, the monitoring data can be used to detect the risk outcome of each production well in the oil and gas field or gas storage facility at the next moment. The risk outcome can include normal and abnormal results. Abnormalities can be various abnormal situations that may lead to safety accidents, equipment damage, or production interruptions, such as well blowouts, well leakage, or equipment damage. If the risk outcome is normal, no changes are made. If the risk outcome is abnormal, the well control strategy is modified promptly, such as suspending well control operations or modifying well control parameters, to ensure the safety of well control operations for each production well in the oil and gas field or gas storage facility.
[0055] In this embodiment, the environmental model and reinforcement learning model can be evaluated and optimized according to a set period. This embodiment does not impose specific limitations on the set period; for example, it can be one week, one month, etc. This embodiment, by periodically evaluating and optimizing the environmental model and reinforcement learning model, adapts to the long-term production changes and operational environment changes of each production well in the oil and gas field or gas storage facility, ensuring that the reinforcement learning model can continuously provide the best well control strategy and respond to environmental changes.
[0056] In this embodiment, the reinforcement learning model and the environment model can also periodically receive guidance data from an external expert system. The guidance data is based on the latest oil and gas field or gas storage operation research technology and is used to correct and enhance the quality of the reinforcement learning model and the environment model.
[0057] The technical solution of this application embodiment acquires real-time production data of each production well in an oil and gas field or gas storage facility. The production data includes production volume, water injection volume, and well pressure. A pre-trained reinforcement learning model, combined with the real-time production data, determines the well control strategy for each production well, thereby optimizing the production of the oil and gas field or gas storage facility based on the well control strategy. The reinforcement learning model is obtained through iterative training using a pre-built environmental model. The well control strategy includes a water injection volume adjustment strategy, a well pressure adjustment strategy, and a production well on / off state adjustment strategy. This embodiment of the application, by using a reinforcement learning model and real-time production data to determine the well control strategy for each production well in real time, can optimize the production of each production well in an oil and gas field or gas storage facility in real time, thereby improving the overall production of the oil and gas field or gas storage facility.
[0058] The well control method provided in this embodiment can be applied to actual production, and production data after implementation can be collected to verify its effectiveness. The effectiveness can be verified by comparing key indicators such as the total production of the oil and gas field or gas storage facility, single-well production, and production costs before and after implementation.
[0059] For example, the well control method provided in this embodiment can be used to optimize the production of oil and gas field A or gas storage facility. The main steps are as follows: First, data collection is carried out by installing pressure and temperature sensors on each production well to monitor well pressure and ground temperature in real time, and recording the daily production and water injection volume of each production well through flow meters.
[0060] Using deep learning technology, an environmental model of oil and gas field A or its storage facility was constructed. The model training used the first production training data from the past five years, achieving an accuracy rate of 92% through iterative learning.
[0061] A reinforcement learning model was designed and a simulated A oil and gas field or gas storage environment (i.e., an environment model) was used. 100,000 PPO training iterations were carried out to find the optimal well control strategy and obtain the trained reinforcement learning model.
[0062] Using the trained reinforcement learning model and real-time production data for each production well, a well control strategy is obtained for each production well. During the application of the well control strategy, the well control parameters are adjusted in real time. For example, the water injection rate of 10 key production wells is adjusted, the well pressure of 5 production wells is increased, and 3 inefficient production wells are shut down.
[0063] At the same time, the enhanced model and environmental model are updated regularly to adapt to real-time changes in oil and gas field or gas storage production.
[0064] For example, one month later, the reinforcement learning model predicted that production wells in a certain area might experience a decline in production due to insufficient water injection, and suggested a well control strategy to increase water injection. After verification through small-scale trials, the water injection adjustment strategy was fully implemented. Before and after implementing the new well control strategy, the total production, single-well production, and production costs of the oil and gas field or gas storage facility A were compared. Total production increased by 15%, single-well production increased by an average of 10%, and production costs decreased by 5%. Based on the implementation results, the environmental model and reinforcement learning model were fine-tuned to further improve prediction accuracy and optimization efficiency. Production data and environmental changes are continuously monitored, and the environmental model and reinforcement learning model are comprehensively evaluated and optimized quarterly to adapt to long-term changes in the oil and gas field or gas storage facility.
[0065] Figure 2 This is a schematic diagram of a well control device structure provided in an embodiment of this application, as shown below. Figure 2 As shown, the device includes: a production data acquisition module 210 and a well control strategy determination module 220;
[0066] The production data acquisition module 210 is used to acquire real-time production data of each production well in an oil and gas field or gas storage facility; wherein, the production data includes production output, water injection volume, and well pressure;
[0067] The well control strategy determination module 220 is used to determine the well control strategy for each production well by combining the real-time production data with a pre-trained reinforcement learning model, so as to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy; wherein, the reinforcement learning model is obtained by iterative training using a pre-built environment model; the well control strategy includes the adjusted water injection volume, the adjusted well pressure, and the on / off status of the production well.
[0068] The technical solution of this application embodiment acquires real-time production data of each production well in an oil and gas field or gas storage facility through a production data acquisition module. The production data includes production volume, water injection volume, and well pressure. A well control strategy determination module determines a well control strategy for each production well using a pre-trained reinforcement learning model combined with the real-time production data, thereby optimizing the production of the oil and gas field or gas storage facility based on the well control strategy. The reinforcement learning model is obtained through iterative training using a pre-built environmental model. The well control strategy includes a water injection volume adjustment strategy, a well pressure adjustment strategy, and a production well on / off state adjustment strategy. This embodiment of the application, by using a reinforcement learning model and real-time production data to determine the well control strategy for each production well in real time, can optimize the production of each production well in an oil and gas field or gas storage facility in real time, thereby improving the overall production of the oil and gas field or gas storage facility.
[0069] The environment model includes an input layer, an output layer, and multiple hidden layers; the hidden layers use convolutional neural networks and recurrent neural networks.
[0070] The training process of the environmental model is as follows: acquiring the first production training data of each production well in the oil and gas field or the gas storage facility; preprocessing the first production training data of each production well; and training the basic environmental model based on the preprocessed first production training data of each production well to obtain the environmental model.
[0071] The reinforcement learning model includes multiple agents. The training method for the reinforcement learning model is as follows: each iteration of the training process is as follows: based on the second production training data corresponding to the first time step of each production well, the state training data corresponding to the second time step of the multiple agents and the reward value corresponding to the first time step are determined; wherein, the state training data of the multiple agents is the output of the environment model; the second production training data corresponding to the first time step includes the production training data corresponding to the first time step and the action training data corresponding to the first time step; wherein, the action training data corresponding to the first time step is obtained by the multiple agents based on the state training data corresponding to the first time step; based on the state training data corresponding to the second time step of the multiple agents, combined with the reward value corresponding to the first time step, the action training data corresponding to the second time step of the multiple agents is determined; wherein, the action training data represents the well control strategy of each production well during the training phase; when the set iteration training termination condition is met, training is stopped, and the trained reinforcement learning model is obtained.
[0072] The reward value corresponding to the first time point is obtained based on the total production revenue and total production cost generated by all production wells at the first time point; wherein the total production revenue and total production cost are obtained based on the second production training data corresponding to the first time point.
[0073] Optionally, the above-mentioned device further includes an adjustment module, used to acquire first actual production data of each production well in the oil and gas field or gas storage facility after optimization according to the well control strategy; acquire first predicted production data obtained based on the real-time production data and the environmental model; determine the evaluation result based on the first actual production data and the first predicted production data; and adjust the reinforcement learning model according to the evaluation result; wherein the adjustment method includes using an adaptive learning rate adjustment mechanism.
[0074] Optionally, the above-mentioned device further includes an update module for acquiring environmental data and second actual production data of each production well in real time; and updating the environmental model based on the environmental data and the second actual production data.
[0075] Optionally, the above-mentioned device further includes a monitoring module, used to acquire monitoring data during the implementation of the well control strategy; determine risk results based on the monitoring data; and adjust the well control strategy based on the risk results.
[0076] The well control device provided in this application embodiment can execute the well control method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0077] Figure 3A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0078] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0079] Multiple components in electronic device 10 are connected to input / output (I / O) interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0080] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as method well control.
[0081] In some embodiments, the method control may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the method control described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the method control by any other suitable means (e.g., by means of firmware).
[0082] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0083] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0084] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0085] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0086] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0087] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0088] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the well control method provided in any embodiment of this application.
[0089] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0090] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.
Claims
1. A well control method, characterized in that, include: Real-time production data of each production well in an oil and gas field or gas storage facility is obtained; wherein, the production data includes production, water injection volume, and well pressure; A well control strategy for each production well is determined by combining a pre-trained reinforcement learning model with the real-time production data, so as to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy; wherein, the reinforcement learning model is obtained by iterative training using a pre-built environment model; the well control strategy includes a water injection adjustment strategy, a well pressure adjustment strategy, and a production well on / off state adjustment strategy.
2. The method according to claim 1, characterized in that, in, The environment model includes an input layer, an output layer, and multiple hidden layers; the hidden layers use convolutional neural networks and recurrent neural networks.
3. The method according to claim 1, characterized in that, The training process of the environment model is as follows: Obtain the first production training data for each production well in the oil and gas field or the gas storage facility; The first production training data for each production well is preprocessed; The basic environmental model is trained based on the preprocessed first production training data of each production well to obtain the environmental model.
4. The method according to claim 1, characterized in that, in, The reinforcement learning model includes multiple agents; the training method of the reinforcement learning model is as follows: The training process for each iteration is as follows: Based on the second production training data corresponding to the first moment of each production well, the state training data corresponding to the second moment of the multi-agent and the reward value corresponding to the first moment are determined; wherein, the state training data of the multi-agent is the output of the environment model; the second production training data corresponding to the first moment includes the production training data corresponding to the first moment and the action training data corresponding to the first moment; wherein, the action training data corresponding to the first moment is obtained by the multi-agent based on the state training data corresponding to the first moment. Based on the state training data corresponding to the second time step of the multi-agent, and combined with the reward value corresponding to the first time step, the action training data corresponding to the second time step of the multi-agent is determined; wherein, the action training data represents the well control strategy of each production well during the training phase; When the set termination condition for iterative training is met, training stops, and the trained reinforcement learning model is obtained.
5. The method according to claim 4, characterized in that, in, The reward value corresponding to the first time point is obtained based on the total production revenue and total production cost generated by all production wells at the first time point; wherein, the total production revenue and total production cost are obtained based on the second production training data corresponding to the first time point.
6. The method according to claim 1, characterized in that, After determining the well control strategy for each production well using a pre-trained reinforcement learning model combined with the real-time production data, and optimizing the production of the oil and gas field or the gas storage facility based on the well control strategy, the method further includes: Obtain the first actual production data of each production well in the oil and gas field or gas storage facility after optimization according to the well control strategy; Obtain first predicted production data based on the real-time production data and the environmental model; The evaluation results are determined based on the first actual production data and the first predicted production data. The reinforcement learning model is adjusted based on the evaluation results; wherein the adjustment method includes using an adaptive learning rate adjustment mechanism.
7. The method according to claim 1, characterized in that, After determining the well control strategy for each production well using a pre-trained reinforcement learning model combined with the real-time production data, and optimizing the production of the oil and gas field or the gas storage facility based on the well control strategy, the method further includes: Real-time acquisition of environmental data and second actual production data for each production well; The environmental model is updated based on the environmental data and the second actual production data.
8. The method according to claim 1, characterized in that, After determining the well control strategy for each production well using a pre-trained reinforcement learning model combined with the real-time production data, and optimizing the production of the oil and gas field or the gas storage facility based on the well control strategy, the method further includes: Acquire monitoring data during the implementation of the well control strategy; The risk outcome is determined based on the monitoring data; The well control strategy is adjusted based on the risk results.
9. A well control device, characterized in that, include: The production data acquisition module is used to acquire real-time production data of each production well in an oil and gas field or gas storage facility; wherein, the production data includes production output, water injection volume, and well pressure; The well control strategy determination module is used to determine the well control strategy for each production well by combining the real-time production data with a pre-trained reinforcement learning model, so as to optimize the production of the oil and gas field or the gas storage facility according to the well control strategy; wherein, the reinforcement learning model is obtained by iterative training using a pre-built environment model; the well control strategy includes the adjusted water injection volume, the adjusted well pressure, and the on / off status of the production well.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the well control method as described in any one of claims 1-8.
11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the well control method as described in any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the well control method as described in any one of claims 1-8.