Oil and gas well throttle valve regulation and control method based on deep reinforcement learning

By constructing a multilayer perceptron model and Markov decision process through deep reinforcement learning, the problem of insufficient control accuracy of traditional oil and gas well throttle valves is solved, realizing dynamic response to complex working conditions and maximizing production efficiency.

CN120798249APending Publication Date: 2025-10-17SICHUAN ZHENDAO TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511142236.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional throttle valve control methods for oil and gas wells lack precision under complex and variable operating conditions, making it difficult to adapt to new well types and unconventional processes. Furthermore, intelligent technologies struggle to achieve precise and continuous control and safety constraints when data is sparse and communication bandwidth is limited.

Method used

A deep reinforcement learning-based approach is adopted to construct a multilayer perceptron model to fit the throttle valve opening-flow characteristics. Combined with a Markov decision process, the throttle valve opening is controlled by an agent. The adjustment strategy is optimized using Actor and Critic networks to achieve real-time dynamic response.

Benefits of technology

It improves the accuracy and adaptability of throttle valve regulation, meets personalized needs, maximizes production efficiency, dynamically responds to complex working conditions, and improves production safety and equipment stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120798249A_ABST
    Figure CN120798249A_ABST
Patent Text Reader

Abstract

The invention relates to an oil and gas well throttle valve regulation and control method based on deep reinforcement learning, and belongs to the technical field of oil exploitation, and the method comprises the following steps: S1, fitting a throttle valve opening-flow characteristic model by using a multilayer perceptron according to detection data of a valve; s2, intelligent agent construction of deep reinforcement learning is carried out on the fitted throttle valve opening-flow characteristic model based on Markov decision, and one valve corresponds to one intelligent agent; and S3, solving a specific valve opening degree by utilizing the intelligent agent, so as to carry out intelligent valve regulation and control according to a production target of a user. According to the method, the logic explanation module of state-action mapping is constructed, the influence of key parameters on the throttle valve opening decision can be determined, the flow is accurately regulated and controlled, and the production benefit is maximized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of oil exploitation, and in particular to a choke valve control method for oil and gas wells based on deep reinforcement learning. BACKGROUND

[0002] In the production process of oil and gas wells, the choke valve is the core equipment for regulating wellhead pressure and flow rate, and its control accuracy directly affects the oil and gas recovery efficiency, equipment safety and formation protection. The traditional control method faces severe challenges under complex and variable oil and gas well conditions: for example, PID control is prone to cause regulation oscillation when the gas-liquid ratio changes suddenly or the wellhead pressure fluctuates sharply, resulting in fracture closure or oil pipe rupture (such as frequent manual intervention due to sudden drop of gas-liquid ratio in the late production stage of shale gas wells); model predictive control (MPC) relies on accurate multiphase flow dynamics model, but it is difficult to model nonlinear phenomena such as gas-liquid slip and hydrate formation in actual working conditions, and model inaccuracy caused by temperature and pressure changes in deep-sea oil and gas wells may cause choke valve ice blocking; the rule base based on expert experience is difficult to adapt to new development well types (such as ultra-deep wells and heavy oil thermal recovery wells) or unconventional processes (such as CO2 flooding), and fixed rules in the steam huff and puff stage may cause energy waste and steam channeling.

[0003] Existing intelligent technologies also have bottlenecks: supervised learning requires a large amount of "state-optimal action" labeled data, but the oil and gas well full life cycle data is sparse and the working conditions are diverse, resulting in insufficient model generalization ability; traditional reinforcement learning (RL) cannot achieve fine continuous control due to the design of discrete action space, and lacks embedding of safety constraints (such as minimum opening to prevent blocking and pressure drop rate to prevent sand production), and exploratory actions may cause production accidents. In addition, the downhole communication bandwidth and computing power are limited, and the traditional cloud training mode is difficult to support millisecond-level real-time response. SUMMARY

[0004] To solve the above technical problems, the present application provides a choke valve control method for oil and gas wells based on deep reinforcement learning.

[0005] A choke valve control method for oil and gas wells based on deep reinforcement learning, comprising the following steps:

[0006] S1. Using a multilayer perceptron to fit a choke valve opening-flow characteristic model according to the detection data of the valve;

[0007] S2. Constructing an agent based on deep reinforcement learning of the fitted choke valve opening-flow characteristic model based on Markov decision, one valve corresponding to one agent;

[0008] S3. Solving the specific valve opening using the agent, thereby intelligently regulating the valve according to the user's production target.

[0009] Furthermore, the flow formula of the throttle valve opening-flow characteristic model is:

[0010]

[0011] Where Q is the measured water flow in cubic meters per hour; Δp v is the net differential pressure of the valve, in kPa; ρ is the density of water, in kilograms per cubic meter; ρ0 is the density of water at 15°C, in kilograms per cubic meter.

[0012] Furthermore, the fitting process of the multilayer perceptron in step S1 is specifically as follows:

[0013] The formula for the multi-layer perceptron is:

[0014] z (l) =xW (l) +b (l) ,

[0015] Where x is the input vector, l is the number of layers, in the first layer, x refers to the valve opening; W (l) is the hidden layer weight matrix, W (l) ∈R n×d , in the lth layer, n=l, d is the number of neurons in the hidden layer; z is the output of the lth layer, and R is the domain of all real numbers;

[0016] Use ReLU as the activation value of the hidden layer:

[0017] a (l) =σ(z (l) ),

[0018] Among them, σ represents the activation function ReLU(x)=max(0,x), a (l) is the output of the activation layer;

[0019] For the lth layer, the forward propagation formula is:

[0020] z (l) =a (l-1) W (l) +b (l) ,a (l) =σ(z (l) ),

[0021] The loss function is the mean square error MSE:

[0022]

[0023] Among them, t is the actual measured flow characteristic value, and y is the output of the multi-layer perceptron;

[0024] right Find the gradient:

[0025]

[0026] where, ⊙ represents element-wise multiplication, and σ' is the derivative of the activation function;

[0027] Gradient update is performed on the weight parameters:

[0028]

[0029] where, α is a learning rate, and is 1e-3.

[0030] Further, the step S2 specifically comprises:

[0031] S201. Constructing a throttle-regulated Markov chain, including construction of an action space, a state space, and a reward function;

[0032] Action space: defining the action of the agent at the t-th moment as The action space is the throttle opening;

[0033] State space: defining the state observation of the agent at the t-th moment as Including wellhead pressure, pressure after throttling, flow rate, and valve opening, denoted as: s t = [P 井口 ,P 节流后 ,Q 流量 ,a 开度 ], where a 开度 ∈ [0, 1] represents the current valve opening;

[0034] Reward function: comprehensive control accuracy, segmented control reward function designed for the control process: reward based on error setting:

[0035] r 误差 = w·(ΔQ 前一个 - ΔQ 当前

[0036]

[0037] R = r 误差 + r 达成 ,

[0038] S202. Constructing a single-agent deep reinforcement learning training regulation framework, in the training phase, the agent learns the state-action value function by interacting with the environment, and optimizes the regulation strategy; in the execution phase, the agent independently decides the throttle opening according to the real-time observed local state information.

[0039] Further, the training regulation framework in S202 includes an Actor network and a Critic network, the Actor network is responsible for generating action strategy, its input is the environment state s t , and the output is the action probability distribution π θ (a|s t ), that is, the mean and standard deviation of the output action, the goal is to update the parameters θ by the policy gradient method to maximize the cumulative reward; the input of the Critic network is the environment state s t , and the output is the state value estimation V φ (s t ), the goal is to update the parameters φ by the mean square error to minimize the difference between the predicted value and the actual return.

[0040] Further, the Actor network and the Critic network both use a multi-layer perceptron as the policy network, and the training process is as follows:

[0041] Randomize the parameters of the Actor and the Critic;

[0042] Then run multiple rounds in the set throttle environment, for each time step t, perform: select action a according to t ;

[0043] Record state s t , action a t , reward r t , next state s t+1 , and termination flag done t ; store the action probability of the old policy

[0044] After sampling, calculate the Monte Carlo return:

[0045]

[0046] Where γ is the decay coefficient, and r is the immediate reward obtained at the t+k step;

[0047] Estimate the state value V φ (s t ) using the Critic network;

[0048] Calculate the generalized advantage estimation:

[0049] δ t =r t +γV φ (s t+1 )-V φ (s t ),

[0050]

[0051] V φ (s t ) is the output of the value network, which refers to the expected cumulative return that can be obtained by following the current policy from s t at the beginning; r t is the immediate reward at time step t, and λ is a parameter used to calculate the advantage function; A t represents the evaluation value of the action;

[0052] The target value is calculated as:

[0053]

[0054] where, The target value function at time step t, which is the target value of the value function update.

[0055] Further, the policy gradient method estimates the gradient of the new policy by the sample of the old policy The construction process of the total target function is:

[0056] Calculate the probability ratio r t (θ):

[0057]

[0058] Get the Actor network target function:

[0059] L CLIP (θ) = E t [min(r t (θ)A t , clip(r t (θ), 1-∈, 1+∈)A t )],

[0060] Where L CLIP (θ) represents the target function of policy parameter θ; clip() represents the directional constraint of policy update; ∈ is a hyperparameter, usually taking values 0.1-0.3; E t represents the expectation operator; For the target function of Critic network:

[0061]

[0062] Entropy regularization term:

[0063] S[π θ ] = -∑ a π θ (a|s t )logπ θ (a|s t),

[0064] Total objective function:

[0065] L Total = L CLIP -c1L VF +c2E t [S[π θ ]],

[0066] After defining the loss function, the parameters of the Actor and Critic networks are updated respectively; for the Actor network, the goal is to maximize L CLIP , so gradient ascent is used to update the parameters θ:

[0067]

[0068] For the Critic network, the goal is to minimize L VF , so gradient descent is used to update the parameters φ:

[0069]

[0070] At the same time, the old policy parameters are updated:

[0071] θ old ← θ,

[0072] where θ represents the set of all trainable parameters in the Actor network, φ represents the set of all trainable parameters in the Critic network, and θ old represents the old policy parameters.

[0073] The beneficial effects of the present application are reflected in:

[0074] 1. The throttle valve intelligent control method of the present application can meet the individual needs of users;

[0075] 2. The present application carries out throttle valve simulation modeling, which fully reflects the production conditions on site;

[0076] 3. The present application is based on deep reinforcement learning, fully utilizes production data, and constructs a valve model to provide technical support for the development of smart oilfields;

[0077] 4. The present application generates complex conditions such as dynamic response to gas-liquid ratio mutation, sand production, and hydrate through real-time interaction between the deep reinforcement learning agent and the throttle valve flow model, and the regulation accuracy is significantly improved;

[0078] 5. The present application realizes the maximization of production efficiency through precise flow regulation. BRIEF DESCRIPTION OF DRAWINGS

[0079] Figure 1 is the flowchart of the method of the present application.

[0080] Figure 2 The process graph of the intelligent agent of the present application is regulated. DETAILED DESCRIPTION

[0081] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0082] In the present embodiment, referring to Figure 1 Fig. 1, a throttling valve regulation method for oil and gas wells based on deep reinforcement learning in an embodiment of the present application comprises:

[0083] First, a valve opening-flow characteristic model is constructed: according to the detection data of the valve, a multi-layer perceptron is used to fit an opening-flow characteristic curve model, wherein the formula of the multi-layer perceptron is as follows:

[0084] z (l) =xW (l) +b (l) , (1)

[0085] In (1), x is an input vector, and l is the number of layers. In the first layer, x refers to the valve opening. W (1) is a hidden layer weight matrix, W (l) ∈R n×d , n=1 in the first layer, and d is the number of hidden layer neurons; R is the entire real number field; z (l) is the output of the lth layer.

[0086] ReLU is used as the hidden layer activation value,

[0087] a (l) =σ(z (l) ), (2)

[0088] In (2), σ is ReLU(x)=max(0,x)

[0089] wherein, for the lth layer, the forward propagation formula is:

[0090] z (l) =a (l-1) W (l) +b (l) ,a (l) =σ(z (l) ), (3)

[0091] The loss function is mean square error (MSE):

[0092]

[0093] where t is the actual measured flow characteristic value, and y is the output of the multi-layer perception.

[0094] The gradient is calculated as:

[0095]

[0096] where is the element-wise multiplication, and σ' is the derivative of the activation function

[0097] The gradient update of the weight parameters is:

[0098]

[0099] α in (6) is the learning rate, which is 1e-3;

[0100] The flow coefficient calculation formula is as follows:

[0101]

[0102] Transforming (7) to obtain the flow formula as follows:

[0103]

[0104] In (8): Q is the measured water flow, with a unit of cubic meters per hour (m 3 / h); Δp v is the net differential pressure of the valve, with a unit of kilopascal (kPa); ρ is the density of water, with a unit of kilograms per cubic meter (kg / m 3 ); ρ0 is the water density at 15℃, with a unit of kilograms per cubic meter (kg / m 3 );

[0105] According to the above formula, the throttle valve model is constructed.

[0106] A Markov chain for regulating the production throttle valve of an oil and gas field is constructed, including the construction of an action space, a state space, and a reward function.

[0107] A single-agent deep reinforcement learning training and regulation framework is constructed. In the training phase, the agent learns the state-action value function by interacting with the environment and optimizes the regulation strategy. In the execution phase, the agent independently decides the throttle valve opening degree based on the real-time observed local state information, without relying on historical or other external information.

[0108] The state space construction process is as follows: the state observation of the agent at the t-th time is defined as ​The wellhead pressure, the throttling pressure, the flow rate and the valve opening degree are denoted as:

[0109] s t = [P 井口 , P 节流后 , Q 流量 , a 开度 ],

[0110] where a 开度 ∈ [0, 1] represents the current valve opening degree (0 for closed and 1 for fully open) ;

[0111] The action space construction process is to define the action of the agent at the t-th moment as The action space is the throttle opening degree, and the adjustment range is limited to [0.05, 0.95], that is: a t ∈ [0.05, 0.95].

[0112] The reward function design process is to comprehensively control the control accuracy, and the segmented control reward function is designed:

[0113] The reward based on the error is set as:

[0114] r 误差 = w · (ΔQ 前一个 - ΔQ 当前 ),

[0115]

[0116] R = r 误差 + r 达成 ,

[0117] For the construction of the training framework, the proximal policy optimization is selected as the main algorithm, and the principle is as follows:

[0118] First, there are two neural networks, which are Actor network and Critic network. The Actor is responsible for generating the action policy, and its input is the environment state s t , and the output is the action probability distribution π θ (a|s t ), that is, the mean and standard deviation of the output action. Its goal is to update the parameters θ by the policy gradient method to maximize the cumulative reward. For the Critic network, its input is the environment state s t , and the output is the state value estimate V φ (s t), the goal is to update the parameters φ by minimizing the difference between the predicted value and the actual return via mean squared error (MSE). Actor and Critic are fully separated and parameters are not shared. In this application, both Actor and Critic use multi-layer perceptron as policy network. The specific training steps are as follows:

[0119] Randomize the parameters of Actor and Critic;

[0120] Then run multiple episodes in the set throttle environment, for each time step t, perform: according to Select action a t .

[0121] Record state s t , action a t , reward r t , next state s t+1 , termination flag done t .

[0122] Store the action probability of the old policy

[0123] After sampling, calculate the Monte Carlo return:

[0124]

[0125] Estimate the state value V φ (s t ) using the Critic network;

[0126] Calculate the generalized advantage estimate:

[0127] δ t = r t + γV φ (s t+1 ) - V φ (s t ),

[0128]

[0129] V φ (s t ) is the output of the value network, which refers to the expected cumulative return that can be obtained starting from s t , following the current policy. r t is the immediate reward at time step t, and λ is a parameter used to calculate the advantage function;

[0130] Calculate the target value:

[0131]

[0132] wherein, The target value function at time step t, which is the target value for the value function update.

[0133] For policy gradient methods, the core is to directly optimize the policy parameters θ to maximize the expected return. However, using the new policy π θ Generating samples is inefficient, so it is desirable to reuse old policy samples to estimate the gradient of the new policy. Therefore, the objective function is constructed as follows:

[0134] Calculate the probability ratio:

[0135]

[0136] Objective function:

[0137] L CLIP (θ)=E t [min(r t (θ)A t ,clip(r t (θ),1-∈,1+∈)A t )],

[0138] Objective function for Critic network:

[0139]

[0140] Entropy regularization term:

[0141] S[π θ ]=-∑ a π θ (a|s t )logπ θ (a|s t ),

[0142] Total objective function:

[0143] L Total =L CLIP -c1L VF +c2E t [S[π θ ]],

[0144] After defining the loss function, update the parameters of the Actor and Critic networks respectively; for the Actor network, the goal is to maximize L CLIP , so use gradient ascent to update the parameters θ:

[0145] For the Critic network, the goal is to minimize L VF , so use gradient descent to update the parameters φ:

[0146] At the same time, update the old policy parameters: θ old ←θ; where θ represents the set of all trainable parameters in the Actor network, φ represents the set of all trainable parameters in the Critic network, and θ old Indicates the old policy parameters.

[0147] The hyperparameter settings are shown in Table 1:

[0148] Table 1

[0149] Parameter Value Total number of sampling steps 5e+5 Learning rate a 1e-3 Gamma coefficient 0.99 Advantage coefficient 0.95 Iteration rounds 10 Shear coefficient 0.3 Entropy coefficient 0.1

[0150] After the intelligent agent is trained, the system is regulated. The environmental parameters include the set opening, feedback opening, target flow, and pressure difference before and after throttling. During the regulation process, the target flow is set to Q target ∈[0,50]. Set the upper limit of valve action per day to 48 times, and get the valve opening every half hour based on the demand side response. The intelligent agent control process is as follows Figure 2 shown.

[0151] This solution solves the problem of poor interpretability of deep reinforcement learning (DRL) models by constructing a logical interpretation module for state-action mapping and clarifying the impact of key parameters (such as pressure gradient) on throttle valve opening decisions.

[0152] In the description of the embodiments of the present invention, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly specifying the number of the technical features indicated. Therefore, a feature specified as "first," "second," "third," or "fourth" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0153] In describing the embodiments of the present invention, the term "and / or" is used herein to describe the association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " is generally used herein to indicate that the associated objects are in an "or" relationship.

[0154] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A throttle valve control method for oil and gas wells based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Use a multi-layer perceptron to fit the throttle valve opening-flow characteristic model based on the valve detection data. S2. Build an intelligent agent that performs deep reinforcement learning on the fitted throttle valve opening-flow characteristic model based on Markov decision making, with one intelligent agent corresponding to each valve; S3. Use the intelligent agent to solve the specific valve opening, so as to perform intelligent valve control according to the user's production goals.

2. The oil and gas well throttle valve control method based on deep reinforcement learning according to claim 1 is characterized in that: The flow formula of the throttle valve opening-flow characteristic model is: Where Q is the measured water flow in cubic meters per hour; Δp v is the net differential pressure of the valve, in kPa; ρ is the density of water, in kilograms per cubic meter; ρ0 is the density of water at 15°C, in kilograms per cubic meter.

3. The oil and gas well throttle valve control method based on deep reinforcement learning according to claim 1 is characterized in that: The specific fitting process of the multilayer perceptron in step S1 is as follows: The formula for the multi-layer perceptron is: z (l) =xW (l) +b (l) , Where x is the input vector, l is the number of layers, in the first layer, x refers to the valve opening; W (l) is the hidden layer weight matrix, W (l) ∈R n×d , in the lth layer, n=l, d is the number of neurons in the hidden layer; z is the output of the lth layer, and R is the domain of all real numbers; Use ReLU as the activation value of the hidden layer: a (l) =σ(z (l) ), Among them, σ represents the activation function ReLU(x)=max(0,x), a (l) is the output of the activation layer; For the lth layer, the forward propagation formula is: With (l) =a (l-1) IN (l) +b (l) ,and (l) =σ(z (l) ), The loss function is the mean square error MSE: Among them, t is the actual measured flow characteristic value, and y is the output of the multi-layer perceptron; right Find the gradient: Among them, ⊙ represents element-by-element multiplication, σ ' is the derivative of the activation function; Perform gradient updates on weight parameters: Among them, α is the learning rate, which is 1e-3.

4. The oil and gas well throttle valve control method based on deep reinforcement learning according to claim 1 or 2, characterized in that: The step S2 specifically includes: S201. Construct a Markov chain for throttle regulation, including the construction of action space, state space, and reward function; Action space: Define the action of the agent at time t as The action space is the throttle valve opening; State space: Define the state observation of the agent at time t as Including wellhead pressure, pressure after throttling, flow rate and valve opening, recorded as: s t =[P 井口 ,P 节流后 ,Q 流量 ,a 开度 ], where a 开度 ∈[0,1] represents the current valve opening; Reward function: comprehensive control accuracy, segmented control of control process design Reward function: reward based on error setting: r 误差 =w·(ΔQ 前一个 -ΔQ 当前 ) R=r 误差 +r 达成 , S202. Construct a single-agent deep reinforcement learning training and control framework. During the training phase, the agent learns the state-action value function by interacting with the environment and optimizes the adjustment strategy. During the execution phase, the agent independently decides the throttle valve opening based on the local state information observed in real time.

5. The oil and gas well throttle valve control method based on deep reinforcement learning according to claim 4 is characterized in that: The training control framework in S202 includes an Actor network and a Critic network. The Actor network is responsible for generating action strategies, and its input is the environment state s. t , the output is the action probability distribution π θ (a|s t ) is the mean and standard deviation of the output action, and its goal is to update the parameter θ through the policy gradient method to maximize the cumulative reward; the input of the Critic network is the environment state s t , the output is the state value estimate V φ (s t ), the goal is to update the parameter φ by the mean squared error to minimize the difference between the predicted value and the actual return.

6. The oil and gas well throttle valve control method based on deep reinforcement learning according to claim 5, characterized in that: The Actor network and Critic network both use multi-layer perceptrons as policy networks, and their training process is as follows: Randomize the parameters of Actor and Critic; Then run multiple rounds in the set throttling environment, and for each time step t, execute: Select action a t ; Record Status t , action a t , reward r t , the next state s t+1 , termination flag done t ;Store the action probability of the old strategy After sampling, calculate the Monte Carlo return: Where γ is the decay coefficient, r is the immediate reward obtained at step t+k; Use the Critic network to estimate the state value V φ (s t ); Compute the generalized odds estimate: δ t =r t +γV φ (s t+1 )-V φ (s t ), V φ (s t ) is the output of the value network, which refers to the t Initially, the expected cumulative return that can be obtained by following the current strategy; r t is the immediate reward at time step t, λ is the parameter used to calculate the advantage function; A t Indicates the evaluation value of the quality of the action; Calculate the target value: in, The target value function at time step t, which is the target value of the value function update.

7. The oil and gas well throttle valve control method based on deep reinforcement learning according to claim 6, characterized in that: The policy gradient method passes the old policy The sample is used to estimate the gradient of the new strategy. The construction process of the total objective function is: Calculate the probability ratio r t (θ): Get the Actor network objective function: L CLIP (θ)=E t [min(r t (i)A t ,clip(r t (θ),1-∈,1+∈)A t )], Among them, L CLIP (θ) represents the objective function of the policy parameter θ; clip() represents the directional constraint of the policy update; ∈ is a hyperparameter, usually with a value of 0.1 to 0.3; E t Represents the expectation operator; for the objective function of the Critic network: Entropy regularization term: S[π θ ]=-∑ a p θ (a|s t )logπ θ (a|s t ), Overall objective function: IT Total =L CLIP -c1L VF +c2E t [S[π θ ]], After defining the loss function, the parameters of the Actor and Critic networks are updated separately; for the Actor network, the goal is to maximize L CLIP , so we use gradient ascent to update the parameter θ: For the Critic network, the goal is to minimize L VF , so gradient descent is used to update the parameter φ: At the same time, update the old policy parameters: i old ←θ, Among them, θ represents the set of all trainable parameters in the Actor network, φ represents the set of all trainable parameters in the Critic network, and θ old Indicates the old policy parameters.

Citation Information

Cited By

  • Gas well group production pressure difference control method based on reinforcement learning and safety constraint projection

    CN121277006A