Well drilling key parameter automatic regulation and control method based on improved DDPG algorithm

By applying a reinforced learning model based on DDPG algorithm during drilling, dynamically adjusting the drilling pressure and rotation speed, the problem that the existing technology is difficult to take into account multiple indicators in complex formation environments, and the improvement of drilling efficiency and safety is achieved.

CN119937305APending Publication Date: 2025-05-06XI'AN PETROLEUM UNIVERSITY
View PDF 0 Cites 19 Cited by

Patent Information

Application Number
CN202510012266.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing drilling parameter control methods are difficult to take into account multiple indicators such as drilling speed, drill bit life and wellbore quality in complex formation environments, and have poor real-time and insufficient adaptability.

Method used

The automated regulation method of key drilling parameters based on the depth deterministic strategy gradient (DDPG) algorithm is adopted. By building a reinforcement learning model, drilling pressure and speed are dynamically adjusted, and multiple target needs such as drilling efficiency, drilling bit life and wellbore quality are comprehensively considered.

Benefits of technology

It significantly improves the intelligence level and economic benefits of the drilling process, realizes the optimization of drilling speed, ensures the quality and operational safety of the wellbore, and adapts to the needs of complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937305A_ABST
    Figure CN119937305A_ABST
Patent Text Reader

Abstract

The invention discloses a drilling key parameter automatic regulation and control method based on an improved DDPG algorithm. The method comprises the following steps: collecting and preprocessing drilling data, and constructing a complete data set for training; a drilling speed (ROP) time sequence prediction model is constructed and trained on the basis of a self-attention mechanism (self-attention); building a simulation environment by using the trained ROP prediction model, and defining a state space, an action space and a multi-target reward function of a reinforcement learning DDPG algorithm model; a deep deterministic strategy gradient algorithm (DDPG) is adopted to train a reinforcement learning model, through an Actor-Critic network structure, an intelligent agent learns an optimal bit pressure (WOB) and rotating speed (RPM) regulation strategy in a simulation environment, and dynamic optimization of the drilling speed (ROP) is achieved. According to the method, through automatic and intelligent parameter regulation and control, the efficiency and stability of the drilling process are improved, the drill bit wear rate and the operation cost are reduced, and the method has good adaptability and multi-target optimization capacity and is suitable for drilling operation under the complex stratum condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of petroleum engineering, in particular to a key parameter control method in a drilling process. Background Art

[0002] In drilling engineering, weight on bit (WOB) and rotation speed (RPM) are key parameters that affect the rate of penetration (ROP). Engineers usually manually adjust these parameters based on real-time drilling feedback to optimize drilling efficiency. However, the traditional manual control method has the following problems: first, it relies on the experience of the operator and is difficult to maintain a stable optimization effect under complex formation conditions; second, the adjustment process lacks real-time and precision, which may lead to a decrease in drilling efficiency or the occurrence of complex situations downhole.

[0003] With the development of intelligent drilling technology, automated control systems have gradually been introduced into drilling sites. Existing automated control methods are mainly based on rule models or optimization algorithms. For example, Self, R proposed using the PSO algorithm to find the best parameter combination in the paper "Reducing drilling cost by finding optimal operational parameters using particle swarm algorithm". These methods can achieve certain optimization effects in specific scenarios, but when faced with complex formation environments and multi-objective constraints, it is difficult to take into account multiple indicators such as drilling speed, drill bit life and wellbore quality. In addition, traditional intelligent optimization algorithms require recalculation every time the parameters are adjusted, and cannot effectively use previous learning experience for rapid response, resulting in poor real-time performance and insufficient adaptability in rapidly changing drilling environments.

[0004] In recent years, the rapid development of artificial intelligence technology has provided new solutions for the automated control of drilling parameters. Reinforcement Learning (RL), as a data-driven optimization technology, can learn the optimal strategy in real time through the continuous interaction between intelligent agents and the environment. Its core advantage is that it can directly integrate real-time drilling feedback into the decision-making process and dynamically optimize parameter settings. However, when applied to drilling parameter control, the existing reinforcement learning methods have not fully considered the combination of engineering constraints and formation complexity, resulting in the lack of operability of their optimization results in actual engineering, and no research inventions in related fields have been formed.

[0005] Based on this, there is an urgent need for an intelligent method that can integrate multi-objective constraints and dynamically adjust drilling pressure and rotation speed to improve drilling efficiency, optimize drilling speed, and ensure the safety and economy of downhole operations. Summary of the invention

[0006] In order to overcome the defects of poor real-time performance, insufficient adaptability and limited multi-objective optimization capability of existing drilling parameter control methods, this paper proposes an automatic control method for key drilling parameters based on the Deep Deterministic Policy Gradient (DDPG) algorithm, aiming to achieve dynamic optimization and control of key parameters such as bit weight (WOB) and rotation speed (RPM) through intelligent algorithms to improve the drilling rate (ROP) and ensure wellbore quality and operational safety. This method constructs a reinforcement learning DDPG algorithm model, dynamically adjusts drilling parameters, and comprehensively considers multi-objective requirements such as drilling efficiency, drill bit life and wellbore quality, significantly improving the intelligence level and economic benefits of the drilling process.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0008] An automatic control method for key drilling parameters based on an improved DDPG algorithm comprises the following steps:

[0009] 1) Collect and preprocess drilling data to construct a training data set for training the ROP time series prediction model;

[0010] 2) Build and train a ROP time series prediction model to provide accurate dynamic response, which is used as the environmental feedback basis for the reinforcement learning DDPG algorithm model;

[0011] 3) Build a simulation environment based on the ROP prediction model in step 2), define the state space, action space and reward function of the reinforcement learning DDPG algorithm model, and design an improved DDPG algorithm model;

[0012] 4) Use the simulation environment to train the DDPG algorithm model and learn the optimal parameter control strategy through intelligent agents.

[0013] The drilling parameters in step 1) include weight on bit (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), mud inflow (MFI) and large trench load (HKL), etc., and the target parameter is rate of penetration (ROP). The data preprocessing process includes:

[0014] 1.1) Perform preliminary cleaning of the collected data, remove data entries with missing value ratio exceeding 10%, remove abnormal data exceeding three times the standard deviation, and perform linear interpolation on the missing data;

[0015] 1.2) Normalize the data to the maximum and minimum values, the formula is:

[0016]

[0017] Among them, X normalized is the normalized data; X is the original data; min(X) is the minimum value of the original data; max(X) is the maximum value of the original data.

[0018] In the step 2), a deep temporal neural network model is constructed using the training set data in step S1 to predict the ROP value, which specifically includes the following sub-steps:

[0019] 2.1) Construct a deep neural network model, which uses the self-attention framework as the core architecture, taking advantage of its advantages in processing time series data to capture the dynamic laws of drilling parameters and formation characteristics over time. The model includes a neural network architecture of an encoder and a decoder based on the self-attention mechanism to extract multi-level time series features, and sets a fully connected layer at the output to generate the predicted ROP value.

[0020] 2.2) Set input features and output targets. The input features of the model include weight on bit (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), mud inflow (MFI) and large trench load (HKL); the output target is the rate of penetration (ROP) at the corresponding moment. The input features are normalized to accelerate training and improve the generalization performance of the model.

[0021] 2.3) Time series prediction model loss function and optimizer settings. The model training process uses mean square error (MSE) as the loss function, and the formula is as follows:

[0022]

[0023] Among them, y i is the real ROP value, is the ROP value predicted by the model, and N is the number of samples. During the training optimization process, the Adam optimizer is used to update the parameters.

[0024] 2.4) When the test set evaluation results meet expectations, save the optimal state of the model.

[0025] In step 3), based on the ROP timing prediction model trained in step 2), a simulation environment is constructed, and the key elements of the reinforcement learning DDPG algorithm model are defined, including the following:

[0026] 3.1) Construct a simulation environment with the ROP prediction model as the core to simulate the impact of weight on bit (WOB) and rotation speed (RPM) on the rate of penetration (ROP). The input of the simulation environment includes the underlying characteristic parameters such as weight on bit (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), mud inflow (MFI) and large trench load (HKL). Since these underlying characteristic parameters cannot be directly manipulated by reinforcement learning, they are used as the state input of the environment to affect the prediction of the rate of penetration in the simulation environment and the subsequent optimization process. The simulation environment calculates the rate of penetration (ROP) through the ROP prediction model and provides feedback indicators of the rate of penetration under the reward and punishment mechanism for the optimization of the DDPG algorithm model.

[0027] The constraints in the simulation environment include the physical operating limitations of drilling pressure and rotation speed (such as the maximum allowable drilling pressure and rotation speed of the equipment), the limit of drill bit wear rate and the requirements of wellbore quality. These constraints ensure that the control scheme generated by the model is feasible and can effectively improve the drilling speed in actual operation.

[0028] 3.2) Define the key elements of the DDPG algorithm model, which include presetting the action space and setting the reward function. The action space includes the adjustment of drilling pressure and speed, which are the controllable targets of the DDPG algorithm model. The model optimizes the drilling rate (ROP) by increasing, decreasing or maintaining these parameters unchanged.

[0029] The reward function is designed as a multi-objective optimization, taking into account ROP improvement, adjustment cost and satisfaction of constraints. The specific reward function form is:

[0030] Reward=α·ROP gain -β·Adjustment cost -γ·Constraint penalty

[0031] Among them, ROP gain It indicates the increase in drilling speed at the current time step t compared to the previous time step t-1, multiplied by the weight coefficient α, to encourage the increase in drilling speed; cost Represents the adjustment range of drilling pressure and speed at the current time step, multiplied by the weight coefficient β, to punish excessive parameter adjustments and reduce operating costs and equipment wear; Constraint penaltyIt is used to measure whether the engineering constraints (such as physical limitations of drilling pressure and speed, drill bit wear rate, wellbore quality, etc.) are violated, and multiplied by a higher weight coefficient γ to ensure strict compliance with safety and performance standards. Through this reward mechanism, the intelligent agent is not only incentivized to increase the drilling speed, but also constrained to operate within a reasonable adjustment range, and severely punishes any behavior that violates the engineering constraints, thereby achieving a balanced optimization of efficiency, safety and economic benefits during the drilling process.

[0032] 3.3) Construct the architecture of the reinforcement learning model, which uses the Deep Deterministic Policy Gradient (DDPG) algorithm and adopts the structure of the policy network (Actor) and the value network (Critic) to optimize the control strategy of the drilling pressure (WOB) and the rotation speed (RPM). Compared with the traditional DDPG algorithm model, the present invention is optimized and improved in the following aspects:

[0033] 3.3.1) This model adds an extra fully connected layer to the traditional policy network, making the network depth reach four layers. The introduction of this extra layer enables the network to better understand and process multi-dimensional data in the drilling process, improving the accuracy and stability of the strategy;

[0034] 3.3.2) In the value network design, the present invention introduces residual connections. Specifically, by adding two linear layers (residual1 and residual2), the input features are directly added to the output of the hidden layer. This residual connection effectively alleviates the gradient vanishing problem in deep networks, promotes a more stable and efficient training process, and improves the accuracy of value assessment;

[0035] 3.3.3) The model uses normal distribution (mean 0, standard deviation 0.1) for weight initialization in each layer of the policy network and value network. This initialization method helps to avoid gradient disappearance or explosion, allowing the network to maintain good performance in the early stages of training;

[0036] 3.3.4) In order to prevent the model from overfitting and improve the generalization ability, the present invention adds a Dropout layer (p=0.1) after the first hidden layer of the strategy network and the key hidden layer of the value network. The Dropout layer reduces the network's dependence on specific neurons by randomly discarding the output of some neurons, thereby enhancing the adaptability and robustness of the model in different drilling environments;

[0037] 3.3.5) The output of the policy network is limited to the range of [-1, 1] through the tanh activation function, and then the action value is scaled to the actual action space range of drilling pressure [WOB_low, WOB_high] and speed [RPM_low, RPM_high] through linear transformation. The specific formula is:

[0038]

[0039] Action represents the value of the current action. This design ensures that the output drilling pressure and rotation speed are within a reasonable and operable range, avoiding the generation of invalid or extreme action values, thereby improving the feasibility and safety of the control strategy.

[0040] In step 4), the simulation environment constructed in step 3) and the reinforcement learning DDPG algorithm model are used to train the model to learn the optimal parameter control strategy, which specifically includes the following contents:

[0041] 4.1) Construct the training process and use the deep reinforcement learning deep deterministic policy gradient algorithm (DDPG) algorithm to train the model. During the training process, the intelligent agent interacts with the simulation environment and at each time step t, according to the current state s t Select an action t , and get the next state s t+1 and instant rewards t . Accumulate experience data by constantly interacting with the environment.

[0042] 4.2) Establishment and update of experience pool: Experience Replay Buffer is introduced to store the interaction data of intelligent agents, including the four-tuple of state, action, reward and next state (s t , a t , r t ,s t+1 ). Each time the strategy is updated, a mini-batch of data is randomly extracted from the experience pool for training to break the data correlation and improve the stability and efficiency of training.

[0043] 4.3) Update model parameters, use the policy network and value network to update model parameters. The goal of the value network is to estimate the expected cumulative reward under a given state and action. The loss function of the value network is defined as:

[0044]

[0045] in,

[0046] y i =r i +γQ′(si+1 ,μ′(s i+1 |θ μ′ )|θ Q′ )

[0047] γ is the discount factor, Q′ and μ′ are the parameters of the target value network and the target policy network respectively. By minimizing the loss function L Q , use gradient descent to update the value network parameters θ Q .

[0048] The role of the policy network is to output the optimal action (i.e., the adjustment value of drilling pressure and rotation speed) according to the current state, with the goal of maximizing the cumulative reward of the intelligent agent in the environment. The policy network directly determines the control strategy that the intelligent agent should adopt in different states by learning the mapping relationship from state to action.

[0049] The update goal of the policy network is to improve the value of the actions selected in various states. The parameter update of the policy network is achieved through the policy gradient method, and its gradient calculation formula is:

[0050]

[0051] where θ μ is the parameter of the policy network, μ(s i |θ μ ) indicates that the policy network is in state s i The output action, Q(s, a|θ Q ) is the value network, which evaluates the value of a in performing an action in state s. By using the policy gradient method, the policy network parameters θ are updated. μ , so that the policy network can output actions that maximize the expected cumulative reward in a given state.

[0052] 4.4) Model convergence and strategy optimization. During the training process, the model convergence is judged by monitoring the maximum cumulative reward and the stability of the strategy. When the cumulative reward reaches the expected level on the validation set and the strategy output tends to be stable, the model is considered to have converged. At this time, the optimal model parameters are saved for subsequent practical applications.

[0053] The trained DDPG algorithm model can output the optimal weight on bit (WOB) and rotation speed (RPM) adjustment strategy according to the current drilling status, optimize the rate of penetration (ROP), improve drilling efficiency and ensure drilling safety.

[0054] Compared with the prior art, the advantages of the present invention are mainly reflected in:

[0055] (1) Realize intelligent control of key parameters. The present invention realizes automatic control of bit weight (WOB) and rotation speed (RPM) through an improved DDPG algorithm model, and can dynamically adjust key parameters according to the real-time drilling status without manual intervention, greatly improving the automation level of drilling operations.

[0056] (2) Improve the accuracy and efficiency of ROP optimization. Combining the ROP time series prediction model and the DDPG algorithm, the present invention can predict the ROP in real time under different working conditions, and maximize the ROP improvement effect through intelligent regulation, significantly improving drilling efficiency.

[0057] (3) Multi-objective optimization capability. The present invention designs a multi-objective reward function that comprehensively considers drilling speed improvement, parameter adjustment cost, and safety constraints. It can improve drilling speed while ensuring drill bit life and wellbore quality, thus meeting a variety of engineering requirements.

[0058] (4) Enhanced adaptability to complex working conditions: By constructing a simulation environment and improving the dynamic optimization of the DDPG algorithm model, the present invention can operate stably under different formation conditions, equipment configurations, and operating constraints, and adapt to complex drilling conditions.

[0059] (5) Reduce manual intervention and operational risks. In traditional drilling operations, the adjustment of drilling pressure and rotation speed depends on manual experience, which is prone to misoperation and low efficiency. The present invention significantly reduces the need for manual intervention and reduces operational risks through the intelligent decision-making ability of the improved reinforcement learning DDPG algorithm model.

[0060] In summary, the present invention provides an efficient, stable and safe method for automatic control of key drilling parameters through intelligent algorithms and reinforcement learning DDPG algorithm models, which can significantly improve drilling efficiency, reduce operational risks, and adapt to complex working conditions, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flow chart of the automatic control method of drilling parameters based on the improved DDPG algorithm.

[0062] Figure 2 Schematic diagram of the structure of a neural network based on self-attention.

[0063] Figure 3 Schematic diagram of the neural network structure of the reinforcement learning DDPG algorithm.

[0064] Figure 4 Comparison chart of ROP results for improving the DDPG algorithm. DETAILED DESCRIPTION

[0065] Example 1

[0066] See also Figure 1 The automatic control method of key drilling parameters based on the improved DDPG algorithm of the present invention comprises the following steps:

[0067] 1) Collect and preprocess drilling data to construct a training data set for training the ROP time series prediction model;

[0068] 2) Build and train a ROP time series prediction model, and input the data in step 1) into the model to provide accurate dynamic response, which is used as the environmental feedback basis for the reinforcement learning DDPG algorithm model;

[0069] 3) Build a simulation environment based on the ROP prediction model in step 2), define the state space, action space and reward function of the reinforcement learning DDPG algorithm model, and design an improved DDPG algorithm model;

[0070] 4) Use the ROP simulation environment in step 3) to train the reinforcement learning DDPG algorithm model, learn the optimal parameter control strategy through the intelligent agent, and finally save the model for application.

[0071] In this embodiment, in step 1), the key parameter data related to the drilling process are collected in real time through sensors and monitoring systems at the drilling site. The drilling operation parameters include: bit pressure (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), drilling time (BDT), mud inflow (MFI) and large trench load (HKL); the target parameter is drilling rate (ROP). The collected data includes historical data and real-time drilling data, which are used to form a complete training data set. The collected data is then preprocessed, and the preprocessing process is as follows:

[0072] 1.1) Perform preliminary cleaning on the collected data, remove data entries with missing value ratio exceeding 10%, remove abnormal data exceeding three times the standard deviation, and ensure the quality and consistency of the remaining data. Table 1 shows some data samples after processing:

[0073] Table 1 ROP logging sample data example

[0074]

[0075] 1.2) Perform data interpolation on subsequent data, use linear interpolation to fill missing values, and use data from adjacent moments to interpolate missing data within a specific time interval.

[0076] 1.3) Perform Min-Max Normalization on the data. The normalization formula is as follows:

[0077]

[0078] Among them, X normalized is the normalized data; X is the original data; min(X) is the minimum value of the original data; max(X) is the maximum value of the original data.

[0079] In this embodiment, in step 2, a self-attention-based ROP time series prediction model is constructed using the training set data to capture the dynamic laws of drilling parameters and formation characteristics changing over time. Figure 2 As shown in the figure, the main process of model construction is as follows:

[0080] 2.1) The prediction model structure design includes: the encoder layer and the decoder layer with self-attention mechanism as the core architecture, and the fully connected output layer: mapping the extracted features to the prediction target, i.e., the drilling rate (ROP).

[0081] 2.2) Loss function and model optimization settings. The mean square error (MSE) is used as the loss function in model training, and its formula is:

[0082]

[0083] Among them, y i is the real ROP value, is the ROP value predicted by the model, and N is the number of samples. During the training optimization process, the Adam optimizer was used to update the parameters, with an initial learning rate of 0.001, and the learning rate was adjusted by a decay factor of 0.9 after every 10 training cycles.

[0084] 2.3) Training process design: The model was trained for 100 rounds using the small batch gradient descent method (batch size = 32), and the ROP value after automatic parameter adjustment was improved accordingly on the validation set.

[0085] In this embodiment, in step 3), based on the ROP timing prediction model trained in step 2), a simulation environment is constructed for training and optimizing the DDPG model, referring to Figure 3 , specifically including the following:

[0086] 3.1) Construction of simulation environment. The simulation environment takes the trained ROP prediction model as the core and simulates the dynamic effects of key parameters such as bit pressure (WOB) and rotation speed (RPM) on the drilling rate (ROP). The input features of the simulation environment include the current bit pressure (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), mud inflow (MFI) and large trench load (HKL) and other underlying characteristic parameters. These underlying characteristic parameters, as the state input of the environment, directly affect the drilling rate prediction results in the simulation environment, and provide a basis for the strategy optimization of the DDPG algorithm model. Since these underlying characteristic parameters cannot be directly controlled by the DDPG model, the reinforcement learning DDPG algorithm mainly learns and optimizes the adjustment strategies of bit pressure and rotation speed.

[0087] 3.2) Based on the simulation environment, the basic conditions of the reinforcement learning process are defined, including the state space, action space, and reward function. The state space (State, s) includes the current drilling pressure (WOB) and rotation speed (RPM), two adjustable characteristic parameters, other unadjustable parameters, and historical drilling speed (WOB) data. Through the definition of the state space, the reinforcement learning DDPG algorithm model can fully perceive the current drilling conditions and provide sufficient information for subsequent decision-making.

[0088] The action space (Action, a) is defined as the adjustment value of drilling pressure and speed, including increasing, decreasing or keeping drilling pressure and speed unchanged, and imposing physical operation restrictions on drilling pressure and speed. The reinforcement learning DDPG algorithm model outputs the optimal action under different states through learning, thereby optimizing the drilling rate (ROP) and ensuring the wellbore quality and drill bit life. The reward function (Reward, r) is designed as a multi-objective optimization function, which comprehensively considers the drilling rate increase (ROP gain), parameter adjustment cost (Adjustment cost) and constraint satisfaction (Constraint penalty). The specific form of the reward function is as follows:

[0089] Reward=α·ROP gain -β·Adjustment cost -γ·Constraint penalty

[0090] Among them, ROP gain Adjustment is the increase in drilling speed. cost is the adjustment cost, Constraint penalty is the penalty term for violating the constraint, and α, β, and γ are weight coefficients used to balance the objectives.

[0091] 3.3) The reinforcement learning model uses the deep reinforcement learning algorithm DDPG (Deep Deterministic Policy Gradient), and its architecture includes a policy network (Actor) and a value network (Critic). The policy network is used to output the optimal adjustment action (i.e., the adjustment value of drilling pressure and speed) according to the current state, while the value network is used to evaluate the cumulative reward of each state-action pair to help the policy network optimize the control strategy.

[0092] Both the policy network and the value network adopt a fully connected layer structure, and the activation function is ReLU (Rectified Linear Unit), and the calculation formula is:

[0093] ReLU(x)=max(0,x)

[0094] In the model architecture, the policy network receives the current state information of the simulation environment (such as drilling pressure, rotation speed, mud flow, wellbore depth and historical ROP value, etc.), and outputs two continuous values, corresponding to the adjustment values ​​of drilling pressure and rotation speed. The value network receives the current state and the action generated by the policy network, and calculates the value estimate of the state-action pair (i.e., cumulative reward). The policy network and the value network continuously cooperate and optimize with each other, so that the model can gradually learn the optimal parameter control strategy.

[0095] In this embodiment, in step 4), the simulation environment and the reinforcement learning DDPG algorithm model are used for training. The simulation environment and the reinforcement learning DDPG algorithm model constructed in step S3 are used to train the model to learn the optimal pressure on bit (WOB) and speed (RPM) control strategy, thereby achieving dynamic optimization of the drilling speed (ROP). The specific implementation process is as follows:

[0096] 4.1) The training process uses the deep deterministic policy gradient algorithm (DDPG) to gradually learn the optimal parameter control strategy through the continuous interaction between the reinforcement learning agent and the simulation environment. At each time step t, the agent adjusts the state s according to the current state. t Generate an action a through the policy network t , which corresponds to the adjustment value of drilling pressure and speed. The simulation environment is based on the input action a t and the current state s t Update status to s t+1 , and generate an immediate reward r t The calculation of the reward value takes into account the drilling speed increase, parameter adjustment cost, and the satisfaction of the constraints, thereby guiding the behavior optimization of the intelligent agent. Through this continuous interaction, the intelligent agent accumulates experience data and provides training samples for subsequent model optimization.

[0097] 4.2) Establishment and update of experience pool During the training process, the experience replay buffer is introduced to store the interaction data between the intelligent agent and the simulation environment. The data stored in the experience pool is recorded in the form of four tuples, including state, action, immediate reward and next state (s t , a t , r t ,s t+1 ). Each time the model parameters are updated, a mini-batch of data is randomly extracted from the experience pool for training the policy network and the value network. By breaking data correlation through random sampling, the experience pool improves the stability and efficiency of training. The capacity of the experience pool is usually set to a fixed size, such as 10,000 interaction records, and the latest data will overwrite the oldest data to ensure that training is always based on the latest interaction information.

[0098] 4.3) Update setting of model parameters. During the training process, the parameters of the policy network and the value network are updated through different optimization objectives. The value network aims to estimate the cumulative reward under the current state and action, and the loss function is defined as:

[0099]

[0100] in,

[0101] y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ )

[0102] s is the state at a certain moment, and α is the action in the current state. i Denotes the target cumulative reward γ as the discount factor, Q′ and μ′ are the parameters of the target value network and the target policy network, respectively. By minimizing the loss function L Q , use gradient descent to update the value network parameters θ Q .

[0103] The policy network is responsible for generating the optimal action a based on the current state s. Its optimization goal is to maximize the cumulative reward of the intelligent agent. The policy network updates parameters through the policy gradient method, and its gradient formula is:

[0104]

[0105] where θ μ is the parameter of the policy network, μ(s i |θ μ ) indicates that the policy network is in state s iThe output action, Q(s, a|θ Q ) is the value network, which evaluates the value of a in performing an action in state s. By using the policy gradient method, the policy network parameters θ are updated. μ , so that the policy network can output actions that maximize the expected cumulative reward in a given state.

[0106] 4.4) Model Convergence and Strategy Optimization During the training process, the convergence of the model is judged by monitoring the stability of the cumulative rewards and strategies. When the performance of the intelligent agent in the simulation environment tends to be stable, the cumulative reward curve reaches the expected level, and the action output of the policy network changes little in multiple rounds of interaction, it means that the model has converged. In this case, the optimal parameters of the trained policy network and value network are saved for subsequent practical applications.

[0107] 4.5) In the current example, the model generates the optimal WOB and RPM adjustment strategy based on the real-time drilling status (such as current drilling pressure, rotation speed, mud flow and other parameters), thereby optimizing the drilling rate (ROP). Figure 4 ,Through continuous regulation under different working conditions, the improved DDPG algorithm model of the present invention can significantly improve the drilling efficiency.

Claims

1. A method for automatic control of key drilling parameters based on an improved DDPG algorithm, characterized in that: The following steps are involved: 1) Collect and preprocess drilling data to build a complete data set for training; 2) Build and train a self-attention-based neural network structure drilling rate (ROP) time series prediction model; 3) Build a simulation environment based on the ROP time series prediction model in 2), define the state space, action space and reward function of the reinforcement learning DDPG algorithm model, and design an improved DDPG algorithm network structure model; 4) Using the simulation environment of 3) to interactively learn the reinforcement learning agent decision model, the optimal weight on bit (WOB) and rotation speed (RPM) control strategy is learned by the intelligent agent to achieve dynamic optimization of the drilling rate (ROP).

2. The method according to claim 1, characterized in that The pretreatment in step 1) includes: 1.1) Perform preliminary cleaning of the collected data, remove data entries with missing value ratio exceeding 10%, remove abnormal data exceeding three times the standard deviation, and perform linear interpolation on the missing data; 1.2) Normalize the data to the maximum and minimum values, the formula is: Among them, X normalized is the normalized data; X is the original data; min(X) is the minimum value of the original data; max(X) is the maximum value of the original data.

3. The method according to claim 1, characterized in that The step 2) further comprises: 2.1) Design the prediction model structure, including the encoder and decoder layers based on the self-attention mechanism to extract multi-level time series features, and the fully connected output layer to generate the predicted drilling rate (ROP) value; 2.2) Set the loss function to mean square error (MSE), and use Adam optimizer to optimize model parameters. The initial learning rate is set to 0.001, and the learning rate is adjusted by a decay factor of 0.9 after every 10 training cycles: 2.3) The model was trained for 100 rounds using mini-batch gradient descent (batch size = 32), and the model performance was evaluated on the validation set, saving the optimal model state during the training process.

4. The method according to claim 1, characterized in that The step 3) further comprises: 3.1) Constructing a simulation environment, which is based on the trained ROP prediction model and whose input features include weight on bit (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), drilling time (BDT), mud inflow (MFI) and large trench load (HKL); 3.2) Define the key elements of the reinforcement learning model, including state space, action space and reward function. The state space includes: weight on bit (WOB), rotation speed (RPM), standpipe pressure (SPP), torque (TQ), drilling time (BDT), mud inflow (MFI), large trench load (HKL) and historical drilling rate (ROP) data; the action space is defined as the increase, decrease or remain unchanged of the weight on bit and rotation speed to constrain the constraints; the reward function is a multi-objective optimization function in the form of: Reward=α·ROP gain -β·Adjustment cost -γ·Constraint penalty Among them, ROP gain It indicates the increase in drilling speed at the current time step t compared to the previous time step t-1, multiplied by the weight coefficient α, to encourage the increase in drilling speed; cost Represents the adjustment range of drilling pressure and speed at the current time step, multiplied by the weight coefficient β, to punish excessive parameter adjustments and reduce operating costs and equipment wear; Constraint penalty It is used to measure whether the engineering constraints (such as physical limitations of drilling pressure and rotation speed, drill bit wear rate, wellbore quality, etc.) are violated, multiplied by a higher weight coefficient γ to ensure strict compliance with safety and performance standards; 3.3) Create a reinforcement learning network structure model, which uses the deep reinforcement learning algorithm DDPG (Deep Deterministic Policy Gradient), including a policy network (Actor) and a value network (Critic), where the network structure uses a fully connected layer and the activation function is ReLU. The policy network in the reinforcement learning model outputs continuous drilling pressure and speed adjustment values ​​according to the current state, and the value network evaluates the cumulative reward of each state-action pair to assist the policy network in optimizing the control strategy. The characteristics of its network structure further include: 3.3.1) This model adds an extra fully connected layer to the traditional policy network, making the network depth reach four layers. The introduction of this extra layer enables the network to better understand and process multi-dimensional data in the drilling process, improving the accuracy and stability of the strategy; 3.3.2) In the value network design, the present invention introduces residual connections. Specifically, by adding two linear layers (residual1 and residual2), the input features are directly added to the output of the hidden layer. This residual connection effectively alleviates the gradient vanishing problem in deep networks, promotes a more stable and efficient training process, and improves the accuracy of value assessment; 3.3.3) The model uses a normal distribution (mean 0, standard deviation 0.1) for weight initialization in each layer of the policy network and value network. This initialization method helps to avoid gradient disappearance or explosion; 3.3.4) In order to prevent the model from overfitting and improve the generalization ability, the present invention adds a Dropout layer (p=0.1) after the first hidden layer of the strategy network and the key hidden layer of the value network. The Dropout layer reduces the network's dependence on specific neurons by randomly discarding the output of some neurons, thereby enhancing the adaptability and robustness of the model in different drilling environments; 3.3.5) The output of the policy network is limited to the range of [-1, 1] through the tanh activation function, and then the action value is scaled to the actual action space range of drilling pressure [WOB_low, WOB_high] and speed [RPM_low, RPM_high] through linear transformation. The specific formula is: Action represents the value of the current action. This design ensures that the output drilling pressure and rotation speed are within a reasonable and operable range, avoiding the generation of invalid or extreme action values, thereby improving the feasibility and safety of the control strategy.

5. The method according to claim 4, characterized in that The constraints in the simulation environment include the real physical operation limits of drilling pressure and rotation speed during the drilling process, ensuring the feasibility of the generated control scheme in actual operation.

6. The method according to claim 1, characterized in that The step 4) further comprises: 4.1) The intelligent agent continuously interacts with the simulation environment and at each time step t, it t Generate action a through the policy network t , the simulation environment is based on action a t and status t Update the state to the next state s t+1 and generate instant rewards; 4.2) Establish and update the Experience Replay Buffer to store interaction data (s t ,a t ,r t ,s t+1 ), each time the model parameters are updated, a small batch of data is randomly extracted from the experience pool for training; 4.3) By minimizing the loss function of the value network and updating the policy network parameters using the policy gradient method, ensure that the policy network can output actions that maximize the expected cumulative reward in different states, where the loss function of the value network is defined as: in, y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ) s is the state at a certain moment, α is the action in the current state. yi represents the target cumulative reward, γ is the discount factor, Q′ and μ′ are the parameters of the target value network and the target strategy network respectively. By minimizing the loss function L Q , use gradient descent method to update the value network parameters θ; 4.4) Monitor the cumulative rewards and strategy stability during training, determine whether the model has converged, and save the optimal strategy network and value network parameters after convergence; 4.5) Apply the trained reinforcement learning model to actual drilling operations, generate the optimal weight on bit (WOB) and rotation speed (RPM) adjustment strategy according to the real-time drilling status, and achieve maximum control of the rate of penetration (ROP).

Citation Information

Cited By

  • Dynamic optimization control method for tin smelting process

    CN120215281A

  • Analog circuit parameter determination method and device, medium and product

    CN120337840A

  • Underground tool face dynamic control method and system based on reinforcement learning

    CN120487037A

  • Downhole tool face dynamic control method and system based on reinforcement learning

    CN120487037B

  • Large cylinder forging simulation friction coefficient calibration method based on reinforcement learning

    CN120542277A