A drilling trajectory deviation correction control method and system

CN122728554APending Publication Date: 2026-09-11SICHUAN YINMA OILFIELD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611069033.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0005]本发明针对现有技术中钻井轨迹控制依赖专家经验进行分段纠偏、难以适应地层突变、纠偏路径缺乏全局优化容易产生急弯段,以及井眼曲率约束与工具造斜率极限难以实时满足的技术问题,提供一种钻孔轨迹纠偏控制方法及系统

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122728554A_ABST
    Figure CN122728554A_ABST
Patent Text Reader

Abstract

The application discloses a drilling trajectory deviation correction control method and system, and relates to the technical field of well drilling engineering.The method comprises the following steps: obtaining drilling digital twin environment data; constructing a trajectory planning mathematical model under a Markov decision process framework to obtain a state space, an action space and a reward function; deploying a deep reinforcement learning agent model in the drilling digital twin environment, performing offline pre-training and online fine-tuning, and obtaining an intelligent planning agent model; calling the intelligent planning agent model to output a control instruction in an actual drilling process, generating an optimal action sequence by using a rolling optimization strategy, and executing the first action; and detecting the deviation between an actual trajectory and an expected trajectory, and automatically generating a transition trajectory under the conditions of meeting a wellbore curvature constraint and a tool build-up rate limit when the deviation exceeds a preset deviation threshold.The application realizes the global optimization and real-time deviation correction of a drilling trajectory, and reduces the downhole operation risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drilling engineering technology, specifically to a borehole trajectory correction control method and system. Background Technology

[0002] Currently, oilfield drilling trajectory control generally employs offline pre-designed trajectories using geometric analytical methods, followed by segmented corrections through manual adjustment of the tool face. This approach relies on expert experience, making it difficult to adapt to sudden formation changes, and the correction path lacks global optimization, easily resulting in sharp bends that increase friction and downhole vibration. With the improvement of drilling automation, both domestic and international efforts have begun to explore the introduction of digital twin and deep reinforcement learning technologies into the field of trajectory control.

[0003] For example, patent CN121031307A discloses a geological support intelligent planning and mining system based on deep reinforcement learning, but it is mainly aimed at mine mining planning, and neither the state space nor the action space involves the wellbore curvature constraints and tool build-up rate limits specific to drilling engineering. Another patent CN113687659B discloses an optimal trajectory generation method based on digital twins, which uses a reinforcement learning surrogate model to generate an initial trajectory and then optimizes iteratively by a planner. However, this method is designed for industrial robot motion planning and does not consider the real-time updates of downhole measurement-while-drilling data and the need for reservoir encounter rate optimization.

[0004] Therefore, there is an urgent need for an intelligent method that integrates deep reinforcement learning with digital twin technology and is specifically adapted to dynamic planning and correction of oilfield drilling trajectories. Summary of the Invention

[0005] This invention addresses the technical problems in existing technologies, such as the reliance on expert experience for segmented correction of drilling trajectory, difficulty in adapting to sudden formation changes, lack of global optimization of correction paths leading to sharp bends, and difficulty in meeting wellbore curvature constraints and tool build-up rate limits in real time. It provides a drilling trajectory correction control method and system.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0007] In a first aspect, the present invention provides a borehole trajectory correction control method, comprising:

[0008] Acquire drilling digital twin environment data;

[0009] Based on the drilling digital twin environment data, a trajectory planning mathematical model under the Markov decision process framework is constructed to obtain the state space, action space and reward function;

[0010] A deep reinforcement learning agent model is deployed in the drilling digital twin environment. The trajectory planning mathematical model is used to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model to obtain an intelligent planning agent model.

[0011] During actual drilling, the intelligent planning agent model is invoked to output the current control command, and the optimal action sequence is generated using a rolling optimization strategy, and the first action in the optimal action sequence is executed.

[0012] The deviation between the actual trajectory and the expected trajectory is detected. When the deviation exceeds a preset deviation threshold, the intelligent planning agent model automatically generates a transition trajectory under the conditions of satisfying the wellbore curvature constraint and the tool build-up rate limit.

[0013] Secondly, the present invention provides a borehole trajectory correction control system, comprising:

[0014] The environment acquisition module is used to acquire drilling digital twin environment data;

[0015] The framework construction module is used to construct a trajectory planning mathematical model under the Markov decision process framework based on the drilling digital twin environment data, and obtain the state space, action space and reward function;

[0016] The agent training module is used to deploy a deep reinforcement learning agent model in the drilling digital twin environment, and to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model using the trajectory planning mathematical model to obtain an intelligent planning agent model.

[0017] The instruction execution module is used to call the intelligent planning agent model to output the current control instruction during the actual drilling process, and to generate the optimal action sequence using a rolling optimization strategy, and execute the first action in the optimal action sequence.

[0018] The deviation correction module is used to detect the deviation between the actual trajectory and the expected trajectory. When the deviation exceeds a preset deviation threshold, the intelligent planning agent model automatically generates a transition trajectory under the conditions of satisfying the wellbore curvature constraint and the tool build-up rate limit.

[0019] The beneficial effects of this invention are:

[0020] Compared to existing technologies, this invention first integrates digital twins and deep reinforcement learning, constructing a Markov decision process framework using drilling digital twin environment data. It embeds wellbore curvature constraints and tool build-up rate limits into the action space and reward function, enabling the surrogate model to grasp safety boundaries during the training phase. Secondly, it employs a two-stage mechanism of offline pre-training and online fine-tuning. In the offline stage, it uses historical data and random formation models for batch training to obtain basic navigation capabilities. In the online stage, it uses real-time drilling measurement data for incremental gradient updates, rapidly adapting to changes in actual downhole conditions. Thirdly, it introduces a rolling optimization control strategy. Within each drilling step length, the surrogate model plans the optimal action sequence in a preset prediction time domain. After executing the first action, it replans for new observation states, achieving closed-loop rolling correction of the trajectory. Finally, when the trajectory deviation exceeds the limit, the surrogate model automatically generates a smooth transition trajectory under dual mechanical constraints, completing path replanning with the goal of minimizing additional footage and friction. This invention achieves global optimization and real-time correction of the drilling trajectory. Attached Figure Description

[0021] Figure 1 A flowchart illustrating a borehole trajectory correction control method provided by the present invention;

[0022] Figure 2 This is a schematic diagram of the overall logic of a borehole trajectory correction control method provided by the present invention;

[0023] Figure 3 This is a schematic diagram of a borehole trajectory correction control system provided by the present invention.

[0024] In the attached diagram, the components represented by each number are as follows:

[0025] Module 11 for environment acquisition, Module 12 for framework construction, Module 13 for agent training, Module 14 for instruction execution, and Module 15 for deviation correction. Detailed Implementation

[0026] Example 1, as Figure 1 , Figure 2 As shown, this embodiment of the invention provides a borehole trajectory correction control method, including:

[0027] S10: Acquire drilling digital twin environment data;

[0028] In oil drilling, the drill bit operates thousands of meters underground. Operators cannot directly observe the downhole conditions and can only rely on limited parameters such as well inclination angle, azimuth angle, and drill bit position collected by measurement-while-drilling (MWD) tools to determine the trajectory. Due to the significant heterogeneity and uncertainty of underground rock formations, the drillability of rocks varies greatly at different depths and in different areas. The drill bit may exhibit unpredictable deviations when traversing different formations. Relying solely on surface experience for manual correction can easily lead to risks such as trajectory deviation from the target formation, sudden changes in wellbore curvature causing drill string jamming, or tool damage.

[0029] Therefore, it is necessary to construct a digital mirror system that can reflect the real downhole physical environment, geological features, and the interaction between the drill string and the formation—that is, a drilling digital twin environment. This drilling digital twin environment can dynamically update and simulate downhole conditions by integrating initial geological model data, real-time measurement-while-drilling data, logging-while-drilling data, and drilling engineering parameter data, providing a high-fidelity virtual verification platform for subsequent trajectory planning and deviation correction decisions.

[0030] Specifically, the purpose of this step is to acquire the basic data used to build and drive the digital twin environment, covering the entire process from geological modeling to real-time data acquisition, so as to ensure that the digital twin system can synchronously reflect the drilling process in time and accurately represent the distribution of formation properties in space.

[0031] Specifically, acquiring drilling digital twin environment data includes:

[0032] Obtain initial geological model data, which includes target stratigraphic position, thickness, dip angle, and physical property distribution;

[0033] Real-time acquisition of measurement-while-drilling data, logging-while-drilling data, and drilling engineering parameter data during the drilling process;

[0034] A data cleaning algorithm combining Kalman filtering and neural networks is used to perform real-time repair and interpolation on the measurement-while-drilling data, the logging-while-drilling data, and the drilling engineering parameter data.

[0035] Based on the Bayesian inversion method, the local three-dimensional geological model is dynamically updated using the repaired and interpolated data, and the updated local three-dimensional geological model is fused with the initial geological model data to obtain the drilling digital twin environment data.

[0036] First, initial geological model data is obtained. This initial geological model data is a three-dimensional geological model established before drilling operations begin, based on seismic exploration data, adjacent well logging data, and regional geological research findings. This initial geological model data includes the target stratigraphic level, i.e., the oil and gas-bearing strata planned to be encountered; thickness, i.e., the vertical extent of the target stratigraphic level; dip angle, i.e., the angle of inclination of the stratigraphic layer relative to the horizontal plane; and physical property distribution, i.e., the distribution of physical properties such as rock density, porosity, and permeability in three-dimensional space. The initial geological model data is used to provide a macroscopic geological reference framework for drilling trajectory planning.

[0037] Secondly, during the drilling process, measurement-while-drilling (MWD) data, logging-while-drilling (LOD) data, and drilling engineering parameter data are acquired in real time. Specifically, MWD data includes the three-dimensional coordinates of the drill bit's current position, inclination angle, azimuth angle, and tool face angle, used to describe the drill bit's spatial position and attitude in real time. LOD data includes physical property parameters such as natural gamma ray, resistivity, density, and neutron porosity, used to identify changes in the lithology and physical properties of the formation surrounding the drill bit. Drilling engineering parameter data includes drilling pressure, rotational speed, torque, pump pressure, displacement, and drilling speed, used to reflect the stress state of the drilling tools and drilling efficiency.

[0038] Optionally, the above data is collected in real time by sensors installed near the drill bit and on ground equipment at a frequency of one sampling point per second or per meter to form a time-series data stream.

[0039] Furthermore, a data cleaning algorithm combining Kalman filtering and neural networks is used to perform real-time repair and interpolation of measurement-while-drilling (MWD) data, logging-while-drilling (LOD) data, and drilling engineering parameter data. Downhole data acquisition is affected by high temperature, high pressure, and strong vibration environments, making it prone to sensor malfunctions, data transmission interruptions, or signal noise interference, resulting in outliers or missing values ​​in the data sequence. Existing single-stage processing methods have significant limitations: Kalman filtering alone can remove Gaussian noise but has limited ability to repair non-Gaussian noise caused by transient sensor failures and continuous multi-point missing values; neural networks alone can predict nonlinear sequences but lack sufficient ability to suppress measurement noise. Therefore, this step employs a two-stage data cleaning mechanism combining neural networks and Kalman filtering, leveraging the advantages of both.

[0040] In an optional embodiment, during operation, the real-time acquired drilling data stream is first input into a pre-trained Long Short-Term Memory (LSTM) network. This LSM network is a variant of a recurrent neural network suitable for time series forecasting, controlling the memory and forgetting of information through three gating units: a forget gate, an input gate, and an output gate. During the offline preparation phase, complete sensor records from historical drilling data are used as training samples, with continuous data windows before missing or anomaly events as input, and actual values ​​or completed samples as supervision labels, to train the LSM network to learn nonlinear time-series patterns in the data sequence.

[0041] For example, the Long Short-Term Memory (LSTM) network employs a three-layer stacked structure. The first layer contains 128 hidden units, the second layer contains 128 hidden units, and the third layer contains 128 hidden units. The input layer receives data with a time step length of 10, meaning it uses data from 10 consecutive sampling points prior to the current time step to predict the value of the current time step or subsequent missing segments. Each LSM layer contains three gating units: a forget gate, an input gate, and an output gate. This gating mechanism controls the retention and forgetting of information, preventing gradient vanishing or exploding. The forget gate determines the proportion of the unit state from the previous time step that needs to be retained until the current time step. The input gate determines the proportion of the input data from the current time step that needs to be written into the unit state. The output gate determines the proportion of the current unit state that needs to be output to the next layer or the output layer. The network output layer is a fully connected layer, and the number of neurons in the output layer is consistent with the dimension of the single data to be predicted; that is, it predicts the value of one drilling parameter at a time, so the number of neurons in the output layer is set to 1. The output layer activation function is a linear activation function. During training, mean squared error was used as the loss function, Adam was used as the optimizer, the initial learning rate was set to 0.001, the batch size was set to 32, and the number of training epochs was set to 200. Training was terminated when the loss function value on the validation set no longer decreased after 10 consecutive training epochs.

[0042] During the online operation phase, after the real-time data stream enters the Long Short-Term Memory (LSTM) network, the network predicts the data value at the current moment based on several consecutive sampling points prior to the current moment, such as the values ​​of the previous 10 sampling points, and simultaneously predicts the values ​​of each sampling point in the missing data segment. For discontinuous outliers caused by momentary sensor malfunctions, the LSM network performs nonlinear interpolation using windows of normal data before and after the sensor. For multiple consecutive missing data points, the LSM network uses learned temporal patterns to recursively predict and fill in the missing data point by point. Through neural network compensation, a preliminarily repaired data sequence is output, in which non-Gaussian noise is effectively corrected and missing values ​​are appropriately filled.

[0043] Subsequently, the preliminary repaired data sequence output by the Long Short-Term Memory (LSTM) network is input into the Kalman filter. The Kalman filter is an optimal estimation algorithm based on a state-space model, and its operation consists of two recursive steps: prediction and update. The prediction step uses the state estimate from the previous time step to deduce the prior state estimate and the prior estimation error covariance at the current time step according to the system state transition equation. The update step uses the observations at the current time step to calculate the Kalman gain and uses the observations to weight and correct the prior estimate to obtain the posterior state estimate.

[0044] Specifically, taking drilling pressure data in drilling engineering parameters as an example, the system state variable is defined as the actual value of drilling pressure, and the observed variable is the sensor measurement value. The Kalman filter, through recursive calculation, smooths the data sequence after neural network compensation, further suppressing residual Gaussian measurement noise, eliminating minor fluctuations that may be introduced into the neural network prediction, and outputting the optimal estimate of the state variable. Through the smoothing process of the Kalman filter, a continuous, stable, and low-noise high-fidelity data sequence can be obtained.

[0045] The sequence in which the neural network and Kalman filter work together is as follows: first, the neural network performs nonlinear error compensation and missing value prediction; then, the Kalman filter performs optimal estimation and smoothing / denoising. Their roles are clearly defined: the neural network handles the prediction and filling of nonlinear anomalies and long missing segments, while the Kalman filter handles Gaussian noise filtering and overall smoothing. The final output is a high-fidelity real-time data sequence, providing accurate time alignment and numerically accurate input data for subsequent geological modeling.

[0046] Finally, based on the Bayesian inversion method, the local 3D geological model is dynamically updated using the repaired and interpolated data. The updated local 3D geological model is then fused with the initial geological model data to obtain the drilling digital twin environment data. The Bayesian inversion method is a probabilistic statistical back-inference algorithm that combines the prior distribution of geological parameters with the likelihood function of logging-while-drilling data. The posterior probability distribution is calculated using the Bayesian formula to infer the actual properties of the formation around the drill bit. Since the initial geological model is based on seismic and adjacent well data and has a certain degree of uncertainty, while the newly acquired logging-while-drilling data reflects the true formation information at the current drill bit location, this step uses the Bayesian inversion method to fuse prior geological information with measured logging-while-drilling data. The posterior formation parameters are output in the form of a probability distribution, thus preserving the macroscopic trend of the initial geological model while gradually correcting the geological properties of local areas with real-time data, continuously improving the accuracy of the digital twin environment data as the drilling progresses.

[0047] Specifically, the Bayesian inversion method uses Bayes' theorem to multiply the prior probability provided by the initial geological model by the likelihood function of the logging-while-drilling data to calculate the posterior probability distribution. The maximum value of this distribution is then used as the estimate of the formation parameters. The formula is: the posterior probability density is proportional to the prior probability density multiplied by the likelihood function. Here, the prior probability density is provided by the initial geological model and represents the initial understanding of the formation parameters before obtaining new measurement data; the likelihood function represents the probability of observing actual logging-while-drilling data under the currently assumed formation parameters.

[0048] Optimal estimates of formation parameters can be obtained by maximizing the posterior probability density or sampling the posterior distribution. For example, taking the spatial distribution of natural gamma values ​​as an example, the gamma values ​​predicted in the initial geological model are used as the prior distribution, and the gamma values ​​measured during drilling are used as the observation data. The posterior distribution of gamma values ​​within a certain range near the drill bit is calculated through Bayesian inversion, and the expected value of this distribution is used as the updated local model parameters. Using the repaired and interpolated logging-while-drilling data, a Bayesian inversion is performed after each drilling distance, continuously correcting the formation properties within a few meters in front of the drill bit in the initial geological model, thus achieving dynamic updates of the local three-dimensional geological model.

[0049] Finally, the updated local 3D geological model is fused with the initial geological model. Optionally, the fusion method is as follows: for drilled areas, the model updated with logging-while-drilling data is used; for undrilled areas, the predicted values ​​of the initial geological model are retained. The final fused model is the drilling digital twin environment data.

[0050] Specifically, the resulting drilling digital twin environment data has real-time temporal properties and high spatial resolution, which can be used to provide accurate environmental input for subsequent trajectory planning mathematical models.

[0051] S20: Based on the drilling digital twin environment data, construct a trajectory planning mathematical model under the Markov decision process framework to obtain the state space, action space and reward function;

[0052] Furthermore, after obtaining the drilling digital twin environment data, the drilling trajectory control problem needs to be transformed into a sequential decision-making problem suitable for deep reinforcement learning. Drilling trajectory control exhibits typical sequential decision-making characteristics: within each control step, the surrogate model selects an action based on the current downhole state. After this action is applied to the drilling system, the state transitions, the surrogate model receives an immediate reward, and then proceeds to the next decision step. This process perfectly aligns with the fundamental assumptions of Markov decision processes, namely that the state at the next moment depends only on the current state and the current action, and is independent of historical states. Therefore, this step constructs a trajectory planning mathematical model within the Markov decision process framework, defining three core elements: the state space, the action space, and the reward function.

[0053] Specifically, a trajectory planning mathematical model under the Markov decision process framework is constructed, including obtaining the state space, action space, reward function, and state transition probability model; the state transition probability model combines the bottom drill string assembly mechanics model and the formation drillability model, and uses a long short-term memory network to predict the probability distribution of the next state.

[0054] The state space includes the three-dimensional coordinates of the current drill bit position, the current inclination angle, the current azimuth angle, the current wellbore curvature, the vertical distance from the top boundary of the target layer, the feature vector of the formation attributes ahead, and the current remaining directional drilling capacity.

[0055] The action space includes the target well inclination angle, target azimuth angle, drill bit guiding force direction and magnitude, and drilling speed parameter combination for the next drilling step;

[0056] The reward function adopts a multi-objective weighted design, including trajectory tracking reward, reservoir encounter reward, engineering safety reward, drilling efficiency reward and energy consumption penalty, wherein the weight of each sub-reward adopts an adaptive adjustment mechanism.

[0057] The state space describes all available information upon which the agent model bases its decisions. In drilling trajectory control scenarios, the state space needs to contain all parameters characterizing the current drill bit position, attitude, wellbore geometry, and formation conditions ahead. These parameters are extracted from drilling digital twin environment data and normalized to form a fixed-dimensional state vector.

[0058] The action space defines the set of operations that the surrogate model can execute within each decision step. In drilling trajectory control, the action space includes adjustable variables such as the target wellbore inclination angle, target azimuth angle, direction and magnitude of the drill bit steering force, and combinations of drilling rate parameters for the next drilling step. The design of the action space must consider the physical constraints of the downhole actuators, such as an upper limit on the tool build-up rate and a wellbore curvature that cannot exceed a safety threshold. These constraints are ensured through pruning or scaling mechanisms during the action output phase.

[0059] The reward function is used to evaluate the quality of actions taken by the surrogate model and is a quantitative expression of the optimization objectives of reinforcement learning. In drilling trajectory control, the reward function needs to consider multiple optimization objectives simultaneously: trajectory tracking accuracy, i.e., the degree of deviation between the actual trajectory and the designed trajectory; reservoir encounter rate, i.e., the proportion of the drill bit that penetrates within the target formation; engineering safety, i.e., whether the wellbore curvature exceeds the safe range and whether the drill string friction is too high; drilling efficiency, i.e., drilling speed and footage cost; and energy consumption penalty, i.e., the additional energy consumption caused by excessive adjustment of the directional tool. These sub-objectives are competitive; for example, increasing the drilling speed may sacrifice trajectory accuracy, and excessive deviation may increase friction. Therefore, the reward function adopts a multi-objective weighted design, and the weight of each sub-reward can be dynamically adjusted according to actual operational needs.

[0060] Specifically, in a preferred embodiment, the state space includes the three-dimensional coordinates of the current drill bit position, the current inclination angle, the current azimuth angle, the current wellbore curvature, the vertical distance from the top boundary of the target layer, the formation attribute feature vector ahead, and the current remaining directional drilling capacity. The three-dimensional coordinates of the current drill bit position refer to the drill bit's east, north, and vertical depth in space, used to determine the drill bit's spatial location; the current inclination angle is the angle between the drill bit axis and the vertical line, ranging from 0 to 180 degrees, with 0 degrees representing vertically downwards; the current azimuth angle is the angle between the projection of the drill bit axis onto the horizontal plane and the due north direction, ranging from 0 to 360 degrees; the current wellbore curvature is the change in inclination angle or azimuth angle per unit length of well section, used to assess whether the wellbore curvature exceeds the safe range; The vertical distance from the top boundary of the target layer refers to the vertical distance between the current position of the drill bit and the top interface of the target oil and gas layer. A positive value indicates that the drill bit is above the target layer, and a negative value indicates that it has entered the target layer. The formation attribute feature vector ahead refers to a multi-dimensional vector composed of physical property parameters such as rock type, porosity, and permeability within a few meters ahead of the drill bit, extracted from the drilling digital twin environment data. The current remaining directional drilling capacity refers to the remaining tool directional drilling capacity that can be used to adjust the wellbore trajectory under the current drill string configuration, calculated from the tool design parameters and the amount of directional drilling already used. The above state variables are extracted in real time from the drilling digital twin environment data, normalized, and then concatenated into a fixed-dimensional state vector, which is then input into the surrogate model.

[0061] The action space includes the target wellbore inclination angle, target azimuth angle, direction and magnitude of the drill bit steering force, and the combination of drilling speed parameters for the next drilling step. The target wellbore inclination angle refers to the desired wellbore inclination angle at the end of the next drilling step; the target azimuth angle refers to the desired azimuth angle at the end of the next drilling step; the direction of the drill bit steering force refers to the direction angle of the lateral force applied to the drill bit on the bottom plane, and the magnitude of the drill bit steering force refers to the amplitude of this lateral force. Both together determine the direction and intensity of the drill bit's build-up; the combination of drilling speed parameters includes the set values ​​of the drill bit rotation speed and the mechanical drilling speed. These action variables are coupled; for example, the change in the target wellbore inclination angle is limited by the tool's maximum build-up rate, and the change in the target azimuth angle is limited by the tool's azimuth drift characteristics. The action values ​​output by the surrogate model need to be post-processed and trimmed to ensure they fall within the range allowed by the actuator.

[0062] Furthermore, the reward function employs a multi-objective weighted design, including trajectory tracking reward, reservoir encounter reward, engineering safety reward, drilling efficiency reward, and energy consumption penalty. The trajectory tracking reward guides the surrogate model to drill along the planned path. Positive incentive values ​​are calculated based on the positional and angular deviations between the actual and desired trajectories. Positional deviation is the straight-line distance between the actual drill bit position and the corresponding point on the desired trajectory; angular deviation is the absolute value of the difference between the actual well inclination angle, azimuth angle, and the desired value. Smaller positional and angular deviations result in higher trajectory tracking rewards. Specifically, a negative weighted root mean square form can be used to minimize trajectory deviation during the learning process of the surrogate model.

[0063] Reservoir encounter rewards ensure the drill bit always travels within high-quality reservoirs. Real-time natural gamma ray and resistivity values ​​are obtained from logging-while-drilling data and compared to the characteristic ranges of the target reservoir. When both the real-time gamma ray and resistivity values ​​fall within the target reservoir's gamma characteristic range, the drill bit is considered to be within a high-quality reservoir, and a positive reward is given. Conversely, when either the real-time gamma ray or resistivity value deviates from the characteristic range, the drill bit is considered to have exited the reservoir, and a negative penalty is imposed. This reward mechanism guides the surrogate model to prioritize maintaining the drill bit's drilling path within the reservoir, avoiding drilling into non-reservoir or surrounding rock.

[0064] The engineering safety reward is used to constrain safety indicators such as wellbore curvature and friction torque to prevent sharp turns or stuck pipe accidents. Wellbore curvature is calculated by dividing the change in inclination and azimuth angles between two adjacent measurement points by the well section length. When the wellbore curvature is below 80% of the preset safety threshold, the engineering safety reward is a full positive value; when the wellbore curvature is between 80% and the threshold, the reward decreases linearly with increasing curvature; when the wellbore curvature exceeds the preset safety threshold, a large negative reward is given. Friction torque is estimated by the difference between the surface torque sensor reading and the downhole torque measurement; a negative reward is also applied when it exceeds the safety threshold. This reward mechanism guides the proxy model to actively avoid operations exceeding safety boundaries while meeting trajectory requirements.

[0065] Drilling efficiency rewards are used to improve operational efficiency while meeting trajectory accuracy and engineering safety requirements. The rate of drilling (RBP) is the increment of drill bit footage per unit time, calculated by differentiating depth values ​​from depth sensors. When wellbore curvature and trajectory deviation are within reasonable ranges, a higher RBP results in a greater drilling efficiency reward. Conversely, if wellbore curvature or trajectory deviation exceeds allowable limits, no efficiency reward is given even with a high RBP, preventing the proxy model from sacrificing trajectory accuracy or safety constraints to increase drilling speed.

[0066] Energy consumption penalties are used to encourage smooth control and reduce unnecessary energy consumption. The variation in control commands between adjacent decision steps is monitored, including changes in the target well inclination angle, target azimuth angle, drill bit steering force, and drilling rate parameters. If the variation in a control command exceeds a preset threshold, a small negative reward proportional to the variation is applied; if multiple control commands change significantly simultaneously, the penalty value is accumulated. This penalty mechanism guides the surrogate model to output a continuous and smooth control sequence, avoiding frequent and large adjustments to the steering tool, thus reducing tool wear and energy consumption.

[0067] The five sub-rewards mentioned above are weighted and summed according to preset weights to obtain the comprehensive reward function. The weights of each sub-reward employ an adaptive adjustment mechanism: when the actual trajectory deviates significantly from the designed trajectory, the weight of the trajectory tracking reward is increased; when the drill bit frequently encounters the target layer, the weight of the reservoir encounter reward is increased; and when the wellbore curvature continuously approaches the safety limit, the weight of the engineering safety reward is increased. This adaptive adjustment mechanism dynamically calculates the weight coefficients by monitoring the real-time deviation of each sub-objective, enabling the surrogate model to automatically focus on the most critical optimization objective at different drilling stages. The weighted summation of all sub-rewards yields the comprehensive reward value, which serves as a monitoring signal for the surrogate model's optimization.

[0068] Furthermore, state transition probability models are used to describe the probability distribution of the system's state at the next moment after taking an action in the current state. In drilling trajectory control problems, state transitions are influenced by a variety of complex factors, such as the interaction mechanism between the drill bit and the formation, the mechanical response of the bottom drill string assembly, and changes in formation drillability. These factors result in high nonlinearity and uncertainty, making them difficult to describe with simple analytical expressions.

[0069] Therefore, this step employs a Long Short-Term Memory (LSTM) network to model the state transition probability. Specifically, prior information provided by the bottom drill string assembly mechanics model and the formation drillability model is used as supplementary input features, which are input into the LSM network along with the current state and actions. The bottom drill string assembly mechanics model describes the lateral forces and build-up responses generated by the drill bit under different drilling pressures, rotational speeds, and guiding forces. Structured constraints for state transitions can be obtained based on the finite element analysis results or empirical formulas of the drill string assembly. The formation drillability model describes the mechanical drilling rate and deflection trends of the drill bit under different rock types, formation dip angles, and hardness conditions. Statistical relationships can be established based on data from adjacent wells or regional geological patterns.

[0070] The input layer of the Long Short-Term Memory (LSTM) network receives vectors including the current state's three-dimensional coordinates of the drill bit position, inclination angle, azimuth angle, wellbore curvature, vertical distance from the top boundary of the target formation, the preceding formation attribute feature vector, and the current remaining directional drilling capacity; the current action's target inclination angle, target azimuth angle, steering force direction and magnitude, drilling rate parameter combination; and prior features extracted from the bottom-of-the-line assembly mechanics model and formation drillability model. The input sequence length is a fixed historical window step size, for example, 5 steps, to capture the temporal dependencies of state transitions. The network output layer outputs the probability distribution of the next state through a Softmax function or a Gaussian distribution parameterized form, i.e., the mean and variance of each possible state variable.

[0071] By using state transition samples from a large amount of historical drilling data as training data and the actual next state as the supervision label, a Long Short-Term Memory (LSTM) network is trained under supervision. This allows the network to learn the probability distribution of the next state given the current state and an action. After training convergence, this state transition probability model can be used as an environmental dynamic model within a Markov decision process framework for trajectory prediction and policy evaluation during surrogate model training. When the surrogate model needs to evaluate the long-term reward of a certain action sequence, it can call this state transition probability model to sample possible state transition paths multiple times and calculate the expected cumulative reward from multiple possible outcomes, enabling the surrogate model to make more robust decisions in the face of downhole uncertainty.

[0072] In summary, by constructing a trajectory planning mathematical model within the Markov decision process framework, the continuous dynamic process of drilling trajectory control can be abstracted into a discrete-time-step sequential decision problem, providing a clear learning objective and interactive environment for training the deep reinforcement learning surrogate model. The quality of this trajectory planning mathematical model directly determines the learning efficiency and final control effect of the subsequent surrogate model.

[0073] S30: Deploy a deep reinforcement learning agent model in the drilling digital twin environment, and use the trajectory planning mathematical model to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model to obtain an intelligent planning agent model;

[0074] Furthermore, after constructing the trajectory planning mathematical model within the Markov decision process framework, it is necessary to train an intelligent planning agent model capable of outputting the optimal action based on the current state. Since drilling operations are a high-risk, high-cost, continuous process, directly conducting trial-and-error training in a real drilling environment would lead to significant economic losses and safety risks. Therefore, this step uses a digital twin environment as a virtual training ground to pre-train the agent model in a safe and controllable simulation environment. Incremental online fine-tuning is then performed using real-time data collected during actual drilling operations, allowing the agent model to gradually adapt to real downhole conditions.

[0075] Specifically, a deep reinforcement learning agent model is deployed within the drilling digital twin environment, and the deep reinforcement learning agent model is pre-trained offline and fine-tuned online using the trajectory planning mathematical model, including:

[0076] A deep deterministic policy gradient algorithm is used as the basic framework to deploy an action network and an evaluation network.

[0077] In the offline pre-training phase, historical drilling data and simulated data generated by random formation models are used to train the deep reinforcement learning agent model in batches, using a priority experience replay mechanism.

[0078] During the online fine-tuning phase, in the actual drilling process, after each drilling progress is completed, the status is updated using real-time collected data, and the new status and reward value after actual execution are stored in the experience pool as online experience. An incremental gradient update is performed every preset time interval.

[0079] This deep reinforcement learning surrogate model employs a deep deterministic policy gradient algorithm (DPR) as its basic framework. The DPR is a deep reinforcement learning algorithm specifically designed for continuous action spaces, suitable for control problems where action variables such as well inclination angle, azimuth angle, and guide force magnitude take values ​​in continuous intervals. Specifically, the DPR is used as the basic framework, deploying an action network and an evaluation network. The action network directly outputs the optimal control command based on the current state, which includes continuous variables such as the target well inclination angle, target azimuth angle, drill bit guide force direction and magnitude, and drilling speed parameter combinations. The evaluation network receives the current state and the action output from the action network, evaluates the value of the action in the current state, and outputs a scalar evaluation value. The two networks are trained collaboratively: the action network aims to generate actions that can obtain high evaluation values ​​from the evaluation network; the evaluation network aims to accurately estimate the true value of the actions output by the action network. Through the alternating optimization of the two networks, the DPR can learn the optimal drilling trajectory control strategy in a continuous action space.

[0080] In the offline pre-training phase, historical drilling data and simulated data generated by a random formation model are used to train the deep reinforcement learning agent model in batches, employing a priority experience replay mechanism. Specifically, historical drilling data comes from complete drilling processes recorded in previous drilling projects, including state sequences at different depths, sequences of actions performed, and corresponding reward values. The random formation model is generated by applying random perturbations to the parameter space of the initial geological model. For example, the target layer thickness is randomly varied within ±20%, the formation dip angle within ±5 degrees, and the rock drillability coefficient within ±15%. Through this randomization process, a large amount of simulated data with different geological characteristics can be generated, expanding the coverage of the training samples.

[0081] During training, a priority experience replay mechanism is employed. In the training sample pool, sampling weights are assigned based on the temporal difference error of each sample. The temporal difference error is the difference between the network's predicted value and the actual reward plus the next-state prediction. A larger error indicates a less accurate current network assessment of the sample, and consequently, a higher learning value for that sample. The priority experience replay mechanism assigns higher sampling weights to trajectory experiences with high temporal difference errors, allowing them to be used more frequently in model training, thereby accelerating the surrogate model's learning from failures.

[0082] The offline pre-training continues until the average reward of the surrogate model in the verification scenario reaches a preset threshold, such as an average reward of no less than -10 over 100 consecutive rounds. At this point, a pre-trained surrogate model with basic trajectory planning capabilities is obtained.

[0083] Further, the process moves into the online fine-tuning phase. Specifically, during actual drilling, after each drilling step, the state is updated using real-time collected data. The new state and reward value are stored as online experience in an experience pool, and incremental gradient updates are performed every preset time interval. The drilling step is typically set to a measurement interval, such as every 5 meters or every 10 meters. After completing a drilling step, sensors collect data, calculate the new state, and calculate the reward value based on the action execution results output by the intelligent planning agent model. This experience sample is stored in the experience pool. Every preset time step, for example, after every five drilling steps, the most recently accumulated online experience samples are extracted from the experience pool for a small-batch gradient update. The learning rate in the online fine-tuning phase is set to one-tenth of the initial learning rate in the offline pre-training phase to ensure that the agent model adapts to new working conditions without forgetting the knowledge learned in the offline pre-training phase. Through online fine-tuning, the deep reinforcement learning agent model continuously adapts to the current actual formation conditions downhole, corrects the deviation between the digital twin environment and the actual environment, and continuously optimizes the control strategy.

[0084] Finally, through a two-stage training strategy of offline pre-training and online fine-tuning, an intelligent planning agent model was obtained. This intelligent planning agent model not only has the basic ability to plan trajectories under typical formation conditions, but also can quickly adapt to changes in formation and downhole conditions during the drilling process based on real-time feedback.

[0085] S40: During the actual drilling process, the intelligent planning agent model is invoked to output the current control command, and the optimal action sequence is generated using a rolling optimization strategy, and the first action in the optimal action sequence is executed;

[0086] Furthermore, in actual drilling operations, within each control step, the current state vector is input into the action network of the intelligent planning surrogate model. After forward computation, the action network outputs the optimal control command for the current state. This control command includes action variables such as the target wellbore inclination angle, target azimuth angle, drill bit steering force direction and magnitude, and drilling speed parameter combinations. If a single-step decision-making mode is adopted, i.e., only predicting the optimal action for the next step at a time, the surrogate model cannot foresee the potential risks of several future steps. For example, although the currently selected angle adjustment may produce a wellbore curvature within a safe range within the current step, the cumulative curvature over several consecutive steps may lead to an excess of wellbore curvature. Therefore, this step introduces a rolling optimization strategy to generate a set of optimal action sequences within each control step, rather than simply outputting a single action.

[0087] Specifically, a rolling optimization strategy is used to generate the optimal action sequence, including:

[0088] Within each control step, the intelligent planning agent model generates a series of optimal action sequences in a preset prediction time domain.

[0089] After executing the first action in the optimal action sequence, the intelligent planning agent model is called again for rolling optimization in the next step.

[0090] The specific execution method of this rolling optimization strategy is as follows: Based on the current state, the intelligent planning agent model generates a series of optimal action sequences within a preset prediction time domain. The prediction time domain refers to the number of time steps predicted forward, for example, it can be set to 5 drilling steps, each corresponding to a fixed drilling distance, such as 10 meters. Within each prediction step, the intelligent planning agent model outputs the corresponding action based on the current prediction state. Internally, the intelligent planning agent model estimates the cumulative reward of each action sequence within the prediction time domain and selects the action sequence with the highest cumulative reward as the optimal action sequence. This optimal action sequence contains the target control instructions to be executed at each step from the current step forward for several future drilling steps.

[0091] Subsequently, after generating the optimal action sequence, the system executes the first action in that sequence, which is the control command corresponding to the current drilling step length. After the drill bit executes drilling according to this command, the actual state changes, and the sensors collect new measurement-while-drilling data to update the state vector. When the next control step length arrives, the intelligent planning agent model is invoked again, and rolling optimization is performed again based on the updated state vector to generate a new optimal action sequence, and the first action in the new sequence is executed again. This process is repeated to achieve progressive rolling optimization control of the trajectory.

[0092] Specifically, the core advantage of this rolling optimization strategy lies in the fact that within each control step, the surrogate model considers not only the reward within the current step but also the cumulative reward over multiple future steps, thus avoiding subsequent trajectory loss of control due to short-sighted decision-making. Simultaneously, the strategy of executing only the first action and discarding the remaining actions in the sequence allows the system to retain its responsiveness to new external information. When actual downhole conditions deviate from the prediction model, the system will replan based on the latest measured conditions before the next control step begins, promptly correcting trajectory deviations caused by model errors or sudden formation changes.

[0093] Through the aforementioned rolling optimization strategy, the intelligent planning agent model can output the globally optimal drilling trajectory control sequence while satisfying the wellbore curvature constraint and tool build-up rate limit.

[0094] S50: Detect the deviation between the actual trajectory and the expected trajectory. When the deviation exceeds a preset deviation threshold, the intelligent planning agent model automatically generates a transition trajectory under the conditions of satisfying the wellbore curvature constraint and the tool build-up rate limit.

[0095] In actual drilling operations, even with a rolling optimization strategy, the actual drill bit trajectory may still deviate from the expected trajectory due to the unpredictability of downhole formation changes, the inherent delay of measurement-while-drilling data, and the error between the actual drill string response and model predictions. If not corrected in time, the accumulated deviation will cause the drill bit to miss the target reservoir or generate wellbore curvature exceeding safe limits. Therefore, this step performs real-time detection of the deviation between the actual and expected trajectories within each control cycle, and triggers an automatic transition trajectory generation mechanism when the deviation exceeds a preset threshold.

[0096] The specific method for deviation detection is as follows: Based on the three-dimensional coordinates of the drill bit position collected by the measurement-while-drilling system at intervals of one drilling progress, the spatial distance between this position and the corresponding depth point on the desired trajectory is calculated as the positional deviation. Simultaneously, the absolute value of the difference between the actual wellbore inclination angle and actual azimuth angle of the current drill bit and the desired wellbore inclination angle and desired azimuth angle of the corresponding point on the desired trajectory is calculated as the angular deviation. The positional deviation and angular deviation are then weighted and fused to obtain a comprehensive deviation value. Optionally, during the weighted fusion process, the weighting coefficients are set according to the different requirements for positional and directional accuracy at different drilling stages. When the drill bit approaches the top boundary of the target reservoir, the positional deviation directly affects whether it can accurately enter the reservoir; at this time, the positional deviation weight can be set to 0.7, and the angular deviation weight to 0.3. After the drill bit enters the reservoir, the directional deviation affects the stability of its movement within the reservoir; at this time, the positional deviation weight can be set to 0.4, and the angular deviation weight to 0.6.

[0097] The preset deviation threshold can be set according to reservoir thickness, wellbore size, and engineering safety requirements. For example, the positional deviation threshold can be set to 0.5 meters, and the angle deviation threshold can be set to 2 degrees. When the overall deviation value exceeds the preset deviation threshold, the intelligent planning agent model is invoked to automatically generate a transition trajectory.

[0098] Specifically, the transition trajectory is automatically generated, including:

[0099] Take the current actual location as the starting point and a point on the corrected target trajectory as the ending point;

[0100] Under the constraints that the wellbore curvature does not exceed a preset safety value and the tool build-up rate does not exceed a limit value, the intelligent planning agent model generates a transition trajectory that minimizes the sum of additional footage and friction.

[0101] Specifically, the transition trajectory generation process is as follows: the current actual drill bit position is taken as the starting point of the transition trajectory, and a point on the corrected target trajectory is taken as the ending point of the transition trajectory. The corrected target trajectory refers to the corrected path that starts from the current actual position, undergoes a smooth transition, and intersects with the original desired trajectory. The principle for selecting the ending point is to minimize the length of the transition trajectory and the angle of convergence with the original desired trajectory.

[0102] The intelligent planning agent model is subject to two hard constraints in the process of generating the transition trajectory: the first is the wellbore curvature constraint, and the second is the tool build-up rate limit constraint.

[0103] Wellbore curvature refers to the rate of change of the wellbore orientation angle per unit length of well section, usually measured in degrees per 30 meters or per 100 meters. Wellbore curvature is used to constrain the smoothness of the transition trajectory and prevent sharp bends. When the wellbore curvature is too large, the frictional resistance of the drill string in the curved section increases significantly, making it difficult to effectively transmit the drilling pressure applied from the surface to the drill bit. Simultaneously, the drill string is subjected to alternating bending stress in the curved section, accelerating tool fatigue failure. In severe cases, the drill string may become stuck in the curved section, causing a stuck pipe accident. Therefore, the wellbore curvature must be strictly controlled within a preset safety value, such as no more than 6 degrees per 100 meters or no more than 2 degrees per 30 meters. This safety value is determined comprehensively based on the bending strength of the drill string, the stiffness of the casing, and the fatigue life of the downhole tools.

[0104] Tool build-up rate refers to the inherent bending capability of a bottom-end assembly (BWA) to change the wellbore orientation under specific drilling parameters, typically measured in degrees per 30 meters. Different BWAs have different build-up rate characteristics. For example, a rotary steerable tool can change the inclination angle by 15 degrees per 30 meters under maximum steerable force, while the build-up rate of a conventional curved-shell screw drill string is generally 3 to 8 degrees per 30 meters. Tool build-up rate constraints mean that the change in inclination angle or azimuth angle between any two adjacent measurement points on the transition trajectory cannot exceed the maximum change that the currently used BWA can actually achieve within that drilling length. This constraint is determined by the mechanical design parameters and actual drilling capacity of the BWA and must be strictly satisfied during transition trajectory generation to ensure that the generated trajectory is physically achievable with existing tools.

[0105] Under the aforementioned two constraints, the intelligent planning agent model generates a transition trajectory that minimizes the sum of extra footage and frictional resistance. Extra footage refers to the portion of the actual drilling length that exceeds the required length for the ideal, unbiased trajectory during trajectory correction. Each additional meter of extra footage increases drilling time and cost; therefore, extra footage is a core indicator for evaluating the economic efficiency of the transition trajectory. Frictional resistance refers to the frictional resistance generated between the drill string and the wellbore wall as the drill string moves through the well. When the wellbore curvature is too large or the well section is too long, frictional resistance increases significantly, potentially leading to insufficient transmission of drilling pressure to the drill bit, resulting in pressure build-up. In severe cases, this can cause drill string jamming or drill string fatigue failure; therefore, frictional resistance is a core indicator for evaluating the operational safety of the transition trajectory. By minimizing the weighted sum of extra footage and frictional resistance during transition trajectory generation, both economic efficiency and operational safety can be considered while correcting trajectory deviations, ensuring that the generated transition trajectory neither excessively increases drilling costs nor excessively worsens the downhole stress state.

[0106] Optionally, the additional footage and friction resistance can be dimensionlessly processed. For example, by dividing each by its respective reference value, dimensionless additional footage coefficient and friction resistance coefficient can be obtained. Then, these are weighted and summed according to preset weights to obtain the comprehensive optimization target. The preset weights are set comprehensively based on the emphasis on economic benefits and downhole safety in drilling operations. For example, when the formation stability in the operating area is good and the drilling rig rental cost is high, the weight of additional footage is set to 0.7 and the weight of friction resistance is set to 0.3, prioritizing the reduction of footage costs caused by deviation correction. When the formation is fractured and prone to stuck pipe accidents, the weight of additional footage is set to 0.3 and the weight of friction resistance is set to 0.7, prioritizing the reduction of friction resistance to ensure downhole safety. Under normal operating conditions, both weights are set to 0.5 to achieve a balanced optimization of economic benefits and safety. Through the above weight settings, the optimization target of the transition trajectory can be flexibly adjusted according to actual operational needs.

[0107] Under the constraints of wellbore curvature and tool build-up rate, the intelligent planning agent model searches for the comprehensive optimization objective value of different transition trajectory candidate schemes through a value network, or directly outputs the transition trajectory parameter sequence that minimizes the comprehensive optimization objective value through a strategy network, thus generating the optimal transition trajectory. Through this optimization mechanism, the generated transition trajectory can minimize economic losses and downhole friction resistance while ensuring safety and tool feasibility.

[0108] Furthermore, this embodiment also includes a step of retaining the manual intervention interface, requesting manual takeover and recording intervention samples under preset risk conditions, including:

[0109] The interface for manual intervention is retained, allowing for requests for manual takeover and recording of intervention samples under preset risk conditions:

[0110] When the wellbore curvature prediction value of the control command output by the intelligent planning agent model exceeds the preset safety value in a preset number of consecutive cycles, or when the reward value given by the reward function decreases continuously within a preset window and the decrease exceeds the preset threshold, the system actively generates a prompt to request manual intervention and outputs the confidence level of the current recommended action.

[0111] All states and actions during manual intervention are recorded as high-quality samples and stored in an experience pool for subsequent online fine-tuning.

[0112] During drilling operations, while the intelligent planning agent model can handle most routine conditions and some abnormal situations, it may output high-risk control commands in certain extremely complex scenarios, such as encountering unforeseen large faults, sudden downhole tool failures, or severe sensor data conflicts. To ensure downhole safety, the system also retains a manual intervention interface, allowing experienced drilling engineers to take over control when necessary.

[0113] Specifically, the criteria for judging the preset risk conditions include two categories. The first category is when the wellbore curvature prediction value of the control commands output by the intelligent planning agent model exceeds a preset safety value within a preset number of consecutive output commands. For example, if the agent model predicts wellbore curvature exceeding a safety threshold in three consecutive output commands (the safety threshold can be set at 4 degrees per 30 meters), multiple consecutive exceedances indicate that the agent model is continuously generating high-risk actions under the current state, requiring manual intervention. The second category is when the reward value given by the reward function continuously decreases within a preset window, and the decrease exceeds a preset threshold. The reward value reflects the overall quality of the agent model's current decisions. If the reward value continuously decreases within five consecutive drilling steps, and the cumulative decrease exceeds 30% of the initial reward value, it indicates that the control effect of the agent model is continuously deteriorating, requiring manual intervention to analyze the causes and take over.

[0114] When any of the aforementioned preset risk conditions is triggered, the system proactively generates a prompt requesting manual intervention, issues an audible and visual alarm through the ground monitoring system's display interface, and outputs the confidence level of the currently recommended action. Specifically, the confidence level of the currently recommended action refers to the agent model's estimate of the reliability of its output control commands, which can be quantified by evaluating the value score of the network output or the probability distribution of the action network output. The confidence level ranges from 0 to 1; the closer the value is to 1, the more confident the agent model is in the action, while the closer the value is to 0, the greater the uncertainty of the model. The output confidence level can provide a reference for manual decision-making.

[0115] Simultaneously, the system records all states and actions during human intervention as high-quality samples. When an engineer takes over control, their decisions based on experience, along with the corresponding states and reward values, are fully recorded. These samples represent expert decision-making experience in complex and risky scenarios, possessing extremely high learning value. After recording, these high-quality samples are stored in an experience pool for subsequent online fine-tuning. In the subsequent online fine-tuning phase, the system extracts human intervention samples from the experience pool with higher sampling weights for incremental gradient updates, enabling the intelligent planning agent model to learn the decision-making patterns of human experts in extremely complex scenarios and continuously improve the model's control capabilities under edge conditions.

[0116] Through the aforementioned manual intervention interface and sample recording mechanism, the system maintains the efficiency of automated control while being compatible with expert decision-making, forming a human-machine collaborative drilling trajectory control system. As drilling operations continue, the number of high-quality manual intervention samples accumulated in the experience pool increases continuously, and the intelligent planning agent model evolves continuously through online learning, gradually improving its adaptability to complex and abnormal working conditions.

[0117] Furthermore, this embodiment also includes a reward function weight adaptive adjustment step based on historical statistical data, including:

[0118] Obtain the reservoir encounter failure rate due to trajectory deviation and the tool failure rate due to excessive deviation correction in historical drilling data;

[0119] The ratio of the reservoir drilling failure rate to the tool failure rate is calculated and set as the reward function weight adjustment factor.

[0120] The weight ratio of reservoir drilling reward to engineering safety reward in the reward function is automatically adjusted according to the reward function weight adjustment factor.

[0121] In drilling operations, there is an inherent competition between reservoir encounter rewards and engineering safety rewards. Improving the reservoir encounter rate requires the drill bit to stay within the target reservoir as much as possible, necessitating frequent trajectory adjustments to cope with formation undulations. However, frequent trajectory adjustments increase wellbore curvature and drill string friction, thereby increasing the risk of tool failure. Conversely, overly conservatively limiting wellbore curvature and tool load may reduce the reservoir encounter rate, leading to the drill bit breaking out of the reservoir and reducing oil production. Due to differences in geological conditions and equipment status in different well areas, the relative importance of these two indicators varies. Therefore, adaptive weight adjustments based on historical statistical data are necessary.

[0122] First, two indicators were obtained from the historical drilling database: the reservoir encounter failure rate due to trajectory deviation and the tool failure rate due to over-correction. Specifically, the reservoir encounter failure rate was calculated as follows: the ratio of the actual length of the drilled reservoir to the total length of the target reservoir in historical drillings was calculated, and the complementary value of this ratio was taken as the failure rate. For example, if the total length of the target reservoir in a well is 200 meters, and the actual length of the drilled reservoir is 160 meters, then the encounter failure rate is 40 meters divided by 200 meters, which equals 20%. The tool failure rate was calculated as follows: the number of tool failures caused by over-correction (i.e., the wellbore curvature frequently approaching or exceeding the safety threshold, or the directional tool operating at high intensity for extended periods) in historical drillings was counted, and this number was divided by the total number of wells drilled. For example, if 2 out of 10 wells experienced tool failures related to over-correction, then the tool failure rate is 20%.

[0123] Secondly, the ratio of reservoir drilling failure rate to tool failure rate is calculated, and this ratio is set as the weight adjustment factor of the reward function. Let the reservoir drilling failure rate be R and the tool failure rate be T, then the weight adjustment factor λ equals R divided by T. If the reservoir drilling failure rate is higher than the tool failure rate, then λ is greater than 1, indicating that the main problem in the current operating area is the reservoir drilling issue, and the weight of the reservoir drilling reward should be appropriately increased; if the tool failure rate is higher than the reservoir drilling failure rate, then λ is less than 1, indicating that the main problem in the current operating area is the tool safety issue, and the weight of the engineering safety reward should be appropriately increased.

[0124] Finally, based on the reward function weight adjustment factor, the weight ratio of reservoir drilling reward to engineering safety reward in the reward function is automatically adjusted. Specifically, let the original weight of the reservoir drilling reward be W. a The original weight of the engineering safety reward is W. b The adjusted reservoir encounter reward weight is W. a Multiply by λ, the adjusted engineering safety reward weight is W. b Divide by λ. After the above adjustment, the ratio of the two weights is changed from W... a W b Change to Wa Multiplied by the square of λ, W b After normalization, the sum of the two adjusted weights is still equal to the sum of the original two weights, which is 1, thus keeping the total weight unchanged.

[0125] Through the above adaptive adjustment mechanism, the weight configuration of the reward function can be continuously optimized based on historical data feedback. This guides the intelligent planning surrogate model to more accurately balance the contradiction between reservoir drilling rate and engineering safety in subsequent drilling operations, ensuring that the surrogate model's control strategy matches the actual conditions of the operating area. As historical data accumulates, the statistics of the weight adjustment factors become increasingly accurate, the weight configuration of the reward function becomes more reasonable, and the control performance of the surrogate model continues to improve.

[0126] In summary, the embodiments of this application have at least the following technical effects:

[0127] This invention first transforms the drilling trajectory control problem into a sequential decision-making problem by acquiring drilling digital twin environment data and constructing a trajectory planning mathematical model under the Markov decision process framework. This provides a structured state space, action space, and reward function for training the intelligent agent model. Second, a deep reinforcement learning agent model is deployed within the digital twin environment. Combining offline pre-training and online fine-tuning, the agent model can fully learn the optimal trajectory planning strategy in the virtual environment and continuously optimize using real-time drilling data, improving its adaptability to formation changes. Third, a rolling optimization strategy is adopted during actual drilling. Within each control step, the agent model generates the optimal action sequence within a preset prediction time domain and executes only the first action, achieving progressive optimization control of the trajectory.

[0128] Finally, when the deviation between the actual trajectory and the expected trajectory exceeds a preset threshold, the surrogate model automatically generates a transition trajectory under strict wellbore curvature constraints and tool build-up rate limits, effectively avoiding sharp bends and reducing friction and downhole vibration risks. This invention solves the technical problems of existing technologies that rely on expert experience for segmented correction, are difficult to adapt to formation changes, and lack global optimization of the correction path.

[0129] Example 2, as Figure 3 As shown, based on the same inventive concept as the borehole trajectory correction control method provided in Embodiment 1, this embodiment of the invention also provides a borehole trajectory correction control system, including:

[0130] Environment acquisition module 11 is used to acquire drilling digital twin environment data;

[0131] Framework construction module 12 is used to construct a trajectory planning mathematical model under the Markov decision process framework based on the drilling digital twin environment data, and obtain the state space, action space and reward function;

[0132] Agent training module 13 is used to deploy a deep reinforcement learning agent model in the drilling digital twin environment, and to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model using the trajectory planning mathematical model to obtain an intelligent planning agent model;

[0133] The instruction execution module 14 is used to call the intelligent planning agent model to output the current control instruction during the actual drilling process, and to generate the optimal action sequence using a rolling optimization strategy, and execute the first action in the optimal action sequence.

[0134] The deviation correction module 15 is used to detect the deviation between the actual trajectory and the expected trajectory. When the deviation exceeds the preset deviation threshold, the intelligent planning agent model automatically generates a transition trajectory under the condition of satisfying the wellbore curvature constraint and the tool build-up rate limit.

[0135] The environment acquisition module 11 is specifically used for:

[0136] Specifically, acquiring drilling digital twin environment data includes:

[0137] Obtain initial geological model data, which includes target stratigraphic position, thickness, dip angle, and physical property distribution;

[0138] Real-time acquisition of measurement-while-drilling data, logging-while-drilling data, and drilling engineering parameter data during the drilling process;

[0139] A data cleaning algorithm combining Kalman filtering and neural networks is used to perform real-time repair and interpolation on the measurement-while-drilling data, the logging-while-drilling data, and the drilling engineering parameter data.

[0140] Based on the Bayesian inversion method, the local three-dimensional geological model is dynamically updated using the repaired and interpolated data, and the updated local three-dimensional geological model is fused with the initial geological model data to obtain the drilling digital twin environment data.

[0141] Specifically, the framework construction module 12 is used for:

[0142] A mathematical model for trajectory planning under the Markov decision process framework is constructed, including obtaining the state space, action space, reward function, and state transition probability model. The state transition probability model combines the bottom drill string assembly mechanics model and the formation drillability model, and uses a long short-term memory network to predict the probability distribution of the next state.

[0143] Specifically, the state space includes the three-dimensional coordinates of the current drill bit position, the current inclination angle, the current azimuth angle, the current wellbore curvature, the vertical distance from the top boundary of the target layer, the feature vector of the formation attributes ahead, and the current remaining directional drilling capacity.

[0144] The action space includes the target well inclination angle, target azimuth angle, drill bit guiding force direction and magnitude, and drilling speed parameter combination for the next drilling step;

[0145] The reward function adopts a multi-objective weighted design, including trajectory tracking reward, reservoir encounter reward, engineering safety reward, drilling efficiency reward and energy consumption penalty, wherein the weight of each sub-reward adopts an adaptive adjustment mechanism.

[0146] The agent training module 13 is specifically used for:

[0147] Deploying a deep reinforcement learning agent model within the drilling digital twin environment, and using the trajectory planning mathematical model to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model, including:

[0148] A deep deterministic policy gradient algorithm is used as the basic framework to deploy an action network and an evaluation network.

[0149] In the offline pre-training phase, historical drilling data and simulated data generated by random formation models are used to train the deep reinforcement learning agent model in batches, using a priority experience replay mechanism.

[0150] During the online fine-tuning phase, in the actual drilling process, after each drilling progress is completed, the status is updated using real-time collected data, and the new status and reward value after actual execution are stored in the experience pool as online experience. An incremental gradient update is performed every preset time interval.

[0151] Specifically, the instruction execution module 14 is used for:

[0152] The optimal action sequence is generated using a rolling optimization strategy, including:

[0153] Within each control step, the intelligent planning agent model generates a series of optimal action sequences in a preset prediction time domain.

[0154] After executing the first action in the optimal action sequence, the intelligent planning agent model is called again for rolling optimization in the next step.

[0155] Specifically, the deviation correction module 15 is used for:

[0156] First, the transition trajectory is automatically generated, including:

[0157] Take the current actual location as the starting point and a point on the corrected target trajectory as the ending point;

[0158] Under the constraints that the wellbore curvature does not exceed a preset safety value and the tool build-up rate does not exceed a limit value, the intelligent planning agent model generates a transition trajectory that minimizes the sum of additional footage and friction.

[0159] In addition, it also includes:

[0160] The interface for manual intervention is retained, allowing for requests for manual takeover and recording of intervention samples under preset risk conditions:

[0161] When the wellbore curvature prediction value of the control command output by the intelligent planning agent model exceeds the preset safety value in a preset number of consecutive cycles, or when the reward value given by the reward function decreases continuously within a preset window and the decrease exceeds the preset threshold, the system actively generates a prompt to request manual intervention and outputs the confidence level of the current recommended action.

[0162] All states and actions during manual intervention are recorded as high-quality samples and stored in an experience pool for subsequent online fine-tuning.

[0163] In addition, it also includes:

[0164] Obtain the reservoir encounter failure rate due to trajectory deviation and the tool failure rate due to excessive deviation correction in historical drilling data;

[0165] The ratio of the reservoir drilling failure rate to the tool failure rate is calculated and set as the reward function weight adjustment factor.

[0166] The weight ratio of reservoir drilling reward to engineering safety reward in the reward function is automatically adjusted according to the reward function weight adjustment factor.

Claims

1. A method for controlling borehole trajectory correction, characterized in that, include: Acquire drilling digital twin environment data; Based on the drilling digital twin environment data, a trajectory planning mathematical model under the Markov decision process framework is constructed to obtain the state space, action space and reward function; A deep reinforcement learning agent model is deployed in the drilling digital twin environment. The trajectory planning mathematical model is used to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model to obtain an intelligent planning agent model. During actual drilling, the intelligent planning agent model is invoked to output the current control command, and the optimal action sequence is generated using a rolling optimization strategy, and the first action in the optimal action sequence is executed. The deviation between the actual trajectory and the expected trajectory is detected. When the deviation exceeds a preset deviation threshold, the intelligent planning agent model automatically generates a transition trajectory under the conditions of satisfying the wellbore curvature constraint and the tool build-up rate limit.

2. The method as described in claim 1, characterized in that, Acquire drilling digital twin environment data, including: Obtain initial geological model data, which includes target stratigraphic position, thickness, dip angle, and physical property distribution; Real-time acquisition of measurement-while-drilling data, logging-while-drilling data, and drilling engineering parameter data during the drilling process; A data cleaning algorithm combining Kalman filtering and neural networks is used to perform real-time repair and interpolation on the measurement-while-drilling data, the logging-while-drilling data, and the drilling engineering parameter data. Based on the Bayesian inversion method, the local three-dimensional geological model is dynamically updated using the repaired and interpolated data, and the updated local three-dimensional geological model is fused with the initial geological model data to obtain the drilling digital twin environment data.

3. The method as described in claim 1, characterized in that, A mathematical model for trajectory planning under the Markov decision process framework is constructed, including obtaining the state space, action space, reward function, and state transition probability model. The state transition probability model combines the bottom drill string assembly mechanics model and the formation drillability model, and uses a long short-term memory network to predict the probability distribution of the next state.

4. The method as described in claim 1, characterized in that, The state space includes the three-dimensional coordinates of the current drill bit position, the current inclination angle, the current azimuth angle, the current wellbore curvature, the vertical distance from the top boundary of the target layer, the feature vector of the formation attributes ahead, and the current remaining directional drilling capacity. The action space includes the target well inclination angle, target azimuth angle, drill bit guiding force direction and magnitude, and drilling speed parameter combination for the next drilling step; The reward function adopts a multi-objective weighted design, including trajectory tracking reward, reservoir encounter reward, engineering safety reward, drilling efficiency reward and energy consumption penalty, wherein the weight of each sub-reward adopts an adaptive adjustment mechanism.

5. The method as described in claim 1, characterized in that, Deploying a deep reinforcement learning agent model within the drilling digital twin environment, and using the trajectory planning mathematical model to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model, including: A deep deterministic policy gradient algorithm is used as the basic framework to deploy an action network and an evaluation network. In the offline pre-training phase, historical drilling data and simulated data generated by random formation models are used to train the deep reinforcement learning agent model in batches, using a priority experience replay mechanism. During the online fine-tuning phase, in the actual drilling process, after each drilling progress is completed, the status is updated using real-time collected data, and the new status and reward value after actual execution are stored in the experience pool as online experience. An incremental gradient update is performed every preset time interval.

6. The method as described in claim 1, characterized in that, The optimal action sequence is generated using a rolling optimization strategy, including: Within each control step, the intelligent planning agent model generates a series of optimal action sequences in a preset prediction time domain. After executing the first action in the optimal action sequence, the intelligent planning agent model is called again for rolling optimization in the next step.

7. The method as described in claim 1, characterized in that, Automatically generate transition trajectories, including: Take the current actual location as the starting point and a point on the corrected target trajectory as the ending point; Under the constraints that the wellbore curvature does not exceed a preset safety value and the tool build-up rate does not exceed a limit value, the intelligent planning agent model generates a transition trajectory that minimizes the sum of additional footage and friction.

8. The method as described in claim 1, characterized in that, Also includes: The manual intervention interface is retained, allowing for requests for manual takeover and recording of intervention samples under preset risk conditions: When the wellbore curvature prediction value of the control command output by the intelligent planning agent model exceeds the preset safety value in a preset number of consecutive cycles, or when the reward value given by the reward function decreases continuously within a preset window and the decrease exceeds the preset threshold, the system actively generates a prompt to request manual intervention and outputs the confidence level of the current recommended action. All states and actions during manual intervention are recorded as high-quality samples and stored in an experience pool for subsequent online fine-tuning.

9. The method as described in claim 1, characterized in that, Also includes: Obtain the reservoir encounter failure rate due to trajectory deviation and the tool failure rate due to excessive deviation correction in historical drilling data; The ratio of the reservoir drilling failure rate to the tool failure rate is calculated and set as the reward function weight adjustment factor. Based on the weight adjustment factor of the reward function, the weight ratio of reservoir drilling reward to engineering safety reward in the reward function is automatically adjusted.

10. A drilling trajectory correction control system, characterized in that, A drilling trajectory correction control method according to any one of claims 1-9, comprising: The environment acquisition module is used to acquire drilling digital twin environment data; The framework construction module is used to construct a trajectory planning mathematical model under the Markov decision process framework based on the drilling digital twin environment data, and obtain the state space, action space and reward function; The agent training module is used to deploy a deep reinforcement learning agent model in the drilling digital twin environment, and to perform offline pre-training and online fine-tuning of the deep reinforcement learning agent model using the trajectory planning mathematical model to obtain an intelligent planning agent model. The instruction execution module is used to call the intelligent planning agent model to output the current control instruction during the actual drilling process, and to generate the optimal action sequence using a rolling optimization strategy, and execute the first action in the optimal action sequence. The deviation correction module is used to detect the deviation between the actual trajectory and the expected trajectory. When the deviation exceeds a preset deviation threshold, the intelligent planning agent model automatically generates a transition trajectory under the conditions of satisfying the wellbore curvature constraint and the tool build-up rate limit.

Citation Information

Patent Citations

  • A method and system for generating optimal trajectories based on digital twins

    CN113687659B

  • Geological guarantee intelligent planning mining system based on deep reinforcement learning

    CN121031307A