Well drilling overflow early recognition method based on deep reinforcement learning
Through the deep reinforcement learning algorithm, the drilling parameter threshold is optimized, combined with the near-end strategy optimization algorithm and the Reward function, the empirical dependence and transparency problems in drilling overflow monitoring are solved, and the overflow recognition with high accuracy and fast response is achieved, reducing false alarm rates and ensuring drilling safety.
Patent Information
- Application Number
- CN202510742669.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing drilling overflow monitoring technology relies on engineer experience, has a delay in response and a high false alarm rate, making it difficult to adapt to complex geological conditions, and machine learning methods lack transparency, which affects the timely identification and control of overflows.
The early identification method of drilling overflow based on deep reinforcement learning is adopted. Through data collection and preprocessing, a near-end strategy optimization algorithm model is built, and the drilling parameter threshold is optimized in combination with the Reward function, and the feature quantity changes are monitored in real time, and the overflow pattern is identified by self-learning historical data, providing transparent decision-making logic.
It realizes high accuracy and fast response overflow identification, reduces false alarm rates, enhances engineer trust, ensures drilling safety, and reduces economic losses.
Smart Images

Figure CN120257060A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of oil drilling engineering, and particularly relates to a method for early identification of drilling overflow based on deep reinforcement learning. Background Technique
[0002] In the field of oil drilling, overflow is an extremely serious problem. Once drilling overflow occurs, it will bring various harms. From the perspective of downhole conditions, this means that the downhole pressure management is out of control. It may not only directly damage the reservoir, but the intrusion of pollutants will also affect the long-term production efficiency of the reservoir. In severe cases, it may even cause permanent damage to the reservoir, resulting in the loss of energy resources. Economically, dealing with overflow accidents requires a large amount of manpower, material resources and financial resources, bringing a heavy burden to the drilling company. More seriously, if it is not properly handled in the initial stage of overflow, it may quickly evolve into a well kick and even a blowout. As an extreme well control failure event, a blowout will cause serious consequences such as toxic gas leakage, fire, explosion, etc. It not only directly threatens the lives of the drilling platform operators, but also has a long-term negative impact on the surrounding ecological environment. Therefore, the early identification and timely treatment of overflow are crucial for ensuring personnel safety, reducing economic losses and preventing the occurrence of more serious accidents.
[0003] At present, the drilling overflow monitoring technologies mainly include traditional threshold methods and machine learning methods. The traditional threshold method is to set a series of pre-determined safety thresholds to monitor key drilling parameters in real time, such as the flow difference between the inlet and outlet, standpipe pressure, methane, ethane, pump strokes, and hook height. When the parameter change exceeds the threshold, the monitoring system will issue an alarm. The setting of the threshold usually depends on historical data analysis, geological condition assessment and the experience of engineers. Although this method provides a basic means for drilling overflow monitoring, there are many problems. It highly depends on the personal experience of engineers, and different engineers may have different judgments on the same situation. When the drilling environment changes rapidly, the response to real-time data will be delayed, affecting the timely control of overflow. In addition, the geological conditions are complex and changeable, and a single threshold judgment is prone to false alarms or missed alarms, and it is difficult to adapt to all drilling environments, especially in the exploration and development of unconventional oil and gas resources, the problem is more prominent.
[0004] In recent years, the application of machine learning methods in the field of drilling overflow monitoring has gradually increased. Common techniques include support vector machines, decision trees, random forests, and deep learning networks. These methods can analyze a large amount of drilling data and accurately identify overflow signs in a timely manner. For example, Shi Xiaoyan et al. proposed a solution based on the random forest algorithm, which improved the accuracy and real-time performance of overflow detection; Sun Baojiang et al. proposed an integrated pattern recognition model and successfully diagnosed gas rush faults; Li Yufei et al. proposed an intelligent early overflow identification method based on SVM and D-S evidence theory, which improved the reliability of monitoring. However, most machine learning methods operate independently of traditional threshold methods. For front-line drilling workers and on-site engineers, they are usually regarded as "black box" operations, and the internal working mechanisms and decision-making logics lack transparency, making it difficult for front-line drilling workers to understand the basis of early warnings and affecting the trust in the system and the timeliness of early warning responses. When the algorithm prediction does not match the worker's experience, it may lead to hesitation in decision-making or even ignoring the algorithm early warning, bringing safety risks in emergency situations.
[0005] To solve the above problems, the present invention proposes an early identification method for drilling overflow based on deep reinforcement learning. Summary of the Invention
[0006] The purpose of the present invention is to provide an early identification method for drilling overflow based on deep reinforcement learning, aiming to solve the problems raised in the above background technology.
[0007] The purpose of the present invention is achieved through the following technical solutions: An early identification method for drilling overflow based on deep reinforcement learning includes the following steps: Step 1: Data collection and preprocessing; Clean the collected data using real drilling historical data; construct samples using the time window method, select six characteristic quantities including the flow difference between the outlet and inlet, standpipe pressure, methane, ethane, pump strokes, and hook height, calculate the mean value, slope, and variance of the characteristic quantities within the time window, and determine whether it is an overflow by judging whether the change in the statistical quantities between adjacent windows exceeds the threshold and record the label; Step 2: Construction and optimization of an early identification model for drilling overflow based on the proximal policy optimization algorithm, including parameter initialization, main iteration loop, policy parameter update, value function learning, and iteration loop steps; Step 3: Early identification of drilling overflow; Preprocess the real-time data and input it into the early identification model for drilling overflow based on the proximal policy optimization algorithm; the model uses the optimized parameters of the proximal policy optimization algorithm to analyze the changes in the flow difference between the inlet and outlet, standpipe pressure, methane, ethane, pump strokes, and hook height, and determine whether there are overflow signs by discriminating the overflow drilling process flow; if overflow signs are detected, issue an early warning according to the policy determined by the Reward function.
[0008] Furthermore, in step 1, when cleaning the data, the invalid data points are replaced with zeros and the features with incomplete information are removed; when constructing the samples using the time window method, a time series matrix is constructed with 380 time slices as a unit, and the sliding step size is 1.
[0009] Furthermore, the specific process of step 2 is as follows: Step 21: Parameter initialization; Set the hyperparameter truncation factor , the number of sub-iterations for policy update M and the number of sub-iterations for value function update B ; Initialize the policy parameters θ and the initial value function parameters ; Step 22: Main iteration loop; Starting from the first iteration, that is , each iteration performs operations in the following order: Trajectory collection and reward calculation: Execute the current policy in the drilling environment π θk , collect the trajectory dataset D k ={ τ i}, and calculate the cumulative reward at each time step G t , where τ i is the collected trajectory data; Advantage function estimation: Based on the current value function , use the advantage estimation method to calculate the advantage value of each state-action pair ; Step 23: Policy parameter update; Enter the policy optimization stage, and update the policy parameters through the following sub-iteration process: For execute in sequence: Calculate the probability ratio of the new and old policies: ; In the formula, is the probability ratio of the new and old policies; are the policies before and after update; is the policy before update; θ old are the policy parameters before update; are the policy parameters after update; is the advantage value of the state-action pair; is the time step t where the environment state is located; The Adam optimizer is adopted, and the truncated objective function of PPO is maximized by the stochastic gradient ascent method: ; In the formula, is the policy network parameter after the k +1-th policy optimization; is the number of iterations; is the trajectory point; is the trajectory set; is the time step; is the advantage value of the state-action pair; is the probability ratio of the old and new policies; is the clipping function; is the hyperparameter truncation factor; Step 24: Value function learning; After completing the update of the policy parameters, enter the value function optimization stage, and update the parameters through the following sub-iteration process: For execute in sequence: The mean squared error of the value function is minimized by the gradient descent method: ; In the formula, is the value network parameter after the k +1-th policy optimization; is the estimation of the value network for the state ; is the cumulative reward for each time step; Step 25: Iterative loop; Repeat Step 22 to Step 24 until the algorithm converges or reaches the preset termination condition to obtain the optimized model.
[0010] Furthermore, the Reward function construction step is also included in Step 2, and the specific process is as follows: Set a "sweet zone", assign the highest score to the marked points within the sweet zone, assign negative scores to the marked points outside the sweet zone but in the normal state before and after the overflow occurs, and assign positive scores much lower than the scores of the marked points within the sweet zone to the points where the overflow actually occurs; Define the structured action space: ; In the formula, is the action space, that is, the set of all possible actions; is the action vector; is each dimension of the action vector; is the set of real numbers, is the 18-dimensional set of real numbers; The scoring function formula is as follows: ; ; In the formula, is the scoring function; represents the number of marked points in the sweet spot; represents the number of true overflow points among the marked points outside the sweet spot; represents the number of non-overflow points among the marked points outside the sweet spot; represents the number of that meet the conditions in the square brackets; is the coordinate of the true overflow point; is the position of the i th overflow point predicted by the model in the time series.
[0011] Furthermore, in the said step 3, the specific process for discriminating the overflow drilling process is as follows: First, check whether the flow difference between the inlet and outlet or the standpipe pressure increases. If it does not increase, the current flow change may be normal fluctuation. If it increases, check whether the methane or ethane content increases. If the methane or ethane content increases, it is judged that there is an overflow sign. If it does not increase, check whether the pump strokes increase. If the pump strokes increase, the current flow change may be normal fluctuation. If it does not increase, check whether the hook height decreases. If the hook height decreases, the current flow change may be normal fluctuation. If it does not decrease, it is judged that there is an overflow sign.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. High accuracy and high detection rate: The algorithm proposed by the present invention has a detection rate of 100% in the previous time window of overflow, and the overall accuracy is 79%. Compared with traditional machine learning algorithms such as support vector machine (SVM), random forest (RF), and backpropagation neural network (BP-NN), it has a higher correct recognition rate and a lower false alarm rate, can provide timely and accurate early warning for drilling engineers, help them take preventive measures to ensure the safety of the wellbore, reduce economic losses, and has outstanding practical application value.
[0013] 2. Fast response and precise positioning: The processing speed of the algorithm proposed by the present invention exceeds that of the comparison algorithm and reacts faster to overflow events in practical applications. The nearest distance (ND) to the initial overflow point is only 5, and the average distance (AD) to the initial overflow point is only 193. It can accurately locate the starting point of the overflow, gain precious time for preventing blowout accidents, greatly improve the response speed and positioning accuracy to the overflow, and is conducive to timely controlling the harm of the overflow.
[0014] 3. Effectively reduce the false alarm rate: By constructing a reasonable Reward function, the present invention distinguishes normal operations from overflow events, assigns scores to marked points in different regions, and reduces the possibility of false alarms. In drilling operations, false alarms are avoided from interfering with production order, increasing costs and risks, making the monitoring system more reliable, and assisting engineers in making accurate decisions based on early warning information.
[0015] 4. Ensure operation safety and reduce economic losses: The algorithm proposed in the present invention monitors overflow in advance with a low false alarm rate, provides timely early warnings for engineers, enables them to adjust drilling parameters, implement well killing operations, etc., avoids the deterioration of overflow, ensures the safety of personnel's lives, reduces the input of human, material and financial resources for dealing with overflow accidents, reduces financial costs, and plays a key role in enhancing the safety and economic benefits of drilling operations.
[0016] In summary, the present invention combines the traditional threshold method with machine learning, develops an automated model that does not rely on personal experience, optimizes the drilling parameter threshold using the deep reinforcement learning algorithm, monitors changes before overflow by capturing multiple characteristic quantities such as the flow difference at the inlet and outlet, standpipe pressure, and methane in real time, and can also self-learn historical data to identify overflow patterns, reducing the dependence on on-site operators. In addition, the present invention focuses on demonstrating the decision-making process and logical basis, enhancing engineers' understanding and trust in the model, providing a more reliable overflow monitoring guarantee for drilling operations, further reducing the risk of overflow accidents, and ensuring wellbore safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is the method flow chart of the present invention.
[0018] Figure 2 is the target function scoring flow chart for the starting center area of overflow.
[0019] Figure 3 is the mean change curve of the hook height, standpipe pressure, and pump strokes in the data set.
[0020] Figure 4 is the mean change curve of methane, ethane, and the flow difference at the inlet and outlet in the data set.
[0021] Figure 5 is the drilling process flow chart for discriminating overflow.
[0022] Figure 6 is the overflow monitoring situation of data set 2 under a window size of 380, a step size of 380, and a scoring function.
[0023] Figure 7 is the overflow monitoring situation of data set 7 under a window size of 380, a step size of 380, and a scoring function. DETAILED DESCRIPTION OF THE INVENTION
[0024] To have a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the following provides a detailed description of the technical solution of the present invention, but it should not be construed as a limitation on the scope of implementation of the present invention.
[0025] The present invention provides a method for early identification of drilling overflow based on deep reinforcement learning. The method includes the following steps (as Figure 1 shown): Step 1: Data collection and preprocessing; Using the real drilling historical data of a certain well, the collected data is cleaned, invalid data points are replaced with zeros, and features with incomplete information are removed. The time window method is used to construct samples, and a time series matrix is constructed with 380 time slices as a unit, and the sliding step size is set to 1. Six feature quantities, namely the difference between the outlet and inlet flow rates, standpipe pressure, methane, ethane, pump strokes, and hook height, are selected. The mean value, slope, and variance of the feature quantities within the time window are calculated, and it is determined whether there is an overflow by judging whether the change in the statistical quantity between adjacent windows exceeds the threshold, and the label is recorded. Figure 3 And Figure 4 shows the change curve of the mean value of the feature quantities of part of the data set.
[0026] Step 2: Construction and optimization of the early identification model of drilling overflow based on the Proximal Policy Optimization (PPO) algorithm; The core of this model lies in accurately identifying the mutation of feature quantities during the drilling process, and judging the change point by comparing the changes in the mean value, slope, and variance of the feature quantities within the front and back time windows. Since the drilling data collected on-site has vibration noise, it is necessary to set appropriate thresholds to distinguish whether the change in the feature quantity is a change point. The setting of the threshold should take into account both capturing early signs and avoiding false alarms. The selection of the time window and the determination of the step size also affect the algorithm performance. A too large window will lead to untimely response, and a too small window will be sensitive to noise; the step size determines the sampling density and affects the fineness of feature extraction. To optimize these parameters (threshold, time window size, and step size), the PPO algorithm is used for parameter optimization. The PPO algorithm introduces a strategy update method of the clipping probability ratio, which can reduce the performance instability during the training process and improve the learning efficiency, so as to determine the optimal threshold, time window size, and step size, realize the efficient identification of the change points of drilling feature quantities, and improve the accuracy and reliability of overflow monitoring.
[0027] The specific steps are as follows: Step 21: Parameter initialization; Set the hyperparameters truncation factor , the number of sub-iterations for policy update M and the number of sub-iterations for value function update B ; Initialize the policy parameter θ and the initial value function parameter .
[0028] Step 22: Main Iteration Loop Starting from the first iteration (i.e., ), the following operations are performed in each iteration in the following order: Trajectory Collection and Reward Calculation: Execute the current policy in the drilling environment π θk to collect a trajectory dataset D k ={ τ i}, and calculate the cumulative reward for each time step G t , where τ i is the collected trajectory data.
[0029] Advantage Function Estimation: Based on the current value function , use an advantage estimation method (such as GAE) to calculate the advantage values of each state-action pair .
[0030] Step 23: Policy Parameter Update Enter the policy optimization phase and update the policy parameters through the following sub-iteration process: For , perform the following operations in sequence: Calculate the probability ratio of the old and new policies: ; where is the probability ratio of the old and new policies; are the policies before and after the update; is the policy before the update; θ old are the policy parameters before the update; are the policy parameters after the update; is the advantage value of the state-action pair; is the time step t in the environment state.
[0031] Use the Adam optimizer to maximize the truncated objective function of PPO through stochastic gradient ascent: ; where is the policy network parameter after the k +1-th policy optimization; is the iteration number; is the trajectory point; is the trajectory set; is the time step; is the advantage value of the state-action pair; is the probability ratio of the old and new policies; is a clipping function; is a hyperparameter truncation factor; Step 24: Value function learning; After completing the policy parameter update, enter the value function optimization stage, and update the parameters through the following sub-iteration process: For Execute in sequence: Use the gradient descent method to minimize the mean square error of the value function: ; In the formula, is the value network parameter after the k +1-th policy optimization; is the valuation of the state by the value network; is the cumulative reward for each time step; By continuous updating, the value function can more accurately fit the cumulative reward.
[0032] Step 25: Iterative loop; Repeat Step 22 to Step 24 until the algorithm converges or reaches the preset termination condition, so as to obtain the optimized model.
[0033] Reward function construction: Construct a scoring function for precisely controlling the position of the recognition point, and use this as the Reward function for the PPO algorithm to determine the starting center area of the overflow. Set a "sweet spot" (the predicted overflow point is within the time window before and after the real overflow point). Based on the early recognition and warning requirements of the overflow, assign the highest score (such as adding 1000 points) to the marked points within the sweet spot, so that they can be given priority attention; assign negative numbers (such as deducting 50 points) to the marked points in the normal state before and after the overflow but outside the sweet spot (the marked points outside the sweet spot are non-overflow points), so as to reduce false alarms; give a positive but relatively low score (such as adding 100 points) to the marked points where the overflow actually occurs (the marked points outside the sweet spot are real overflow points), and encourage the algorithm to concentrate the recognition points near the starting point of the overflow. Define the structured action space: ; In the formula, is the action space, that is, the set of all possible actions; is the action vector, which is an 18-dimensional real vector; is each dimension of the action vector; is the set of real numbers, is the 18-dimensional set of real numbers. Among them, each dimension of the action vector a is restricted within a very small range to ensure that the fine changes in the action can precisely control the position of the recognition point, so that the PPO algorithm can optimize the action selection strategy. The formula of the scoring function is as follows: ; ; In the formula, is the scoring function; represents the number of marked points in the sweet spot; represents the number of true overflow points among the marked points outside the sweet spot; represents the number of non-overflow points among the marked points outside the sweet spot; represents the number of that meet the conditions in the square brackets; is the coordinate of the true overflow point; is the position of the i th overflow point predicted by the model in the time series.
[0034] The score is determined by calculating the number of marked points in the sweet spot, the number of true overflow points among the marked points outside the sweet spot, and the number of non-overflow points among the marked points outside the sweet spot. Figure 2 Intuitively demonstrates the working mechanism of this function.
[0035] Step 3: Early identification of drilling fluid overflow; During the drilling operation, drilling data is collected in real time. After preprocessing the collected data, it is input into the early identification model of drilling fluid overflow based on the Proximal Policy Optimization (PPO) algorithm. The model uses the parameters optimized by the PPO algorithm to analyze the changes in characteristic quantities such as the flow rate difference between the inlet and outlet, standpipe pressure, methane, ethane, pump strokes, and hook height. By judging the drilling process of fluid overflow, it determines whether there are signs of fluid overflow. If signs of fluid overflow are detected, a warning is issued according to the strategy determined by the Reward function to remind the staff to take corresponding measures to prevent blowout accidents and ensure the safety of drilling operations.
[0036] The drilling process for judging fluid overflow is as follows: As Figure 5 shown, when monitoring data, first check whether the flow rate difference between the inlet and outlet or the standpipe pressure increases. If it does not increase, the current flow rate change may be normal fluctuations. If it increases, check whether the methane or ethane content increases; if the methane or ethane content increases, it is judged that there are signs of fluid overflow. If it does not increase, check whether the pump strokes increase; if the pump strokes increase, the current flow rate change may be normal fluctuations. If it does not increase, check whether the hook height decreases; if the hook height decreases, the current flow rate change may be normal fluctuations. If it does not decrease, it is judged that there are signs of fluid overflow.
[0037] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.
[0038] Example 1: Overflow monitoring of dataset 2; Use dataset 2 in the real drilling historical data of a certain well from October 10 to November 10, 2015. This data contains the data 60 minutes before the occurrence of overflow and during the overflow, and the sampling rate is slightly lower than 1 Hz (50 samples per minute). When constructing samples, the time window method is adopted, and a time series matrix is constructed in units of 380 time slices, with a sliding step of 380. Six characteristic quantities, namely the difference between the outlet and inlet flow rates, standpipe pressure, methane, ethane, pump strokes, and hook height, are selected, and a drilling overflow early identification model based on the Proximal Policy Optimization (PPO) algorithm is used, combined with the Reward function for monitoring and analysis. The model will analyze the changes in the six characteristic quantities according to the parameters optimized by the PPO algorithm, and judge whether there are signs of overflow by discriminating the overflow drilling process flow. If there are signs of overflow, an early warning will be issued according to the strategy determined by the Reward function.
[0039] As can be seen from Figure 6 , the vertical dotted line represents the actual initial overflow point, the blue line shows the dynamic change of the difference between the inlet and outlet flow rates over time, and the red dots represent the overflow points predicted by the model. The model can accurately predict the occurrence time of the overflow event. The red predicted points are close to the vertical dotted line, indicating that the model has a high prediction accuracy. At the same time, the change of the blue line reflects that the model is sensitive to the flow rate change and has a fast response speed. Moreover, the red dots also show that the model has the ability to give early warnings before the actual overflow occurs, and the false alarm rate is low, which strongly proves the effectiveness of the model in overflow monitoring.
[0040] Example 2: Overflow monitoring of dataset 7; In this example, dataset 7 in the real drilling historical data during the same time period as in Example 1 is used. The data processing method, sample construction method, and the monitoring model and function used are all the same as in Example 1.
[0041] As Figure 7 shown, under the conditions of the optimal time window size, optimal step, and scoring function, the model performs excellently. The actual initial overflow point is close to the red predicted point, indicating that the model can accurately predict the time of overflow occurrence. The change of the curve of the difference between the inlet and outlet flow rates reflects that the model is very sensitive to the flow rate change monitoring. The predicted points can not only accurately mark the overflow time, but also give early warnings in advance, and the false alarm rate is low, which verifies the reliability and effectiveness of the model in practical applications again.
[0042] Example 3: Performance evaluation; The proposed algorithm is compared with the Support Vector Machine (SVM) algorithm, the Random Forest (RF) algorithm which is an ensemble of decision trees, and the traditional Backpropagation Neural Network (BP-NN) algorithm through comparative experiments. The same dataset is used in the experiments, and the same preprocessing steps are applied to all algorithms. Metrics such as accuracy, the nearest distance to the initial overflow point, precision, recall, the average distance to the initial overflow point, and the number of monitored overflow points within one time window before the initial overflow point are used to comprehensively evaluate the performance of the algorithms. These metrics are closely related to the actual industrial requirements and can provide guidance for algorithm improvement and practical applications.
[0043] (1) Accuracy (ACC) measures the ability of the algorithm to correctly identify overflow events, which is the ratio of the total number of correctly identified positive and negative examples to the total number of all test samples. The formula is as follows: ; where, (True Positives) is the number of samples correctly predicted as positive, (True Negatives) is the number of samples correctly predicted as negative, FP (False Positives) is the number of samples wrongly predicted as positive, FN (FalseNegatives) is the number of samples wrongly predicted as negative.
[0044] (2) The Nearest Distance (ND) to the initial overflow point focuses on the distance between the overflow point identified by the algorithm and the actual overflow starting point. The closer the distance, the higher the positioning accuracy of the algorithm. The formula is as follows: ; where, is the position of the i th predicted overflow point, is the position of the actual overflow starting point.
[0045] (3) Precision (P) refers to the proportion of actually positive samples among all samples predicted as positive, which reflects the accuracy of the algorithm. The formula is as follows: ; (4) Recall (R), also known as the true positive rate, refers to the proportion of samples correctly predicted as positive among all actually positive samples, which reflects the completeness of the algorithm. The formula is as follows: ; (5) The average distance (AD) from the initial overflow point provides a statistical perspective. By calculating the average of the distances between all predicted overflow points and the actual starting point of the overflow, it evaluates the overall positioning ability of the algorithm. The formula is as follows: ; Where, N is the total number of predicted overflow points, is the position of the i -th predicted overflow point, is the position of the actual starting point of the overflow.
[0046] (6) The number of monitored overflow points (Counts) within one time window before the initial overflow point: This criterion evaluates the number of overflow symptom points that the algorithm can monitor within the set time window before the actual occurrence of the overflow. This indicator helps to measure the early warning ability of the algorithm for upcoming overflow events. The formula is as follows: ; Where, is the position of the actual starting point of the overflow, W is the length of the time window, is the position of the i -th predicted overflow point, and N is the number of all overflow points predicted by the model.
[0047] The comparison results are shown in Table 1: Table 1 Comparison of this algorithm with other algorithms
[0048] It can be seen from the data in Table 1 that this algorithm shows advantages in multiple evaluation indicators on the test set. In terms of accuracy, it reaches 0.79. Compared with other algorithms, it has a higher correct recognition rate and a lower false alarm rate; it is faster in processing speed and can respond to overflow events more quickly; it performs better in indicators such as the closest distance to the initial overflow point (only 5) and the average distance from the initial overflow point (only 193), meaning higher positioning accuracy. Considering all indicators, this algorithm has excellent performance in the overflow monitoring task, with higher reliability and effectiveness, providing a better solution for drilling overflow monitoring.
[0049] The above is only the preferred implementation mode of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
Claims
1. A method for early identification of drilling overflow based on deep reinforcement learning, characterized in that, The following steps are involved: Step 1: Data collection and preprocessing; The collected data are cleaned using real historical drilling data; the time window method is used to construct samples, and six characteristic quantities, namely, the difference between the outlet and inlet flow rate, vertical pressure, methane, ethane, pump stroke and hook height, are selected to calculate the mean, slope and variance of the characteristic quantities within the time window. By judging whether the change of the statistics of adjacent windows exceeds the threshold, it is determined whether it is overflow and the label is recorded; Step 2: Construction and optimization of the early identification model of drilling overflow based on the proximal strategy optimization algorithm, including parameter initialization, main iteration loop, strategy parameter update, value function learning and iteration loop steps; Step 3: Early identification of drilling overflow; After preprocessing, the real-time data is input into the early identification model of drilling overflow based on the proximal strategy optimization algorithm. The model uses the parameters optimized by the proximal strategy optimization algorithm to analyze the changes in inlet and outlet flow difference, vertical pressure, methane, ethane, pump stroke and hook height, and determines whether there are signs of overflow by identifying the overflow drilling process. If signs of overflow are detected, an early warning is issued according to the strategy determined by the Reward function.
2. The early identification method for drilling overflow based on deep reinforcement learning according to claim 1, characterized in that In step 1, invalid data points are replaced with zeros during data cleaning and features with incomplete information are removed; When the time window method is used to construct samples, the time series matrix is constructed in units of 380 time slices, and the sliding step size is 1.
3. The early identification method of drilling overflow based on deep reinforcement learning according to claim 1, characterized in that The specific process of step 2 is as follows: Step 21: Parameter initialization; Set the hyperparameter truncation factor , the number of sub-iterations for policy update M and the number of sub-iterations for value function update B ; Initialize the policy parameters θ and the initial value function parameters ; Step 22: Main iteration loop; Starting from the first iteration, that is , each iteration performs operations in the following order: Trajectory Collection and Reward Calculation: Execute the current policy in the drilling environment π θk , collect a trajectory dataset D k ={ τ i}, and calculate the cumulative reward for each time step G t , where τ i is the collected trajectory data; Advantage function estimation: Based on the current value function , the advantage value of each state-action pair is calculated using the advantage estimation method ; Step 23: Update strategy parameters; Entering the strategy optimization phase, the strategy parameters are updated through the following sub-iteration process: For Perform in sequence: Calculate the probability ratio of the new and old strategies: ; wherein, is the probability ratio of the new and old policies; are the policies before and after update; is the policy before update; θ old is the policy parameter before update; is the policy parameter after update; is the advantage value of the state-action pair; is the time step t is the environmental state at which it is located; The Adam optimizer is used to maximize the truncated objective function of PPO through the stochastic gradient ascent method: ; Wherein, is the policy network parameter after the k +1-th policy optimization; is the number of iterations; is the trajectory point; is the trajectory set; is the time step; is the advantage value of the state-action pair; is the probability ratio of the old and new policies; is the clipping function; is the hyperparameter truncation factor; Step 24: Value function learning; After completing the policy parameter update, enter the value function optimization phase and update the parameters through the following sub-iteration process: To Execute in sequence: Use gradient descent to minimize the mean square error of the value function: ; In the formula, is the value network parameter after the k +1 - th policy optimization; is the valuation of the value network for the state ; is the cumulative reward at each time step; Step 25: Iteration loop; Repeat steps 22 to 24 until the algorithm converges or reaches a preset termination condition to obtain an optimized model.
4. The early identification method for drilling overflow based on deep reinforcement learning according to claim 1, wherein The step 2 also includes a Reward function construction step, and the specific process is as follows: Set a "sweet spot" and give the highest score to the marked points in the sweet spot, give negative scores to the marked points outside the sweet spot but in normal state before and after the overflow occurs, and give positive scores much lower than the scores of the marked points in the sweet spot to the points where the overflow actually occurs; Define the structured action space: ; In the formula, is the action space, that is, the set of all possible actions; is the action vector; is each dimension of the action vector; is the set of real numbers, is the 18-dimensional set of real numbers; The scoring function formula is as follows: ; ; In the formula, is the scoring function; represents the number of marked points in the sweet zone; represents the number of marked points outside the sweet zone that are true overflow points; represents the number of marked points outside the sweet zone that are non-overflow points; represents the number that meets the conditions within the square brackets; is the coordinate of the true overflow point; is the position of the i th overflow point predicted by the model in the time series.
5. The early identification method for drilling overflow based on deep reinforcement learning according to claim 1, wherein In step 3, the specific process flow of identifying overflow drilling is as follows: First, check whether the inlet and outlet flow difference or vertical pressure increases. If not, the current flow change may be a normal fluctuation. If it increases, check whether the methane or ethane content increases. If the methane or ethane content increases, it is judged that there are signs of overflow. If not, check whether the pump stroke increases. If the pump stroke increases, the current flow change may be a normal fluctuation. If not, check whether the hook height decreases. If the hook height decreases, the current flow change may be a normal fluctuation. If not, it is judged that there are signs of overflow.
Citation Information
Patent Citations
Petroleum drilling well leakage and overflow intelligent detection method and device and electronic equipment
CN111414955A
Drainage system real-time control method and device
CN112068420A
Overflow working condition prediction model training method and device and overflow working condition prediction method
CN114997485A
Regional multi-system collaborative development evaluation method and device, electronic equipment and medium
CN115545589A
Training and predicting method for intelligent prediction model of well drilling overflow working condition and underground overflow risk probability prediction system
CN116796647A
Cited By
Exploration drilling parameter control method and system
CN120575835A
Early overflow intelligent identification method and system based on knowledge guidance and residual enhancement
CN122333112A
Early overflow intelligent identification method and system based on knowledge guidance and residual enhancement
CN122333112B