Mechanical drilling speed prediction method and system based on particle swarm and random forest
By combining particle swarms and random forest methods in mechanical drilling speed prediction, the hyperparameters of the random forest model are optimized, and the problem of cumbersome prediction process and relying on manual parameter adjustment in the existing technology is solved, and accurate and reliable prediction in complex drilling environments is achieved, which improves drilling efficiency and reduces costs.
Patent Information
- Application Number
- CN202510159511.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-13
AI Technical Summary
When predicting mechanical drilling speed, the prior art relies on a single model or requires manual adjustment of hyperparameters, which leads to cumbersome and time-consuming process and relies on domain knowledge, making it difficult to provide accurate and reliable prediction results in complex drilling environments.
Using a method based on particle swarm and random forest, the hyperparameters of the random forest model are optimized through the inertial weight parameter adaptive strategy and the particle swarm algorithm of the Cauchy variant operator, the initial mechanical drilling speed prediction model is constructed, and the prediction results are obtained through training.
It realizes automated search of hyperparameter space, avoids the burden of manual parameter adjustment, provides accurate and reliable mechanical drilling speed prediction in complex drilling environments, improves drilling efficiency and reduces costs.
Smart Images

Figure CN120145135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model prediction technology, and in particular to a mechanical drilling rate prediction method and system based on particle swarm and random forest. Background Art
[0002] Oil drilling is a project with huge investment and high risks. Among them, the mechanical rate of penetration (ROP) is one of the key factors affecting drilling efficiency and an important economic indicator of oil engineering drilling operations. Therefore, how to effectively improve the mechanical drilling rate, reduce cost consumption and avoid drilling risks has become the core of solving the problem of energy shortage. The level of ROP is mainly related to the formation characteristics and drilling technical measures. However, the characteristics of the formation cannot be changed artificially, so we can only optimize the drilling technical measures to improve efficiency. The technical variables that affect drilling efficiency include drill bit type, mud performance, drilling pressure, drilling rate, torque, drilling fluid viscosity and hydraulic parameters. Traditionally, drilling engineers obtain the optimal drilling rate value through empirical reasoning and indoor field core drillability tests. However, these methods are time-consuming and labor-intensive and have limitations, and it is difficult to meet the needs of field engineering.
[0003] With the development of big data, more useful information can be obtained in the analysis and processing of data. The application of artificial intelligence in accurately predicting mechanical drilling rate is an important indicator for measuring drilling performance. In recent years, it has aroused great interest in oil and gas well drilling operations. Ahmed et al. compared the prediction accuracy of four models: artificial neural network (ANN), extreme learning machine (ELM), support vector regression (SVR) and least squares support vector regression (LS-SVR). All four computational intelligence technologies showed performance within an acceptable accuracy range. Alsaihati et al. introduced an integrated model based on random forest (RF), in which artificial neural network (ANN) and adaptive neural fuzzy inference system (ANFIS) are the basic learner models to predict the ROP of different rock formations. Zuo Diyi used random forest regression method and gradient boosting tree regression method to establish a mechanical drilling rate prediction model that meets the characteristics of various wells, and optimized and verified the model, achieving good application results.
[0004] However, existing technologies often use a single model or require manual adjustment of the model's hyperparameters, which is a cumbersome and time-consuming process that relies on domain knowledge and has low drilling efficiency. Therefore, how to search the hyperparameter space, find the optimal hyperparameter configuration, avoid the burden and time consumption of manual parameter adjustment, and provide accurate and reliable prediction results in the face of complex drilling environments, massive data, and variable drilling parameters is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of this, it is necessary to provide a mechanical drilling rate prediction method and system based on particle swarm and random forest to determine the optimal hyperparameters of the model, avoid the burden and time consumption of manual parameter tuning, and provide accurate and reliable prediction results in the face of complex drilling environments, massive data, and variable drilling parameters.
[0006] To achieve the above object, in a first aspect, the present invention provides a mechanical drilling rate prediction method based on particle swarm and random forest, including:
[0007] Collecting original logging data and extracting feature data from the original logging data;
[0008] Optimizing the hyperparameters in the random forest model based on the inertia weight parameter adaptive strategy and the particle swarm algorithm introducing the Cauchy mutation operator to obtain an optimized random forest model, and constructing an initial mechanical drilling rate prediction model based on the optimized random forest model;
[0009] Training the initial mechanical drilling rate prediction model based on the feature data to obtain a fully trained mechanical drilling rate prediction model, and using the fully trained mechanical drilling rate prediction model to output the predicted mechanical drilling rate.
[0010] In a possible implementation manner, the original logging data includes drilling depth, weight on bit, drilling rate, and mud properties.
[0011] In a possible implementation manner, extracting the feature data from the original logging data includes:
[0012] Processing the outliers, missing values, and duplicate values in the original logging data to obtain first logging data;
[0013] Performing linear normalization processing on the first logging data to obtain the preprocessed logging data;
[0014] Selecting the features related to the target attribute from the preprocessed logging data to obtain the feature data.
[0015] In a possible implementation manner, the particle swarm algorithm based on the inertia weight parameter adaptive strategy and introducing the Cauchy mutation operator includes:
[0016] Adapting and adjusting the inertia weight parameter in the particle swarm algorithm;
[0017] Introducing the Cauchy mutation operator to guide the particles in the particle swarm to search.
[0018] In a possible implementation manner, the adapting and adjusting the inertia weight parameter in the particle swarm algorithm includes:
[0019] Linearly adaptively adjust the inertia weight parameter:
[0020]
[0021] Or adaptively adjust the inertia weight parameter in a curve:
[0022]
[0023] Where, w t represents the adjusted inertia weight parameter, w max represents the maximum inertia weight value, w min represents the minimum inertia weight value, iter represents the number of iterations, and iter_max represents the maximum number of iterations.
[0024] In a possible implementation manner, the Cauchy mutation perturbation formula of the Cauchy mutation operator includes:
[0025] p' gd = p gd +(x max (d)-x min (d))·Cauchy(o, s)
[0026] Where, p' gd represents the mutated Cauchy operator p gd represents the original Cauchy operator, x max (d) represents the maximum value of the particle in dimension d, x min (d) represents the minimum value of the particle in dimension d, and s represents the proportionality parameter.
[0027] In a possible implementation manner, the particle swarm optimization algorithm based on the inertia weight parameter adaptive strategy and the introduction of the Cauchy mutation operator optimizes the hyperparameters in the random forest model to obtain an optimized random forest model, including:
[0028] Determine the solution space range of the random forest model hyperparameters;
[0029] Based on the inertia weight parameter adaptive strategy, dynamically adjust the inertia weight and learning factor in the particle swarm, and perform a global search on the particle swarm based on the Cauchy mutation operator, so that the particles in the particle swarm perform iterative updates of velocity and position within the solution space range, and determine the fitness function values under different hyperparameter combinations;
[0030] Search for the hyperparameter combination with the minimum fitness function value, and use the hyperparameter combination with the minimum fitness function value as the optimal hyperparameter combination of the random forest model to obtain an optimized random forest model.
[0031] In a possible implementation manner, training the initial mechanical drilling rate prediction model based on the feature data to obtain a completely trained mechanical drilling rate prediction model includes:
[0032] Using the features in the feature data that affect the mechanical drilling rate of the oil and gas well as input feature vectors and the mechanical drilling rate in the feature data as output feature variables to train the initial mechanical drilling rate prediction model, so as to obtain a completely trained mechanical drilling rate prediction model.
[0033] In a possible implementation manner, the method further includes:
[0034] Determining the predicted value and the actual value of the mechanical drilling rate;
[0035] Analyzing the predicted value and the actual value based on the Pearson correlation coefficient to determine the correlation between the predicted value and the actual value.
[0036] In a second aspect, the present invention further provides a mechanical drilling rate prediction system based on particle swarm and random forest, including:
[0037] A data preprocessing module, configured to collect original logging data and extract feature data from the original logging data;
[0038] A model optimization module, configured to optimize the hyperparameters in the random forest model based on an inertia weight parameter adaptive strategy and a particle swarm algorithm introducing a Cauchy mutation operator to obtain an optimized random forest model, and construct an initial mechanical drilling rate prediction model based on the optimized random forest model;
[0039] A mechanical drilling rate prediction module, configured to train the initial mechanical drilling rate prediction model based on the feature data to obtain a completely trained mechanical drilling rate prediction model, and use the completely trained mechanical drilling rate prediction model to output the predicted mechanical drilling rate.
[0040] The beneficial effects of the present invention are:
[0041] The present invention first analyzes the original logging data to extract the features affecting the mechanical drilling rate, then uses a machine learning algorithm to construct a prediction model, tunes the parameters of the random forest through an improved particle swarm optimization algorithm, accurately predicts the mechanical drilling rate, and further helps oil well technicians optimize the drilling cycle, adjust parameters to increase the drilling rate, realize the optimization of drilling equipment resources, and reduce the drilling cost. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic flow chart of a particle swarm optimization algorithm provided by an embodiment of the present invention;
[0044] Figure 2 It is a schematic flow chart of a random forest algorithm provided by an embodiment of the present invention;
[0045] Figure 3 It is a schematic flow chart of an embodiment of a mechanical drilling rate prediction method based on particle swarm and random forest provided by the present invention.
[0046] Figure 4 It is a schematic structural diagram of an embodiment of a mechanical drilling rate prediction system based on particle swarm and random forest provided by the present invention. Detailed implementation manners
[0047] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings. Among them, the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.
[0048] The descriptions such as "first", "second", etc. involved in the embodiments of the present invention are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0049] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0050] To better understand the present invention, the original particle swarm algorithm and random forest algorithm are described below.
[0051] In 1995, Kennedy and Eberhart proposed the particle swarm algorithm, as Figure 1 shown, Figure 1The flowchart of a particle swarm optimization algorithm provided by an embodiment of the present invention. Its basic principle is to let a randomly distributed particle swarm search for the optimal solution under the guidance of individual extreme values and population extreme values. The particle swarm algorithm has strong versatility, and its steps are as follows:
[0052] a. Initialization: In a D-dimensional space, N particles are initialized, and these particles are initialized with the following attributes: the position xi of the i-th particle, the velocity vi of the i-th particle, the best position pbest that the i-th particle has passed through, the best position Gbest that the entire particle swarm has passed through, adding limits xlimiti to the positions of all particles, adding limits vlimiti to the velocities of all particles, setting the number of iterations iter, setting the self-learning factor c1, the population learning factor c2, and the inertia weight w for the particles during each iteration process.
[0053] b. Calculate the fitness of the particle: The fitness function of the particle is f(x) = -(x - 10) 2 + x × sin(x)cos(2x) - 5x × sin(3x). Substituting the current position xi of the i-th particle can obtain the current fitness f(xi) of this particle.
[0054] c. Update the individual extreme value and the global optimal solution: Update the individual best fitness fpbest(xi) of the i-th particle and the best fitness fgbest of the entire population, then update the best position pbesti of the particle according to fpbest(xi), and then find a best position gbest of the population from these best positions pbesti, which is called the global best position of this iteration.
[0055] d. Update the velocity and position of the individual: The update formula is as follows: v i = v i × w + c 1 × r 1 × (p besti - x i ) + c 2 × r 2 × (g best - x i ), x i = x i + β × v i , where r1 and r2 are two independent random numbers generated between [0, 1], xi is the current position of the particle, vi is the velocity of the particle, and β is the constraint factor. If during the iteration process, the position of the i-th particle exceeds the boundary, it is adjusted to Xmin or Xmax.
[0056] e. Set termination conditions: Generally, there are two termination conditions: reaching the number of iterations or the difference between a certain index and the ideal target meeting a certain minimum threshold. If the termination condition is not met, continue to update the fitness of the particles.
[0057] A specific embodiment of the present invention is as Figure 2 shown Figure 2 A mechanical drilling rate prediction method based on particle swarm and random forest provided by the present invention includes:
[0058] S201: Collect the original logging data and extract the characteristic data from the original logging data;
[0059] S202: Optimize the hyperparameters in the random forest model based on the inertia weight parameter adaptive strategy and the particle swarm algorithm introducing the Cauchy mutation operator to obtain an optimized random forest model, and construct an initial mechanical drilling rate prediction model based on the optimized random forest model;
[0060] S203: Train the initial mechanical drilling rate prediction model based on the characteristic data to obtain a well-trained mechanical drilling rate prediction model, and use the well-trained mechanical drilling rate prediction model to output the predicted mechanical drilling rate.
[0061] First of all, it should be noted that the mechanical drilling rate prediction method based on particle swarm and random forest provided by the embodiments of the present invention can be applied to a mechanical drilling rate prediction system based on particle swarm and random forest. The mechanical drilling rate prediction system can be a software system running on a terminal device. The terminal device can be a server, a tablet computer, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a mobile phone and other terminal devices. The specific type of the terminal device is not limited in the embodiments of the present application.
[0062] The present invention analyzes the original logging data, extracts the characteristics affecting the mechanical drilling rate, then constructs a prediction model using a machine learning algorithm, tunes the parameters of the random forest through an improved particle swarm optimization algorithm, accurately predicts the mechanical drilling rate, and further helps oil well technicians optimize the drilling cycle, adjust parameters to increase the drilling rate, realize the optimization of drilling equipment resources, and reduce the drilling cost.
[0063] In an embodiment of the present invention, extracting the characteristic data from the original logging data includes:
[0064] Process the outliers, missing values, and duplicate values in the original logging data to obtain the first logging data;
[0065] Perform linear normalization on the first logging data to obtain the preprocessed logging data;
[0066] Select the features related to the target attribute from the preprocessed logging data to obtain the feature data.
[0067] It can be understood that the original logging data includes drilling depth, weight on bit, penetration rate, and mud properties. By exporting the original logging data from logging equipment or databases, such as data on drilling depth, weight on bit, penetration rate, mud properties, etc., and organizing them into csv or txt files, they are used as the original data for establishing the mechanical penetration rate prediction model.
[0068] Then, preprocess the data, initially screen out the main geological feature parameters in the logging data and the actual drilling parameter features of the core formations in logging, and extract the key factors affecting the mechanical penetration rate through steps such as well location clustering analysis, correlation analysis, and data dimensionality reduction.
[0069] Specifically, in the environment of the Python language, use the pandas library to preprocess the original data and construct an initial sample data set.
[0070] 1) Process outliers: First, analyze whether there are outliers in the data. If there are, use the direct elimination method for processing.
[0071] 2) Process missing values: For row records with more missing attributes that will significantly affect the model accuracy, directly ignore the entire row data record. For other missing situations, the mean method or median can be used to fill in the missing values.
[0072] 3) Process duplicate values: Remove duplicate values and only retain the one data record that can best show the relationship of feature variables to avoid redundancy in model training.
[0073] 4) Data transformation: Perform linear normalization on the data to limit each feature variable within the same range (such as [0, 1] or [-1, 1]) to ensure the same scale for different features.
[0074] 5) Data reduction: Through feature selection, screen out the features that are truly relevant to the target attribute from the original features to optimize the input data set of the model.
[0075] In an embodiment of the present invention, the particle swarm optimization algorithm based on the adaptive strategy of the inertia weight parameter and introducing the Cauchy mutation operator includes:
[0076] Adaptive adjustment of the inertia weight parameter in the particle swarm optimization algorithm;
[0077] Introduce the Cauchy mutation operator to guide the particles in the particle swarm for search.
[0078] It can be understood that the parameters of the standard particle swarm algorithm are fixed. The inertia weight w describes the "inertia" of the particles. It should be larger in the early stage of the algorithm to ensure that each particle flies independently and fully explores the space, and smaller in the later stage to learn more from other particles. In addition, c1 and c2 respectively control the flight step sizes of the particles towards the individual optimal position and the global optimal position. c1 should be larger in the early stage, and c2 should be larger in the later stage, so as to balance the global search ability and the local search ability of the particles. These three parameters jointly affect the flight direction of the particles, resulting in the situation that even if other particles find a better position, but the current particle has too large inertia and still cannot fly quickly to that position. The standard particle swarm optimization algorithm only allows the global optimal particle to provide information, and there is a lack of information sharing among particles. Each update only points to the global optimal solution and is not affected by other particles. Therefore, the particle swarm has problems such as poor local search ability, being easily trapped in local optima, and low search accuracy.
[0079] The improved particle swarm algorithm of the present invention has made the following improvements:
[0080] Adaptive adjustment of the inertia weight parameter in the particle swarm algorithm:
[0081] In practical applications, the particle swarm optimization algorithm is prone to falling into local extreme points. To address this problem, the control idea of adaptive adjustment is introduced to improve the inertia weight. A larger inertia weight leads to global exploration, while a smaller inertia weight leads to local exploration, and the current exploration area can be fine-tuned. At the same time, the current iteration number and the population size during the algorithm update iteration are also considered. The formula for the inertia weight w is as follows:
[0082]
[0083] where w t represents the adjusted inertia weight parameter, w max represents the maximum inertia weight value, w minLet \(w_{min}\) denote the minimum inertia weight value, iter denote the number of iterations, and \(iter_{max}\) denote the maximum number of iterations. The inertia weight parameter \(w\) of the particle swarm decreases as the number of iterations increases, which is beneficial for the particles to move near the target value, thus enabling the optimal solution of the algorithm to gradually converge to the target value. This is to control and balance the transition of the particle swarm algorithm from exploration to exploitation during the entire optimization process. Therefore, the inertia weight parameter \(w\) of the particle swarm is a very important control parameter in the standard particle swarm algorithm and has a significant impact on the performance of the particle swarm algorithm. The inertia weight parameter \(w\) of the particle swarm adopts linear adaptive reduction to control the exploration and exploitation capabilities of the particle swarm algorithm. However, the rapid decrease of the parameter \(w\) in the initial stage of calculation will lead to insufficient exploration in the early stage, resulting in the particle swarm algorithm being unable to search for the global optimum. If the convergence in the later stage of the algorithm is too slow, it means that the algorithm has not been sufficiently exploited in the later stage, resulting in the particle swarm algorithm being unable to accurately search in the local area.
[0084] According to this phenomenon, the inertia weight parameter of the particle swarm w also adopts a method of curve adaptive adjustment, which can effectively solve the problem. Using the adaptive value method, the inertia weight parameter of the particle swarm w is expressed as follows:
[0085]
[0086] The cosine curve adaptive strategy is used to update the inertia weight parameter of the particle swarm instead of the above linear adaptation w method. The selection of the inertia weight parameter of the particle swarm w is related to the search range of the particles: it will affect the global and local search capabilities of the particle swarm algorithm. The larger \(w\) is, the stronger the capabilities are. The difference between curve adaptation and linear adaptation lies in the decreasing trend of the parameter \(w\). Linear adaptation decreases at the same speed, while the cosine adaptive strategy is slow at the beginning and end of the decrease, which helps to enhance the global search ability of the particles at the beginning and the local search ability in the later stage. In addition, the value of the parameter \(w\) is greater than that of linear adaptation in the early stage of cosine adaptation, which can expand the search and enhance the particle detection ability; the later cosine is less than linear adaptation, and linear adaptation can narrow the particle search range and improve the exploration ability of the particle swarm.
[0087] Introduce the Cauchy mutation operator to guide other particles to search:
[0088] It is understandable that at the end stage of the algorithm, it usually enters the local search phase. At this time, if the target value is deviated from in the early stage, it is difficult to find the global optimal solution in local optimization. In the particle swarm algorithm, particles search for better positions by sharing information. Therefore, during the particle iteration process, if there are good particles to guide the search of other particles, the entire particle swarm can move towards a better position, thereby improving the search performance and jumping out of the local optimum. To obtain better particles, mutation operators can be introduced, such as Cauchy mutation, differential mutation, etc. Cauchy mutation can enable particles to search within a larger range, increasing the probability of particles finding better positions, so as to check whether there are better target solutions and ultimately achieve the effect of global optimization. The probability density function of the Cauchy distribution:
[0089]
[0090] x' i = x i + ηδ
[0091] where t is a proportionality parameter greater than 0. Different from the common Cauchy mutation perturbation strategy, η is the perturbation amplitude parameter, and σ is a random variable conforming to the Cauchy distribution.
[0092] Adopting the Cauchy mutation perturbation strategy with linearly decreasing proportionality parameter, the main idea is to use the difference between the maximum and minimum values of the particle positions as the mutation perturbation amplitude, and then perform mutation operations on the particle positions. The better particles after mutation are used as leading particles to lead the rest of the particles to perform iterative search, improving the particle convergence performance and jumping out of the local optimal solution. The Cauchy mutation perturbation formula for the mutated particle positions is:
[0093] p' gd = p gd +(x max (d) - x min (d))·Cauchy(o, s)
[0094] where x max (d) and x min (d) represent the current maximum and minimum values of the particle in dimension d, s is the proportionality parameter, and it changes linearly in the decreasing direction using the following formula.
[0095]
[0096] where tmax is the maximum number of iterations, and the improved Cauchy mutation perturbation strategy can enable the example to have strong search and optimization capabilities.
[0097] Therefore, by dynamically adjusting the inertia weight ω and adding an improved Cauchy mutation operator when calculating the particle position, the inertia weight ω can better achieve global search in the early stage and local search in the later stage. At the same time, it can also increase the diversity of the population, improve the search and optimization ability of particles, avoid particles falling into local optima, and thus improve the accuracy and convergence speed of the algorithm.
[0098] In an embodiment of the present invention, based on the particle swarm optimization algorithm with an adaptive strategy for inertia weight parameters and the introduction of a Cauchy mutation operator, the hyperparameters in the random forest model are optimized to obtain an optimized random forest model, including:
[0099] Determine the solution space range of the hyperparameters of the random forest model;
[0100] Dynamically adjust the inertia weight and learning factor in the particle swarm based on the adaptive strategy for inertia weight parameters, and perform global search on the particle swarm based on the Cauchy mutation operator, so that the particles in the particle swarm perform iterative updates of velocity and position within the solution space range to determine the fitness function values under different hyperparameter combinations;
[0101] Search for the hyperparameter combination with the minimum fitness function value, and use the hyperparameter combination with the minimum fitness function value as the optimal hyperparameter combination of the random forest model to obtain an optimized random forest model.
[0102] First of all, it should be noted that the random forest algorithm is an algorithm that integrates multiple decision trees through the idea of ensemble learning. As Figure 3 shown, Figure 3 is a schematic structural diagram of a random forest model provided by an embodiment of the present invention. Data is randomly sampled from the dataset as the training set of the decision tree, and feature nodes are randomly selected from the feature data to build a decision tree. After repeating the operation, a forest is formed. On this basis, the values obtained from all the trees are selected, and the one selected the most is the final output result. The following are the detailed steps of this algorithm:
[0103] a. Establish a dataset for training individual decision trees: To construct a mechanical drilling rate model based on a random forest, first, it is necessary to construct corresponding base learners (decision trees). Each decision tree corresponds to a subset of feature data, and then a random sampling method without weights and with replacement is used to generate the corresponding subset. This method first puts the original dataset into a black box model, and then randomly samples the data. After each sampling, the data is put back into the black box.
[0104] b. Construct a decision tree: The construction of a decision tree mainly includes two steps: node splitting and randomly selecting feature variables. The node splitting algorithm, also known as the attribute selection algorithm, is the core of generating a decision tree. A regression tree is constructed by gradually splitting nodes, and finally a random forest regression algorithm model is formed. To construct a complete decision tree, appropriate feature parameters need to be selected, and these features are used for node splitting. To improve the generalization ability during the learning process of the random forest model, more different combinations of feature parameters need to be found to enhance the randomness of the data, thereby constructing a model with better performance. The random forest algorithm model uses the Gini index to select feature attributes. For a single decision tree model, each time it will split according to the smallest Gini index, that is, the best feature.
[0105] c. Generate a random forest model and give the result: A large number of different decision trees are generated through the above steps and then combined into a random forest. Finally, the mean value of the mechanical drilling rate predicted by each decision tree is used as the final prediction result.
[0106] The random forest has the following advantages in this prediction method:
[0107] The random forest can handle high-dimensional data without feature selection, adapting to the large amount of high-dimensional data involved in the drilling process; the training speed is very fast, supports parallelization, and can significantly improve the calculation efficiency in the prediction of mechanical drilling rate with a large amount of data; by integrating multiple decision trees, the accuracy and stability of the model are improved, effectively reducing the risks of bias and overfitting; it can capture the complex non-linear relationships in the prediction of mechanical drilling rate, adapt to various influencing factors, such as formation characteristics, drilling parameters, bit types, etc., and provide more accurate prediction results; compared with black-box models such as deep learning, the random forest has higher interpretability, can analyze the specific impacts of each feature on the drilling rate, and provide a scientific basis for parameter adjustment in actual operations. For example, researchers can analyze through the random forest model that the density of the drilling fluid has a greater impact on the mechanical drilling rate, and thus adjust the drilling fluid parameters in actual operations to optimize the drilling rate. Therefore, the present invention can provide reliable and accurate prediction results in the face of complex drilling environments, massive data, and changing drilling parameters, thereby optimizing drilling efficiency, reducing drilling costs and risks.
[0108] It can be understood that in the Python language environment, the scikit-learn library is used to initialize the random forest model and define the initial hyperparameters, such as the number of trees, the maximum depth, etc. Then, the improved particle swarm optimization is used to optimize the parameters of the initial random forest model. In the Python language environment, the improved particle swarm algorithm is implemented through the pyswarm or deap library to optimize the hyperparameters in the random forest model. The improved particle swarm algorithm optimizes the hyperparameters of the random forest through two aspects: the cosine curve adaptive parameter and the Cauchy mutation operator. Specifically, each particle in the particle swarm algorithm represents a combination of random forest hyperparameters, such as the number of trees, the maximum depth of the tree, the minimum number of samples for splitting each tree, etc. Through the search of the particle swarm, the goal is to find a set of hyperparameters that make the prediction performance of the model optimal.
[0109] Among them, the cosine curve adaptive strategy plays a key role in the optimization process. Particle swarm optimization usually needs to balance global search and local search. The cosine curve strategy helps the particle swarm have strong exploration ability in the initial stage of the search by dynamically adjusting the inertia weight and learning factor in the particle swarm, and gradually transitions to refined local optimization in the later stage. The inertia weight is larger in the initial stage of the optimization, which helps the particles widely explore the entire search space. As the iteration progresses, the weight gradually decreases, making the particles more focused on local search when approaching the optimal solution, thus improving the search accuracy. Similarly, the learning factor will also be adjusted according to the change of the cosine curve to ensure that the particles can effectively approach the global optimal solution and maintain a certain exploration ability to avoid falling into the local optimal solution prematurely.
[0110] The Cauchy mutation operator is mainly used to enhance the global search ability of the particle swarm and prevent the particles from falling into the local optimal solution. By randomly perturbing the search positions of the particles, the Cauchy mutation operator enables the particles to jump out of the current local optimal region and explore a wider search space. This mutation operation can provide sufficient diversity especially when the particle swarm is close to the global optimal solution, thus effectively avoiding the problem of too fast convergence and premature termination in the search process and maintaining the exploration ability of the particle swarm.
[0111] Through these two improvement strategies, the particle swarm optimization algorithm can perform efficient searches in a wider hyperparameter space, thereby optimizing the key hyperparameters of the random forest. Finally, through the fine-tuning of these hyperparameters, the random forest model can better adapt to the data characteristics and improve the accuracy and robustness of the mechanical drilling rate prediction.
[0112] For example, the parameters of the initialized particle swarm are shown in the following table:
[0113]
[0114]
[0115] The fitness function is: f(x) = -(x - 10) 2 + x×sin(x)cos(2x) - 5x×sin(3x). Substituting the current position xi of the i-th particle into it can obtain the current fitness f(xi) of the particle; determining the solution space range of the hyperparameters of the random forest model; setting particles within the solution space range through the improved particle swarm algorithm to perform iterative updates of velocity and position. Update the particle velocity. First, perform velocity boundary detection. Generally, v(v > vmax) = vmax, and the same applies to the position. For each particle, compare the fitness value of its current position with the fitness value corresponding to its historical best position Pbest. If the fitness value of the current position is higher, then update the historical best position with the current position, otherwise do not update. Among them, the update formula is:
[0116] v i = v i × w + c 1 × rand() × (pbes i t - x i ) + c 2 × rand() × (gbest - x i ), x i = x i + v i
[0117] In the formula, rand() is a random number function that generates a random number between [0, 1]. If during the iteration process, the position of the i-th particle exceeds the boundary, it is adjusted to Xmin or Xmax; then update the velocity and position of the particle, and continuously perform iterative search until the stop condition is reached, such as iterating a certain number of times. Then determine the fitness function values under different hyperparameter combinations, search for the hyperparameter combination with the minimum fitness function value, and use the hyperparameter combination with the minimum fitness function value as the optimal hyperparameter combination of the random forest model to obtain the optimized random forest model.
[0118] In addition, during the process of constructing the initial mechanical drilling rate prediction model based on the optimized random forest model, the random forest method is used for modeling. In the Python language environment, the scikit-learn library is used to initialize the random forest model and define the initial hyperparameters. The random forest improves the prediction ability of the model through the ensemble learning method. When generating each decision tree, first, a subset is randomly sampled from the original training data by the Bootstrap sampling method (usually sampling with replacement). Each tree is trained on a different subset, which can increase the diversity between each tree and thus reduce the overfitting phenomenon. During the construction of each tree, by continuously splitting the sample data and selecting a feature as the decision node until the set stopping conditions are reached, such as the maximum depth of the tree, the minimum number of samples, or complete fitting of the data. At the same time, when splitting each node of the decision tree, the random forest does not use all features, but randomly selects a subset from the feature set. This process is achieved by restricting the number of features that each node can select. This can reduce the similarity between each tree, increase the diversity of the model, and improve the stability and generalization ability of the model. After the training of each decision tree is completed, the random forest model makes predictions through the ensemble learning method. When predicting a new sample, the random forest will take the average of the prediction results of all trees. Each tree votes on a prediction result, and the final prediction result is determined by the comprehensive prediction results of all trees. Through this method, the random forest can effectively reduce the overfitting or bias problems that may occur in a single tree and provide more robust and accurate predictions.
[0119] In one embodiment of the present invention, the initial mechanical drilling rate prediction model is trained based on the feature data to obtain a fully trained mechanical drilling rate prediction model, including:
[0120] The features in the feature data that affect the mechanical drilling rate of the oil and gas well are used as the input feature vector, and the mechanical drilling rate in the feature data is used as the output feature variable to train the initial mechanical drilling rate prediction model, obtaining a fully trained mechanical drilling rate prediction model.
[0121] It can be understood that during the training of the initial mechanical drilling rate prediction model based on feature data, the training set and the test set are first divided. Specifically, according to the scale of the feature data and a preset standard, the feature data is randomly divided into a training set and a test set. Among them, the ratio of the number of training data to the number of test data is 9:1. Then, a feature vector and a target vector are constructed. The attributes in the logging data that affect the mechanical drilling rate of the oil and gas well are used as the input feature vector, the data of each feature attribute is used as the input variable x, and the mechanical drilling rate is used as the output variable y for the training of the prediction model. Then, in combination with the characteristics of the random forest model, each decision tree corresponds to a subset of feature data, and a sampling method of random sampling without weights and with replacement is used to generate the corresponding subset. The random forest model uses the Gini index to select feature attributes. For a single decision tree model, each time it will split according to the smallest Gini index, that is, the best feature. By generating a large number of different decision trees and then combining them into a random forest, finally, the average value of the mechanical drilling rates predicted by each decision tree is calculated as the final prediction result.
[0122] In an embodiment of the present invention, the above method further includes:
[0123] Determine the predicted value and the actual value of the mechanical drilling rate;
[0124] Based on the Pearson correlation coefficient, analyze the predicted value and the actual value to determine the correlation between the predicted value and the actual value.
[0125] It can be understood that the present invention uses the Pearson correlation coefficient for analysis and evaluation. If there is a high correlation between the predicted value of the mechanical drilling rate and the actual value, and the error between the two is small, it indicates that the random forest model optimized by the improved particle swarm optimization can be effectively used for the classification prediction of the mechanical drilling rate.
[0126] The mechanical drilling rate prediction method based on particle swarm and random forest provided by the present invention can optimize drilling parameters such as weight on bit, drilling rate, and torque by accurately predicting the mechanical drilling rate, so as to maximize the drilling speed, thereby shortening the drilling time and improving the drilling efficiency; accurately predicting the mechanical drilling rate helps to reasonably allocate resources and optimize the use of drilling tools, reduce the wear and replacement frequency of drill bits, thereby reducing the maintenance cost of drilling equipment and the usage amount of drilling fluid, and finally reducing the overall drilling cost; and it can identify potential drilling risks such as lost circulation and blowout during the prediction process. Anticipating risks in advance can take corresponding measures to avoid accidents and improve the safety of drilling operations; finally, compared with the traditional mechanical drilling rate prediction method, the mechanical drilling rate prediction method of the present invention has higher prediction accuracy in dealing with multi-variable and non-linear data, making the prediction result closer to the actual drilling situation. The above technical effects can help drilling engineers make more scientific and reasonable decisions during operations, thus achieving a win-win situation in improving economic benefits and ensuring safety.
[0127] In order to better implement the mechanical drilling rate prediction method based on particle swarm and random forest in the embodiments of the present invention, correspondingly, based on the mechanical drilling rate prediction method based on particle swarm and random forest, as Figure 4 shown, the embodiments of the present invention also provide a mechanical drilling rate prediction system based on particle swarm and random forest. The mechanical drilling rate prediction system 400 based on particle swarm and random forest includes:
[0128] A data preprocessing module 401, configured to collect original logging data, preprocess the original logging data to obtain preprocessed logging data, and extract feature data from the logging data;
[0129] A model optimization module 402, configured to optimize the hyperparameters in the random forest model based on the inertia weight parameter adaptive strategy and the particle swarm algorithm introducing the Cauchy mutation operator to obtain an optimized random forest model, and construct an initial mechanical drilling rate prediction model based on the optimized random forest model;
[0130] A mechanical drilling rate prediction module 403, configured to train the initial mechanical drilling rate prediction model based on the feature data to obtain a fully trained mechanical drilling rate prediction model, and use the fully trained mechanical drilling rate prediction model to output the predicted mechanical drilling rate.
[0131] The mechanical drilling rate prediction system 400 based on particle swarm and random forest provided in the above embodiments can implement the technical solutions described in the embodiments of the mechanical drilling rate prediction method based on particle swarm and random forest. For the specific implementation principles of the above modules or units, reference can be made to the corresponding content in the embodiments of the mechanical drilling rate prediction method based on particle swarm and random forest, which will not be elaborated here.
[0132] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, a random access memory, etc.
[0133] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A mechanical drilling speed prediction method based on particle swarm and random forest, characterized in that: include: Collecting original logging data and extracting characteristic data from the original logging data; Based on the inertia weight parameter adaptive strategy and the particle swarm algorithm with the Cauchy mutation operator, the hyperparameters in the random forest model are optimized to obtain the optimized random forest model, and the initial mechanical drilling speed prediction model is constructed based on the optimized random forest model. The initial mechanical drilling speed prediction model is trained based on the characteristic data to obtain a fully trained mechanical drilling speed prediction model, and the predicted mechanical drilling speed is output using the fully trained mechanical drilling speed prediction model.
2. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 1, characterized in that: The original logging data includes drilling depth, drilling pressure, drilling speed and mud properties.
3. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 1, characterized in that: Extracting characteristic data from the original logging data includes: Processing abnormal values, missing values and duplicate values in the original logging data to obtain first logging data; Performing linear normalization processing on the first logging data to obtain the pre-processed logging data; Features related to target attributes are selected from the preprocessed logging data to obtain the feature data.
4. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 1, characterized in that: The particle swarm algorithm based on the inertia weight parameter adaptive strategy and the introduction of the Cauchy mutation operator includes: Adaptively adjust the inertia weight parameters in the particle swarm algorithm; The Cauchy mutation operator is introduced to guide the particles in the particle swarm to search.
5. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 4, characterized in that: The adaptive adjustment of the inertia weight parameter in the particle swarm algorithm includes: Perform linear adaptive adjustment on the inertia weight parameter: Or make curve adaptive adjustment to the inertia weight parameter: Among them, w t represents the adjusted inertia weight parameter, w max Indicates the maximum inertia weight value, w min Represents the minimum inertia weight value, iter represents the number of iterations, and iter_max represents the maximum number of iterations.
6. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 4, characterized in that: The Cauchy mutation perturbation formula of the Cauchy mutation operator includes: p' gd =p gd +(x max (d)-x min (d))·Cauchy(o,s) Among them, p' gd represents the mutated Cauchy operator; p gd represents the original Cauchy operator, x max (d) represents the maximum value of the particle in d dimension, x min (d) represents the minimum value of the particle in d dimension, and s represents the scale parameter.
7. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 4, characterized in that: The particle swarm algorithm based on the inertia weight parameter adaptive strategy and the introduction of the Cauchy mutation operator optimizes the hyperparameters in the random forest model to obtain an optimized random forest model, including: Determine the solution space range of the hyperparameters of the random forest model; The inertia weight and learning factor in the particle swarm are dynamically adjusted based on the inertia weight parameter adaptive strategy, and the particle swarm is globally searched based on the Cauchy mutation operator, so that the particles in the particle swarm are iteratively updated in the speed and position within the solution space, and the fitness function value under different hyperparameter combinations is determined; The hyperparameter combination with the smallest fitness function value is searched, and the hyperparameter combination with the smallest fitness function value is used as the optimal hyperparameter combination of the random forest model to obtain an optimized random forest model.
8. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 1, characterized in that: The initial mechanical drilling rate prediction model is trained based on the characteristic data to obtain a fully trained mechanical drilling rate prediction model, including: The features in the feature data that affect the mechanical drilling speed of oil and gas wells are used as input feature vectors and the mechanical drilling speed in the feature data is used as output feature variables to train the initial mechanical drilling speed prediction model to obtain a fully trained mechanical drilling speed prediction model.
9. The method for predicting mechanical drilling speed based on particle swarm and random forest according to claim 1, characterized in that: The method further comprises: Determine predicted and actual ROP values; The predicted value and the actual value are analyzed based on the Pearson correlation coefficient to determine the correlation between the predicted value and the actual value.
10. A mechanical drilling speed prediction system based on particle swarm and random forest, characterized in that: include: A data preprocessing module, used for collecting raw logging data and extracting characteristic data from the raw logging data; The model optimization module is used to optimize the hyperparameters in the random forest model based on the inertia weight parameter adaptive strategy and the particle swarm algorithm with the Cauchy mutation operator, obtain the optimized random forest model, and build the initial mechanical drilling speed prediction model based on the optimized random forest model; The mechanical drilling speed prediction module is used to train the initial mechanical drilling speed prediction model based on the characteristic data to obtain a fully trained mechanical drilling speed prediction model, and output a predicted mechanical drilling speed using the fully trained mechanical drilling speed prediction model.
Citation Information
Cited By
Intelligent selecting and transmitting device for oil and gas well and control method
CN121279892A