A real-time lithology identification method based on drilling data

CN122839079APending Publication Date: 2026-09-29HUNAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611219974.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-12
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]本发明的目的是提出一种基于钻井数据的实时岩性识别方法,该方法通过提取钻井信号的短时能量特征、构建分层特征、引入数据驱动正则化机制,解决了现有技术中岩性特征微弱、特征耦合性强以及非线性关系复杂的问题,显著提高了岩性识别的准确性和鲁棒性,为复杂地质钻进过程的实时监控和智能决策提供了有力支持

Benefits of technology

[0021]本发明通过提取钻井信号的短时能量特征,能够有效增强微弱岩性信号、解耦强相关特征、处理复杂非线性关系,从而获得较高的岩性识别精度和较强的跨井泛化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839079A_ABST
    Figure CN122839079A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of intelligent control, specifically relating to a real-time lithology identification method based on drilling data. The method includes: data acquisition and denoising; enriching the lithological features in the original data to enhance sensitivity to lithological changes, including using an adaptive sliding window to extract short-time energy features as intermediate parameters; modeling: constructing a stacked autoencoder network, i.e., using an improved SAE algorithm to establish a lithology identification model for hierarchical feature extraction of the input data; in the layer-by-layer pre-training stage of the stacked autoencoder, the maximum information coefficient (MIC) between each input feature and the lithology label is introduced as a regularization term into the loss function to guide the network in learning features highly correlated with lithology, forming an MR-SAE model; fine-tuning and application of the model. This invention has been validated on a multi-well real-world drilling dataset, demonstrating excellent performance not only in single-well identification tasks but also strong generalization ability and robustness in cross-well prediction tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control of complex geological drilling processes, and specifically relates to a real-time lithology identification method based on drilling data. Background Technology

[0002] Real-time geological condition sensing is fundamental to building intelligent drilling systems, as formation changes directly affect drilling efficiency, operational safety, and process stability. Lithology identification refers to the process of classifying lithological units based on rock characteristics, which is of great significance for improving oil and gas exploration efficiency and reducing drilling risks.

[0003] Traditional methods rely heavily on statistical analysis of well logging data. However, well logging data is costly, time-consuming, and carries a high risk of tool failure in complex geological environments. Many drilling sites are even lacking well logging tools. To achieve cost-effective and efficient real-time lithology identification, researchers have begun to use measurement-while-drilling (MWD) parameters as an alternative data source.

[0004] Drilling-while-drilling (DSD) data is time-series data continuously acquired during the drilling process through surface and downhole sensing systems, characterized by high real-time performance and continuity. These parameters characterize the mechanical response of the drilling process to formation features, providing indirect evidence for related research.

[0005] In recent years, machine learning techniques have been widely applied to lithology identification based on drilling data. Existing methods include artificial neural networks, long short-term memory networks, Kalman filtering combined with the K-nearest neighbor algorithm, and models combining wavelet transform and attention mechanisms. Some studies introduce intermediate parameters (such as mechanical specific energy) to improve the interpretability of the model.

[0006] However, reliable lithology identification based on drilling data still faces challenges. First, strong noise in the drilling data may mask signal fluctuations caused by lithological changes, leading to blurred formation variation characteristics. Second, the strong coupling and complex nonlinear relationships among drilling parameters limit further improvements in classification accuracy. Therefore, how to effectively extract and construct robust feature representations to connect drilling parameters with lithology, enhance data reliability, and suppress noise interference is a pressing technical problem to be solved in this field. Summary of the Invention

[0007] The purpose of this invention is to propose a real-time lithology identification method based on drilling data. This method solves the problems of weak lithology features, strong feature coupling, and complex nonlinear relationships in the prior art by extracting short-time energy features of drilling signals, constructing layered features, and introducing a data-driven regularization mechanism. It significantly improves the accuracy and robustness of lithology identification and provides strong support for real-time monitoring and intelligent decision-making in complex geological drilling processes.

[0008] This invention provides a real-time lithology identification method based on drilling data. The method includes the following steps: Step A, data acquisition and denoising: real-time drilling data during the drilling process is acquired and denoised using wavelet filtering to construct a reliable clean dataset; the real-time drilling data is the raw data, which includes at least drilling depth, drilling speed, drilling pressure, rotational speed, and torque; Step B, enriching the lithological features in the raw data to enhance sensitivity to lithological changes: based on the denoised real-time drilling data, appropriate drilling parameters are selected and their short-time energy features are extracted as intermediate parameters using an adaptive sliding window; the appropriate drilling parameters include drilling speed and torque; the adaptive sliding window is the calculated optimal sliding window; Step C, modeling: constructing a stacked autoencoder network. A network is used to perform hierarchical feature extraction on the input data, which includes the original data from the clean dataset in step A and the intermediate parameters from step B. In the layer-by-layer pre-training stage of the stacked autoencoder, the maximum information coefficient (MIC) between each input feature and the lithology label is introduced into the loss function as a regularization term to guide the network to learn features that are highly correlated with lithology, forming an MIC-regularized stacked autoencoder model based on energy feature extraction, named the MR-SAE model. Step D: Fine-tuning and application of the model: The pre-trained MR-SAE model is fine-tuned in a supervised manner using a weighted cross-entropy loss function to complete the establishment of the real-time lithology identification model. The obtained real-time lithology identification model is then used to perform real-time lithology identification on the input drilling data.

[0009] In this invention, the lithologies identified include, for example, granite, dolomite, diorite, limestone, siltstone, and shale.

[0010] In one specific implementation, in step A, the real-time drilling data further includes hook load, stand pressure, flow rate, pump pressure, and traveling block speed. When using wavelet filtering for noise reduction, the dmey function is selected as the wavelet basis function, and two-level wavelet decomposition is used to suppress high-frequency noise and retain the overall trend of the data.

[0011] In one specific implementation, step B includes the following steps: Step B1, for the drilling speed or torque signal, its short-time energy characteristic E n The calculation formula is:

[0012] Among them, E n This represents the short-time energy characteristic corresponding to the nth sampling position, where k is the index of the signal sampling point within the sliding window, and x 2(k) represents the square of the amplitude of the corresponding sampling point, N represents the sliding window length, i.e., the number of sampling points included in each short-time energy calculation, and n represents the current position of the short-time energy calculation; the short-time energy sequence calculated by this formula is used to effectively characterize the energy characteristics of drilling data; Step B2: Define a candidate window set. ={50,60,70,80,90,100}; Step B3: For each candidate window size... Calculate the corresponding short-time energy characteristic E n And calculate the maximum information coefficient ρ between the sequence and the lithological tag Y. w =MIC(E n (Y), select the window that maximizes the correlation, thereby achieving adaptive extraction of energy features.

[0013] In one specific implementation, step C includes the following steps: Step C1, calculate the maximum information coefficient MIC(X,Y) between each input feature X, including the original data and intermediate parameters, and the lithological label Y.

[0014] Where a and b represent the number of grids in the horizontal and vertical directions, respectively; B controls the size of the maximum grid division, set to the power of 0.6 of the data volume; I(X;Y) measures the amount of information shared between variables X and Y; the maximum information coefficient MIC(X,Y) ranges from (0,1], where 0 indicates that the variables are independent and 1 indicates that they are perfectly correlated.

[0015] Step C2: Incorporate the calculated maximum information coefficient value into the loss function J(W) of the stacked autoencoder pre-training stage;

[0016] Where N is the number of input samples, i is the index of the i-th sample, and x i It is the i-th original input sample. It is the output after reconstruction by the autoencoder, α is the regularization coefficient, and N x N represents the number of neurons in the input layer. h MIC(x) represents the number of hidden layer neurons, n is the nth input feature, m is the nth hidden layer neuron, and MIC(x) represents the nth hidden layer neuron. n W(y) is the maximum information coefficient between the input features and the lithological label. nm Input the connection weights between feature n and hidden neuron m into the weight parameters;

[0017] Step C3: The penalty term uses L2 regularization. Taking the weight update of the encoding part as an example, its gradient descent function becomes:

[0018] in, This represents the weights from the nth input feature to the mth hidden neuron at the t-th iteration. These are the updated weights, where η is the learning rate, and ▽ W J(W) is the partial derivative of the loss function with respect to the weight parameters. From the above equation, it can be seen that when MIC(x) n When y increases, The weight decay of the nth drilling parameter, which includes both raw data and intermediate parameters, decreases as the iteration is completed. The network tends to retain larger weights, thus enabling the adaptive decay mechanism to effectively guide the network to prioritize features rich in lithological information.

[0019] In one specific implementation, step D includes: supervised fine-tuning using a weighted cross-entropy loss function to address the imbalance in the number of samples from different lithological categories; the weighted cross-entropy loss function L... W The expression is:

[0020] Among them, w c y represents the weight of class c. c The actual label representing the category, p c is the predicted probability of the c-th class output by the softmax layer of the model, and K is the total number of lithology classes.

[0021] This invention extracts the short-time energy characteristics of drilling signals, which can effectively enhance weak lithology signals, decouple strong correlation features, and handle complex nonlinear relationships, thereby achieving high lithology identification accuracy and strong cross-well generalization ability.

[0022] This invention addresses the problems of low identification accuracy and poor generalization ability caused by weak lithological features, strong feature coupling, and strong nonlinear relationships in drilling data. The beneficial effects of this invention include at least the following: 1) This invention provides a real-time lithology identification method based on drilling data. First, wavelet filtering is used to denoise the original drilling data, effectively suppressing high-frequency noise and laying a good foundation for subsequent modeling work; 2) This invention extracts short-time energy features from drilling data and uses a maximum information coefficient adaptive optimization sliding window, effectively amplifying local signal fluctuations caused by lithological changes, significantly enhancing the sensitivity of conventional drilling parameters to lithological changes, and achieving rapid response; 3) This invention introduces a maximum information coefficient-based method during the pre-training stage of the stacked autoencoder. The regularization term enables the model to prioritize learning features with strong nonlinear correlation to lithology labels (including raw data and short-time energy features), effectively reducing feature coupling and suppressing noise interference; 4) This invention integrates energy feature extraction with a data-driven regularization mechanism to form a unified intelligent lithology identification architecture with strong model interpretability and high identification accuracy; 5) The method described in this invention has been validated on multiple real drilling datasets, demonstrating excellent performance not only in single-well identification tasks but also in cross-well prediction tasks, showcasing strong generalization ability and robustness, which is beneficial for the application of this invention in actual production. Attached Figure Description

[0023] Figure 1 The figures show the original data and the filtered results, where (a), (b), (c), and (d) correspond to the conditions of wells a, b, c, and d, respectively.

[0024] Figure 2 The sliding window plot is used to extract energy features, where (a), (b), (c), and (d) correspond to the conditions of wells a, b, c, and d, respectively.

[0025] Figure 3 Box plots for different algorithm models are shown, where (a) and (b) correspond to the cases of wells a and b, respectively.

[0026] Figure 4 The image shows a comparison of model accuracy before and after adding short-term energy, where (a) and (b) correspond to the cases of wells a and b, respectively.

[0027] Figure 5 A visualization of the lithology identification results for well d, trained based on data from well c.

[0028] Figure 6 A visualization of the lithology identification results for well c, based on data training from well d. Detailed Implementation

[0029] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Example

[0031] A real-time lithology identification method based on drilling data is proposed. First, the real drilling datasets from four different regions (wells a, b, c, and d) used in this embodiment are denoised. Wells a and b are located in different regions, while wells c and d are both located in Gedian Town, Ezhou City, Hubei Province. Next, the drilling data is analyzed, and data closely related to lithology are selected. An adaptive sliding window is used to extract energy features (short-time energy) from drilling parameters such as drilling rate and torque as intermediate parameters. Finally, the selected drilling datasets and the extracted energy features are used as inputs. Based on the obtained sample datasets, normalization processing is performed. An improved maximum information coefficient (MIC) regularized stacked autoencoder model based on energy feature extraction is used to establish a lithology identification model, and simulation verification is performed using actual production data. The specific steps are as follows.

[0032] Table 1 shows some of the original drilling data. Drilling data from four wells were collected, mainly including drilling speed and torque. Wavelet filtering was used to denoise the original data. The dmey function was selected as the wavelet basis function, and two-level wavelet decomposition was applied to obtain the denoised drilling speed and torque data. The denoising effect is shown in the table below. Figure 1 As shown. Figure 1 The image shows the original data and the filtered results. Figure 1 In the figure, (a), (b), (c), and (d) correspond to the conditions of wells a, b, c, and d, respectively. In the figure, Original represents the original data line (red line), and Wavelet filtering represents the filtered data line (blue line). Figure 1 It is evident that the filtering method effectively suppresses high-frequency noise while preserving the overall trend of the original data.

[0033] Table 1

[0034] Extracting short-time energy features: Energy features (short-time energy) are calculated for the denoised drilling speed and torque data. First, a candidate window set W = {50, 60, 70, 80, 90, 100} is defined. For each candidate window size, the corresponding short-time energy sequence is calculated, and the maximum information coefficient between this sequence and the lithology label is calculated. The window that maximizes the maximum information coefficient is selected as the optimal sliding window, thus achieving adaptive extraction of energy features. The sliding window optimization results for four wells are shown below. Figure 2 As shown, Figure 2 A sliding window plot for extracting energy features. Figure 2 (a), (b), (c), and (d) correspond to the situations of wells a, b, c, and d, respectively. Figure 2For well a, the optimal short-term energy for drilling speed is at a window size of 100, and the optimal short-term energy for torque is also at a window size of 100. For well b, the optimal short-term energy for drilling speed is at a window size of 70, and the optimal short-term energy for torque is at a window size of 60. For well c, the optimal short-term energy for drilling speed is at a window size of 100, and the optimal short-term energy for torque is at a window size of 50. For well d, the optimal short-term energy for drilling speed is at a window size of 80, and the optimal short-term energy for torque is at a window size of 90. Figure 2 The results show that the optimal window size varies for different wells and parameters, proving the necessity of the adaptive extraction strategy. Energy characteristics are recalculated based on the optimal window and used as input along with the drilling data obtained in step A.

[0035] Based on the data obtained from the above steps, the first 80% was used as the dataset and the last 20% as the test set. A MIC regularization-driven stacked autoencoder model based on energy feature extraction was constructed. The model hyperparameters were optimized using the Optuna framework, and the optimal parameters are shown in Table 2. Table 2 shows the hyperparameter tuning results of the MR-SAE algorithm provided in this invention.

[0036] Table 2

[0037] Based on the model established in the previous step, it is compared with Random Forest (RF), K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Grid Search-Artificial Neural Network (GS-ANN), Wavelet Transform Self-Cross-Attention Model (SACWT), the original SAE model (SAE, Stacked Autoencoder), and the SAE model (N1) that only uses short-time energy without using maximum information coefficient regularization. The comparison results are as follows: Figure 3 , Figure 4 As shown in Table 3. Figure 3 Box plots for different algorithm models. Figure 3 (a) and (b) correspond to the situations of wells a and b, respectively. Figure 4 A comparison chart showing the model accuracy before and after adding short-time energy. Figure 4 (a) and (b) correspond to the situations of wells a and b, respectively. Figure 3 and Figure 4 The "Proposed" entries refer to data corresponding to the method of this invention. Table 3 shows the performance comparison results of different lithology identification models for wells a and b. In Table 3, accuracy, precision, recall, and F1 score are used to express the excellence of the model. Higher values ​​for these four indicators are better. Based on the evaluation of these four indicators, the method described in this invention is the optimal one. Figure 3 , Figure 4As shown in Table 3, the method proposed in this invention outperforms other methods in both prediction accuracy and generalization ability by effectively extracting the energy characteristics of drilling data and integrating MIC regularization.

[0038] Table 3

[0039] Based on the model established through the above steps, due to variations in formation thickness and highly uneven distribution of samples from different lithological types, the Matthews correlation coefficient was used for performance evaluation. Table 4 shows the results of different methods on this index. Table 4 shows the Matthews correlation coefficients of well a and well b using different methods. As can be seen from Table 4, the MR-SAE algorithm proposed in this invention, based on energy feature extraction, outperforms all other methods on this performance index.

[0040] Table 4

[0041] Based on the model established in the above steps, the generalization ability and robustness of the model under different geological conditions were further evaluated. Cross-well prediction experiments were conducted, and two cross-well prediction scenarios were designed. Scenario 1 used data from well c for training and data from well d for testing. Scenario 2 used data from well d for training and data from well c for testing. The results were recorded in [the relevant documentation / documentation]. Figure 5 , Figure 6 And in Table 5. Figure 5 A visualization of the lithology identification results for well d, trained based on data from well c. Figure 6 A visualization of the lithology identification results for well c, based on data training from well d. Figure 5 and Figure 6 The "Proposed" entries refer to data from the corresponding methods of this invention. Table 5 shows the performance comparison results of the cross-well model verification.

[0042] Figure 5 , Figure 6 The experimental results corresponding to Table 5 show that the method of the present invention enhances the model's adaptability to different drilling data distributions through energy feature extraction, and still maintains a high recognition accuracy in cross-well tasks.

[0043] Table 5

[0044] In summary, this invention relates to a real-time lithology identification method based on drilling data, comprising: data acquisition and denoising; enriching the lithological features in the original data to enhance sensitivity to lithological changes, including using an adaptive sliding window to extract short-time energy features as intermediate parameters; modeling: constructing a stacked autoencoder network, i.e., using an improved SAE algorithm to establish a lithology identification model for hierarchical feature extraction of input data; in the layer-by-layer pre-training stage of the stacked autoencoder, the maximum information coefficient (MIC) between each input feature and the lithology label is introduced as a regularization term into the loss function to guide the network to learn features highly correlated with lithology, forming an MR-SAE model; fine-tuning the model and applying it. This invention, by combining drilling data energy feature extraction with hierarchical feature enhancement and data-driven regularization, effectively suppresses noise interference and decouples feature coupling, significantly improving the accuracy of lithology identification and cross-well generalization ability, laying a solid foundation for intelligent control of complex geological drilling processes. This invention has been validated on multiple real drilling datasets, demonstrating excellent performance not only in single-well identification tasks but also in cross-well prediction tasks, exhibiting strong generalization ability and robustness.

[0045] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A real-time lithology identification method based on drilling data, characterized in that, The method includes the following steps. Step A: Data Acquisition and Denoising: Real-time drilling data during the drilling process is acquired and denoised using wavelet filtering to construct a reliable and clean dataset; the real-time drilling data refers to the raw data, which includes at least drilling depth, drilling speed, drilling pressure, rotational speed, and torque; Step B: Enrich the lithological features in the original data and enhance the sensitivity to lithological changes: Based on the denoised real-time drilling data, select appropriate drilling parameters and use an adaptive sliding window to extract their short-time energy characteristics as intermediate parameters; the appropriate drilling parameters include drilling speed and torque; the adaptive sliding window is the calculated optimal sliding window; Step C, Modeling: Construct a stacked autoencoder network to perform hierarchical feature extraction on the input data, which includes the original data from the clean dataset in Step A and the intermediate parameters from Step B. In the layer-by-layer pre-training stage of the stacked autoencoder, the maximum information coefficient (MIC) between each input feature and the lithology label is introduced into the loss function as a regularization term to guide the network to learn features that are highly correlated with lithology, forming an MIC-regularized stacked autoencoder model based on energy feature extraction, named the MR-SAE model. Step D: Fine-tuning and application of the model: The pre-trained MR-SAE model is fine-tuned in a supervised manner using the weighted cross-entropy loss function to complete the establishment of the real-time lithology identification model; and the obtained real-time lithology identification model is used to perform real-time lithology identification on the input drilling data.

2. The real-time lithology identification method based on drilling data according to claim 1, characterized in that, In step A, the real-time drilling data also includes hook load, stand pressure, flow rate, pump pressure and traveling block speed. When using wavelet filtering to denoise, the dmey function is selected as the wavelet basis function, and two-level wavelet decomposition is used to suppress high-frequency noise and retain the overall trend of the data.

3. The real-time lithology identification method based on drilling data according to claim 1, characterized in that, Step B includes the following steps: Step B1: For drilling speed or torque signals, their short-time energy characteristics E n The calculation formula is: Among them, E n This represents the short-time energy characteristic corresponding to the nth sampling position, where k is the index of the signal sampling point within the sliding window, and x 2 (k) represents the square of the amplitude of the corresponding sampling point, N represents the sliding window length, i.e. the number of sampling points included in each short-time energy calculation, and n represents the current position of the short-time energy calculation; the short-time energy sequence calculated by this formula is used to effectively characterize the energy characteristics of drilling data; Step B2: Define a candidate window set W = {50, 60, 70, 80, 90, 100}; Step B3: For each candidate window size, Calculate the corresponding short-time energy characteristic E n And calculate the maximum information coefficient ρ between the sequence and the lithological tag Y. w =MIC(E n (Y), select the window that maximizes the correlation, thereby achieving adaptive extraction of energy features.

4. The real-time lithology identification method based on drilling data according to claim 3, characterized in that, Step C includes the following steps: Step C1: Calculate the maximum information coefficient MIC(X,Y) between each input feature X, including the original data and intermediate parameters, and the lithology label Y. Where a and b represent the number of grids in the horizontal and vertical directions, respectively; B controls the size of the maximum grid division, set to the power of 0.6 of the data volume; I(X;Y) measures the amount of information shared between variables X and Y; the maximum information coefficient MIC(X,Y) ranges from (0,1], where 0 indicates that the variables are independent and 1 indicates that they are perfectly correlated. Step C2: Incorporate the calculated maximum information coefficient value into the loss function J(W) of the stacked autoencoder pre-training stage; Where N is the number of input samples, i is the index of the i-th sample, and x i It is the i-th original input sample. This is the output after reconstruction by the autoencoder, where α is the regularization coefficient and N is the value of N. x N represents the number of neurons in the input layer. h MIC(x) represents the number of hidden layer neurons, n is the nth input feature, m is the nth hidden layer neuron, and MIC(x) represents the nth hidden layer neuron. n W(y) is the maximum information coefficient between the input features and the lithological label. nm Input the connection weights between feature n and hidden neuron m into the weight parameters; Step C3: The penalty term uses L2 regularization. Taking the weight update of the encoding part as an example, its gradient descent function becomes: in, This represents the weights from the nth input feature to the mth hidden neuron at the t-th iteration. These are the updated weights, where η is the learning rate, and ▽ W J(W) is the partial derivative of the loss function with respect to the weight parameters. From the above equation, it can be seen that when MIC(x) n When y increases, The weight decay of the nth drilling parameter, which includes both raw data and intermediate parameters, decreases as the iteration is completed. The network tends to retain larger weights, thus enabling the adaptive decay mechanism to effectively guide the network to prioritize features rich in lithological information.

5. The real-time lithology identification method based on drilling data according to claim 4, characterized in that, Step D includes: supervised fine-tuning using a weighted cross-entropy loss function to address the imbalance in the number of samples from different lithological categories; the weighted cross-entropy loss function L... W The expression is: Among them, w c y represents the weight of class c. c The actual label representing the category, p c is the predicted probability of the c-th class output by the softmax layer of the model, and K is the total number of lithology classes.