A lithology prediction method based on piecewise weighted spearman algorithm and gradient ensemble learning
By using a piecewise weighted Spearman algorithm to screen key logging curves and integrating a gradient algorithm to construct a lithology prediction model, the problem of insufficient accuracy and generalization ability of traditional models in lithology identification is solved, and higher accuracy lithology prediction and reservoir evaluation are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEAST GASOLINEEUM UNIV
- Filing Date
- 2026-03-12
- Publication Date
- 2026-06-16
Smart Images

Figure CN122220833A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas field exploration well technology, and in particular to a lithology prediction method based on the piecewise weighted Spearman algorithm and gradient ensemble learning. Background Technology
[0002] Currently, in the field of well logging interpretation, the accuracy of lithology identification is severely limited by problems such as inconsistent quality of well logging curves, feature redundancy, and significant data gaps. The complex correlation between various well logging curves and lithology makes it difficult for traditional single-prediction models to establish universal mapping relationships, resulting in limited lithology prediction accuracy and insufficient model generalization ability. Existing methods cannot effectively integrate multi-source heterogeneous well logging information, restricting the accuracy and efficiency of reservoir evaluation and oil and gas discovery. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a lithology prediction method based on the piecewise weighted Spearman algorithm and gradient ensemble learning.
[0004] The objective of this invention is achieved through the following technical solution: a lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning, comprising the following steps:
[0005] S1: Collect well logging curve data and perform preprocessing;
[0006] S2: Screening well logging curves based on piecewise weighted Spearman algorithm;
[0007] S3: Construct a lithology prediction model;
[0008] S4: Evaluate and optimize the lithology prediction model;
[0009] S5: Predict the lithology of the entire well section.
[0010] Preferably, in step S1, the various curves are subjected to a unified dimension processing:
[0011] ;
[0012] in, For standardized well logging curve data, For the initial curve data, For curves Median sorted from smallest to largest For well logging curves The median of the second half of the sorted data from smallest to largest. For well logging curves The median of the first half of the data after sorting from smallest to largest.
[0013] Preferably, step S2 further includes the following step:
[0014] S21: Convert the logging curves and specific lithology values into their respective ranking orders, and calculate the rank difference at each depth point.
[0015] ;
[0016] in, For the first The first logging curve rank difference at each depth point For the first The first logging curve Rank of the logging curve after conversion at each depth point The order of lithology;
[0017] S22: Calculate the weighted sum of squares ,
[0018] ;
[0019] in, For the first The weighted sum of squares of the logging curves. This represents the total number of data points. For the first The first logging curve Signal-to-noise ratio at each depth point For the first The first logging curve Integrity factor at each depth point For the first The first logging curve rank difference at each depth point For the first The first logging curve Comprehensive adaptive weights at each depth point;
[0020] S23: Calculate the piecewise weighted Spearman correlation coefficient.
[0021] ;
[0022] in, For the first Adaptive weighted Spearman correlation coefficient of the well logging curves.
[0023] If the correlation coefficient of the curve is greater than the screening threshold Then it will participate in the training of subsequent prediction models.
[0024] ;
[0025] in, This represents the median of the absolute values of the correlation coefficients of all well logging curves. This represents the difference between the upper and lower quartiles of the absolute values of the correlation coefficients of all well logging curves.
[0026] Preferably, step S3 further includes the following step:
[0027] S31: Construct a conventional gradient boosting decision tree
[0028] ;
[0029] in, From decision tree ensemble Find a specific decision tree , This is the set of well logging curves after screening. For loss function, For the first The lithology value corresponding to each depth point for The set of well logging curves input after round of iteration At depth point Lithological prediction values at the location, The candidate decision tree being evaluated;
[0030] S32: Construct an extreme gradient boosting decision tree
[0031] ;
[0032] ;
[0033] in, The first derivative of the loss function. The second derivative of the loss function. For the candidate decision trees being evaluated, For regularization terms, The total number of leaf nodes in the candidate decision tree. leaf node The weights;
[0034] S33: Divided equally throughout the entire well section There are several intervals, and samples are randomly drawn from each interval. Constructing a training subset In each iteration, the decision tree Only in the training subset Go to training,
[0035] ;
[0036] S34: In each iteration, the decision tree construction strategies are fused.
[0037] ;
[0038] ;
[0039] ;
[0040] ;
[0041] in, For the first The predicted lithology output of the ensemble model after multiple iterations. For the first The predicted lithology output of the ensemble model after multiple iterations. For learning rate, The predicted correction value is obtained by fusing the decision trees from the three gradient algorithms. For the first Algorithm in the 1st The weights corresponding to the rounds, For the first Algorithm in the 1st The contribution of the wheel.
[0042] Preferably, step S4 further includes the following step:
[0043] S41: Based on the coefficient of determination and root mean square error Evaluation model,
[0044] ;
[0045] ;
[0046] in, Total number of data points For the first The true lithological value at each depth point For the corresponding predicted value, This is the average value for all actual lithologies;
[0047] S42: If the model's evaluation result is optimal, then determine it as the final prediction model; otherwise, return to step S3, readjust the model's parameters and structure, and retrain it. After training, evaluate it again until the model's evaluation result is optimal.
[0048] The present invention has the following advantages: The present invention uses the piecewise weighted Spearman algorithm to screen key logging curve features and integrates three types of gradient algorithms to construct a lithology prediction model, thereby effectively depicting the mapping relationship between logging curves and lithology, significantly improving the accuracy of lithology identification and the generalization ability of the model, and providing reliable support for the fine evaluation of reservoirs. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the lithology prediction method based on the piecewise weighted Spearman algorithm and gradient ensemble learning.
[0050] Figure 2 This is a schematic diagram of some well logging curves;
[0051] Figure 3 A schematic diagram showing the sorting of the weighted sum of squares calculation results for well logging curves;
[0052] Figure 4 A schematic diagram of the segmented weighted Spearman correlation coefficient sorting of well logging curves;
[0053] Figure 5 This is a schematic diagram of the gradient ensemble algorithm process;
[0054] Figure 6 A diagram showing the performance metrics of each model compared;
[0055] Figure 7 This is a schematic diagram comparing the actual lithology with the predicted lithology of well X. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0057] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0058] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.
[0059] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0060] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0061] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0062] In this embodiment, as Figure 1 As shown, a lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning includes the following steps:
[0063] S1: Collect well logging curve data and perform preprocessing;
[0064] S2: Screening well logging curves based on piecewise weighted Spearman algorithm;
[0065] S3: Construct a lithology prediction model;
[0066] S4: Evaluate and optimize the lithology prediction model;
[0067] S5: Lithology prediction for the entire well section. A lithology prediction model is constructed by screening key logging curve features using a piecewise weighted Spearman algorithm and integrating three types of gradient algorithms. This effectively characterizes the mapping relationship between logging curves and lithology, significantly improving lithology identification accuracy and model generalization ability, and providing reliable support for fine reservoir evaluation.
[0068] Furthermore, in step S1, the logging data includes sonic transit time, natural gamma ray, porosity, resistivity, formation bulk density, and clay content, as shown in the schematic diagram of the logging curve. Figure 2 As shown. Furthermore, addressing the significant differences in dimensions among different logging curves, a standardization method is employed to unify the dimensions of various curves, specifically:
[0069] ;
[0070] in, For standardized well logging curve data, For the initial curve data, For curves Median sorted from smallest to largest For well logging curves The median of the second half of the sorted data from smallest to largest. For well logging curves The median of the first half of the data after sorting from smallest to largest.
[0071] In this embodiment, step S2 further includes the following step:
[0072] S21: Convert the logging curves and specific lithology values into their respective ranking orders, and calculate the rank difference at each depth point.
[0073] ;
[0074] in, For the first The first logging curve rank difference at each depth point For the first The first logging curve Rank of the logging curve after conversion at each depth point The order of lithology;
[0075] S22: Calculate the weighted sum of squares ,
[0076] ;
[0077] in, For the first The weighted sum of squares of the logging curves. This represents the total number of data points. For the first The first logging curve Signal-to-noise ratio at each depth point For the first The first logging curve Integrity factor at each depth point For the first The first logging curve rank difference at each depth point For the first The first logging curve Comprehensive adaptive weights at each depth point;
[0078] S23: Calculate the piecewise weighted Spearman correlation coefficient.
[0079] ;
[0080] in, For the first The adaptive weighted Spearman correlation coefficient of the logging curves ranges from -1 to 1. The closer the absolute value is to 1, the stronger the monotonic correlation.
[0081] If the correlation coefficient of the curve is greater than the screening threshold Then it will participate in the training of subsequent prediction models.
[0082] ;
[0083] in, This represents the median of the absolute values of the correlation coefficients of all well logging curves. The difference between the upper and lower quartiles of the absolute values of the correlation coefficients of all well logging curves is used to measure the dispersion of their distribution. Specifically, in step S21, taking the sonic transit time logging curve as an example, Table 1 shows the calculation results of the rank difference for some depth segments.
[0084] Table 1
[0085] depth Preprocessed data Lithology rank of curve Lithological Rank Rank difference 3101.75 -0.911232763 5 227 4810 -4583 3101.875 -0.920633821 5 212 4810 -4598 3102 -0.863673111 2 283 1813.5 -1530.5 3102.125 -0.744854826 2 483 1813.5 -1330.5 3102.25 -0.614418035 2 770 1813.5 -1043.5 3102.375 -0.59356016 2 826 1813.5 -987.5 3102.5 -0.677615319 2 611 1813.5 -1202.5 3102.625 -0.77730811 4 435 4106 -3671 3102.75 -0.797565426 4 403 4106 -3703 3102.875 -0.779294574 4 430 4106 -3676 3103 -0.778601622 4 434 4106 -3672 3103.125 -0.817984432 4 358.5 4106 -3747.5 3103.25 -0.88681773 1 249 679.5 -430.5 3103.375 -0.933314855 1 206 679.5 -473.5
[0086] In step S22, the weighted sum of squares represents the cumulative difference in rank between the logging curve and the lithology. A larger sum indicates lower monotonicity between the two. The weighted sum of squares calculation results for each logging curve are as follows: Figure 3 As shown in the figure; in step S23, the piecewise weighted Spearman correlation coefficients of each logging curve are shown in Table 2, and the corresponding screening threshold is 0.4490. The logging curves that show strong correlation with lithology are POR (porosity), RT (resistivity), GR (natural gamma), AC (acoustic transit time), and DEN (formation bulk density), as shown in the figure. Figure 4 As shown,
[0087] Table 2
[0088] Well logging curves AC CAL DEN GR Correlation coefficient 0.5183 0.1764 -0.5031 -0.5462 Well logging curves POR QT RT SH Correlation coefficient 0.6378 0.3147 0.5916 -0.2975
[0089] Furthermore, such as Figure 5As shown, step S3 also includes the following steps:
[0090] S31: Construct a conventional gradient boosting decision tree
[0091] ;
[0092] in, From decision tree ensemble Find a specific decision tree This minimizes the subsequent summation expression. This is the set of well logging curves after screening. The loss function quantifies the difference between the predicted lithology value and the actual lithology value. For the first The lithology value corresponding to each depth point for The set of well logging curves input after round of iteration At depth point Lithological prediction values at the location, The candidate decision tree to be evaluated is the set of well logging curves given the input. At depth point The predicted correction value at the location;
[0093] S32: Construct an extreme gradient boosting decision tree
[0094] ;
[0095] ;
[0096] in, The first derivative of the loss function. The second derivative of the loss function. For the candidate decision trees being evaluated, This is a regularization term to prevent overfitting. The total number of leaf nodes in the candidate decision tree. leaf node The weight of the lithology prediction is the lithology prediction correction value output by that node;
[0097] S33: Divided equally throughout the entire well section There are several intervals, and samples are randomly drawn from each interval. Constructing a training subset In each iteration, the decision tree Only in the training subset Go to training,
[0098] ;
[0099] S34: In each iteration, the decision tree construction strategies are fused.
[0100] ;
[0101] ;
[0102] ;
[0103] ;
[0104] in, For the first The predicted lithology output of the ensemble model after multiple iterations. For the first The predicted lithology output of the ensemble model after rounds of iterations, and the current gradient ensemble model. It is 500. The learning rate can be adjusted to control the magnitude of corrections in each round; this is the current gradient ensemble model's... It is 0.1. The predicted correction value is obtained by fusing the decision trees from the three gradient algorithms. For the first Algorithm in the 1st The weights corresponding to the rounds, For the first Algorithm in the 1st The contribution of each step. Specifically, a lithology prediction model integrating three different gradient boosting algorithms is constructed. The core of the gradient boosting algorithm is to correct the prediction error of the preceding model by progressively adding decision trees. Different gradient boosting algorithms optimize the decision trees. The respective focuses and advantages are reflected in this way. Specifically, in step S31, the conventional gradient boosting algorithm directly constructs a decision tree using the error of the previous prediction as the fitting target. It focuses on the progressive learning of the complex nonlinear relationship between various logging curves and lithology; in step S32, the extreme gradient boosting algorithm (XGBoost) is used to construct... The regularization term was explicitly added at that time. To prevent overfitting and to enhance the generalization ability of the second-order Taylor expansion model for lithology prediction across different well sections, in step S33, the Lightweight Gradient Boosting (LightGBM) algorithm employs a layered uniform sampling strategy for well sections to handle large-scale data. In step S34, in each iteration, the decision tree construction strategies of the three algorithms are fused, and the actual contribution of each algorithm to the improvement of prediction accuracy in this round is considered. Size, dynamically assigning different weights Contribution It is a quantification of the degree to which prediction error is reduced.
[0105] Furthermore, step S4 also includes the following steps:
[0106] S41: Based on the coefficient of determination and root mean square error Evaluation model,
[0107] ;
[0108] ;
[0109] in, Total number of data points For the first The true lithological value at each depth point For the corresponding predicted value, This is the average value for all actual lithologies;
[0110] S42: If the model's evaluation result is optimal, then it is determined as the final prediction model; otherwise, return to step S3, readjust the model's parameters and structure, and retrain it. After training, evaluate it again until the model's evaluation result is optimal. Specifically, the coefficient of determination... The root mean square error is used to measure the degree of agreement between predicted and actual lithology. A value closer to 1 indicates a more reliable model. This is used to quantify the average deviation between predicted and actual lithology; the smaller the value, the higher the prediction accuracy. The model performance evaluation results are shown in Table 3. The gradient ensemble model exhibits the best overall performance.
[0111] Table 3
[0112] Algorithm Model Gradient Boosting XGBoost LightGBM Ensemble Coefficient of determination 0.330737 0.334863 0.259334 0.38348 Root mean square error 1.058059 1.054793 1.113071 0.76378
[0113] Figure 6 A direct comparison of the various models in terms of coefficient of determination was made. With root mean square error The score differences verified the advantages of the gradient ensemble algorithm. Using the gradient ensemble model finally determined in step S42, lithology prediction was performed on the target well section, achieving accurate lithology identification for the entire well section. Taking well X, which did not participate in the training, as an example, its lithology was predicted based on the relevant logging curves of that well, such as... Figure 7 As shown, the gradient integrated lithology prediction model demonstrates accurate lithology prediction performance and proves its generalization ability, showing promising application prospects.
[0114] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning, characterized in that: Includes the following steps: S1: Collect well logging curve data and perform preprocessing; S2: Screening well logging curves based on piecewise weighted Spearman algorithm; S3: Construct a lithology prediction model; S4: Evaluate and optimize the lithology prediction model; S5: Predict the lithology of the entire well section.
2. The lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning according to claim 1, characterized in that: In step S1, the various curves are subjected to unified dimension processing: ; in, For standardized well logging curve data, For the initial curve data, For curves Median sorted from smallest to largest For well logging curves The median of the second half of the sorted data from smallest to largest. For well logging curves The median of the first half of the data after sorting from smallest to largest.
3. The lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning according to claim 2, characterized in that: Step S2 further includes the following steps: S21: Convert the logging curves and specific lithology values into their respective ranking orders, and calculate the rank difference at each depth point. ; in, For the first The first logging curve rank difference at each depth point For the first The first logging curve Rank of the logging curve after conversion at each depth point The order of lithology; S22: Calculate the weighted sum of squares , ; in, For the first The weighted sum of squares of the logging curves. This represents the total number of data points. For the first The first logging curve Signal-to-noise ratio at each depth point For the first The first logging curve Integrity factor at each depth point For the first The first logging curve rank difference at each depth point For the first The first logging curve Comprehensive adaptive weights at each depth point; S23: Calculate the piecewise weighted Spearman correlation coefficient. ; in, For the first Adaptive weighted Spearman correlation coefficient of the well logging curves. If the correlation coefficient of the curve is greater than the screening threshold Then it will participate in the training of subsequent prediction models. ; in, This represents the median of the absolute values of the correlation coefficients of all well logging curves. This represents the difference between the upper and lower quartiles of the absolute values of the correlation coefficients of all well logging curves.
4. The lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning according to claim 3, characterized in that: Step S3 further includes the following steps: S31: Construct a conventional gradient boosting decision tree ; in, From decision tree ensemble Find a specific decision tree , This is the set of well logging curves after screening. For loss function, For the first The lithology value corresponding to each depth point for The set of well logging curves input after round of iteration At depth point Lithological prediction values at the location, The candidate decision tree being evaluated; S32: Construct an extreme gradient boosting decision tree ; ; in, The first derivative of the loss function. The second derivative of the loss function. For the candidate decision trees being evaluated, For regularization terms, The total number of leaf nodes in the candidate decision tree. leaf node The weights; S33: Divided equally throughout the entire well section There are several intervals, and samples are randomly drawn from each interval. Constructing a training subset In each iteration, the decision tree Only in the training subset Go to training, ; S34: In each iteration, the decision tree construction strategies are fused. ; ; ; ; in, For the first The predicted lithology output of the ensemble model after multiple iterations. For the first The predicted lithology output of the ensemble model after multiple iterations. For learning rate, The predicted correction value is obtained by fusing the decision trees from the three gradient algorithms. For the first Algorithm in the 1st The weights corresponding to the rounds, For the first Algorithm in the 1st The contribution of the wheel.
5. The lithology prediction method based on piecewise weighted Spearman algorithm and gradient ensemble learning according to claim 4, characterized in that: Step S4 also includes the following steps: S41: Based on the coefficient of determination and root mean square error Evaluation model, ; ; in, Total number of data points For the first The true lithological value at each depth point For the corresponding predicted value, This is the average value for all actual lithologies; S42: If the model's evaluation result is optimal, then determine it as the final prediction model; otherwise, return to step S3, readjust the model's parameters and structure, and retrain it. After training, evaluate it again until the model's evaluation result is optimal.