A lithology classification method, device, storage medium and electronic device

Through the stacking model combining basic models such as gradient enhancement trees, random forests and XGboost and logistic regression, the problems of high time consumption and uncertain results in the existing lithologic classification methods are solved, and higher classification accuracy and generalization capabilities are achieved.

CN116821815BActive Publication Date: 2025-07-18CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210265238.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-07-18
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The existing lithologic classification methods have problems such as long time, large workload, and uncertain classification results when using well logging data. Machine learning algorithms such as K-nearest neighbor algorithm, support vector machines and artificial neural networks have shortcomings in data mining and classification effects, especially when the number of lithologic samples is unbalanced.

Method used

A stacked model that fuses different single basis models for lithogenetic category prediction, including the combination of basic models such as gradient boosting trees, random forests, and XGboost and logistic regression models. Through dimensionless, equalization processing and noise addition, standard data sets are constructed to improve classification accuracy and generalization capabilities.

Benefits of technology

It improves the accuracy and generalization ability of lithology classification, can better utilize logging data information, solves the problems of large time and workload and uncertain results in lithology classification, and improves the stability and accuracy of classification algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821815B_ABST
    Figure CN116821815B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a lithology classification method, apparatus, storage medium, and electronic device. The lithology classification method includes: obtaining at least one well logging data of a target well section; inputting the at least one well logging data into a pre-trained stacked model to predict the lithology category of the target well section; wherein, the stacked model includes at least one base model and a logistic regression model, the input of the at least one base model is the at least one well logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one well logging data. The present invention uses a stacked model that combines different single base models to predict the lithology category, fully utilizes the well logging data information, and has better effects in classification accuracy and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lithology classification and identification, and particularly to a lithology classification method, device, storage medium and electronic device. Background Art

[0002] Lithology classification is a very basic and crucial link in log interpretation and is the basis for reservoir evaluation, reservoir description and other work. Accurately and clearly understanding the lithology of formation rocks is of great significance for oil exploration and other research work. At present, the most direct and reliable method for lithology classification is core analysis. However, due to the small number of cores, it is impossible to take cores for the entire well section, and it is costly and time-consuming for analysis, making it difficult to promote. Therefore, in actual exploration and development, using rich and high-precision logging data to indirectly classify the lithology of the target reservoir has become an important means for studying the reservoir.

[0003] The lithology, physical properties and geological characteristics of underground reservoirs can be obtained through different logging technologies. Traditional methods for lithology identification using logging data include the description of logging curves, cross plots and the establishment of linear logging response equations, etc. These methods have problems while making full use of expert experience, such as long time consumption, large workload, strong uncertainty in classification results, etc. In recent years, machine learning lithology classification based on logging data has been developed and widely applied. Common machine learning algorithms include the K-Nearest Neighbor (KNN) algorithm, Artificial Neural Network (ANN), Support Vector Machine (SVM), etc. However, the K-Nearest Neighbor algorithm and the Support Vector Machine (SVM) cannot fully mine the information in logging data; the Artificial Neural Network (ANN) has disadvantages such as slow convergence speed and being easily trapped in local optima. At the same time, considering the problem of unbalanced sample numbers of different lithologies in the reservoir, oversampling algorithms have also been applied. Unbalanced sample numbers will affect the classification effect. When the sample numbers of different lithologies are close to being balanced, the classification effect of the classification algorithm can be better manifested; when the sample numbers of different lithologies are unbalanced, the lithology with fewer samples cannot exert the classification ability of the classification algorithm, resulting in a poor classification effect, and the classification algorithm cannot obtain excellent generalization performance, restricting the application scope.

[0004] Therefore, how to make full use of data information to improve the classification accuracy is an urgent problem to be solved in this field. Summary of the Invention

[0005] Embodiments of the present invention provide a lithology classification method, device, storage medium and electronic device, which use a stacking model that combines different single base models to predict lithology categories, make full use of logging data information, and have better effects in terms of classification accuracy and generalization ability.

[0006] In a first aspect, an embodiment of the present invention provides a lithology classification method, including:

[0007] Obtain at least one logging data of a target well section;

[0008] Input the at least one logging data into a pre-trained stacked model to predict the lithology category of the target well section;

[0009] Wherein, the stacked model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one logging data.

[0010] In some embodiments, in the above lithology classification method, before inputting the at least one logging data into the pre-trained stacked model to predict the lithology category of the target well section, it further includes:

[0011] Perform dimensionless processing on the at least one logging data.

[0012] In some embodiments, in the above lithology classification method, the at least one logging data includes logging data obtained by at least one of the spontaneous potential method, natural gamma method, acoustic travel time method, compensated density method, and deep lateral resistivity method.

[0013] In some embodiments, the at least one base model includes at least one of gradient boosting tree, random forest, and XGboost.

[0014] In some embodiments, in the above lithology classification method, the training process of the stacked model includes:

[0015] Obtain logging data by different logging methods to construct an original logging data set;

[0016] Obtain core data marked with lithology labels to construct a marked logging data set;

[0017] Determine whether the quantity is balanced according to the quantity distribution of logging data of different lithologies, and perform balancing processing on the data set with unbalanced quantity to form a standard data set, where the standard data set includes a training set, a validation set, and a test set;

[0018] Add Gaussian noise to the standard data set to form a standard data set with noise;

[0019] Build and train a stacking model including at least one base model and a logistic regression model. During training, the training data input to the at least one base model comes from the standard data set, and the training data input to the logistic regression model comes from the output of the at least one base model;

[0020] Determine the final stacking model and model parameters.

[0021] In some embodiments, in the above lithology classification method, before determining whether the quantity is balanced according to the quantity distribution of different lithology logging data, it further includes:

[0022] Perform dimensionless processing on the original logging data set and the labeled logging data set;

[0023] Screen out the logging data with a linear relationship lower than the set threshold in the original logging data set and the labeled logging data set.

[0024] In some embodiments, in the above lithology classification method, the performing dimensionless processing on the original logging data set and the labeled logging data set includes:

[0025] Scale the logging data in the original logging data set and the labeled logging data set to the range of [0, 1] through min-max standardization.

[0026] In a second aspect, an embodiment of the present invention provides a lithology classification device, including:

[0027] A data acquisition module for acquiring at least one type of logging data of a target well section;

[0028] A classification prediction module for inputting the at least one type of logging data into a pre-trained stacking model to predict the lithology category of the target well section;

[0029] Wherein, the stacking model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one type of logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one type of logging data.

[0030] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, including: a computer program is stored on the computer-readable storage medium, and when the computer program is executed by one or more processors, it implements the lithology classification method as described in the first aspect.

[0031] Fourthly, an embodiment of the present invention provides an electronic device, including: a memory and one or more processors, where a computer program is stored on the memory, and when the computer program is executed by the one or more processors, the lithology classification method described in the first aspect is implemented.

[0032] Compared with the prior art, one or more embodiments of the present invention at least have the following beneficial effects:

[0033] In the present invention, at least one well logging data of a target well section is acquired, and the at least one well logging data is input into a pre-trained stacking model to predict the lithology category of the target well section; since the stacking model includes at least one base model and a logistic regression model, the advantages of different single base models can be fused, and the well logging data can be fully utilized to obtain an accurate lithology classification result. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope.

[0035] Figure 1 It is a flowchart of a lithology classification method provided by an embodiment of the present invention;

[0036] Figure 2 It is a schematic diagram of a stacking model provided by an embodiment of the present invention;

[0037] Figure 3a It is a schematic diagram of the dataset situation before equalization processing provided by an embodiment of the present invention;

[0038] Figure 3b It is a schematic diagram of the dataset situation after equalization processing provided by an embodiment of the present invention;

[0039] Figure 4 It is an example of a stacking model structure provided by an embodiment of the present invention;

[0040] Figure 5 It is a comparison chart of F1 scores of different models provided by an embodiment of the present invention;

[0041] Figure 6 It is a block diagram of a lithology classification device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0043] Lithology classification is a very basic and crucial link in well logging interpretation and is the basis for reservoir evaluation, reservoir description, and other work. An accurate and clear understanding of the lithology of formation rocks is of great significance for petroleum exploration and other research work. Currently, the most direct and reliable method for lithology classification is core analysis. However, due to the small number of cores, it is impossible to core the entire well section, and it is costly and time-consuming for analysis, making it difficult to promote. Therefore, in actual exploration and development, using rich and high-precision well logging data to indirectly classify the lithology of the target reservoir has become an important means for studying reservoirs.

[0044] Through different well logging technologies, the lithology, physical properties, and geological characteristics of underground reservoirs can be obtained. Traditional methods for lithology identification using well logging data include the description of well logging curves, cross plots, and the establishment of linear well logging response equations, etc. The above methods have problems while making full use of expert experience, such as long time consumption, large workload, and strong uncertainty in classification results. In recent years, machine learning lithology classification based on well logging data has been developed and widely applied.

[0045] Common machine learning algorithms include the K-Nearest Neighbor (KNN) algorithm, Artificial Neural Network (ANN), Support Vector Machine (SVM), etc. However, the K-Nearest Neighbor algorithm and the Support Vector Machine (SVM) cannot fully mine the information in well logging data; the Artificial Neural Network (ANN) has disadvantages such as slow convergence speed and being easily trapped in local optima. At the same time, considering the problem of unbalanced sample numbers of different lithologies in reservoirs, oversampling algorithms have also been applied. Unbalanced sample numbers will affect the classification effect. When the sample numbers of different lithologies are close to balance, the classification effect of the classification algorithm can be better manifested; when the sample numbers of different lithologies are unbalanced, the lithology with fewer samples cannot exert the classification ability of the classification algorithm, resulting in a poor classification effect, which causes the classification algorithm to be unable to obtain excellent generalization performance and limits the application scope.

[0046] Therefore, how to make full use of data information to improve the classification accuracy is an urgent problem to be solved in this field. The present invention provides a lithology classification method, device, storage medium and electronic device, which uses a stacked model that combines different single base models to predict the lithology category, makes full use of well logging data information, and has better effects in classification accuracy and generalization ability.

[0047] Example 1

[0048] Figure 1 shows a flowchart of a lithology classification method, as Figure 1 shown, this embodiment provides a lithology classification method, including steps S101 to S102:

[0049] Step S101, obtain at least one well logging data of the target well section.

[0050] In some embodiments, the at least one well logging data includes well logging data obtained by at least one of natural potential method, natural gamma method, acoustic travel time method, compensated density method, and deep lateral resistivity method.

[0051] In practical applications, each well logging data can be embodied in the form of a corresponding well logging curve.

[0052] Step S102, input the at least one well logging data into a pre-trained stacked model, and predict the lithology category of the target well section.

[0053] Among them, the stacked model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one well logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one well logging data.

[0054] In some embodiments, the at least one base model includes at least one of gradient boosting tree, random forest, and XGboost.

[0055] In practical applications, the lithology category may include but is not limited to mudstone, dolomitic mudstone, siltstone, dolomitic siltstone, micritic dolomite.

[0056] Since the data measured by different well logging methods have different dimensions and attribute value orders of magnitude, if they are directly input into the stacked model, the influence degrees of different types of well logging data on the prediction result may be different. In order to eliminate this systematic error, this embodiment performs dimensionless processing on different types of well logging data, so that different types of well logging data are dimensionless and the prediction result is more accurate. Therefore, in some embodiments, before inputting the at least one well logging data into a pre-trained stacked model to obtain the lithology category of the target well section, it further includes:

[0057] Nondimensionalize at least one well logging data.

[0058] In some implementations, the nondimensionalization process can be carried out in the following manner:

[0059] Scale the well logging data to the range [0, 1] through maximum-minimum normalization. The function used is:

[0060]

[0061] where X new is the data after nondimensionalization processing;

[0062] X max is the maximum value in the well logging data;

[0063] X min is the minimum value in the well logging data.

[0064] In practical applications, before using the stacking model to predict the lithology category of the target well section, it is necessary to pre-construct and train the stacking model. In some embodiments, the training process of the stacking model includes:

[0065] Step a: Obtain well logging data through different well logging methods to construct an original well logging data set.

[0066] At the target well location in the study area, obtain conventional well logging curves (well logging data) through different well logging methods. The well logging methods can include spontaneous potential method, natural gamma method, acoustic transit time method, compensated density method, and deep lateral resistivity method, so as to construct an original well logging data set.

[0067] Step b: Obtain core data marked with lithology labels to construct a labeled well logging data set.

[0068] Combined with core experiments, manually add lithology labels to each sample point to obtain a labeled well logging data set.

[0069] Step c: Perform nondimensionalization processing on the original well logging data set and the labeled well logging data set.

[0070] Since the data measured by different well logging methods have different dimensions and magnitude orders of attribute values, if they are directly used as training data, the influence degrees of different well logging curves on the results are different. To eliminate this systematic error, in this embodiment, the well logging data is nondimensionalized.

[0071] In some embodiments, performing nondimensionalization processing on the original well logging data set and the labeled well logging data set includes:

[0072] Step c1: Scale the logging data in the original logging dataset and the labeled logging dataset to the range of [0, 1] through min-max normalization.

[0073] The function used is:

[0074]

[0075] where X new is the dimensionless data;

[0076] X max is the maximum value in the original logging dataset or the labeled logging dataset;

[0077] X min is the minimum value in the original logging dataset or the labeled logging dataset.

[0078] Step d: Screen out the logging data in the original logging dataset and the labeled logging dataset with a linear relationship lower than the set threshold.

[0079] In some implementation manners, collinearity analysis is performed on the original logging dataset and the labeled logging dataset to screen out the logging data in the original logging dataset and the labeled logging dataset with a linear relationship lower than the set threshold, including:

[0080] Step d1: According to the model requirements, set the correlation coefficient threshold and select the logging curves (logging data) lower than the correlation coefficient threshold.

[0081] If there is a strong linear relationship between different types of logging data, it will not only lead to overfitting of the model, but also reduce the stability and accuracy of the model. Therefore, considering the limited types of commonly available logging data, in this embodiment, the correlation coefficient method is used to perform collinearity analysis on the original logging dataset and the labeled logging dataset.

[0082] The correlation coefficient is a statistical index that reflects the degree of closeness of the relationship between variables. The value range of the correlation coefficient is between 1 and -1. Among them, 1 indicates that the two variables are completely linearly correlated, -1 indicates that the two variables are completely negatively correlated, and 0 indicates that the two variables are not correlated. The closer the data is to 0, the weaker the correlation relationship.

[0083] The correlation coefficient calculation formula is as follows:

[0084]

[0085]

[0086]

[0087]

[0088] Among them, r xy represents the correlation coefficient between sample x and sample y;

[0089] S xy represents the covariance between sample x and sample y;

[0090] S x represents the standard deviation of sample x;

[0091] S y represents the standard deviation of sample y;

[0092] n represents the type of sample;

[0093] It should be understood that, for the samples in the above calculation formula, in this embodiment, they are well logging data. n represents the type of well logging data. Taking the well logging data including the well logging data obtained by the spontaneous potential method, natural gamma method, acoustic travel time method, compensated density method, and deep lateral resistivity method as an example, n = 5.

[0094] Step e: Determine whether the quantity is balanced according to the quantity distribution of well logging data of different lithologies, and perform balancing processing on the unbalanced data set to form a standard data set.

[0095] Among them, the standard data set includes a training set, a validation set, and a test set.

[0096] In this embodiment, the balance of the quantity is determined by analyzing the class balance of well logging data of different lithologies, analyzing the quantity distribution in the standard data set corresponding to different lithology classes, and determining whether the well logging data of each lithology is balanced.

[0097] For the unbalanced data set, the synthetic minority over-sampling technique is used for balancing processing. In one example, the comparison before and after the balancing processing is as Figure 3a and Figure 3b shown. It can be seen that the original data corresponding to the four lithologies A, B, C, and D is unbalanced. After the balancing processing, the data of each lithology reaches balance.

[0098] In one example, the standard data set is divided into three parts: 70% is the training set, 10% is the validation set, and 20% is the test set. Among them, the training set is used to train the stacked model, the validation set is used to optimize the parameters of the stacked model, and the test set is used to evaluate the classification effect of the final stacked model.

[0099] Step f: Add Gaussian noise to the standard data set to form a noisy standard data set.

[0100] In this embodiment, in order to enable the trained stacked model to accurately predict the lithology category, Gaussian noise processing is performed on the training set and the validation set to form a noisy training set and a noisy validation set.

[0101] In one example, the mean of the added Gaussian noise is 0 and the variance is 0.12.

[0102] Step g: Establish a stacked model including at least one base model and a logistic regression model and train it. During the training, the training data input to at least one base model comes from the standard data set, and the training data input to the logistic regression model comes from the outputs of at least one base model.

[0103] A stacked model is a hierarchical ensemble model, and its main idea is to train a model to learn the prediction results of the underlying base models.

[0104] Step f: Determine the final stacked model and its model parameters.

[0105] In this embodiment, the validation set is used to optimize the stacked model, and the model is repeatedly optimized using the validation set. The optimization process is the same as the training process.

[0106] The following indicators can be used as the optimization indicators:

[0107] Accuracy, precision (for one class), recall, and F1Score.

[0108]

[0109]

[0110]

[0111]

[0112]

[0113] TP: Positive samples predicted as positive by the model

[0114] TN: Negative samples predicted as negative by the model

[0115] FP: Negative samples predicted as positive by the model

[0116] FN: Positive samples predicted as negative by the model.

[0117] In some cases, the grid search method and / or the particle swarm optimization algorithm can be used to adjust the parameters of the stacked model, and K-fold cross-validation is used to train the base model and the LR model in the stacked model.

[0118] In one example, the value of K is taken as 10.

[0119] When conducting model evaluation, the final stacked model is evaluated using the test set and the noisy test set to obtain the final classification evaluation result.

[0120] Since the gradient boosting tree and XGBoost are both highly accurate but prone to overfitting, the accuracy of the random forest is slightly lower than the former two but has stronger generalization ability. The stacked model can integrate the advantages of different single models with differences, and the models with differences can play their respective advantages. The stacked model has a simple architecture, fully utilizes the logging data information, is more balanced in terms of classification accuracy, generalization ability, and noise resistance, and has better overall performance. Further, the synthetic minority over-sampling technique is used to amplify the minority class to further improve the classification accuracy.

[0121] The method of this embodiment obtains at least one logging data of the target well section and inputs the at least one logging data into a pre-trained stacked model to predict the lithology category of the target well section; since the stacked model includes at least one base model and a logistic regression model, it can integrate the advantages of different single base models, fully utilize the logging data, and obtain accurate lithology classification results.

[0122] Example 2

[0123] This embodiment provides an application example of a lithology classification method:

[0124] In step S101, at least one logging data of the target well section is obtained. Among them, the at least one logging data includes logging data obtained by natural potential method, natural gamma method, acoustic time difference method, compensated density method, and deep lateral resistivity method.

[0125] The following method is used to perform dimensionless processing on the logging data obtained by natural potential method, natural gamma method, acoustic time difference method, compensated density method, and deep lateral resistivity method:

[0126] The logging data is scaled to the range [0, 1] through maximum-minimum normalization, and the function used is:

[0127]

[0128] where X new is the data after dimensionless processing;

[0129] X max is the maximum value in the logging data;

[0130] X min is the minimum value in the logging data.

[0131] In step S102, the stacking model includes three base models, namely gradient boosting tree, random forest, and XGboost, and a Logistic Regression (LR) model, as Figure 2 shown. Lithology categories can include mudstone, dolomitic mudstone, siltstone, dolomitic siltstone, and micritic dolomite.

[0132] When pre - constructing and training the stacking model, the following steps are included:

[0133] Step a: On the target well positions in the study area, logging data are obtained through different logging methods to construct an original logging data set.

[0134] Step b: Core data marked with lithology labels are obtained to construct a labeled logging data set.

[0135] Step c: Dimensionless processing is performed on the original logging data set and the labeled logging data set: The logging data in the original logging data set and the labeled logging data set are scaled to the range [0, 1] through min - max normalization. The function used is:

[0136]

[0137] where X new is the data after dimensionless processing;

[0138] X max is the maximum value in the original logging data set or the labeled logging data set;

[0139] X min is the minimum value in the original logging data set or the labeled logging data set.

[0140] Step d: Logging data with a linear relationship lower than the set threshold in the original logging data set and the labeled logging data set are filtered out: According to the model requirements, a correlation coefficient threshold is set, and logging curves lower than the correlation coefficient threshold are selected.

[0141] Step e: Determine whether the quantity is balanced according to the quantity distribution of logging data for different lithologies. For data sets with unbalanced quantities, balancing processing is performed to form a standard data set. The standard data set is divided into three parts: 70% is the training set, 10% is the validation set, and 20% is the test set. Among them, the training set is used to train the stacking model, the validation set is used to optimize the parameters of the stacking model, and the test set is used to evaluate the classification effect of the final stacking model.

[0142] Step f: Gaussian noise with a mean of 0 and a variance of 0.12 is added to the standard data set to form a noisy standard data set.

[0143] Step g: Establish a stacking model that includes two base models (XGboost and Random Forest RF) and a Logistic Regression model LR. The model structure is as shown in Figure 4 . Using the Logistic Regression model LR at the secondary level can avoid overfitting of the stacking model.

[0144] Step f: Determine the final stacking model and its model parameters.

[0145] In this embodiment, the validation set is used to optimize the stacking model. The model is repeatedly optimized using the validation set, and the optimization process is the same as the training process.

[0146] The following metrics can be used as optimization metrics:

[0147] Accuracy, Precision (for one class), Recall, and F1 Score.

[0148]

[0149]

[0150]

[0151]

[0152] Taking the F1 score as an example, multiple models including the model of this method are used for lithology classification and identification of mudstone, dolomitic mudstone, siltstone, dolomitic siltstone, and micritic dolomite. The F1 scores obtained by the models are shown in Figure 5 .

[0153] Use the grid search method and / or the particle swarm optimization algorithm to adjust the parameters of the stacking model, and use K-fold cross-validation to train the base models and the LR model in the stacking model, where the value of K is 10.

[0154] When evaluating the model, use the test set and the noisy test set to evaluate the final stacking model to obtain the final classification evaluation result.

[0155] Since both Gradient Boosting Tree and XGBoost have high accuracy but are prone to overfitting, the accuracy of Random Forest is slightly lower than the former two but has strong generalization ability. The stacking model can integrate the advantages of different single models with differences. The models with differences can give full play to their respective advantages. The stacking model has a simple architecture, makes full use of well logging data information, is more balanced in terms of classification accuracy, generalization ability, and anti-noise ability, and has better overall performance. Further, the Synthetic Minority Over-sampling Technique is used to amplify the minority class to further improve the classification accuracy.

[0156] Example 3

[0157] Figure 6 A block diagram of a lithology classification device is shown. This embodiment provides a lithology classification device, including:

[0158] A data acquisition module 601, configured to acquire at least one well logging data of a target well section.

[0159] In some embodiments, the at least one well logging data includes well logging data obtained by at least one of the spontaneous potential method, the natural gamma method, the acoustic transit time method, the compensated density method, and the deep lateral resistivity method.

[0160] In practical applications, each well logging data can be embodied in the form of a corresponding well logging curve.

[0161] A classification prediction module 602, configured to input the at least one well logging data into a pre-trained stacking model, and predict the lithology category of the target well section.

[0162] Wherein, the stacking model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one well logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one well logging data.

[0163] In some embodiments, the at least one base model includes at least one of a gradient boosting tree, a random forest, and XGboost.

[0164] In some cases, the stacking model includes three base models, namely a gradient boosting tree, a random forest, and XGboost, and a logistic regression model (Logistic Regression, LR), as Figure 2 shown.

[0165] In practical applications, the lithology category may include, but is not limited to, mudstone, dolomitic mudstone, siltstone, dolomitic siltstone, and micritic dolomite.

[0166] Since the data measured by different well logging methods have different dimensions and attribute value orders of magnitude, if they are directly input into the stacking model, the influence degrees of different types of well logging data on the prediction result may be different. To eliminate this systematic error, this embodiment performs dimensionless processing on different types of well logging data, so that different types of well logging data are dimensionless and the prediction result is more accurate. Therefore, in some embodiments, before inputting the at least one well logging data into the pre-trained stacking model to obtain the lithology category of the target well section, the classification prediction module 602 is further configured to: perform dimensionless processing on the at least one well logging data.

[0167] In some implementation manners, the dimensionless processing can be performed in the following manner:

[0168] The logging data is scaled to the range of [0, 1] through min-max normalization, and the function used is:

[0169]

[0170] where X new is the data after dimensionless processing;

[0171] X max is the maximum value in the logging data;

[0172] X min is the minimum value in the logging data.

[0173] In some embodiments, the above lithology classification device further includes: a model training module, which is used to pre-construct and train a stacked model before predicting the lithology category of the target well section using the stacked model. In some embodiments, the training process of the stacked model includes:

[0174] Step a: Obtain logging data through different logging methods and construct an original logging data set.

[0175] At the target well location in the study area, conventional logging curves (logging data) are obtained through different logging methods. The logging methods may include spontaneous potential method, natural gamma method, acoustic transit time method, compensated density method, and deep lateral resistivity method, so as to construct an original logging data set.

[0176] Step b: Obtain core data marked with lithology labels and construct a labeled logging data set.

[0177] Combined with core experiments, artificial lithology labels are added to each sample point to obtain a labeled logging data set.

[0178] Step c: Perform dimensionless processing on the original logging data set and the labeled logging data set.

[0179] Since the data measured by different logging methods have different dimensions and magnitude of attribute values, if they are directly used as training data, the influence degrees of different logging curves on the results are different. To eliminate this systematic error, dimensionless processing is performed on the logging data in this embodiment.

[0180] In some embodiments, performing dimensionless processing on the original logging data set and the labeled logging data set includes:

[0181] Step c1: Scale the logging data in the original logging data set and the labeled logging data set to the range of [0, 1] through min-max normalization.

[0182] The function used is:

[0183]

[0184] Among them, X new is the data after dimensionless processing;

[0185] X max is the maximum value in the original logging dataset or the labeled logging dataset;

[0186] X min is the minimum value in the original logging dataset or the labeled logging dataset.

[0187] Step d: Screen out the logging data with a linear relationship lower than the set threshold in the original logging dataset and the labeled logging dataset.

[0188] In some implementation manners, by performing collinearity analysis on the original logging dataset and the labeled logging dataset to screen out the logging data with a linear relationship lower than the set threshold in the original logging dataset and the labeled logging dataset, it includes:

[0189] Step d1: According to the model requirements, set the correlation coefficient threshold, and select the logging curves (logging data) lower than the correlation coefficient threshold.

[0190] Since if there is a strong linear relationship between different types of logging data, it will not only lead to overfitting of the model, but also reduce the stability and accuracy of the model. Therefore, in this embodiment, considering that the types of commonly available logging data are limited, the correlation coefficient method is used to perform collinearity analysis on the original logging dataset and the labeled logging dataset.

[0191] The correlation coefficient is a statistical index that reflects the closeness of the relationship between variables. The value range of the correlation coefficient is between 1 and -1. Among them, 1 means that the two variables are completely linearly correlated, -1 means that the two variables are completely negatively correlated, and 0 means that the two variables are not correlated. The closer the data is to 0, the weaker the correlation relationship.

[0192] The calculation formula of the correlation coefficient is as follows:

[0193]

[0194]

[0195]

[0196]

[0197] Among them, r xy represents the correlation coefficient between sample x and sample y;

[0198] Sxy represents the covariance between sample x and sample y;

[0199] S x represents the standard deviation of sample x;

[0200] S y represents the standard deviation of sample y;

[0201] n represents the type of samples;

[0202] It should be understood that, for the samples in the above calculation formula, in this embodiment, they are logging data. n represents the type of logging data. Taking the logging data including the logging data obtained by the spontaneous potential method, natural gamma method, acoustic travel time method, compensated density method, and deep lateral resistivity method as an example, n = 5.

[0203] Step e: Determine whether the quantity is balanced according to the quantity distribution of logging data of different lithologies, and perform balancing processing on the unbalanced data set to form a standard data set.

[0204] Among them, the standard data set includes a training set, a validation set, and a test set.

[0205] In this embodiment, the quantity balance is determined by analyzing the class balance of logging data of different lithologies, analyzing the quantity distribution in the standard data set corresponding to different lithology classes, and determining whether the logging data of each lithology is balanced.

[0206] For the unbalanced data set, the synthetic minority over-sampling technique is used for balancing processing.

[0207] Step f: Add Gaussian noise to the standard data set to form a noisy standard data set.

[0208] In this embodiment, in order to enable the trained stacked model to accurately predict the lithology class, Gaussian noise processing is performed on the training set and the validation set to form a noisy training set and a noisy validation set.

[0209] Step g: Establish a stacked model including at least one base model and a logistic regression model and train it. During the training, the training data input to at least one base model comes from the standard data set, and the training data input to the logistic regression model comes from the output of at least one base model.

[0210] Step f: Determine the final stacked model and model parameters.

[0211] In this embodiment, the validation set is used to optimize the stacked model, and the model is repeatedly optimized using the validation set. The optimization process is the same as the training process.

[0212] The optimization index can adopt the following index:

[0213] Accuracy, precision (for one category), recall, and F1 Score.

[0214] In some cases, the grid search method and / or the particle swarm optimization algorithm can be used to tune the parameters of the stacked model, and K-fold cross-validation is used to train the base model and the LR model in the stacked model.

[0215] When evaluating the model, the final stacked model is evaluated using the test set and the noisy test set to obtain the final classification evaluation results.

[0216] Since both the gradient boosting tree and XGBoost have high accuracy but are prone to overfitting, the accuracy of the random forest is slightly lower than the former two but has stronger generalization ability. The stacked model can integrate the advantages of different single models with differences, and the models with differences can play their respective advantages. The stacked model has a simple architecture, makes full use of well logging data information, is more balanced in terms of classification accuracy, generalization ability, and anti-noise ability, and has better overall performance. Further, the synthetic minority over-sampling technique is used to amplify the minority class to further improve the classification accuracy.

[0217] The device of this embodiment obtains at least one well logging data of the target well section, and inputs the at least one well logging data into a pre-trained stacked model to predict the lithology category of the target well section; since the stacked model includes at least one base model and a logistic regression model, it can integrate the advantages of different single base models, make full use of well logging data, and obtain accurate lithology classification results.

[0218] It should be understood that the device of this embodiment has all the beneficial effects of the method embodiment.

[0219] Those skilled in the art should understand that the above-mentioned modules or steps can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. The present invention is not limited to any specific combination of hardware and software.

[0220] Example 4

[0221] This embodiment provides a computer-readable storage medium, including: a computer program is stored on the computer-readable storage medium, and when the computer program is executed by one or more processors, it implements the lithology classification method of Embodiment 1.

[0222] In this embodiment, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disc.

[0223] When the computer program is executed by one or more processors, the implemented method includes steps S101 to S102:

[0224] Step S101, obtain at least one well logging data of the target well section.

[0225] In some embodiments, the at least one well logging data includes well logging data obtained by using at least one of the spontaneous potential method, natural gamma method, acoustic travel time method, compensated density method, and deep lateral resistivity method.

[0226] In practical applications, each well logging data can be embodied in the form of a corresponding well logging curve.

[0227] Step S102, input the at least one well logging data into a pre-trained stacked model, and predict the lithology category of the target well section.

[0228] Among them, the stacked model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one well logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one well logging data.

[0229] In some embodiments, the at least one base model includes at least one of gradient boosting tree, random forest, and XGboost.

[0230] In practical applications, the lithology category can include, but is not limited to, mudstone, dolomitic mudstone, siltstone, dolomitic siltstone, micritic dolomite.

[0231] In some embodiments, before inputting the at least one well logging data into the pre-trained stacked model to obtain the lithology category of the target well section, it further includes: dimensionless normalization of the at least one well logging data.

[0232] In some implementations, the following method can be used for dimensionless processing: The logging data is scaled to the range of [0, 1] through min-max normalization.

[0233] In practical applications, before using the stacked model to predict the lithology category of the target well section, it is necessary to pre-construct and train the stacked model. In some embodiments, the training process of the stacked model includes:

[0234] Step a: Obtain logging data through different logging methods to construct an original logging data set.

[0235] Step b: Obtain core data marked with lithology labels to construct a labeled logging data set.

[0236] Step c: Perform dimensionless processing on the original logging data set and the labeled logging data set.

[0237] In some embodiments, performing dimensionless processing on the original logging data set and the labeled logging data set includes:

[0238] Step c1: Scale the logging data in the original logging data set and the labeled logging data set to the range of [0, 1] through min-max normalization.

[0239] Step d: Screen out the logging data in the original logging data set and the labeled logging data set with a linear relationship lower than a set threshold.

[0240] In some implementations, screening out the logging data in the original logging data set and the labeled logging data set with a linear relationship lower than a set threshold through collinearity analysis of the original logging data set and the labeled logging data set includes:

[0241] Step d1: According to the model requirements, set the correlation coefficient threshold, and select the logging curves (logging data) lower than the correlation coefficient threshold.

[0242] If there is a strong linear relationship between different types of logging data, it will not only lead to overfitting of the model, but also reduce the stability and accuracy of the model. Therefore, in this embodiment, considering the limited types of commonly used logging data that can be obtained, the correlation coefficient method is used to perform collinearity analysis on the original logging data set and the labeled logging data set.

[0243] Step e: Determine whether the quantity is balanced according to the quantity distribution of logging data for different lithologies, and perform balancing processing on the data set with unbalanced quantity to form a standard data set.

[0244] Among them, the standard data set includes a training set, a validation set, and a test set.

[0245] In this embodiment, the balance of the quantity is determined by analyzing the category balance of well logging data of different lithologies, the quantity distribution in the standard data set corresponding to different lithology categories is analyzed, and whether the well logging data of each type of lithology is balanced is determined.

[0246] For the unbalanced data set, the Synthetic Minority Over-sampling Technique (SMOTE) is used for balancing processing.

[0247] Step f: Add Gaussian noise to the standard data set to form a noisy standard data set.

[0248] In this embodiment, in order to enable the trained stacking model to accurately predict the lithology category, Gaussian noise processing is performed on the training set and the validation set to form a noisy training set and a noisy validation set.

[0249] Step g: Establish and train a stacking model including at least one base model and a logistic regression model. During the training, the training data input to at least one base model comes from the standard data set, and the training data input to the logistic regression model comes from the outputs of at least one base model.

[0250] Step f: Determine the final stacking model and model parameters.

[0251] In this embodiment, the validation set is used to optimize the stacking model, and the model is repeatedly optimized using the validation set. The optimization process is the same as the training process.

[0252] The optimization metrics can adopt the following metrics:

[0253] Accuracy, precision (for one category), recall, and F1 Score.

[0254] In some cases, the grid search method and / or the particle swarm optimization algorithm can be used to adjust the parameters of the stacking model, and K-fold cross-validation is used to train the base model and the LR model in the stacking model.

[0255] When evaluating the model, the final stacking model is evaluated using the test set and the noisy test set to obtain the final classification evaluation result.

[0256] Example 5

[0257] This embodiment provides an electronic device, including: a memory and one or more processors. A computer program is stored on the memory, and when the computer program is executed by the one or more processors, the lithology classification method of Embodiment 1 is implemented.

[0258] In practical applications, the electronic device can be a terminal device such as a mobile phone or a tablet computer. In this embodiment, the processor can be implemented by an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the method in the above embodiment.

[0259] When the computer program running on the processor is executed, the implemented method includes steps S101 to S102:

[0260] Step S101, obtain at least one well logging data of the target well section.

[0261] In some embodiments, the at least one well logging data includes well logging data obtained by at least one of the spontaneous potential method, the natural gamma method, the acoustic travel time method, the compensated density method, and the deep lateral resistivity method.

[0262] In practical applications, each well logging data can be embodied in the form of a corresponding well logging curve.

[0263] Step S102, input the at least one well logging data into a pre-trained stacked model, and predict the lithology category of the target well section.

[0264] Among them, the stacked model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one well logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one well logging data.

[0265] In some embodiments, the at least one base model includes at least one of a gradient boosting tree, a random forest, and XGboost.

[0266] In practical applications, the lithology category can include, but is not limited to, mudstone, dolomitic mudstone, siltstone, dolomitic siltstone, and micritic dolomite.

[0267] In some embodiments, before inputting at least one well logging data into a pre-trained stacked model to obtain the lithology category of the target well section, it further includes: dimensionless processing of the at least one well logging data.

[0268] In some implementation manners, the dimensionless processing can be performed in the following manner: The well logging data is scaled to the range of [0, 1] through min-max standardization.

[0269] In practical applications, before using the stacked model to predict the lithology category of the target well section, it is necessary to pre-construct and train the stacked model. In some embodiments, the training process of the stacked model includes:

[0270] Step a: Obtain well logging data through different well logging methods to construct an original well logging data set.

[0271] Step b: Obtain core data marked with lithology labels to construct a marked well logging data set.

[0272] Step c: Perform dimensionless processing on the original well logging data set and the marked well logging data set.

[0273] In some embodiments, performing dimensionless processing on the original well logging data set and the marked well logging data set includes:

[0274] Step c1: Scale the well logging data in the original well logging data set and the marked well logging data set to the range of [0, 1] through min-max standardization.

[0275] Step d: Screen out the well logging data with a linear relationship lower than a set threshold in the original well logging data set and the marked well logging data set.

[0276] In some implementation manners, screening out the well logging data with a linear relationship lower than a set threshold in the original well logging data set and the marked well logging data set through collinearity analysis of the original well logging data set and the marked well logging data set includes:

[0277] Step d1: According to the model requirements, set the correlation coefficient threshold, and select the well logging curves (well logging data) lower than the correlation coefficient threshold.

[0278] If there is a strong linear relationship between different types of well logging data, it will not only lead to overfitting of the model, but also reduce the stability and accuracy of the model. Therefore, in this embodiment, considering that the types of commonly used well logging data that can be obtained are limited, the correlation coefficient method is used to perform collinearity analysis on the original well logging data set and the marked well logging data set.

[0279] Step e: Determine whether the quantity is balanced according to the quantity distribution of well logging data of different lithologies, and perform balancing processing on the data set with unbalanced quantity to form a standard data set.

[0280] Among them, the standard data set includes a training set, a validation set, and a test set.

[0281] In this embodiment, the balance of the quantity is determined by analyzing the class balance of the logging data of different lithologies, the quantity distribution in the standard data set corresponding to different lithology classes is analyzed, and whether the logging data of each lithology is balanced is determined.

[0282] For the unbalanced data set, the Synthetic Minority Over-sampling Technique (SMOTE) is used for balancing processing.

[0283] Step f: Add Gaussian noise to the standard data set to form a noisy standard data set.

[0284] In this embodiment, in order to enable the trained stacking model to accurately predict the lithology class, Gaussian noise processing is performed on the training set and the validation set to form a noisy training set and a noisy validation set.

[0285] Step g: Establish a stacking model including at least one base model and a logistic regression model and train it. During the training, the training data input to at least one base model comes from the standard data set, and the training data input to the logistic regression model comes from the output of at least one base model.

[0286] Step f: Determine the final stacking model and model parameters.

[0287] In this embodiment, the validation set is used to optimize the stacking model, and the model is repeatedly optimized using the validation set. The optimization process is the same as the training process.

[0288] The optimization metrics can adopt the following metrics:

[0289] Accuracy, precision (for one class), recall, and F1 Score.

[0290] In some cases, the grid search method and / or the particle swarm optimization algorithm can be used to tune the parameters of the stacking model, and K-fold cross-validation is used to train the base model and the LR model in the stacking model.

[0291] When evaluating the model, the test set and the noisy test set are used to evaluate the final stacking model to obtain the final classification evaluation result.

[0292] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system and method embodiments described above are only illustrative.

[0293] It should be noted that in this text, the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0294] Although the disclosed embodiments of the present invention are as above, the above content is only an embodiment adopted for the convenience of understanding the present invention, and is not intended to limit the present invention. Any person skilled in the art within the technical field to which the present invention pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by the present invention. However, the scope of patent protection of the present invention shall still be subject to the scope defined by the appended claims.

Claims

1. A lithology classification method, characterized in that, Including: Obtain at least one logging data of the target well section; Input the at least one logging data into a pre-trained stacking model to predict the lithology category of the target well section; Wherein, the stacking model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one logging data; The training process of the stacking model includes: Obtain logging data through different logging methods to construct an original logging data set; Obtain core data marked with lithology labels to construct a marked logging data set; Perform dimensionless processing on the original logging data set and the marked logging data set; Screen out the logging data with a linear relationship lower than a set threshold in the original logging data set and the marked logging data set; Determine whether the quantity is balanced according to the quantity distribution of logging data of different lithologies, and perform balancing processing on the data set with unbalanced quantity to form a standard data set, where the standard data set includes a training set, a validation set and a test set; Add Gaussian noise to the standard data set to form a standard data set with noise; Establish and train a stacking model including at least one base model and a logistic regression model. During training, the training data input by the at least one base model comes from the standard data set, and the training data input by the logistic regression model comes from the output of the at least one base model; Determine the final stacking model and model parameters.

2. The lithology classification method according to claim 1, characterized in that, Before inputting the at least one logging data into a pre-trained stacking model to predict the lithology category of the target well section, it further includes: Perform dimensionless processing on the at least one logging data.

3. The lithology classification method according to claim 1, characterized in that, The at least one logging data includes logging data obtained by at least one of the spontaneous potential method, natural gamma method, acoustic travel time method, compensated density method, and deep lateral resistivity method.

4. The lithology classification method according to claim 1, characterized in that, The at least one base model includes at least one of gradient boosting tree, random forest, and XGboost.

5. The lithology classification method according to claim 1, characterized in that The performing dimensionless processing on the original logging data set and the marked logging data set includes: Scale the logging data in the original logging data set and the marked logging data set to the range of [0, 1] through maximum-minimum normalization.

6. A lithology classification device, characterized in that, Including: A data acquisition module for obtaining at least one logging data of the target well section; A classification prediction module for inputting the at least one logging data into a pre-trained stacking model to predict the lithology category of the target well section; Wherein, the stacking model includes at least one base model and a logistic regression model. The input of the at least one base model is the at least one logging data, the output of the at least one base model is used as the input of the logistic regression model, and the output of the logistic regression model is the lithology category predicted according to the at least one logging data; The training process of the stacking model includes: Obtain logging data through different logging methods to construct an original logging data set; Obtain core data marked with lithology labels to construct a marked logging data set; Perform dimensionless processing on the original logging data set and the marked logging data set; Filter out the original logging data set and mark the logging data in the logging data set with a linear relationship lower than the set threshold; Determine whether the quantity is balanced according to the quantity distribution of logging data of different lithologies, and perform balancing processing on the data set with unbalanced quantity to form a standard data set, where the standard data set includes a training set, a validation set, and a test set; Add Gaussian noise to the standard data set to form a standard data set with noise; Build and train a stacking model including at least one base model and a logistic regression model. During training, the training data input to the at least one base model comes from the standard data set, and the training data input to the logistic regression model comes from the output of the at least one base model; Determine the final stacking model and model parameters.

7. A computer-readable storage medium, characterized in that, Including: A computer program is stored on the computer-readable storage medium. When the computer program is executed by one or more processors, the lithology classification method according to any one of claims 1 to 5 is implemented.

8. An electronic device, characterized in that, Including: Including a memory and one or more processors. A computer program is stored on the memory. When the computer program is executed by the one or more processors, the lithology classification method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • A multi-well complex lithology intelligent identification method and system based on logging data

    CN109919184A

  • Advertisement click rate prediction method and device, electronic equipment and readable storage medium

    CN111507765A