Deep strong heterogeneous sandstone reservoir lithofacies logging identification method
Through the KmeanSMOTE method and Bayesian inversion combined with Markov transfer matrix and probability calibration, the problems of data imbalance and insufficient geological constraints in the lithophase logging identification of deep strong heterogeneous sandstone reservoirs were solved, and more accurate lithophase recognition was achieved.
Patent Information
- Application Number
- CN202311701710.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
There are data imbalances and insufficient geological constraints in the identification of lithophagomatic well logging in deep strong heterogeneous sandstone reservoirs, which leads to the impact of model classification performance and makes it difficult to achieve accurate lithophagomatic predictions.
The KmeanSMOTE method is used for clustering, filtering and oversampling, the Markov transfer matrix is constructed, Bayesian inversion is performed, and the posterior distribution is improved through probability calibration, and resampling is combined with core data to improve the recognition accuracy of the model.
Through the improved method, the lithophase type of deep strong heterogeneous reservoirs can be more accurately identified, which improves the recognition accuracy and solves the problems of data imbalance and insufficient geological constraints.
Smart Images

Figure CN120145178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of lithofacies logging identification, and particularly to a method for identifying lithofacies of deep and strongly heterogeneous sandstone reservoirs by logging. Background Art
[0002] Lithofacies can form a bridge between the microscopic and mesoscopic characteristics of reservoirs, and thus provide a basis for geological research at the macroscopic scale. Therefore, lithofacies parameters are very important in reservoir modeling and reservoir characterization. Compared with conventional reservoirs in the middle and shallow layers, the mineralogical composition of most deep and low-permeability clastic rock reservoirs is more complex, and the lithofacies change rapidly. The strong reservoir heterogeneity poses great risks to oil and gas exploration. Accurately describing the lithofacies distribution of such reservoirs is crucial for the success or failure of oil and gas exploration. Due to the high cost of coring in deep wells and the possible drilling safety risks during the coring process, the available core data and the lithofacies distribution information provided by them are often very limited. Therefore, logging lithofacies identification is the most important method for obtaining rock information of deep reservoirs.
[0003] Traditional lithofacies logging interpretation methods rely on interpretation models, such as using T-S diagrams to interpret sandstone types and logging curve superposition, etc. In addition, specific parameters (such as petrophysical parameters) are introduced during the lithofacies identification process to assist in rock facies identification. These methods have good applications for specific geological problems, but often require simplification of geological attributes. In recent years, a large number of scholars have tried to apply data-driven logging interpretation methods to lithofacies identification of reservoirs in different regions and of different types, which are mainly divided into unsupervised interpretation methods and supervised interpretation methods. The supervised interpretation method pays more attention to the correlation between geological attributes and logging responses. Among them, the Bayesian method can apply different prior frameworks and likelihood models to avoid geologically or petrophysically unlikely transitions between lithofacies.
[0004] Although the previous research work has greatly promoted the rapid development of big data identification of logging lithofacies, there are still some problems when applying big data methods to lithofacies classification of strongly heterogeneous reservoirs at present. The most prominent problem is that the proportion of different types of lithofacies corresponding to logging data in the vertical direction varies greatly. This data imbalance problem will affect the classification performance of the model. When considering this problem, previous studies generally merged the minority lithofacies, but these minority lithofacies often have a more important impact on oil and gas reservoir characterization and modeling. For example, mudstone in thick sandstone can act as an interlayer to hinder fluid flow. At the same time, high-quality lithofacies often account for a small proportion in the formation and are also vulnerable to data imbalance. Obviously, these minority lithofacies cannot be ignored. Therefore, it is necessary to improve the data imbalance problem in the logging dataset.
[0005] Another problem is that the traditional lithofacies logging identification method has too little geological constraint, the samples are independent of each other, and the identification is carried out point by point vertically. Therefore, the interpretation results do not conform to the actual formation situation, and even some lithofacies sequences that are less likely geologically may appear. Previous researchers have realized that more geological constraints should be introduced, specifically, accurate geological priors should be introduced during the identification process. In fact, in response to these problems, some researchers have conducted preliminary explorations on seismic data volumes, and proved that applying the Markov chain method in the Bayesian framework can obtain a prior model that conforms to the formation lithofacies distribution. These research works provide valuable references for how to introduce geological constraints in logging lithofacies identification.
[0006] The logging lithofacies identification of strongly heterogeneous reservoirs is a complex non-linear classification problem, and it is difficult for current methods to accurately predict the heterogeneous lithofacies of deep reservoirs. Summary of the Invention
[0007] In view of the above problems, the present invention is proposed to provide a logging lithofacies identification method for deep strongly heterogeneous sandstone reservoirs that overcomes the above problems or at least partially solves the above problems.
[0008] According to one aspect of the present invention, there is provided a logging lithofacies identification method for deep strongly heterogeneous sandstone reservoirs, and the identification method includes:
[0009] Step S1: successively perform clustering, filtering, and oversampling using the KmeanSMOTE method;
[0010] Step S2: resample according to the number of core data samples in the study area with reference to the parameters in Step S1;
[0011] Step S3: construct a Markov transition matrix;
[0012] Step S4: perform Bayesian inversion;
[0013] Step S5: further improve the posterior distribution using probability calibration.
[0014] Optionally, in Step S1: when successively performing clustering using the KmeanSMOTE method, the k-means clustering method is used to cluster the input space into k clusters.
[0015] Optionally, the filtering in Step S1 specifically includes:
[0016] Select the clusters to be oversampled and determine how many samples to generate in each cluster, and use the imbalance rate IR and the sampling weight α os for control;
[0017] The calculation formula of the imbalance rate IR is as follows,
[0018]
[0019] In the formula, c is the clustering cluster label, majorityCount is the majority class sample count of the cluster, minorityCount is the minority class sample count of the cluster, and the imbalance rate IR is the ratio of the majority class sample count to the minority class sample count of a certain cluster c;
[0020] Sampling weight α os The calculation formula is as follows:
[0021] α os = N rm / N M
[0022] In the formula, N rm is the number of samples in the minority class after resampling, and N M is the number of samples in the majority class after resampling. The sampling weight corresponds to the ratio of the number of samples in the minority class to the number of samples in the majority class after resampling;
[0023] When performing multi-class classification, it is necessary to distinguish between the majority class and the minority class.
[0024] Optionally, the oversampling specifically includes:
[0025] Applying SMOTE in each selected cluster to increase the ratio α of the minority samples to the majority samples os ;
[0026] SMOTE has a process: First, select a random minority class observation X i , and among its k nearest minority class neighbors, select the sample X zi ;
[0027] Generate a new sample through the following formula:
[0028] x new = x i + λ × (x zi - x i )
[0029] In the formula, λ is a random weight in [0,1], and X i , X zi are the original samples of the minority class in the selected cluster.
[0030] Optionally, the step S3: constructing the Markov transition matrix specifically includes: constructing the Markov transition matrix by using the equal-spacing stratigraphic unit method.
[0031] Optionally, after the step S3: constructing the Markov transition matrix, it further includes:
[0032] After constructing the Markov transition matrix P, a random initial lithofacies distribution probability π(0) is given, and then the probability π(t) of different states at time t is
[0033] π(t - 1)P = …… = π(0)Pt. After long-term iteration n times, π(t) of different iteration processes is obtained;
[0034] When n is large enough, π(t) will be very close to the true lithofacies distribution, thus determining the prior probability distribution:
[0035]
[0036] where i and j are lithofacies categories, n is the number of iterations, the state space I = {0, 1, 2, …} represents the total lithofacies categories, and π j is the probability of lithofacies j approximated through iteration.
[0037] Optionally, the step S4: performing Bayesian inversion specifically includes:
[0038] When performing Bayesian inversion, the prior probability is determined according to the Markov chain, and the likelihood is assumed to be a Gaussian distribution, and its mean and variance are estimated according to the maximum likelihood method.
[0039] Optionally, the step S5: further improving the posterior distribution using probability calibration specifically includes:
[0040] Further improve the posterior distribution using probability calibration, map the posterior probability output by the classifier to [0, 1] to obtain the calibrated probability. For the output f of the classifier for a given sample i , the calibrator predicts p(y i = 1|f i ).
[0041] Optionally, for the multi-classification problem in the classifier, the one-vs-rest classifier strategy is adopted, fitting a classifier for each class and fitting it with all the remaining classes.
[0042] Optionally, the probability calibration in the step S5 adopts isotonic regression. Given a predicted probability f i , the true probability y i is expressed as: y i = m(f i ) + ε;
[0043] where m represents a monotonically increasing function and is also the objective function to be fitted:
[0044] Then the decision function is corrected to
[0045] where the likelihood function uses the function Make corrections:
[0046]
[0047] Since should be a monotonically increasing function, the above equation simplifies to:
[0048]
[0049] where the weight w i is a strictly positive number, and the result is a piecewise linear function. The corrected likelihood function is considered to be able to reflect the data characteristic distribution after oversampling.
[0050] Optionally, in the step S1: using the KmeanSMOTE method to perform clustering in sequence, the k-means clustering method is used to cluster the input space into k clusters.
[0051] A method for identifying lithofacies of deep strong heterogeneous sandstone reservoirs provided by the present invention, the identification method includes: step S1: using the KmeanSMOTE method to perform clustering, filtering, and oversampling in sequence; step S2: resampling according to the number of core data samples in the study area with reference to the parameters in step S1; step S3: constructing a Markov transition matrix; step S4: performing Bayesian inversion; step S5: further improving the posterior distribution using probability calibration. Accurately identify the lithofacies of strong heterogeneous reservoirs using conventional logging data.
[0052] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a lithofacies identification flowchart based on logging and core data provided by an embodiment of the present invention;
[0055] Figure 2 It is a schematic diagram of oversampling the safe area using the K-means SMOTE method provided by an embodiment of the present invention and eliminating the within-class imbalance phenomenon;
[0056] Figure 3Schematic diagram of clustering input data provided by the embodiments of the present invention;
[0057] Figure 4 Schematic diagram of the change in the distribution of lithofacies data before and after KmeanSMOTE processing provided by the embodiments of the present invention;
[0058] Figure 5 Schematic diagram of the process of sedimentary modeling of the study area based on Markov chain provided by the embodiments of the present invention;
[0059] Figure 6 Schematic diagram of the iterative process of prior probability provided by the embodiments of the present invention. a is a schematic diagram of the sedimentary sequence of a single well (taking a partial well section of Well Zhuang 101 as an example), b is a color phase diagram of the transition matrix, and c is a schematic diagram of the change in the probability of a single lithofacies during the iterative process;
[0060] Figure 7 Schematic diagram of the comparison between the naive Bayes recognition of the second member of the Sangonghe Formation in the Moxizhuang area and two probability correction curves provided by the embodiments of the present invention (taking the probability of fine sandstone as an example);
[0061] Figure 8 Schematic diagram of the comparison of the confusion matrix before and after improvement provided by the embodiments of the present invention. Detailed implementation manners
[0062] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0063] The terms "comprising" and "having" and any variations thereof in the description, embodiments and claims of the present invention are intended to cover non-exclusive inclusion. For example, a series of steps or units are included.
[0064] Hereinafter, with reference to the accompanying drawings and embodiments, the technical solutions of the present invention will be further described in detail.
[0065] Embodiment 1
[0066] The present invention provides a method for identifying lithofacies from well logs in a deep and strongly heterogeneous sandstone reservoir based on improved Bayesian inversion, aiming to accurately identify the lithofacies of strongly heterogeneous reservoirs using conventional well logging data and solve the above technical problems. The specific objectives include: (1) solving the data imbalance problem in well log lithofacies identification of heterogeneous reservoirs due to the relatively large difference in lithofacies proportions; (2) constructing a sedimentary prior using the Markov chain method to provide geological constraints for lithofacies distribution prediction; (3) proposing a set of Bayesian inversion processes applicable to strongly deep heterogeneous reservoirs.
[0067] A method for identifying lithofacies from well logs in a deep and strongly heterogeneous sandstone reservoir provided by the present invention specifically includes the following steps:
[0068] Step S1, perform oversampling using the KmeanSMOTE method;
[0069] Step S2, perform resampling according to the number of core data samples in the study area;
[0070] Step S3, construct a Markov transition matrix;
[0071] Step S4, perform Bayesian inversion;
[0072] Step S5, further improve the posterior distribution using probability calibration.
[0073] The improvement of the present invention is based on the Naive Bayes method. The Naive Bayes method is a supervised learning algorithm based on Bayes' theorem, which assumes that the feature parameters are independent of each other. Since the well logging data are not strictly independent, it is not easy to directly use them as the input data of the classifier. Therefore, the present invention uses the method of principal component analysis (PCA) to perform singular value decomposition on the original well logging data, project it into a low-dimensional space, and perform a linear transformation to obtain uncorrelated variables X{x 0 ,x 1 ......,x n-1}.
[0074] The mathematical expression of Bayes' theorem is as follows:
[0075]
[0076] Among them, y is the class label; x 1 ,…,x n represents the feature parameters; P(y|x 1 ,…,x n ) is the conditional probability of a set of observation results corresponding to a specific class y, that is, the posterior distribution; P(y) is the prior probability of class y, and the prior information of lithofacies can be obtained according to the core observation results; P(x 1 ,…,x n |y) is the likelihood function.
[0077] Since the Naive Bayes assumes that the feature parameters are independent of each other, P(x 1 , …, x n ) is a constant, and Equation (2-1) is equivalent to:
[0078]
[0079] where is the joint probability of the input variable X under the condition of class y. Based on this, for each group of X, when making a decision, the class that maximizes P(y|x 1 , ……, x n ) is selected as the output class, that is:
[0080]
[0081] According to the principle of the Bayesian method, the key to this method lies in obtaining the appropriate prior function P(y) and likelihood function p(x i |y). Usually, the maximum a posteriori (MAP) and Markov chain Monte Carlo (MCMC) methods are used to obtain them. However, Bayesian inference allows adjusting the prior by applying the Bayesian rule in different datasets (Dymarski, 2011). Therefore, the process of obtaining the prior probability can be separated, and a specific method can be used to accurately estimate the prior.
[0082] Due to different assumptions of the likelihood function, the classifiers also vary. Since the input parameter characteristics of each lithofacies show an obvious normal distribution, the Gaussian likelihood function is adopted in this paper to obtain the joint distribution. The calculation formula for each feature is as follows:
[0083]
[0084] where μ yk,i represents the mean of the feature x k in the samples of class y i ; represents the variance of the feature x k in the samples of class y i . The parameters for each class y can all be estimated by the maximum likelihood method.
[0085] The research of modern geology has found that the superposition of deterministic relationships such as the rhythm, cycle, and period of rock formations and random relationships during the sedimentation process often endows the stratigraphic sequence with the properties of a Markov chain. This Markov property is reflected in the lithofacies section, which can help accurately estimate the probability P(y) of the distribution of different lithofacies in the reservoir within the working area.
[0086] This property is expressed as: given the state at the current time, the conditional distribution of the state at a future time is only related to the state at the current time, that is:
[0087] π{Y(t + 1) = i n+1 |Y(t) = i n , Y(t - 1) = i n+1 , Y(1) = i 1 , Y(0) = i 0}} = π{Y(t + 1) = j|Y(t) = i} (5)
[0089] In the formula, π is the probability, Y(t) is the state at time t on the Markov chain, and i is the value of its state. There is a linear correspondence between time and depth in the strata, so Y(t) = Y(h) in the formula, representing the state at the hth position on the Markov chain.
[0090] Considering that the shallower the stratum depth, the later the corresponding lithofacies appears, for the single - well lithofacies sequence, an upward Markov chain should be established. When inferring the state of the next position, it needs to be realized through the transition matrix. The probability P of transitioning from state i to j is:
[0091] p(i, j) = p(i → j) = π{Y t+1 = j|Y t = i} (6)
[0092] Calculate the number of transitions of the entire classification from bottom to top based on core data and normalize it to estimate the transition probability, then the upward transition probability matrix P can be obtained; obtain the random - sequence transition matrix S according to the lithofacies proportion, and determine the lower - limit probability of the transition of each lithofacies based on this, and then correct the transition matrix P.
[0093] After constructing the Markov transition matrix P, randomly give an initial lithofacies distribution probability π(0), then the probability π(t) of different states at time t is π(t) = π(t - 1)P = …… = π(0)P t , after long - term iteration n times, the π(t) of different iteration processes can be obtained. When n is large enough, π(t) will be very close to the true lithofacies distribution, thus determining the prior probability distribution:
[0094]
[0095] where \(i\) and \(j\) are lithofacies categories, \(n\) is the number of iterations, the state space \(I=\{0,1,2,\cdots\}\) represents the total lithofacies categories, and \(\pi\) j is the probability of lithofacies \(j\) approximated through iteration.
[0096] In addition, when calculating the prior probability \(P(y)\), it is necessary to consider how to calculate the transition matrix of the Markov chain of multiple wells, and the prior probability of the work area should be obtained by weighting according to the depth. Because essentially the probability represented by \(P(y)\) is not the frequency of lithofacies distribution, but the thickness of various lithofacies distributions under the condition of a determined formation. The former is greatly affected by the sampling interval, while the latter is an inherent property of the formation. Compared with the conventional Bayesian method, the spatial relationship of lithofacies distribution between different depth points (characterized by the transition matrix) is considered in this prior model, thus being more in line with geological laws.
[0097] When estimating the parameters of the likelihood function, due to the influence of data set imbalance, the likelihood function cannot accurately describe the distribution of the minority class, and the model will have bias. At the same time, the data characteristics of the minority class samples are not sufficient enough for the classifier to fully learn. One way to solve this problem is to generate new samples in the underrepresented classes, that is, oversampling.
[0098] The oversampling method adopted in this paper is K-means SMOTE, that is, K-means Synthetic Minority Over-sampling Technique. The K-means algorithm is used to cluster the input data set, and minority class oversampling is combined within the clusters with more minority class samples, so as to avoid the generation of noise. K-means SMOTE consists of three steps: clustering, filtering, and oversampling, as Figure 2 shown.
[0099] In the clustering step, the k-means clustering method is used to cluster the input space into \(k\) clusters. \(k\) is the most important hyperparameter in the K-means method. Finding a suitable value for \(k\) is crucial for the effectiveness of K-means SMOTE because it will affect the number of minority clusters found in the filtering step.
[0100] The filtering step selects the clusters to be oversampled and determines how many samples to generate in each cluster. The purpose is to oversample only the clusters dominated by the minority class, so as to avoid generating noise as much as possible, and at the same time allocate more of the newly generated samples to the sparse minority class clusters rather than the dense ones. It can be controlled by the imbalance ratio and sampling weights:
[0101]
[0102] Among them, c is the clustering cluster label, majorityCount is the majority class sample count of the cluster, minorityCount is the minority class sample count of the cluster, and the imbalance rate IR is the ratio of the majority class sample count to the minority class sample count of a certain cluster c.
[0103] α os = N rm / N M (9)
[0104] Among them, N rm is the number of samples in the minority class after resampling, N M is the number of samples in the majority class after resampling, and the sampling weight corresponds to the ratio of the number of samples in the minority class to the number of samples in the majority class after resampling. When performing multi-class classification, it is necessary to distinguish between the majority class and the minority class.
[0105] In the oversampling step, SMOTE is applied to each selected cluster to increase the ratio α of the minority samples to the majority samples. os . The process of SMOTE is as follows: First, select a random minority class observation x i , and among its k nearest minority class neighbors, select the sample X zi ; generate a new sample through the following formula:
[0106] x new = x i + λ × (x zi - x i ) (10)
[0107] In the formula, λ is a random weight in [0,1], and X i , X zi are the original samples of the minority class in the selected cluster.
[0108] In fact, applying the oversampling method may cause some other problems to the model, such as increased variance (due to the change in the number of samples) and distorted posterior distribution (due to the impact on the prior probability). The first problem can be reduced by the averaging strategy, while the second problem requires calibrating the new prior probability and updating the likelihood function. According to Bayes' principle, when ensuring data independence, the posterior probability obtained by inverting with part of the samples can also be regarded as the prior probability after calibration and then inversion. Therefore, this error can be corrected by calibrating the posterior probability of the initial inversion result.
[0109] The probability calibration method not only uses the maximum posterior probability for discrimination, but also includes a fitting regressor that maps the posterior probability output by the classifier to [0,1] to obtain the calibrated probability, that is, for the output f i of the classifier for a given sample, the calibrator predicts p(y i= 1f i )(Two-class classification). For multi-class classification problems, the one-vs.-rest classifier strategy is adopted, that is, a classifier is fitted for each class, and this class is fitted against all other classes. Isotonic regression is a commonly used probability calibration method and belongs to a non-parametric regression model. It assumes that the function space is monotonically increasing. Given a predicted probability f i , the true probability y i is expressed as:
[0110] y i = m(f i ) + ε (11)
[0111] where m represents a monotonically increasing function and is also the target function to be fitted:
[0112] Then the decision function is corrected to
[0113]
[0114] where the likelihood function is corrected using the function for correction:
[0115]
[0116] Since should be a monotonically increasing function, the above formula can be simplified to:
[0117]
[0118] where the weight w i is a strictly positive number, and the result should be a piecewise linear function. The corrected likelihood function can be considered to be able to reflect the data characteristic distribution after oversampling.
[0119] The flow of applying the whole method to lithofacies classification is as Figure 1 shown. Through step-by-step calibration and verification, the lithofacies information propagates from the core sampling interval to the logging coverage interval, and the final lithologies obtained from different wells are geologically matchable. This lithofacies parameter can be used for facies interpretation and reservoir modeling.
[0120] The present invention uses a confusion matrix to evaluate the classification accuracy. By definition, each column of the confusion matrix represents the predicted class, and its total represents the data predicted as this class; each row represents the true attribution class of the data, and its total represents the number of instances of this class. Based on the confusion matrix, the precision, recall rate, and F1 value can be easily calculated. To avoid underfitting and overfitting, the present invention uses the K-fold cross-validation method to calculate the precision, recall rate, and F1 value.
[0121] Example 2
[0122] The present invention first performs oversampling using the KmeanSMOTE method, with the main parameters being IR (imbalance ratio) and α. os The sampling weight, where the former controls the algorithm to discriminate unbalanced clusters, and the latter controls the specific resampling number for each class. The clustering step uses the mini-batch K-means algorithm to reduce the calculation time. The size of the mini-batch (i.e., the subset of the input data) is 1024; in the oversampling step, the nearest neighbor algorithm is used to judge nearby samples, and the number of neighbors used for the algorithm query is 5. The Euclidean metric is used to calculate the neighbor distance.
[0123] For the clustering step, generally, the number of clustering clusters k should be close to the actual classes. As Figure 3 shown, it can be seen that when the number of clusters is set to 8, the clustering centers often deviate from the clusters. Therefore, the number of clusters should be increased to make the clusters more homogeneous. After continuous testing, when the number of clusters is 15, the clustering centers can well represent the distribution range of the clusters, and at this time, the sum of the squares of the distances from the samples to the nearest neighbors is the smallest.
[0124] Regarding the sampling weight α os , since it is a multi-class problem, the target classes for resampling are specified as all classes except the majority class, and the ratio is set to make the number of samples in different classes equal. Regarding the imbalance ratio IR, when the imbalance rate threshold is large, the cluster selection is more selective, and a higher proportion of minority samples is required to select a cluster. However, the proportion of some minority classes in this paper is too sparse (less than 5% of the total number of samples), and the commonly set threshold (usually 1) will cause these classes to be unable to find a suitable cluster for oversampling, and no cluster dominated by minority classes can be formed after clustering, so the standard needs to be relaxed. Therefore, in this paper, continuous testing is carried out according to the data characteristics, and the threshold is set to 0.25, allowing the selection of clusters with a higher majority ratio. Clusters above this threshold are considered unbalanced and need to be sampled. The sampling quantity of each lithofacies determined according to the sampling weight is: [(0, 640), (1, 630), (2, 687), (3, 93), (5, 669), (6, 666), (7, 710)].
[0125] According to the number of core data samples in the study area, resampling is carried out according to the above parameters. The original number of samples before processing is 2017, and the total number of samples after processing is 6128. The distribution of the newly generated samples is as Figure 4 shown. Since the dimension of the observed variables is relatively high and not conducive to observation, the characteristic variables are converted into two-dimensional variables using the principal component analysis method for visualization. The x-axis is the first principal component, and the y-axis is the second principal component. The cumulative variance contribution rate is 73%. There are no sample points distributed in some characteristic spaces, which may be related to variance loss. The first and second principal components cannot fully cover the variable space; it is also possible that the characteristics of the data points are limited and comprehensive data has not been collected.
[0126] The left figure shows the distribution of the original dataset without oversampling on two-dimensional variables, and the right figure shows the data distribution after oversampling under two-dimensional variables. It can be found that the data distribution is more balanced after processing. At the same time, some discrete data points are not encrypted after screening, such as 3 (fine sandstone) and 4 (medium sandstone). In addition, 7 (fine conglomerate) is the sparest, so it is supplemented most fully. However, its data distribution is already somewhat different from the original data distribution. Some points in the feature space are relatively dense, so oversampling has a certain impact on the prior probability. At the same time, the algorithm itself also distorts the probability estimation of the lithofacies distribution for each class, and probability calibration is required.
[0127] In this paper, an equal-spacing stratigraphic unit method is used to calculate the transition matrix, with a spacing of 0.125 m, which is consistent with the sampling interval of the logging curve. Based on the vertical distribution of lithofacies obtained from core observations, the labels at each interval point are determined, and the sedimentary sequence is established after statistically analyzing the lithofacies data of each well section in the study area. The specific processing flow is as Figure 5 shown. In a single well, the transition times between different lithofacies are sequentially counted from bottom to top to form a single-well transition times matrix N. For the entire work area, wells with continuous cores and strong representativeness, such as Well Zheng 11, Well Zhuang 3, Well Zhuang 101, Well Zhuang 105, and Well Zhuang 102, are selected as representative wells. The transition times matrices of these wells are correspondingly added according to depth weighting to obtain the transition times matrix of the work area, forming the transition matrix P, as shown in Table 1.
[0128] Taking the first row of the transition matrix obtained from all core observation data of the second member of the Sangonghe Formation in the Moxizhuang area in Table 1 as an example, the characteristics of this transition matrix are briefly introduced. In core observation, lithofacies 5, 6, and 7 will not directly convert to lithofacies 0. From a sedimentary perspective, it means that mudstone facies will not directly develop above coarse-grained sandstone (coarse sandstone, gravel-bearing sandstone, and conglomerate). In contrast, lithofacies 1, 2, 3, and 4 can undergo a transition to 0, and their probabilities are 0.0313, 0.0313, 0.0141, and 0.0795 respectively.
[0129] On the basis of constructing the Markov transition matrix, the average value of each lithofacies statistically obtained from the core observation data is set as the initial probability, and the transition matrix is continuously multiplied iteratively until a stationary distribution is obtained, that is, the prior probability corresponding to the lithofacies distribution is obtained. Among them, the number of iterations is set to 100,000 times, and the stationary error is set to 10 -6 , and once the difference between the current matrix and the previous iteration matrix is less than the stationary error or the number of iterations reaches 100,000 times, this matrix is accepted as the stationary distribution.
[0130] Table 1 Transition matrix of all core observation results of the second member of the Sangonghe Formation in the Moxizhuang area
[0131]
[0132]
[0133] Figure 6 An example of the process of obtaining the prior probability using a Markov chain is given. In the figure, the probabilities of different lithofacies tend to stabilize when the iteration step reaches 80. It can be seen from the transfer matrix color phase diagram that the mudstone facies is relatively stable and not easily transformed into other lithofacies, while the gravel-bearing sandstone and muddy gravel-bearing sandstone are easily transformed into other lithofacies, which is also consistent with their geological significance. Mudstone is relatively stable in the sedimentary sequence and is quite different from other lithofacies in structure; while the gravel-bearing sandstone and muddy gravel-bearing sandstone have a wide grain size range and are more distinguished from other lithofacies in terms of composition, so they are easy to transform.
[0134] The final calculation results show that the prior probability values of the lithofacies are 0.095 (mudstone), 0.102 (very fine sandstone), 0.086 (muddy gravel-bearing sandstone), 0.211 (fine sandstone), 0.231 (medium sandstone), 0.101 (coarse sandstone), 0.173 (gravel-bearing sandstone), and 0.087 (fine conglomerate), respectively. That is, in the second member of the Sangonghe Formation, fine sandstone, medium sandstone, and gravel-bearing sandstone play a dominant role, and the probability of their occurrence in the strata is about 20%, while mudstone, muddy gravel-bearing sandstone, coarse sandstone, very fine sandstone, and conglomerate are less, generally about 10%. In fact, there are significant differences between this prior probability and that obtained by the MCMC method. The Py obtained by MCMC is [0.085, 0.101, 0.043, 0.291, 0.335, 0.058, 0.046, 0.041]. The distribution probabilities of some lithofacies in the latter are relatively extreme and do not match well with the distribution of the reservoir lithofacies in the second member of the Sangonghe Formation in the work area.
[0135] When performing Bayesian inversion, the prior is determined according to the Markov chain, and the likelihood is assumed to be a Gaussian distribution, and its mean and variance are estimated according to the maximum likelihood method. The results are shown in Table 2.
[0136] Table 2 Statistical table of Gaussian characteristic parameter values of input data corresponding to different lithofacies in the second member of the Sangonghe Formation in Moxizhuang area
[0137]
[0138] On this basis, the posterior distribution is further improved using probability calibration. Figure 7It shows the differences between the naive Bayesian prediction probabilities for lithology 3 and the prediction probabilities optimized by two probability calibrations (isotonic regression and sigmoid function calibration) and the true distribution probabilities at different probability distribution stages (partial points are shown). It can be seen from the figure that for different lithofacies, the probability calibration method optimizes the prediction probabilities of the naive Bayesian on the basis of the likelihood function. This calibration is particularly obvious in lithofacies 2, 3, 5, and 7, which also correspond to the lithofacies with the worst recognition effects in the naive Bayesian. The Brier score is shown in the lower right corner of the figure, and it is obvious that the gap between the predicted probability and the true probability is further reduced after probability calibration. The diagonal line in the figure represents the case where the predicted probability exactly matches the true probability. It can be seen that the probability predicted after isotonic regression calibration corresponding to the yellow line in the figure is closest to the probability of the true sample. Isotonic regression itself is directly optimized for logarithmic loss, so the training process itself is to make the predicted probability as close as possible to the true label.
[0139] Based on the above improvements, a set of improved Bayesian inversion processes was formed, and the confusion matrix results were obtained through training and discrimination in 7 wells in the work area. The comparison with the original naive Bayesian classification results is as Figure 8 shown. The right figure is the recognition result after improvement. Compared with the classifier before improvement, the degree of confusion is greatly reduced, especially for the easily confused classes 3, 4, and 6. Based on the confusion matrix, the precision, recall, and f1score of the recognition result can be obtained. As shown in Table 3, the recognition accuracy can reach about 0.84, which basically meets the requirements for lithofacies discrimination.
[0140] Table 3 Recognition results of multi-well data after improvement
[0141]
[0142] In this embodiment, using the conventional logging data of the deep reservoirs in the second member of the Sangonghe Formation in the Moxizhuang area and taking the core data as a constraint, the solution methods for the key problems faced in the process of lithofacies logging identification using machine learning methods are explored, and a set of Bayesian inversion prediction processes for lithofacies suitable for strongly heterogeneous reservoirs is established.
[0143] Beneficial effects: (1) The lithofacies change frequently in deep and strongly heterogeneous reservoirs, which poses challenges to traditional point-by-point recognition machine learning methods. The sedimentary prior constructed based on the Markov chain can better constrain the vertical distribution of lithofacies during the sedimentation process, making the predicted vertical distribution results of lithofacies conform to geological laws.
[0144] (2) Strongly heterogeneous reservoirs are characterized by a relatively large disparity in lithofacies proportions. This data imbalance problem causes well logging data to be unable to comprehensively reflect the characteristics of minority categories. By using the KmeamSMOTE and probability calibration methods to correct the likelihood function, the data imbalance problem in lithofacies well logging identification of strongly heterogeneous reservoirs can be solved.
[0145] (3) An improved Bayesian lithofacies inversion process was established and applied to lithofacies identification and prediction of multiple wells in the work area. The application of the fluvial-delta facies reservoirs in the Sangonghe Formation in the Moxizhuang area shows that, compared with traditional machine learning methods, the new method improves the identification accuracy by 20%, and can more accurately identify the lithofacies types of deep strongly heterogeneous reservoirs.
[0146] The above specific implementation manners have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs, characterized in that, the identification method includes: Step S1: Successively perform clustering, filtering, and oversampling using the KmeanSMOTE method; Step S2: Resample according to the number of core data samples in the study area with reference to the parameters in Step S1; Step S3: Construct a Markov transition matrix; Step S4: Perform Bayesian inversion; Step S5: Further improve the posterior distribution using probability calibration.
2. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, in Step S1: When using the KmeanSMOTE method to perform clustering successively, the k-means clustering method is used to cluster the input space into k clusters.
3. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, the filtering in Step S1 specifically includes: Select the clusters to be oversampled and determine how many samples to generate in each cluster, controlled by the imbalance rate IR and the sampling weight α os for control; The calculation formula of the imbalance rate IR is as follows, where c is the clustering cluster label, majorityCount is the majority class sample count of the cluster, minorityCount is the minority class sample count of the cluster, and the imbalance rate IR is the ratio of the majority class sample count to the minority class sample count of a certain cluster c; Sampling weight α os The calculation formula is as follows: α os = N rm / N M Where N rm is the number of samples in the minority class after resampling, and N M is the number of samples in the majority class after resampling. The sampling weight corresponds to the ratio of the number of samples in the minority class to the number of samples in the majority class after resampling; When performing multi-class classification, it is necessary to distinguish between the majority class and the minority class.
4. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, the oversampling specifically includes: Apply SMOTE in each selected cluster to increase the ratio α of minority samples to majority samples os ; SMOTE has the following process: First, select a random minority-class observation X i , among its k nearest minority-class neighbors, select the sample X zi ; Generate new samples through the following formula: x new = x i + λ × (x zi - x i ) where λ is a random weight in [0, 1], X i , X zi are the original samples of the minority class in the selected cluster.
5. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, Step S3: The specific steps for constructing a Markov transition matrix include: Construct a Markov transition matrix using the method of equal-spacing stratigraphic units.
6. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, after Step S3: Constructing a Markov transition matrix, it further includes: After constructing the Markov transition matrix P, randomly give an initial lithofacies distribution probability π(0), then the probability of different states at time t is π(t) = π(t - 1)P = …… = π(0)Pt. After long-term iteration n times, π(t) of different iteration processes is obtained; When n is large enough, π(t) will be very close to the true lithofacies distribution, thereby determining the prior probability distribution: where \(i\) and \(j\) are lithofacies categories, \(n\) is the number of iterations, the state space \(I=\{0,1,2,\cdots\}\) represents the total lithofacies categories, and \(\pi\) j is the probability of lithofacies \(j\) approximated by iteration.
7. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, Step S4: The specific steps for performing Bayesian inversion include: When performing Bayesian inversion, determine the prior probability according to the Markov chain, and the likelihood is assumed to be a Gaussian distribution, and its mean and variance are estimated according to the maximum likelihood method.
8. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that, Step S5: The specific steps for further improving the posterior distribution using probability calibration include: Further improve the posterior distribution using probability calibration, map the posterior probability output by the classifier to [0, 1] to obtain the calibrated probability, for the classifier output f of a given sample i , the calibrator predicts p(y i = 1|f i ).
9. The method for well logging identification of lithofacies in deep and strongly heterogeneous sandstone reservoirs according to claim 8, characterized in that, For the multi-classification problem in the classifier, the one-vs.-rest classifier strategy is adopted, fitting a classifier for each class and fitting it against all the remaining classes.
10. A method for identifying lithofacies of deep and strongly heterogeneous sandstone reservoirs according to claim 1, characterized in that The probability calibration in step S5 adopts isotonic regression. Given a predicted probability f i , the true probability y i is expressed as: y i = m(f i ) + ε; where m represents a monotonically increasing function, which is also the objective function to be fitted: The decision function is then corrected to Among them, the likelihood function is corrected using the function for correction: Since should be a monotonically increasing function, the above equation simplifies to: Among them, the weight w i is a strictly positive number, and the result is a piecewise linear function. The corrected likelihood function is considered to be able to reflect the data characteristic distribution after oversampling.