An asymptotically optimal lithology prediction method based on multi-threshold quantization output observations

By constructing a multi-threshold quantification lithology identification model and a weighted quasi-Newton algorithm for the information matrix, the problems of accuracy and interpretability in lithology identification are solved, achieving efficient and accurate lithology prediction, which is applicable to oilfield exploration and development.

CN115204277BActive Publication Date: 2025-12-02ACAD OF MATHEMATICS & SYSTEMS SCIENCE - CHINESE ACAD OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210736022.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-12-02
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

Existing machine learning-based lithology identification methods are insufficient in terms of accuracy and interpretability, making it difficult to meet the needs of actual industrial production. Furthermore, traditional well logging lithology identification methods are costly, slow, and have a significant impact.

Method used

An asymptotically optimal lithology prediction method based on multi-threshold quantization output observation is adopted. By constructing a weighted quasi-Newton algorithm based on the information matrix and combining it with well logging curve data, a multi-threshold quantization lithology identification model is built to achieve interpretability and efficiency in lithology identification.

Benefits of technology

It achieves improved accuracy and interpretability of lithology identification while saving manpower and resources, with asymptotically optimal convergence and a recall rate of over 85%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115204277B_ABST
    Figure CN115204277B_ABST
Patent Text Reader

Abstract

This invention proposes an asymptotically optimal lithology prediction method based on multi-threshold quantized output observations, comprising the following steps: 1. Constructing input samples for the model based on well logging curve data and corresponding lithology information, and dividing the total samples into a training set and a test set; 2. Constructing and executing an asymptotically optimal identification algorithm for the multi-threshold quantized lithology identification model based on the input samples and corresponding lithology category information in the training set; 3. Predicting lithology using the lithology identification model based on the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix, and determining which lithology category a sample in the test set belongs to. This invention is the first to propose a multi-threshold quantized lithology identification model to solve the problem of well logging lithology identification, achieving interpretability of the lithology identification model while saving manpower and resources. The weighted quasi-Newton algorithm based on the information matrix proposed in this invention has asymptotic optimality, meaning its convergence performance is superior to previous quantization identification algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to problems related to lithology identification in geophysical exploration, specifically to an asymptotically optimal lithology prediction method based on multi-threshold quantized output observations. Background Technology

[0002] Accurate lithology identification results can provide a reliable basis for oilfield exploration and development, playing a significant role in finding oil and gas resources and assessing resource reserves.

[0003] Generally, data sources for obtaining subsurface lithology information mainly fall into two categories: core data and well logging curves. Core data is collected directly during drilling and analyzed by geologists to obtain relatively accurate lithology information; however, its high cost makes it difficult to widely apply in actual oilfield development. Well logging curves, on the other hand, have advantages such as high vertical resolution, good continuity, and convenient data acquisition, and are often used in lithology identification research. Traditional well logging lithology identification methods... [1] Low accuracy, slow speed, and significant influence from human factors have made the use of computers for faster and more accurate lithology identification a hot research topic for well logging researchers both domestically and internationally. [2-3] However, existing machine learning-based lithology identification methods prioritize high accuracy, neglecting the interpretability of the lithology identification model. Interpretability is crucial for real-world industrial production models where risk is a significant concern. Therefore, this invention presents a lithology prediction method based on a multi-threshold quantization lithology identification model.

[0004] The references are as follows:

[0005] [1] Honarkhah M, and Caers J, Direct pattern-based simulation of non-stationary geostatistical models, Mathematical Geosciences, 2012, 44(6): 651-672.

[0006] [2] Kang Qiankun, Lu Laijun, Application of Random Forest Algorithm in Well Logging Lithology Classification, World Geology, 2020, 39(02):398-405.

[0007] [3] Ma Longfei, Xiao Hanmin, Tao Jingwei, Zhang Fan, Luo Yongcheng, Zhang Haiqin, Research and application of lithological classification based on deep learning, Science, Technology and Engineering, 2022, 22(07):2609-2617. Summary of the Invention

[0008] To address the aforementioned problems, this invention provides an asymptotically optimal lithology prediction method based on multi-threshold quantization output observations, which can effectively predict lithology categories. To solve the above problems, this invention adopts the following technical solution:

[0009] An asymptotically optimal lithology prediction method based on multi-threshold quantized output observations includes the following steps:

[0010] Step 1: Construct the model input sample based on well logging curve data and corresponding lithological information. Specifically, this includes constructing the model input sample based on well logging curve data and corresponding lithological information: forming an 8-dimensional sample feature vector from data such as depth, sonic logging, caliper logging, compensated neutron logging, gamma logging, spontaneous potential logging, 2.5m bottom gradient resistivity logging, and density logging, and dividing the total sample into a training set and a test set.

[0011] Step 2: Based on the input samples and corresponding lithology category information in the training set, construct and execute an asymptotically optimal identification algorithm for the multi-threshold quantized lithology identification model. Specifically, this includes: first, constructing a lithology identification model based on the input samples and corresponding lithology information in the training set; then, designing a weighted quasi-Newton algorithm based on the information matrix to identify unknown parameters in the lithology identification model; and evaluating the convergence characteristics and optimality of the algorithm.

[0012] Step 3: Based on the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix, use the lithology identification model to predict the lithology and determine which lithology category a sample in the test set belongs to.

[0013] Advantages and beneficial effects of the present invention:

[0014] 1. This invention is the first to propose a multi-threshold quantitative lithology identification model to solve the problem of well logging lithology identification. While saving manpower and material resources, it achieves the interpretability of the lithology identification model.

[0015] 2. Compared with previous parameter identification algorithms based on quantization output measurement, the weighted quasi-Newton algorithm based on information matrix proposed in this invention has asymptotic optimality, that is, its convergence effect is better than that of previous quantization identification algorithms. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention.

[0017] Figure 2 This is a block diagram of the lithology identification model of the present invention.

[0018] Figure 3 This is a flowchart of the weighted quasi-Newton algorithm based on the information matrix in this invention.

[0019] Figure 4 This is a structural diagram of the lithology prediction model of this invention. Detailed Implementation

[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the protection scope of the present invention.

[0021] To make the technical solution of the asymptotically optimal lithology prediction method based on multi-threshold quantization output observation proposed in this invention clearer and more complete, the specific steps are now described (the flowchart of this invention is shown below). Figure 1 (As shown) Detailed description is as follows:

[0022] Step 1: Based on the well logging curve data and corresponding lithology information, construct the input samples for the model, and divide the total samples into training set and test set.

[0023] Based on well logging data and corresponding lithological information, an 8-dimensional sample input is constructed using depth (DEPTH), acoustic logging (AC), caliper logging (CAL), compensated neutron logging (CNL), gamma logging (GR), spontaneous potential logging (SP), 2.5m bottom gradient resistivity logging (R25), and density logging (DEN). This sample is denoted as... in, Represents the n-dimensional real number field. The sample labels q∈{0,1,…,m} represent lithology categories, including: mudstone, coarse sandstone, fine sandstone, and conglomerate. Assume there are a total of N samples, and the training and test sets are separated according to a 7:3 ratio. Let N be the number of training set samples. s There are N test set samples. t strip.

[0024] Step 2: Based on the input samples and corresponding lithology category information in the training set, construct and execute the asymptotic optimal identification algorithm of the multi-threshold quantization lithology identification model.

[0025] First, a lithology identification model with unknown parameters is constructed based on the input samples and corresponding lithology information in the training set (the block diagram of the lithology identification model of this invention is shown below). Figure 2 As shown), the details are as follows:

[0026]

[0027] Where, φ k θ represents the n-dimensional sample features of the k-th lithological sample, where T represents the transpose of the vector; θ is the n-dimensional unknown steady parameter vector of the lithology identification model; d kThis is white noise from the k-th lithological sample. Errors are unavoidable during data acquisition and processing; therefore, the addition of random noise is necessary. Based on the central limit theorem, we assume the noise follows a pattern with a mean of 0 and a variance of σ. 2 The normal distribution of y is given by distribution function F(·) and density function f(·), respectively; k This is the output of the lithology identification model, which can only classify lithology into m+1 lithology categories based on m lithology classification thresholds, where the lithology classification thresholds are C1, C2, ..., C... m And the threshold satisfies -∞ <C1<C2<…<C m <∞. The quantization process output by this lithology identification model can be represented as follows:

[0028]

[0029] In fact, q k This represents the lithology category of the k-th lithology sample.

[0030] The asymptotic optimal identification problem of this lithology identification model needs to be based on the quantization model (1)-(2), and an asymptotic optimal identification algorithm based on the input samples and corresponding lithology category information in the training set should be designed—a weighted quasi-Newton algorithm based on the information matrix, and the convergence characteristics and optimality of the algorithm should be evaluated.

[0031] like Figure 3 As shown, a weighted quasi-Newton algorithm based on the information matrix is ​​constructed:

[0032] Estimating initial values ​​for any n-dimensional parameter and n-dimensional positive definite matrix The k-th (k≥1) iteration of the algorithm is performed as follows:

[0033] a) Parameter estimates calculated based on the (k-1)th iteration Adaptively calculate the quantization output weight coefficients respectively Covariance weighting coefficient

[0034]

[0035] For i = 1, ..., m+1, and These are the probability density and probability, respectively, of the predicted k-th sample belonging to the i-th lithological category. and The noise density function f(·), the noise distribution function F(·), and the lithology classification threshold C can be used as the basis for classification. i The sample characteristics φ of the k-th lithological sample k The parameter estimates corresponding to the (k-1)th sample The calculations are as follows:

[0036]

[0037] Where, C0=-∞,C m+1 =+∞.

[0038] b) Based on the quantization output weight coefficients calculated in step a). Calculate the quantized output q k Weighted transformations k :

[0039]

[0040] The purpose of formula (4) is to adjust the quantization output weights so that the algorithm can achieve better recognition results.

[0041] c) Based on the quantization output weight coefficients calculated in step a). Covariance weighting coefficient The quantized output weighted transformation s calculated in step b) k Calculate the estimates of unknown parameters in the lithology identification model:

[0042]

[0043] Formula (5) is used to utilize the sample characteristics φ of the kth lithological sample. k The parameter estimates given by the (k-1)th lithological sample The covariance matrix calculated from the sample characteristics of the first k-1 lithological samples The weighted transformation s of the quantization output corresponding to the k-th lithology sample k and the weighting coefficients given in step a). and estimated probability The calculated weighted transformation estimate To calculate estimates of unknown parameters in the lithology identification model.

[0044] Formula (6) is used to calculate the covariance weights given in step a). The sample characteristics φ of the k-th lithological sample k and the covariance matrix calculated from the sample characteristics of the first k-1 lithological samples. To calculate the covariance matrix formed by the sample characteristics of the first k lithological samples.

[0045] d) Evaluate the convergence properties and optimality of the weighted quasi-Newton algorithm based on the information matrix:

[0046] In fact, the weighted quasi-Newton algorithm based on the information matrix (as shown in formulas (3)-(6)) can achieve the following property: if the characteristic sequence {φ} of the lithological sample k} is a bounded, continuous incentive, that is And there exists a positive integer h such that for all k, (I n (Referring to the n-dimensional identity matrix). Therefore, the weighted quasi-Newton algorithm (3)-(6) based on the information matrix is ​​almost everywhere convergent, mean-square convergent, and convergent to higher-order matrices, and the mean-square convergence speed of the algorithm is...

[0047]

[0048] in Representing mathematical expectation, O is an asymptotic symbol indicating that the two variables are of the same order of magnitude. That is, b = O(a) means... Less than or equal to a non-zero constant.

[0049] Furthermore, the weighted quasi-Newton type algorithms (3)-(6) based on the information matrix are also asymptotically optimal, that is, the covariance matrix of the algorithm's estimation error asymptotically approaches the Cramero lower bound of the lithology identification model, i.e.:

[0050]

[0051] in, This is the lower bound of Clamer-Rhodes for lithological identification models (1)-(2), where and

[0052] In fact, the Cramerlow lower bound represents the infimum of the parameter estimation error covariance. Therefore, the weighted quasi-Newton algorithm based on the information matrix that can reach this lower bound is the asymptotically optimal identification algorithm.

[0053] Step 3: Based on the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix, use the lithology identification model to predict the lithology and determine which lithology category a sample in the test set belongs to.

[0054] Based on the sample information in the training set and the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix in step 2. Unknown parameters in the lithology identification model can be determined. Where N s This refers to the number of samples in the training set. Then, based on the following lithology identification model...

[0055]

[0056] Calculate the lithology category corresponding to all sample data in the test set. For example... Figure 4The diagram shown is a structural diagram of the lithology prediction model of this invention.

[0057] According to one application embodiment of the present invention, the lithology identification method of the present invention is used to perform lithology identification work on a certain well in an oil field. The results show that on an imbalanced dataset, the lithology identification method of the present invention has excellent prediction results, with a recall rate greater than 85% for each lithology category, as detailed below:

[0058] 1. Data preparation and preprocessing:

[0059] A total of 10,110 lithology samples were used from well logging. The sample numbers for mudstone, coarse sandstone, fine sandstone, and conglomerate were 4,800, 2,900, 1,300, and 1,110 respectively. These 10,110 samples were divided into a training set and a test set according to a 7:3 ratio, with the four lithologies having the same proportion in both sets to ensure lithology balance. The training set then contained N sample data points. s = 7077 records, N test set sample data t =3033 samples. Each feature of the samples has a specific physical meaning and different orders of magnitude. To avoid the influence of data format on the establishment of the lithology model, the same feature of all samples is normalized, and the value is normalized to the range [0,1], completing the normalization process for all 8-dimensional sample feature values. In addition, the sample data are labeled according to lithology, with mudstone set to 0, coarse sandstone to 1, fine sandstone to 2, and conglomerate to 3.

[0060] 2. Identify unknown parameters in the lithology identification model using a weighted quasi-Newton algorithm based on the information matrix:

[0061] Based on the sample features and lithological sample category information of the training set, a quantitative lithology identification model is constructed as follows.

[0062]

[0063] Where φ k θ is the 8-dimensional sample feature of the k-th (k=1,…,7077) lithological sample in the training set, θ is the 8-dimensional unknown parameter vector of the lithological identification model, and d k q represents the noise from the k-th lithological sample, following a normal distribution with a mean of 0 and a variance of 0.35. k ∈{0,1,2,3} represents the lithology category of the k-th lithology sample.

[0064] Subsequently, based on the asymptotically optimal identification algorithm of the multi-threshold quantization lithology identification model—the weighted quasi-Newton algorithm based on the information matrix (3)-(6)—the sample information in the training set (i.e., the sample feature vector φ) was used. k and sample lithology category qk (k = 1, ..., 7077)), select initial values ​​for parameter estimation. and 8-dimensional positive definite matrix

[0065]

[0066] As shown in equations (3)-(6), the weighted quasi-Newton algorithm based on the information matrix is ​​iteratively executed to obtain estimates of the unknown parameters in the lithology identification model:

[0067]

[0068] 3. Based on the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix. A lithology identification model is used to predict the lithology of samples in the test set, determining which lithology category a given sample belongs to. The specific calculation is as follows:

[0069]

[0070] Therefore, the sample features φ of the test set are utilized. k (k=1,…,3033) and the parameter estimates obtained in step 2 The lithology of all sample data in the test set is calculated using the formula above. Table 1 shows the results of the lithology prediction of this invention.

[0071] Table 1

[0072]

Claims

1. A method for asymptotically optimal lithology prediction based on multi-threshold quantization output observations, characterized in that: Step 1: Construct the input samples for the model based on the well logging curve data and corresponding lithological information; Step 2: Based on the input samples and corresponding lithology category information in the training set, construct and execute the asymptotic optimal identification algorithm for the multi-threshold quantization lithology identification model; design a weighted quasi-Newton algorithm based on the information matrix to identify unknown parameters in the lithology identification model, and evaluate the convergence characteristics and optimality of the algorithm. Step 3: Based on the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix, use the lithology identification model to predict the lithology and determine which lithology category a sample in the test set belongs to. The formula for the lithology identification model is as follows: Where, φ k θ represents the sample characteristics of the k-th lithological sample, where T represents transpose; θ is the unknown parameter of the lithology identification model; d k This is white noise from the k-th lithological sample, assuming the noise follows a mean of 0 and a variance of σ. 2 The normal distribution of y is given by distribution function F(·) and density function f(·), respectively; k This is the output of the lithology identification model, which can classify lithology into m+1 lithology categories based on m lithology classification thresholds, where the lithology classification thresholds are C1, C2, ..., C... m And the threshold satisfies -∞ <C1<C2<…<C m <∞; The quantization process output by this lithology identification model is represented as follows: q k Represents the lithological category of the k-th lithological sample; The weighted quasi-Newton algorithm based on the information matrix is ​​constructed as follows: Estimate initial values ​​for any parameter and positive definite matrix The k-th iteration is performed as follows, where k≥1: a) Parameter estimates calculated based on the (k-1)th iteration Adaptively calculate the quantization output weight coefficients respectively Covariance weighting coefficient i = 1, ..., m+1; For i = 1, ..., m+1, and These are the probability density and probability, respectively, of the predicted k-th sample being the i-th lithology category; Where, C0=-∞,C m+1 =+∞; C i This is the threshold for lithological classification; b) Calculate the quantized output q k Weighted transformations k : c) Calculate estimates of unknown parameters in the lithology identification model: Where, φ k The sample characteristics of the k-th lithological sample are as follows: The parameter estimates given for the (k-1)th lithological sample are as follows: The covariance matrix calculated for the sample characteristics of the first k-1 lithological samples. This is a weighted transformation estimate. These are estimated values ​​of unknown parameters in the lithology identification model. For covariance weights, The covariance matrix is ​​formed by the sample characteristics of the first k lithological samples; d) Evaluate the convergence properties and optimality of the weighted quasi-Newton algorithm based on the information matrix: If the characteristic sequence of the lithological sample is {φ k } is a bounded, continuous incentive, that is And there exists a positive integer h such that for all k, I n Refers to the identity matrix; the mean square convergence rate of the weighted quasi-Newton algorithm based on the information matrix is: in, Representing mathematical expectation, O is an asymptotic symbol indicating that the two variables are of the same order of magnitude. That is, b = O(a) means... Less than or equal to a non-zero constant; The weighted quasi-Newton algorithm based on the information matrix is ​​asymptotically optimal, that is: in, This is the lower bound of Cramer-Rao for formulas (1)-(2), where and 2. The asymptotically optimal lithology prediction method based on multi-threshold quantization output observation as described in claim 1, characterized in that: In step 3: based on the sample information in the training set and the parameter estimates given by the weighted quasi-Newton algorithm based on the information matrix in step 2. Determine the unknown parameters in the lithology identification model Where N s This represents the number of samples in the training set; then, based on the following lithology identification model, the lithology category corresponding to all sample data in the test set is calculated.

Citation Information

Patent Citations

  • Lithology identification method and system

    CN112766309A

  • Complex lithology identification method and identification system, electronic equipment and storage medium

    CN114528746A