Natural gamma logging lithology real-time identification method based on Bayesian theory

By using a dual-mode prior knowledge base based on Bayesian theory and a Naive Bayes classifier, the identification problem of natural gamma logging under complex geological conditions is solved, realizing quantitative lithology probability output and powerful decision support capabilities, which is suitable for the lightweight requirements of logging while drilling.

CN121454629APending Publication Date: 2026-02-03CHINA COAL TECH & ENG GRP CHONGQING RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511656713.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing natural gamma logging lithology identification methods lack adaptability and reliability under complex geological conditions, lack confidence assessment of identification results, cannot effectively integrate prior geological knowledge, and experience performance degradation under single gamma input scenarios.

Method used

A dual-mode prior knowledge base is constructed based on Bayesian theory. By combining statistical parameters from adjacent wells and geological rule parameters, a posterior probability of lithology is calculated using a Naive Bayes classifier, and quantitative lithology probability results are output, thus integrating prior knowledge with real-time data.

Benefits of technology

It improves the accuracy and geological consistency of lithology identification, provides quantitative confidence assessment, and meets the real-time and lightweight requirements of logging while drilling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121454629A_ABST
    Figure CN121454629A_ABST
Patent Text Reader

Abstract

The invention provides a natural gamma logging lithology real-time identification method based on the Bayesian theory. According to the natural gamma logging lithology real-time identification method based on the Bayesian theory, the adjacent well statistical law and geological rule priori knowledge can be fused, the lithology posterior probability is calculated in real time based on the Bayesian theory, the identification accuracy and the geological conformity under the complex stratum condition are remarkably improved, and the identification accuracy and the geological conformity are improved. And meanwhile, probabilistic output with clear statistical significance is provided, and the reliability and risk management and control capability of geosteering decision are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geological steered drilling and formation evaluation technology, and in particular to a real-time lithology identification method based on Bayesian theory using natural gamma logging. Background Technology

[0002] Natural gamma-ray logging (NVR) is a core tool for formation assessment during drilling and is widely used in coal mining and geological engineering. This technology identifies lithology by measuring the decay intensity of radioactive elements in the formation. Its technical system covers the entire process from data acquisition to decision support, including key aspects such as modeling the correlation between gamma values ​​and lithological characteristics, real-time data processing, and multi-source information fusion. With the development of NVR technology, existing methods mainly rely on gamma thresholding or deep learning models (such as LSTM-CRF) for lithological classification. Gamma thresholding achieves rapid judgment by setting fixed numerical limits, while deep learning methods utilize the sequential characteristics of multiple logging curves (such as sonic transit time and resistivity) for nonlinear fitting. Both attempt to solve the lithology identification problem through different technical paths. However, existing technologies have systematic defects in single gamma input scenarios, making it difficult to meet engineering requirements in terms of adaptability and reliability under complex geological conditions.

[0003] Existing gamma thresholding methods directly map lithology to gamma values ​​using linear physical assumptions, failing to consider the heterogeneity of formations and the complexity of radioactive element distribution. This can lead to technical problems such as misclassifying radioactive sandstone as mudstone and blurring the boundaries of argillaceous coal seams. Furthermore, these methods output deterministic conclusions, lacking confidence assessment of the identification results, making it difficult to quantify risks in geologically guided decision-making. While deep learning-based LSTM-CRF models improve identification accuracy through sequence labeling, they heavily rely on multiple parameter inputs (such as density and acoustic waves), resulting in significant performance degradation in drilling scenarios with only natural gamma data. Moreover, these models lack a geologically logical mathematical framework to integrate data from neighboring wells or expert experience, leading to insufficient robustness in data-sparse regions. Since existing technologies have not established a collaborative system of multi-source prior knowledge and real-time observation data, the geological consistency and decision support capabilities of their outputs are fundamentally limited. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the related art.

[0005] Therefore, the first objective of this invention is to propose a real-time lithology identification method based on Bayesian theory in natural gamma logging.

[0006] The second objective of this invention is to propose a real-time lithology identification device for natural gamma logging based on Bayesian theory.

[0007] The third objective of this invention is to provide an electronic device.

[0008] The fourth objective of this invention is to provide a computer-readable storage medium.

[0009] The fifth objective of this invention is to provide a computer program product.

[0010] To achieve the above objectives, a first aspect of the present invention proposes a real-time lithology identification method based on Bayesian theory in natural gamma logging, comprising: A dual-mode prior knowledge base is constructed, which includes statistical parameters of adjacent wells and geological rule parameters. The statistical parameters of adjacent wells are calculated based on the data of adjacent wells in the target work area, and the geological rule parameters are generated by parameterization through geological literature and expert experience. Real-time acquisition of natural gamma logging data while drilling and preprocessing of the data include converting the gamma data into a series of equal depth intervals, using a moving average filter for noise suppression, and verifying the validity of the data according to a preset physical range. Based on the dual-mode prior knowledge base and the preprocessed gamma data, the posterior probability distribution of lithology corresponding to the adjacent well statistical model and the geological rule model are calculated respectively. Then, the posterior probability distribution of the adjacent well statistical model and the geological rule model are weighted and averaged according to the preset fusion weight to generate the fused lithology probabilistic identification result.

[0011] Optionally, the construction of a bimodal prior knowledge base including adjacent well statistical parameters and geological rule parameters further includes: The adjacent well statistical parameters are obtained by statistically analyzing the total number of occurrences of each lithological category in the adjacent well data of the target work area. Total number of points for all lithologies Calculate the prior probability using the ratio ; The geological rule parameters are defined by geological literature and expert experience as typical gamma response ranges. Mapped to the mean of a Gaussian distribution and standard deviation ,in The midpoint of the interval, For the interval span times.

[0012] Optionally, the real-time acquisition and preprocessing of natural gamma-ray logging data during drilling also includes: A moving average filter is used to perform real-time smoothing of gamma data. The window length of the moving average filter is... The calculation formula is: ; For missing gamma values, nearest neighbor interpolation is used, and the interpolation formula is as follows: .

[0013] Optionally, the step of calculating the posterior probability distribution of lithology corresponding to the adjacent well statistical model and the geological rule model based on the dual-mode prior knowledge base and the preprocessed gamma data further includes: The posterior probability calculation of the adjacent well statistical model uses the likelihood function of a Gaussian distribution. ; The posterior probability calculation of the geological rule model uses a Gaussian likelihood function. .

[0014] Optionally, the step of weighting the posterior probability distributions of the adjacent well statistical model and the geological rule model according to preset fusion weights further includes: The fusion weight and Satisfy normalization conditions ; The weighted average formula is as follows: .

[0015] Optionally, it also includes: dynamically updating the dual-mode prior knowledge base based on real-time recognition results, specifically including: The identification results of the current depth point are compared with the data from neighboring wells to calculate the updated prior probability. ,in It is a smoothing factor; Refit the likelihood function parameters based on the updated prior probabilities. and And adjust the fusion weights and The value of .

[0016] To achieve the above objectives, a second aspect of the present invention provides a real-time lithology identification device for natural gamma logging based on Bayesian theory, comprising: The dual-mode prior knowledge base construction module is used to construct a dual-mode prior knowledge base containing adjacent well statistical parameters and geological rule parameters. The adjacent well statistical parameters are calculated based on adjacent well data of the target work area, and the geological rule parameters are generated parameterized through geological literature and expert experience. The gamma data preprocessing module is used to acquire and preprocess natural gamma logging data while drilling in real time. The preprocessing includes converting the gamma data into an equal depth interval sequence, using a moving average filter for noise suppression, and verifying the validity of the data according to a preset physical range. The lithology posterior probability and probability fusion identification calculation module is used to calculate the lithology posterior probability distribution corresponding to the adjacent well statistical model and the geological rule model based on the dual-mode prior knowledge base and the preprocessed gamma data, and to perform a weighted average of the posterior probability distribution of the adjacent well statistical model and the geological rule model according to the preset fusion weights to generate the fused lithology probabilistic identification result.

[0017] To achieve the above objectives, a third aspect of the present invention provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of the first aspects.

[0018] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of the first aspects.

[0019] To achieve the above objectives, a fifth aspect of the present invention provides a computer program product that, when executed by a processor, implements the method described in any one of the first aspects.

[0020] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a real-time lithology identification method based on Bayesian theory for natural gamma logging provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a real-time lithology identification device based on Bayesian theory for natural gamma logging, provided in an embodiment of the present invention. Detailed Implementation

[0022] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0023] Logging while drilling (LoWWD) is a core technology for acquiring real-time formation information during drilling, and it is crucial for fields such as coal mining and geological engineering. Among its applications, accurate real-time lithology identification is fundamental for geological steering, reservoir evaluation, drilling parameter optimization, and safety risk early warning.

[0024] Among various logging-while-drilling (LOD) methods, natural gamma-ray logging (NGR) has become the most basic and widely used real-time formation assessment tool due to its sensitivity to lithology, high instrument reliability, and relatively low cost. NGR distinguishes different lithologies by measuring the decay intensity of radioactive elements in the formation. Typically, mudstone exhibits high gamma values ​​due to its high potassium content, while sandstone, carbonate rocks, and coal seams show low gamma values ​​due to their low radioactive content.

[0025] Currently, the most mainstream real-time lithology identification method in practical applications of drilling geological steering and formation evaluation is the gamma threshold method. This method classifies real-time acquired gamma data by pre-setting one or more gamma value thresholds. For example, when the gamma value is below a certain threshold, it is identified as a target reservoir (such as a coal seam); when it is above another threshold, it is identified as a non-reservoir (such as mudstone or roof and floor). This method is simple, intuitive, and fast, meeting basic real-time requirements. However, the gamma threshold method has inherent and insurmountable drawbacks: (1) Challenges of geological complexity: Geological conditions are far more complex than simple threshold models. For example, “radioactive sandstone” containing feldspar or heavy minerals may exhibit high gamma values ​​and thus be misclassified as mudstone. Conversely, coal seams or argillaceous sandstones with high mud content may have gamma values ​​that differ significantly from those of pure coal seams or sandstones, leading to blurred identification boundaries. In addition, the physical properties of rocks vary significantly in different regions and strata, making it difficult for fixed threshold models to be universally applicable.

[0026] (2) Low information utilization: This method is a "memoryless" instantaneous judgment that cannot systematically and automatically integrate valuable prior geological knowledge, such as regional geological patterns, seismic data interpretation results, and logging and well logging data from adjacent wells. Each drilling operation is like an independent exploration, and historical data and expert geological knowledge cannot be effectively utilized in the real-time identification model, resulting in low information utilization efficiency.

[0027] (3) The identification results are deterministic, lacking uncertainty assessment: The threshold method provides a deterministic "yes" or "no" conclusion, which cannot quantify the credibility of the identification results. In geologically guided decision-making, a lithology judgment with 80% confidence is obviously more valuable than a judgment with only 50% confidence. The lack of probabilistic information makes it difficult for decision-makers to assess risks, especially at critical moments such as complex geological structures and thin reservoir drilling.

[0028] To overcome the shortcomings of traditional gamma thresholding methods, existing technologies have developed intelligent lithology identification methods based on artificial intelligence algorithms such as deep learning. A typical example of an implementation scheme most similar to this invention is the patent "An Intelligent Identification Method for Formation Lithology Based on Well Logging Data" (Application Publication No.: CN 119739992 A) published by Beijing Zhongmei Mining Engineering Co., Ltd. The core technology of this patent is to construct a deep learning model (LSTM-CRF) combining a Long Short-Term Memory (LSTM) network and a Conditional Random Field (CRF). Its implementation is as follows: First, multiple well logging curves, including natural gamma, sonic transit time, resistivity, and density, are collected as input features for the model; then, the well logging data sequence is input into the LSTM network to capture the long-term dependencies of the data in the depth sequence; finally, the output of the LSTM is fed into the CRF layer for sequence labeling to consider the transfer relationships between adjacent lithologies, ultimately outputting the globally optimal lithology sequence. Compared to the traditional thresholding method, this scheme fully utilizes sequence information. Through LSTM and CRF models, it effectively leverages the correlation information between upper and lower depth points on the logging curve, making the identification results more consistent with geological patterns in terms of formation contact relationships. Simultaneously, this method achieves higher identification accuracy. Leveraging the powerful nonlinear fitting capabilities of deep learning models, it can automatically learn and extract complex features from multi-dimensional logging data, thus achieving higher identification accuracy than traditional methods under data-rich conditions. However, existing technologies, represented by LSTM-CRF models, still have the following fundamental limitations when addressing the specific problems addressed in this invention, especially in real-time drilling applications: (1) It heavily relies on multiple parameter inputs, which limits its real-time application: The effectiveness of this scheme is based on obtaining multiple logging curves such as natural gamma, density, sonic, and resistivity. However, in cost-sensitive or technology-limited logging-while-drilling environments, only natural gamma is often the core data available in real time. Under the condition of a single gamma input, the performance of such complex models will be greatly reduced, and their applicability will be greatly limited.

[0029] (2) Inability to effectively integrate prior geological knowledge: The LSTM-CRF model is a purely data-driven model, and all its "knowledge" is learned from labeled training data. It lacks a direct and logical mathematical framework to integrate prior knowledge provided by geological experts or neighboring well data. The Bayesian theory on which this invention is based can organically integrate this prior knowledge (prior probability) with real-time observation data (likelihood), making the model more consistent with regional geological characteristics.

[0030] (3) Unclear probability output and insufficient decision support capability: Although CRF itself is a probabilistic graphical model, the final output of the entire LSTM-CRF scheme is an optimal deterministic lithology sequence. It cannot intuitively answer the question that decision-makers care about most: "At the current depth, what is the probability that the lithology is a coal seam? What is the probability that it is mudstone?" In contrast, this invention is based on Bayesian theory, and its natural output is the posterior probability of each lithology category, which can provide a quantitative and statistically significant confidence assessment for geological guidance, and has a stronger decision support capability.

[0031] In summary, while existing technologies, such as CN119739992A, utilize advanced deep learning algorithms, their approach is multi-parameter and purely data-driven, making it difficult to effectively integrate prior geological knowledge. Furthermore, they suffer from limitations in applicability and capabilities in real-time logging-while-drilling scenarios where data types are limited and decision-making information requirements are high. Therefore, there is an urgent need in this field for a novel real-time lithology identification method specifically designed for gamma-ray logging-while-drilling, capable of integrating prior knowledge within a probabilistic framework and providing explicit probabilistic results.

[0032] The core technology of the traditional gamma thresholding method is based on an oversimplified linear physical assumption: that there is a simple, fixed linear correspondence between lithology and gamma value. This assumption contradicts the complex heterogeneity of strata. Therefore, when encountering complex strata where the gamma response does not match the conventional lithology (such as potassium-bearing feldspar sandstone or peat-bearing seams), the method inevitably leads to logical confusion, resulting in frequent misjudgments and low accuracy in lithology identification. Since this method only performs "greater than / less than" numerical comparisons, its output is essentially a deterministic Boolean value (yes / no), fundamentally failing to provide the probabilistic information or confidence assessment necessary for decision-making. Because this method is a memoryless, instantaneous judgment mechanism, it logically excludes the possibility of integrating historical data, information from adjacent wells, or prior expert knowledge, thus wasting valuable geological information and rendering the model's intelligence and adaptability meaningless.

[0033] For deep learning methods represented by CN 119739992 A patent, the core technology is a pure data-driven "black box" learning paradigm. The performance of the model is highly dependent on a large-scale, multi-dimensional, labeled training dataset, and its design intention is to pursue the global optimal solution of sequence labeling. The limitations of such methods are as follows: the direct consequences (for real-time drilling scenarios): (1) All the knowledge of the model is learned from the data. There is no logical port for receiving external knowledge. Therefore, in terms of framework, it is impossible to effectively and systematically integrate "prior knowledge" with clear geological significance (such as experts judging the probability of a certain lithology appearing in this area) with real-time data, resulting in insufficient predictive robustness of the model when facing newly encountered strata or sparse data areas. (2) Its CRF core objective is to decode the entire lithology sequence with the highest probability, rather than to evaluate the lithology distribution at a single depth point. Therefore, it cannot directly provide confidence information, resulting in insufficient support for real-time geological-guided risk assessment.

[0034] In view of this, in order to solve the fundamental defects in the existing technology, the purpose of this invention is to provide a real-time lithology identification method based on Bayesian theory in natural gamma logging. This method aims to overcome the oversimplification of the traditional threshold method and the excessive reliance on data in the deep learning method, and to create a new paradigm of intelligent lithology identification that is particularly suitable for real-time drilling scenarios, lightweight, capable of combining prior information and having strong decision support capabilities. The main contributions are as follows: (1) Establishing a fusion framework of prior knowledge and real-time data: Using the mathematical framework of Bayesian theory, prior knowledge (represented as prior probability) that carries regional geological laws, neighboring well information and expert experience is organically integrated with real-time acquired gamma logging data (represented as likelihood probability), so that the identification model can not only "see" the current data, but also "learn from" historical experience, which greatly improves the accuracy and geological consistency of the identification results. (2) Providing quantitative probabilistic results with clear geological significance: Overcoming the defects of the existing technology in outputting single results or unclear probabilistic significance, this invention aims to output the posterior probability of the current depth point belonging to each candidate lithology. This probabilistic result can provide geological steering engineers with a clear confidence assessment and provide strong data support for risk decision-making. (3) Ensure the real-time performance and lightweight nature of the method: The proposed method should have low computational complexity in its mathematical model and be able to be embedded into the well site processing system or drilling instrument to meet the stringent requirements of logging-while-drilling for real-time data processing.

[0035] Figure 1 This is a flowchart illustrating a real-time lithology identification method based on Bayesian theory for natural gamma logging, provided as an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: S1. Construct a dual-mode prior knowledge base containing statistical parameters of adjacent wells and geological rule parameters. The statistical parameters of adjacent wells are calculated based on the data of adjacent wells in the target work area, and the geological rule parameters are generated by parameterization through geological literature and expert experience.

[0036] This step, completed before drilling operations begin, aims to construct a multi-layered, adaptive prior knowledge base. This knowledge base not only encapsulates regional geological patterns but also incorporates universal geological common sense, real-time wellhead information, and the inherent continuity of stratigraphic sequences, thereby providing a dynamic, accurate, and robust prior foundation for the Bayesian model.

[0037] This invention proposes to construct a system incorporating prior knowledge from the following aspects: (1) Prior based on regional neighbor well statistics.

[0038] This step is completed offline before drilling operations begin. Its core purpose is to systematically and quantitatively transform existing, authoritative geological knowledge and logging data within the target work area into a set of probability parameters that can be directly used by the Bayesian model—that is, a priori knowledge base. This knowledge base is the prior foundation and source of knowledge for the intelligent recognition achieved in this invention. Specifically, it includes the following sub-steps: Step 1: Dataset creation and lithology category definition.

[0039] First, collect a comprehensive geological and logging dataset of one or more adjacent wells in the target work area or a neighboring area with a similar geological background. This dataset must contain continuous natural gamma logging curves and authoritative lithology labeling data. Then, based on the geological task and the lithological composition of the target interval, define a complete and mutually exclusive set of candidate lithology categories L, denoted as:

[0040] in, This represents the total number of lithology categories to be identified. For example, in a coal seam exploration scenario, this set could be defined as... .

[0041] Step 2, Prior Probability The calculation.

[0042] Prior probability This represents the determination of lithology at any depth point as a category based on regional geological patterns, without any prior knowledge of the target strata. The inherent possibility. It is a mathematical quantification of the geological understanding of "what lithology is more common in the region". The calculation method is as follows: First, statistical analysis was performed on authoritative lithology label data to obtain the lithology category. The total number of occurrences in all adjacent wells is denoted as . The total number of points for all lithologies is denoted as So, lithological categories Prior probability The calculation formula is:

[0043] in, The prior probabilities of all lithological categories constitute a prior probability distribution, and must satisfy the probability completeness condition:

[0044] Step 3: Modeling and parameterizing the likelihood function.

[0045] Likelihood function The description states that "in determining the current lithology as..." Under the premise of [specific conditions], what probability distribution will its natural gamma logging value (GR) exhibit? It establishes a causal relationship model from "lithology" to "logging response." The modeling method is as follows: For each lithology category... First, filter out all lithology labels from the dataset. The depth points are determined; then, all natural gamma logging values ​​corresponding to these depth points are extracted to form a gamma dataset for this lithology; finally, the probability distribution of this gamma dataset is fitted to establish a continuous probability density function. This function is the likelihood function. .

[0046] In a preferred embodiment of the invention, it is assumed that the gamma value distribution of each lithology approximately follows a Gaussian distribution (normal distribution). (For more complex geological conditions, the likelihood function...) Non-parametric methods such as kernel density estimation or Gaussian mixture models can also be used for modeling (to accommodate probability distributions of arbitrary shapes), which is a reasonable and efficient approximation in many geological environments. Therefore, it is necessary to calculate the mean of the gamma distribution for each lithology. and standard deviation Likelihood function The mathematical expression for is the probability density function of the Gaussian distribution:

[0047] in: The input is the natural gamma value. Lithology The sample mean of the gamma logging values. Lithology The sample standard deviation of gamma logging values.

[0048] (2) Based on prior knowledge of geological literature and common sense.

[0049] First, based on the geological background of the target area and generally accepted knowledge of rock physics, a set of criteria for candidate lithologies is established. ) and natural gamma ( Qualitative rules for response relationships.

[0050] Rule 1: Coal seam ( The natural gamma value of ) is usually the lowest.

[0051] Rule 2: Pure sandstone ( The gamma value is lower, but higher than that of coal seams.

[0052] Rule 3: Mudstone ( Because it is rich in radioactive elements such as potassium, thorium, and uranium, its gamma value is usually the highest.

[0053] Rule 4: Contains mudstone or carbonaceous mudstone ( The gamma value of ) is between that of sandstone and mudstone.

[0054] Secondly, the aforementioned qualitative rules are mapped onto a specific gamma value (e.g., 0-200 API) coordinate axis that conforms to the instrument measurement range of the region. For each lithology... Define a typical gamma response range The definition of this interval can be based on publicly published geological literature or by consulting senior geological experts. When uncertainty is high, a larger standard deviation is set to represent "uncertainty," thus allowing subsequent real-time data (likelihood) to have a greater weight in the Bayesian calculation.

[0055] (3) Construct a dual-mode prior knowledge base and define fusion weights.

[0056] To construct a more robust and comprehensive lithology identification engine that incorporates both data-driven objective laws and expert experience and geological rules, this invention proposes a strategy based on dual-model parallel reasoning. This step aims to construct a knowledge base containing two sets of independent prior parameters and set their fusion weights in the final decision.

[0057] First, construct and store two independent sets of prior parameters.

[0058] Adjacent Well Statistical Parameter Set: Through statistical analysis of adjacent well data, a set of statistical parameters is generated for each lithology. Calculate and store its prior probability And the parameters of its likelihood function (Gaussian distribution), namely the mean. and standard deviation This parameter set represents the objective patterns learned from historical data.

[0059] Geological rule parameter set: Through parameterization of geological literature, petrological knowledge, and expert experience, a set of parameters is provided for each lithology. Estimate and store its prior probability And the parameters of its likelihood function (Gaussian distribution), namely the mean. and standard deviation This set of parameters represents universal geological insights and expert knowledge.

[0060] Secondly, fusion weights are set, assigning confidence weights to the "models" represented by the two parameter sets respectively. and .in This represents the degree of confidence in the statistical data from neighboring wells, and both must meet the normalization condition: When there is no data from adjacent wells When there is a high degree of trust in the data from neighboring wells, or when the data is very close to the well being measured, Take the larger value. This knowledge base, containing two sets of parameters and a set of fusion weights, provides a solid foundation for subsequent real-time Bayesian model averaging calculations.

[0061] S2, real-time acquisition of natural gamma logging data while drilling and preprocessing, the preprocessing includes converting the gamma data into an equal depth interval sequence, using a moving average filter for noise suppression, and verifying the validity of the data according to a preset physical range.

[0062] This process is accomplished using a natural gamma-ray logging instrument deployed in the drill string assembly. Through a series of signal acquisitions and processing, the gamma-ray logging instrument ultimately collects a sequence of raw gamma-ray logging values ​​with depth information. Due to the harsh downhole environment, limited data transmission channel bandwidth, and interference from drilling activities, the raw gamma-ray data directly decoded from the surface system often contains noise, outliers, and is a depth-based sequence, requiring a series of preprocessing steps before it can be used for accurate lithology identification.

[0063] The preprocessing process is as follows: Step 1: Convert the gamma intensity information into equal depth intervals.

[0064] Gamma-ray logging instruments typically acquire signals at equal time points, while the goal of this invention is to predict formation lithology based on depth. Therefore, it is first converted into a gamma-ray intensity sequence with equal depth intervals. In this invention, 0.1 meters is used as the smallest interval unit. If multiple gamma-ray intensity values ​​are available at the same depth, the first value is taken to represent the gamma-ray intensity at that depth. When a gamma-ray value at a certain depth is missing, nearest neighbor interpolation is used.

[0065] in These represent the depths of the interpolation point, the preceding adjacent point, and the following adjacent point, respectively. This represents the gamma value of the interpolation point, the point immediately preceding it, and the point immediately following it.

[0066] Step 2: Noise filtering and smoothing.

[0067] Raw gamma sequences often contain high-frequency random noise, which can cause unnecessary fluctuations in subsequent probability calculations and even lead to misjudgments. To suppress this noise while preserving the true trend reflecting stratigraphic changes, this invention employs a moving average filter to perform real-time smoothing of the data.

[0068] Specifically, a fixed length is set as The time window. At each new gamma data point Upon arrival, calculate the point and its preceding points. The arithmetic mean of the points is used as the current depth. Filtered gamma value:

[0069] Step 3: Data validity check.

[0070] To prevent invalid data from contaminating the computational model due to signal transmission errors or temporary instrument malfunctions, it is necessary to perform matching on the data points. Perform a validity check. Check if the gamma value is within a preset reasonable physical range (0 ~ 400 API in this invention). Values ​​outside the range are considered outliers.

[0071] After the above acquisition and preprocessing steps, the final output of this step is a clean, smooth gamma-ray logging data stream that accurately corresponds to depth. This data stream is... The data pairs are fed into the Bayesian computation core in real time and point by point, serving as direct evidence for the model to make lithological posterior probability judgments.

[0072] S3. Based on the dual-mode prior knowledge base and the preprocessed gamma data, calculate the posterior probability distribution of lithology corresponding to the adjacent well statistical model and the geological rule model respectively, and perform a weighted average of the posterior probability distribution of the adjacent well statistical model and the geological rule model according to the preset fusion weight to generate the fused lithology probabilistic identification result.

[0073] The core computational idea of ​​this step is to transform the lithology identification problem from a traditional, deterministic judgment problem based on fixed thresholds into a probabilistic reasoning problem based on Bayesian theory. Its essence lies in fusing two types of information to make the optimal judgment: (a) Prior information: that is, the knowledge we already know about the probability of different lithologies based on historical data and geological patterns before we see any real-time data.

[0074] (b) Observational evidence: namely, the preprocessed natural gamma value measured in real time at the current depth point.

[0075] In this embodiment of the invention, the Naive Bayes classifier is an ideal model for realizing this idea. It not only perfectly aligns with Bayes' theorem mathematically, but its "naive" assumption of feature independence naturally holds true in the single-input feature (only natural gamma value) scenario of this invention, making it a theoretically sound, accurate, and efficient inference engine. This idea aims to achieve the following goals: (a) Probabilistic quantification of results: The final output is no longer a "yes / no" conclusion, but a quantitative probability that the current stratum belongs to each candidate lithology, providing confidence-based information support for geologically guided decision-making.

[0076] (b) Lightweight computation: The entire computation process involves only basic algebraic operations and does not involve complex iterative optimization, which fully meets the stringent real-time requirements of logging while drilling for data processing, making it potential to be deployed in well site computing units or even downhole instruments.

[0077] It should be noted that the Naive Bayes classifier used in this invention can be decomposed into the following four core parts:

[0078] Posterior probability The model's output. It represents the observed current natural gamma value. Under these conditions, the lithology at this depth point is classified as category [missing information]. The final probability. This is a "new understanding" that integrates prior knowledge and real-time data, and it is also the direct basis for decision-making.

[0079] Likelihood probability This serves as a bridge connecting "lithology" and "observational values." It is provided by the Gaussian probability density function (likelihood function) in the prior knowledge base constructed in step 1. This component describes "if the current lithology is indeed..." Then its observed gamma value is "How likely is it?" is the data-driven part of the model.

[0080] Prior probability This forms the knowledge base of the model. It is directly provided by the fused prior knowledge base constructed in step S1. This component represents the lithology at any point, judged based on regional geological patterns and expert knowledge, in the absence of any gamma observation data. The inherent probability is the knowledge-driven part of the model.

[0081] Evidence factors (Evidence): The model's normalization constant. It represents the observed gamma value. The total probability. In a single calculation, it is a constant for all lithologies, its function being to ensure that the sum of the posterior probabilities of all candidate lithologies is 1. Its calculation formula is:

[0082] The model architecture effectively integrates the prior knowledge base from step S1 with the real-time data stream from step S2 through the organic combination of these four parts, thereby completing probabilistic reasoning.

[0083] Regarding the calculation process, the core of this invention is to employ the idea of ​​Bayesian model averaging to dynamically weight and fuse the judgment results from the "data statistical model" and the "geological rule model." When a valid piece of preprocessed data from step S2... When (depth, gamma value) is input into this calculation module, the system will strictly follow the following algorithm in real time.

[0084] Input: Current depth Gamma .

[0085] Step A: Calculate the posterior probability based on the "adjacent well statistical model".

[0086] Extract parameters: Query and retrieve all lithology categories from the knowledge base. "Corresponding set of statistical parameters for adjacent wells": .

[0087] b. Calculation: For each lithology, calculate its nonnormalized posterior value:

[0088] c. Normalization: Calculate the total score:

[0089] Then the posterior probability distribution under this model is obtained:

[0090] Step B: Calculate the posterior probability based on the "geological rule model".

[0091] Extract parameters: Query and retrieve all lithology categories from the geological rule parameter set. "Corresponding set of statistical parameters for adjacent wells": .

[0092] Calculation: For each lithology, calculate its nonnormalized posterior value:

[0093] c. Normalization: Calculate the total score:

[0094] Then the posterior probability distribution under this model is obtained:

[0095] Step C: Weighted fusion of the final posterior probabilities.

[0096] Extract weights: Obtain preset fusion weights from the knowledge base. and .

[0097] Weighted average: For each lithology, the posterior probabilities obtained in the first two steps are weighted and averaged to obtain the final posterior probability that integrates multi-source knowledge.

[0098] Output: A complete probability distribution of the current depth h and its corresponding final posterior probabilities for all lithologies. The data is combined to form structured records for subsequent judgment and visualization.

[0099] Furthermore, embodiments of the present invention also support dynamic updating of the dual-mode prior knowledge base based on real-time identification results. This step feeds back the lithological posterior probability information obtained during real-time identification to the prior knowledge base, thereby enabling dynamic correction of the statistical model and geological rule model of adjacent wells, and thus improving the identification accuracy and robustness of the model under new working area or new drilling conditions.

[0100] Specifically, the process first compares the identification result of the current depth point with the data from neighboring wells, and then calculates the updated prior probability. ,in As a smoothing factor, the likelihood function parameters are then refitted based on the updated prior probabilities. and And adjust the fusion weights and The value of . Its technical effect lies in that, by dynamically updating the likelihood function parameters and adjusting the model fusion weights, the Bayesian inference model can adapt to stratigraphic changes in real time, improve the robustness and accuracy of the identification results, and enhance the model's adaptability to new areas, providing continuously optimized decision support for geological guidance.

[0101] To achieve the above embodiments, the present invention also proposes a real-time lithology identification device for natural gamma logging based on Bayesian theory. Figure 2 This invention provides a schematic diagram of a real-time lithology identification device based on Bayesian theory for natural gamma logging. The device includes: The dual-mode prior knowledge base construction module 100 is used to construct a dual-mode prior knowledge base containing adjacent well statistical parameters and geological rule parameters. The adjacent well statistical parameters are calculated based on adjacent well data of the target work area, and the geological rule parameters are generated parameterized through geological literature and expert experience. The gamma data preprocessing module 200 is used to acquire and preprocess natural gamma logging data while drilling in real time. The preprocessing includes converting the gamma data into an equal depth interval sequence, using a moving average filter for noise suppression, and verifying the validity of the data according to a preset physical range. The lithology posterior probability and probability fusion identification calculation module 300 is used to calculate the lithology posterior probability distribution corresponding to the adjacent well statistical model and the geological rule model based on the dual-mode prior knowledge base and the preprocessed gamma data, and to perform a weighted average of the posterior probability distribution of the adjacent well statistical model and the geological rule model according to the preset fusion weights to generate the fused lithology probabilistic identification result.

[0102] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0103] To implement the above embodiments, the present invention also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0104] To implement the above embodiments, the present invention also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0105] To implement the above embodiments, the present invention also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0106] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0107] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0108] This invention is intended to provide implementation schemes for users to selectively prevent the use or access to personal information data. That is, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information can be de-identified to protect user privacy.

[0109] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0110] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0111] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of the invention pertain.

[0112] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0113] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0114] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0115] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0116] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0117] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0118] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A real-time lithology identification method based on Bayesian theory in natural gamma logging, characterized in that, include: A dual-mode prior knowledge base is constructed, which includes statistical parameters of adjacent wells and geological rule parameters. The statistical parameters of adjacent wells are calculated based on the data of adjacent wells in the target work area, and the geological rule parameters are generated by parameterization through geological literature and expert experience. Real-time acquisition of natural gamma logging data while drilling and preprocessing of the data include converting the gamma data into a series of equal depth intervals, using a moving average filter for noise suppression, and verifying the validity of the data according to a preset physical range. Based on the dual-mode prior knowledge base and the preprocessed gamma data, the posterior probability distribution of lithology corresponding to the adjacent well statistical model and the geological rule model are calculated respectively. Then, the posterior probability distribution of the adjacent well statistical model and the geological rule model are weighted and averaged according to the preset fusion weight to generate the fused lithology probabilistic identification result.

2. The method as described in claim 1, characterized in that, The construction of the dual-mode prior knowledge base, which includes statistical parameters of adjacent wells and geological rule parameters, also includes: The adjacent well statistical parameters are obtained by statistically analyzing the total number of occurrences of each lithological category in the adjacent well data of the target work area. Total number of points for all lithologies Calculate the prior probability using the ratio ; The geological rule parameters are defined by geological literature and expert experience as typical gamma response ranges. Mapped to the mean of a Gaussian distribution and standard deviation ,in The midpoint of the interval, For the interval span times.

3. The method as described in claim 1, characterized in that, The real-time acquisition and preprocessing of natural gamma ray logging data during drilling also includes: A moving average filter is used to perform real-time smoothing of gamma data. The window length of the moving average filter is... The calculation formula is: ; For missing gamma values, nearest neighbor interpolation is used, and the interpolation formula is as follows: .

4. The method as described in claim 1, characterized in that, The step of calculating the posterior probability distributions of lithology corresponding to the adjacent well statistical model and the geological rule model based on the dual-mode prior knowledge base and the preprocessed gamma data also includes: The posterior probability calculation of the adjacent well statistical model uses the likelihood function of a Gaussian distribution. ; The posterior probability calculation of the geological rule model uses a Gaussian distribution likelihood function. .

5. The method as described in claim 1, characterized in that, The step of weighting the posterior probability distributions of the adjacent well statistical model and the geological rule model according to the preset fusion weights also includes: The fusion weight and Satisfy normalization conditions ; The weighted average formula is as follows: .

6. The method as described in claim 1, characterized in that, Also includes: The dual-mode prior knowledge base is dynamically updated based on real-time recognition results, specifically including: The identification results of the current depth point are compared with the data from neighboring wells to calculate the updated prior probability. ,in It is a smoothing factor; Refit the likelihood function parameters based on the updated prior probabilities. and And adjust the fusion weights and The value of .

7. A real-time lithology identification device for natural gamma logging based on Bayesian theory, characterized in that, include: The dual-mode prior knowledge base construction module is used to construct a dual-mode prior knowledge base containing adjacent well statistical parameters and geological rule parameters. The adjacent well statistical parameters are calculated based on adjacent well data of the target work area, and the geological rule parameters are generated parameterized through geological literature and expert experience. The gamma data preprocessing module is used to acquire and preprocess natural gamma logging data while drilling in real time. The preprocessing includes converting the gamma data into an equal depth interval sequence, using a moving average filter for noise suppression, and verifying the validity of the data according to a preset physical range. The lithology posterior probability and probability fusion identification calculation module is used to calculate the lithology posterior probability distribution corresponding to the adjacent well statistical model and the geological rule model based on the dual-mode prior knowledge base and the preprocessed gamma data, and to perform a weighted average of the posterior probability distribution of the adjacent well statistical model and the geological rule model according to the preset fusion weights to generate the fused lithology probabilistic identification result.

8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Intelligent stratum lithology identification method based on logging data

    CN119739992A