A complex lithology intelligent identification method based on geophysical constraint feature enhancement

CN122546343BActive Publication Date: 2026-09-15JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611035676.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-15
Estimated Expiration
2046-07-13

AI Technical Summary

Technical Problem

但这类方法通常依赖较大规模训练数据,且需要复杂的参数调优,模型性能受样本规模、参数设置影响较大

Benefits of technology

[0047]Compared with existing technologies, the beneficial effects of this invention are as follows: The invention introduces geophysical constraint features based on the original logging curves, which can fully utilize empirical knowledge from logging interpretation, allowing for a more complete expression of logging response information such as natural gamma, sonic transit time, density, neutron density, and resistivity, thereby enhancing the distinguishability between different lithologies; the invention automatically generates multiple combined features using a depth feature synthesis method, which can further explore the nonlinear combination relationships between different logging curves, compensating for the insufficient ability of a single original logging curve to characterize complex lithologies, and increasing the contribution of features to lithology classification; the invention employs a cross-validation recursive feature elimination algorithm... This invention filters generated features to remove redundant and noisy features, retaining key features with strong discriminative power, thereby improving the stability of the feature set and reducing the interference of invalid features on the model classification results. Utilizing the TabPFN model for complex lithology classification, this invention achieves good recognition results even with limited sample sizes, reduces the complex hyperparameter tuning process in traditional machine learning models, and improves model training and prediction efficiency. This invention enhances the accuracy, stability, and generalization ability of complex strata lithology identification, providing a more efficient and reliable technical means for lithological classification, reservoir evaluation, and comprehensive interpretation of oil and gas reservoirs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122546343B_ABST
    Figure CN122546343B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of well logging interpretation and reservoir evaluation in oil and gas exploration and development, and particularly relates to a complex lithology intelligent identification method based on geophysical constraint feature enhancement. The method comprises the following steps: S1, well logging and core data matching and sample pretreatment; S2, geophysical constraint feature enhancement; S3, depth feature synthesis; S4, redundant feature elimination and feature selection; and S5, lithology identification based on a TabPFN model. The present application introduces geophysical constraint features on the basis of original well logging curves, can fully utilize the experience in well logging interpretation, makes the well logging response information such as natural gamma, acoustic time difference, density, neutron and resistivity more fully expressed, and thus enhances the distinguishability between different lithologies; the present application automatically generates various combined features by combining with a depth feature synthesis method, can further mine the nonlinear combination relationship between different well logging curves, and improves the contribution of features to lithology classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of well logging interpretation and reservoir evaluation technology in oil and gas exploration and development, specifically to a method for intelligent identification of complex lithology based on enhanced geophysical constraint characteristics. Background Technology

[0002] Lithology identification is a crucial foundation for reservoir evaluation and hydrocarbon layer interpretation in oil and gas exploration and development. Well logging data, with its advantages of high continuity and vertical resolution, is frequently used for formation lithology identification. For conventional formations, different lithologies typically exhibit certain differences in well logging curves for natural gamma, sonic transit time, density, neutron density, and resistivity. Traditional cross plots, empirical discrimination, and statistical analysis methods can be used to classify the lithology of conventional formations.

[0003] However, in complex strata, due to the variable mineral composition, complex pore structure, intense diagenesis, and the development of fractures, alteration, or heterogeneity, the logging responses of different lithologies often overlap significantly, making it difficult to obtain accurate identification results using traditional methods. Igneous strata are a typical example of complex lithology identification, characterized by diverse lithological types, significant variations in mineral composition and structure, and frequent influences from alteration, fractures, and differences in pore structure, resulting in indistinct differences in the responses of different igneous lithologies on conventional logging curves.

[0004] In recent years, methods such as random forests, gradient boosting trees, support vector machines, and deep learning have been used for lithology classification, improving identification accuracy to some extent. However, these methods typically rely on large-scale training data and require complex parameter tuning; model performance is significantly affected by sample size and parameter settings.

[0005] Furthermore, existing methods often directly use raw well logging curves or simple mathematical transformation features, failing to fully utilize geophysical experience information in well logging interpretation. Directly generating a large number of combined features can easily introduce redundancy and noise, affecting model stability. Therefore, there is an urgent need for a complex lithology identification method that can integrate geophysical constraints, is suitable for small sample conditions, and reduces the complexity of parameter tuning. Summary of the Invention

[0006] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0007] To address the aforementioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution:

[0008] A method for intelligent identification of complex lithology based on enhanced geophysical constraint features includes the following steps:

[0009] S1 logging and core data matching and sample preprocessing:

[0010] S1.1 Collect well logging data for the study area;

[0011] S1.2 Collect core lithology identification data;

[0012] S1.3 performs well logging data matching with lithology labels;

[0013] S1.4 preprocesses the original lithology identification dataset;

[0014] Enhanced S2 geophysical constraint characteristics:

[0015] S2.1 Construct geophysical constraint features with lithological indication significance;

[0016] S2.2 Combines geophysical constraint features with the original well logging curves to form a geophysical constraint dataset;

[0017] S3 deep feature synthesis:

[0018] S3.1 uses the original well logging curves and geophysical constraint features as the basic input features and sets the mathematical transformation method;

[0019] S3.2 uses a deep feature synthesis method to automatically combine basic features to generate combined feature and extended feature datasets;

[0020] S4 Redundant Feature Removal and Feature Selection:

[0021] S4.1 Calculate the Pearson correlation coefficient between each feature in the extended feature dataset, and delete highly correlated redundant features according to the set threshold;

[0022] S4.2 uses RFECV for feature selection;

[0023] S4.3 uses the selected features as the final input variables to construct a feature selection dataset for subsequent lithology classification model training and prediction;

[0024] S5 uses the TabPFN model for lithology identification:

[0025] S5.1 uses the feature selection dataset as the model input data and the core identification lithology as the output label to form the tabular classification data required for the TabPFN model;

[0026] S5.2 uses a fixed random seed to divide the dataset into training and test sets;

[0027] S5.3 Input the TabPFN model for test set prediction and validation;

[0028] S5.4 Input actual well section samples for lithology prediction;

[0029] S5.5 outputs lithology identification results.

[0030] As a preferred embodiment of the intelligent identification method for complex lithology based on enhanced geophysical constraint features described in this invention, the well logging data in S1.1 includes sonic transit time, density, neutron porosity, natural gamma, deep lateral resistivity, shallow lateral resistivity, and microresistivity curves.

[0031] As a preferred embodiment of the intelligent identification method for complex lithology based on geophysical constraint features as described in this invention, the specific method of S1.3 is to match well logging curve data with core lithology labels according to the depth correspondence to form an original lithology identification dataset with well logging curves as variables and core lithology as labels.

[0032] As a preferred embodiment of the intelligent identification method for complex lithology based on geophysical constraint feature enhancement described in this invention, in step S1.4, the dataset preprocessing includes missing sample removal, abnormal logging value checking, and logarithmic transformation of resistivity curves to reduce the impact of invalid data and extreme values ​​on subsequent feature construction and model training.

[0033] As a preferred embodiment of the intelligent identification method for complex lithology based on enhanced geophysical constraints described in this invention, in step S2.1, the geophysically derived features include apparent limestone density and porosity. Apparent limestone density and porosity With neutron porosity average Apparent limestone density and porosity With neutron porosity difference M and N values ​​in the M–N cross plot and skeletal acoustic transit time in the MID lithology identification map. Skeletal density value .

[0034] As a preferred embodiment of the intelligent identification method for complex lithology based on enhanced geophysical constraint features described in this invention, the specific calculation formula for the derived features with geophysical significance is as follows:

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042] in The density of the limestone skeleton. For pore fluid density, This is the logging density value. For the acoustic transit time of pore fluid, This represents the sonic transit time value in well logging. This represents the neutron logging value for pore fluid.

[0043] As a preferred embodiment of the complex lithology intelligent identification method based on geophysical constraint feature enhancement described in this invention, the specific method of S4.2 is as follows: based on the features after correlation screening, a recursive feature elimination combined with cross-validation method is used to further screen features. By recursively training the model, evaluating the feature contribution, and gradually eliminating features with lower importance, a feature combination with lithology discrimination capability is determined.

[0044] As a preferred embodiment of the intelligent identification method for complex lithology based on enhanced geophysical constraint features described in this invention, in step S5.2, multiple random divisions are performed, and the average identification effect is statistically analyzed.

[0045] As a preferred embodiment of the intelligent identification method for complex lithology based on geophysical constraint features as described in this invention, the specific method of S5.3 is as follows: the training samples and their lithology labels are input into the TabPFN model as context information, and the test samples are input into the model as prediction objects to obtain the lithology prediction results corresponding to the test samples. Subsequently, the model prediction results are compared with the core identification lithology corresponding to the test samples to verify the lithology identification effect of the model.

[0046] As a preferred embodiment of the complex lithology intelligent identification method based on geophysical constraint feature enhancement described in this invention, the specific method of S5.5 is to arrange the lithology prediction results corresponding to each depth point of the actual well section in depth order to form a continuous well section lithology identification profile. The lithology identification profile is displayed in correspondence with conventional logging curves, core photos and core identification lithology, and is used for continuous identification of lithology in the target well section and reservoir evaluation.

[0047] Compared with existing technologies, the beneficial effects of this invention are as follows: The invention introduces geophysical constraint features based on the original logging curves, which can fully utilize empirical knowledge from logging interpretation, allowing for a more complete expression of logging response information such as natural gamma, sonic transit time, density, neutron density, and resistivity, thereby enhancing the distinguishability between different lithologies; the invention automatically generates multiple combined features using a depth feature synthesis method, which can further explore the nonlinear combination relationships between different logging curves, compensating for the insufficient ability of a single original logging curve to characterize complex lithologies, and increasing the contribution of features to lithology classification; the invention employs a cross-validation recursive feature elimination algorithm... This invention filters generated features to remove redundant and noisy features, retaining key features with strong discriminative power, thereby improving the stability of the feature set and reducing the interference of invalid features on the model classification results. Utilizing the TabPFN model for complex lithology classification, this invention achieves good recognition results even with limited sample sizes, reduces the complex hyperparameter tuning process in traditional machine learning models, and improves model training and prediction efficiency. This invention enhances the accuracy, stability, and generalization ability of complex strata lithology identification, providing a more efficient and reliable technical means for lithological classification, reservoir evaluation, and comprehensive interpretation of oil and gas reservoirs. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0049] Figure 1 This is a flowchart of an embodiment of a complex lithology intelligent identification method based on geophysical constraint feature enhancement according to the present invention;

[0050] Figure 2 This is a comparison chart of the lithology identification performance of different machine learning models in an embodiment of the intelligent identification method for complex lithology based on geophysical constraint features of the present invention. Among them, (a) is a comparison chart of Accuracy index; (b) is a comparison chart of F1_weighted index; and (c) is a comparison chart of F1_macro index.

[0051] Figure 3 This is a TabPFN prediction result diagram of various lithologies in the actual formation of Well A in an embodiment of the intelligent identification method for complex lithologies based on geophysical constraint features of the present invention. Among them, (a) is the lithology identification result diagram of the 3190–3210 m section of Well A; (b) is the lithology identification result diagram of the 3350–3380 m section of Well A. Detailed Implementation

[0052] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0053] Secondly, the present invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not according to the usual scale. Furthermore, the schematic diagrams are merely examples and should not limit the scope of protection of the present invention. In addition, actual fabrication should include three-dimensional spatial dimensions of length, width, and depth.

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0055] This invention introduces empirical logging formulas to construct geophysical constraint features, enabling a more comprehensive expression of the response characteristics of logging data to different lithologies. By combining depth feature synthesis methods to automatically generate multiple combined features, and employing a cross-validation recursive feature elimination algorithm to screen out key features with strong discriminative power, this invention can reduce the impact of redundant and noisy features on classification results, improving the stability and effectiveness of the feature set. Based on this, the TabPFN model is used for complex lithology classification, which can reduce the complex parameter tuning process under small sample conditions, improving the accuracy, stability, and generalization ability of lithology identification. This provides an efficient and reliable technical method for lithological division, reservoir evaluation, and comprehensive interpretation of oil and gas reservoirs in complex formations.

[0056] Specifically, this invention proposes a lithology intelligent identification method based on geophysical constraint feature enhancement and the TabPFN model. This method mainly includes the following five steps:

[0057] S1: Matching of well logging and core data and sample preprocessing:

[0058] S1.1 Collect logging data for the study area: Collect conventional logging data from multiple wells in the study area, including curves for acoustic transit time (AC), density (DEN), neutron porosity (CNL), natural gamma (GR), deep lateral resistivity (RLLD), shallow lateral resistivity (RLLS), and microresistivity (RMLL).

[0059] S1.2 Collect core lithology identification data: Based on the results of core observation, thin section identification or core description, determine the true lithology type corresponding to each depth point in the study section, and use it as label data for training and validation of machine learning models.

[0060] S1.3 Matching well logging data with lithology labels: According to the depth correspondence, the well logging curve data is matched with the core lithology labels to form an original lithology identification dataset with well logging curves as variables and core lithology as labels.

[0061] S1.4 Data preprocessing: The matched original lithology identification dataset is preprocessed, including missing sample removal, abnormal logging value checking, and logarithmic transformation of resistivity curves, in order to reduce the impact of invalid data and extreme values ​​on subsequent feature construction and model training.

[0062] S2: Enhanced Geophysical Constraint Features:

[0063] S2.1 Constructing Geophysical Constraint Features with Lithological Indication Significance: Combining well logging interpretation theory, based on conventional well logging curves such as sonic transit time, density, and neutron porosity, constructing derived features such as apparent limestone density porosity, density porosity and neutron porosity combination parameters, M-N cross plot parameters, and MID lithology identification parameters to enhance the differences in well logging responses between different lithologies.

[0064] Specifically, the following seven derived features with geophysical significance were constructed:

[0065] Apparent limestone density and porosity It reflects the size of the pores and helps to distinguish lithologies with different porosity distributions.

[0066] Apparent limestone density and porosity With neutron porosity average This approach takes into account the differences in the response of the two logging methods to porosity, thereby enhancing the robustness of porosity estimation.

[0067] Apparent limestone density and porosity With neutron porosity difference It is also a commonly used indicator in the field of well logging to distinguish lithology.

[0068] The M and N values ​​in the M–N cross plot are classic lithology identification parameters.

[0069] Skeletal acoustic transit time in MID lithology identification maps , and skeleton density value .

[0070] The specific calculation formula is as follows:

[0071] (1)

[0072] (2)

[0073] (3)

[0074] (4)

[0075] (5)

[0076] (6)

[0077] (7)

[0078] in The density of the limestone skeleton. For pore fluid density, This is the logging density value. For the acoustic transit time of pore fluid, This represents the sonic transit time value in well logging. This represents the neutron logging value for pore fluid.

[0079] S2.2 Forming a Geophysical Constraint Dataset: The constructed geophysical constraint features are combined with the original well logging curves to form a geophysical constraint dataset, so that the model input simultaneously includes the original well logging response information and professional well logging features with lithological interpretation significance, providing a foundation for subsequent feature synthesis and lithological classification.

[0080] S3: Deep Feature Synthesis

[0081] S3.1 Determine the basic features and combination methods: The original well logging curves and geophysical constraint features are used as the basic input features, and mathematical transformation methods such as addition, subtraction, multiplication, absolute value and negation are set to provide a computational basis for subsequent feature combination expansion.

[0082] S3.2 Generate combined features and construct extended feature dataset: The basic features are automatically combined using the deep feature synthesis method to generate an extended feature dataset consisting of the original well logging curves, geophysical constraint features and their synthesized features, in order to explore the potential correlations between different well logging curves and between well logging curves and geophysical constraint features, and enhance the model's ability to express complex lithological differences.

[0083] S4: Redundant Feature Removal and Feature Selection

[0084] S4.1 Perform correlation screening: Calculate the Pearson correlation coefficient between each feature in the extended feature dataset, and delete highly correlated redundant features according to the set threshold to reduce the impact of duplicate information and multicollinearity on model training.

[0085] S4.2 Feature selection using RFECV: Based on the features screened by correlation, a method of recursive feature elimination combined with cross-validation is used to further screen features. By recursively training the model, evaluating the feature contribution, and gradually eliminating features with lower importance, feature combinations with strong lithology discrimination ability are determined.

[0086] S4.3 Constructing a Feature Selection Dataset: Using the selected features as the final input variables, construct a feature selection dataset for subsequent lithology classification model training and prediction.

[0087] S5: Lithology identification based on the TabPFN model:

[0088] S5.1 Constructing Model Input Data: The feature selection dataset obtained in S4 is used as the model input data, and the lithology of the core identification is used as the output label to form the tabular classification data required for the TabPFN model.

[0089] S5.2 Splitting the Training and Test Sets: The dataset is split into training and test sets using a fixed random seed. To improve the stability of the results, multiple random splits can be performed, and the average recognition performance is calculated.

[0090] S5.3 Inputting the TabPFN model for test set prediction and validation: The training samples and their lithology labels are input into the TabPFN model as contextual information, and the test samples are input into the model as prediction objects to obtain the lithology prediction results corresponding to the test samples. Subsequently, the model prediction results are compared with the core identification lithology corresponding to the test samples to verify the lithology identification effect of the model.

[0091] S5.4 Input actual well section samples for lithology prediction: After completing the model validation, input the actual well section samples to be identified as the prediction objects into the TabPFN model to obtain the lithology prediction results corresponding to each depth point of the actual well section.

[0092] S5.5 Output Lithology Identification Results: The lithology prediction results corresponding to each depth point in the actual well section are arranged in depth order to form a continuous well section lithology identification profile. The lithology identification profile can be displayed in correspondence with conventional logging curves, core photographs, and core-identified lithology, and is used for continuous identification of lithology in the target well section and reservoir evaluation.

[0093] Terminology Explanation

[0094] TabPFN: A pre-trained model for tabular data classification that can use known samples and their labels to classify and predict samples to be identified, and is suitable for small sample classification tasks.

[0095] RFECV: Recursive Feature Elimination combined with Cross-Validation is a method used to progressively eliminate features with low contribution and determine a better subset of features.

[0096] Pearson correlation coefficient: A statistical indicator used to measure the degree of linear correlation between two variables. In this invention, it is used to determine whether there is a high correlation and information redundancy between features.

[0097] Original lithology identification dataset: refers to the dataset formed by matching the original well logging curves with the core lithology labels according to the depth correspondence.

[0098] Geophysically constrained datasets: These are datasets formed by adding geophysically significant derived features to the original lithology identification dataset.

[0099] Extended feature dataset: refers to a dataset formed by synthesizing combined features through deep feature synthesis based on a geophysical constraint dataset.

[0100] Feature selection dataset: refers to the dataset formed by retaining key discriminative features after performing relevance filtering and RFECV feature selection on the extended feature dataset.

[0101] Example:

[0102] Taking the igneous strata of a certain study area as an example, the specific implementation process of the method of the present invention is shown in the flowchart below. Figure 1 As shown.

[0103] Step 1: Matching well logging and core data and sample preprocessing:

[0104] The data used in this embodiment comes from geophysical logging data and core lithology identification data from 16 wells. The geophysical logging data includes conventional logging curves such as acoustic transit time (AC), neutron porosity (CNL), density (DEN), natural gamma ray (GR), deep lateral resistivity (RLLD), shallow lateral resistivity (RLLS), and microresistivity (RMLL).

[0105] A correspondence was established between core sampling depth and well logging curve depth. The core identification results were used as lithology tags, and well logging response data at the corresponding depth points were extracted to construct an original lithology identification dataset. The lithology tags include six types of igneous rocks: trachyte, trachyte, diabase, volcanic breccia, breccia lava, and basalt.

[0106] This embodiment yielded 1537 valid samples, including 411 trachyte, 380 basalt, 341 diabase, 163 trachyandesite, 122 breccia lava, and 120 volcanic breccia. Through the above processing, a raw lithology identification dataset was formed, consisting of well logging curve data and corresponding lithology labels. This dataset provides a data foundation for subsequent geophysical constraint feature construction, depth feature synthesis, feature selection, and TabPFN lithology identification.

[0107] Step 2: Enhancement of Geophysical Constraint Features:

[0108] Based on the original lithology identification dataset constructed in step one, a geophysical constraint dataset is further constructed according to the geophysical response relationship between well logging response and lithology, mineral composition, pore structure, and fluid properties. The geophysical constraint features are used to enhance the differences in well logging responses between different igneous lithologies, thereby improving the ability of subsequent models to distinguish complex igneous lithologies.

[0109] Specifically, based on the original logging curves such as acoustic transit time AC, neutron porosity CNL, density DEN, natural gamma ray GR, deep lateral resistivity RLLD, shallow lateral resistivity RLLS, and microresistivity RMLL, several derived features reflecting lithological response, electrical differences, pore structure characteristics, and multi-curve coupling response are constructed and denoted as C1, C2, C3, C4, C5, C6, and C7, respectively. Their specific calculation methods are shown in formulas (1)-(7).

[0110] Wherein, C1 represents the apparent density porosity of limestone, C2 represents the average porosity, C3 represents the porosity difference, C4 represents the M value in the M-N cross plot, C5 represents the N value in the M-N cross plot, C6 represents the matrix sonic transit time, and C7 represents the matrix density. By constructing the above geophysical constraint dataset, the model input can include not only the original well logging response information but also well logging interpretation features with lithological indication significance, providing a foundation for subsequent depth feature synthesis and lithology identification.

[0111] Step 3: Deep Feature Synthesis

[0112] Based on the original well logging curves obtained in step one and the geophysical constraint dataset constructed in step two, depth features are synthesized to expand the input feature space of the lithology identification model. In this embodiment, the original well logging curves AC, CNL, DEN, GR, RLLD, RLLS, and RMLL, together with the geophysical constraint features C1, C2, C3, C4, C5, C6, and C7, are used as the basic feature set.

[0113] By employing feature transformation methods such as addition, subtraction, multiplication, absolute value, and inversion, the basic feature set is automatically combined and calculated to generate a variety of combined features.

[0114] After deep feature synthesis, a total of 316 combined features were generated. These combined features were then merged with the original well logging curves and geophysical constraint features to form an extended feature dataset. In this embodiment, the extended feature dataset includes several candidate features for subsequent correlation screening, recursive feature elimination, and feature selection dataset construction.

[0115] Step 4: Redundant Feature Removal and Feature Selection

[0116] Based on the extended feature dataset obtained in step three, feature selection and key feature set construction are performed. Since deep feature synthesis generates a large number of combined features, some of which may have strong correlations or information redundancy, directly inputting all features into the model could increase model complexity and reduce the stability and generalization ability of the model's recognition results. Therefore, this embodiment first performs correlation analysis on the extended feature dataset, calculating the Pearson correlation coefficient between each candidate feature; when the absolute value of the correlation coefficient between any two features is greater than a preset threshold of 0.95, one redundant feature is removed to reduce the impact of duplicate information and multicollinearity on subsequent model recognition.

[0117] After removing redundant features, a recursive feature elimination combined with cross-validation method is used to further filter the remaining features. Specifically, a classification model that can output feature importance or feature weights is used as the base learner, such as random forest, extreme random trees, support vector machines, or gradient boosting trees. The model is repeatedly trained and the contribution of each feature to the lithology classification result is evaluated. Features with lower contributions are gradually eliminated until a better feature subset is obtained. To improve the reliability of the feature selection results, cross-validation can be used to evaluate the recognition effect corresponding to different feature subsets, and the final retained features are determined based on the model performance.

[0118] After the above feature selection process, this embodiment ultimately obtained 42 key features for lithology identification. These key features include original well logging curves, geophysical constraint features, absolute value features, additive combination features, difference combination features, and multiplicative combination features. The final selected features are shown in Table 1.

[0119] Table 1: Key Features for RFECV Screening

[0120]

[0121] Step 5: Lithology identification based on the TabPFN model:

[0122] The features selected in step four are used as model input, and core lithology identification is used as labels to construct the TabPFN feature selection dataset. Samples with core lithology labels are used as known samples and input into the TabPFN model for lithology classification inference, thus validating the model.

[0123] like Figure 2As shown, this embodiment compares and verifies the lithology identification performance of five classification models: CatBoost (category feature gradient boosting model), LightGBM (lightweight gradient boosting model), Random Forest (random forest model), TabPFN (tab prior fitting network model), and XGBoost (extreme gradient boosting model). Evaluation metrics include accuracy, weighted F1 score, and macro-average F1 score. The results show that all models have good overall identification performance, with TabPFN performing better in all three metrics, indicating its good classification performance and applicability in the lithology identification task of this study.

[0124] After model validation, Well A in the study area was selected as the actual application well, and the logging characteristics of the target well section of Well A were input into the TabPFN model to obtain the lithology prediction results corresponding to each depth point of the target well section of Well A. Subsequently, the lithology prediction results were arranged in depth order to form a continuous well section lithology identification profile, which is used to characterize the vertical variation characteristics of the igneous lithology of the target well section of Well A.

[0125] like Figure 3 The figure shows conventional logging curves, including CAL, GR, SP, AC, CNL, DEN, RMLL, RLLS, and RLLD. The core lithology channel represents the core identification lithology results, and the TabPFN predicted lithology channel represents the predicted lithology from the TabPFN model. The core photograph channel shows the core photograph and its corresponding lithology and depth, corresponding to the depth of the red line in the figure. The predicted lithology of the actual well section in Well A generally shows good consistency with the core identification results. Near 3193.08m and 3360.5m, the model prediction results correspond to the core identification results, indicating that this method can achieve continuous identification of igneous rock lithology in the actual well section.

[0126] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for intelligent identification of complex lithology based on geophysical constraint feature enhancement, characterized in that, Includes the following steps: S1 logging and core data matching and sample preprocessing: S1.1 Collect well logging data for the study area; S1.2 Collect core lithology identification data; S1.3 performs well logging data matching with lithology labels; S1.4 preprocesses the original lithology identification dataset; Enhanced S2 geophysical constraint characteristics: S2.1 Construct geophysical constraint features with lithological indication significance; S2.2 Combines geophysical constraint features with the original well logging curves to form a geophysical constraint dataset; S3 deep feature synthesis: S3.1 uses the original well logging curves and geophysical constraint features as the basic input features and sets the mathematical transformation method; S3.2 uses a deep feature synthesis method to automatically combine basic features to generate combined feature and extended feature datasets; S4 Redundant Feature Removal and Feature Selection: S4.1 Calculate the Pearson correlation coefficient between each feature in the extended feature dataset, and delete highly correlated redundant features according to the set threshold; S4.2 uses RFECV for feature selection; S4.3 uses the selected features as the final input variables to construct a feature selection dataset for subsequent lithology classification model training and prediction; S5 uses the TabPFN model for lithology identification: S5.1 uses the feature selection dataset as the model input data and the core identification lithology as the output label to form the tabular classification data required for the TabPFN model; S5.2 uses a fixed random seed to divide the dataset into training and test sets; S5.3 Input the TabPFN model for test set prediction and validation; S5.4 Input actual well section samples for lithology prediction; S5.5 outputs lithology identification results.

2. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, The logging data in S1.1 includes sonic transit time, density, neutron porosity, natural gamma, deep lateral resistivity, shallow lateral resistivity, and microresistivity curves.

3. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, The specific method of S1.3 is to match the logging curve data with the core lithology label according to the depth correspondence to form an original lithology identification dataset with the logging curve as the variable and the core lithology as the label.

4. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, In S1.4, dataset preprocessing includes missing sample removal, abnormal logging value checking, and logarithmic transformation of resistivity curves to reduce the impact of invalid data and extreme values ​​on subsequent feature construction and model training.

5. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, In S2.1, the geophysical derivative features include apparent limestone density and porosity. Apparent limestone density and porosity With neutron porosity average Apparent limestone density and porosity With neutron porosity difference M and N values ​​in the M–N cross plot and skeletal acoustic transit time in the MID lithology identification map. Skeletal density value .

6. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 5, characterized in that, The specific calculation formula for the derived features with geophysical significance is as follows: in The density of the limestone skeleton. For pore fluid density, This is the logging density value. For the acoustic transit time of pore fluid, This represents the sonic transit time value in well logging. This represents the neutron logging value for pore fluid.

7. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, The specific method of S4.2 is as follows: based on the features after correlation screening, a recursive feature elimination combined with cross-validation method is used to further screen features. The feature combination with lithology discrimination ability is determined by recursively training the model, evaluating the feature contribution and gradually eliminating features with low importance.

8. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, In step S5.2, multiple random divisions are performed, and the average recognition effect is calculated.

9. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, The specific method of S5.3 is as follows: the training samples and their lithology labels are input into the TabPFN model as context information, and the test samples are input into the model as prediction objects to obtain the lithology prediction results corresponding to the test samples. Then, the model prediction results are compared with the core identification lithology corresponding to the test samples to verify the lithology identification effect of the model.

10. The intelligent identification method for complex lithology based on geophysical constraint feature enhancement according to claim 1, characterized in that, The specific method of S5.5 is to arrange the lithology prediction results corresponding to each depth point of the actual well section in order of depth to form a continuous well section lithology identification profile. The lithology identification profile is displayed in correspondence with conventional logging curves, core photos and core identification lithology, and is used for continuous identification of lithology of the target well section and reservoir evaluation.

Citation Information

Patent Citations

  • Inematic rock lithology identification method based on engineering optimization cascade forest

    CN119475117A

  • Intelligent reservoir logging evaluation method driven by double models

    CN119807888A