Low-resistance layer intelligent identification method based on large language model

CN122451594BActive Publication Date: 2026-09-08CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610911853.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-08
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

[0005]为了解决现有油气田开发中低阻层与水层测井响应相近导致识别精度低、边界样本油水判别不稳定、不同区块和层系适用性差以及误判结果缺乏可追溯解释的问题,本发明提供一种基于大语言模型的低阻层智能识别方法

Benefits of technology

[0054] 1. This invention unifies well logging curve data, static geological parameters, stratigraphic information, top and bottom depth information of small layers, and true labels of oil and water layers into small-layer samples. Compared with traditional identification methods that rely on only a single well logging curve or a small number of parameters, this invention can simultaneously characterize electrical response, reservoir properties, and geological background in scenarios where the electrical responses of low-resistivity layers and water layers are similar, thereby improving the completeness and consistency of the data expression required for identifying low-resistivity potential layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451594B_ABST
    Figure CN122451594B_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent identification technology for oil and gas reservoirs, and in particular to an intelligent identification method for low-resistivity layers based on a large language model. The method includes: acquiring multi-source data of low-resistivity layers at small levels in a target oilfield, and preprocessing well logging curves, static geological parameters, and oil-water layer labels; extracting time-series statistical features of well logging curves and fusing them with parameters such as porosity, permeability, clay content, and layer thickness; inputting the fused features into multiple large language models to obtain oil-water layer prediction results, and constructing an error sample set and sample difficulty labels; constructing a supervised comparative learning model based on the sample difficulty labels to form a model cognitive map; clustering and extracting rules from error sample clusters to generate an error pattern library and a diagnostic rule library, and injecting them into the large language model for knowledge-enhanced reasoning to improve the accuracy and interpretability of low-resistivity layer identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent interpretation technology for oil and gas exploration and development and well logging, and in particular to an intelligent identification method for low-resistivity layers based on a large language model. Background Technology

[0002] Low-resistivity reservoirs are crucial potential reservoirs for tapping remaining oil in the later stages of oilfield development. Their identification directly impacts the degree of reserve utilization, the formulation of development adjustment plans, and the effectiveness of stabilizing oil production and increasing efficiency. Low-resistivity oil reservoirs typically exhibit characteristics such as weak resistivity response, indistinct electrical differences between oil and water layers, and ambiguous logging interpretation boundaries. Under complex geological conditions such as high clay content, high bound water saturation, development of thin interbedded sandstone and mudstone layers, low porosity and permeability, or high formation water salinity, the logging response of low-resistivity oil reservoirs can easily resemble that of water layers, dry layers, or low-production layers. This can lead to misjudgment or omission of potential oil reservoirs during conventional interpretation, affecting subsequent oil and gas development outcomes.

[0003] Traditional methods for identifying low-resistivity layers mainly include curve overlay, intersection plotting, empirical discrimination, multi-parameter comprehensive interpretation, and machine learning, deep learning, and large language model-assisted identification methods. Traditional methods rely on human experience, are highly subjective, and lack cross-block adaptability. Machine learning and deep learning methods focus primarily on the final classification result, lacking analysis of misclassified samples and their underlying causes. While large language models can assist in understanding geological semantic information, existing methods mostly remain at the direct prediction level, failing to fully utilize the differences in multi-model predictions, misclassification results, and reasoning to enhance the identification process.

[0004] Especially in complex scenarios such as thin layers of high-muddy soil, low porosity and low permeability, thin interlayers of sandstone and mudstone, mixed resistivity responses, pseudo-low resistivity with smooth curves, and insufficient minority class samples, the logging responses of low-resistivity layers and water layers are similar, making the model prone to stable misjudgments. Existing intelligent low-resistivity layer identification technologies have not fully integrated small-level multi-source data fusion, logging curve time-series feature extraction, multi-language model error analysis, sample difficulty modeling, supervised comparative learning, error pattern mining, and knowledge-enhanced reasoning. Problems still exist, including unclear identification boundaries, difficulty in explaining the causes of misjudgments, and the inability to feedback and utilize erroneous experience. Therefore, a low-resistivity layer intelligent identification method based on a large language model is needed at this stage. Summary of the Invention

[0005] To address the problems in existing oil and gas field development, such as low identification accuracy due to similar logging responses between low-resistivity layers and water layers, unstable oil-water discrimination of boundary samples, poor applicability to different blocks and formations, and lack of traceable explanations for misjudgments, this invention provides an intelligent identification method for low-resistivity layers based on a large language model.

[0006] In a first aspect, the present invention provides a low-resistivity intelligent recognition method based on a large language model, which adopts the following technical solution:

[0007] A low-resistivity intelligent recognition method based on a large language model, characterized by comprising:

[0008] Acquire multi-source data of small-level layers related to low resistivity layers in the target oilfield, and preprocess the multi-source data of small-level layers, which includes well logging curve data, static geological parameters and true labels of oil and water layers;

[0009] Based on the preprocessed logging curve data, logging curve segments within the corresponding depth range are extracted according to the top and bottom depths of the sub-layer, and time-series features are extracted from the logging curve segments.

[0010] The extracted well logging time-series features are fused with static geological parameters to obtain a small-level fused feature vector;

[0011] The feature vectors of the small-level fusion are transformed into text description information and input into multiple large language models to obtain the oil-water layer prediction results of the corresponding small-level samples.

[0012] Based on the difference between the actual labels of the oil-water layer and the predicted oil-water layer, a set of erroneous samples is constructed, and sample difficulty labels are generated.

[0013] A supervised contrastive learning model is constructed based on the fusion feature vectors of small layers and the sample difficulty labels. The small-layer samples are mapped to a low-dimensional embedding space to form a model cognitive map that represents the distribution of recognition difficulty of large language models.

[0014] Cluster analysis and rule extraction are performed on error sample clustering regions in the low-dimensional embedding space to generate a low-resistivity layer error pattern library and a diagnostic rule library;

[0015] The low-resistivity layer error pattern library and diagnostic rule library are injected into the large language model to perform knowledge-enhanced reasoning and output the low-resistivity layer recognition results.

[0016] Furthermore, the preprocessing of the multi-source data at the sub-layer level includes: truncating the logging curve data into sub-layer intervals using the top and bottom depths of the sub-layer as boundaries; associating the truncated logging curve data with the static geological parameters and true labels of the oil and water layers of the corresponding sub-layer; and performing missing value imputation, outlier handling, and standardization on the logging curves and static geological parameters to obtain unified multi-source sample data at the sub-layer level. The multi-source data at the sub-layer level is constructed in the following form:

[0017] ,

[0018] in, This represents the i-th sub-layer sample. Indicates the well number information. This indicates sub-layer or layer information. This indicates the top and bottom depth information of a small layer. Represents a set of static geological parameters. Indicates the true label of the oil-water layer; the set of static geological parameters It includes at least one or more of the following: porosity, permeability, oil saturation, water saturation, clay content, sublayer thickness, or lithology.

[0019] Further, the well logging curve segments within corresponding depth intervals are extracted according to the top and bottom depths of the sub-layers. This is obtained by extracting curve data within the corresponding depth intervals from the pre-processed continuous well logging curves. The well logging curves include one or more of the following: resistivity curves, natural gamma curves, sonic transit time curves, density curves, and neutron curves. Through sub-layer interval extraction, the original continuous well logging curves are converted into well logging curve segments corresponding one-to-one with the sub-layers. The formula for extracting well logging curve segments is as follows:

[0020] ,

[0021] in, This represents the logging curve segment corresponding to the i-th sub-layer sample. Indicates the depth of the top floor of the small-scale building. Indicates the depth of the bottom layer. This represents the logging response value at depth d.

[0022] Furthermore, the extraction of time-series features from the logging curve segments includes truncating the logging curve segments with the top and bottom depths of the sub-layers as boundaries, calculating the statistical features, distribution features, trend features, and fluctuation features of the logging curve segments using the Tsfresh feature extraction method, and filtering them based on feature missing rate, variance, correlation, and the degree of association with the true labels of oil and water layers to obtain a logging time-series feature vector for low-resistivity layer identification. The formula for generating the logging time-series feature vector is:

[0023] ,

[0024] in, Let φ(·) represent the logging time-series feature vector of the i-th sub-layer sample, and let φ(·) represent the time-series feature extraction function. The p-th type of time series feature is represented; the time series feature includes one or more of the following: mean, variance, extreme values, quantiles, autocorrelation coefficient, Fourier coefficient, frequency domain energy, kurtosis, skewness, number of local extrema, and fluctuation amplitude.

[0025] Furthermore, the process of fusing the extracted well logging time-series features with static geological parameters to obtain a small-level fused feature vector includes using small-layer samples as a unified object, associating the well logging time-series features extracted from well logging curve segments with one or more static geological parameters such as porosity, permeability, oil saturation, water saturation, clay content, small-layer thickness, or lithology. Field unification and numerical range correction are performed on features from different sources to form a small-level fused feature vector that can be used simultaneously for natural language description generation, large language model prediction error analysis, and supervised contrastive learning modeling. The formula for constructing the small-level fused feature vector is as follows:

[0026] ,

[0027] in, This represents the sub-layer fusion feature vector of the i-th sub-layer sample. Represents the well logging time series feature vector. This represents a static geological parameter vector, and [;] represents the feature concatenation operation. This indicates standardization or normalization.

[0028] Furthermore, the process of converting the sub-level fusion feature vector into textual description information includes parsing the well logging statistical features and static geological parameters in the sub-level fusion feature vector according to a preset text template, and converting continuous numerical parameters into corresponding geological grade descriptions. The geological grade descriptions include one or more of the following: low, medium, high, good, poor, large fluctuation, or small fluctuation. The field parsing results and geological grade descriptions are combined to form natural language descriptions and structured JSON descriptions, enabling the large language model to understand the electrical response, reservoir properties, and oil-water characteristics of the sub-layers based on the textual description information without receiving the actual labels of the oil-water layers.

[0029] Further, obtaining the oil-water layer prediction results for the corresponding sub-layer samples includes inputting the textual description information of the same sub-layer sample into various language models, and obtaining the prediction results for oil and water layer categories under the same task instructions, the same output format constraints, and the same inference parameters; formatting and parsing the outputs of various language models to extract prediction categories, prediction rationales, and confidence information; and forming a multi-model prediction result set for constructing erroneous samples and evaluating sample difficulty based on the prediction consistency and prediction differences among multiple large language models. The formula for the oil-water layer prediction result set is:

[0030] ,

[0031] in, This represents the set of multi-model prediction results for the i-th sub-layer sample. This represents the predicted category of the m-th large language model. Indicates the prediction confidence level. Indicate the reasons for the prediction. This indicates the number of large language models involved in the prediction.

[0032] Furthermore, based on the difference between the true labels of the oil-water layer and the predicted results, an error sample set is constructed. This includes marking the sub-layer sample as an error sample when at least one large language model makes a prediction error, and recording the misjudged model number, misjudgment category, prediction reason, and confidence information; when multiple large language models misjudge the same sub-layer sample, the sub-layer sample is further marked as a common misjudged sample or an extremely difficult sample, thereby forming a hierarchical error sample set for error pattern mining and sample difficulty evaluation. The error sample set is constructed using the following error labeling function:

[0033] ,

[0034] in, This represents the error weight of the m-th large language model for the i-th small-layer sample. Indicates the true label, Indicates the prediction confidence level. This indicates the degree of inconsistency between the reasoning for the prediction and the diagnostic rules. and The total set of erroneous samples is represented by the weighting coefficients:

[0035] ,

[0036] Where E represents the set of erroneous samples misclassified by at least one large language model. Let M represent the number of large language models involved in the prediction, where M is the i-th small-layer sample. Indicates an indicator function.

[0037] Furthermore, the sample difficulty label is generated based on the number of misjudgments by multiple large language models for each sub-layer sample, the divergence of prediction results, and the stability of misjudgment categories. The sub-layer samples are divided into easily identifiable samples, high-risk samples, and extremely difficult samples. This sample difficulty label serves as a supervisory signal for the supervised contrastive learning model, characterizing the recognition difficulty and misjudgment risk of the large language model in the low-resistivity layer recognition task. The formula for calculating the sample difficulty label is as follows:

[0038] ,

[0039] ,

[0040] in, Let represent the difficulty score of the i-th sub-layer sample. This indicates the degree of discrepancy between the predictions of multiple large language models. This indicates the stability of the misjudged category. The labels indicate the difficulty level of the samples, with 0, 1, and 2 corresponding to easy-to-identify samples, high-risk samples, and extremely difficult samples, respectively. , , For adjustment coefficients, and Thresholds are used to define the difficulty level.

[0041] Furthermore, the formation of the model cognitive map for characterizing the difficulty distribution of large language model recognition includes a supervised contrastive learning model that uses small-level fused feature vectors as input and sample difficulty labels as supervision signals. An encoder maps the small-level samples into low-dimensional embedding representations, and ensures that small-level samples of different difficulty categories form a distinguishable distribution in the embedding space, resulting in a model cognitive map for characterizing the recognition boundaries and risk regions of the large language model. The low-dimensional embedding representation is as follows:

[0042] ,

[0043] in, Let represent the low-dimensional embedding vector of the i-th sub-layer sample. This represents an encoder with parameter θ.

[0044] Furthermore, the supervised contrastive learning model employs a hard sample weighted training mechanism. In each training batch, positive and negative samples of anchor point samples are determined according to the sample difficulty label, and the similarity of well logging time-series features, static geological parameters, and large language model misjudgment difference between the anchor point sample and the candidate negative sample is calculated. When the feature similarity between the candidate negative sample and the anchor point sample reaches a preset similarity threshold but the difficulty label is different or the misjudgment difference threshold exceeds the preset threshold, the separation weight of the candidate negative sample is written into the hard sample weighted contrastive learning loss. The loss function formula of the supervised contrastive learning model is:

[0045] ,

[0046]

[0047] in, Weighted supervised contrastive learning loss representing error difficulty perception. This represents the set of positive samples that have the same difficulty label or the same misclassification category as the anchor sample. Indicates the sample-level difficulty weight. Indicates the clustering weight of positive samples. Indicates the separating weight of the difficult-to-bear samples. This indicates the degree of confusion between negative samples and anchor samples. Indicates error pattern similarity. This represents the temperature coefficient.

[0048] Furthermore, the generation of the low-resistivity layer error pattern library and diagnostic rule library includes obtaining them through cluster analysis of high-risk areas or extremely difficult sample clusters in the model cognitive map. Based on the well logging time-series characteristics, static geological parameters, true oil-water layer labels, large language model prediction results, and misjudgment types of samples within the clusters, error patterns and diagnostic rules with geological interpretation significance are extracted. Candidate rules are then selected based on rule support, rule confidence, and geological interpretation constraints to form the low-resistivity layer error pattern library and diagnostic rule library, resulting in several error pattern clusters. The formula for the low-resistivity layer error pattern library is:

[0049] ,

[0050] The formula for the diagnostic rule base is:

[0051] ,

[0052] in, This represents the difficulty-weighted cluster center of the k-th error pattern cluster. This represents the k-th error pattern cluster. This represents the error pattern cluster number to which the i-th erroneous sample belongs. This represents the set of diagnostic rules corresponding to the k-th error mode. and These represent rule support and rule confidence, respectively. Indicates whether the rule satisfies geological interpretation constraints.

[0053] In summary, the present invention has the following beneficial technical effects:

[0054] 1. This invention unifies well logging curve data, static geological parameters, stratigraphic information, top and bottom depth information of small layers, and true labels of oil and water layers into small-layer samples. Compared with traditional identification methods that rely on only a single well logging curve or a small number of parameters, this invention can simultaneously characterize electrical response, reservoir properties, and geological background in scenarios where the electrical responses of low-resistivity layers and water layers are similar, thereby improving the completeness and consistency of the data expression required for identifying low-resistivity potential layers.

[0055] 2. This invention addresses the problems of weak logging response, subtle curve changes, and strong intra-layer heterogeneity in low-resistivity reservoirs. It extracts time-series features from logging curve segments within a small depth range and integrates them with static parameters such as porosity, permeability, clay content, oil saturation, and small layer thickness. Compared to directly using the original curve or a single statistical parameter, this invention can more fully capture the comprehensive response of low-resistivity reservoirs in terms of electrical properties, physical properties, and curve morphology, thereby improving the accuracy of identifying complex low-resistivity reservoirs.

[0056] 3. This invention uses multiple large language models to predict samples in the same small layer, and constructs a set of incorrect samples and sample difficulty labels based on the difference between the prediction results and the true labels. Compared with traditional intelligent recognition methods that only output the final recognition results, it can discover easily misjudged layers in complex scenarios such as high clay thin layers, low porosity and low permeability, thin interlayers of sandstone and mudstone, and mixed resistivity responses, providing a basis for low resistivity layer risk identification and potential layer verification.

[0057] 4. This invention forms a low-resistivity layer error pattern library and a diagnostic rule library by comparing and learning modeling, clustering analysis and extracting diagnostic rules for high-risk samples and extremely difficult samples. These are then used for subsequent knowledge-enhanced reasoning. Compared with the direct prediction method of ordinary large language models, the discovered misjudgment patterns can be reversed and used in the new small-layer recognition process, improving the stability, traceability and engineering application value of the low-resistivity layer recognition results. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the overall architecture of a low-resistivity intelligent recognition method based on a large language model according to an embodiment of the present invention.

[0059] Figure 2 This is a schematic diagram of the small-level multi-source data preprocessing and fusion feature construction in an embodiment of the present invention.

[0060] Figure 3 This is a schematic diagram illustrating the multi-language model prediction and error sample construction in an embodiment of the present invention.

[0061] Figure 4 This is a schematic diagram of the cognitive map construction of the supervised contrastive learning model according to an embodiment of the present invention.

[0062] Figure 5 This is a schematic diagram of the error pattern library, diagnostic rule library, and knowledge-enhanced reasoning mechanism in an embodiment of the present invention. Detailed Implementation

[0063] Example 1

[0064] Reference Figure 1 This embodiment of a low-resistivity intelligent recognition method based on a large language model includes:

[0065] S1. Obtain multi-source data of small layers related to the low resistivity layer of the target oilfield, and preprocess the logging curve data, static geological parameters, layer information, top and bottom depth information of small layers, and true labels of oil and water layers.

[0066] S2. Based on the preprocessed logging curve data, logging curve segments within the corresponding depth intervals are extracted according to the top and bottom depths of the sub-layers, and statistical features of the logging curves are extracted using the time-series feature extraction method.

[0067] S3. Fuse well logging time series features with static geological parameters to generate a small-level fused feature vector, and convert the small-level fused feature vector into text description information;

[0068] S4. Input the text description information into multiple large language models, and obtain the oil-water layer prediction results of the corresponding small layer samples under the same task instructions and the same inference parameters;

[0069] S5. Construct an error sample set based on the differences between the prediction results of multiple large language models and the true labels of oil and water layers, and generate sample difficulty scores and sample difficulty labels.

[0070] S6. A supervised contrastive learning model is constructed based on the fusion feature vectors of small layers and the sample difficulty labels. The small-layer samples are mapped to a low-dimensional embedding space to form a model cognitive map for representing the distribution of recognition difficulty of large language models.

[0071] S7. Perform cluster analysis and rule extraction on the error sample clustering areas in the model cognitive map, and construct a low-resistivity layer error pattern library and diagnostic rule library;

[0072] S8. Inject the low-resistivity layer error pattern library and diagnostic rule library into the large language model, perform knowledge-enhanced reasoning, and output the low-resistivity layer recognition results, recognition confidence, and explanation information.

[0073] Specifically, a low-resistivity intelligent recognition method based on a large language model includes the following:

[0074] S1. Obtain multi-source data of small-level layers related to the low resistivity layer of the target oilfield, and preprocess the multi-source data of small-level layers.

[0075] like Figure 1 , Figure 2 As shown, in the task of intelligent identification of low-resistivity layers, a single resistivity curve or a small number of static parameters is insufficient to fully reflect the electrical response, reservoir properties, and geological background of the low-resistivity layer. This embodiment uses the sub-layer as the basic identification unit, and unifies and aligns well logging curve data, static geological parameters, stratigraphic information, sub-layer top and bottom depth information, and real labels of oil and water layers to form a multi-source sample at the sub-layer level.

[0076] For the The sample structure of a small layer of samples is represented as follows:

[0077] ,

[0078] in, Indicates the first Small layer samples; Indicates the well number information; Indicates sub-layer or layer information; Indicates the top and bottom depth information of the sub-layer; Represents a set of static geological parameters; This indicates the true label of the oil-water layer.

[0079] The depth information of the top and bottom of the sublayer is represented as follows:

[0080] ,

[0081] in, This represents the top depth of the i-th sublayer sample. This represents the depth of the sublayer of the i-th sublayer sample.

[0082] The set of static geological parameters includes one or more of the following: porosity, permeability, oil saturation, water saturation, clay content, sublayer thickness, lithology, sedimentary facies, or dynamic development parameters:

[0083] ,

[0084] in, Let q represent the j-th static geological parameter of the i-th sublayer sample, and q represent the number of static geological parameters.

[0085] During preprocessing, the following steps are taken: First, well numbers, sub-layer numbers, layer names, top depths, bottom depths, logging sampling depths, and oil-water layer labels are standardized. Second, sub-layer alignment is performed on data from different sources. Third, missing values, outliers, and data with inconsistent units are processed. For missing values, interpolation of adjacent sub-layers within the same well, filling with the mean of the same layer, or filling with expert rules can be used. For outliers, they can be identified and corrected by combining box plots, quantile ranges, or geological experience thresholds. For different dimensional characteristics, standardization or normalization processing can be performed.

[0086] S2. Based on the preprocessed logging curve data, perform small-layer interval truncation and extract logging time-series features.

[0087] like Figure 2 As shown, in this embodiment, based on the top and bottom depths of each sub-layer sample, well logging curve segments within the corresponding depth range are extracted from the original well logging curves. The well logging curve data includes one or more of the following: resistivity curves, natural gamma curves, sonic transit time curves, density curves, and neutron curves. For the first... The logging curve segments of a small layer sample are represented as follows:

[0088] ,

[0089] in, This represents a segment of the well logging curve corresponding to i sub-layer samples; Indicates depth The logging response value at the location.

[0090] After obtaining the small-layer logging curve segment, the Tsfresh feature extraction method is used to automatically extract time-series statistical features from the curve segment. These time-series features include one or more of the following: mean, variance, extreme values, quantiles, autocorrelation coefficient, Fourier coefficients, frequency domain energy, kurtosis, skewness, number of local extrema, and fluctuation amplitude. This step converts the original high-dimensional continuous curve into structured time-series features, reducing the impact of curve noise and sampling differences on subsequent identification results. The extracted logging time-series feature vector is represented as follows:

[0091] ,

[0092] in, This represents the logging time-series feature vector of the i-th sub-layer sample; This represents the time-series feature extraction function; This represents the function for calculating the time series characteristics of the p-th class.

[0093] S3. Integrate well logging time-series characteristics with static geological parameters and convert them into textual description information.

[0094] like Figure 2 , Figure 3 As shown, the extracted well logging time-series features need to be combined with static geological parameters to characterize the sub-layer samples. In this embodiment, the well logging time-series feature vector and the static geological parameter vector are concatenated to obtain the sub-layer fused feature vector, which is represented as follows:

[0095] ,

[0096] in, This represents the sub-layer fusion feature vector of the i-th sub-layer sample; Represents the well logging time series feature vector; Represents a vector of static geological parameters; This indicates a feature concatenation operation. This indicates standardization or normalization.

[0097] After fusion, the numerical fusion features are converted into textual descriptions. These textual descriptions can be in one or more formats: natural language, JSON, or CSV. The text conversion process is represented as follows:

[0098] ,

[0099] in, The text description information represents the i-th sub-layer sample; or This represents a textual transformation function for the features. The textual description information includes at least basic information about the sublayer, porosity, permeability, oil saturation, clay content, sublayer thickness, statistical characteristics of logging curves, and instructions for oil-water layer discrimination.

[0100] S4. Input the text description information into multiple large language models to obtain oil-water layer prediction results.

[0101] like Figure 3 As shown, this embodiment employs multiple large language models to perform parallel predictions on the same small-layer samples. Different large language models output oil-water layer prediction results under the same textual description information, the same task instructions, and the same inference parameters, thereby reducing the impact of accidental misjudgments by a single model on the construction of erroneous samples.

[0102] No. The prediction results of the large language model for the i-th small-layer sample are expressed as follows:

[0103] ,

[0104] in, This represents the oil-water layer prediction result of the m-th large language model for the i-th small layer sample; This represents the m-th large language model; The text description information represents the i-th sub-layer sample; This represents a unified set of inference parameters. The formula for the oil-water layer prediction result set is:

[0105] ,

[0106] in, This represents the set of multi-model prediction results for the i-th sub-layer sample. This represents the predicted category of the m-th large language model. Indicates the prediction confidence level. The reason for the prediction is indicated by M, which represents the number of large language models involved in the prediction.

[0107] By analyzing the outputs of multiple large language models, we can identify low-level samples that are prone to misjudgment from the perspective of consistency and differences between models, thus providing a foundation for the subsequent construction of erroneous samples and the evaluation of sample difficulty.

[0108] S5. Construct an error sample set based on the prediction results and the true labels, and generate sample difficulty labels.

[0109] like Figure 3 As shown, this embodiment compares the prediction results of each large language model with the true labels of the oil-water layer to generate corresponding error labels. The error label function is expressed as:

[0110] ,

[0111] in, This represents the error weight of the m-th large language model for the i-th small-layer sample. Indicates the true label, Indicates the prediction confidence level. This indicates the degree of inconsistency between the reasoning for the prediction and the diagnostic rules. and The total set of erroneous samples is represented by the weighting coefficients:

[0112] ,

[0113] Where E represents the set of erroneous samples misclassified by at least one large language model. M represents the number of large language models involved in the prediction, where M represents the i-th small-layer sample.

[0114] After generating the set of erroneous samples, the number of times each sub-layer sample was misclassified by multiple large language models is further counted. The sample difficulty score is then calculated by combining the divergence between the prediction results of multiple large language models and the stability of the misclassified categories. The formula for calculating the sample difficulty label is as follows:

[0115] ,

[0116] The sample difficulty label generated based on the sample difficulty score is represented as follows:

[0117] ,

[0118] in, Let represent the difficulty score of the i-th sub-layer sample. This indicates the degree of discrepancy between the predictions of multiple large language models. This indicates the stability of the misjudged category. The labels indicate the difficulty level of the samples, with 0, 1, and 2 corresponding to easy-to-identify samples, high-risk samples, and extremely difficult samples, respectively. , , For adjustment coefficients, and Difficulty thresholds are set. Indicates an indicator function.

[0119] S6. Construct a supervised contrastive learning model based on sample difficulty labels and generate a model cognitive map.

[0120] like Figure 4As shown, this embodiment uses the fused feature vectors of small layers as input and the sample difficulty labels as supervision signals to construct a supervised contrastive learning model. The model can use a multilayer perceptron, a Transformer encoder, a graph neural network encoder, or a combination thereof as the encoder, and maps each small-layer sample to a low-dimensional embedding representation.

[0121] The low-dimensional embedding representation formula is:

[0122] ,

[0123] in, Let represent the low-dimensional embedding vector of the i-th sub-layer sample. This represents an encoder with parameter θ.

[0124] In each training batch, positive and negative samples of anchor point samples are determined according to the sample difficulty label. Samples with the same difficulty label or the same misclassification category as anchor point samples are considered positive samples, and samples with different difficulty labels are considered negative samples. For candidate negative samples, the similarity of well logging time series features, static geological parameters, and large language model misclassification differences between them and anchor point samples are further calculated. In this embodiment, the similarity of well logging time series features, static geological parameters, and large language model misclassification differences are all normalized to a value range of 0 to 1. The comprehensive feature similarity between candidate negative samples and anchor point samples is determined based on the similarity of well logging time series features and static geological parameters. The preset similarity threshold can be set to 0.75 to 0.90, preferably 0.80; the preset threshold corresponding to the misclassification difference can be set to 0.40 to 0.60, preferably 0.50. When the feature similarity between a candidate negative sample and an anchor sample reaches a preset similarity threshold, and the difficulty labels of the two samples are different, or the difference in misjudgment exceeds a preset threshold, the candidate negative sample is identified as a difficult negative sample, and its separation weight is written into the difficult sample weighted contrastive learning loss, so that the training process focuses on distinguishing easily confused samples.

[0125] The loss function formula for the supervised contrastive learning model is:

[0126] ,

[0127]

[0128] in, Weighted supervised contrastive learning loss representing error difficulty perception. This represents the set of positive samples that have the same difficulty label or the same misclassification category as the anchor sample. Indicates the sample-level difficulty weight. Indicates the clustering weight of positive samples. Indicates the separating weight of the difficult-to-bear samples. This indicates the degree of confusion between negative samples and anchor samples. Indicates error pattern similarity. This represents the temperature coefficient.

[0129] During training, the encoder parameters are iteratively updated according to the aforementioned loss function, and the low-dimensional embedding vectors after training are visualized in two or three dimensions to form a model cognitive map. In the model cognitive map, easily identifiable samples, high-risk samples, and extremely difficult samples form different distribution regions, which are used to characterize the safe region, risk region, and extremely difficult sample clustering region in low-resistivity layer identification.

[0130] S7. Perform cluster analysis and rule extraction on the error sample clustering areas in the model's cognitive map.

[0131] like Figure 5 As shown, this embodiment performs cluster analysis on high-risk areas or clusters of extremely difficult samples in the model's cognitive map. The clustering objects are the low-dimensional embedding vectors corresponding to the error samples or high-risk samples. To ensure that extremely difficult and high-risk samples have a higher influence in the error pattern formation process, this embodiment uses a difficulty-weighted method to calculate the cluster centers of error pattern clusters.

[0132] The formula for the low-resistivity layer error mode library is:

[0133] ,

[0134] After clustering, for each error pattern cluster, the well logging time series characteristics, static geological parameters, true labels, large language model prediction results, misjudgment types and prediction reasons of the samples within the cluster are statistically analyzed. Then, decision trees, rule induction models, threshold screening or expert review are used to extract diagnostic rules with geological interpretation significance.

[0135] The formula for the diagnostic rule base is:

[0136] ,

[0137] in, This represents the difficulty-weighted cluster center of the k-th error pattern cluster. This represents the k-th error pattern cluster. This represents the error pattern cluster number to which the i-th erroneous sample belongs. This represents the set of diagnostic rules corresponding to the k-th error mode. and These represent rule support and rule confidence, respectively. Indicates whether the rule satisfies geological interpretation constraints.

[0138] Through the above processing, a low-resistivity layer error pattern library and a diagnostic rule library can be formed. The error patterns include one or more of the following: high-muddy thin-layer low-resistivity misjudgment patterns, low-porosity and low-permeability low-resistivity misjudgment patterns, mixed resistivity misjudgment patterns of thin interbedded sandstone and mudstone layers, curve-smooth pseudo-low-resistivity misjudgment patterns, misjudgment patterns of indistinct electrical differences between oil and water layers, or missed judgment patterns for a few types of water layers. Each type of error pattern corresponds to several geologically meaningful diagnostic rules.

[0139] S8. Inject the low-resistivity layer error pattern library and diagnostic rule library into the large language model to perform knowledge-enhanced reasoning.

[0140] like Figure 5 As shown, after obtaining the low-resistivity layer error pattern library and diagnostic rule library, this embodiment converts them into prompt information, retrieval enhancement knowledge fragments, or rule constraint text that can be called by the large language model. For the small-layer sample to be identified, the corresponding small-layer fusion features and text description information are generated according to steps S1 to S3, and it is determined whether its position in the model's cognitive map is close to a high-risk area or an area where extremely difficult samples are clustered.

[0141] When a sample in a small layer to be identified meets the combination of high-risk features in the diagnostic rule base, the corresponding error pattern, diagnostic rules, and risk warnings are added to the inference input of the large language model, so that the large language model can refer to historical misjudgment patterns and geological interpretation rules when distinguishing between oil and water layers.

[0142] During the knowledge-enhanced reasoning process, the large language model combines the textual descriptions of the sub-layer samples to be identified, the low-resistivity layer error pattern library, the diagnostic rule library, and risk warnings to output the low-resistivity layer identification results, identification confidence, misjudgment risk warnings, and traceable geological interpretation information. If the sub-layer sample to be identified is close to a high-risk area in the model's cognitive map, a review suggestion is also provided in the output to assist interpreters in further evaluating potential low-resistivity layers.

[0143] Example 2

[0144] The difference between this embodiment and Embodiment 1 is that this embodiment further provides model verification and prediction effect analysis to illustrate the recognition accuracy, stability and engineering applicability of the low-resistivity layer intelligent recognition method based on error sample contrast learning and knowledge enhancement reasoning in actual small layer samples.

[0145] This embodiment selects small-scale samples from the target oilfield as test data. The logging time-series features, static geological parameters, and true oil-water layer labels corresponding to each small-scale sample are used as input. Comparative verification is performed using a basic large language model without contrastive learning, a conventional machine learning baseline model, and the large language model integrating contrastive learning and knowledge-enhanced reasoning as described in this invention. Here, Baseline LLM represents the basic large language model, ML Base represents the conventional machine learning baseline model, and CL-KGE LLM represents the large language model integrating contrastive learning and knowledge-enhanced reasoning. Evaluation metrics include Accuracy, Macro-F1, and Oil Recall.

[0146] Table 1. Experimental Results Comparing Model Recognition Performance

[0147] Accuracy 0.8446 0.9092 0.9208 Macro-F1 0.7950 0.8214 0.8823 Oil Recall 0.9341 0.9464 0.9528

[0148] As shown in Table 1, the CL-KGE LLM in this embodiment achieves an accuracy of 0.9208, higher than the baseline LLM's 0.8446 and ML Base's 0.9092; the Macro-F1 score improves from 0.7950 and 0.8214 to 0.8823; and the oil layer recall reaches 0.9528. These results demonstrate that the supervised contrastive learning model trained based on a set of erroneous samples and sample difficulty labels can enhance the ability to distinguish between low-resistivity oil layers and water layer boundary samples, and improve the overall recognition stability of minority classes and easily confused samples.

[0149] Furthermore, the CL-KGE LLM incorporates misjudgment patterns extracted from the model's cognitive map, including high-muddy thin-layer low-resistivity misjudgment patterns, low-porosity and low-permeability low-resistivity misjudgment patterns, mixed resistivity misjudgment patterns of sandstone and mudstone thin interbedded layers, and curve-smooth pseudo-low-resistivity misjudgment patterns, into the inference process. This allows the identification of the sub-layer to provide misjudgment risk warnings and traceable explanations while outputting the oil-water layer category. Therefore, this embodiment can verify the accuracy, stability, and engineering applicability of this embodiment in the intelligent identification task of low-resistivity layers.

[0150] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent changes made to the structure, method, model construction, rule extraction, or knowledge-enhanced reasoning methods of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A low-resistivity intelligent recognition method based on a large language model, characterized in that, include: Acquire multi-source data of small-level layers related to low resistivity layers in the target oilfield, and preprocess the multi-source data of small-level layers, which includes well logging curve data, static geological parameters and true labels of oil and water layers; Based on the preprocessed logging curve data, logging curve segments within the corresponding depth range are extracted according to the top and bottom depths of the sub-layer, and time-series features are extracted from the logging curve segments. The extracted well logging time-series features are fused with static geological parameters to obtain a small-level fused feature vector; The feature vectors of the small-level fusion are transformed into text description information and input into multiple large language models to obtain the oil-water layer prediction results of the corresponding small-level samples. Based on the difference between the actual labels of the oil-water layer and the predicted oil-water layer, a set of erroneous samples is constructed, and sample difficulty labels are generated. The construction of the error sample set includes marking a sub-level sample as an error sample when at least one large language model makes a prediction error, and recording the misjudged model number, misjudgment category, prediction reason, and confidence information; when multiple large language models misjudge the same sub-level sample, the sub-level sample is further marked as a common misjudged sample or an extremely difficult sample, thereby forming a hierarchical error sample set for error pattern mining and sample difficulty evaluation. The error sample set is constructed using the following error labeling function: , in, This represents the error weight of the m-th large language model for the i-th small-layer sample. Indicates the true label, Indicates the prediction confidence level. This indicates the degree of inconsistency between the reasoning behind the prediction and the diagnostic rules. and The total set of erroneous samples is represented by the weighting coefficients: , Where E represents the set of erroneous samples misclassified by at least one large language model. M represents the number of large language models involved in the prediction, where M represents the i-th small-layer sample. The sample difficulty label is generated based on the number of misclassifications by multiple large language models for each sub-layer sample, the divergence of prediction results, and the stability of misclassification categories. Sub-layer samples are divided into easily identifiable samples, high-risk samples, and extremely difficult samples. This sample difficulty label serves as a supervisory signal for the supervised contrastive learning model, characterizing the recognition difficulty and misclassification risk of the large language model in the low-resistivity layer recognition task. The formula for calculating the sample difficulty label is as follows: , , in, Let represent the difficulty score of the i-th sub-layer sample. This indicates the degree of discrepancy between predictions from multiple large language models. This indicates the stability of the misjudged category. The labels indicate the difficulty level of the samples, with 0, 1, and 2 corresponding to easy-to-identify samples, high-risk samples, and extremely difficult samples, respectively. , , For adjustment coefficients, and Difficulty thresholds are set. Indicates an indicator function; A supervised contrastive learning model is constructed based on the fusion feature vectors of small layers and the sample difficulty labels. The small-layer samples are mapped to a low-dimensional embedding space to form a model cognitive map that represents the distribution of recognition difficulty of large language models. Cluster analysis and rule extraction are performed on error sample clustering regions in the low-dimensional embedding space to generate a low-resistivity layer error pattern library and a diagnostic rule library; The low-resistivity layer error pattern library and diagnostic rule library are injected into the large language model to perform knowledge-enhanced reasoning and output the low-resistivity layer recognition results.

2. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The preprocessing of the multi-source data at the sub-layer level includes: truncating the logging curve data into sub-layer intervals using the top and bottom depths of the sub-layer as boundaries; associating the truncated logging curve data with the static geological parameters and true labels of the oil and water layers of the corresponding sub-layer; and performing missing value imputation, outlier handling, and standardization on the logging curves and static geological parameters to obtain unified multi-source sample data at the sub-layer level. The multi-source data at the sub-layer level is constructed in the following form: , in, This represents the i-th sub-layer sample. Indicates the hash number information. This indicates sub-layer or layer information. This indicates the top and bottom depth information of a small layer. Represents a set of static geological parameters. This indicates the true label of the oil-water layer.

3. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The method involves extracting logging curve segments within corresponding depth intervals based on the top and bottom depths of the sub-layers. This is achieved by extracting curve data within the corresponding depth intervals from the pre-processed continuous logging curves. The logging curves include one or more of the following: resistivity curves, natural gamma curves, sonic transit time curves, density curves, and neutron curves. By extracting segments within sub-layers, the original continuous logging curves are converted into logging curve segments corresponding one-to-one with the sub-layers. The formula for extracting these logging curve segments is as follows: , in, Indicates the first The logging curve segments corresponding to each small layer sample Indicates the depth of the top of the small layer. Indicates the depth of the bottom layer. Indicates depth The logging response value at the location.

4. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The step of extracting time-series features from well logging curve segments includes truncating the well logging curve segments with the top and bottom depths of the sub-layers as boundaries, using the Tsfresh feature extraction method to calculate the statistical features, distribution features, trend features, and fluctuation features of the well logging curve segments, and filtering them based on feature missing rate, variance, correlation, and the degree of association with the true labels of oil and water layers to obtain a well logging time-series feature vector for low-resistivity layer identification. The formula for generating the well logging time-series feature vector is as follows: , in, Let φ(·) represent the logging time-series feature vector of the i-th sub-layer sample, and let φ(·) represent the time-series feature extraction function. This represents the p-th type of time series feature; the time series feature includes one or more of the following: mean, variance, extreme values, quantiles, autocorrelation coefficient, Fourier coefficient, frequency domain energy, kurtosis, skewness, number of local extrema, and fluctuation amplitude.

5. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The process of fusing the extracted well logging time-series features with static geological parameters to obtain a small-level fused feature vector includes: taking small-layer samples as a unified object, associating the well logging time-series features extracted from well logging curve segments with one or more static geological parameters from porosity, permeability, oil saturation, water saturation, clay content, small-layer thickness, or lithology; unifying fields and correcting numerical ranges for features from different sources; and forming a small-level fused feature vector that can be used simultaneously for natural language description generation, large language model prediction error analysis, and supervised contrastive learning modeling. The formula for constructing the small-level fused feature vector is as follows: , in, This represents the sub-layer fusion feature vector of the i-th sub-layer sample. Represents the well logging time series feature vector. This represents a static geological parameter vector, and [;] represents a feature concatenation operation. This indicates standardization or normalization.

6. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The process of converting the sub-level fusion feature vector into text description information includes parsing the well logging statistical features and static geological parameters in the sub-level fusion feature vector according to a preset text template, and converting continuous numerical parameters into corresponding geological grade descriptions. The geological grade descriptions include one or more of the following: low, medium, high, good, poor, large fluctuation, or small fluctuation. The field parsing results and geological grade descriptions are combined to form natural language descriptions and structured JSON descriptions, enabling the large language model to understand the electrical response, reservoir properties, and oil-water characteristics of the sub-layers based on the text description information without receiving the actual labels of the oil and water layers.

7. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The process of obtaining the oil and water layer prediction results for the corresponding sub-layer sample includes inputting the text description information of the same sub-layer sample into various language models, and obtaining the prediction results of oil and water layer categories under the same task instructions, the same output format constraints, and the same inference parameters. The outputs of various language models are formatted and parsed to extract prediction categories, prediction rationales, and confidence information. Based on the prediction consistency and differences among multiple language models, a multi-model prediction result set is formed for constructing erroneous samples and evaluating sample difficulty. The formula for the oil-water layer prediction result set is: , in, This represents the set of multi-model prediction results for the i-th sub-layer sample. This represents the predicted category of the m-th large language model. Indicates the prediction confidence level. Indicate the reasons for the prediction. This indicates the number of large language models involved in the prediction.

8. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The formation of the model cognitive map for characterizing the recognition difficulty distribution of a large language model includes a supervised contrastive learning model that uses low-level fusion feature vectors as input and sample difficulty labels as supervision signals. An encoder maps low-level samples to low-dimensional embedding representations, and ensures that low-level samples of different difficulty categories form a discriminative distribution in the embedding space, resulting in a model cognitive map for characterizing the recognition boundaries and risk regions of a large language model. The low-dimensional embedding representation is as follows: , in, Let represent the low-dimensional embedding vector of the i-th sub-layer sample. This represents an encoder with parameter θ.

9. The low-resistivity intelligent recognition method based on a large language model according to claim 8, characterized in that, The supervised contrastive learning model employs a hard sample weighted training mechanism. In each training batch, positive and negative samples of anchor point samples are determined according to the sample difficulty label, and the similarity of well logging time-series features, static geological parameters, and large language model misjudgment difference between the anchor point sample and the candidate negative sample is calculated. When the feature similarity between the candidate negative sample and the anchor point sample reaches a preset similarity threshold but the difficulty label is different or the misjudgment difference threshold exceeds the preset threshold, the separation weight of the candidate negative sample is written into the hard sample weighted contrastive learning loss. The loss function formula of the supervised contrastive learning model is as follows: , , in, Weighted supervised contrastive learning loss representing error difficulty perception. This represents the set of positive samples that have the same difficulty label or the same misclassification category as the anchor sample. Indicates the sample-level difficulty weight. Indicates the clustering weight of positive samples. Indicates the separating weight of the difficult-to-bear samples. This indicates the degree of confusion between negative samples and anchor samples. Indicates error pattern similarity. This represents the temperature coefficient.

10. The low-resistivity intelligent recognition method based on a large language model according to claim 1, characterized in that, The generation of the low-resistivity layer error pattern library and diagnostic rule library includes clustering analysis of high-risk areas or extremely difficult sample clusters in the model cognitive map. Based on the well logging time-series characteristics, static geological parameters, true oil-water layer labels, large language model prediction results, and misjudgment types of samples within the clusters, error patterns and diagnostic rules with geological interpretation significance are extracted. Candidate rules are then selected based on rule support, rule confidence, and geological interpretation constraints to form the low-resistivity layer error pattern library and diagnostic rule library, resulting in several error pattern clusters. The formula for the low-resistivity layer error pattern library is: , The formula for the diagnostic rule base is: , in, This represents the difficulty-weighted cluster center of the k-th error pattern cluster. This represents the k-th error pattern cluster. This represents the error pattern cluster number to which the i-th erroneous sample belongs. This represents the set of diagnostic rules corresponding to the k-th error mode. and These represent rule support and rule confidence, respectively. Indicates whether the rule satisfies geological interpretation constraints.

Citation Information

Patent Citations

  • Knowledge tracking method for student and question joint modeling based on large language model

    CN121327154A

  • Test question knowledge point labeling method and system based on multi-agent collaboration

    CN121706964A