Thin sandstone layer logging lithofacies prediction method based on machine learning

By constructing a well logging lithofacies prediction model for thin sandstone layers based on convolutional neural networks and Transformers, the problem of insufficient thin-layer identification capability in lithofacies prediction of thin sandstone layers is solved, the accuracy of inter-well prediction and the robustness of the model are improved, the geological rationality and interpretability are enhanced, and the accuracy of oil and gas exploration and development is improved.

CN121919681APending Publication Date: 2026-04-24CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (WUHAN)
Filing Date
2025-11-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing thin sandstone lithofacies prediction technologies suffer from problems such as insufficient thin-layer identification ability, scarce and unbalanced samples, insufficient model generalization ability, insufficient interpretability and utilization of geological prior constraints, and insufficient practical deployment and robustness. In particular, it is difficult to balance accuracy and robustness under complex reservoir conditions.

Method used

A convolutional neural network is used as a local feature encoder, combined with a Transformer encoder. Through adaptive high-frequency enhancement processing and dynamic channel gating to correct wellbore anomalies, geological prior attention is introduced to construct a well logging lithofacies prediction model for thin sandstone layers, thereby improving the saliency of thin layer features and vertical smoothness.

Benefits of technology

It effectively solves the problem of insufficient thin-layer identification capability, improves the accuracy of inter-well prediction and the robustness of the model, reduces cross-phase misjudgment, provides geological rationality and interpretability, and enhances the accuracy of oil and gas exploration and development in thin sandstone layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121919681A_ABST
    Figure CN121919681A_ABST
Patent Text Reader

Abstract

The invention discloses a thin sandstone layer well logging lithofacies prediction method based on machine learning, and relates to the technical field of well logging lithofacies prediction.The method comprises the steps that well diameter abnormal data are recognized based on a CAL well logging curve, and the feature contribution degree of the well diameter abnormal data is corrected; performing self-adaptive high-frequency enhancement processing on the high-frequency response features in the corrected logging feature data set to obtain the logging feature data set; a thin sandstone layer logging lithofacies prediction model based on machine learning is constructed, and a convolutional neural network is used as a local feature encoder; a Transform encoder is used as a global feature encoder, and geological prior of logging depth is introduced into attention calculation; a logging data set is input into the local feature encoder to extract local features, the feature fusion module fuses the local features and then inputs the fused local features into the global feature encoder to establish a global dependency relationship of logging data, and the classifier outputs logging lithofacies prediction of the thin sandstone layer. According to the method, the thin sandstone layer logging lithofacies identification capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of well logging facies prediction technology, and in particular to a well logging facies prediction method for thin sandstone layers based on machine learning. Background Technology

[0002] Thin sandstone layers, as important reservoir units in marine and continental oil and gas reservoirs, are characterized by their thin thickness, poor continuity, rapid lithological changes, and strong reservoir heterogeneity, making them crucial for oil and gas exploration and development. Lithofacies prediction is a key step in reservoir evaluation and sweet spot identification, effectively improving the discovery rate and development efficiency of oil and gas in thin sandstone layers. Currently, lithofacies interpretation mainly relies on data from well logging, core samples, and conventional logging curves. However, core samples are costly to obtain, sparsely distributed between wells, and logging data is subjective; therefore, logging response characteristics have become an important basis for lithofacies identification. For thin sandstone layers with a thickness of only tens of centimeters to several meters, limitations in logging instrument resolution, interference from clay interlayers, and curve noise lead to blurred lithological interfaces and weakened lithofacies characteristics, resulting in low accuracy of traditional interpretation methods.

[0003] Extensive research has been conducted both domestically and internationally on the development of lithofacies prediction technology. Early methods primarily relied on manual experience and statistical approaches, such as distinguishing lithologies like sandstone, mudstone, and carbonate rocks through core thin-section comparison, well logging curve interpretation, and cross-plot analysis. However, in thin interbedded layers or complex sedimentary blocks, the phenomenon of "different well logs for the same rock and the same well logs for different rocks" frequently occurs, leading to high misclassification rates and insufficient accuracy of manual methods. With the improvement of computing power and data volume, machine learning algorithms such as Support Vector Machine (SVM), Random Forest (RF), and K-Nearest Neighbors (KNN) have been introduced into lithofacies classification, achieving recognition rates of approximately 75%–85% in specific regions or on high-quality datasets. These methods can automatically discover some nonlinear features, but they rely on manual feature selection, are sensitive to data noise and missing data, and have poor generalization effects across wells or strata.

[0004] In recent years, deep learning models such as convolutional neural networks (CNNs), recurrent neural networks (such as LSTMs), and hybrid structures (such as CNN-Transformer and CNN-BiLSTM) have been gradually applied to the prediction of lithofacies or reservoir parameters. For example, in the prediction of porosity in carbonate reservoirs, the CNN-Transformer structure can significantly reduce prediction errors compared to a single CNN or LSTM (Long Short Term Memory) recurrent neural network, improving the coefficient of determination R² by 0.2–0.3. Some studies have also explored the long-range dependency modeling capabilities based on Transformer to improve the recognition of the continuity of curves throughout the well. However, in complex reservoirs, there are no publicly available cases of systematically applying the CNN-Transformer architecture for lithofacies classification. Existing research mainly focuses on reservoir attribute prediction (such as porosity and permeability) rather than end-to-end lithofacies prediction.

[0005] Despite some progress in existing technologies, the following prominent problems still exist in the prediction of lithofacies in thin sandstone layers: (1) Insufficient thin-layer identification capability: Due to the limitation of logging curve resolution, thin-layer lithofacies features are easily smoothed or mixed, and the model is difficult to accurately depict its subtle changes.

[0006] (2) Sample scarcity and class imbalance: Core data and thin section annotations are limited, and the sample size of rare lithofacies (such as locally developed thin sandstone) is too small, resulting in the model's weak ability to identify minority classes.

[0007] (3) Insufficient generalization ability of the model: There are significant differences in geological background between different wells. CNN and other models are sensitive to local features but lack long-distance dependence capture. There is almost no research on Transformer generalization and transfer learning across wells and strata. The model performance degrades on untrained wells.

[0008] (4) Insufficient interpretability and utilization of geological prior constraints: Although depth models can improve accuracy, they are difficult to explain which logging curves or depth segments play a dominant role in classification and lack effective integration with sedimentary laws and sequence stratigraphic characteristics.

[0009] (5) Insufficient practical deployment and robustness: When real-time prediction is performed downhole or logging conditions are complex, the stability of the model has not been verified and there is a lack of deployable solutions for engineering purposes.

[0010] In summary, existing lithofacies prediction technologies for thin sandstone layers struggle to balance accuracy and robustness, especially under complex reservoir conditions, where significant room for improvement remains. With the development of machine learning and geological modeling techniques, there is an urgent need for an intelligent lithofacies prediction method that can enhance the identification capability of thin sandstone layers and improve inter-well prediction accuracy, in order to better guide oil and gas exploration and development. Summary of the Invention

[0011] The purpose of this invention is to address the insufficient thin-layer identification capability of existing well logging facies prediction technologies, and to propose a machine learning-based well logging facies prediction method for thin sandstone layers, comprising the following steps: S1. Acquire multiple well logging data and preprocess them to construct a well logging feature dataset; S2. Identify well diameter anomaly data in the well logging feature dataset based on CAL logging curves, and correct the feature contribution of the well diameter anomaly data to obtain the corrected well logging feature dataset. S3. Adaptive high-frequency enhancement processing is performed on the high-frequency response features in the corrected well logging feature dataset to obtain the final well logging feature dataset. S4. Construct a machine learning-based well logging lithofacies prediction model for thin sandstone layers, including: a local feature encoder, a feature fusion module, a global feature encoder, and a classifier; Among them, a convolutional neural network is used as a local feature encoder; and a Transformer encoder is used as a global feature encoder. The final logging dataset is input into the local feature encoder to extract local features. The feature fusion module fuses the local features and inputs them into the global feature encoder to establish the global dependency relationship of the logging data. The classifier outputs the logging lithofacies prediction of the thin sandstone layer. S5. Predict the facies of the thin sandstone layer using the well logging facies prediction model.

[0012] Furthermore, logging characteristics include: natural gamma, spontaneous potential, well diameter, density, sonic transit time, formation true resistivity, and resistivity of the flushed zone.

[0013] Furthermore, the preprocessing includes: The well logging data is depth aligned, and the sampling step size is unified through interpolation. The lithofacies type label is also determined. The textual descriptions of lithofacies types are converted into digital labels, and a mapping relationship between lithofacies categories and digital labels is established. Statistical analysis of the sample distribution of various lithofacies was performed, and rare categories with fewer than a set threshold of samples were removed. The IQR method is used to identify and remove outlier data points; Each logging parameter was standardized to have a mean of 0 and a variance of 1.

[0014] Furthermore, based on CAL logging curves, well diameter anomaly data in the logging feature dataset are identified, and a dynamic channel gating strategy is used to correct the feature contribution of the well diameter anomaly data, resulting in the corrected logging feature dataset, represented as:

[0015]

[0016] in, This represents the i-th corrected well logging feature data. This represents the i-th well logging feature data. This indicates depth-wise, channel-wise broadcast multiplication. Indicates the gating weight, Let W represent the sigmoid activation function, and W be a learnable vector. It is the well diameter threshold. This represents the well diameter displayed in the CAL logging curve for the i-th logging feature data.

[0017] Furthermore, adaptive high-frequency enhancement processing is performed on the high-frequency response features in the corrected well logging feature dataset to obtain the final well logging feature dataset, represented as:

[0018]

[0019] in, This represents the i-th final well logging feature data. This represents the i-th corrected well logging feature data. Represents the learnable coefficient. Indicates high-frequency residuals, This represents a low-frequency smoothing filter with a kernel size of k.

[0020] Furthermore, a geological prior to the logging depth is introduced into the attention calculation of the Transformer encoder. The attention based on this prior is as follows:

[0021]

[0022] Where τ is the attenuation factor, This represents the depth difference between adjacent depth points i and j in well logging.

[0023] The present invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described machine learning-based method for predicting lithofacies in thin sandstone layers.

[0024] The present invention also proposes an electronic device, including a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program including computer-readable instructions, and the processor is configured to invoke the computer-readable instructions to execute the above-described machine learning-based logging lithofacies prediction method for thin sandstone layers.

[0025] The present invention also proposes a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-described machine learning-based method for predicting lithofacies in thin sandstone layers.

[0026] The beneficial effects of the technical solution provided by this invention are: This invention provides a machine learning-based method for predicting the lithofacies of thin sandstone layers in well logging. It constructs a prediction model using a convolutional neural network as a local morphological feature encoder and a Transformer encoder as a global sequence dependency modeler. Channel gating corrects the feature contribution of anomalous wellbore data, effectively reducing interference from wellbore collapse and wellbore enlargement on the identification of thin sandstone layer boundaries. Adaptive high-frequency enhancement processing is applied to the unique high-frequency response features of thin sandstone layers, effectively preserving the peak responses and subtle texture features of the thin sandstone bodies and improving the saliency of thin-layer features. The geological prior of the vertical depth of the well logging is introduced into the attention calculation of the Transformer encoder, significantly reducing "cross-facies" misjudgments in lithofacies identification and improving the vertical smoothness and geological rationality of the prediction results. This invention effectively solves the problem of insufficient thin-layer identification capability in existing well logging lithofacies prediction technologies. Attached Figure Description

[0027] Figure 1 This is a flowchart of a machine learning-based well logging lithofacies prediction method for thin sandstone layers according to an embodiment of the present invention; Figure 2 This is a statistical map of lithofacies distribution in the training dataset of this invention embodiment; Figure 3 It is a distribution characteristic map of training set logging parameters among different lithofacies; Figure 4 It is a heatmap of the Pearson correlation between logging parameters in the training set; Figure 5 It is a graph showing the changes in loss and accuracy during the model training and validation process; Figure 6 It is a test set of lithofacies classification confusion matrix diagram; Figure 7 This is a block diagram of an electronic device according to an exemplary embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0029] The flowchart of the machine learning-based well logging lithofacies prediction method for thin sandstone layers in this embodiment of the invention is as follows: Figure 1 Specifically, it includes the following steps: S1. Acquire multiple well logging data and preprocess them to construct a well logging feature dataset.

[0030] Well logging characteristic data include: natural gamma, spontaneous potential, well diameter, density, sonic transit time, formation true resistivity, and flushed zone resistivity.

[0031] Preprocessing includes: The well logging data were subjected to depth alignment and step size unification interpolation to unify the well logging data with different vertical sampling intervals to a fixed sampling interval, so that the data from different wells have a consistent vertical sampling scale. Combined with existing core lithofacies description data, the main lithofacies types and their corresponding well logging response characteristics were determined, providing a geological basis for the construction of the training dataset.

[0032] Tag encoding process converts the lithofacies types described in the text into digital tags, establishing a mapping relationship between lithofacies categories and digital tags.

[0033] Rare category processing involves statistically analyzing the sample distribution of various lithofacies, removing rare categories with fewer than a set threshold of samples, and recoding the retained categories.

[0034] Outlier detection and removal utilizes the IQR (Interquartile Range) method to identify outliers, and its application is illustrated with a Python code snippet. This method identifies and removes outlier data points, ensuring data quality.

[0035] Data standardization processing involves standardizing each logging parameter to a mean of 0 and a variance of 1, ensuring that different logging parameters participate in model training on the same scale.

[0036] S2. Based on the CAL logging curve, identify the well diameter anomaly data in the logging feature dataset, and correct the feature contribution of the well diameter anomaly data to obtain the corrected logging feature dataset.

[0037] A wellbore anomaly identification mechanism is established based on CAL logging curves. An upper limit threshold for wellbore diameter is set. When the wellbore diameter value in the CAL logging curve exceeds the threshold range, it is identified as a wellbore enlargement anomaly segment. For the identified anomaly segment samples, a dynamic channel gating strategy is adopted in the subsequent feature extraction process. A learnable gating weight matrix is ​​used to correct the feature contribution of the anomaly channel, thereby effectively reducing the interference caused by wellbore collapse and wellbore enlargement on the identification of thin sand layer boundaries, improving the model's signal-to-noise ratio and discrimination ability for thin layer boundaries in anomaly segments, and enhancing the signal-to-noise ratio. This can be expressed by the formula:

[0038]

[0039] in, This represents the i-th corrected well logging feature data. This represents the i-th well logging feature data. This represents depth-wise, channel-wise broadcast multiplication, with a weight matrix. It will automatically expand along the depth direction and the characteristic channel direction, in conjunction with well logging characteristic data. Perform element-wise multiplication. Indicates the gating weight, Let W represent the sigmoid activation function, and W be a learnable vector. It is the well diameter threshold. This represents the well diameter displayed in the CAL logging curve for the i-th logging feature data.

[0040] S3. Adaptive high-frequency enhancement processing is performed on the high-frequency response features in the corrected well logging feature dataset to obtain the final well logging feature dataset.

[0041] Thin sand layers often exhibit high-frequency spikes at their interfaces, information that traditional convolution and downsampling operations easily lose. To address the unique high-frequency response characteristics of thin sand layers in corrected logging feature data, adaptive high-frequency enhancement processing is performed within a local short window scale. Specifically, the difference between the logging curve and its local moving average is calculated to extract the high-frequency residual component. An adjustable amplification factor is used to enhance the high-frequency signal, and the enhanced high-frequency residual signal is weighted and fused with the corrected logging feature data as input features for subsequent networks. This effectively preserves the spike response and subtle texture features of thin sand bodies, avoiding excessive smoothing of thin-layer information due to increased receptive field during the forward propagation of the convolutional neural network. The formula is:

[0042]

[0043] in, This represents the i-th final well logging feature data. This represents the i-th corrected well logging feature data. Represents the learnable coefficient. Indicates high-frequency residuals, This represents a low-frequency smoothing filter with a kernel size of k.

[0044] Design principle: Through feature calculation Its low-frequency smoothing version residual To obtain high-frequency information, and through learnable coefficients The residual signals are weighted and fused. This module ensures responsiveness to rapidly changing boundaries of thin layers during subsequent convolutional encoding.

[0045] S4. Construct a machine learning-based well logging lithofacies prediction model for thin sandstone layers, including: a local feature encoder, a feature fusion module, a global feature encoder, and a classifier.

[0046] In this study, a convolutional neural network is used as the local feature encoder. First, a multi-scale Conv1D convolutional group (kernel size k∈{3,5,7}) is employed to extract features at each scale:

[0047] in, denoted as the local logging feature tensor extracted using an m-scale convolution kernel, used to characterize lithological response patterns under different spatial receptive fields, and f denotes the feature mapping function resulting from linear convolution of the input sequence and superposition with nonlinear activation. Indicates that the convolution kernel is Convolution operation, This represents a local window sequence of logging curves centered at depth i and with a length of n, which serves as the input to the local feature encoder.

[0048] Features at each scale are assigned learnable weights Fusion:

[0049] in, This represents the fused features of the output of the convolutional neural network. express The weight.

[0050] A Transformer encoder is used as the global feature encoder, and a geological prior of vertical well logging depth is introduced into the attention calculation of the Transformer encoder. By assigning higher attention weights to adjacent depth points and suppressing excessive attention to distant non-adjacent points, the model prioritizes short-distance sequence correlations when capturing stratigraphic sequence features, reducing misclassification of skipped layers. The introduced prior attention is as follows:

[0051]

[0052] in, It is the attention score, Q is the query vector, and K is the key vector. τ is the scaling factor, and τ is the decay factor. As a depth continuity factor, This represents the depth difference between adjacent depth points i and j in well logging. Design principle: Incorporate the depth continuity factor... The logarithm of the value is added as a bias term to the standard attention score, forcing the model to prioritize neighboring depth points.

[0053]

[0054]

[0055]

[0056]

[0057] in, Indicates attention output, These are the weight matrices for Q, K, and V, respectively.

[0058] The final logging dataset is input to the local feature encoder to extract local features. The feature fusion module fuses the local features and then inputs them into the global feature encoder to establish global dependencies of the logging data. The global features obtained by the global feature encoder are input into the classifier, and the classifier outputs the logging lithofacies prediction for thin sandstone layers.

[0059] S5. Predict the facies of the thin sandstone layer using the well logging facies prediction model.

[0060] During model training, the AdamW optimizer is used for parameter updates, and the learning rate is dynamically adjusted using the OneCycleLR learning rate scheduling strategy to improve model convergence efficiency and stability. To address the issue of insufficient samples in the thin sandstone category, the Focal Loss loss function is introduced. By adjusting the class weight parameters and focus factor parameters, the impact of class imbalance on model training is effectively mitigated. At the same time, a Bayesian optimization framework (e.g., using the Optuna tool) is used to automatically search for key hyperparameters of the model, including the number of Transformer heads, hidden layer dimensions, and learning rate, to obtain the optimal model configuration.

[0061] A hierarchical cross-validation strategy was adopted to ensure that the proportion of various lithofacies in each validation set remained consistent with the overall dataset. The model performance was comprehensively evaluated using multi-dimensional metrics such as accuracy, macro-average F1 score, and confusion matrix. To improve the model's generalization ability across wells, single-well Z-score normalization was employed, combined with cross-well domain adversarial data augmentation techniques to enhance the model's adaptability to geological differences in different wells. Finally, independent testing was conducted using blind well data that was not involved in training to generate continuous lithofacies prediction profiles, validating the model's practicality and reliability.

[0062] The trained and optimized model is applied to new logging data to generate lithofacies prediction results in batches. The model outputs complete lithology category probability distribution, prediction confidence score, and thin-layer boundary confidence interval for each depth point. It provides professional visualization output, including lithofacies profiles, probability distribution curves, and boundary marker maps, which facilitates comparison and verification with logging data and core description results. At the same time, combined with the well network distribution and geological background knowledge of the work area, it realizes accurate identification of single well lithofacies types, comparison of well profiles, and prediction of planar distribution patterns, providing reliable technical support for oilfield exploration and development.

[0063] To verify the effectiveness of the method of this invention, conventional logging curves from multiple production wells in the study area were collected, including key measurement curves such as natural gamma (GR), spontaneous potential (SP), wellbore diameter (CAL), density (DEN), sonic transit time (AC), formation true resistivity (RT), and flushed zone resistivity (RXO). To ensure the spatial scale consistency of the samples, the sampling interval for different wells was uniformly set to 0.125m, and depth alignment was achieved through linear interpolation or other appropriate interpolation methods. Subsequently, single-well Z-Score normalization was applied to each well. Normalization was performed to reduce the impact of inter-well instrument calibration differences on model training. After preprocessing, the total number of data samples was 11,944, containing 10 types of lithological labels. The labels were derived from core and logging interpretation (example label distribution: mudstone, argillaceous limestone, argillaceous siltstone, limestone, gypsum rock, siltstone, fine sandstone, breccia, calcareous argillaceous siltstone, calcareous sandstone).

[0064] To ensure data quality, an isolated forest (contamination=0.01, random_state=42) was used to detect and remove outliers; linear interpolation was used to repair missing segments. These processes minimize the impact of curve noise and outliers on the model, resulting in a multi-dimensional feature vector for each depth point as model input.

[0065] Training employs a well-sharing strategy to evaluate generalization ability: the training set and blind well set are divided into multiple wells (e.g., 10 training wells and 3 blind wells for independent validation). The dataset is divided into a training set (9555 samples), a validation set (1194 samples), and a test set (1195 samples) in an 8:1:1 ratio. Training uses the AdamW optimizer with a period of 50 epochs and an early stopping threshold of 0.0001. The lithofacies distribution statistics of the training dataset in this embodiment are referenced. Figure 2Among them, there were 2044 mudstone samples, 283 argillaceous limestone samples, 1452 argillaceous siltstone samples, 1559 limestone samples, 1497 gypsum rock samples, 1233 siltstone samples, 1401 fine sandstone samples, 234 breccia samples, 1100 calcareous argillaceous siltstone samples, and 1196 calcareous sandstone samples.

[0066] Reference diagram of the distribution characteristics of training set logging parameters among different lithofacies Figure 3 The distribution characteristics of well logging parameters (including natural gamma ray spectroscopy (GR), spontaneous potential (SP), borehole diameter (CAL), density (DEN), sonic transit time (AC), resistivity (RT), and compensated resistivity (RXO)) were visualized and analyzed. The variation curves of each parameter at different depths and between lithofacies showed significant differences. These distributions provided guidance for feature selection, indicating that parameters such as GR, DEN, and AC are highly sensitive to lithofacies differentiation.

[0067] Pearson correlation heatmap reference between training set logging parameters Figure 4 , Figure 4 Correlation analysis of well logging parameters was performed, and the associations between parameters were quantified using correlation coefficient thermograms. The results showed that the correlation coefficient between GR and SP was 26%, between DEN and AC was -66%, and between RT and RXO was 23%. Positive correlations (e.g., 8% between GR and CAL) and negative correlations (e.g., -41% between CAL and SP) revealed intrinsic relationships between parameters; for example, the strong negative correlation (-66%) between density and sonic transit time indicates that they can be used complementaryly in lithofacies identification. This analysis helps reduce redundant features and improve model efficiency.

[0068] Reference graphs showing the changes in loss and accuracy during model training and validation Figure 5 The training loss rapidly decreased from the initial 0.0053 to 0.0000, and the validation loss decreased from 0.0035 to 0.0000; the training accuracy improved from 0.0374 to 0.9684, and the validation accuracy improved from 0.0343 to 0.9581; the training F1 score increased from 0.0255 to 0.9704, and the validation F1 score increased from 0.0249 to 0.9615. The curves indicate that the model tends to converge after the 20th epoch, with no significant overfitting.

[0069] After training, the test set is evaluated. The lithofacies classification confusion matrix for the test set is shown in the reference diagram. Figure 6The test set confusion matrix shows that diagonal values ​​are relatively high, for example, mudstone (205 / 205 correct) and limestone (156 / 156 correct), but there is slight confusion between argillaceous siltstone and siltstone (e.g., 129 correct out of 146 samples). The overall accuracy is 0.9590, and the macro-average F1 score is 0.9617. The test set classification report is shown in Table 1.

[0070] Table 1

[0071] The output consists of the probability distribution of 10 lithofacies categories at each depth point and the corresponding confidence score (confidence can be obtained from the maximum probability or entropy measure). The model uses a combination of cross-entropy and Focal Loss to address class imbalance.

[0072]

[0073] Where L is the loss function, γ can be selected as needed (e.g., γ=2), and α is dynamically adjusted based on the category count to improve the ability to identify rare thin sandy types (accounting for about 10% of the total sample). This represents the model's predicted probability of the true class t for the i-th sample (i.e., the probability of the true label in the softmax output).

[0074] Hyperparameters are adaptively searched using the Optuna Bayesian optimization framework. Typical search spaces include: num_heads, num_layers, d_model, dim_feedforward, dropout_rate, learning_rate, batch_size, conv_channels, and kernel_size. Table 2 shows the optimal model hyperparameters adaptively searched using the Optuna Bayesian optimization framework.

[0075] Table 2

[0076] Refer to Table 3 for a comparison of the basic performance of the Transformer model and the model of this invention.

[0077] Table 3

[0078] This invention outperforms the Transformer model in all metrics, especially in the F1-Score, demonstrating its superior ability to identify rare samples such as thin sand layers.

[0079] The model trained using this method shows a significant improvement in the recognition performance of thin sand layers (especially siltstone). Experimental comparison results show that, compared with the baseline CNN-Transformer model, the complete model using the three improved modules improves the recall rate by about 8%-12% and the precision rate by about 6%-10% in the thin sandstone category; the overall lithofacies prediction accuracy in blind well testing reaches over 92%, and the vertical continuity of the prediction results is significantly improved, with a significant reduction in misclassification and layer skipping phenomena.

[0080] In one exemplary embodiment, a computer-readable storage medium is included, which stores a computer program that, when executed by a processor, implements the above-described machine learning-based method for predicting lithofacies in thin sandstone layers.

[0081] Please see Figure 7 In one exemplary embodiment, the device further includes an electronic device including at least one processor, at least one memory, and at least one communication bus.

[0082] The memory stores a computer program, which includes computer-readable instructions. The processor calls the computer-readable instructions stored in the memory through the communication bus to execute the aforementioned machine learning-based logging lithofacies prediction method for thin sandstone layers.

[0083] In one exemplary embodiment, a computer program product is proposed, including a computer program / instruction that, when executed by a processor, implements the steps of the above-described machine learning-based method for predicting lithofacies in thin sandstone layers.

[0084] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A machine learning-based well logging method for predicting lithofacies in thin sandstone layers, characterized in that, Includes the following steps: S1. Acquire multiple well logging data and preprocess them to construct a well logging feature dataset; S2. Identify well diameter anomaly data in the well logging feature dataset based on CAL logging curves, and correct the feature contribution of the well diameter anomaly data to obtain the corrected well logging feature dataset. S3. Adaptive high-frequency enhancement processing is performed on the high-frequency response features in the corrected well logging feature dataset to obtain the final well logging feature dataset. S4. Construct a machine learning-based well logging lithofacies prediction model for thin sandstone layers, including: a local feature encoder, a feature fusion module, a global feature encoder, and a classifier; Among them, a convolutional neural network is used as a local feature encoder, and a Transformer encoder is used as a global feature encoder; The final logging dataset is input into the local feature encoder to extract local features. The feature fusion module fuses the local features and inputs them into the global feature encoder to establish the global dependency relationship of the logging data. The classifier outputs the logging lithofacies prediction of the thin sandstone layer. S5. Predict the facies of the thin sandstone layer using the well logging facies prediction model.

2. The method for predicting lithofacies in thin sandstone layers based on machine learning according to claim 1, characterized in that, Well logging characteristics include: natural gamma, spontaneous potential, well diameter, density, sonic transit time, formation true resistivity, and resistivity of the flushed zone.

3. The method for predicting lithofacies in thin sandstone layers based on machine learning according to claim 1, characterized in that, Preprocessing includes: The well logging data was subjected to depth alignment and step size standardization interpolation, and lithofacies type labels were determined. Convert the lithofacies types described in the text into numerical labels and establish a mapping relationship between lithofacies categories and numerical labels; Statistical analysis of the sample distribution of various lithofacies was performed, and rare categories with fewer than a set threshold of samples were removed. The IQR method is used to identify and remove outlier data points; Each logging parameter was standardized to have a mean of 0 and a variance of 1.

4. The method for predicting lithofacies in thin sandstone layers based on machine learning according to claim 1, characterized in that, Based on the identification of well diameter anomalies in the well logging feature dataset using CAL logging curves, and employing a dynamic channel gating strategy to correct the feature contribution of these anomalies, the corrected well logging feature dataset is obtained, as follows: in, This represents the i-th corrected well logging feature data. This represents the i-th well logging feature data. This indicates depth-wise, channel-wise broadcast multiplication. Indicates the gating weight, Let W represent the sigmoid activation function, and W be a learnable vector. It is the well diameter threshold. This represents the well diameter displayed in the CAL logging curve for the i-th logging feature data.

5. The method for predicting lithofacies in thin sandstone layers based on machine learning according to claim 1, characterized in that, Adaptive high-frequency enhancement processing is applied to the high-frequency response features in the corrected well logging feature dataset to obtain the final well logging feature dataset, represented as: in, This represents the i-th final well logging feature data. This represents the i-th corrected well logging feature data. Represents the learnable coefficient. Indicates high-frequency residuals, This represents a low-frequency smoothing filter with a kernel size of k.

6. The method for predicting lithofacies in thin sandstone layers based on machine learning according to claim 1, characterized in that, In the attention calculation of the Transformer encoder, a geological prior of well logging depth is introduced. The attention based on this prior is as follows: Where τ is the attenuation factor, This represents the depth difference between adjacent depth points i and j in well logging.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

8. An electronic device, characterized in that, The device includes a processor and a memory, the processor being interconnected with the memory, wherein the memory is used to store a computer program, the computer program including computer-readable instructions, and the processor is configured to invoke the computer-readable instructions to perform the method as described in any one of claims 1-6.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-6.