Logging curve abnormal section reconstruction method based on attribute co-occurrence relation

By using a logging curve reconstruction method based on attribute co-occurrence relationships, using the Spearman correlation coefficient and Savitzky-Golay filter to extract trend features, combined with polynomial regression and the ACF-Informer model, the problem of accurately reconstructing abnormal sections of logging curves in complex geological environments is solved, achieving high-precision logging data recovery and formation evaluation.

CN120653901APending Publication Date: 2025-09-16CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510824292.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies have difficulty accurately identifying and correcting abnormal sections of logging curves in complex geological environments, resulting in inaccurate reservoir evaluation and production capacity prediction. Traditional methods are easily influenced by subjective experience, and simple interpolation methods are difficult to reflect the actual geological characteristics. Machine learning algorithms are not effective when processing high-dimensional and nonlinear data.

Method used

The Spearman correlation coefficient was used to screen co-occurring attributes, and the Savitzky-Golay filter was used to extract trend features. Polynomial regression modeling was adopted and the ACF-Informer model was constructed. The co-occurrence relationship features of the attributes were integrated, and the abnormal segments of the logging curves were reconstructed using the ensemble learning method.

Benefits of technology

The reconstruction accuracy of abnormal sections of logging curves is improved, the mean absolute error of the reconstruction results is reduced, the retention of geological characteristics and the robustness of the model are enhanced, and the spatial variability of complex geological environments is adapted. The error of stratigraphic interface identification is controlled within 0.3m.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653901A_ABST
    Figure CN120653901A_ABST
Patent Text Reader

Abstract

The invention discloses a logging curve abnormal section reconstruction method based on an attribute co-occurrence relation. The logging curve abnormal section reconstruction method comprises the steps of research area data preprocessing, attribute correlation analysis, curve trend characteristic extraction, modeling of a nonlinear correlation relation between curves, generation of attribute co-occurrence relation characteristics of a missing section and construction of an ACF-Informer curve reconstruction model. According to the method, co-occurrence attributes are screened by using Spearman correlation coefficients, nonlinear equations among the attributes in all periods are modeled through polynomial regression, and ACR features are fused to an Informer encoder-decoder architecture in stages as priori knowledge; the multi-strategy ACR features are fused by adopting integrated learning, so that the model can adapt to the spatial variability of the high-heterogeneity geological environment; a sub-model is constructed through a plurality of feature fusion strategies, and final output is optimized by adopting an integrated learning technology, so that the robustness and prediction precision of the model in a complex geological environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for reconstructing anomaly segments of a well logging curve, in particular to a method for reconstructing anomaly segments of a well logging curve based on attribute co-occurrence relationships, and belongs to the technical field of geophysical data processing. Background Art

[0002] During geological exploration and oil and gas field development, well logging technology is a crucial tool for assessing formation properties. With increasing drilling depths and increasingly complex geological structures, anomalous data segments often appear in logging curves due to instrument failure, wellbore damage, mud intrusion, and other factors. These anomalous segments can interfere with the accuracy of reservoir evaluation and production capacity prediction, rendering them ineffective for subsequent interpretation and analysis.

[0003] Traditional methods typically rely on manual or single-dimensional statistical discrimination to identify and correct abnormal sections, which are easily influenced by subjective experience and have low accuracy in formation environments with significant multi-attribute coupling. In addition, some abnormal data filling methods based on simple interpolation are difficult to reflect the true geological characteristics, often leading to confusion between high-frequency and low-frequency information, which in turn affects the results of lithologic discrimination and stratigraphic division. For example, prior art 1): the abnormal section identification method based on local statistical thresholds proposed by Wang et al. in "Petroleum Geophysical Exploration" Vol. 55, No. 2, 2020) has obvious limitations in complex geological environments with multi-attribute coupling: it is highly subjective and difficult to capture nonlinear correlations. Simple interpolation filling (such as linear interpolation or moving average) easily confuses high-frequency characteristics of the formation with low-frequency trends. For example, prior art 2): the multi-scale decomposition and reconstruction method described in "Well Logging Technology" Vol. 42, No. 4, 2018 by Li et al. leads to a decrease in the accuracy of lithologic discrimination and stratigraphic division.

[0004] Machine learning algorithms (including deep learning algorithms) have gained widespread recognition for their powerful ability to fit complex nonlinear mapping relationships. While existing machine learning and deep learning methods can learn complex relationships between attributes from known well logs by building data-driven models and adapting to complex geological conditions by leveraging the nonlinear nature of data and the synergistic effects of multiple variables (e.g., prior art 3): Smith et al. "Log Data Reconstruction Using LSTM Networks," SPE Annual Technical Conference, 2021, these methods often fall short in real-world applications. This is due to the complexity of the geological environment and the limitations of the models themselves. Geological environments are characterized by high nonlinearity, heterogeneity, and spatial variability. In such complex geological environments, data are often high-dimensional, and complex nonlinear interactions exist between geological parameters, posing significant challenges to model fitting. Even the most advanced machine learning algorithms often struggle to fully capture all key factors when dealing with high-dimensional, nonlinear, and dynamically changing data. Existing ML / DL models (such as CNNs and RNNs) struggle to fully learn the complex synergistic mechanisms between multiple attributes. Especially when abnormal segments cause the co-occurrence relationship of key attributes to be broken, pure data-driven models often ignore the inherent correlation laws between geological attributes. The reconstruction results under complex geological conditions have oscillation distortion or geological logic contradictions, which restricts its reliability in industrial scenarios. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for reconstructing abnormal segments of well logging curves based on attribute co-occurrence relationships in order to solve at least one of the above technical problems.

[0006] The present invention achieves the above-mentioned purpose through the following technical solutions: a method for reconstructing abnormal segments of well logging curves based on attribute co-occurrence relationships, the method for reconstructing abnormal segments of well logging curves comprising the following steps:

[0007] S1. Data preprocessing in the study area: preprocessing the well logging data through data cleaning, data stratification and data normalization to eliminate shallow abnormal data, achieve single point outlier replacement, equal length processing, and stratigraphic division and normalization;

[0008] S2, attribute correlation analysis, using the Spearman correlation coefficient to analyze the correlation between logging attributes and determine the co-occurring attribute pairs;

[0009] S3, extracting the trend characteristics of the curve, using Savitzky-Golay filter to smooth the logging curve and extract the trend characteristics;

[0010] S4. Model the nonlinear correlation between curves, segment the curves based on extreme points, and use polynomial regression method to establish a nonlinear regression model between co-occurring attribute pairs;

[0011] S5. Generate attribute co-occurrence relationship features of the missing segment, substitute the co-occurrence attribute values ​​of the curve segment to be restored into the regression model, and generate attribute co-occurrence relationship (ACR) features of the target attribute;

[0012] S6. Construct an ACF-Informer curve reconstruction model, fuse the ACR features on the basis of the Informer model, and construct an ACF-Informer curve reconstruction model based on different feature fusion strategies. Use an ensemble learning method to combine the outputs of each strategy to achieve high-precision reconstruction of the abnormal segment of the logging curve.

[0013] As a further solution of the present invention: data cleaning includes: removing shallow logging data, using the shortest curve as the standard, truncating other curves accordingly to achieve alignment, and selecting the mean of the previous and next data for replacement to correct these erroneous values ​​and ensure the continuity and accuracy of the data;

[0014] Data stratification involves constructing appropriate data sets to address the impact of differences in data characteristics across different layers on model predictions. This involves initially dividing the raw data by geological age before modeling, helping to separate the attributes of different strata, allowing the model to focus more on modeling relationships within a single layer, thereby improving prediction accuracy and reliability.

[0015] Data normalization processing includes: unifying data dimensions, eliminating scale differences between different logging parameters, accelerating model convergence, improving numerical stability, and reducing the risk of numerical overflow. Min-max normalization is used to scale the data to the range of [0, 1]. The calculation method is as follows:

[0016]

[0017] Among them, x′ is the normalized value, x is the original data value, and x max is the maximum value in the data set, x min is the minimum value in the data set.

[0018] As a further solution of the present invention, the Spearman correlation coefficient analysis of the correlation between the logging attributes specifically includes:

[0019] The Spearman correlation coefficient is a non-parametric statistical indicator used to measure the strength of the monotonic correlation between two sets of data. It is calculated by calculating the correlation between the rank values ​​of the two sets of data after sorting. It is applicable to linear and nonlinear but monotonic relationships. The calculation method is as follows:

[0020]

[0021] Among them, ρ is the Spearman correlation coefficient, rank(X) and rank(Y) are the ranks of the two sets of data, σ rank(X) and σ rank(Y) is the standard deviation of the ranks.

[0022] As a further solution of the present invention, the smoothing process of the well logging curve by the Savitzky-Golay filter specifically includes:

[0023] Savitzky-Golay filter is used for data smoothing and noise removal. It smoothes the data while retaining the high-frequency characteristics of the signal. It smoothes the data based on local polynomial regression. Suppose the data sequence is (x i ,y i ), where i∈[1,N], the goal is to fit the local data with a p-order polynomial, and the formula used is:

[0024] P(x)=c0+c1x+c2x 2 +…+c p x p

[0025] Where n is the order of the polynomial, c i are the polynomial coefficients;

[0026] For the center at x k The coefficient vector [a0,a1,...,a p ], so that the polynomial can best fit the data points in the window, the formula used is:

[0027]

[0028] Among them, 2m+1 represents the window size, and the window center is point x k ; It can effectively reduce random noise in data while maintaining the shape and characteristics of the data;

[0029] When applying Savitzky-Golay filtering, set the order of the polynomial fitted within the sliding window to which the filter is applied. This polynomial is used to describe the relationship between data points and to smooth the data. First-order polynomials are suitable for situations where data changes are relatively gentle and approximately linear. Second-order polynomials can better capture curve changes in the data. Third-order or even higher-order polynomials can capture more complex changes. Based on the characteristics of the logging curve morphology and the focus on the logging curve trend due to attribute correlation, it is necessary to ignore details and complex data fluctuations, so a second-order polynomial is selected to fit the logging data.

[0030] As a further solution of the present invention: when a nonlinear regression model between co-occurrence attribute pairs is established using a polynomial regression method, polynomial regression is used to extract ACR features, specifically including:

[0031] Polynomial regression is used to analyze the linear relationship between the target variable and multiple independent variables. The relationship between the modeled predictor variable x and the response variable y is an n-order polynomial. By introducing higher-order terms of x, it can fit nonlinear patterns in the data and provide a more complex and flexible relationship description than the linear model. The formula used is:

[0032] y=β0+β1x+β2x 2 +…+β n x n +∈

[0033] Among them, y is the dependent variable, x is the independent variable, β0,β1,…,β n are model parameters, n is the order of the polynomial, and ∈ is the error term, which is usually assumed to be normally distributed;

[0034] In the logging curve trend fitting, it is possible to integrate the relationship between logging parameters, comprehensively analyze the impact of characteristic parameters on the target curve, extract the main trend and remove noise;

[0035] In the process of polynomial regression analysis of the well logging curve, the curve is further subdivided into several easy-to-handle segments. In the previous section, Savitzky-Golay filtering was used to fit the original curve using a quadratic polynomial. On this basis, a curve segmentation strategy based on extreme points was proposed to adapt to the non-periodic characteristics and morphological differences of the curve.

[0036] The "restored feature polynomial" of the attribute is obtained through polynomial regression. The co-occurrence attribute curve of the borehole to be reconstructed for abnormal attributes is cut according to the extreme value points, and the curve data of each cut segment is input into the corresponding restored feature polynomial to extract the ACR features of each segment. These features are then used as conditional constraints to enhance the training effect of the Informer model.

[0037] As a further solution of the present invention: The first strategy for constructing the ACF-Informer curve reconstruction model is to add the ACR feature as a column of "input" data to the original input data before the Encoder. The specific method is to add the original data and ACR features t (|1≤t≤L) is concatenated by column, where t is the tth time step, N is the number of features for each time step, and L is the sequence length), and the enhanced input data matrix is ​​obtained. The resulting model is called "Encoder Pre-Informer," or "EP-Informer" for short. The advantage of this strategy is that new features begin to influence the model from the data input stage, allowing the model to incorporate information from new features early on, enhancing the expressiveness of the input data and thus helping to improve the overall feature extraction quality.

[0038] As a further solution of the present invention: The second strategy for constructing the ACF-Informer curve reconstruction model is to fuse the new features with the output of the encoder or the initial input of the decoder. After the encoder processes the original data, the output data dimension is (L / 2, 512), where L is the length of the input sequence and 512 is the feature dimension of each time step; the dimension of the ACR feature is (L, 1), and each corresponding time step contains only one attribute relationship value. Since the two feature sets do not match in dimension, dimensional conversion is required. Specifically, the following steps are performed:

[0039] First, the ACR feature is subjected to the maximum pooling operation, which is a downsampling operation used to reduce the feature dimension and retain the most important information; let X∈R L×1 Represents the ACR feature sequence; where L is the time step, 1 represents single-channel data, and the maximum pooling operation uses a window size k = 2 and a step size s = 2 for downsampling:

[0040]

[0041] The pooling result X pool With shape X pool ∈R (L / 2)×1 ;

[0042] It is then mapped to the same dimension as the encoder output through linear embedding; this is done through a fully connected layer that linearly transforms the pooled features using a weight matrix:

[0043] X embed =X pool W+b

[0044] Where W∈R 1×512 is the weight matrix, b∈R 512 is the bias term. The final transformed ACR feature has shape X embed ∈R (L / 2)×512 ;

[0045] After completing the dimension adjustment, the encoder output features are fused with the ACR features through the cross-attention mechanism. The cross-attention mechanism can automatically learn the interdependence and importance between the two features. The specific operation is as follows:

[0046] The encoder output is (B, T, D), and the ACR feature is (B, T′, F). The ACR feature (B, T′, F) is used as the key and value, and is projected to the appropriate dimension (B, T′, D) through the learned linear mapping. The encoder output (B, T, D) is used as the query, and the linear mapping is used to obtain (B, T, D). The weight is obtained by calculating the cross attention:

[0047]

[0048] Among them, Q comes from the Encoder output, K and V come from ACR features, d k is the dimension of the Key;

[0049] The calculated attention weights are used to weight the Value, thereby obtaining a weighted feature representation for each time step. In this way, the model can adjust the representation of the encoder output features according to the importance of the ACR features, so that the model can simultaneously focus on the temporal information of the time series and the dependencies between different attributes.

[0050] The resulting model is called "EncoderDecodeFusion-Informer". This strategy ensures that new features can take effect in the middle layer of the model, helping the model to better understand and adjust the representation of the original features. It is particularly suitable for models that need to dynamically adjust features.

[0051] As a further solution of the present invention: The third strategy for constructing the ACF-Informer curve reconstruction model is to add new features to the features of each time step output by the Decoder and then input them together into the fully connected layer. The data dimension of the Decoder output is (L / 2,512), and the features output by the Decoder are fused with the ACR features in the same way as the second strategy; the resulting model is called "DecoderPost-Informer". This strategy adds new features before the final output layer so that the new features have a direct impact on the final prediction results. It is suitable for scenarios where the final output needs to be fine-tuned or specific output dimensions need to be optimized. It can further improve the accuracy of the final prediction while ensuring feature expression.

[0052] The beneficial effects of the present invention are:

[0053] 1) The present invention uses the Spearman correlation coefficient to screen co-occurrence attributes and combines Savitzky-Golay filtering with periodic decomposition strategy to extract nonlinear trend features, so that the ACR feature can accurately characterize the association law of geological attributes missing in the abnormal segment, and the mean absolute error (MAE) between the reconstructed segment and the real data is reduced;

[0054] 2) This paper uses polynomial regression to model the nonlinear equations between attributes within each period and integrates the ACR features as prior knowledge into the informer encoder-decoder architecture in stages, effectively avoiding the high / low frequency information confusion caused by traditional interpolation methods and enhancing the preservation of geological features;

[0055] 3) This paper uses ensemble learning to integrate multi-strategy ACR features, enabling the model to adapt to the spatial variability of highly heterogeneous geological environments. In the reconstruction results of fault-developed areas, the error in identifying stratum interfaces is controlled within 0.3 m, achieving adaptive optimization for complex strata.

[0056] 4) EP-Informer, EDF-Informer, and DP-Informer sub-models are constructed through multiple feature fusion strategies, and ensemble learning technology is used to optimize the final output, thereby improving the robustness and prediction accuracy of the model in complex geological environments. The present invention can be widely used in logging data recovery and quality improvement in oil and gas exploration and formation evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0058] Figure 2 This is a Spearman correlation coefficient analysis diagram between attributes of the present invention;

[0059] Figure 3 This is a visualization diagram of the SG filtering results of DZL performed by the present invention;

[0060] Figure 4 This is a visualization diagram of the SG filtering results of ZRDW in the present invention;

[0061] Figure 5 This is a visualization diagram of the results of the present invention dividing DZL into "extreme points";

[0062] Figure 6 This is a visualization diagram of the results of dividing ZRDW by "extreme points" in the present invention;

[0063] Figure 7 This is a polynomial equation diagram of the DZL and ZRDW fitting of the part after segmentation by extreme points in the present invention;

[0064] Figure 8 This is the structural diagram of the EP-Informer model of the present invention;

[0065] Figure 9 It is the ACR feature dimension transformation diagram of the present invention;

[0066] Figure 10 This is the structural diagram of the EDF-Informer model of the present invention;

[0067] Figure 11 This is the structural diagram of the DP-Informer model of the present invention;

[0068] Figure 12 This is the structural diagram of the ACR-Informer model of the present invention. DETAILED DESCRIPTION

[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0070] Example 1, as Figures 1 to 12 As shown, a method for reconstructing anomaly segments of well logging curves based on attribute co-occurrence relationships includes the following steps:

[0071] S1. Data preprocessing in the study area: preprocessing the well logging data through data cleaning, data stratification and data normalization to eliminate shallow abnormal data, achieve single point outlier replacement, equal length processing, and stratigraphic division and normalization;

[0072] S2, attribute correlation analysis, using the Spearman correlation coefficient to analyze the correlation between logging attributes and determine the co-occurring attribute pairs;

[0073] S3, extracting the trend characteristics of the curve, using Savitzky-Golay filter to smooth the logging curve and extract the trend characteristics;

[0074] S4. Model the nonlinear correlation between curves, segment the curves based on extreme points, and use polynomial regression method to establish a nonlinear regression model between co-occurring attribute pairs;

[0075] S5. Generate attribute co-occurrence relationship features of the missing segment, substitute the co-occurrence attribute values ​​of the curve segment to be restored into the regression model, and generate attribute co-occurrence relationship (ACR) features of the target attribute;

[0076] S6. Construct an ACF-Informer curve reconstruction model, fuse ACR features on the basis of the Informer model, and construct an ACF-Informer curve reconstruction model based on different feature fusion strategies. Use the ensemble learning method to combine the outputs of each strategy to achieve high-precision reconstruction of the abnormal section of the logging curve.

[0077] In addition to all the technical features of the first embodiment, the second embodiment also includes:

[0078] Data cleaning includes: First, due to the combined effects of ground environmental interference, insufficient instrument stability, wellbore fluid conditions and formation characteristics, shallow logging data may be distorted. Therefore, this study, as recommended by geophysical exploration experts, excludes the first 20 meters of shallow logging data to avoid these influences; Second, deep logging data may be misaligned in the last few dozen meters due to the combined effects of depth accumulation error, signal delay, formation condition complexity, wellbore environment and instrument performance. Therefore, the shortest curve is used as the standard, and the other curves are truncated accordingly to achieve alignment; Finally, due to instrument problems or improper data processing, erroneous values ​​may appear in the middle of the curve (such as single-point mutations or negative density values, etc.). To this end, the mean of the previous and next data is used for replacement to correct these erroneous values ​​and ensure the continuity and accuracy of the data.

[0079] Data stratification involves the fact that strata from different geological ages exhibit distinct characteristics in well logging curves due to differences in their depositional environments and physical properties. By constructing appropriate datasets, we can address the impact of these differences in data characteristics across different layers on model predictions. Preliminary segmentation of raw data by geological age before modeling helps isolate the attributes of different strata, allowing the model to focus more on modeling relationships within a single layer, thereby improving prediction accuracy and reliability.

[0080] Data normalization processing includes: Well logging data usually contains multiple different types of curves, and the dimensions and value ranges of these curves may vary greatly. If this data is directly used for modeling, the impact of certain features on the model may be amplified or ignored, thereby affecting the performance of the model. Normalization processing can unify the data dimensions, eliminate the scale differences between different well logging parameters, accelerate model convergence, improve numerical stability, and reduce the risk of numerical overflow. Use minimum-maximum normalization to scale the data to the range of [0,1]. The calculation method is as follows:

[0081]

[0082] Among them, x′ is the normalized value, x is the original data value, and x max is the maximum value in the data set, x min is the minimum value in the data set.

[0083] The Spearman correlation coefficient analysis of the correlation between logging attributes specifically includes:

[0084] The Spearman correlation coefficient is a nonparametric statistical indicator used to measure the strength of the monotonic correlation between two sets of data. It is calculated by calculating the correlation between the ranked values ​​of the two sets of data. It is applicable to linear and nonlinear but monotonic relationships. The calculation method is as follows:

[0085]

[0086] Among them, ρ is the Spearman correlation coefficient, rank(X) and rank(Y) are the ranks of the two sets of data, σ rank(X) and σ rank(Y) is the standard deviation of the ranks.

[0087] The relationship between each attribute and the corresponding serial number is detailed in Table 1.

[0088] Table 1 shows the serial numbers and corresponding attributes

[0089]

[0090]

[0091] The Spearman correlation coefficient analysis was performed on the eight selected logging attributes. The strategy fully utilized the correlation between the attributes and provided a scientific basis for the screening of homogeneous section data.

[0092] The Savitzky-Golay filter smoothing process for logging curves specifically includes:

[0093] Savitzky-Golay filter is used for data smoothing and noise removal. It can smooth the data while retaining the high-frequency features of the signal, such as peaks and other sharp change points. It can smooth the data based on local polynomial regression (usually second-order or third-order polynomial). Suppose the data sequence is (x i ,y i ), where i∈[1,N], the goal is to fit the local data with a p-order polynomial, and the formula used is:

[0094] P(x)=c0+c1x+c2x 2 +…+c p x p

[0095] Where n is the order of the polynomial, c i are the polynomial coefficients;

[0096] For the center at x k The coefficient vector [a0,a1,...,a p ], so that the polynomial can best fit the data points in the window, the formula used is:

[0097]

[0098] Among them, 2m+1 represents the window size, and the window center is point x k ; It can effectively reduce random noise in the data while maintaining the shape and characteristics of the data, such as peaks and inflection points.

[0099] When applying Savitzky-Golay filtering, set the order of the polynomial fitted within the sliding window to which the filter is applied. This polynomial is used to describe the relationship between data points and to smooth the data. First-order polynomials (i.e., linear regression) are suitable for situations where data changes are relatively gentle and approximately linear. Second-order polynomials can better capture curve changes in the data, such as parabolic fluctuations. Third-order or even higher-order polynomials can capture more complex changes. Based on the characteristics of the logging curve morphology and the focus on the logging curve trend due to attribute correlation, it is necessary to ignore details and complex data fluctuations, so a second-order polynomial is selected to fit the logging data.

[0100] Polynomial regression is used to extract ACR features when establishing a nonlinear regression model between co-occurrence attribute pairs, specifically including:

[0101] Polynomial regression is used to analyze the linear relationship between the target variable and multiple independent variables. The relationship between the modeled predictor variable x and the response variable y is an n-order polynomial. By introducing higher-order terms of x, it can fit nonlinear patterns in the data and provide a more complex and flexible relationship description than the linear model. The formula used is:

[0102] y=β0+β1x+β2x 2 +…+β n x n +∈

[0103] Among them, y is the dependent variable (response variable), x is the independent variable (predictor variable), β0, β1,…, β s are model parameters, n is the order of the polynomial, and ∈ is the error term, which is usually assumed to be normally distributed.

[0104] In the trend fitting of logging curves, the relationship between logging parameters (such as density and natural gamma, resistivity and natural potential, etc.) can be integrated, the impact of characteristic parameters on the target curve can be comprehensively analyzed, the main trends can be extracted and noise can be removed.

[0105] A major challenge in performing polynomial regression analysis on well log curves is that the curves often do not exhibit strict periodicity and their morphology can vary significantly from period to period. This characteristic makes directly fitting the entire dataset not always the most effective strategy. Therefore, to improve the manageability and accuracy of the analysis, the curves are further subdivided into manageable segments. In the previous section, we applied Savitzky-Golay filtering and fitted the original curves with quadratic polynomials. Building on this, we propose a curve segmentation strategy based on extreme points to accommodate the non-periodic characteristics and morphological differences of the curves.

[0106] Although this method may destroy the integrity of the parabola, it can lead to a simpler model when fitting the relationship between attributes. The advantage of this strategy is that it simplifies the complexity of the model, making the fitting process and its results easier to understand and apply.

[0107] The "restored feature polynomial" of the attribute is obtained through polynomial regression. The co-occurrence attribute curve of the borehole to be reconstructed for abnormal attributes is cut according to the extreme value points, and the curve data of each cut segment is input into the corresponding restored feature polynomial to extract the ACR features of each segment. These features are then used as conditional constraints to enhance the training effect of the Informer model.

[0108] In addition to all the technical features of Example 1, this embodiment also includes: The first strategy for constructing the ACF-Informer curve reconstruction model is to add the ACR feature as a column of "input" data to the original input data before the encoder. The specific approach is:

[0109] The original data and ACR features t (|1≤t≤L) is concatenated by column, where t is the tth time step, N is the number of features for each time step, and L is the sequence length), and the enhanced input data matrix is ​​obtained. The resulting model is called "Encoder Pre-Informer," or "EP-Informer" for short. The advantage of this strategy is that new features begin to influence the model from the data input stage, allowing the model to incorporate information from new features early on, enhancing the expressiveness of the input data and thus helping to improve the overall feature extraction quality.

[0110] The second strategy for constructing the ACF-Informer curve reconstruction model is to fuse the new features with the output of the encoder or the initial input of the decoder. After the encoder processes the original data, the output data dimension is (L / 2,512), where L is the length of the input sequence and 512 is the feature dimension of each time step; the dimension of the ACR feature is (L,1), and each corresponding time step contains only one attribute relationship value.

[0111] Since the two feature sets do not match in dimension, dimension conversion is required, including:

[0112] First, the ACR feature is subjected to the maximum pooling operation, which is a downsampling operation used to reduce the feature dimension and retain the most important information; let X∈R L×1 Represents the ACR feature sequence; where L is the time step, 1 represents single-channel data, and the maximum pooling operation uses a window size k = 2 and a step size s = 2 for downsampling:

[0113]

[0114] The pooling result X pool With shape X pool ∈R (L / 2)×1 ;

[0115] It is then mapped to the same dimension as the encoder output through linear embedding, which is done through a fully connected layer that linearly transforms the pooled features using a weight matrix:

[0116] X embed =X pool W+b

[0117] Where W∈R 1×512 is the weight matrix, b∈R 512 is the bias term. The final transformed ACR feature has shape X embed ∈R (L / 2)×512 ;

[0118] After completing the dimension adjustment, the encoder output features are fused with the ACR features through the cross-attention mechanism. The cross-attention mechanism can automatically learn the interdependence and importance between the two features. The specific operation is as follows:

[0119] The encoder output is (B, T, D), and the ACR feature is (B, T′, F). The ACR feature (B, T′, F) is used as the key and value, and is projected to the appropriate dimension (B, T′, D) through the learned linear mapping. The encoder output (B, T, D) is used as the query, and the linear mapping is used to obtain (B, T, D). The weight is obtained by calculating the cross attention:

[0120]

[0121] Among them, Q comes from the Encoder output, K and V come from ACR features, d k is the dimension of Key (usually F); Value is weighted by the calculated attention weight to obtain a weighted feature representation for each time step; in this way, the model can adjust the representation of the encoder output features according to the importance of the ACR features, so that the model can simultaneously focus on the temporal information of the time series and the dependencies between different attributes; the resulting model is called "EncoderDecodeFusion-Informer". This strategy ensures that the new features can play a role in the middle layer of the model, helping the model to better understand and adjust the representation of the original features, and is particularly suitable for models that require dynamic feature adjustment.

[0122] The third strategy for constructing the ACF-Informer curve reconstruction model is to add new features to the features of each time step output by the decoder and then input them together into the fully connected layer. The data dimension of the decoder output is (L / 2,512). The features output by the decoder are fused with the ACR features in the same way as the second strategy. The resulting model is called "DecoderPost-Informer". This strategy adds new features before the final output layer, so that the new features have a direct impact on the final prediction results. It is suitable for scenarios that require fine-tuning the final output or optimizing specific output dimensions. It can further improve the accuracy of the final prediction while ensuring feature expression.

[0123] Stacking is an advanced ensemble learning technique designed to improve overall prediction accuracy by combining the predictions of multiple different "base models" (also called first-level models). The key to the stacking method is the use of a "meta-model" (also called a second-level model or meta-learner), which learns how to optimally integrate the outputs of the underlying models to form the final prediction results. This approach often significantly improves model performance. Because the three feature fusion strategies mentioned above have complementary properties, they are used as base models. A fully connected neural network is used to fuse the predictions of the three and train a meta-model. The resulting model is called "ACR-Informer."

[0124] Combining domain knowledge with deep learning models, this method analyzes the correlations between well logging attributes, extracts attribute co-occurrence relationship (ACR) features, and integrates them with the informer model to achieve high-precision reconstruction of anomalous segments in well logging curves. EP-Informer, EDF-Informer, and DP-Informer sub-models are constructed through multiple feature fusion strategies, and ensemble learning techniques are used to optimize the final output, improving the model's robustness and prediction accuracy in complex geological environments. This method can be widely used for well logging data recovery and quality improvement in oil and gas exploration and formation evaluation.

[0125] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

[0126] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A method for reconstructing abnormal segments of well logging curves based on attribute co-occurrence relationships, characterized in that: The method for reconstructing abnormal segments of well logging curves comprises the following steps: S1. Data preprocessing in the study area: preprocessing the well logging data through data cleaning, data stratification and data normalization to eliminate shallow abnormal data, achieve single point outlier replacement, equal length processing, and stratigraphic division and normalization; S2, attribute correlation analysis, using the Spearman correlation coefficient to analyze the correlation between logging attributes and determine the co-occurring attribute pairs; S3, extracting the trend characteristics of the curve, using Savitzky-Golay filter to smooth the logging curve and extract the trend characteristics; S4. Model the nonlinear correlation between curves, segment the curves based on extreme points, and use polynomial regression method to establish a nonlinear regression model between co-occurring attribute pairs; S5. Generate attribute co-occurrence relationship features of the missing segment, substitute the co-occurrence attribute values ​​of the curve segment to be restored into the regression model, and generate attribute co-occurrence relationship features of the target attribute; S6. Construct an ACF-Informer curve reconstruction model, integrate the attribute co-occurrence relationship features on the basis of the Informer model, and construct an ACF-Informer curve reconstruction model based on different feature fusion strategies. Use the integrated learning method to combine the outputs of each strategy to achieve high-precision reconstruction of the abnormal segment of the logging curve.

2. The method for reconstructing abnormal segments of well logging curves according to claim 1, characterized in that: In S1, data cleaning includes: removing shallow logging data, using the shortest curve as the standard, truncating other curves accordingly to achieve alignment, and selecting the mean of the previous and next data for replacement to correct data errors and ensure data continuity and accuracy; Data stratification involves constructing data sets to address the impact of differences in data characteristics across different layers on model predictions. This involves initially dividing the raw data by geological age before modeling, helping to separate the attributes of different strata, allowing the model to focus more on modeling relationships within a single layer, thereby improving prediction accuracy and reliability. Data normalization processing includes: unifying data dimensions, eliminating scale differences between different logging parameters, accelerating model convergence, improving numerical stability, and reducing the risk of numerical overflow. Min-max normalization is used to scale the data to the range of [0, 1]. The calculation method is as follows: Among them, x′ is the normalized value, x is the original data value, and x max is the maximum value in the data set, x min is the minimum value in the data set.

3. The method for reconstructing abnormal segments of well logging curves according to claim 1, characterized in that: In S2, the Spearman correlation coefficient analysis of the correlation between logging attributes specifically includes: The Spearman correlation coefficient is a non-parametric statistical indicator used to measure the strength of the monotonic correlation between two sets of data. It is calculated by calculating the correlation between the rank values ​​of the two sets of data after sorting. The calculation method is as follows: Among them, ρ is the Spearman correlation coefficient, rank(X) and rank(Y) are the ranks of the two sets of data, σ rank(X) and σ rank(Y) is the standard deviation of the ranks.

4. The method for reconstructing abnormal segments of well logging curves according to claim 1, characterized in that: In S3, the smoothing process of the logging curve by the Savitzky-Golay filter specifically includes: Smooth the data based on local polynomial regression, assuming the data sequence is (x i ,y i ), where i∈[1,N], the goal is to fit the local data with a p-order polynomial, and the formula used is: P(x)=c0+c1x+c2x 2 +…+c p x p Where n is the order of the polynomial, c i are the polynomial coefficients; For the center at x k The coefficient vector [a0,a1,...,a p ], so that the polynomial can best fit the data points in the window, the formula used is: Among them, 2m+1 represents the window size, and the window center is point x k .

5. The method for reconstructing abnormal segments of well logging curves according to claim 1, characterized in that: In S4, when a polynomial regression method is used to establish a nonlinear regression model between co-occurring attribute pairs, polynomial regression is used to extract attribute co-occurrence relationship features, specifically including: The relationship between the modeled predictor variable x and the response variable y is an n-order polynomial. By introducing higher-order terms of x, it is possible to fit the nonlinear pattern in the data. The formula used is: y=β0+β1x+β2x 2 +…+b n x n +∈ Among them, y is the dependent variable, x is the independent variable, β0,β1,…,β n are model parameters, n is the order of the polynomial, and ∈ is the error term, which is usually assumed to be normally distributed.

6. The method for reconstructing abnormal segments of well logging curves according to claim 1, characterized in that: In S6, the first strategy for constructing the ACF-Informer curve reconstruction model is to add the attribute co-occurrence relationship feature as a column of "input" data to the original input data before the encoder. The specific approach is: The original data and ACR features t (|1≤t≤CL) is concatenated by column, where t is the tth time step, N is the number of features for each time step, and L is the sequence length), and the enhanced input data matrix is ​​obtained. The resulting model is called "EncoderPre-Informer".

7. The method for reconstructing abnormal segments of well logging curves according to claim 6, characterized in that: In S6, the second strategy for constructing the ACF-Informer curve reconstruction model is to fuse the new features with the output of the encoder or the initial input of the decoder. After the encoder processes the original data, the output data dimension is (L / 2, 512), where L is the length of the input sequence and 512 is the feature dimension of each time step. The dimension of the ACR feature is (L, 1), and each time step contains only one attribute relationship value; Since the two feature sets do not match in dimension, dimension conversion is required. Specifically, the following steps are needed: First, perform the maximum pooling operation on the attribute co-occurrence relationship features, and set X∈R L×1 Represents the ACR feature sequence; where L is the time step, 1 represents single-channel data, and the maximum pooling operation uses a window size k = 2 and a step size s = 2 for downsampling: The pooling result X pool With shape X pool ∈R (L / 2)×1 ; It is then mapped to the same dimension as the encoder output via linear embedding and completed by a fully connected layer, where the pooled features are linearly transformed using a weight matrix: X embed =X pool ·W+b Where W∈R 1×512 is the weight matrix, b∈R 512 is the bias term. The final transformed ACR feature has shape X embed ∈R (L / 2)×512 ; After completing the dimension adjustment, the output features of the encoder are fused with the attribute co-occurrence relationship features through the cross attention mechanism. The specific operations are as follows: The encoder output is (B, T, D), and the ACR feature is (B, T′, F). The ACR feature (B, T′, F) is used as the key and value, and is projected to the appropriate dimension (B, T′, D) through the learned linear mapping. The encoder output (B, T, D) is used as the query, and the linear mapping is used to obtain (B, T, D). The weight is obtained by calculating the cross attention: Among them, Q comes from the Encoder output, K and V come from ACR features, d k is the dimension of the Key; The calculated attention weights are used to weight the Values, resulting in a weighted feature representation for each time step. The model can adjust the representation of the encoder output features based on the importance of the attribute co-occurrence relationship features, allowing the model to simultaneously focus on the temporal information of the time series and the dependencies between different attributes. The resulting model is called "EncoderDecodeFusion-Informer".

8. The method for reconstructing abnormal segments of well logging curves according to claim 7, characterized in that: In S6, the third strategy for constructing the ACF-Informer curve reconstruction model is to add new features to the features of each time step output by the decoder and then input them together into the fully connected layer. The data dimension of the decoder output is (L / 2,512). The features output by the decoder are fused with the attribute co-occurrence relationship features in the same way as the second strategy. The resulting model is called "DecoderPost-Informer".

Citation Information

Cited By

  • Intelligent algorithm-based inter-salt and under-salt stratum distortion curve reconstruction method and system

    CN122239188A