Diagenetic facies identification method based on multi-modal fusion features and extreme random tree model
By processing logging curves using multimodal fusion features and extreme random tree models, the low efficiency and low accuracy of traditional methods for identifying diagenetic facies in tight sandstone reservoirs are solved, achieving faster and more accurate diagenetic facies identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWEST PETROLEUM UNIV
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies struggle to efficiently and accurately identify the diagenetic facies of tight sandstone reservoirs using well logging data. Traditional methods are also inadequate for capturing the nonlinear relationship between well logging data and lithology when processing complex geological data, and are costly and time-consuming.
By employing multimodal fusion features and an extreme random tree model, multi-scale features are extracted from well logging curves through preprocessing and multi-view feature transformation, and an extreme random tree model is constructed for diagenetic facies identification.
It enables faster and more accurate prediction of diagenetic facies types, improves identification accuracy, reduces costs, and allows for continuous identification of diagenetic facies at vertical depths in a single well.
Smart Images

Figure CN122413166A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tight sandstone oil and gas exploration and development technology, and in particular to a diagenetic facies identification method based on multimodal fusion features and an extreme random tree model. Background Technology
[0002] Compared to conventional sandstone reservoirs, tight sandstone reservoirs are characterized by complex pore structures, extremely low porosity and permeability, and prominent heterogeneity. Their formation is primarily attributed to intense and complex diagenetic alteration. Given a clear understanding of hydrocarbon sources, caprock conditions, and tectonic setting, diagenetic facies is one of the key factors controlling the effectiveness of tight sandstone reservoirs. Therefore, accurately identifying diagenetic facies is of great significance for guiding exploration site selection and resource assessment of tight oil and gas.
[0003] Early diagenetic facies identification mainly relied on experimental analysis methods such as core casting thin sections, cathodoluminescence, and scanning electron microscopy. However, these methods are costly, time-consuming, and limited by the acquisition of core data, making it difficult to achieve continuous identification of diagenetic facies at the vertical depth of a single well, which has a certain constraint on the evaluation of tight sandstone reservoirs.
[0004] With advancements in well logging technology, conventional well logging curves have become a standard method for lithology identification due to their ease of acquisition, low cost, and coverage of rich subsurface reservoir lithology information. For example, cross-plot methods distinguish rock types by combining different well logging curves. Subsequently, many scholars have attempted to use well logging data for diagenetic facies identification, achieving considerable progress. However, due to the complexity of the rock physical response in well logging, traditional cross-plot methods often exhibit low resolution and identification accuracy in lithology differentiation.
[0005] In recent years, with the rapid development of artificial intelligence technology, machine learning has been increasingly widely applied in the field of geosciences, showing significant potential in diagenetic facies identification. Traditional methods are often limited in processing complex geological data, making it difficult to effectively capture the complex nonlinear relationship between well logging data and lithology. Machine learning has unique advantages in processing large-scale and multi-dimensional data. However, research on diagenetic facies identification of tight sandstone based on machine learning is still relatively limited.
[0006] Existing diagenetic facies identification methods based on well logging curves typically use the original curves or simple combinations thereof (such as cross plots) as input features, failing to fully exploit the rich geological information contained in well logging data. Well logging curves not only contain instantaneous values of depth points but also contain the vertical variation patterns, spectral characteristics, and morphological features of strata. How to extract multi-view, multi-scale features from single well logging data and effectively fuse these features to improve the accuracy of diagenetic facies identification is a current research hotspot and challenge. Summary of the Invention
[0007] To address the aforementioned problems, this invention aims to provide a diagenetic facies identification method based on multimodal fusion features and an extreme random tree model.
[0008] The technical solution of the present invention is as follows: A diagenetic facies identification method based on multimodal fusion features and an extreme random tree model includes the following steps: S1: The lithofacies of the study area are divided into strongly compacted facies, calcareous cemented facies, siliceous cemented-medium dissolution facies, and strongly dissolution facies; S2: Obtain the logging curves of the study area, and determine the logging facies corresponding to each rock based on the logging curves; S3: Preprocess the logging curves and perform multi-view feature transformation to obtain a multimodal fusion feature set, and combine the multimodal fusion feature set with its corresponding lithofacies type to form a sample set; S4: Construct an extreme random tree model and train it using the sample set to obtain a diagenetic facies identification model; S5: Obtain the logging curve of the well to be logged, perform preprocessing and multi-view feature transformation on it, and then use the diagenetic facies identification model to predict its lithofacies type.
[0009] As a preferred option, in step S1, when classifying the lithofacies of the study area, the core data is used as a basis to establish a fine classification standard for the lithofacies of tight sandstone by determining the typical diagenetic facies markers of the study area, and lithofacies are classified according to it.
[0010] Preferably, the criteria for finely classifying the diagenetic facies of the dense sandstone are as follows: The pore types of the strongly compacted phase are mainly primary pores and gravel edge microcracks, with a matrix content of 2% to 14% and a porosity of <7%. The pore type of the calcareous cement phase is mainly primary pores, the content of heterogeneous matrix is 1%~5%, the average content of carbonate cement is 3.9%, and the porosity is <7%. The pore type of the silica cement-intermediate etched phase is mainly secondary pores, with a heterostructure content of 1%~4% and a porosity of 7%~10%. The pore type of the strongly eroded phase is mainly secondary pores, with a cement content of 0%~2%, a matrix content of 1%~3%, and a porosity >10%.
[0011] Preferably, in step S2, the logging curves include six types: GR, CNL, AC, RLLD, RLLS, and DEN.
[0012] Preferably, in step S3, the preprocessing includes outlier removal and standardization.
[0013] Preferably, step S3, performing multi-view feature transformation, specifically includes the following sub-steps: S31: Perform fractional derivative transformation on the preprocessed logging curve, and extract statistical features on the curve after fractional derivative transformation based on the sliding window to obtain the statistical feature mode enhanced by fractional derivative. S32: Perform wavelet transform or Fourier transform on the preprocessed logging curves to extract transform domain features and obtain transform domain feature modes; S33: Extract geometric morphological parameters from the preprocessed logging curves to obtain morphological feature modes; S34: Perform feature-level fusion of the fractional derivative-enhanced statistical feature mode, transform domain feature mode, and morphological feature mode to obtain the multimodal fusion feature set.
[0014] Preferably, in step S31, when performing fractional derivative, the order of the derivative is in the range of 0.2 to 0.8.
[0015] Preferably, in step S31, the statistical features include any one or more of the following: mean, variance, skewness, and kurtosis; in step S32, the transform domain features include any one or more of the following: frequency domain energy and wavelet coefficients; and in step S33, the geometric morphology parameters include any one or more of the following: slope, curvature, and concavity / convexity.
[0016] Preferably, in step S4, during training, the sample set is divided into a training set and a test set.
[0017] Preferably, in step S4, during training, the hyperparameters of the extreme random tree model are optimized using a Bayesian optimization method.
[0018] The beneficial effects of this invention are: This invention extracts multi-scale derived features from well logging curves by performing multi-view feature transformation, and uses the obtained multi-modal fusion feature set to train a model. This enables the obtained diagenetic facies identification model to use the vertical variation law, spectral features and morphological features in the well logging curves, thereby predicting diagenetic facies types faster and more accurately. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the diagenetic facies identification method based on multimodal fusion features and an extreme random tree model of the present invention. Figure 2 Here are typical photographs of various diagenetic facies types in a specific embodiment; Figure 3 This is a schematic diagram of the well logging intersection of diagenetic facies in a specific embodiment; Figure 4 This is a schematic diagram of an extreme random tree model in a specific embodiment; Figure 5 This is a schematic diagram illustrating the effect of the extreme random tree model of well Y1 on the identification of sandstone diagenetic facies in a specific embodiment. Detailed Implementation
[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and technical features described in this application can be combined with each other. It should also be pointed out that, unless otherwise indicated, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terms "comprising" or "including" and similar words used in this invention refer to elements or objects preceding the word that encompass the elements or objects listed following the word and their equivalents, without excluding other elements or objects.
[0022] like Figure 1 As shown, this invention provides a diagenetic facies identification method based on multimodal fusion features and an extreme random tree model, comprising the following steps: S1: The lithofacies of the study area are divided into strongly compacted facies, calcareous cemented facies, siliceous cemented-medium dissolution facies, and strongly dissolution facies.
[0023] In a specific embodiment, when delineating the lithofacies of the study area, core data is used as a basis to establish a fine-grained diagenetic facies classification standard for tight sandstone by identifying typical diagenetic facies markers of the study area, and lithofacies classification is performed accordingly; the fine-grained diagenetic facies classification standard for tight sandstone is as follows: The pore types of the strongly compacted phase are mainly primary pores and gravel edge microcracks, with a matrix content of 2% to 14% and a porosity of <7%; it is mainly found in sandstones with poor sorting, fine grain size and high matrix content. The pore type of the calcareous cement phase is mainly primary pores, with a matrix content of 1% to 5%, an average carbonate cement content of 3.9%, and a porosity of <7%. It is mostly found at lithological interfaces, especially near sandstone and mudstone interfaces. It can also be seen as basal or porous calcareous cement after weak or moderate compaction, and the later dissolution effect is weak. The pore type of the siliceous cement-medium dissolution phase is mainly secondary pores, with a matrix content of 1% to 4% and a porosity of 7% to 10%. It is mainly found in sandstones with good composition and structural maturity. Siliceous cement or a small amount of calcareous cement can also be seen after moderate or strong compaction, with a moderate degree of dissolution. The pore type of the strongly dissolved phase is mainly secondary pores, with cement content of 0%~2%, matrix content of 1~3%, and porosity >10%. It is mainly found in sandstone with good composition and structure maturity, with a small amount of siliceous and calcareous cementation. Dissolution is common, mainly intergranular and intragranular dissolution.
[0024] In the above embodiments, the types and characteristics of each diagenetic facies obtained according to the fine classification criteria for diagenetic facies of tight sandstone are shown in Table 1: Table 1. Diagenetic facies types and their characteristics
[0025] Taking the Huagang Formation reservoir of the YY structure as an example, typical photographs of various diagenetic facies types are as follows: Figure 2 As shown, Figure 2 In the diagram, a represents well YY2, H44, 4619.92m, strongly compacted phase; b represents well YY4, 4310.31m, calcareous cemented phase; c represents well YY2, H44, 4613.02m, siliceous cemented-medium dissolved phase; and d represents well YY2, H44, 4616.32m, strongly dissolved phase.
[0026] S2: Obtain the logging curves of the study area, and determine the logging facies corresponding to each rock based on the logging curves.
[0027] In one specific embodiment, the logging curves include six types of logging curves: GR (natural gamma), CNL (compensated neutron), AC (longitudinal wave transit time), RLLD (deep lateral resistivity), RLLS (shallow lateral resistivity), and DEN (density).
[0028] In the above embodiments, the GR can well reflect the grain size and clay content characteristics of the reservoir, while CNL and DEN are more sensitive to changes in porosity and can reflect the differences in porosity caused by different compaction strength, dissolution strength and cementation strength.
[0029] S3: Preprocess the well logging curves and perform multi-view feature transformation to obtain a multimodal fusion feature set, and combine the multimodal fusion feature set with its corresponding lithofacies type to form a sample set.
[0030] In one specific embodiment, the preprocessing includes outlier removal and standardization. It should be noted that outlier removal and standardization are existing technologies, and the specific processing methods will not be described in detail here.
[0031] In a specific embodiment, performing multi-view feature transformation includes the following sub-steps: S31: Perform fractional derivative transformation on the preprocessed logging curve, and extract statistical features on the curve after fractional derivative transformation based on the sliding window to obtain the statistical feature mode enhanced by fractional derivative. S32: Perform wavelet transform or Fourier transform on the preprocessed logging curves to extract transform domain features and obtain transform domain feature modes; S33: Extract geometric morphological parameters from the preprocessed logging curves to obtain morphological feature modes; S34: Perform feature-level fusion of the fractional derivative-enhanced statistical feature mode, transform domain feature mode, and morphological feature mode to obtain the multimodal fusion feature set.
[0032] In a specific embodiment, in step S31, when performing fractional derivative, the value range of the derivative order is 0.2 to 0.8. Optionally, the value range of the derivative order is 0.3 to 0.5. It should be noted that fractional derivative is a generalized form of integer derivative, which can enhance the local details of the curve while preserving the overall trend, and its derivative order can be adaptively selected according to the curve shape.
[0033] In a specific embodiment, in step S31, the statistical features include any one or more of the following: mean, variance, skewness, and kurtosis; in step S32, the transform domain features include any one or more of the following: frequency domain energy and wavelet coefficients; and in step S33, the geometric morphology parameters include any one or more of the following: slope, curvature, and concavity / convexity.
[0034] S4: Construct an extreme random tree model and train it using the sample set to obtain a diagenetic facies identification model.
[0035] In one specific embodiment, during training, the sample set is divided into a training set and a test set. Optionally, the ratio of the training set to the test set is 8:2.
[0036] In one specific embodiment, during training, the hyperparameters of the extreme random tree model are optimized using a Bayesian optimization method. Optionally, the hyperparameters include the number of trees, maximum depth, and minimum number of sample splits. In this embodiment, optimization using a Bayesian optimization method can yield a diagenetic facies identification model with higher prediction accuracy.
[0037] S5: Obtain the logging curve of the well to be logged, perform preprocessing and multi-view feature transformation on it, and then use the diagenetic facies identification model to predict its lithofacies type.
[0038] In a specific embodiment, taking the Huagang Formation reservoir in region Y as an example, the lithofacies identification method based on multimodal fusion features and extreme random tree model described in this invention is used to identify the lithofacies, including the following steps: (1) Based on core data, the typical diagenetic facies markers of the study area were determined, and the diagenetic facies of the tight sandstone in the study area were refined. The facies of the core section were divided into four types: strongly compacted facies, calcareous cemented facies, siliceous cemented-medium dissolution facies, and strongly dissolution facies.
[0039] (2) Taking advantage of the good vertical continuity and resolution of well logging curves, and combining well logging curve data, a well logging facies corresponding to the diagenetic facies type of the core was constructed. The results are as follows: Figure 3 As shown in Table 2: Table 2. Pearson correlation coefficients between diagenetic facies and well logging curves
[0040] From Table 2 and Figure 3 It can be seen that the greater the compaction strength, the greater the DEN, the smaller the AC, and the higher the CNL and GR. The depth range with greater dissolution strength generally shows medium CNL, low DEN and low AC, and the CNL porosity is generally higher than the AC porosity. Calcium cement has lower GR and AC, and higher DEN. Siliceous cement has lower GR, medium to low AC, and higher RLLS, RLLD and high DEN.
[0041] (3) The logging curves are preprocessed and multi-view feature transformations are performed to obtain a multimodal fusion feature set, and the multimodal fusion feature set and its corresponding lithofacies type are combined to form a sample set; In this embodiment, conventional logging curve data, after environmental correction and depth merging, were collected from multiple key wells in the study area to form the original dataset, which mainly includes six types of logging curves: GR, CNL, AC, RLLD, RLLS, and DEN. Simultaneously, lithological coding data precisely calibrated from core analysis data was loaded to establish a lithological tag library. The coding rules are as follows: 1 represents strongly compacted facies, 2 represents calcareous cemented facies, 3 represents siliceous cemented-medium dissolution facies, and 4 represents strongly dissolution facies.
[0042] During data preprocessing, environmental correction and depth merging were first performed on all curves, and cubic spline interpolation was used to standardize the sampling interval to 0.125m. After outliers were removed, normalization was performed. Then, fractional-order differential transformation was performed on each standardized logging curve. In this embodiment, considering the characteristics of the Huagang Formation logging curves in the Y region, the Grunwald-Letnikov definition was used to perform fractional-order differential transformations of order α=0.35 on the six curves: GR, CNL, AC, RLLD, RLLS, and DEN. The discrete calculation formula is as follows: (1) In the formula: For fractional differential operators (representing the differential operator with respect to the function) beg (order derivative) The order of the differential; The differentiable function (such as GR, CNL, etc.); The sampling interval; The cut-off length; For summation index; For depth.
[0043] In the implementation of this embodiment, it was found that when the sampling interval h = 0.125 m and the truncation length n = 10, the differentiated curve exhibited more significant fluctuations near the lithofacies interface, and the differences in curve morphology between different diagenetic facies were nonlinearly amplified. Comparative experiments showed that the fractional derivative order α was most effective in the range of 0.3–0.5; the enhancement effect was not significant when α < 0.2, and excessive noise amplification occurred when α > 0.8. After extracting the statistical features of the fractional derivative curve, the subtle differences in slope between the strongly compacted phase and the calcareous cemented phase, which were originally difficult to distinguish on the original density curve, were transformed into significant separations in skewness and kurtosis, increasing the classification interval by approximately 40%.
[0044] When acquiring the multimodal fusion feature set, the statistical feature modes are as follows: A sliding window size of w is set to 5, 11, or 21, and the mean, standard deviation, skewness, and kurtosis within the window are calculated for each logging curve at each depth point. The transform domain feature modes are also analyzed: a discrete wavelet transform is performed on each curve, using the db4 wavelet basis, and the decomposition is performed in three layers, extracting the energy and entropy of the detail coefficients in each layer. The morphological feature model is further analyzed: the first and second derivatives of each curve are calculated, and the slope, curvature, local maximum density, and local minimum density are extracted. These three modes are then concatenated to obtain the initial fusion features for each depth point, thus acquiring the multimodal fusion feature set. This multimodal fusion feature set and the corresponding lithofacies are then used as the sample set.
[0045] (4) Construct an extreme random tree model and train it using the sample set to obtain a diagenetic facies identification model; In this embodiment, the sample set is divided into a training set and a test set in an 8:2 ratio; basic model parameters are set, including an initial number of decision trees of 100-500, an initial maximum depth of decision trees of None, an initial minimum number of samples for node splits of 2, and an initial minimum number of samples for leaf nodes of 1; when tuning model parameters, 5-fold cross-validation is used, and Bayesian optimization is employed to find the optimal model parameters based on accuracy changes; the training set is then input into an extreme random tree model (such as...). Figure 4 As shown in the figure, a diagenetic facies identification model is obtained through training.
[0046] (5) Evaluate the diagenetic facies identification model and predict the lithofacies of the well to be logged; In this embodiment, the accuracy of the diagenetic facies identification model is evaluated using precision, recall, and F1 score. The results are as follows: Figure 5 As shown in Table 3; where, based on the "error-ambiguity decomposition" theory, it is assumed that the generalization error of the individual learner is E. i The weighted value of the generalization error of the learner for: (2) In the formula: The total number of individual learners; Let these be the weights of the i-th learner; Assume the divergence value of the individual learner is A. i Then the weighted divergence value of the learner for: (3) Generalization error after integration It can be represented as: (4) Table 3. Training performance of the model incorporating multimodal fusion features in this invention.
[0047] In addition, the extreme random tree model was trained using the original well logging curve data, and the training results are shown in Table 4: Table 4. Model training results using raw well logging data
[0048] Comparing Tables 3 and 4, it can be seen that the model accuracy improved by approximately 11 percentage points after introducing multi-attribute fusion features, resulting in a significant improvement in recognition precision. Except for the strongly eroded phase, which may have overfitted due to insufficient training samples, the recognition accuracy of the other three phases reached over 85%, and the average model accuracy was 87.1%.
[0049] When the diagenetic facies identification model described in this invention was used to predict the diagenetic facies of well Y1, the well logging interpretation results showed good correlation with the core interpretation, with an accuracy rate of over 85%, proving that this invention can accurately predict diagenesis using only well logging curves.
[0050] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A diagenetic facies identification method based on multimodal fusion features and an extreme random tree model, characterized in that, Includes the following steps: S1: The lithofacies of the study area are divided into strongly compacted facies, calcareous cemented facies, siliceous cemented-medium dissolution facies, and strongly dissolution facies; S2: Obtain the logging curves of the study area, and determine the logging facies corresponding to each rock based on the logging curves; S3: Preprocess the logging curves and perform multi-view feature transformation to obtain a multimodal fusion feature set, and combine the multimodal fusion feature set with its corresponding lithofacies type to form a sample set; S4: Construct an extreme random tree model and train it using the sample set to obtain a diagenetic facies identification model; S5: Obtain the logging curve of the well to be logged, perform preprocessing and multi-view feature transformation on it, and then use the diagenetic facies identification model to predict its lithofacies type.
2. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 1, characterized in that, In step S1, when classifying the lithofacies of the study area, the core data is used as a basis to establish a fine classification standard for the lithofacies of tight sandstone by determining the typical diagenetic facies markers of the study area, and lithofacies are classified according to it.
3. The diagenetic facies identification method based on multimodal fusion features and end-random tree model according to claim 2, characterized in that, The criteria for finely classifying the diagenetic facies of the dense sandstone are as follows: The pore types of the strongly compacted phase are mainly primary pores and gravel edge microcracks, with a matrix content of 2% to 14% and a porosity of <7%. The pore type of the calcareous cement phase is mainly primary pores, the content of heterogeneous matrix is 1%~5%, the average content of carbonate cement is 3.9%, and the porosity is <7%. The pore type of the silica cement-intermediate etched phase is mainly secondary pores, with a heterostructure content of 1%~4% and a porosity of 7%~10%. The pore type of the strongly eroded phase is mainly secondary pores, with a cement content of 0%~2%, a matrix content of 1%~3%, and a porosity >10%.
4. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 1, characterized in that, In step S2, the logging curves include six types: GR, CNL, AC, RLLD, RLLS, and DEN.
5. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 1, characterized in that, In step S3, the preprocessing includes outlier removal and standardization.
6. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 1, characterized in that, Step S3, the multi-view feature transformation specifically includes the following sub-steps: S31: Perform fractional derivative transformation on the preprocessed logging curve, and extract statistical features on the curve after fractional derivative transformation based on the sliding window to obtain the statistical feature mode enhanced by fractional derivative. S32: Perform wavelet transform or Fourier transform on the preprocessed logging curves to extract transform domain features and obtain transform domain feature modes; S33: Extract geometric morphological parameters from the preprocessed logging curves to obtain morphological feature modes; S34: Perform feature-level fusion of the fractional derivative-enhanced statistical feature mode, transform domain feature mode, and morphological feature mode to obtain the multimodal fusion feature set.
7. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 6, characterized in that, In step S31, when performing fractional derivative, the order of the derivative ranges from 0.2 to 0.
8.
8. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 6, characterized in that, In step S31, the statistical features include any one or more of the following: mean, variance, skewness, and kurtosis; in step S32, the transform domain features include any one or more of the following: frequency domain energy and wavelet coefficients; in step S33, the geometric morphology parameters include any one or more of the following: slope, curvature, and concavity / convexity.
9. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to claim 1, characterized in that, In step S4, during training, the sample set is divided into a training set and a test set.
10. The diagenetic facies identification method based on multimodal fusion features and an extreme random tree model according to any one of claims 1-9, characterized in that, In step S4, during training, the hyperparameters of the extreme random tree model are optimized using the Bayesian optimization method.