Oil saturation prediction method based on data driving and mechanism model fusion
By combining rock electrical experiments and machine learning, and utilizing Archie's formula and dynamic committee ensemble model, the complexity of oil saturation prediction caused by the heterogeneity of reservoir properties in Block M of the Junggar Basin was solved, achieving high-precision oil saturation prediction and providing a reliable basis for oil and gas exploration and development.
Patent Information
- Application Number
- CN202511043123.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-25
AI Technical Summary
The reservoir properties of Block M in the Junggar Basin are highly heterogeneous, making it difficult to predict oil saturation and thus difficult to accurately assess with existing technologies.
Combining rock electrical experiments and machine learning, oil saturation was calculated using Archie's formula. A machine learning model was trained using parameters such as porosity and resistivity, and a dynamic committee ensemble model was constructed to dynamically adjust the model weights to adapt to different geological conditions.
It effectively overcomes the problem of limited core sample quantity, improves the accuracy and generalization ability of oil saturation prediction, and provides strong support for oil and gas exploration and development.
Smart Images

Figure CN121009301A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geological exploration technology, and in particular to a method for predicting oil saturation based on data-driven and mechanistic model fusion. Background Technology
[0002] Oil saturation is a percentage of the effective pore volume of an oil-bearing reservoir to the effective pore volume of the rock. Changes in oil saturation can lead to increased heterogeneity within the reservoir. Accurate assessment of oil saturation can provide a better understanding of the reservoir's internal structural characteristics and optimize stratified exploitation strategies.
[0003] For Block M in the Junggar Basin, the strata are mainly lacustrine deposits with abundant interbedded conglomerate and sandstone-mudstone. This sedimentary environment leads to significant heterogeneity in reservoir properties. The physical properties of different rock types in this block vary considerably, and the region has experienced multiple tectonic movements, resulting in strong heterogeneity. This increases the complexity of predicting reservoir oil saturation and makes reservoir productivity prediction difficult. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a method for predicting oil saturation based on the fusion of data-driven and mechanistic models.
[0005] This invention is achieved through the following technical solution: The oil saturation prediction method based on the fusion of data-driven and mechanistic models provided by this invention includes the following steps: The porosity cementation index, saturation cementation index, proportionality coefficient between formation factors and porosity, and proportionality coefficient between resistivity increase coefficient and oil saturation were determined by rock electrical experiments for each core corresponding well section. Based on the porosity cementation index, saturation cementation index, the proportionality coefficient between formation factors and porosity, and the proportionality coefficient between resistivity increase coefficient and oil saturation, the oil saturation of the corresponding well section of each well is calculated using the Archie formula. The oil saturation values at different depths in each well section are used as the label values to be predicted. The main controlling factors affecting the oil saturation at the corresponding depth are used as input parameters to train the machine learning prediction model and obtain the trained machine learning prediction model. The oil saturation is predicted using a trained machine learning prediction model.
[0006] Optionally, based on the porosity value of each core, the formation water resistivity under full saturation, and the simulated formation water resistivity of each well, a functional relationship between the porosity of each core and formation factors can be fitted. The porosity cementation index and porosity cementation index are determined based on the functional relationship between each core sample and formation factors. By configuring simulated formation water to set different oil saturation values for each core and measuring the corresponding core resistivity, the functional relationship between different oil saturation values and resistivity increase coefficient for each core was fitted. Based on the relationship between different oil saturation values and resistivity increase coefficients of each core, the saturation cementation index and the proportionality coefficient between resistivity increase coefficient and oil saturation are determined.
[0007] The functional relationship between the porosity of each core and formation factors can be expressed as follows: ; In the above formula, F represents the formation factors calculated from each core sample; The measured porosity of each core sample is shown.
[0008] The oil saturation is calculated using the following formula: ; ; ; In the above formula, For oil saturation, Here, φ represents the formation water resistivity, and φ represents the porosity. Let m be the resistivity value when the rock pore space contains oil and gas, m be the porosity cementation index, and n be the... Saturation cementation index, where F represents formation factors. Let be the resistivity of formation water when the rock is fully saturated, 'a' be the proportionality coefficient between formation factors and porosity, 'I' be the resistivity increase coefficient, and 'b' be the proportionality coefficient between the resistivity increase coefficient and oil saturation.
[0009] Preferably, the main controlling factors affecting oil saturation are deep resistivity, acoustic transit time, formation density, compensated neutrons, resistivity of the flushed zone, and porosity.
[0010] Optionally, the following steps are also included before model training: Obtain the raw dataset, which includes oil saturation, porosity, GR, AC, CNL, DEN, RT, RXO, SP, CALC, CALI, depth, and lithological labels; Identify the main controlling factors affecting oil saturation.
[0011] Preferably, the method for determining the main controlling factor includes the following steps: using the Pearson correlation coefficient and the two rank correlation coefficients of Kendall's and Spearman to calculate the linear and nonlinear correlation between oil saturation and the corresponding influencing parameters respectively; taking the absolute value of the calculated values of the three correlation coefficients and performing a weighted average; if the weighted average of the absolute value of the correlation coefficient of one of the influencing parameters is higher than the mean of the weighted average of the absolute values of the correlation coefficients of all the influencing parameters, then the influencing parameter is considered to be the main controlling factor.
[0012] Optionally, the machine learning prediction model is one of BP neural network, LightGBM, GRNN, or Dilateformer.
[0013] Alternatively, the machine learning prediction model may be a dynamic committee ensemble model.
[0014] Optionally, the method for constructing a dynamic committee integration model includes the following steps: Fuzzy C-means clustering is used as a gate network for dynamic integration to divide the data base into subsets, dividing the dataset into K clusters, and the optimal matching relationship between each expert model and the subset is determined by the fuzzy C-means clustering method. Assign weights to the expert models within each cluster; For each sample in the dataset, calculate its membership degree in each cluster; Based on the membership degree of the sample in each cluster and the model weight of the corresponding cluster, the membership degree of the sample and the model weight value of the corresponding cluster are multiplied together to obtain the weight allocation result of each model to each sample in each cluster. Furthermore, the weight values of each model to a certain sample in all clusters are accumulated to obtain the dynamic weight of the expert model for that sample. The dynamic weights of each expert model on the samples are assigned, and the prediction results of each expert model are weighted and integrated to form a dynamic committee ensemble model.
[0015] Preferably, four expert models—BP neural network, LightGBM, GRNN, and Dilateformer—are used to construct the dynamic committee ensemble model.
[0016] Compared with the prior art, this application has at least the following beneficial effects: 1. This application is based on a data-driven and mechanism model fusion method for predicting oil saturation. By calculating the rock electrical parameters of different well sections, the rock electrical parameters of the entire well section are characterized, which effectively overcomes the difficulty in fully covering the geological conditions of a large well section due to the limited number of core samples. 2. The dynamic committee integration model constructed in this application can efficiently and accurately predict oil saturation, which can provide favorable support for oil and gas exploration and development; and the model can dynamically adjust the weights of each model according to different geological conditions and data characteristics, which can improve the prediction accuracy and generalization of the model. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a cross-plot of porosity and formation factors in the embodiments; Figure 2 This is a cross-plot of saturation and resistivity increase coefficient in the embodiment; Figure 3 This is a graph showing the correlation between oil saturation in the examples; Figure 4 The diagram shows the BP neural network structure in the embodiment. Figure 5 This is an example diagram of the LightGBM histogram algorithm in the embodiment; Figure 6 The GRNN network structure is shown in the example. Figure 7 This is a schematic diagram of multi-scale extended attention MSDA in the embodiment; Figure 8 This is a structural diagram of the multi-scale dilated attention mechanism in the embodiment; Figure 9 This is a flowchart illustrating the training process of the oil saturation model in this embodiment. Figure 10 This is a comparison chart of the oil saturation prediction results of each model in the embodiments; Figure 11 This is a flowchart of the prediction process for the dynamic committee integrated model in the embodiment; Figure 12 This is a framework diagram for predicting oil saturation based on a precision-weighted dynamic committee ensemble model in the embodiment. Figure 13 This is a graph showing the oil saturation prediction results based on the dynamic committee ensemble model in the example. Figure 14 This is a comprehensive interpretation diagram of reservoir geological parameters in well X1 in the embodiment; Figure 15 This is a comprehensive interpretation diagram of reservoir geological parameters in well X2 in the example. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other. It should also be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0021] The oil saturation prediction method based on the fusion of data-driven and mechanistic models disclosed in this embodiment includes the following steps: 1. Data collection and data preprocessing (1) Since the porosity of the study area directly determines the size of the space in the reservoir that can store fluids, and also indirectly affects the reserves of movable oil in the reservoir. The pore structure also affects the distribution and flow path of fluids. Larger pores are conducive to the storage and flow of oil, thereby increasing the oil saturation, while smaller or more complex pore structures may make it difficult for oil to be effectively stored and flowed, thus reducing the oil saturation. Previous statistical results show that there is a certain correlation between the oil level and physical properties of core samples. Generally, the higher the oil level, the better the physical properties. That is, in this embodiment, based on the nine logging curves of GR, AC, CNL, DEN, RT, RXO, SP, CALC, and CALI, as well as depth and lithology labels as characteristic factors affecting oil saturation, porosity is also regarded as a characteristic factor affecting oil saturation, so as to obtain an oil saturation prediction dataset.
[0022] Among them, GR represents natural gamma, reflecting the total content of radioactive elements in the formation; AC represents acoustic transit time, which is a measure of the difference between the speed of sound propagation in rock and the speed of sound propagation in a standard medium. It is often used to identify the lithology and thickness of formations and to identify oil and gas-bearing layers; CNL represents compensated neutrons, which refers to the thermal neutron flux caused by a neutron source measured along the well profile, used to estimate the hydrogen index and porosity of the formation; DEN represents formation density, which refers to the density value of different formations corresponding to different lithologies, used to determine the thickness and lithology of the formation; RT represents deep resistivity, which refers to the resistivity of the formation before drilling fluid invasion, used to determine the oil-bearing nature of the formation; RXO represents flush zone resistivity, which refers to the resistivity in the flush zone after drilling fluid invasion, used to determine the degree of influence of drilling fluid on the formation; SP represents spontaneous potential, which refers to the potential difference generated by the resistivity of drilling fluid adjacent to porous formations, used to delineate permeable layers, determine lithology and formation water properties; CALC represents well diameter difference, which refers to the difference between the actual well diameter and the standard well diameter. It is used to identify wellbore irregularities and provide a basis for formation evaluation; CALI stands for wellbore diameter, which usually refers to the diameter of the wellbore and is used to determine lithology, check casing condition, etc.
[0023] (2) In this embodiment, the reservoir geological parameter calculation model established in the same study area based on core data and well logging response characteristic analysis is used to calculate and compare the oil saturation data collected this time.
[0024] Considering that Archie's formula can effectively describe the relationship between rock resistivity, porosity and oil saturation, its calculation formula is as shown in (1). However, when calculating oil saturation based on Archie's formula, the selection of rock electrical parameters needs to be considered reasonably. Since only 5 rock cores were used to determine oil saturation in the data collected in this embodiment, the analysis results of these 5 rock cores were used as a reference, and the relationship between formation factors and porosity, water saturation and resistivity increase coefficient was clarified by combining formulas (2) and (3).
[0025] (1); (2); (3); In the above formula, Oil saturation; φ represents the resistivity of formation water; φ represents porosity. denoted as ρ, where ρ is the resistivity value when the rock pore space contains oil and gas; m is the porosity cementation index; n is the saturation cementation index; and F is the formation factor. denoted as , where is the resistivity of formation water when the rock is fully saturated; 'a' is the proportionality coefficient between formation factors and porosity; 'I' is the resistivity increase coefficient; and 'b' is the proportionality coefficient between the resistivity increase coefficient and oil saturation.
[0026] The rock electrical parameters of each core corresponding to the well section were determined by rock electrical experiments: Specifically, based on the experimentally measured porosity value of each core, the formation water resistivity under fully saturated conditions, and the simulated formation water resistivity of each well after configuring simulated formation water, the functional relationship between the porosity of each core and formation factors was fitted, and the rock electrical parameters a and m were determined by comparing with formula (2). For each core, different oil saturation values were set by configuring simulated formation water and the corresponding core resistivity was measured, and the functional relationship between different oil saturation values and resistivity increase coefficient of each core was fitted, and the rock electrical parameters b and n were determined by comparing with formula (3).
[0027] The functional relationship between the porosity of each core and formation factors can be expressed as follows: ; In the above formula, F represents the formation factors calculated from each core sample; The measured porosity of each core sample is shown.
[0028] By fitting the formula (3) above, the functional relationship between oil saturation and resistivity increase coefficient of different core samples from different wells is obtained in sequence, so as to determine the saturation cementation index n and the proportionality coefficient b between resistivity increase coefficient and oil saturation for each well. In this embodiment, the functional relationship between oil saturation and resistivity increase coefficient of five core samples from different wells is as follows: Well a core 1: ; Core sample 2 from well b: ; Core sample 3 from well b: ; Core 4 from well c: ; Core 5 from well d: ; In the above formula, I is the resistivity increase coefficient of each well core, and S w This represents the oil saturation of the core samples from each well.
[0029] Since there are two core samples for the corresponding well section of well b with fitting results of oil saturation and resistivity increase coefficient, the values of the saturation cementation index, resistivity increase coefficient and the proportional coefficient between oil saturation obtained from fitting the two core samples are weighted and averaged to characterize the saturation cementation index, resistivity increase coefficient and the proportional coefficient between oil saturation of well b.
[0030] Based on the obtained rock electrical parameters a, b, m, and n, the Archie formula is used as the oil saturation calculation model to calculate the oil saturation of the corresponding well section. Figure 1 and Figure 2(a)-(d) show the specific results of rock electrical testing. Table 1 shows the statistical results of rock electrical parameters measured in full-diameter rock samples from the study area.
[0031] Table 1: Statistical Table of Rock Electrical Parameters of Full-Diameter Rock Samples
[0032] The above experimental analysis results show that: the correlation coefficient b of the diameter cores of the Upper Urho Formation ranges from 1.18 to 1.41, with an average lithological correlation coefficient of 1.25; the core saturation index n ranges from 1.18 to 1.41, with an average lithological correlation coefficient of 1.25; the core saturation index n ranges from 2.20 to 3.37, with an average saturation index of 2.71; the lithological correlation coefficient a = 0.96, and the cementation index m = 1.53. Based on the calculated values of a, b, m, and n, corresponding well sections were selected. Combined with the Rt, Rw, and porosity values measured by rock electrical experiments at corresponding depths, the Archie formula was used to calculate the oil saturation at corresponding depths in each well section. This data set for predicting oil saturation in this embodiment was then established by combining the corresponding logging parameters, depth, lithological labels, and porosity values.
[0033] It is worth noting that in this embodiment, the a, b, m, and n values of a single well are obtained based on experimental data to calculate the data for that well. However, when used for the entire block, due to the influence of heterogeneity, the data from a single well may contain errors when applied to the entire block. Therefore, this calculated data is used as the input source for predicting the oil saturation of other wells.
[0034] The experimental dataset for this embodiment comes from measured core samples from the Urho Formation in Block M of the study area. Only five full-diameter core samples were collected for rock electrical experiments, and the corresponding well electrical parameters were obtained from these five core samples. Due to the limited number of core samples, it is difficult to comprehensively cover the geological conditions of a large well section. Furthermore, the study area exhibits strong heterogeneity; even within the same stratum, variations in lithology, inconsistent sedimentary patterns, interlayer distribution, and fracture development can lead to significant differences in lithology, physical properties, and oil content at different depths, resulting in inconsistent rock electrical parameters. Machine learning models excel at handling the nonlinear relationships between various logging and well logging parameters and oil saturation, learning complex geological patterns from limited samples and generalizing to unseen logging and well logging datasets. Therefore, this embodiment first determines the rock electrical parameters of the corresponding wells based on the five core samples through rock electrical experiments, and then uses Archie's formula as the oil saturation calculation model to calculate the oil saturation of each well. The calculated oil saturation values at different depths in each well section were input into the machine learning prediction model as the label values to be predicted. The depth, lithology labels, and logging parameters such as GR and AC at the corresponding depths were selected as the main controlling factors. These selected main controlling factors were then used as the input parameters for the machine learning prediction model, thus establishing an oil saturation prediction method based on the fusion of data-driven and mechanistic models. The trained machine learning prediction model was then applied to the oil saturation prediction of other wells. Table 2 shows some data from well D used for oil saturation prediction.
[0035] Table 2: Partial Data Table of SW Forecast
[0036] During the calculation of oil saturation in each well, a dataset of 3845 oil saturation data points was obtained. This dataset, used for predicting oil saturation, was divided into training and testing sets in a 7:3 ratio. Due to measurement errors and differences in evaluation metrics during data acquisition, a maximum-minimum normalization method was used to ensure that the distribution of each feature data point remained within a reasonable range.
[0037] 3. Analysis of Controlling Factors The Pearson correlation coefficient is a linear correlation function used to analyze the linear correlation between two features, given two feature sets X and Y.
[0038] in, , ,but The formula for calculating the Pearson correlation coefficient is: (4); In the above formula, and represents the average values of x and y, respectively, and r represents the Pearson correlation coefficient.
[0039] Kendall's correlation coefficient is a rank correlation coefficient used to measure the consistency of changes in the rank of two variables. For any two adjacent data points (Xi, Yi) and (Xj, Yj) at positions i and j in two feature sets X and Y, if the relationship between Xi and Xj is consistent with the relationship between Yi and Yj, then this set of data points is called a consistent pair; otherwise, it is called an inconsistent pair. Based on the ratio of consistent to inconsistent pairs in the two sets of data, Kendall's rank correlation coefficient is calculated to measure the consistency of changes in the rank of the two sets of variables. The formula is: (5); In the above formula, Kendall's correlation coefficient. For the number of consistent pairs, The number of inconsistent pairs, For the total logarithm, Let X be the number of parallel pairs. Let Y be the number of parallel pairs.
[0040] The Spearman correlation coefficient is used to measure the monotonic correlation between two variables. Unlike the Pearson correlation coefficient, which requires the data to follow a normal distribution, the Spearman correlation coefficient does not have strict requirements on the data distribution. When calculating the Spearman correlation coefficient of two sets of data, the data in the two feature sets X and Y are first assigned ranks in descending order. For each data point i, the difference between its ranks in X and Y is calculated, as shown in formula (13). The Spearman correlation coefficient of the two sets of data is then calculated based on the rank difference, as shown in formula (14).
[0041] (6); (7); In the above formula, , For data points Rank the data in the two sets; For data The difference in rank between the two sets of data.
[0042] The range of all three correlation coefficients is [-1, 1], where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation. In practical research, the strength of the correlation is usually graded based on the absolute value of the correlation coefficient. Specifically, 0-0.3 represents a weak or negligible correlation; 0.3-0.5 represents a low correlation; 0.5-0.9 represents a moderate to high correlation between the two variables; and values above 0.9 represent an extremely strong correlation between the variables.
[0043] The linear and nonlinear correlations between 11 parameters (GR, AC, CNL, DEN, RT, RXO, SP, CALC, CALI, Depth, and lithology) and oil saturation were calculated using Pearson correlation coefficient and two-rank correlation coefficients from Kendall's and Spearman, respectively. A weighted average was then applied to the results from the three correlation coefficient calculation methods to combine their performance under different data characteristics, resulting in a more comprehensive and robust correlation assessment. The correlation calculation results for the parameters are shown below. Figure 3 As shown in the figure. During the calculation process, all correlation parameters are taken as absolute values. The larger the mean absolute value of the weighted correlation coefficient of a certain parameter with oil saturation, the greater the influence of that parameter on oil saturation prediction.
[0044] The standard for selecting the main controlling factors of geological parameters of each reservoir is: if, for oil saturation, the weighted average of the absolute value of the correlation coefficient of one of the influencing parameters is higher than the average of the weighted average of the absolute values of the correlation coefficients of all influencing parameters under oil saturation, then the influencing parameter is considered to be the main controlling factor of oil saturation.
[0045] Depend on Figure 3 It can be seen that RT, AC, DEN, CNL, RXO, and porosity have a significant impact on the prediction of oil saturation. Therefore, these influencing factors are selected as the main controlling factors for oil saturation prediction.
[0046] 5. Build machine learning prediction models 5.1 Model Selection In this embodiment, four machine learning models—BP neural network, LightGBM model, GRNN neural network, and Dilateformer model—are selected to predict oil saturation.
[0047] (1) The full name of BP neural network is artificial neural network with backpropagation of error, such as Figure 4 As shown, the basic structure of a BP neural network consists of an input layer, hidden layers, and an output layer. Because BP neural networks have powerful nonlinear modeling capabilities, they can capture complex nonlinear relationships in reservoir geological parameters, which is very useful for understanding and predicting complex geological features.
[0048] (2) LIGHTGBM is a machine learning algorithm based on Gradient Boosting Tree (GBDT). It fits new decision trees sequentially along the negative direction of the decision tree loss function. Using a histogram algorithm, each continuous feature value is divided into a finite number of histograms, such as... Figure 5 As shown, the optimal split point is selected from the candidate set of each feature, and the corresponding bin is selected as the gain of the split point to accelerate model training. Furthermore, in Gradient One-Sided Sampling (GOSS), information gain is calculated only for samples with large gradients, and only a portion of samples with smaller gradients are randomly retained, thus preserving the distribution of the original data as much as possible. By bundling and merging mutually exclusive features, training efficiency is improved, ensuring good generalization ability while enhancing model accuracy.
[0049] Compared to traditional machine learning models such as linear regression and SVM, LIGHTGBM is more flexible, better able to capture complex mapping relationships in data, and can automatically select features from reservoir and well logging data that affect reservoir porosity and permeability prediction methods, reducing the need for manual intervention. Compared to convolutional neural networks (CNN), recurrent neural networks (RNN), and long short-term memory networks (LSTM), LIGHTGBM is more lightweight and suitable for resource-constrained environments. Deep learning models often have a black-box nature, with internal decision-making processes difficult to interpret, while LIGHTGBM can better reveal the influence of internal features on reservoir geological parameters, and can more intuitively analyze feature importance to guide further well logging evaluation studies.
[0050] Because the prediction of geological parameters for the sandstone and conglomerate reservoirs in Block M involves multiple geological and logging features across various dimensions, LightGBM employs a leaf node splitting strategy. It selects the largest leaf node for splitting at each step to fully discover and utilize the complex patterns and interactions within these multidimensional features, and can effectively leverage these features for pre-training. Given LightGBM's high efficiency in data processing and its strong ability to learn complex nonlinear relationships between features, it is worth considering combining it with other network models to construct a diverse and high-precision reservoir geological parameter prediction system, providing reliable geological data for oil and gas exploration and development in Block M.
[0051] (3) GRNN neural network, also known as generalized regression neural network, has the following structure: Figure 6As shown, this network is trained using nonparametric kernel regression. The regression value of the dependent variable relative to the independent variable is calculated by determining the correlation density function between the independent and dependent variables from the training samples. The GRNN neural network consists of four layers of neurons. The first layer is the input layer, with the same number of neurons as the dimension of the input vector in the training samples. Each neuron can be viewed as a simple distributed neural network. The second layer is the pattern layer, with the same number of neurons as the number of training samples. The summation layer sums the results from all training samples, and the output layer outputs the predicted result.
[0052] Because GRNN can quickly adapt to data distribution without requiring complex parameter tuning, it is suitable for complex and variable well logging and geological data. Furthermore, by adjusting the smoothing factor of the network, it can provide smooth prediction results, making it suitable for prediction tasks involving continuous variables with deep sequential characteristics. Compared to traditional machine learning methods, GRNN is more flexible, adept at capturing complex relationships in the data, and its parameter tuning is simpler and computationally more efficient than deep learning models.
[0053] (4) Dilateformer is a multi-scale dilated attention mechanism neural network. The Dilateformer network introduces dilated convolution operations into its multi-head attention mechanism module to capture the dependencies between features at different scales, thereby better understanding the complex structure and patterns of the input features. Furthermore, it improves the model's running efficiency through sparse selection. The Dilateformer model generally presents a four-layer pyramid structure, as shown below. Figure 7 As shown, the first two layers are Multi-Scale Extended Attention (MSDA). These two layers can be understood as performing a sliding window operation within a multi-head attention mechanism. The execution process is as follows: Figure 8 As shown. The last two layers are standard Multi-Head Self-Attention (MHSA). Overlapping markers are used in each layer, with zero-padding achieved through multiple overlapping 3×3 convolutional modules. The output resolution is adjusted by alternately controlling the stride of the convolutional kernel to 1 or 2. A convolutional overlap downsampling unit with an overlap kernel size of 3 and a stride of 2 is used to adapt each layer to inputs at different resolutions. Furthermore, whenever feature information is input to MSDA or MHSA, conditional embedding (CPE) is used to positionally encode the input and depth logging curves, enabling the model to better adapt to feature variations at different depth locations.
[0054] In MSDA, sliding window extended attention (SWDA) is performed using different inflation coefficients (r). Specifically, in the attention mechanism, the keys (K) and values (V) of each head are sparsely selected as feature elements within a sliding window centered on the red query element (Q). Self-attention is then calculated for these representative elements. The calculation formula is as follows: (8); In the above formula, X is the calculated self-attention score; Q, K, and V represent the query, key, and value matrices, respectively, with each row of each matrix representing a query / key / value feature vector; r is the inflation rate corresponding to different sliding windows; and in the original feature map... When querying at a location, SWDA will use... Self-attention computation is performed by sparsely selecting keys and values within a sliding window of size w×w centered on the network. This represents the x-coordinate value corresponding to the center point of the sliding window. This represents the ordinate value corresponding to the center point of the sliding window. The dimensions are the length and width of the rectangular sliding window.
[0055] Because reservoir geological parameters typically exhibit complex depth sequence relationships, the multi-head attention mechanism in the Dilateformer model effectively captures long-distance dependencies between logging data, aiding in understanding potential relationships between different locations within the reservoir. Furthermore, by adjusting the dilation rate, dilated convolution can extract features at different scales, increasing the amount of input data that each neuron in the neural network can perceive. This allows the model to cover a larger depth sequence region, something traditional machine learning models (such as linear regression and support vector machines) struggle to achieve. Unlike LSTM and RNN models, Dilateformer does not rely on the sequential information of the depth sequence, enabling more flexible processing of logging data across longer depth ranges.
[0056] 5.2 Calculation Model Evaluation Indicators This embodiment uses the coefficient of determination (COP). The root mean square error (RMSE) and mean absolute error (MAE) are used as evaluation indicators to measure the accuracy of each reservoir geological prediction model. The MAE represents the average of the sum of the absolute values of all prediction errors, and the RMSE represents the square root of the mean of the squared prediction errors for all data. The formulas for calculating these evaluation metrics are as follows: (9); (10); (12); in, This represents the actual values of reservoir geological parameters. ` represents the average value of the actual values of reservoir geological parameters. This represents the predicted value of the reservoir geological parameters, where i is the sample number. This represents the total number of samples in the dataset.
[0057] 5.3 Introduction to Model Hyperparameters (1) The main parameters and optimization range settings in the BP neural network prediction model are as follows: Hidden_layers_sizes: This refers to the number of hidden layer neurons in the network. Setting this parameter affects the model's expressive power. Setting too many layers increases training time and can lead to overfitting, while setting too few layers can prevent the model from fully capturing the feature representations between data points, potentially resulting in underfitting. In this embodiment, the hyperparameter optimization range for the BP neural network is set to (5, 30).
[0058] Learning_rate: In this embodiment, the learning rate hyperparameter optimization range of the BP neural network is set to (0.001, 0.1).
[0059] Maximum number of iterations: In this embodiment, the hyperparameter range for the maximum number of iterations is set to (10, 500).
[0060] Regularization coefficient: This parameter is used to adjust the complexity of the model during training. Setting it too high will lead to oversimplification of the model, resulting in underfitting; setting it too low will result in insufficient model complexity, leading to overfitting. In this embodiment, the optimization range of the regularization coefficient hyperparameter is set to (0.0001, 0.01).
[0061] (2) In the LightGBM model, the main hyperparameters and optimization range are set as follows: N_estimators: This refers to the number of decision trees, which affects the model's complexity. Too many trees may lead to overfitting, while too few trees may lead to underfitting, preventing the model from fully capturing data features. In this embodiment, the decision tree hyperparameter optimization range is set to (0, 300).
[0062] Max Depth: This limits the maximum growth depth of each decision tree. Its purpose is to regulate the complexity of the model training process to prevent overfitting. In this embodiment, the maximum growth depth hyperparameter is set to (1, 10).
[0063] Min Child Samples: This refers to the minimum number of leaf nodes in the tree. This parameter limits the growth of the decision tree and the complexity of the model. In this embodiment, the optimization range for the minimum number of leaf nodes is set to (1, 20).
[0064] Lambda_L1 and Lambda_L2 refer to the large gradient sample sampling rate and the large gradient random sample sampling rate, respectively. These two parameters allow the model to randomly select a subset of samples and features for training during a single iteration. Random sampling improves the model's generalization ability. The optimization range for both Lambda_L1 and Lambda_L2 is set to (0, 1).
[0065] (3) The prediction results of the GRNN neural network are related to the smoothing factor. If this parameter is set too large, the prediction result will be close to the mean of the sample data. If it is set too small, the model will be too sensitive to local changes in the input samples, resulting in an unsmooth prediction curve. In this embodiment, the optimization range of the smoothing factor is set to (0.1, 10).
[0066] (4) In the Dilateformer model, the main hyperparameters and optimization range are set as follows: Droupout refers to a technique used in neural networks to randomly disable neurons during training. Specifically, during forward propagation, a certain proportion of neurons are deactivated, effectively preventing overfitting and improving the model's generalization ability. In this embodiment, the optimization range for Droupout is set to (0.1, 1). The optimization ranges for other hyperparameters are as follows: Hidden_layers_sizes: (5, 30), Num_Layers: (0, 5), Learning_rate: (0.001, 0.1).
[0067] 6. Prediction of oil saturation Data anomalies and missing values were handled in the collected reservoir geological parameter dataset. The oil saturation dataset constructed using the Alfvé formula as the calculation model was normalized to eliminate dimensional differences between different features.
[0068] When predicting oil saturation parameters, considering that oil saturation is the ratio of oil volume to pore volume, in actual exploration, formations contain a certain amount of oil to varying degrees. Even if the oil content is extremely low, the oil saturation should be greater than 0. However, formation pores contain not only oil but also water, gas, other fluids, and solid components of the rock itself. That is, oil cannot occupy all the space in the pores; therefore, the oil saturation should be less than 1. Thus, this embodiment uses Archie's formula as the mechanism model for calculating oil saturation, combining actual lithology and well logging data to obtain an oil saturation dataset. When using this data-driven and mechanism model-based oil saturation prediction method, 0 < The value <1 is added as a boundary constraint to the loss functions of three neural networks: BP, GRNN, and Dilateformer, making the model training process more consistent with actual physical laws. The model training process for the oil saturation dataset is as follows: Figure 9 As shown.
[0069] The design idea of the overall loss function is as follows: First, design the original regression loss, and select the mean square error (MSE) during prediction as the basic loss function. Second, add a boundary constraint penalty term. When the predicted value of oil saturation is not within the range of [0,1], impose a penalty on the exceeded part to guide the predicted value of the model to return to a reasonable interval. The loss function with the boundary constraint term is shown in formula (12). That is, when the predicted value of oil saturation is, the penalty is ( -1), so that the model tends to reduce such predictions exceeding the upper limit during the training process, forcing the predicted value to approach 1; when is, the penalty is (- ), and the value of the penalty term is set to the opposite of the predicted value, encouraging the model to adjust its prediction result to make it closer to 0, ensuring that the model can meet the boundary constraint conditions during subsequent training. Use a penalty term in the form of a square to enhance the stability of the gradient change.
[0070] (12); In the above formula, is the total loss function considering physical constraints, is the loss function of the original mean square error, is the predicted value of a single sample, is a hyperparameter for controlling the penalty intensity, used to control the weight of the boundary constraint penalty term in the total loss function.
[0071] Divide the oil saturation dataset into a training set and a test set in a ratio of 7:3. Use four models to predict the oil saturation values calculated based on Archie's formula respectively. During the prediction process of the neural network, add 0 < Sw < 1 as a boundary condition to the loss function, making the training process of oil saturation more in line with actual physical laws. Among them, the prediction accuracy of the BP neural network is 0.83, the prediction accuracy of LightGBM is 0.89, the prediction accuracy of GRNN is 0.81, and the prediction accuracy of Dilateformer is 0.90. That is, the four models have different learning abilities for the general feature representation relationship within the oil saturation dataset. All four models can accurately predict the oil saturation values under boundary constraint conditions. Among them, the Dilatformer and LightGBM models can more accurately learn the non-linear relationship between the oil saturation value and the input features. The comparison results of the prediction results of each model are shown as Figure 10 (a)-(d). Table 3 shows the hyperparameter optimization results when each model predicts oil saturation.
[0072] Table 3: Hyperparameter Optimization Results Table for Oil Saturation Prediction
[0073] Table 4 records the prediction and evaluation indicators of oil saturation for the four models, and compares the prediction results of the four models with the traditional reservoir geological parameter calculation methods.
[0074] Table 4: Comparison of Prediction and Evaluation Indicators for Each Parameter
[0075] The comparison chart and table above show that the prediction error of the oil saturation is relatively small because the calculated oil saturation values are distributed in (0, 1). However, the overall prediction effect of oil saturation still needs to be improved.
[0076] 7. A method for predicting oil saturation based on data-driven and mechanistic model fusion 7.1 Committee Machine Prediction Model Committee machine learning is an ensemble neural network algorithm, a machine learning technique that improves overall prediction performance by combining multiple experts (i.e., individual models). Committee models are divided into static committee models and dynamic committee models. Static committee models use a fixed weighted average of the predictions from multiple experts, assigning fixed weights to each expert model. For data with multiple different patterns or distributions, fixed weights may not fully capture the features of all patterns, making it difficult to improve the training performance of individual experts. Dynamic committee models, on the other hand, use a gate network to divide the input data into multiple sub-tasks and assign them to the corresponding expert models. They then integrate the results from multiple different models to output the final prediction result. Fuzzy C-means clustering is typically used as the gate network for dynamic ensemble weight allocation. This method divides the data into sub-data based on the principle that the inter-class differences are sufficiently large and the intra-class differences are sufficiently small. The fuzzy C-means clustering calculation formula is: (13); In the above formula, is the loss function for fuzzy clustering, used to measure the quality of the current clustering result. U is the membership matrix, V is the cluster centroid, X is the input well logging data, M is the number of data entries, C is the number of clusters, and q is the fuzzy coefficient. For the i-th data point in the sample dataset, For the k-th cluster center, Let i represent the degree to which data point i belongs to cluster k, where i is the index of the base sample point and k is the index of the cluster center.
[0077] The construction method of the dynamic committee integration model includes the following steps: Fuzzy C-means clustering is used as a gate network for dynamic integration to divide the data base into subsets, dividing the dataset into K clusters, and the optimal matching relationship between each expert model and the subset is determined by the fuzzy C-means clustering method. Assign weights to the expert models within each cluster; For each sample in the dataset, calculate its membership degree in each cluster; Based on the membership degree of the sample in each cluster and the model weight of the corresponding cluster, the membership degree of the sample and the model weight value of the corresponding cluster are multiplied together to obtain the weight allocation result of each model to each sample in each cluster. Furthermore, the weight values of each model to a certain sample in all clusters are accumulated to obtain the dynamic weight of the expert model for that sample. The dynamic weights of each expert model on the samples are assigned, and the prediction results of each expert model are weighted and integrated to form a dynamic committee ensemble model.
[0078] The specific steps for implementing dynamic weight allocation using fuzzy C-means clustering are as follows: Step 1, Feature Clustering: Use the features of the training set to perform FCM clustering to divide the data into K clusters.
[0079] Step 2, evaluate model performance: Within each cluster, evaluate the performance of each base learner by weighted mean square error (WeightedMSE), as shown in Equation (14).
[0080] (14) In the above formula, It is a cluster The number of samples; It is an expert model For the sample The predicted value, It is a sample The true value, Let be the membership degree of sample j to cluster K.
[0081] Step 3, Weight Allocation: Based on the model's performance within each cluster Assign weights to the model for each cluster. For example, the calculation formula (15) suggests that the better the model performs, the higher its weight.
[0082] (15); In the above formula, ε is a small constant, introduced to avoid the situation where a certain model... When ε is 0, the denominator becomes 0, resulting in a division-by-zero error. Typically, ε takes values much smaller than the order of magnitude of the MSE, such as when the predicted value's MSE is on the order of magnitude of... - In between, MSE is in - Random numbers are generated between the given values.
[0083] Step 4, Prediction Phase: First, for each test sample... Calculate its membership degree to each cluster, that is, the degree to which the sample belongs to each cluster. If the data is divided into K clusters, then k takes values from 1 to K, and satisfies the following condition: Calculate the final weight of each model based on its membership degree and the model weight of its corresponding cluster. As shown in formula (16). The prediction results of the base learners are weighted based on dynamic weights, and the final ensemble prediction value is obtained by formula (17). .
[0084] (16); (17); By using a dynamic ensemble approach, the weights of each base expert are dynamically adjusted based on the characteristics of the input samples, the optimal matching relationship between each expert and the subset is determined, and more accurate prediction results are generated.
[0085] Directly combining samples and models to obtain prediction results can weaken the relationship between the model and samples to some extent. Therefore, this embodiment uses a dynamic weighted ensemble method to achieve weight allocation. First, weights are allocated to the model. For a single sample... Calculate its membership degree to each cluster. Furthermore, for the membership degree of the same sample in different clusters, the weighted sum of all membership degrees must be 1. Then, calculate the membership degree of each expert learner. For the sample final weight Based on the calculated dynamic weights, the optimal matching relationship between each expert and the subset of data is determined. The various expert models are then combined to form the final prediction model. The model training process is as follows: Figure 11 As shown.
[0086] Considering that LightGBM excels at handling complex relationships between logging features, BP neural networks can learn deep couplings between reservoir geological parameters through hidden layer structures, and GRNN is suitable for approximating complex functions, its probability density estimation characteristics can effectively reduce overfitting when the original lithological data calibration is limited. Dilateformer excels at capturing long-distance dependencies between logging data and can effectively capture long-term sedimentary cycles in the vertical direction of logging curves. Due to the complex geological conditions and strong reservoir heterogeneity of Block M, a committee machine model is constructed using the above four models. The weights of each model are dynamically adjusted according to different geological conditions and data characteristics to better adapt to the complex geological conditions of Block M.
[0087] When predicting oil saturation, the ensemble models that contributed the most to the prediction results were Dilateformer, GRNN, Dilateformer, and BP. Table 5 shows the dynamic weight allocation for each model. The overall framework of the improved accuracy-weighted dynamic committee ensemble model for predicting oil saturation is as follows: Figure 12 As shown.
[0088] Table 5: Model's Relationship to Weight Table During Prediction with Different Parameters
[0089] 7.2 Model Prediction Results like Figure 13 As shown, the oil saturation prediction results based on the dynamic committee ensemble model are presented in Table 6, which records the performance evaluation indicators of the ensemble model in predicting oil saturation.
[0090] Table 6: Evaluation Indicators for Oil Saturation Prediction by the Dynamic Committee Model
[0091] It can be seen that the constructed dynamic committee ensemble model significantly improves the prediction performance of oil saturation compared to the individual models. The model maintains a good fit between the predicted and actual oil saturation values, with a relatively uniform error distribution. Based on this model, the results show that the correlation coefficient of the oil saturation prediction is [not specified]. The accuracy reached 0.94, indicating that the integrated model can further improve the prediction accuracy of the model, and can adapt to the geological conditions of the study area more flexibly than the single model, thus improving the generalization ability of the model.
[0092] 7.3 Single-well processing interpretation application examples Based on the dynamic committee integration model established above, comprehensive well logging interpretation of reservoir geological parameters was carried out on X1 and X2, two verification wells in block M that participated in model training. Figure 14 , Figure 15The results show that the dynamic committee model's predicted curves for oil saturation are in high agreement with the core data curves. This indicates that the model maintains high predictive performance even with noise and uncertainty in the data.
[0093] This embodiment characterizes the rock electrical parameters of the entire well section by calculating the rock electrical parameters of one or two core samples from different well sections, effectively overcoming the difficulty in comprehensively covering a large area of geological conditions due to the limited number of core samples. Furthermore, due to the strong heterogeneity of the study area, even within the same strata, variations in lithology and inconsistent sedimentary patterns can lead to significant differences in lithology, physical properties, and oil content at different depths, resulting in inconsistent rock electrical parameters. The machine learning model in this application, compared to traditional formula fitting techniques, can handle the nonlinear relationships between various logging and well logging parameters and oil saturation, learn complex geological patterns from limited samples, and generalize to unseen logging and well logging datasets.
[0094] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for predicting oil saturation based on the fusion of data-driven and mechanistic models, characterized in that, Includes the following steps: The porosity cementation index, saturation cementation index, proportionality coefficient between formation factors and porosity, and proportionality coefficient between resistivity increase coefficient and oil saturation were determined by rock electrical experiments for each core corresponding well section. Based on the porosity cementation index, saturation cementation index, the proportionality coefficient between formation factors and porosity, and the proportionality coefficient between resistivity increase coefficient and oil saturation, the oil saturation of the corresponding well section of each well is calculated using the Archie formula. The oil saturation values at different depths in each well section are used as the label values to be predicted. The main controlling factors affecting the oil saturation at the corresponding depth are used as input parameters to train the machine learning prediction model and obtain the trained machine learning prediction model. The oil saturation of other wells is predicted using a trained machine learning prediction model.
2. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 1, characterized in that, Based on the porosity values of each core, the formation water resistivity under fully saturated conditions, and the simulated formation water resistivity of each well, the functional relationship between the porosity of each core and formation factors was fitted. The porosity cementation index and porosity cementation index are determined based on the functional relationship between each core sample and formation factors. By configuring simulated formation water to set different oil saturation values for each core and measuring the corresponding core resistivity, the functional relationship between different oil saturation values and resistivity increase coefficient for each core was fitted. Based on the relationship between different oil saturation values and resistivity increase coefficients of each core, the saturation cementation index and the proportionality coefficient between resistivity increase coefficient and oil saturation are determined.
3. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 1 or 2, characterized in that, The oil saturation is calculated using the following formula: ; ; ; In the above formula, For oil saturation, Here, φ represents the formation water resistivity, and φ represents the porosity. Let be the resistivity value when the rock pore space contains oil and gas, m be the porosity cementation index, n be the saturation cementation index, and F be the formation factor. Let be the resistivity of formation water when the rock is fully saturated, 'a' be the proportionality coefficient between formation factors and porosity, 'I' be the resistivity increase coefficient, and 'b' be the proportionality coefficient between the resistivity increase coefficient and oil saturation.
4. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 1, characterized in that, The main controlling factors affecting oil saturation are deep resistivity, acoustic transit time, formation density, compensated neutrons, resistivity of the flushed zone, and porosity.
5. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 1, characterized in that, Before training the model, the following steps are also included: Obtain the raw dataset, which includes oil saturation, porosity, GR, AC, CNL, DEN, RT, RXO, SP, CALC, CALI, depth, and lithological labels; Identify the main controlling factors affecting oil saturation.
6. The oil saturation prediction method based on data-driven and mechanistic model fusion according to any one of claims 1, 2, 4 or 5, characterized in that, The method for determining the controlling factors includes the following steps: The linear and nonlinear correlations between oil saturation and the corresponding influencing parameters were calculated using the Pearson correlation coefficient and the two rank correlation coefficients of Kendall's and Spearman, respectively. Take the absolute values of the three correlation coefficients and then perform a weighted average. If the weighted average of the absolute values of the correlation coefficients of one of the influencing parameters is higher than the mean of the weighted average of the absolute values of the correlation coefficients of all the influencing parameters, then that influencing parameter is considered to be the controlling factor.
7. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 1, characterized in that, The machine learning prediction model is one of the following: BP neural network, LightGBM, GRNN, and Dilateformer.
8. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 1, characterized in that, The machine learning prediction model is a dynamic committee ensemble model.
9. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 8, characterized in that, The method for constructing the dynamic committee integration model includes the following steps: Fuzzy C-means clustering is used as a gate network for dynamic integration to divide the data base into subsets, dividing the dataset into K clusters, and the optimal matching relationship between each expert model and the subset is determined by the fuzzy C-means clustering method. Assign weights to the expert models within each cluster; For each sample in the dataset, calculate its membership degree in each cluster; Based on the membership degree of the sample in each cluster and the model weight of the corresponding cluster, the membership degree of the sample and the model weight value of the corresponding cluster are multiplied together to obtain the weight allocation result of each model to each sample in each cluster. Furthermore, the weight values of each model to a certain sample in all clusters are accumulated to obtain the dynamic weight of the expert model for that sample. The dynamic weights of each expert model on the samples are assigned, and the prediction results of each expert model are weighted and integrated to form a dynamic committee ensemble model.
10. The oil saturation prediction method based on data-driven and mechanistic model fusion according to claim 8 or 9, characterized in that, The dynamic committee ensemble model is constructed using four models: BP neural network, LightGBM, GRNN, and Dilateformer.