A prediction method system and storage medium for microfractures in tight sandstone reservoirs
By combining drilling core and imaging logging with the Bayesian optimization random forest classification algorithm, a discrimination model for microfractures in tight sandstone reservoirs was established, which solved the difficult problem of microfracture identification in tight sandstone reservoirs and achieved higher recognition accuracy and simplified operation.
Patent Information
- Application Number
- CN202410927941.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-07-11
AI Technical Summary
Existing technologies make it difficult to effectively identify and predict the causes, development density, and dissolution product combinations of microfractures in tight sandstone reservoirs, resulting in difficulties in accurately predicting the physical property sweet spots of tight sandstone reservoirs.
By combining drilling core, imaging logging and casting thin section image analysis with the random forest classification algorithm using Bayesian optimization, a discriminant model for the origin, development density and combination of microfractures was established. By processing the logging curves through outlier removal, curve splicing and depth regression, a data set was constructed and classified and identified.
It improves the accuracy and systematicness of micro-fracture identification in tight sandstone reservoirs, simplifies the operation process, reduces the dependence on professional technicians, and is suitable for tight sandstone reservoir prediction in different regions.
Smart Images

Figure CN118938348B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of oil and gas exploration and development, and in particular relates to a method, system and storage medium for predicting microcracks in tight sandstone reservoirs. Background Art
[0002] Oil and gas exploration practice has demonstrated that the development of microfractures in tight sandstone reservoirs is a key factor in the development of physical sweet spots. Microfractures not only serve as pathways for fluid migration within tight sandstone reservoirs but also promote the dissolution of unstable minerals, resulting in a distribution of dissolution products with varying assemblages and abundances. Tight sandstone reservoirs undergo complex diagenetic alterations, resulting in complex identification of microfracture genesis and the distribution of dissolution products, severely hindering the accurate prediction of physical sweet spots within these reservoirs. Currently, quantitative prediction methods for the relationship between the genesis and development density of microfractures in tight sandstones and dissolution products remain to be established.
[0003] Currently, the main methods for identifying microfractures in tight sandstone reservoirs include: ① geological identification: identifying microfractures through field outcrop profiles and rock thin sections, and statistically analyzing microfracture-related parameters (such as aperture and filling degree); ② well logging identification: curve fitting based on the response characteristics of conventional well logging curves of microfractures, and establishing empirical formulas for identification; ③ experimental identification: identifying microfractures using core diversion tests, CT scanning, and acoustic emission tests; and ④ seismic identification: the presence of fractures will enhance formation anisotropy, and identification is based on the abnormal seismic reflection characteristics generated in seismic waves.
[0004] The above methods are currently important technical means for microfracture identification, but they all have certain shortcomings: ① The limitation of geological identification is that the characteristic parameters of microfractures are obtained through actual core samples, and the collection of drilling and field outcrop core samples is limited, which cannot meet the actual exploration and development needs; ② Well logging identification can be divided into imaging logging and conventional logging identification. The former method is intuitive and has high resolution, and can clearly determine the characteristic parameters of microfractures, but it is expensive and difficult to obtain in a comprehensive coverage area; the latter needs to consider the accuracy and matching degree of conventional logging curves in identifying microfractures, and requires a large amount of data to train the model; ③ The experimental method can identify the characteristic parameters of microfracture development and determine the main controlling factors of the development of dominant microfractures, but it is difficult to evaluate the impact of microfractures on the heterogeneity of tight sandstone reservoirs, and its guiding significance for oil and gas exploration and development is relatively limited; ④ The resolution of seismic identification is low, and it can often only carry out identification research on macro-scale fractures, resulting in low accuracy and strong subjectivity in the identification results. Therefore, establishing a comprehensive prediction method for the causes, development density and dissolution product combination of microfractures in tight sandstone reservoirs, and effectively improving the credibility of microfracture identification in tight sandstone reservoirs has become a technical issue that is of widespread concern in the industry and urgently needs to be solved. Summary of the Invention
[0005] The purpose of the present invention is to address the above-mentioned deficiencies in the prior art and provide a method, system and storage medium for predicting microcracks in tight sandstone reservoirs, which can essentially solve the current difficulties in the exploration and development of tight sandstone oil and gas reservoirs and can be applied to the identification and prediction of microcracks in tight sandstones in different regions.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The first object of the present invention is to provide a method for identifying microcracks in a tight sandstone reservoir, comprising the following specific steps:
[0008] S1. Obtaining original data
[0009] S11. Obtain the genesis type and development characteristic parameters of microcracks based on analysis of drilling cores, imaging logging images, and casting thin section images;
[0010] S12. Obtaining the characteristics of the dissolution products based on casting thin section image analysis;
[0011] S2. Data Processing
[0012] Based on the three evaluation parameters of microcracks in tight sandstone reservoirs obtained in step S1, namely, the genesis type, development characteristics, and abundance of dissolved products, a quantitative classification standard for each parameter is established, and the raw data with the combination type labeled "microcrack genesis, development density, and dissolved products" is uniformly classified.
[0013] S3, build dataset
[0014] Conventional logging curves were preprocessed by removing outliers, splicing curves, repositioning depths, and standardizing logging curves. Log curves sensitive to the combination of "microfracture genesis type, development density, and dissolution products" were selected. Sensitive logging curve data corresponding to different combinations of "microfracture genesis type, development density, and dissolution products" were obtained, and a logging database for different combinations of "microfracture genesis type, development density, and dissolution products" was established.
[0015] S4, performing missing value processing and outlier detection processing on the well logging database obtained in step S3 to construct a data set, and dividing the data set into a training data set, a validation data set, and a test data set;
[0016] S5. Select the optimal hyperparameters of the random forest classification algorithm through the Bayesian optimization algorithm;
[0017] The Bayesian optimization algorithm optimizes the hyperparameters of the random forest algorithm, including the number of random forest trees and the minimum number of leaf nodes. Bayesian optimization includes the following steps:
[0018] S51. Define a hyperparameter space and determine the hyperparameters to be optimized and their value ranges;
[0019] S52. Define the surrogate model, select an appropriate surrogate model, and use the Gaussian process;
[0020] S53, initialize the sample set, randomly select a set of initial samples in the hyperparameter space, and train the proxy model based on the performance evaluation of these samples;
[0021] S54. Iterative optimization: Based on the predictions of the surrogate model, select the next hyperparameter combination that is most likely to improve performance; evaluate the selected hyperparameter combination on the real model and record its performance; update the surrogate model with the new samples and update the estimate of the objective function in the hyperparameter space; repeat the above steps until the preset number of iterations is reached or the stopping criterion is met;
[0022] S55, output optimal hyperparameters;
[0023] S6. Substitute the selected optimal hyperparameters into the random forest algorithm training model to establish a discriminant model for the combination type of "microcrack genesis, development density and dissolution products";
[0024] S7. Perform cross-validation on the training set and evaluate the model prediction performance on the test set.
[0025] Furthermore, in step S11, the genesis types of the microcracks include tectonic genesis and diagenetic genesis, and the development characteristic parameters of the microcracks include opening, length and surface density. According to the surface density index, they are divided into low-density microcracks combination, surface density <5cm / cm 2 ; Medium density micro crack combination, 5cm / cm 2 <Surface density<10cm / cm 2 ; and high-density microcrack combination, surface density> 10cm / cm 2 .
[0026] Furthermore, in step S12, the characteristics of the dissolution products include product type, product content and development location. The abundance of the dissolution products is quantitatively counted by Image J image analysis software, and divided into a low dissolution product combination, product abundance <3%; a medium dissolution product combination, 3% < product abundance <8%; and a high dissolution product combination, product abundance >8%.
[0027] Furthermore, in step S2, the combination types are divided into 6 categories, specifically: Class I "structural-diagenetic microfractures, high development density, low dissolution products" combination, Class II "structural-diagenetic microfractures, medium development density, medium dissolution products" combination, Class III "structural-diagenetic microfractures, low development density, high dissolution products" combination, Class IV "diagenetic microfractures, high development density + low dissolution products" combination, Class V "diagenetic microfractures, medium development density, medium dissolution products" combination and Class VI "diagenetic microfractures, low development density, high dissolution products" combination.
[0028] Furthermore, in step S3, the logging curves include a natural potential curve, a natural gamma curve, a sonic transit time curve, a supplementary neutron curve, a supplementary density curve, a deep lateral resistivity curve, and a shallow lateral resistivity curve.
[0029] Furthermore, in step S3, the process of removing outliers is to delete the measured values at a certain sample point whose deviation from the mean value exceeds three times the standard deviation;
[0030] The process of curve splicing is to select the logging value of the formation at the same depth measured by two logging curves as the comparison standard, and make the two logging curves overlap by comparing and moving the logging curves;
[0031] The depth homing process is to select a GR curve with high vertical resolution and obvious characteristic marks as the standard curve, and determine the shift of other logging curves relative to the standard curve by comparison to complete the depth homing of the logging curves.
[0032] The process of logging curve standardization is to calibrate the logging curves of different wells, select the standard layer of key wells with homogeneous and stable formation distribution, and determine the correction amount required for each logging curve of other wells through the peak value of the logging curve frequency distribution histogram of the key drilling standard layer to complete the logging curve standardization process.
[0033] Furthermore, in step S4, the process of missing value processing is to use the coding in MATLAB software to traverse the original logging data column data to see if there are more than 20% missing values. If missing, delete them directly; the process of outlier detection processing is to select the sigma detection method to perform outlier detection processing on the data.
[0034] Furthermore, in step S4, the ratio of the training dataset, the validation dataset, and the test dataset is 8:1:1.
[0035] A second object of the present invention is to provide a system for identifying microcracks in tight sandstone reservoirs. The system is used to implement the above-mentioned method for identifying microcracks in tight sandstone reservoirs, and at least includes:
[0036] The data acquisition module is configured to acquire raw data of microfractures labeled with combinations of "microfracture genesis, development density, and dissolution products." Specifically, by using three evaluation parameters—genesis type, development characteristics, and dissolution product abundance—of microfractures in tight sandstone reservoirs, a quantitative classification standard is established for each parameter, thereby uniformly classifying the raw data labeled with combinations of "microfracture genesis, development density, and dissolution products."
[0037] The microcracks have two types of genesis: tectonic and diagenetic; the microcracks have two types of developmental characteristics: opening, length, and surface density; and the microcracks are divided into low-density microcracks, medium-density microcracks, and high-density microcracks according to the surface density index.
[0038] The characteristics of the dissolution products include product type, product content and development location. The abundance of the dissolution products was quantitatively analyzed using Image J image analysis software, and the dissolution products were divided into low dissolution product combination, medium dissolution product combination and high dissolution product combination according to the abundance index of the dissolution products.
[0039] The anomaly rejection module is configured to process well logging curves by removing outliers, splicing curves, repositioning depths, and normalizing well logging curves. It selects well logging curves that are sensitive to the combination of "microfracture genesis type, development density, and dissolution products," obtains sensitive well logging curve data corresponding to different combinations of "microfracture genesis type, development density, and dissolution products," and establishes a well logging database for different combinations of "microfracture genesis type, development density, and dissolution products."
[0040] A data set construction module is configured to perform missing value processing and outlier detection processing on the well logging database to construct a data set, and divide the data set into a training data set, a validation data set and a test data set;
[0041] The classification and discrimination module is configured to use the random forest classification algorithm selected by the Bayesian optimization algorithm as the discrimination model for the combination type of "microfracture genesis, development density, and dissolution products." Sample labels and logging data for different combinations of "microfracture genesis, development density, and dissolution products" are input into the model. The logging data is randomly divided into a training data set, a validation data set, and a test data set in a ratio of 8:1:1. After model training, a model classification confusion matrix is output, and the classification accuracy displayed by the confusion matrix is used to obtain the classification results.
[0042] The recognition module is configured to use the trained discrimination model to identify the combination type of "microcrack genesis, development density and dissolution products" of the microcracks to be tested.
[0043] A third object of the present invention is to provide a computer-readable storage medium storing a program, wherein the program can be executed by one or more processors to implement the above-mentioned method for identifying microfractures in tight sandstone reservoirs.
[0044] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0045] (1) The present invention provides a method for identifying microfractures in tight sandstone reservoirs. By comprehensively using drilling cores, imaging logging, cast thin sections and image analysis software, the method classifies the combination types of "microfracture genesis, development density and dissolution products" of tight sandstone. Conventional logging data is used to establish a discrimination model for different combination types through a Bayesian optimized random forest classification algorithm, thereby improving the recognition accuracy and essentially solving the current difficult problems in the exploration and development of tight sandstone reservoirs. The method makes the microfracture prediction technology of tight sandstone reservoirs more scientific, comprehensive and systematic, and can be applied to the prediction of tight sandstone reservoirs in different oil and gas basins.
[0046] (2) Compared with traditional methods, this method simplifies the operation process, reduces the dependence on professional technicians, and helps to quickly identify microfractures in tight sandstone reservoirs in a wider range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a technical flow chart of a method for identifying micro-fractures in tight sandstone reservoirs provided by the present invention;
[0048] Figure 2 This is a diagram of the genesis and development characteristics of microfractures in the tight sandstone reservoir of the Shihezi Formation in the Ordos Basin, provided by the present invention. In the diagram, (AB) imaging logging, structural microfractures; (CD) core image, structural microfractures; (E) casting thin section image, structural microfractures; (FG) casting thin section image, diagenetic microfractures;
[0049] Figure 3 This is a combined type diagram of "micro-crack genesis, development density and dissolution products" of the tight sandstone of the Shihezi Formation in the Ordos Basin provided by the present invention;
[0050] Figure 4 This is a schematic diagram of establishing a discriminant model based on the random forest algorithm provided by the present invention;
[0051] Figure 5 This is a basic principle diagram of the Bayesian optimization hyperparameter discrimination model provided by the present invention;
[0052] Figure 6a to Figure 6c It is a confusion matrix diagram of the data prediction results of the random forest classification model optimized by Bayesian method provided by the present invention;
[0053] Figure 7 This is a diagram of the identification results of the combination type of "micro-crack genesis, development density and dissolution products" of the tight sandstone of the Shihezi Formation in the Ordos Basin provided by the present invention;
[0054] Figure 8 The present invention provides a structural schematic diagram of a system for identifying micro-fractures in tight sandstone reservoirs. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the present invention more apparent, the following describes the specific embodiments of the present invention in further detail with reference to specific examples and accompanying drawings. Where specific techniques or conditions are not specified in the examples, the techniques or conditions described in the literature in the art or in the product instructions shall prevail.
[0056] like Figure 1 FIG. 1 is a flow chart of a method for identifying microcracks in a tight sandstone reservoir according to the present invention, which comprises the following steps:
[0057] (1) Obtaining raw data
[0058] The genesis type and development characteristic parameters of microcracks are obtained based on analysis of drilling cores, imaging logging images and casting thin section images.
[0059] The characteristics of the corrosion products were obtained based on the analysis of the casting thin section images;
[0060] (2) Data processing
[0061] Based on the three evaluation parameters of microcracks in tight sandstone reservoirs obtained in step S1, namely, the genesis type, development characteristics, and abundance of dissolved products, a quantitative classification standard for each parameter is established, and the raw data with the combination type labeled "microcrack genesis, development density, and dissolved products" is uniformly classified.
[0062] (3) Constructing a dataset
[0063] Conventional logging curves were preprocessed by removing outliers, splicing curves, repositioning depths, and standardizing logging curves. Log curves sensitive to the combination of "microfracture genesis type, development density, and dissolution products" were selected. Sensitive logging curve data corresponding to different combinations of "microfracture genesis type, development density, and dissolution products" were obtained, and a logging database for different combinations of "microfracture genesis type, development density, and dissolution products" was established.
[0064] (4) performing missing value processing and outlier detection processing on the well logging database obtained in step (3) to construct a data set, and dividing the data set into a training data set, a validation data set, and a test data set;
[0065] (5) Select the optimal hyperparameters of the random forest classification algorithm through the Bayesian optimization algorithm;
[0066] The Bayesian optimization algorithm optimizes the hyperparameters of the random forest algorithm, including the number of random forest trees and the minimum number of leaf nodes. Bayesian optimization includes the following steps:
[0067] (51) Define the hyperparameter space and determine the hyperparameters to be optimized and their value ranges;
[0068] (52) Define the surrogate model, select an appropriate surrogate model, and use Gaussian process;
[0069] (53) Initialize the sample set, randomly select a set of initial samples in the hyperparameter space, and train the proxy model based on the performance evaluation of these samples;
[0070] (54) Iterative optimization: based on the prediction of the surrogate model, select the next hyperparameter combination that is most likely to improve performance; evaluate the selected hyperparameter combination on the real model and record its performance; update the surrogate model with the new samples and update the estimate of the objective function in the hyperparameter space; repeat the above steps until the preset number of iterations is reached or the stopping criterion is met;
[0071] (55) Output the optimal hyperparameters;
[0072] (6) Substitute the selected optimal hyperparameters into the random forest algorithm training model to establish a discriminant model for the combination type of “microcrack genesis, development density and dissolution products”;
[0073] (7) Cross-validation was performed on the training set and model prediction performance was evaluated on the test set.
[0074] The technical terms in the present invention are explained as follows:
[0075] Drilling core refers to the use of coring tools to obtain underground rock blocks to the surface during the drilling process. This type of rock block is called a drilling core.
[0076] Development density refers to the width of cracks per unit area;
[0077] The abundance of dissolution products refers to the content of secondary products formed by mineral dissolution in the core sample;
[0078] Well logging curve refers to the curve formed during the logging process, which can reflect the characteristics of different lithologies and layers;
[0079] Sensitive logging curve, in this patent, means a logging curve that is sensitive to the changing characteristics of the combination of microfractures and their dissolution products;
[0080] Well logging data refers to the data values obtained from different well logging curves.
[0081] Example 1
[0082] The specific technical solution of the present invention is illustrated by taking the tight sandstone reservoir of the Shihezi Formation of the Permian System in the Fuxian area of the Ordos Basin as an example.
[0083] Step 1: Development characteristics of microcracks in tight sandstone
[0084] Based on the comprehensive analysis of drilling cores, imaging logging images, and casting thin section images, it is believed that the tight sandstone reservoirs of the Shihezi Formation have two types of microfractures: structural microfractures and diagenetic microfractures (e.g. Figure 2 As shown in Table 1, the opening of the structural microcracks is greater than 100 μm, the length is greater than 10 mm, and the surface density is generally greater than 5 cm / cm 2 The opening of diagenetic microcracks is less than 100 μm, the length is less than 10 mm, and the surface density is generally less than 5 cm / cm 2 According to the micro crack surface density index, it is divided into low density (<5cm / cm 2 ), medium density (5~10cm / cm 2 ) and high density (>10cm / cm 2 ) Microcrack combination.
[0085] Table 1. Statistics of microcracks genesis types and characteristic parameters
[0086] Microcracks genesis type Opening (μm) Length (mm) <![CDATA[Areal density (cm / cm 2 )]]> Tectonic causes >100 >10 >5 Diagenesis <100 <10 <5
[0087] Step 2: Characteristics of Dissolution Products
[0088] Cast thin-section image analysis revealed that the dissolved minerals in the Shihezi Formation's tight sandstone reservoir primarily consist of feldspar and lithic particles, along with minor carbonate cements. Dissolution products primarily consist of quartz, clay, and carbonate minerals. ImageJ analysis software was used to quantitatively analyze the abundance of dissolved products, and the resulting assemblage was categorized into low (<3%), moderate (3%-8%), and high (>8%) dissolved product groups based on their abundance.
[0089] Step 3: Combination of “microcrack genesis, development density, and dissolution products”
[0090] Taking into account the three indicators of micro-fracture genesis, micro-fracture development density and dissolution product abundance of the Shihezi Formation tight sandstone reservoir, the combination types of "micro-fracture genesis, development density and dissolution product" of the Shihezi Formation tight sandstone are divided into 6 categories (such as Figure 3 (shown): ① Type I combination of “structural-diagenetic microfractures, high development density, and low dissolution products” ( Figure 3 Medium A); ② Type II combination of “tectonic-diagenetic microfractures, medium development density, and medium dissolution products” ( Figure 3Middle B); ③ Type III combination of “tectonic-diagenetic microfractures, low development density, and high dissolution products” ( Figure 3 Middle C); ④ Type IV combination of “diagenetic micro-fractures, high development density + low dissolution products” ( Figure 3 Medium CD); ⑤ Type V “diagenetic microcracks, medium development density, medium dissolution products” combination ( Figure 3 Middle E); ⑥ Type VI “diagenetic micro-fractures, low development density, high dissolution products” combination ( Figure 3 Middle F).
[0091] Step 4: Logging curve processing and logging database
[0092] The logging curve data of the tight sandstone reservoir of the Permian Shihezi Formation in the Ordos Basin were processed, mainly including outlier removal, curve splicing, depth regression correction and logging curve standardization. The purpose is to eliminate the depth error and offset error between each logging series and ensure the consistency of core depth and logging depth.
[0093] Well logging curve analysis revealed that: (1) spontaneous potential (SP) and natural gamma ray (GR) curves have good response characteristics to the brittle minerals, plastic minerals, and shale content of tight sandstone reservoirs; (2) acoustic transit time (AC), supplementary neutron (CNL), and supplementary density (DEN) curves can reflect the reservoir porosity and permeability conditions and the degree of microfracture development; (3) the response of deep lateral resistivity (LLD) and shallow lateral resistivity (LLS) curves is closely related to the properties of formation water, oil and gas content, and the degree of dissolution development. Therefore, seven sensitive logging curves, GR, SP, DEN, CNC, AC, LLS, and LLD, were selected and missing value processing and outlier detection were performed.
[0094] Based on core observation, thin section analysis, fracture genesis type and development density, and quantitative statistics of dissolution product abundance of the Shihezi Formation tight sandstone in the Ordos Basin, the logging curve response characteristics of six types of "microfracture genesis, development density and dissolution product" combinations are summarized: ① Type I "structural-diagenetic microfractures, high development density, low dissolution products" combination has the characteristics of "low SP, high GR, low LLD and LLS, high AC, low DEN, and high CNL"; ② Type II "structural-diagenetic microfractures, medium development density, and medium dissolution products" combination has the characteristics of "high SP, low GR, high LLD and LLS, medium AC, medium DEN, and low CNL"; ③ Type III "structural-diagenetic microfractures, low development density, and medium dissolution products" combination has the characteristics of "high SP, low GR, high LLD and LLS, medium AC, medium DEN, and low CNL"; The combination of 4 types of "diagenetic microfractures, high development density and low dissolution products" has the characteristics of "low SP, medium GR, medium LLD and LLS, high AC, medium DEN, and medium CNL"; ④Type IV "diagenetic microfractures, high development density + low dissolution products" has the characteristics of "medium SP, high GR, medium LLD and LLS, low AC, low DEN, and high CNL"; ⑤Type V "diagenetic microfractures, medium development density, and medium dissolution products" has the characteristics of "high SP, low GR, high LLD and LLS, medium AC, high DEN, and low CNL"; ⑥Type VI "diagenetic microfractures, low development density, and high dissolution products" has the characteristics of "medium SP, medium GR, low LLD and LLS, low AC, high DEN, and medium CNL". Through the above logging response characteristic analysis and data extraction, a logging database for different combinations of "microfracture genesis type, development density, and dissolution products" was established.
[0095] Step 5: Discriminant Model
[0096] (1) Using the code in MATLAB software, the original well logging data columns are traversed to see if there are more than 20% missing data. If missing, they are directly deleted, and then the initial data detection is performed on the row data. The traversal data automatically identifies the character data in the original data and converts it into floating point data.
[0097] (2) Select the sigma detection method to detect outliers in the data. The specific steps are as follows: ① Read the data set X, which contains n data points; ② Calculate the mean (μ) and standard deviation (σ) of the data set; ③ According to the properties of the normal distribution, about 68% of the data points fall within the range of ±1σ of the mean, about 95% of the data points fall within the range of ±2σ of the mean, and about 99.7% of the data points fall within the range of ±3σ of the mean; ④ Using the 3σ principle, define data points in the data set that differ from the mean by more than 3 times the standard deviation as outliers; ⑤ Check whether there are outliers in the data set. If so, mark them as outliers.
[0098] (3) Based on the missing value processing and outlier detection processing, MATLAB software is used to build a discriminant model based on the random forest algorithm, and Bayesian optimization processing is performed at the same time. The specific steps are as follows:
[0099] ① Data preparation: The data required for this invention are logging data values of six different combinations of "microfracture genesis, development density and dissolution products" of the Shihezi Formation tight sandstone (as shown in Table 2). In order to test the accuracy of the training model, all the data are randomly divided according to the ratio of "training set: validation set: test set = 8:1:1". Each sample has a set of features and corresponding category labels. These samples will be used to train the classification model.
[0100] Table 2. Well logging database of six combinations of “microfracture genesis, development density, and dissolution products” in the tight sandstone of the Shihezi Formation
[0101]
[0102]
[0103] ② Random forest algorithm model establishment
[0104] The random forest classification model is implemented by calling the TreeBagger function that comes with MATLAB. The random forest algorithm is a collection of multiple unpruned decision trees. The training set of each decision tree is sampled from the original training set using the bagging method. During the construction of each decision tree, each node selects the best split feature based on information gain. Ultimately, each tree votes with equal weight to determine the prediction type for each instance, which is used as the final classification result (e.g. Figure 4 shown).
[0105] ③ Bayesian optimization random forest classification algorithm
[0106] There are some drawbacks when using random forest algorithm for classification: when the correlation between trees is greater, the error rate is greater; when the training data is noisy, it is easy to produce overfitting. To solve the above shortcomings, the Bayesian optimization random forest classification algorithm (such as Figure 5 shown):
[0107] 1) Define the hyperparameter space: determine the hyperparameters to be optimized and their value ranges;
[0108] 2) Define the surrogate model: Select an appropriate surrogate model and use Gaussian process;
[0109] 3) Initialization sample set: Randomly select a set of initial samples in the hyperparameter space and train the proxy model based on the performance evaluation of these samples;
[0110] 4) Iterative Optimization: Based on the surrogate model’s predictions, select the next hyperparameter combination that is most likely to improve performance; evaluate the selected hyperparameter combination on the real model and record its performance; update the surrogate model with the new samples and update the estimate of the objective function in the hyperparameter space; repeat the above steps until the preset number of iterations is reached or the stopping criterion is met;
[0111] 5) Output the best hyperparameters: Based on the optimization results, output the best hyperparameter combination for training the final model;
[0112] 6) The optimization results show that when the number of decision trees is 167 and the minimum number of leaf nodes is 3, the classification result is the best; the results show that all data sets are divided into a ratio of training set: validation set: test set = 8:1:1 and then the classification and discrimination model is trained, validated and tested (such as Figures 6a-6c As shown); the confusion matrix of the training set, each row represents the true category of the data, and each column represents the predicted category given by the model. Taking the first column of the confusion matrix as an example, it can be determined that 16 of the 19 data generated by the classification model are accurate and 3 data categories are wrong. Finally, the accuracy of the model training set is 0.95 (as shown Figure 6a ); The principles of the verification set confusion matrix and the test set confusion matrix are as above, where the verification set accuracy is 0.85 (such as Figure 6b ), the test set accuracy is 0.90 (such as Figure 6c ).
[0113] Step 6: Output and verification of judgment results
[0114] The established model was used to identify the tight sandstone reservoirs of the typical wells of the Shihezi Formation in the Ordos Basin, and the identification results of different combinations of "microfracture genesis type, development density and dissolution products" were obtained. After comparison with the actual drilling core and thin section results, the accuracy of the identification results was 87.3% (such as Figure 7 Therefore, it is believed that this method can be promoted and used.
[0115] Example 2
[0116] This embodiment, based on the design of embodiment 1, discloses a system for identifying micro-fractures in tight sandstone reservoirs. Figure 8 As shown, the system is used to implement the above-mentioned method for identifying micro-fractures in tight sandstone reservoirs, and at least includes the following components built into the system:
[0117] The data acquisition module is configured to acquire raw data of microfractures labeled with combinations of "microfracture genesis, development density, and dissolution products." Specifically, by using three evaluation parameters—genesis type, development characteristics, and dissolution product abundance—of microfractures in tight sandstone reservoirs, a quantitative classification standard is established for each parameter, thereby uniformly classifying the raw data labeled with combinations of "microfracture genesis, development density, and dissolution products."
[0118] The microcracks have two types of genesis: tectonic and diagenetic; the microcracks have two types of developmental characteristics: opening, length, and surface density; and the microcracks are divided into low-density microcracks, medium-density microcracks, and high-density microcracks according to the surface density index.
[0119] The characteristics of the dissolution products include product type, product content and development location. The abundance of the dissolution products was quantitatively analyzed using Image J image analysis software, and the dissolution products were divided into low dissolution product combination, medium dissolution product combination and high dissolution product combination according to the abundance index of the dissolution products.
[0120] The anomaly rejection module is configured to process well logging curves by removing outliers, splicing curves, repositioning depths, and normalizing well logging curves. It selects well logging curves that are sensitive to the combination of "microfracture genesis type, development density, and dissolution products," obtains sensitive well logging curve data corresponding to different combinations of "microfracture genesis type, development density, and dissolution products," and establishes a well logging database for different combinations of "microfracture genesis type, development density, and dissolution products."
[0121] A data set construction module is configured to perform missing value processing and outlier detection processing on the well logging database to construct a data set, and divide the data set into a training data set, a validation data set and a test data set;
[0122] The classification and discrimination module is configured to use the random forest classification algorithm selected by the Bayesian optimization algorithm as the discrimination model for the combination type of "microfracture genesis, development density, and dissolution products." Sample labels and logging data for different combinations of "microfracture genesis, development density, and dissolution products" are input into the model. The logging data is randomly divided into a training data set, a validation data set, and a test data set in a ratio of 8:1:1. After model training, a model classification confusion matrix is output, and the classification accuracy displayed by the confusion matrix is used to obtain the classification results.
[0123] The recognition module is configured to use the trained discrimination model to identify the combination type of "microcrack genesis, development density and dissolution products" of the microcracks to be tested.
[0124] Example 3
[0125] This embodiment, based on the design basis of Embodiments 1 and 2, discloses a computer-readable storage medium storing a program that can be executed by one or more processors to implement the above-mentioned method for identifying microfractures in tight sandstone reservoirs.
[0126] In the absence of conflict, the above embodiments and features in the embodiments may be combined with each other.
[0127] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying microcracks in tight sandstone reservoirs, characterized in that: The specific steps include: S1. Obtaining original data S11. Obtain the genesis type and development characteristic parameters of microcracks based on analysis of drilling cores, imaging logging images, and casting thin section images; S12. Obtaining the characteristics of the dissolution products based on casting thin section image analysis; S2. Data Processing Based on the three evaluation parameters of microfracture genesis, development characteristics, and dissolution product abundance of the tight sandstone reservoir obtained in step S1, a quantitative classification standard for each parameter is established, and the raw data with the combination type labeled "microfracture genesis, development density, and dissolution product" is uniformly classified; S3, build dataset Conventional logging curves were preprocessed by removing outliers, splicing curves, repositioning depths, and standardizing logging curves. Log curves sensitive to the combination of microfracture genesis, development density, and dissolution products were selected. Sensitive logging curve data corresponding to different combinations of microfracture genesis, development density, and dissolution products were obtained, and a logging database for different combinations of microfracture genesis, development density, and dissolution products was established. S4, performing missing value processing and outlier detection processing on the well logging database obtained in step S3 to construct a data set, and dividing the data set into a training data set, a validation data set, and a test data set; S5. Select the optimal hyperparameters of the random forest classification algorithm through the Bayesian optimization algorithm; The Bayesian optimization algorithm optimizes the hyperparameters of the random forest algorithm, including the number of random forest trees and the minimum number of leaf nodes. Bayesian optimization includes the following steps: S51. Define a hyperparameter space and determine the hyperparameters to be optimized and their value ranges; S52, define the proxy model, select the appropriate proxy model, and use the Gaussian process; S53, initialize the sample set, randomly select a set of initial samples in the hyperparameter space, and train the proxy model based on the performance evaluation of these samples; S54. Iterative optimization: Based on the predictions of the surrogate model, select the next hyperparameter combination that is most likely to improve performance; evaluate the selected hyperparameter combination on the real model and record its performance; update the surrogate model with the new samples and update the estimate of the objective function in the hyperparameter space; repeat the above steps until the preset number of iterations is reached or the stopping criterion is met; S55, output optimal hyperparameters; S6. Substitute the selected optimal hyperparameters into the random forest algorithm training model to establish a discriminant model for the combination of "microcrack genesis, development density, and dissolution products"; S7. Perform cross-validation on the training set and evaluate the model prediction performance on the test set.
2. The method for identifying microcracks in a tight sandstone reservoir according to claim 1, wherein: In step S11, the genesis types of the microcracks include tectonic genesis and diagenetic genesis, and the development characteristic parameters of the microcracks include opening, length and surface density. According to the surface density index, they are divided into low-density microcracks combination, surface density <5cm / cm 2 ; Medium density micro crack combination, 5cm / cm 2 <Surface density<10cm / cm 2 ; and high-density microcrack combination, surface density> 10cm / cm 2 .
3. The method for identifying microcracks in a tight sandstone reservoir according to claim 2, wherein: In step S12, the characteristics of the dissolution products include product type, product content and development location. The abundance of the dissolution products is quantitatively counted by Image J image analysis software, and the dissolution products are divided into a low dissolution product combination, with product abundance <3%; a medium dissolution product combination, with 3% < product abundance <8%; and a high dissolution product combination, with product abundance >8%.
4. The method for identifying micro-fractures in a tight sandstone reservoir according to claim 3, wherein: In step S2, the combination types are divided into 6 categories, specifically: Class I "structural-diagenetic microfractures, high development density, low dissolution products" combination, Class II "structural-diagenetic microfractures, medium development density, medium dissolution products" combination, Class III "structural-diagenetic microfractures, low development density, high dissolution products" combination, Class IV "diagenetic microfractures, high development density + low dissolution products" combination, Class V "diagenetic microfractures, medium development density, medium dissolution products" combination and Class VI "diagenetic microfractures, low development density, high dissolution products" combination.
5. The method for identifying micro-fractures in a tight sandstone reservoir according to claim 1, wherein: In step S3, the logging curves include a natural potential curve, a natural gamma curve, a sonic transit time curve, a supplementary neutron curve, a supplementary density curve, a deep lateral resistivity curve, and a shallow lateral resistivity curve.
6. The method for identifying micro-fractures in a tight sandstone reservoir according to claim 5, wherein: In step S3, the process of outlier elimination is to delete the measured values at a certain sample point whose deviation from the mean value exceeds three times the standard deviation; The process of curve splicing is to select the logging value of the formation at the same depth measured by two logging curves as the comparison standard, and make the two logging curves overlap by comparing and moving the logging curves; The depth homing process is to select a GR curve with high vertical resolution and obvious characteristic marks as the standard curve, and determine the shift of other logging curves relative to the standard curve by comparison to complete the depth homing of the logging curves. The process of logging curve standardization is to calibrate the logging curves of different wells, select the standard layer of key wells with homogeneous and stable formation distribution, and determine the correction amount required for each logging curve of other wells through the peak value of the logging curve frequency distribution histogram of the key drilling standard layer to complete the logging curve standardization process.
7. The method for identifying micro-fractures in a tight sandstone reservoir according to claim 1, wherein: In step S4, the process of missing value processing is to use the code in MATLAB software to traverse the original logging data column to see if there are more than 20% missing values. If missing, delete them directly; the process of outlier detection processing is to select the sigma detection method to perform outlier detection processing on the data.
8. The method for identifying micro-fractures in a tight sandstone reservoir according to claim 1, wherein: In step S4, the ratio of the training dataset, the validation dataset, and the test dataset is 8:1:
1.
9. A system for identifying micro-fractures in tight sandstone reservoirs, characterized in that: The system is used to implement the method for identifying microfractures in a tight sandstone reservoir according to any one of claims 1 to 8, and at least includes: The data acquisition module is configured to acquire raw data of microfractures labeled with combinations of "microfracture genesis, development density, and dissolution products." Specifically, by using three evaluation parameters—genesis type, development characteristics, and dissolution product abundance—of microfractures in tight sandstone reservoirs, a quantitative classification standard is established for each parameter, thereby uniformly classifying the raw data labeled with combinations of "microfracture genesis, development density, and dissolution products." The microcracks have two types of genesis: tectonic and diagenetic; the microcracks have two types of developmental characteristics: opening, length, and surface density; and the microcracks are divided into low-density microcracks, medium-density microcracks, and high-density microcracks according to the surface density index. The characteristics of the dissolution products include product type, product content and development location. The abundance of the dissolution products was quantitatively analyzed using Image J image analysis software, and the dissolution products were divided into low dissolution product combination, medium dissolution product combination and high dissolution product combination according to the abundance index of the dissolution products. The anomaly rejection module is configured to process well logging curves by removing outliers, splicing curves, repositioning depths, and normalizing well logging curves. It selects well logging curves that are sensitive to the combination of microfracture genesis, development density, and dissolution products. It then obtains sensitive well logging curve data corresponding to different combinations of microfracture genesis, development density, and dissolution products, and establishes a well logging database for these different combinations. A data set construction module is configured to perform missing value processing and outlier detection processing on the well logging database to construct a data set, and divide the data set into a training data set, a validation data set and a test data set; The classification and discrimination module is configured to use a random forest classification algorithm selected by a Bayesian optimization algorithm as a discriminant model for the combination type of "microfracture genesis, development density, and dissolution products." Sample labels and well logging data for different combinations of "microfracture genesis, development density, and dissolution products" are input into the model. The well logging data is randomly divided into training, validation, and test datasets in a ratio of 8:1:
1. After model training, a model classification confusion matrix is output, and the classification accuracy displayed in the confusion matrix is used to obtain the classification results. The recognition module is configured to use the trained discriminant model to identify the combination type of "microcrack genesis, development density and dissolution products" of the microcracks to be tested.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and the program can be executed by one or more processors to implement the method for identifying microfractures in a tight sandstone reservoir according to any one of claims 1 to 8.
Citation Information
Patent Citations
Deep compact sandstone reservoir evaluation method based on artificial intelligence reservoir diagenesis phase recognition
CN113820754A
Sandstone corrosion type distinguishing method and device, electronic equipment and medium
CN114065828A