An intelligent interpretation method for reservoir diagenetic phase logging
Through the improved SMOTE algorithm and random forest classification model, the problem of insufficient and imbalance in reservoir diagenetic filament log recognition is solved, and efficient and accurate diagenetic facies recognition is achieved, providing technical support for the recognition of high-quality reservoirs of underground oil and gas reservoirs.
Patent Information
- Application Number
- CN202310109592.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-02-08
AI Technical Summary
The existing reservoir diagenetic filament identification methods have problems such as insufficient sample number, relying on manual experience, a lot of workload, multi-solvability, poor classification effect of nonlinear separable situations, and unbalanced samples, resulting in poor identification effect.
Using the improved SMOTE algorithm and random forest classification model, by constructing the Tyson polygon and Delaunay triangle net, new sample data are added to balance the number of samples of each diagenetic facies type, and the hyperparameters of the random forest model are optimized to improve classification accuracy.
It realizes efficient and accurate intelligent interpretation of reservoir diagenetic facies, can identify diagenetic facies types in the reservoir area, improves the accuracy and efficiency of identification, and is suitable for high-quality reservoir identification of underground oil and gas reservoirs.
Smart Images

Figure CN116304921B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of unconventional oil and gas reservoir development geological technology, and in particular to a reservoir diagenetic phase logging intelligent interpretation method. Background Art
[0002] Diagenetic phase is the material manifestation of diagenetic environment. It is the product of sediments undergoing certain diagenesis and evolution stages under the influence of diagenesis and tectonic factors in a specific sedimentary and physicochemical environment. It directly reflects the characteristics of the reservoir and highly summarizes the diagenesis experienced by the sediments. It also comprehensively considers the influence of diagenetic minerals and diagenetic evolution sequence on the reservoir properties. Therefore, the study of diagenetic phase can further determine the favorable diagenetic reservoirs directly related to reservoir properties, which is helpful for the regional evaluation and prediction of reservoirs. In the exploration and production stage of oil and gas reservoirs, with the continuous deepening of exploration and development, the reservoir depth is large, the physical properties change greatly, the microscopic pore throat structure is complex, and the difficulty of production is great. It is urgent to analyze the reservoir diagenesis and diagenetic phase.
[0003] In the research on diagenetic phase identification at home and abroad, most of them are limited to the manual identification of indoor core slices, which cannot realize the continuous evaluation of diagenetic phase in space and is difficult to be applied to underground oil and gas reservoirs. By constructing the response characteristics of logging curves and diagenesis and using the continuity of logging curves, this limitation can be effectively solved. The underground diagenetic phase logging interpretation technology currently published mostly uses logging curve characteristic method, intersection diagram method, spider diagram method, principal component analysis method, discriminant analysis method, Fisher recognition model, etc. In recent years, some scholars have also achieved good results by applying artificial intelligence methods to diagenetic phase logging identification, such as neural network, cluster analysis, integrated analysis and other methods. The logging curve characteristic method analyzes the response characteristics of different diagenetic phases on the logging curve, so as to realize the manual discrimination of the well section of unknown diagenetic phase. The intersection diagram method selects two different types of logging curve parameters to establish a coordinate system, maps the diagenetic phase classification data to the constructed coordinate system, and selects the boundaries of different diagenetic phases to realize diagenetic phase classification. The spider diagram method realizes diagenetic facies classification based on the distribution shape of data on multiple data axes, and the classification effect is slightly better than the cross plot. The principal component analysis method, discriminant analysis method, Fisher recognition model, neural network, cluster analysis, integrated analysis and other methods are essentially mathematical statistical methods, which can achieve the purpose of automatic identification of diagenetic facies by establishing different diagenetic facies discriminant equations. These methods generally have the following problems: First, in actual work, due to the limitation of operation cost, only a few wells are drilled for coring in the target oil and gas reservoir, and the coring section is often only tens of meters, which cannot achieve full coverage of the entire work area and the entire layer, resulting in the difficulty of applying the diagenetic phase based on the identification of indoor core thin sections to the actual underground reservoir; Second, the existing diagenetic phase logging identification method based on the analysis of logging curve response characteristics mainly relies on manual experience, with a large workload, and the analysis results vary from person to person, with a large number of solutions; Third, the intersection diagram method and the spider diagram method have some sample points that overlap and overlap with each other, making it difficult to accurately divide the boundaries manually; Fourth, the principal component analysis method, discriminant analysis method, Fisher identification model and other methods are mostly based on multivariate statistical theory, which are suitable for linearly separable situations, and have poor classification effects for nonlinearly separable situations; Artificial intelligence methods such as neural networks, cluster analysis, and integrated analysis have certain requirements for the number of samples. When the number of samples is too small, the algorithm is prone to overfitting, resulting in poor classification results. However, due to the high cost, long time consumption, and small amount of data of core sampling, diagenetic phase logging identification is a typical case of insufficient sample quantity. The use of oversampling techniques such as SMOTE and MAHAKIL can increase the number of sample points and effectively solve the problem of insufficient sample size. However, it cannot solve the data distribution problem of unbalanced data sets, thus affecting the interpretation of diagenetic facies.
[0004] In view of this, it is urgent to invent an efficient, accurate and practical intelligent interpretation method for reservoir diagenetic phase logging. Summary of the invention
[0005] In view of the above problems, the present invention aims to provide an intelligent interpretation method for reservoir diagenetic phase logging.
[0006] The technical solution of the present invention is as follows:
[0007] A reservoir diagenetic phase logging intelligent interpretation method comprises the following steps:
[0008] S1: collecting core thin section identification results and well logging data, determining the diagenetic facies type based on the core thin section identification results, using the data affecting the diagenetic facies type in the well logging data as the discriminant factor of the diagenetic facies type, and using the diagenetic facies type as the dependent variable;
[0009] S2: Construct Thiessen polygons;
[0010] S3: For diagenetic facies types with less sample data, the improved SMOTE algorithm is used to add new sample data to keep the sample data volume in each diagenetic facies type balanced;
[0011] S4: Based on the newly added data set, a random forest classification model is established;
[0012] S5: Optimizing the hyperparameters of the random forest classification model to obtain an optimal classification model;
[0013] S6: According to the optimal classification model and in combination with the well logging data of the target well, the diagenetic facies type of the target well is obtained.
[0014] Preferably, in step S1, the data influencing the diagenetic facies type in the well logging data include well logging curves and well logging depths.
[0015] Preferably, in step S1, the logging curve is any one or more of a gamma curve, a sonic time difference curve, a neutron curve, and a density curve.
[0016] Preferably, in step S1, the diagenetic phase types include strong dissolution phase, fracture phase, strong mud filling phase, compaction phase and strong calcareous cementation phase.
[0017] Preferably, step S2 specifically includes the following sub-steps: constructing a Delaunay triangulation based on the data of each diagenetic phase type; calculating and constructing the perpendicular bisectors of each side of the Delaunay triangulation, thereby constructing the Thiessen polygon.
[0018] Preferably, in step S3, adding new sample data using the improved SMOTE algorithm specifically includes the following sub-steps:
[0019] S31: Select the diagenetic facies type with the least number of samples as the positive sample, and follow x1, x2, x3, …, xn The sequence number of
[0020] S32: x i Sample number is taken as the target sample point, and the nearest sample point x of the same type as the target sample is found. 近邻 ;
[0021] S33: Determine the nearest sample point x 近邻 Whether the Thiessen polygon is an adjacent Thiessen polygon edge of the target sample, a new sample point x is generated according to the following calculation new :
[0022] x new =x i +rand(0,β)*(x 近邻 -x i ) (1)
[0023]
[0024] Where: d 近邻 Represents x i With x 近邻 The Euclidean distance of 邻接 Represents x i to x i With x 近邻 The connection line with x i The distance between the intersection points of the Thiessen polygons;
[0025] S34: repeating step S33, taking each sample of the diagenetic facies type with the least number of samples in step S31 as a target sample in turn, and generating new sample data;
[0026] S35: merging the new sample data with the initial sample to reconstruct the Thiessen polygon;
[0027] S36: repeating steps S31-S35 until the number of the diagenetic facies type with the least number of samples in step S31 is balanced with the number of the diagenetic facies type with the greatest number of samples;
[0028] S37: Repeat steps S31-S36 until the number of each diagenetic facies type is balanced.
[0029] Preferably, step S4 specifically includes the following sub-steps:
[0030] S41: Sort and divide the newly added data set, using 70% of the data as the training set and 30% of the data as the test set;
[0031] S42: using a random forest algorithm, taking the training set as an input sample set, randomly sampling multiple times with replacement to form a subset, wherein the number of samplings is consistent with the number of samples, and the subset data obtained by sampling is used to construct a decision tree;
[0032] S43: randomly extracting some discriminant factors from the subset data to form a candidate segmentation set, selecting a discriminant factor from the candidate segmentation set as a segmentation point of the decision tree according to the principle of minimum node impurity, and continuing to split using this principle until all samples of the node reach the leaf node, and then ending the splitting;
[0033] S44: Repeat steps S42-S43 to establish multiple decision trees to form the random forest classification model.
[0034] Preferably, step S5 specifically includes the following sub-steps:
[0035] S51: using a grid search method to test different hyperparameters of the random forest classification model to obtain a confusion matrix of true values and predicted values;
[0036] S52: Calculate the accuracy, recall and F1 score of the confusion matrix;
[0037] S53: According to the calculation result of step S52, the optimal hyperparameter is selected to obtain the optimal classification model.
[0038] The beneficial effects of the present invention are:
[0039] The present invention can perform intelligent logging interpretation on the diagenetic phases in the study area, and has the characteristics of high recognition of reservoir areas, accurate evaluation and short time consumption. It can achieve everything from indoor microscope microscopic identification to macroscopic wellbore of oil and gas reservoir drilling, thereby directly reflecting reservoir characteristics and providing effective technical support for the identification of high-quality reservoirs in underground oil and gas reservoirs. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0041] Figure 1 It is a schematic diagram of the process of the reservoir diagenetic phase logging intelligent interpretation method of the present invention;
[0042] Figure 2 A schematic diagram of a process for selecting intelligent interpretation variables for diagenetic facies identification well logging according to a specific embodiment;
[0043] Figure 3 A schematic diagram of Thiessen polygons constructed for an initial sample of a specific embodiment;
[0044] Figure 4 Schematic diagram of Thiessen polygons constructed for a fracture phase sample of a specific embodiment;
[0045] Figure 5 It is a schematic diagram of the diagenetic phase interpretation results of a specific embodiment. DETAILED DESCRIPTION
[0046] The present invention is further described below in conjunction with the accompanying drawings and embodiments. It should be noted that, in the absence of conflict, the embodiments in this application and the technical features in the embodiments can be combined with each other. It should be noted that, unless otherwise specified, all technical and scientific terms used in this application have the same meanings as those generally understood by those of ordinary skill in the art to which this application belongs. The words "including" or "comprising" and the like used in the disclosure of the present invention mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0047] like Figure 1 As shown, the present invention provides a reservoir diagenetic phase logging intelligent interpretation method, comprising the following steps:
[0048] S1: Collect core thin section identification results and logging data, determine the diagenetic facies type based on the core thin section identification results, use the data affecting the diagenetic facies type in the logging data as the discriminant factor of the diagenetic facies type, and use the diagenetic facies type as the dependent variable.
[0049] In a specific embodiment, the data affecting the diagenetic facies type in the well logging data include well logging curves and well logging depths, that is, the well logging curves and well logging depths are used as discriminant factors for the diagenetic facies type, and the well logging curves are any one or more of gamma curves, acoustic time difference curves, neutron curves, and density curves. The diagenetic facies types include strong dissolution phase, fracture phase, strong mud filling phase, compaction phase, and strong calcareous cementation phase.
[0050] It should be noted that the data affecting the diagenetic facies type may be different in different research blocks. The data selected in the above embodiment are only preferred discriminant factors of the present invention. When using the present invention, other influencing data can be selected as discriminant factors of the diagenetic facies type according to the target research block.
[0051] S2: Construct Thiessen polygons; the Thiessen polygons are constructed by the following sub-steps: construct a Delaunay triangulation based on the data of each diagenetic phase type; calculate and construct the perpendicular bisectors of each side of the Delaunay triangulation, thereby constructing the Thiessen polygons.
[0052] S3: For diagenetic facies types with less sample data, the improved SMOTE algorithm is used to add new sample data to keep the sample data volume in each diagenetic facies type balanced.
[0053] It should be noted that in this step, the balance means that the relative error of the sample data in the two types of lithofacies is kept within the balance threshold, which is a manually set value. The larger the balance threshold, the smaller the difference in the sample data of the two types of lithofacies, and the more accurate the result. The balance threshold can be 20%, 15%, 10%, 5%, etc. In a specific embodiment, for example, the number of samples of a certain type of lithofacies is 100 data, and the number of samples of another type of lithofacies is 5 data. At this time, the difference in the data of the two types of lithofacies is too large to provide sufficient data support for subsequent classification. Therefore, it is necessary to use the improved SMOTE algorithm of the present invention to add sample data to the lithofacies type with 5 sample data. When this type of lithofacies type is added to 85 data, the relative error between the two is 15%. It can be considered that the sample data of the two types of lithofacies have reached a balance at this time.
[0054] In a specific embodiment, adding new sample data using the improved SMOTE algorithm specifically includes the following sub-steps:
[0055] S31: Select the diagenetic facies type with the least number of samples as the positive sample, and follow x1, x2, x3, …, x n The sequence number of
[0056] S32: x i Sample number is taken as the target sample point, and the nearest sample point x of the same type as the target sample is found. 近邻 ;
[0057] S33: Determine the nearest sample point x 近邻 Whether the Thiessen polygon is an adjacent Thiessen polygon edge of the target sample, a new sample point x is generated according to the following calculation new :
[0058] x new =x i +rand(0,β)*(x 近邻 -x i ) (1)
[0059]
[0060] Where: d 近邻 Represents x i With x 近邻 The Euclidean distance of 邻接 Represents x i to x iWith x 近邻 The connection line with x i The distance between the intersection points of the Thiessen polygons;
[0061] S34: repeating step S33, taking each sample of the diagenetic facies type with the least number of samples in step S31 as a target sample in turn, and generating new sample data;
[0062] S35: merging the new sample data with the initial sample to reconstruct the Thiessen polygon;
[0063] S36: repeating steps S31-S35 until the number of the diagenetic facies type with the least number of samples in step S31 is balanced with the number of the diagenetic facies type with the greatest number of samples;
[0064] S37: Repeat steps S31-S36 until the number of each diagenetic facies type is balanced.
[0065] The present invention adopts the improved SMOTE algorithm to add small sample data, which can solve the technical problem that the existing research is limited by the imbalance of diagenetic phase samples and improve the classification accuracy of the machine learning algorithm.
[0066] Compared with the classic SMOTE algorithm, the improved SMOTE algorithm described in the present invention merges the newly generated sample points with the initial samples and uses them together as the reference samples added next time, rather than always using the initial sample data as the input of the algorithm. This can effectively solve the problem that the sample points generated by the classic SMOTE algorithm are located on a straight line, and effectively improve the algorithm accuracy.
[0067] At the same time, given that the classic SMOTE algorithm does not consider the problem of the appropriate range of new samples, the improved SMOTE algorithm described in the present invention combines the characteristics of Thiessen polygons in the process of adding new data, and forms new sample points in the Thiessen polygons corresponding to each sample, which can ensure that the newly generated sample point is closest to one of the parent sample points, effectively avoiding the newly generated sample point from falling into the range of other types of samples, so that the newly generated sample data can better maintain the distribution pattern of the initial data, avoiding the problem that the data generated by the classic SMOTE algorithm is on the line connecting the adjacent points, thereby further improving the accuracy of the algorithm.
[0068] In summary, the improved SMOTE algorithm described in the present invention can effectively solve the generalization and marginal problems of data generated by the classic SMOTE algorithm, and provide technical support for improving the classification accuracy of sample types and making the machine learning classification model more accurate.
[0069] S4: Based on the newly added data set, a random forest classification model is established; specifically, the following sub-steps are included:
[0070] S41: Sort and divide the newly added data set, using 70% of the data as the training set and 30% of the data as the test set;
[0071] S42: using a random forest algorithm, taking the training set as an input sample set, randomly sampling multiple times with replacement to form a subset, wherein the number of samplings is consistent with the number of samples, and the subset data obtained by sampling is used to construct a decision tree;
[0072] S43: randomly extracting some discriminant factors from the subset data to form a candidate segmentation set, selecting a discriminant factor from the candidate segmentation set as a segmentation point of the decision tree according to the principle of minimum node impurity, and continuing to split using this principle until all samples of the node reach the leaf node, and then ending the splitting;
[0073] S44: Repeat steps S42-S43 to establish multiple decision trees to form the random forest classification model.
[0074] S5: Optimizing the hyperparameters of the random forest classification model to obtain an optimal classification model; specifically comprising the following sub-steps:
[0075] S51: using a grid search method to test different hyperparameters of the random forest classification model to obtain a confusion matrix of true values and predicted values;
[0076] S52: Calculate the accuracy, recall and F1 score of the confusion matrix;
[0077] S53: According to the calculation result of step S52, the optimal hyperparameter is selected to obtain the optimal classification model.
[0078] The present invention can improve the accuracy of the random forest classification model by optimizing hyperparameters, thereby ensuring the accuracy and reliability of the intelligent interpretation results of the model diagenetic phase logging.
[0079] S6: According to the optimal classification model and in combination with the well logging data of the target well, the diagenetic facies type of the target well is obtained.
[0080] In a specific embodiment, the Huangyan structural area of the Xihu Sag is used as the research area, and the reservoir diagenetic phase logging intelligent interpretation method of the present invention is used to study it and obtain its diagenetic phase type. The strata in the Huangyan structural area of the Xihu Sag are relatively flat, and the main lithologies are sub-lithic sandstone, feldspar lithic sandstone, lithic feldspar sandstone and feldspar sandstone, which are typical low-porosity and low-permeability dense sandstone reservoirs. The entire experimental process includes the following steps:
[0081] (1) Collect core and logging data
[0082] like Figure 2As shown, logging data are collected, including natural gamma ray curve (GR), density curve (ZDEN), acoustic time difference curve (AC), neutron curve (CN), deep resistivity (RD), natural potential curve (SP), logging depth and diagenetic phase type, etc.
[0083] The collected data were analyzed, and four curves, GR, ZDEN, CN, and AC, and logging depth were selected as the discriminant factors for studying the classification of diagenetic phases, and the diagenetic phases were classified into five types: strong dissolution phase, fracture phase, compaction phase, strong mud filling phase, and strong calcareous cementation phase. There are 131 data points of strong dissolution phase, 5 data points of fracture phase, 43 data points of compaction phase, 19 data points of strong mud filling phase, and 21 data points of strong calcareous cementation phase, totaling 219 sample data points.
[0084] (2) Based on the initial sample data obtained in step (1), construct Figure 3 Thiessen polygons shown
[0085] (3) Using the improved SMOTE algorithm to add new sample data
[0086] 1) The data volume of the five types of diagenetic facies is sorted from small to large, which are fracture facies, strong mud filling facies, strong calcareous cementation facies, compaction facies and strong dissolution facies;
[0087] 2) Among the five types, the fracture phase has the least data, only 5, which are taken as positive samples and numbered in the order of x1, x2, x3, x4, and x5; their distribution is as follows: Figure 4 As shown;
[0088] 3) Taking fracture phase sample x1 as an example, sample x1 is taken as the target sample, and the nearest neighbor among the same type of samples is found. The nearest neighbor found is fracture phase x2; according to Figure 4 , determine that the fracture phase x2 is the adjacent Thiessen polygon of sample x1, take β = 0.5, and use formula (1) to generate new sample points;
[0089] 4) Take x2, x3, x4, and x5 as target samples respectively, and use the method in step 3) to generate new sample points;
[0090] 5) Adding the newly added samples to the set L, and merging the set L with the initial samples to form a new set, and reconstructing the Thiessen polygon;
[0091] 6) Repeat steps 2) to 5) until the number of sample data of the fracture phase is relatively balanced with the number of samples of the strong dissolution phase with the largest number of samples, that is, the number of data is increased from the original 5 to 112;
[0092] 7) For the strong mud filling phase, strong calcareous cementation phase and compaction phase, the same method as the fracture phase is used to add data.
[0093] (4) Based on the newly added data set, a random forest classification model is established
[0094] The new data set is sorted and divided, and 70% of the data is used as the training set and 30% of the data is used as the test set; the random forest algorithm is used to take the training set as the input sample set, and a subset is randomly selected multiple times with replacement, and the sampling times are consistent with the number of samples. The subset data obtained by sampling is used to build a decision tree; some discriminant factors are randomly selected from the subset data to form a candidate segmentation set, and a discriminant factor is selected from the candidate segmentation set as the segmentation point of the decision tree according to the principle of minimum node impurity, that is, the principle of minimum Gini coefficient, and the split is continued according to this principle until all samples of the node reach the leaf node, and then the split is ended. In the process of forming the decision tree, each node is split in this way. For example, in any subset data, GR, AC, CN and other discriminant factors are randomly selected, the Gini coefficient is calculated, and GR is selected as the segmentation point until all samples of the node reach the leaf node, and then the split is ended; repeat the above steps to establish multiple decision trees to form a random forest classification model.
[0095] (5) Optimize the hyperparameters of the random forest classification model to obtain the optimal classification model
[0096] By using the grid search method to test different hyperparameters of the random forest classification model, a confusion matrix (Confusion Matrix) about the true value and the predicted value can be obtained, as shown in Table 1:
[0097] Table 1 Confusion matrix
[0098] Confusion Matrix Predicted as positive class Predicted as negative class Positive TP FN Negative FP TN
[0099] The accuracy, recall and F1 Score of the confusion matrix are calculated to select the optimal hyperparameters; the accuracy is the proportion of all correct results of the classification model to the total observations, which is calculated by the following formula:
[0100]
[0101] The recall rate is the proportion of the model predictions among all the results that the model predicts are positive, which is calculated by the following formula:
[0102]
[0103] The F1 Score is an evaluation index that comprehensively evaluates precision and recall, and is calculated using the following formula:
[0104]
[0105] When the number of random forest trees ntree is 350, the accuracy of model prediction reaches 91.06%, which is relatively high, so ntree=350 is set.
[0106] The results newly obtained by using the improved SMOTE algorithm of the present invention are compared with the results newly obtained by using the classic SMOTE algorithm. The results are shown in Table 2:
[0107] Table 2 Comparison of classification results
[0108] New data Accuracy Recall F1SCORE Raw data 0.7727 0.7209 0.6365 Classic SMOTE algorithm 0.8527 0.8497 0.8493 The present invention (Thiessen polygon + improved SMOTE algorithm) 0.9106 0.9032 0.9036
[0109] It can be seen from Table 2 that after the improved SMOTE algorithm is used to add new sample data, the accuracy, recall rate and F1 Score of the present invention are respectively improved by 17.85%, 25.29% and 41.96%. Compared with the classic SMOTE algorithm, the accuracy, recall rate and F1 Score of the present invention are improved, and the generalization and marginal problems of the data generated by the classic SMOTE algorithm can be avoided.
[0110] (6) Model verification
[0111] In order to further verify the reliability of the model, taking the H9 section of Well X as an example, the logging parameters were input into the optimal classification model obtained in step (5). The diagenetic phase logging interpretation results are shown in the figure. Figure 5 As shown. Figure 5 It can be seen that the present invention can accurately predict the drilling diagenetic phase type.
[0112] In summary, the present invention can evaluate the reservoir area by combining relevant geological data, obtain intelligent logging interpretation of diagenetic phases, and provide technical support for the selection of "geological sweet spots" such as tight oil and gas reservoirs and carbonate oil and gas reservoirs. Compared with the prior art, the present invention has significant progress.
[0113] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A reservoir diagenetic phase logging intelligent interpretation method, characterized in that: The following steps are involved: S1: collecting core thin section identification results and well logging data, determining the diagenetic facies type based on the core thin section identification results, using the data affecting the diagenetic facies type in the well logging data as the discriminant factor of the diagenetic facies type, and using the diagenetic facies type as the dependent variable; S2: Construct Thiessen polygons; S3: For diagenetic facies types with less sample data, the improved SMOTE algorithm is used to add new sample data to keep the sample data volume in each diagenetic facies type balanced; Using the improved SMOTE algorithm to add new sample data specifically includes the following sub-steps: S31: Select the diagenetic facies type with the least number of samples as the positive sample, and follow x1, x2, x3, …, x n The sequence number of S32: x i Sample number is taken as the target sample point, and the nearest sample point x of the same type as the target sample is found. 近邻 ; S33: Determine the nearest sample point x 近邻 Whether the Thiessen polygon is an adjacent Thiessen polygon edge of the target sample, a new sample point x is generated according to the following calculation new : x new =x i +rand(0,β)*(x 近邻 -x i ) (1) Where: d 近邻 Represents x i With x 近邻 The Euclidean distance of 邻接 Represents x i to x i With x 近邻 The connection line with x i The distance between the intersection points of the Thiessen polygons; S34: repeating step S33, taking each sample of the diagenetic facies type with the least number of samples in step S31 as a target sample in turn, and generating new sample data; S35: merging the new sample data with the initial sample to reconstruct the Thiessen polygon; S36: repeating steps S31-S35 until the number of the diagenetic facies type with the least number of samples in step S31 is balanced with the number of the diagenetic facies type with the greatest number of samples; S37: repeating steps S31-S36 until the number of each diagenetic facies type is balanced; S4: Based on the newly added data set, a random forest classification model is established; S5: Optimizing the hyperparameters of the random forest classification model to obtain an optimal classification model; S6: According to the optimal classification model and in combination with the well logging data of the target well, the diagenetic facies type of the target well is obtained.
2. The reservoir diagenetic phase logging intelligent interpretation method according to claim 1, characterized in that: In step S1, the data affecting the diagenetic facies type in the well logging data include well logging curves and well logging depths.
3. The reservoir diagenetic facies logging intelligent interpretation method according to claim 2, characterized in that: In step S1, the logging curve is any one or more of a gamma curve, a sonic time difference curve, a neutron curve, and a density curve.
4. The reservoir diagenetic phase logging intelligent interpretation method according to claim 1, characterized in that: In step S1, the diagenetic phase types include strong dissolution phase, fracture phase, strong mud filling phase, compaction phase and strong calcareous cementation phase.
5. The reservoir diagenetic phase logging intelligent interpretation method according to claim 1, characterized in that: Step S2 specifically includes the following sub-steps: constructing a Delaunay triangulation based on the data of each diagenetic phase type; calculating and constructing the perpendicular bisectors of each side of the Delaunay triangulation, thereby constructing the Thiessen polygon.
6. The reservoir diagenetic facies logging intelligent interpretation method according to any one of claims 1 to 5, characterized in that: Step S4 specifically includes the following sub-steps: S41: Sort and divide the newly added data set, using 70% of the data as the training set and 30% of the data as the test set; S42: using a random forest algorithm, taking the training set as an input sample set, randomly sampling multiple times with replacement to form a subset, wherein the number of samplings is consistent with the number of samples, and the subset data obtained by sampling is used to construct a decision tree; S43: randomly extracting some discriminant factors from the subset data to form a candidate segmentation set, selecting a discriminant factor from the candidate segmentation set as a segmentation point of the decision tree according to the principle of minimum node impurity, and continuing to split using this principle until all samples of the node reach the leaf node, and then ending the splitting; S44: Repeat steps S42-S43 to establish multiple decision trees to form the random forest classification model.
7. The reservoir diagenetic facies logging intelligent interpretation method according to any one of claims 1 to 5, characterized in that: Step S5 specifically includes the following sub-steps: S51: using a grid search method to test different hyperparameters of the random forest classification model to obtain a confusion matrix of true values and predicted values; S52: Calculate the accuracy, recall and F1 score of the confusion matrix; S53: According to the calculation result of step S52, the optimal hyperparameters are selected to obtain the optimal classification model.
Citation Information
Patent Citations
Quantitative evaluation method for main control factors of quality of tight sandstone gas reservoir based on random forest
CN113344359A
Intelligent reservoir type division method based on unbalanced samples
CN115422988A