A shale reservoir sweet spot identification method, system and computer equipment
The shale reservoir dessert discrimination model is constructed through the LightGBM integrated learning algorithm, and the logging data is used for discrimination, which solves the problems of large errors and cumbersome steps in the existing technology, and achieves efficient and accurate discrimination of shale reservoir desserts.
Patent Information
- Application Number
- CN202410738728.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-06-07
AI Technical Summary
In the prior art, the shale reservoir dessert logging method has a large error, resulting in low discrimination accuracy and a logging prediction model with multiple parameters is required. The steps are cumbersome, time-consuming and low efficiency.
Directly use logging data, and use LightGBM integrated learning algorithm to construct a shale reservoir dessert discrimination model, select the optimal logging curve as the sample feature, and conduct shale reservoir dessert discrimination, simplify it to directly use logging data for discrimination, reducing complexity and improving accuracy.
It improves the accuracy and efficiency of dessert discrimination in shale reservoirs, simplifies the discrimination process, reduces errors, and achieves efficient dessert discrimination.
Smart Images

Figure CN118655627B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of shale oil and gas exploration, and in particular to a method, system and computer equipment for identifying sweet spots in shale reservoirs. Background Art
[0002] In recent years, shale oil and gas have become a crucial component of my country's energy security. The key to shale oil and gas exploration and development lies in identifying shale reservoir sweet spots. Currently, the main methods for identifying shale reservoir sweet spots include core observation and description, experimental testing and analysis, well logging, and seismic analysis. Shale reservoir sweet spot identification based on core observation and description relies heavily on the experience of geologists, resulting in high subjectivity, low efficiency, and unstable accuracy. Shale reservoir sweet spot evaluation based on experimental testing and analysis requires a large amount of analytical data, resulting in high costs, a long time consumption, and poor vertical continuity. Furthermore, both of these methods require coring of the target interval, making sweet spot identification difficult to generalize to uncored intervals or wells, limiting their scope of application. While seismic data-based sweet spot identification offers good lateral continuity, it suffers from low vertical accuracy, making it difficult to meet the requirements for refined shale reservoir sweet spot identification. Shale reservoir sweet spot identification by logging refers to the identification of shale reservoir sweet spots using logging data. It has the advantages of high vertical accuracy, low cost, and strong scalability. It is currently the main method for identifying shale reservoir sweet spots.
[0003] Chinese patent CN105986816A discloses a method for identifying sweet spots in shale formations using well logging data. The main idea is to use well logging data and a variety of methods to establish prediction models for shale formation geological sweet spot parameters (kerogen volume content, gas porosity, gas saturation, and total organic matter content) and engineering sweet spot parameters (maximum horizontal effective stress value, pore structure index, and brittleness index). Then, based on the ratio of the radar map area of each sweet spot parameter to the baseline radar map area, the geological sweet spot and engineering sweet spot coefficient of the shale formation are obtained respectively. The geological and engineering sweet spot coefficient thresholds are further combined to perform logging identification of the shale formation sweet spot. The journal article "A Logging Evaluation Method for Deep Shale Gas Geological-Engineering Sweet Spot Parameters Based on Rock Physical Phases - A Case Study of the Wufeng-Longmaxi Formations in the LZ Block of the Sichuan Basin" proposes a logging method for identifying sweet spots in shale gas reservoirs. The main idea is to first divide the shale gas reservoir into different rock physical phases. Then, using a random forest regression algorithm, a logging prediction model for multiple geological-engineering sweet spot parameters (TOC, porosity, mineral content such as quartz, feldspar, calcite, dolomite, clay, brittleness index, and free gas content, adsorbed gas content, and total gas content) is established for each rock physical phase. Then, based on the numerical values of the shale reservoir sweet spot parameters, logging is used to identify the sweet spot intervals in the shale gas reservoir. In summary, the main approach to identifying sweet spots in shale reservoirs by logging is to first establish logging prediction models for multiple sweet spot parameters using logging data, and then identify the sweet spot in the shale reservoir by logging based on the shale reservoir sweet spot evaluation criteria.
[0004] However, on the one hand, this approach requires the establishment of a logging evaluation model with multiple parameters, which is cumbersome, time-consuming, and inefficient. On the other hand, the multiple sweet spot parameter logging prediction models established all have certain errors. Combining multiple parameters to identify shale reservoir sweet spots will inevitably accumulate errors, resulting in a reduction in the accuracy of shale reservoir sweet spot identification. Summary of the Invention
[0005] In view of the shortcomings of the existing technology in which the shale reservoir sweet spot logging identification method has large errors, resulting in low accuracy in shale reservoir sweet spot identification, the present invention proposes a shale reservoir sweet spot identification method, system and computer equipment, which directly use logging data to establish a shale reservoir sweet spot identification model to identify shale reservoir sweet spots, without the need to pre-establish logging prediction models of multiple sweet spot parameters to identify shale reservoir sweet spots based on the sweet spot parameter values; thereby solving the problems existing in the existing technology.
[0006] A method for identifying sweet spots in shale reservoirs comprises the following steps:
[0007] Collecting target area shale reservoir core sample data, and performing sweet spot identification on the shale reservoir core sample data to obtain the sweet spot identification type of the shale reservoir core sample;
[0008] Based on the logging data of the shale reservoir in the target area, the best logging curve in the logging data is selected as the sample feature;
[0009] A shale reservoir sweet spot discrimination model was constructed based on the LightGBM ensemble learning algorithm. The specific steps include: initializing a decision tree based on core sample data, building a new decision tree based on sample characteristics, using the sweet spot discrimination type of the shale reservoir core sample as the sample data label, and constructing a shale reservoir sweet spot discrimination model after multiple rounds of update iterations;
[0010] The optimal logging curve of the well to be identified in the target shale reservoir is selected and input into the shale reservoir sweet spot identification model to obtain the identification type corresponding to the shale reservoir sweet spot.
[0011] Furthermore, after collecting the target area shale reservoir core sample data, the discrimination parameters of the core sample sweet spot are obtained. The shale reservoir sweet spot discrimination parameters include total organic carbon content, porosity, total gas content and brittleness index. The acquisition process includes the following steps:
[0012] The total organic carbon content is determined by the combustion method after the core samples are crushed, acid-washed, neutralized and dried;
[0013] After drying the core plug samples, the porosity of the samples was determined using Boyle's law.
[0014] The desorption method was used to measure the desorbed gas content of the bulk sample and the residual gas content of the corresponding powder sample. The USBM direct method was used to calculate the lost gas content. The total gas content of the core sample was obtained by calculating the sum of the desorbed gas content, the residual gas content and the lost gas content.
[0015] After the core samples were crushed, the contents of quartz, feldspar, clay minerals, carbonate minerals and pyrite in the samples were determined using X-ray diffraction technology, and the sample brittleness index was calculated using the following formula:
[0016]
[0017] Among them, BI represents the brittleness index, W Q 、W F 、W CM 、W C and W P They represent the contents of quartz, feldspar, clay minerals, carbonate minerals and pyrite respectively.
[0018] Furthermore, the method also includes preprocessing the sample data, wherein the preprocessing process includes optimizing the well logging curve and normalizing the optimized well logging curve data.
[0019] Furthermore, the selecting of the optimal logging curve from the logging data specifically includes the following steps:
[0020] Obtain shale reservoir logging data for the target layer; specifically, 11 logging curves including GR, KTh, U, K, Th, AC, U / Th, DEN, CNL, LLD and PE;
[0021] By using the recursive feature elimination algorithm to optimize logging parameters, the variation of shale reservoir sweet spot identification accuracy with the number of logging curves was obtained.
[0022] The optimal logging curve is selected according to the change of the shale reservoir sweet spot identification accuracy with the number of logging curves.
[0023] Furthermore, the optimal logging curve is normalized, and its expression is:
[0024]
[0025] Where: X* represents the normalized logging curve data, X represents the original logging curve data, μ represents the mean of the original logging data, and σ represents the standard deviation of the original logging data.
[0026] Furthermore, it also includes tuning the hyperparameters of the shale reservoir sweet spot discrimination model by calling the GridSearchCV module in the Scikit-Learn library; it uses the GridSearchCV module to search through the entire specified hyperparameter value space, exhaustively enumerates the optimal hyperparameter combination, and uses 5-fold cross-validation to evaluate the performance of each hyperparameter combination to ensure that the global optimal solution is found; wherein the hyperparameters include the number of iterations, the learning rate, the maximum tree depth and the number of leaf nodes.
[0027] Furthermore, the shale reservoir sweet spot discrimination model is constructed based on the LightGBM ensemble learning algorithm, including the following steps:
[0028] Load the sample data into memory, initialize the decision tree f0(x), and select the softmax cross entropy function as the error function l:
[0029]
[0030] Among them, M represents the number of categories, y c Represents an indicator variable. When the sample prediction category is consistent with the true category, the value is 1, otherwise it is 0. c Indicates the predicted probability that the sample prediction category belongs to category c;
[0031] For the tth (t=1, 2, ..., T) decision tree: discretize the continuous eigenvalues into k buckets through the piecewise function, traverse all the data, calculate the sum of the first-order gradient and the sum of the second-order gradient of the samples in each bucket, and calculate the sum of the first-order gradient of all buckets G t , the sum of the second-order gradient H t :
[0032]
[0033] Among them, g ti represents the sum of the first-order gradients of the sample data in the i-th bucket, h ti Represents the sum of the second-order gradients of the sample data in the i-th bucket;
[0034] The decision tree is split based on the current node. The sum of the first-order gradients of all buckets is G, the sum of the second-order gradients of all buckets is H, and the initial score is set to 0. For features n = GR, U, U / TH, AC, DEN, CNL, LLD: the left child node is represented by G tl 、H tl , the right child node is represented by G tr 、H tr ;
[0035] Arrange the buckets of sample feature n from small to large, take out the i-th bucket in turn, put the bucket into the left child node, and calculate the sum of the first-order gradients G of the left child node at this time. tl * , the sum of the second-order gradient H tl * And the sum of the first-order gradients of the right child nodes G tr * , the sum of the second-order gradient H tr * :
[0036] G tl * =G tl +g ti , G tr * =GG tl *
[0037] H tl * =H tl +h ti , H tr * =HH tl *
[0038] Update score:
[0039]
[0040] Among them, γ represents the complexity penalty coefficient of the new leaf node, and λ represents the complexity penalty coefficient of the decision tree;
[0041] Select the feature and bucket position corresponding to the maximum score to split the node;
[0042] When the number of leaf nodes or the maximum tree depth reaches the preset value, the current decision tree is established and the strong learner F is updated. t (x):
[0043] F t (x) = F t-1 (x)+f t (x)
[0044] Among them F t-1 (x) represents the strong learner obtained from the t-1th decision tree, f t (x) is the decision tree generated in this round of iteration;
[0045] After T rounds of iteration, the final strong learner F is obtained T (x):
[0046]
[0047] Where I represents the learning rate.
[0048] Furthermore, the method further includes evaluating the shale reservoir sweet spot discrimination model through evaluation indicators; the evaluation indicators include weighted accuracy W A , weighted precision W P , weighted recall W R and kappa coefficient; they are respectively expressed as:
[0049]
[0050]
[0051] Among them, TP j represents the number of true positive examples of dessert category j, FP j represents the number of false positives for dessert category j, FN j represents the number of false negatives for dessert category j, N j and N represent the number of samples and the total number of samples of dessert category j, respectively. o and P e represent the actual agreement rate and the theoretical agreement rate, respectively.
[0052] The present invention also includes a shale reservoir sweet spot identification system, comprising:
[0053] An acquisition module is used to collect target area shale reservoir core sample data, and perform sweet spot identification on the shale reservoir core sample data to obtain the sweet spot identification type of the shale reservoir core sample;
[0054] The optimization module is used to select the best logging curve in the logging data as the sample feature based on the logging data of the shale reservoir in the target area;
[0055] The model building module is used to build a shale reservoir sweet spot discrimination model based on the LightGBM ensemble learning algorithm. The specific steps include: initializing a decision tree based on sample data, building a new decision tree based on sample characteristics, and building a shale reservoir sweet spot discrimination model after multiple rounds of update iterations;
[0056] The discrimination module is used to select the optimal logging curve of the well to be identified in the target area shale reservoir and input it into the shale reservoir sweet spot discrimination model to obtain the discrimination type corresponding to the shale reservoir sweet spot.
[0057] The present invention also includes a computer device for identifying shale reservoir sweet spots, comprising: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the shale reservoir sweet spot identification method are implemented.
[0058] The present invention provides a shale reservoir sweet spot identification method, system and computer equipment, which have the following features:
[0059] Beneficial effects:
[0060] The present invention trains the LightGBM ensemble learning algorithm by taking the sweet spot discrimination type of shale reservoir core samples as the sample data label and selecting the optimal logging curve as the sample data feature, constructs a shale reservoir sweet spot discrimination model, and directly uses logging data to discriminate shale reservoir sweet spots without pre-establishing logging prediction models of multiple sweet spot parameters, thereby reducing the complexity of shale reservoir sweet spot discrimination. By inputting the optimal logging curve of the well to be identified in the target area shale reservoir into the shale reservoir sweet spot discrimination model, the discrimination type corresponding to the shale reservoir sweet spot is obtained, thereby improving the accuracy of shale reservoir sweet spot discrimination. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a flow chart of shale reservoir sweet spot identification by logging in an embodiment of the present invention;
[0062] Figure 2 2. This is a graph showing how the shale reservoir sweet spot identification accuracy changes with the number of logging curves in an embodiment of the present invention;
[0063] Figure 3 This is a graph showing how the model training and test losses change with the number of iterations in an embodiment of the present invention;
[0064] Figure 4 Schematic diagram of the results of distinguishing sweet spots of different types of shale reservoirs in a shale gas well in the southern Sichuan Basin in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0066] The present invention proposes a method for identifying sweet spots in shale reservoirs, comprising the following steps:
[0067] Step 1: Obtain sample data.
[0068] 1.1 Obtaining parameters for identifying sweet spots in shale reservoirs: Core samples of shale reservoirs in the target area are collected and experimentally tested and analyzed to obtain parameters for identifying sweet spots in shale reservoirs: total organic carbon content, porosity, total gas content, and brittleness index. Specific methods: 1) After the core samples are crushed, pickled, neutralized, and dried, the total organic carbon content is determined using the combustion method; 2) After the core plug samples are dried, the porosity of the samples is determined using Boyle's law; 3) The desorption gas content of the block samples and the residual gas content of the corresponding powder samples are measured using the desorption method, and the lost gas content is calculated using the USBM direct method. The sum of the three is the total gas content of the core samples; 4) After the core samples are crushed, the content of quartz, feldspar, clay minerals, carbonate minerals, and pyrite in the samples is determined using X-ray diffraction technology. Then, the sample brittleness index is calculated using the following formula.
[0069]
[0070] Among them, BI represents the brittleness index, W Q 、W F 、W CM 、W C and W P They represent the contents of quartz, feldspar, clay minerals, carbonate minerals and pyrite respectively.
[0071] 1.2 Shale Reservoir Sweet Spot Identification: Shale reservoirs are classified based on single-parameter sweet spot parameter thresholds in Table 1. Then, the weights for each sweet spot parameter in Table 2 are used to comprehensively identify shale reservoir sweet spots. Shale reservoirs are divided into three sweet spot categories: Class I, Class II, and Class III. Class I sweet spots represent the primary shale gas extraction layer under current technology, while Class II and Class III sweet spots represent potential shale gas extraction layers and may become the primary shale gas extraction layer upon further technological breakthroughs. The shale reservoir sweet spot identification results serve as the sample data label.
[0072] Table 1 Classification and standards of shale reservoirs based on single parameters
[0073]
[0074] Table 2 Comprehensive criteria for identifying sweet spots in shale reservoirs
[0075]
[0076] 1.3 Obtaining shale reservoir logging curves: Obtain shale reservoir logging data of the target area's target layer, including 11 logging curves: GR (natural gamma ray), KTh (uranium-free gamma ray), U (uranium), K (potassium), Th (thorium), AC (acoustic time difference), U / Th (uranium-thorium ratio), DEN (compensated density), CNL (compensated neutron), LLD (deep lateral resistivity) and PE (photoelectric absorption cross-section index) as sample data features.
[0077] 1.4 Logging curve standardization: To eliminate the possible systematic errors between logging curves caused by construction batches and other reasons, the Pagoda Formation in the target area, where nodular limestone is stably developed, is used as the standard layer. The logging curves are standardized to help improve modeling accuracy and stability.
[0078] Step 2: Data preprocessing.
[0079] 2.1 Logging Curve Optimization: To eliminate feature redundancy, simplify the model, and improve model training efficiency and accuracy, a recursive feature elimination algorithm was used to optimize logging parameters. Specifically, the LightGBM ensemble learning algorithm was used to train a model, evaluate the importance of each logging curve in the dataset, and then eliminate the logging curve with the least contribution to the model. The model was then retrained based on the remaining logging curves, and so on, to obtain the variation of shale reservoir sweet spot identification accuracy with the number of logging curves. Figure 2 As can be seen, the model achieved the highest discrimination accuracy of 0.8052 when selecting seven well logging curves. As the number of well logging curves continued to increase, the model's discrimination accuracy stopped improving and even began to decline. Based on the importance ranking (see Table 3) (the lower the importance ranking, the greater the contribution of the well logging curve), the top seven well logging curves were selected to obtain the optimal well logging curve subset: GR, U, U / Th, AC, DEN, CNL, and LLD (ranked in no particular order of importance). In summary, the sample data characteristics and label information are shown in Table 4.
[0080] Table 3 Importance ranking of logging parameters
[0081] Logging parameters GR KTh K Th U U / Th AC DEN CNL LLD PE Importance ranking 1 4 5 3 1 1 1 1 1 1 2
[0082] Table 4 Sample data characteristics and label information
[0083]
[0084] 2.2 Data normalization: To eliminate the scale and dimension differences between different logging curve data and accelerate model convergence, all logging curve data are normalized in advance using the following formula. The normalized logging curve data conforms to the standard normal distribution with a mean of 0 and a standard deviation of 1.
[0085]
[0086] Where: X* represents the normalized logging curve data, X represents the original logging curve data, μ represents the mean of the original logging data, and σ represents the standard deviation of the original logging data.
[0087] 2.3 Data division into training set and test set: After data normalization is completed, the sample data is divided into training set and test set in a ratio of 7:3. The training set data is used to build the model, and the test set data is used to evaluate the model performance.
[0088] Step 3: LightGBM model training. LightGBM is an ensemble learning algorithm based on the gradient boosting tree-based framework. It adopts the idea of a forward step-by-step algorithm to continuously generate new decision trees. The final output of the model is the accumulation of the results of each decision tree. In the actual modeling process, each decision tree is built based on the difference between the previous decision tree and the target value, thereby continuously reducing the error between the predicted value and the true value. Because LightGBM adopts algorithmic strategies such as histogram algorithm, leaf-wise leaf growth strategy with depth limit, gradient-based unilateral sampling and mutually exclusive feature bundling strategy, it has the advantages of not easy to overfit, high training accuracy and fast speed, and performs well in dealing with nonlinear classification and regression problems.
[0089] 3.1 Hyperparameter Tuning: To achieve optimal model performance, we used the GridSearchCV module from the Scikit-Learn library to perform hyperparameter tuning based on the LightGBM ensemble learning algorithm. The GridSearchCV module searches the entire specified hyperparameter value space, exhaustively enumerating the optimal hyperparameter combinations. We then used 5-fold cross-validation to evaluate the performance of each hyperparameter combination to ensure the global optimal solution was found. The hyperparameters that significantly impact model performance include the number of iterations, learning rate, maximum tree depth, and number of leaf nodes. The hyperparameter search range, step size, and optimal values are shown in Table 5.
[0090] Table 5 LightGBM ensemble learning algorithm hyperparameter tuning
[0091] Hyperparameters Value range and step size Optimal value Number of iterations 30-100, step size 10 80 Learning rate 0.01、0.02、0.05、0.1、0.15 0.1 Maximum tree depth 3-6, step 1 3 Number of leaf nodes 5-17, step 2 5
[0092] 3.2 Model training: The specific training steps are as follows:
[0093] (1) Load the training set sample data into memory, initialize the decision tree f0(x), and select the softmax cross entropy function as the error function l:
[0094]
[0095] Among them, M represents the number of categories, y c Represents an indicator variable. If the sample prediction category is consistent with the true category, the value is 1, otherwise it is 0. c Indicates the predicted probability that the sample prediction category belongs to category c.
[0096] (2) For the tth (t=1, 2, ..., T) decision tree:
[0097] 1) Discretize the continuous eigenvalues into k (default k = 255) buckets through the piecewise function, traverse all the data, and calculate the first-order gradient (g) of the samples in each bucket respectively. t ) and the second-order gradient (h t ) and calculate the first-order gradient (G t ) and second-order gradient (H t ) and:
[0098]
[0099] Among them, g ti and h ti They represent the sum of the first-order gradient and the second-order gradient of the sample data in the i-th bucket respectively.
[0100] 2) Split the decision tree based on the current node. The sum of the first-order gradient and second-order gradient of all buckets is G and H respectively. The initial score is set to 0. For feature n (n = GR, U, U / Th, AC, DEN, CNL, LLD):
[0101] ①Left child node: G tl =0,H tl =0, right child node: G tr 、H tr .
[0102] ② Arrange the buckets of sample feature n from small to large, take out the i-th bucket in turn, put the bucket into the left child node, and calculate the sum of the first-order gradients of the left child node at this time G tl * , the sum of the second-order gradient H tl * And the sum of the first-order gradients of the right child nodes G tr * , the sum of the second-order gradient H tr * :
[0103] G tl * =G tl +g ti , G tr * =GG tl *
[0104] H tl * =H tl +h ti , H tr * =HH tl *
[0105] ③Update score:
[0106]
[0107] Among them, γ represents the complexity penalty coefficient of the new leaf node, and λ represents the complexity penalty coefficient of the decision tree.
[0108] 3) Select the feature and bucket position corresponding to the maximum score for node splitting.
[0109] 4) When the number of leaf nodes or the maximum tree depth reaches the preset value, the current decision tree is established and the strong learner F is updated. t (x):
[0110] F t (x) = F t-1 (x)+f t (x)
[0111] Among them F t-1 (x) represents the strong learner obtained from the t-1th decision tree, f t (x) is the decision tree generated in this round of iteration.
[0112] (3) After T rounds of iteration, the final strong learner F is obtained T (x):
[0113]
[0114] The model is trained using the above steps and the optimal hyperparameters in Table 5. Figure 3 It can be seen that as the number of iterations increases, the error of the model on the training set and test set data decreases steadily and stabilizes at a low value, and the model as a whole tends to converge.
[0115] Step 4: Model performance evaluation.
[0116] The test set data was used to evaluate the performance of the established shale reservoir sweet spot discrimination model. The evaluation indicators were: weighted accuracy (W A ), weighted precision (W P ), weighted recall (W R ) and the kappa coefficient. Weighted accuracy represents the proportion of all correctly predicted samples to the total number of samples; weighted precision represents the proportion of correctly predicted samples of a certain class, weighted summed according to the proportion of samples in that class; weighted recall represents the proportion of correctly predicted samples of a certain class, weighted summed according to the proportion of samples in that class; and the kappa coefficient represents the reduction in errors compared to completely random classification (0.0-0.2 represents very low consistency, 0.21-0.4 represents fair consistency, 0.41-0.6 represents moderate consistency, 0.61-0.8 represents high consistency, and 0.81-1 represents almost perfect consistency). The larger the values of the four evaluation indicators, the stronger the model's ability to discriminate shale reservoir sweet spots.
[0117]
[0118] Among them, TP j represents the number of true positive examples of dessert category j (j = dessert category I, dessert category II, dessert category III), FP j represents the number of false positives for dessert category j, FN j represents the number of false negatives for dessert category j, N j and N represent the number of samples and the total number of samples of dessert category j, respectively. o and P e represent the actual agreement rate and the theoretical agreement rate, respectively.
[0119] Table 6 shows that, on the test data set, the shale reservoir sweet spot identification model based on the LightGBM ensemble learning algorithm achieved weighted accuracy, weighted precision, and weighted recall of approximately 0.85, with a Kappa coefficient of 0.737. The model's results were highly consistent with the sample. Model training, hyperparameter search, and validation took a total of 67.1 seconds. In summary, the method provided by the present invention offers the advantages of simple procedures, high efficiency, and high accuracy for shale reservoir sweet spot identification.
[0120] Table 6 Performance of shale reservoir sweet spot discrimination model
[0121]
[0122] Step 5. Application of the model: The established shale reservoir sweet spot discrimination model based on the LightGBM ensemble learning algorithm was applied to a shale gas well in the southern Sichuan Basin. There was no coring or experimental test data for this well. Using the well logging data (GR, U, U / Th, AC, DEN, CNL and LLD) of this well, the discrimination of different types of shale reservoir sweet spots in the target layer was achieved. Figure 4 .
[0123] Based on the same inventive concept, the present invention also proposes a shale reservoir sweet spot identification system, comprising:
[0124] The acquisition module is used to collect target area shale reservoir core sample data, and perform sweet spot discrimination on the shale reservoir core sample data to obtain the sweet spot discrimination type of the shale reservoir core sample.
[0125] The optimization module is used to select the optimal logging curve as the sample feature based on the logging data of the shale reservoir in the target area.
[0126] The model building module is used to build a shale reservoir sweet spot discrimination model based on the LightGBM ensemble learning algorithm; specifically, it includes: initializing a decision tree based on sample data, building a new decision tree based on sample characteristics, and obtaining the final strong learner after multiple rounds of update iterations to build a shale reservoir sweet spot discrimination model.
[0127] The discrimination module is used to select the optimal logging curve of the well to be identified in the target area shale reservoir and input it into the shale reservoir sweet spot discrimination model to obtain the discrimination type corresponding to the shale reservoir sweet spot.
[0128] Based on the same inventive concept, the present invention also proposes a computer device for identifying shale reservoir sweet spots, comprising: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the shale reservoir sweet spot identification method are implemented.
[0129] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for identifying sweet spots in shale reservoirs, characterized in that: include: Collecting target area shale reservoir core sample data, and performing sweet spot identification on the shale reservoir core sample data to obtain the sweet spot identification type of the shale reservoir core sample; Based on the target area shale reservoir logging data, the optimal logging curve in the logging data is selected as the sample feature. The specific steps include: obtaining the shale reservoir logging data of the target area target layer, specifically including 11 logging curves of GR, KTh, U, K, Th, AC, U / Th, DEN, CNL, LLD and PE; optimizing the logging parameters by using a recursive feature elimination algorithm to obtain the change of the shale reservoir sweet spot discrimination accuracy with the number of logging curves; and selecting the optimal logging curve based on the change of the shale reservoir sweet spot discrimination accuracy with the number of logging curves. A shale reservoir sweet spot discrimination model was constructed based on the LightGBM ensemble learning algorithm. Specifically, the method includes: initializing a decision tree based on core sample data, building a new decision tree based on sample characteristics, using the sweet spot discrimination type of the shale reservoir core sample as the sample data label, and constructing a shale reservoir sweet spot discrimination model after multiple rounds of update iterations; Select the optimal logging curve of the well to be identified in the target area shale reservoir and input it into the shale reservoir sweet spot identification model to obtain the identification type corresponding to the shale reservoir sweet spot; The shale reservoir sweet spot discrimination model is constructed based on the LightGBM ensemble learning algorithm, which specifically includes the following steps: Load sample data into memory and initialize the decision tree f 0( x ), error function l Select the softmax cross entropy function: ; in, M represents the number of categories, y c Indicates an indicator variable. When the sample prediction category is consistent with the true category, the value is 1, otherwise it is 0. p c Indicates that the sample prediction category belongs to the category c The predicted probability of For the t A decision tree, t =1, 2, ..., T :The continuous eigenvalues are discretized into k Buckets, traverse all data, calculate the sum of the first-order gradient and the sum of the second-order gradient of the samples in each bucket, and calculate the sum of the first-order gradient of all buckets G t , the sum of the second-order gradients H t : ; in, g ti Indicates the i The sum of the first-order gradients of the sample data in the buckets, h ti Indicates the i The sum of the second-order gradients of the sample data in the buckets; The decision tree is split based on the current node, and the sum of the first-order gradients of all buckets is G , the sum of the second-order gradients of all buckets is H , the initial score is set to 0; for the feature n =GR, U, U / TH, AC, DEN, CNL, LLD: The left child node is represented by G tl 、 H tl , the right child node is represented as G tr 、 H tr ; The sample features n The barrels are arranged from small to large, and the first i Bucket, put the bucket into the left child node, and calculate the sum of the first-order gradient of the left child node at this time. G tl * , the sum of the second-order gradients H tl * and the sum of the first-order gradients of the right child nodes G tr * , the sum of the second-order gradients H tr * : G tl * = G tl + g ti , G tr * = G - G tl * ; H tl * = H tl + h ti , H tr * = H - H tl * ; Update score: ; in, γ Represents the complexity penalty coefficient of the new leaf node, Represents the decision tree complexity penalty coefficient; Select the feature and bucket position corresponding to the maximum score to split the node; When the number of leaf nodes or the maximum tree depth reaches the preset value, the current decision tree is established and the strong learner is updated. F t ( x ): ; in F t-1 ( x ) indicates the t -1 strong learner obtained from the decision tree, f t ( x ) is the decision tree generated in this round of iteration; go through T After rounds of iteration, the final strong learner is obtained F T ( x ): ; in, I Represents the learning rate.
2. The method for identifying sweet spots in shale reservoirs according to claim 1, characterized in that: After collecting the target area shale reservoir core sample data, the discrimination parameters of the core sample sweet spot are obtained. The discrimination parameters of the shale reservoir sweet spot include total organic carbon content, porosity, total gas content and brittleness index. The acquisition process includes the following steps: The total organic carbon content is determined by the combustion method after the core samples are crushed, acid-washed, neutralized and dried; After the core plug samples were dried, the porosity of the samples was determined using Boyle's law; The desorption method was used to measure the desorbed gas content of the bulk sample and the residual gas content of the corresponding powder sample. The USBM direct method was used to calculate the lost gas content. The total gas content of the core sample was obtained by calculating the sum of the desorbed gas content, the residual gas content and the lost gas content. After the core samples were crushed, the contents of quartz, feldspar, clay minerals, carbonate minerals and pyrite in the samples were determined using X-ray diffraction technology, and the sample brittleness index was calculated using the following formula: ; in, BI represents the brittleness index, W Q 、 W F 、 W CM 、 W C and W P They represent the contents of quartz, feldspar, clay minerals, carbonate minerals and pyrite respectively.
3. The method for identifying sweet spots in shale reservoirs according to claim 1, characterized in that: The method also includes preprocessing the sample data, wherein the preprocessing process includes optimizing the well logging curve and normalizing the optimized well logging curve data.
4. The method for identifying sweet spots in shale reservoirs according to claim 1, wherein: The optimal logging curve is normalized and its expression is: ; in: X* represents the normalized logging curve data, X Represents the original logging curve data, μ represents the mean of the original logging data, σ Indicates the standard deviation of the original logging data.
5. The method for identifying sweet spots in shale reservoirs according to claim 1, characterized in that: The method also includes tuning the hyperparameters of the shale reservoir sweet spot discrimination model by calling the GridSearchCV module in the Scikit-Learn library; the method uses the GridSearchCV module to search through the entire specified hyperparameter value space, enumerates the optimal hyperparameter combination, and uses 5-fold cross-validation to evaluate the performance of each hyperparameter combination to ensure that the global optimal solution is found; the hyperparameters include the number of iterations, the learning rate, the maximum tree depth, and the number of leaf nodes.
6. The method for identifying sweet spots in shale reservoirs according to claim 1, characterized in that: The method also includes evaluating the shale reservoir sweet spot discrimination model through evaluation indicators; the evaluation indicators include weighted accuracy W A , weighted precision W P , weighted recall W R and kappa Coefficients; they are expressed as: ; ; ; ; in, TP j Dessert category j The number of true positive examples, FP j Dessert category j The number of false positives, FN j Dessert category j The number of false negatives, N j and N Dessert categories j The number of samples and the total number of samples, P o and P e represent the actual agreement rate and the theoretical agreement rate, respectively.
7. A shale reservoir sweet spot identification system, characterized in that: include: An acquisition module is used to collect target area shale reservoir core sample data, and perform sweet spot identification on the shale reservoir core sample data to obtain the sweet spot identification type of the shale reservoir core sample; The optimization module is used to select the optimal logging curve in the logging data as a sample feature based on the logging data of the shale reservoir in the target area. The module specifically includes the following steps: obtaining the logging data of the shale reservoir in the target area, specifically including 11 logging curves of GR, KTh, U, K, Th, AC, U / Th, DEN, CNL, LLD and PE; optimizing the logging parameters by using a recursive feature elimination algorithm to obtain the variation of the shale reservoir sweet spot discrimination accuracy with the number of logging curves; and selecting the optimal logging curve based on the variation of the shale reservoir sweet spot discrimination accuracy with the number of logging curves. The model construction module is used to construct a shale reservoir sweet spot discrimination model based on the LightGBM ensemble learning algorithm; specifically, the following steps are included: initializing a decision tree based on sample data, constructing a new decision tree based on sample characteristics, and constructing a shale reservoir sweet spot discrimination model after multiple rounds of update iterations; wherein, the construction of the shale reservoir sweet spot discrimination model based on the LightGBM ensemble learning algorithm specifically includes the following steps: Load sample data into memory and initialize the decision tree f 0( x ), error function l Select the softmax cross entropy function: ; in, M represents the number of categories, y c Indicates an indicator variable. When the sample prediction category is consistent with the true category, the value is 1, otherwise it is 0. p c Indicates that the sample prediction category belongs to the category c The predicted probability of For the t A decision tree, t =1, 2, ..., T :The continuous eigenvalues are discretized into k Buckets, traverse all data, calculate the sum of the first-order gradient and the sum of the second-order gradient of the samples in each bucket, and calculate the sum of the first-order gradient of all buckets G t , the sum of the second-order gradients H t : ; in, g ti Indicates the i The sum of the first-order gradients of the sample data in the buckets, h ti Indicates the i The sum of the second-order gradients of the sample data in the buckets; The decision tree is split based on the current node, and the sum of the first-order gradients of all buckets is G , the sum of the second-order gradients of all buckets is H , the initial score is set to 0; for the feature n =GR, U, U / TH, AC, DEN, CNL, LLD: The left child node is represented by G tl 、 H tl , the right child node is represented as G tr 、 H tr ; The sample features n The barrels are arranged from small to large, and the first i Bucket, put the bucket into the left child node, and calculate the sum of the first-order gradient of the left child node at this time. G tl * , the sum of the second-order gradients H tl * and the sum of the first-order gradients of the right child nodes G tr * , the sum of the second-order gradients H tr * : G tl * = G tl + g ti , G tr * = G - G tl * ; H tl * = H tl + h ti , H tr * = H - H tl * ; Update score: ; in, γ Represents the complexity penalty coefficient of the new leaf node, Represents the decision tree complexity penalty coefficient; Select the feature and bucket position corresponding to the maximum score to split the node; When the number of leaf nodes or the maximum tree depth reaches the preset value, the current decision tree is established and the strong learner is updated. F t ( x ): ; in F t-1 ( x ) indicates the t -1 strong learner obtained from the decision tree, f t ( x ) is the decision tree generated in this round of iteration; go through T After rounds of iteration, the final strong learner is obtained F T ( x ): ; in, I represents the learning rate; The discrimination module is used to select the optimal logging curve of the well to be identified in the target area shale reservoir and input it into the shale reservoir sweet spot discrimination model to obtain the discrimination type corresponding to the shale reservoir sweet spot.
8. A computer device for identifying sweet spots in shale reservoirs, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of the shale reservoir sweet spot identification method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for recognizing sweet spots in shale stratum
CN105986816A
Method and device for identifying'dessert 'of shale oil and gas reservoir and storage medium
CN114638300A