Machine learning algorithm-based continental facies shale oil reservoir favorable lithofacies prediction method
The flange facies prediction model constructed through machine learning algorithms solves the problem of high cost and poor accuracy of lithophago prediction in continental shale oil reservoirs, and achieves low-cost and high-accuracy facies prediction, which is suitable for areas with low exploration levels.
Patent Information
- Application Number
- CN202510578412.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has high cost and poor accuracy in lithophagometry prediction in terrestrial shale oil reservoirs, especially in areas with low exploration and little physical data, and the imbalance in the logging data of the favorable facies and non-friendly facies leads to large errors in the prediction results.
Using machine learning algorithms, we divide lithophagophyte types by obtaining the mineralogy, sedimentary structure and organic geochemical characteristics of core samples of shale oil reservoirs, and establish a sample set based on logging curve data, build a favorable lithophagophyte prediction model, and use support vector machine algorithm for training to achieve rapid prediction of favorable lithophagophytes.
It realizes low-cost, high-accuracy favorable lithofocal prediction, which is suitable for areas with low exploration, and improves the accuracy and practicality of the prediction.
Smart Images

Figure CN120493071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of continental shale oil and gas exploration and development, and in particular to a method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm. Background Art
[0002] Shale oil, as a typical unconventional oil resource, has huge resource potential.
[0003] Lithofacies encompasses a wide range of characteristics, including rock color, mineral composition, grain size, and sedimentary structure. Lithofacies research plays a crucial role in global shale oil and gas exploration and development. Many researchers have chosen to investigate the mineralogical, sedimentary, and biostructural characteristics of shale. They manually classify lithofacies based on core descriptions and thin section observations, and develop shale lithofacies models in conjunction with well logging and seismic data.
[0004] However, these methods require extensive physical data and relevant support, consume enormous human and material resources, and are costly to produce. Furthermore, they are less applicable to areas with low levels of exploration and limited physical data, resulting in low lithofacies prediction accuracy.
[0005] In addition, the amount of logging data for favorable lithofacies is unbalanced with that for unfavorable lithofacies. Compared with unfavorable lithofacies, the amount of logging data for favorable lithofacies is less, and the corresponding reservoir characteristic sample size is correspondingly smaller, resulting in a large error between the predicted results of favorable lithofacies and the actual situation. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides a method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm.
[0007] The present invention is achieved through the following technical solutions:
[0008] A method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm includes the following steps:
[0009] Obtaining a shale oil reservoir core sample, and determining mineralogical characteristics, sedimentary structural characteristics, and organic geochemical characteristics based on the shale oil reservoir core sample;
[0010] Classify lithofacies types based on the mineralogical characteristics, sedimentary structural characteristics and organic geochemical characteristics;
[0011] Based on the divided lithofacies types, determine the main lithofacies types;
[0012] Based on the main lithofacies types, carry out reservoir evaluation of the corresponding lithofacies from both qualitative and quantitative aspects;
[0013] Based on the reservoir evaluation results, each lithofacies is classified according to reservoir quality to determine favorable lithofacies and unfavorable lithofacies of the shale oil reservoir;
[0014] Establishing a sample set based on the well logging curve data of the favorable lithofacies and unfavorable lithofacies;
[0015] Based on the sample set, constructing new parameter indices based on spontaneous potential logging and neutron logging;
[0016] Construct a favorable lithofacies prediction model based on machine learning, and use the new parameter data to train the favorable lithofacies prediction model to obtain a trained favorable lithofacies prediction model;
[0017] The trained favorable lithofacies prediction model is used to predict the favorable lithofacies of continental shale oil reservoirs.
[0018] Preferably, the qualitative evaluation includes: evaluating the reservoir space type and development degree of the rock based on detailed core observation, microscopic thin section identification and scanning electron microscopy observation; and evaluating the complexity of the rock's pore structure based on nitrogen and carbon dioxide warm adsorption experiments.
[0019] Preferably, the quantitative evaluation includes: obtaining the OSI index based on the total organic carbon content and the pyrolysis hydrocarbon content to evaluate the oil content of the rock; obtaining porosity and permeability data based on physical property tests to evaluate reservoir properties; and calculating the brittleness index based on the ratio of each brittle mineral to the total minerals by calculating the ratio of each brittle mineral to the total minerals based on the mineral composition data of the rock.
[0020] Preferably, the logging curve data includes natural potential logging data, natural gamma ray logging data, acoustic transit time logging data, neutron logging data, density logging data and resistivity logging data.
[0021] Optionally, the method for establishing the sample set includes: collecting well logging data of favorable lithofacies and unfavorable lithofacies, and performing sample expansion processing on a certain type of lithofacies for which the amount of well logging data is too small.
[0022] Preferably, the sample expansion processing method includes: inputting original logging data and labels, the original logging data including natural potential logging data, natural gamma logging data, acoustic time difference logging data, neutron logging data, density logging data and resistivity logging data, and the labels including favorable lithofacies and unfavorable lithofacies; for minority lithofacies samples, generating deterministic candidate samples based on Kriging interpolation, and then superimposing Gaussian noise to form an initial data population; performing fitness calculation and genetic optimization on the initial data population; terminating when the fitness fluctuation is continuously less than a preset value or reaches the maximum number of iterations, outputting the optimal synthetic sample set, and merging the synthetic sample set with the original data to form a new sample set.
[0023] Preferably, the fitness calculation formula is:
[0024] Fit=α·F1+β·G-mean-γ·NR-δ·PV-∈·SC
[0025] In the above formula, Fit represents fitness, F1 and G-mean represent classification indicators, NR represents noise ratio, PV represents rock physical constraint term, SC represents spatial continuity penalty term, and α, β, γ, δ, ∈ are weight coefficients.
[0026] Among them, the genetic optimization includes three operations: selection, crossover and mutation. The selection operation refers to roulette selection, which selects individuals according to probability based on the proportion of individual fitness value to total fitness; the crossover operation refers to the joint crossover of strongly correlated logging parameters and the independent crossover of weakly correlated logging parameters; the mutation operation refers to setting the variation range according to the measurement error range of the logging parameters.
[0027] Wherein, the new parameter index is expressed as:
[0028]
[0029] In the above formula, I sc Refers to the new parameter index, SP max Refers to the maximum value of the spontaneous potential curve, SP min Refers to the minimum value of the natural potential curve, CNL max Refers to the maximum value of the neutron logging curve, CNL min Refers to the minimum value of the neutron logging curve.
[0030] Preferably, the machine learning algorithm is a support vector machine algorithm.
[0031] Compared with the prior art, this application has at least the following beneficial effects:
[0032] 1. This application uses a favorable lithofacies prediction model based on machine learning to quickly predict favorable lithofacies for continental shale oil reservoirs with low prediction cost and high accuracy;
[0033] 2. This application can further improve the accuracy by expanding the data of reservoir characteristic samples with a small amount of samples and increasing the amount of lithofacies logging data. It is suitable for areas with low exploration level and little physical data, and has good practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1 Flowchart of a method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm in an embodiment;
[0036] Figure 2 This is a correlation diagram between TOC and S1 in the embodiment;
[0037] Figure 3 Schematic diagram of the shale oil reservoir lithofacies division scheme in the embodiment;
[0038] Figure 4 This is a statistical diagram of shale oil reservoir lithofacies types in the embodiment;
[0039] Figure 5 This is a standard diagram for shale oil reservoir classification evaluation in the embodiment;
[0040] Figure 6 Schematic diagram of Mahakil genetic algorithm in the embodiment;
[0041] Figure 7 This is a cross-sectional view of the lithofacies prediction results in the embodiment. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other in the absence of conflict. It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other.
[0043] like Figure 1 As shown, the method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm disclosed in this embodiment includes the following steps:
[0044] S1, obtain shale oil reservoir core samples and determine the mineralogical characteristics, sedimentary structural characteristics and organic geochemical characteristics based on the shale oil reservoir samples, including:
[0045] (1) X-ray diffraction experiments were conducted on core samples of oil reservoirs to obtain mineral composition data of the rocks. A classification scheme based on the relative percentage content of clay minerals, felsic minerals, and carbonate minerals was established, with 75% and 50% as the boundaries. For example, mudstone is defined as clay mineral content above 75%; limestone is defined as clay mineral content between 50% and 75% and if the carbonate mineral content is higher than the felsic mineral content, and sandy mudstone is defined as clay mineral content below 50%. Based on the above classification principles, three major categories can be divided according to the mineral composition: limestone, clay mineral, and felsic mineral.
[0046] (2) Based on detailed core observation, microscopic thin section identification and scanning electron microscopy observation, the rock sedimentary structure characteristics are determined according to the thickness and color of the single layer. For example, if the single layer thickness is less than 1 mm, and the core or thin section shows dense superposition of light and dark thin layers, it is determined to be a laminated sedimentary structure; if the single layer thickness is between 1 mm and 10 mm, and the core or thin section shows superposition of light and dark layers, it is determined to be a layered sedimentary structure; if the core or thin section shows no obvious bedding, and the mineral components are evenly distributed, it is determined to be a massive sedimentary structure.
[0047] (3) The total organic carbon content (TOC) and pyrolysis hydrocarbon content (S1) of shale oil reservoir core samples were measured, and the organic matter abundance classification standard was determined based on the correlation between the total organic carbon content (TOC) and the pyrolysis hydrocarbon content (S1).
[0048] For example, Figure 2 As shown in the figure, S1 shows a three-stage characteristic of "slow increase-rapid increase-stable" with the increase of TOC: when the TOC content is low (<0.4%), S1 increases slowly and remains at a low value stage; when the TOC content is between 0.4% and 2%, S1 changes significantly and shows a clear upward trend; when the TOC content is high (>2%), S1 maintains a relatively stable high value. At this stage, the amount of hydrocarbon generated by shale is greater than its own adsorption, and hydrocarbons need to be expelled. Therefore, the TOC content of 2% can be used as the boundary to divide shale into organic-rich shale (≥2%) and organic-bearing shale (<2%).
[0049] S2, based on S1, integrates three levels of organic matter abundance, sedimentary structure and mineral composition, such as Figure 3 As shown in the figure, a division scheme for continental shale oil reservoir lithofacies is established, and the naming rules of lithofacies refer to "organic matter abundance + sedimentary structure + mineral composition + phase".
[0050] S3, based on the lithofacies types classified in S2, theoretically there should be 54 types. However, the redundant lithofacies types are neither in line with the actual exploration of shale oil nor conducive to clarifying the reservoir differences of different lithofacies. Therefore, the main developed lithofacies types can be selected for research based on the actual geological conditions of the study area.
[0051] For example, Figure 4 As shown in the figure, based on the actual test and collection of 230 sample data and core data, a total of 20 lithofacies types are developed in the shale oil reservoirs in the study area, among which the main lithofacies types are organic-rich laminated muddy limestone facies, organic-rich laminated limy mudstone facies, organic-rich laminated muddy limestone facies, organic-rich laminated limy mudstone facies, organic-bearing laminated sandy mudstone facies and organic-bearing massive muddy limestone facies.
[0052] S4: Based on the main lithofacies types determined in S3 and combined with relevant experimental methods, the corresponding lithofacies reservoir evaluation is carried out from both qualitative and quantitative aspects. The qualitative evaluation includes:
[0053] (1) Evaluate the reservoir space type and development degree of the rock based on detailed core observation, thin section identification and scanning electron microscopy observation;
[0054] (2) Based on nitrogen and carbon dioxide isothermal adsorption experiments, the complexity of the rock pore structure was evaluated.
[0055] Quantitative evaluation includes:
[0056] (1) Based on the total organic carbon content (TOC) and pyrolysis hydrocarbon content (S1), the OSI index is obtained to evaluate the oil content of the rock. The OSI index can be expressed as:
[0057]
[0058] In the above formula, S1 represents the free hydrocarbons released by thermal decomposition, mg / g; TOC represents the total organic carbon content, %.
[0059] (2) Based on physical property tests, obtain porosity and permeability data and evaluate reservoir properties;
[0060] (3) Based on the mineral composition data of the rock, the brittleness index is calculated by calculating the ratio of brittle minerals (quartz, feldspar, carbonate minerals) to the total minerals, which can be expressed as:
[0061]
[0062] In the above formula, BI represents the brittleness index, %; I q Indicates the mass percentage of quartz, %; I f Indicates the mass percentage of feldspar, %; I c It represents the mass percentage of carbonate minerals, %; It represents the mass percentage of total minerals, %.
[0063] S5, the reservoir evaluation results based on the qualitative and quantitative combination of S4, such as Figure 5As shown in FIG, according to the shale oil reservoir classification and evaluation standard, each lithofacies is classified to determine the favorable lithofacies and unfavorable lithofacies of the shale oil reservoir.
[0064] S6, collects logging data for favorable and unfavorable lithofacies identified in S5, including spontaneous potential logging data, natural gamma ray logging data, acoustic transit time logging data, neutron logging data, density logging data, and resistivity logging data. The Mahakil genetic algorithm is used to process lithofacies with logging data less than 20% of the total data volume, expanding the sample size and addressing the bias in lithofacies prediction caused by the small amount of logging data for a particular lithofacies type. The process includes:
[0065] (1) Inputting original logging data and labels, the original logging data includes natural potential logging data, natural gamma logging data, acoustic transit time logging data, neutron logging data, density logging data and resistivity logging data, and the labels include favorable lithofacies and unfavorable lithofacies;
[0066] (2) For a few types of lithofacies samples, deterministic candidate samples are generated based on Kriging interpolation, and then Gaussian noise (standard deviation is 10% of the parameter measurement error) is superimposed to form the initial data population;
[0067] (3) Perform fitness calculation and genetic optimization on the initial data population.
[0068] The fitness calculation formula is:
[0069] Fit=α·F1+β·G-mean-γ·NR-δ·PV-∈·SC
[0070] In the above formula, Fit represents fitness, F1 and G-mean represent classification indicators, NR represents noise ratio, PV represents rock physical constraint term, SC represents spatial continuity penalty term, and α, β, γ, δ, ∈ are weight coefficients.
[0071] Genetic optimization includes three operations: selection, crossover, and mutation. The selection operation refers to roulette wheel selection, which selects individuals according to probability based on the proportion of individual fitness value to total fitness; the crossover operation refers to the joint crossover of strongly correlated logging parameters and the independent crossover of weakly correlated logging parameters; the mutation operation refers to setting the variation range according to the measurement error range of the logging parameters.
[0072] (4) When the fitness fluctuation is less than 1% for 20 consecutive generations or the maximum number of iterations (200 generations) is reached, the process is terminated, the optimal synthetic sample set is output, and the synthetic sample set is merged with the original data to form a new sample set.
[0073] The algorithm method diagram is as follows Figure 6 shown.
[0074] S7, using the sample set obtained in S6, constructs new parameter indices based on natural potential logging and neutron logging to distinguish favorable and unfavorable lithofacies. The new parameter indices can be expressed as:
[0075]
[0076] In the above formula, I sc Refers to the new parameter index, SP max Refers to the maximum value of the spontaneous potential curve, SP min Refers to the minimum value of the natural potential curve, CNL max Refers to the maximum value of the neutron logging curve, CNL min Refers to the minimum value of the neutron logging curve.
[0077] S8, performs linear normalization processing on the new parameter data constructed in S7 to obtain new parameter data after linear normalization processing. The linear normalization processing can be expressed as:
[0078]
[0079] Where: x′ ij 、x ij are the values of the i-th sample and j-th variable after and before the quantity processing respectively; x jmin is the minimum value of the jth variable; x jmax is the maximum value of the jth variable.
[0080] S9, divides the data obtained in S8 into training sets and test sets, where the training set accounts for 70% of the total number and is used for model training, and the test set accounts for 30% of the total number and is used to verify the model's prediction ability.
[0081] S10 uses the support vector machine (SVM) algorithm to perform model training on the training set in S9. The compilation environment for model training is Python 3.10. The SVM algorithm uses the relevant interfaces in the Sk-learn library. The graphical interface uses the relevant interfaces in Easygui to implement file or folder opening and saving operations. Data reading and scientific calculations are implemented based on the Pandas and Numpy libraries. In addition, the grid search method (Grid Search) is used to screen the model with the optimal hyperparameters to enhance the generalization ability of the model.
[0082] S11: Based on S10, a favorable lithofacies prediction model is established and applied to the test set to verify the recognition accuracy of favorable lithofacies.
[0083] like Figure 7As shown in the figure, the red circles represent the measured lithofacies and the green lines represent the predicted lithofacies, which intuitively presents the matching of the lithofacies profile. The measured favorable lithofacies in the study area are highly consistent with the predicted favorable lithofacies, and the comprehensive prediction accuracy reaches 90%, among which the prediction accuracy of the organic-rich laminated mudstone limestone facies reaches 92%, and the prediction accuracy of the organic-rich laminated gray mud facies reaches 88%.
[0084] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm, characterized in that: The following steps are involved: Obtaining a shale oil reservoir core sample, and determining mineralogical characteristics, sedimentary structural characteristics, and organic geochemical characteristics based on the shale oil reservoir core sample; Classify lithofacies types based on the mineralogical characteristics, sedimentary structural characteristics and organic geochemical characteristics; Based on the divided lithofacies types, determine the main lithofacies types; Based on the main lithofacies types, carry out reservoir evaluation of the corresponding lithofacies from both qualitative and quantitative aspects; Based on the reservoir evaluation results, each lithofacies is classified according to reservoir quality to determine favorable lithofacies and unfavorable lithofacies of the shale oil reservoir; Establishing a sample set based on the well logging curve data of the favorable lithofacies and unfavorable lithofacies; Based on the sample set, constructing new parameter indices based on spontaneous potential logging and neutron logging; Construct a favorable lithofacies prediction model based on machine learning, and use the new parameter data to train the favorable lithofacies prediction model to obtain a trained favorable lithofacies prediction model; The trained favorable lithofacies prediction model is used to predict the favorable lithofacies of continental shale oil reservoirs.
2. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 1, characterized in that: The qualitative evaluation includes: Evaluate the reservoir space type and development degree of the rock based on detailed core observation, thin section identification and scanning electron microscopy observation; The complexity of the rock pore structure is evaluated based on nitrogen and carbon dioxide warm adsorption experiments.
3. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 1, characterized in that: The quantitative evaluation includes: Based on the total organic carbon content and pyrolysis hydrocarbon content, the OSI index is obtained to evaluate the oil content of the rock; Based on physical property tests, porosity and permeability data are obtained to evaluate reservoir properties; Based on the mineral composition data of the rock, the brittleness index is calculated by calculating the ratio of each brittle mineral to the total minerals.
4. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 3, characterized in that: The OSI index is expressed as: In the above formula, S1 represents the free hydrocarbons released by pyrolysis, mg / g; TOC represents the total organic carbon content, %; The brittleness index is expressed as: In the above formula, BI represents the brittleness index, %; I q Indicates the mass percentage of quartz, %; I f Indicates the mass percentage of feldspar, %; I c It represents the mass percentage of carbonate minerals, %; It represents the mass percentage of total minerals, %.
5. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 1, characterized in that: The logging curve data includes natural potential logging data, natural gamma logging data, acoustic time difference logging data, neutron logging data, density logging data and resistivity logging data.
6. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 1 or 5, characterized in that: The method for establishing the sample set includes: collecting well logging curve data of favorable lithofacies and unfavorable lithofacies, and performing sample expansion processing on a certain type of lithofacies for which the amount of well logging data is too small.
7. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 6, characterized in that: The sample expansion processing method includes: Input the original logging data and labels. The original logging data includes natural potential logging data, natural gamma logging data, acoustic transit time logging data, neutron logging data, density logging data, and resistivity logging data. The labels include favorable lithofacies and unfavorable lithofacies. For a few types of lithofacies samples, deterministic candidate samples are generated based on Kriging interpolation, and then Gaussian noise is superimposed to form the initial data population; Perform fitness calculation and genetic optimization on the initial data population; When the fitness fluctuation is continuously less than the preset value or reaches the maximum number of iterations, the method is terminated and the optimal synthetic sample set is output. The synthetic sample set is merged with the original data to form a new sample set.
8. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 7, characterized in that: The fitness calculation formula is: Fit=α·F1+β·G-mean-γ·NR-δ·PV-∈·SC In the above formula, Fit represents fitness, F1 and G-mean represent classification indicators, NR represents noise ratio, PV represents rock physical constraint term, SC represents spatial continuity penalty term, α, β, γ, δ, ∈ are weight coefficients; The genetic optimization includes three operations: selection, crossover, and mutation. The selection operation refers to roulette wheel selection, which selects individuals according to probability based on the proportion of individual fitness values to total fitness; the crossover operation refers to joint crossover of strongly correlated logging parameters and independent crossover of weakly correlated logging parameters; the mutation operation refers to setting the variation range according to the measurement error range of the logging parameters.
9. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 1, 6 or 7, characterized in that: The new parameter index is expressed as: In the above formula, I sc Refers to the new parameter index, SP max Refers to the maximum value of the spontaneous potential curve, SP min Refers to the minimum value of the natural potential curve, CNL max Refers to the maximum value of the neutron logging curve, CNL min Refers to the minimum value of the neutron logging curve.
10. The method for predicting favorable lithofacies of continental shale oil reservoirs based on a machine learning algorithm according to claim 1, characterized in that: The machine learning algorithm is a support vector machine algorithm.