Shale organic matter enrichment main control factor quantitative characterization method and system based on machine learning
By constructing a multi-source geological parameter system, conducting correlation analysis and collinearity feature identification, and using principal component analysis and gradient boosting tree model, combined with various feature importance assessment methods, the problem of quantitative characterization of the main controlling factors of shale organic matter enrichment was solved, achieving high-precision and efficient resource evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWEST PETROLEUM UNIV
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies for quantitative characterization of the main controlling factors of shale organic matter enrichment suffer from problems such as insufficient multicollinearity handling, inaccurate nonlinear relationship modeling, unsystematic assessment of feature importance, and lack of standardized procedures for quantitative characterization, resulting in low accuracy and efficiency in exploration and development.
A machine learning-based approach is adopted to construct a multi-source geological parameter system, conduct correlation analysis and collinearity feature identification, use principal component analysis for dimensionality reduction, establish nonlinear relationships by combining gradient boosting tree model, and use multiple feature importance assessment methods and weighted fusion to output quantitative characterization results.
It effectively solves the multicollinearity problem, accurately characterizes nonlinear relationships, realizes a systematic assessment of the importance of multidimensional features, establishes a standardized quantitative characterization process, improves the accuracy and efficiency of shale gas resource evaluation, and reduces exploration risks.
Smart Images

Figure CN121997203A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unconventional energy exploration and evaluation technology, and in particular to a method and system for quantitative characterization of the main controlling factors of shale organic matter enrichment based on machine learning. Background Technology
[0002] Shale gas, as an important unconventional energy resource, has its exploration and development value closely related to the degree of organic matter enrichment in shale. Organic matter enrichment is a crucial step in the formation and preservation of shale gas, and its degree is controlled by a combination of geological factors, including paleoclimate, paleoproductivity, redox conditions, terrigenous input intensity, and sedimentation rate. Accurately identifying and quantitatively characterizing the main controlling factors of shale organic matter enrichment is of significant guiding importance for shale gas resource evaluation and exploration deployment.
[0003] Currently, research on the main controlling factors of shale organic matter enrichment mainly relies on geostatistical methods and simple correlation analysis. Traditional methods typically use single indicators or simple linear regression models to analyze the relationship between various factors and organic matter content, which has the following obvious limitations:
[0004] First, multicollinearity is a common problem among geological parameters from multiple sources, making traditional correlation analysis results unreliable. For example, paleoproductivity indicators (such as Ba) bio There are often high correlations between P / Ti and P / Al, and between redox indices such as U / Th and V / (V+Ni), making it difficult to accurately assess the contribution of a single index.
[0005] Secondly, the relationship between shale organic matter enrichment and multiple factors exhibits significant nonlinear characteristics, and traditional linear models struggle to accurately characterize this complex relationship. The influence of geological environmental factors on organic matter enrichment is often not a simple linear superposition, but rather involves complex interactions.
[0006] Third, existing research lacks a systematic mechanism for assessing the importance of features. Single feature importance assessment methods (such as Pearson correlation coefficient) cannot fully reflect the contribution of each factor in the organic matter enrichment process, and there are differences in the results between different assessment methods, lacking an effective integration mechanism.
[0007] Fourth, the lack of standardized procedures for quantitative characterization methods makes it difficult to compare and integrate results from different studies. Current research largely remains at the qualitative descriptive stage, failing to provide a quantitative assessment of the contribution of key controlling factors, thus limiting its application in practical exploration and development.
[0008] In recent years, machine learning methods have been increasingly applied in the field of geology, with algorithms such as random forests and support vector machines being attempted for geological parameter prediction. However, existing research still has the following shortcomings in the quantitative characterization of the main controlling factors of shale organic matter enrichment: 1) It has not effectively addressed the multicollinearity problem among geological parameters; 2) It has not used multiple feature importance assessment methods for systematic analysis; 3) It lacks a scientific integration mechanism for the results of different assessment methods; and 4) It has not established a standardized quantitative characterization process.
[0009] Therefore, there is an urgent need for a quantitative characterization method for the main controlling factors of shale organic matter enrichment that can effectively handle the complex relationships between multiple geological parameters and accurately quantify the contribution of each factor, so as to improve the accuracy and efficiency of shale gas resource evaluation. Summary of the Invention
[0010] This invention addresses the shortcomings of existing technologies by providing a machine learning-based method and system for quantitatively characterizing the main controlling factors of shale organic matter enrichment. It aims to solve problems such as insufficient handling of multicollinearity, inaccurate modeling of nonlinear relationships, unsystematic assessment of feature importance, and lack of standardized procedures for quantitative characterization in existing technologies.
[0011] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:
[0012] A machine learning-based method for quantitative characterization of the main controlling factors of shale organic matter enrichment includes the following steps:
[0013] A parameter system for influencing factors of shale organic matter enrichment was constructed, including paleoclimate indicators, paleoproductivity indicators, redox indicators, terrigenous input intensity indicators, and sedimentation rate variation indicators.
[0014] Correlation analysis and collinearity feature identification were performed on the parameter system of the influencing factors.
[0015] For paleoproductivity and redox indices exhibiting multicollinearity, feature dimensionality reduction was performed. Principal component analysis was then used to reduce the dimensionality of the paleoproductivity and redox indices respectively, yielding principal component variables.
[0016] Using total organic carbon content as a characterization parameter of shale organic matter enrichment, a gradient boosting tree-based machine learning model was constructed based on the principal component variables after dimensionality reduction and environmental indicators that did not participate in the dimensionality reduction process to model the nonlinear relationship between multi-source sedimentation, environmental factors and shale organic matter enrichment.
[0017] A variety of feature importance assessment methods are used to systematically evaluate the role of each input feature parameter in the machine learning model. These methods include statistical analysis methods, information theory analysis methods, model structure-related methods, and model interpretability methods.
[0018] The results obtained from different feature importance assessment methods are normalized and then weighted and fused according to preset weights to obtain the comprehensive influence value of each input feature parameter.
[0019] The comprehensive influence value is converted into a percentage form, and the quantitative characterization results of the main controlling factors of shale organic matter enrichment are output.
[0020] Furthermore, in the step of constructing the parameter system for influencing factors of shale organic matter enrichment, the preferred indicators include: paleoclimate indicators (C-value), paleoproductivity indicators (Babio, P / Ti, P / Al), redox indicators (U / Th, V / (V+Ni)), terrigenous input intensity indicators (Ti / Al), and sedimentary rate variation indicators (La). N / Yb N ).
[0021] Furthermore, in the feature dimensionality reduction step, principal component analysis is used to reduce the dimensionality of the paleoproductivity index and the redox index respectively, and the top principal components that explain 70% to 80% of the cumulative variance are selected as comprehensive variables.
[0022] Furthermore, the machine learning model is a regression model based on the XGBoost algorithm, wherein the maximum depth of the tree is 3 to 10, the learning rate is 0.01 to 0.03, the number of weak learners is 50 to 500, and the sample subsampling ratio is 0.5 to 1.0.
[0023] Furthermore, the multiple feature importance assessment methods include:
[0024] Statistical analysis methods: Pearson correlation coefficient analysis and ANOVA were used;
[0025] Information theory analysis method: mutual information regression analysis is used;
[0026] Model structure-related methods: Calculating decision tree gain values and permutation feature importance based on machine learning models;
[0027] Model interpretation method: SHAP value analysis is used.
[0028] Furthermore, in the step of weighted fusion of the results obtained from different feature importance assessment methods, the SHAP value analysis result has a weight of 30%, the permutation feature importance analysis result has a weight of 25%, the ANOVA variance analysis result has a weight of 20%, the decision tree gain analysis result has a weight of 15%, and the mutual information regression analysis result has a weight of 10%.
[0029] Furthermore, in the step of outputting the quantitative characterization results of the main controlling factors of shale organic matter enrichment, the quantitative characterization results are stored and displayed in the form of data tables, graphical results, or electronic data.
[0030] This invention also discloses a machine learning-based quantitative characterization system for the main controlling factors of shale organic matter enrichment, used to implement the aforementioned machine learning-based quantitative characterization method for the main controlling factors of shale organic matter enrichment, comprising:
[0031] The data acquisition module is used to acquire multi-source sedimentological parameters, geochemical parameters, and mineralogical parameters corresponding to shale samples;
[0032] The parameter filtering and preprocessing module is used to optimize the indicators of the acquired multi-source parameters.
[0033] The feature dimensionality reduction module is used to perform correlation analysis and feature dimensionality reduction on the optimized index parameters in order to construct an input feature parameter set for machine learning analysis.
[0034] The machine learning modeling module is used to construct a machine learning model of shale organic matter enrichment based on the input feature parameter set, and to train and validate the model.
[0035] The feature importance analysis module is used to calculate the influence of each input feature parameter on the enrichment degree of shale organic matter based on multiple feature importance assessment methods;
[0036] The fusion and output module is used to fuse and calculate the assessment results of the importance of different features, convert them into percentage form, and output the quantitative characterization results of the main controlling factors of shale organic matter enrichment.
[0037] Furthermore, the machine learning model used in the machine learning modeling module is a regression model based on the XGBoost algorithm, wherein the maximum depth of the tree is 3 to 10, the learning rate is 0.01 to 0.03, the number of weak learners is 50 to 500, and the sample subsampling ratio is 0.5 to 1.0.
[0038] Furthermore, in the fusion and output module, the weights for weighted fusion of the importance evaluation results of different features are set as follows: SHAP value analysis results have a weight of 30%, permutation feature importance analysis results have a weight of 25%, ANOVA variance analysis results have a weight of 20%, decision tree gain analysis results have a weight of 15%, and mutual information regression analysis results have a weight of 10%.
[0039] Compared with the prior art, the advantages of the present invention are as follows:
[0040] 1. Effectively solves the problem of multicollinearity: By performing principal component analysis (PCA) on paleoproductivity indicators and redox indicators to reduce dimensionality, high explanatory power principal component variables are successfully extracted, significantly reducing the interference of multicollinearity among geological parameters, making the identification results of the main control factors more reliable, and avoiding the risk of misjudgment in traditional correlation analysis.
[0041] 2. Accurate characterization of nonlinear relationships: The machine learning model based on gradient boosting trees can effectively capture the complex nonlinear interactions between shale organic matter enrichment and multi-source sedimentary environmental factors, improving the model's accuracy and adaptability in characterizing geological processes.
[0042] 3. Achieve a multi-dimensional feature importance system assessment: Innovatively integrate four types of feature importance assessment methods: statistical analysis, information theory analysis, model structure correlation, and model interpretability. Through a scientific weighted fusion mechanism, the bias of a single assessment method is completely eliminated, making the quantification results of the contribution of the main control factors more scientific and objective.
[0043] 4. Establish a standardized quantitative characterization process: Convert the comprehensive impact value into a percentage output, forming a repeatable and comparable standardized characterization system. This effectively addresses the industry pain point that existing research is mainly qualitative and difficult to quantify and compare, providing a directly applicable quantitative basis for resource evaluation.
[0044] 5. Improve exploration and development efficiency and decision-making accuracy: Significantly shorten the identification cycle of key control factors, significantly improve the accuracy and efficiency of exploration decisions, reduce exploration risks, and optimize development deployment plans.
[0045] 6. Promoting the upgrading of geological analysis paradigms: Breaking through the limitations of traditional geostatistical methods, it is the first to organically integrate multi-source data fusion, machine learning modeling and feature importance system evaluation, providing a standardized technical path that can be promoted for the evaluation of unconventional energy resources such as shale gas, and promoting the paradigm shift of geological analysis from qualitative description to quantitative characterization. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a flowchart of a method for quantitative characterization of the main controlling factors of shale organic matter enrichment based on machine learning in one embodiment of the present invention;
[0048] Figure 2This is a schematic diagram illustrating the evaluation and fusion of the importance of multiple method features in one embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the correlation analysis of input parameters in one embodiment of the present invention;
[0050] Figure 4 This is a schematic diagram of the quantitative ranking output of the main controlling factors of shale organic matter enrichment in one embodiment of the present invention; wherein, a is a stacked diagram of the relative contribution gradient of the main controlling factors of organic matter enrichment in wells W1 and W2, b is a bar chart of the contribution distribution of the main controlling factors of organic matter enrichment in well W1, c is a bar chart of the contribution distribution of the main controlling factors of organic matter enrichment in well W2, d is a bar chart comparing the contribution of the dominant controlling factors in well W1, and e is a bar chart comparing the contribution of the dominant controlling factors in well W2.
[0051] Figure 5 This is a structural block diagram of a machine learning-based quantitative characterization system for the main controlling factors of shale organic matter enrichment, according to one embodiment of the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] like Figure 1 and 2 As shown in the figure, the present invention provides a method for quantitative characterization of the main controlling factors of shale organic matter enrichment based on machine learning, comprising:
[0054] Step 1: Optimization of paleoenvironmental indicators and construction of parameter system
[0055] To address the issue of shale organic matter enrichment being synergistically controlled by multiple factors and the symbiotic and coupling relationships among various proxies, this study first systematically screens and optimizes geochemical indicators suitable for paleoenvironmental reconstruction, constructing an input parameter system for machine learning analysis. The indicator optimization follows these principles:
[0056] (1) Clear geological significance and mature application basis in paleoenvironment reconstruction: Prioritize the selection of proxy indicators that have been widely used in previous studies to indicate paleoproductivity, redox environment, terrigenous input intensity, sedimentation rate and paleoclimate change, and ensure that each indicator has clear environmental indicative significance.
[0057] (2) Data coverage is complete, stable, reliable and cost-effective: Priority is given to indicators with fewer missing data, fewer outliers and smaller analysis errors in the sample summary data of the study area. At the same time, economic factors can be taken into account, and data indicators that are relatively easy to obtain in isotope testing experiments can be selected to ensure sufficient sample quantity and stable results in the subsequent machine learning modeling process.
[0058] (3) Effectively reduce multicollinearity and signal overlap interference: In the same environmental factor category, parameters that are sensitive to environmental changes but have relatively low redundancy with other indicators are preferred to avoid information duplication due to high correlation between indicators.
[0059] (4) Regional adaptability principle applicable to the sedimentary background of the study area: Based on the characteristics of the sedimentary environment of the study area, the applicability of some traditional indicators was tested multiple times. For example, although the chemical weathering index (CIA) is often used for paleoclimate reconstruction, it has many outliers in the study area due to the influence of high continental input and rapid deposition, and it is difficult to effectively reflect the real climate signal. Therefore, it is not used as the main paleoclimate indicator.
[0060] Based on the above principles, the following indicators were ultimately selected as input parameters for the machine learning model: paleoclimate index (C-value), paleoproductivity index (Babio, P / Ti, P / Al), redox index (U / Th, V / (V+Ni)), terrigenous input intensity index (Ti / Al), sedimentation rate variation index (LaN / YbN), and organic matter enrichment characterization parameter (total organic carbon content, TOC). These indicators have high data completeness in the study area and possess clear environmental indicative significance, reflecting the paleoenvironmental characteristics of the Qiongzhusi Formation deposition period from different perspectives.
[0061] In this embodiment, shale samples from the fifth layer of the Qiongzhusi Formation in wells W1 and W2 in the study area were selected as the research object to obtain multi-source geochemical parameters and total organic carbon (TOC) data corresponding to the samples.
[0062] Step 2: Correlation analysis of indicators and identification of collinearity features.
[0063] like Figure 3 As shown, based on the optimization of indicators, correlation analysis is performed on the selected parameters to quantitatively assess the correlation between different indicators and identify potential symbiotic or coupling characteristics.
[0064] By calculating the correlation coefficients between each pair of indicators and analyzing their overall correlation structure, the results show that each indicator has a certain correlation within the same type of environmental factors, but can still reflect different sedimentary environments between different environmental factors, and has the representativeness to be used as input parameters for comprehensive analysis.
[0065] The results show that:
[0066] There is a strong positive correlation among ancient productivity indicators; there is also a significant correlation among redox indicators; and there is a certain degree of correlation coupling between ancient productivity indicators and redox indicators.
[0067] The above results indicate that directly treating all indicators as independent input variables may introduce multicollinearity interference, thereby affecting the stability of subsequent machine learning models and the reliability of feature importance assessment results.
[0068] Step 3: Multilinearity reduction and feature dimensionality reduction
[0069] Given the strong correlation between ancient productivity indicators and ancient redox condition indicators, and the potential for multicollinearity to interfere with the stability of machine learning models and the evaluation of feature importance, this invention performs dimensionality reduction on the relevant indicators before model construction.
[0070] First, the selected paleoproductivity and redox condition indicators were standardized to eliminate the influence of differences in the dimensions and numerical ranges of different indicators. Based on this, principal component analysis (PCA) was used to reduce the dimensionality of these two types of indicators. By calculating the eigenvalues and explained variance proportions of each component, the top principal components with a cumulative explained variance of ≥70%-80% were selected as comprehensive variables to represent the main variation characteristics of the corresponding environmental factors. The dimensionality-reduced principal component variables, while preserving the original environmental information to the greatest extent, effectively reduced collinearity interference between indicators, enabling subsequent machine learning models to more stably learn the nonlinear relationship between various environmental factors and organic matter enrichment.
[0071] If the paleoproductivity and redox indices exhibiting multicollinearity are not dimensionality-reduced, different input features may repeatedly express the same sedimentary environment information in the machine learning model. This can lead to the amplification or weakening of the model's response to specific environmental factors during training, resulting in instability in the feature importance assessment results and affecting the reliability of the quantitative ranking results of the main control factors.
[0072] This invention integrates the features of highly correlated indicators by introducing principal component analysis, which effectively reduces parameter redundancy while retaining the main environmental change information. This enables subsequent machine learning models to learn the nonlinear mapping relationship between various environmental factors and the enrichment degree of shale organic matter more stably, thereby improving the stability and repeatability of the quantitative characterization results of the main control factors.
[0073] In this embodiment, principal component analysis (PCA) is used to reduce the dimensionality of the two types of indicators. By calculating the eigenvalues of each principal component and the proportion of cumulative explained variance, the top principal components that can explain more than 70% to 80% of the variance of the original data are selected as characteristic variables that comprehensively characterize the ancient productivity conditions and redox conditions.
[0074] Through the dimensionality reduction process described above, the impact of multicollinearity on model training and feature importance evaluation results is effectively reduced while preserving the original environmental information to the greatest extent.
[0075] Step 4: Building a Machine Learning Modeling and Prediction Framework Based on XGBoost
[0076] After completing the index optimization and dimensionality reduction, the total organic carbon content (TOC) is used as the characterization parameter of the enrichment degree of shale organic matter. Based on the dimensionality reduction feature parameter set obtained in step 3, a machine learning model of shale organic matter enrichment is constructed.
[0077] Specifically, the principal component variables (characterized by paleoproductivity and redox conditions) after dimensionality reduction through principal component analysis, along with environmental indicators not involved in the dimensionality reduction process, are used as the model input features. These environmental indicators not involved in the dimensionality reduction process include at least the paleoclimate index C-value, the terrigenous input index Ti / Al, and the sedimentary rate indicator (La). N / Yb N Using TOC as the model output variable, an XGBoost regression model based on gradient boosting trees is constructed to simulate the nonlinear mapping relationship between multi-source sedimentation and environmental factors and the enrichment degree of shale organic matter.
[0078] During model training, a K-fold cross-validation strategy is preferred for training and validating the model, where K is an integer greater than or equal to 3, preferably 5 to 10. The key parameters of the gradient boosting tree-based machine learning model include at least the maximum tree depth, learning rate, number of weak learners, and sample subsampling ratio. The maximum tree depth can be selected from 3 to 10, the learning rate from 0.01 to 0.03, the number of weak learners from 50 to 500, and the sample subsampling ratio from 0.5 to 1.0. By adjusting the model within the above parameter range, the error between the model's predicted output and the measured total organic carbon content of shale is gradually reduced, thereby obtaining a trained shale organic matter enrichment prediction model.
[0079] After simulation training, the model performance was evaluated using a validation dataset to ensure that the model could reasonably characterize the shale organic matter enrichment characteristics under the synergistic effect of multiple factors. By constructing a machine learning model based on gradient boosting trees, a comprehensive modeling of the degree of shale organic matter enrichment by multi-source sedimentation and environmental parameters was achieved, overcoming the problem that existing technologies based on single indicators or linear models are unable to characterize complex control relationships.
[0080] Step 5: Multi-method feature assessment, importance evaluation, fusion, and quantitative characterization of controlling factors
[0081] After the machine learning model constructed in step 4 is trained and its predictive performance is confirmed by the validation dataset, the degree of influence of each input feature parameter in the model is systematically evaluated based on the trained machine learning model, so as to achieve quantitative identification of the main controlling factors of shale organic matter enrichment.
[0082] Specifically, firstly, based on different types of feature importance assessment methods, the relationship between each input feature parameter and shale organic matter enrichment parameters is calculated to obtain multiple sets of importance evaluation results. The feature importance assessment methods include the following categories:
[0083] (1) Statistical analysis methods:
[0084] Pearson correlation coefficient analysis was used to calculate the linear correlation between each input characteristic parameter and the shale organic matter enrichment characterization parameter; at the same time, ANOVA variance analysis was used to evaluate the explanatory power and statistical significance of different characteristic parameters on the shale organic matter enrichment characterization parameter.
[0085] (2) Information theory analysis method:
[0086] Mutual information regression analysis was used to quantify the amount of information provided by each input characteristic parameter to the characterization parameters of shale organic matter enrichment, in order to identify characteristic parameters that have no significant linear correlation but have nonlinear control effects.
[0087] (3) Model structure related methods:
[0088] Based on the internal structural features of the machine learning model, the decision tree gain value and the importance of the permutation features are calculated to evaluate the contribution of each input feature parameter to error reduction during the model prediction process.
[0089] (4) Model interpretation methods:
[0090] Based on SHAP value analysis, the model prediction results are interpreted, the marginal contribution of each input feature parameter in a single sample and the overall sample is calculated, and its positive or negative influence characteristics are statistically analyzed to reveal the trend of the role of different feature parameters in the organic matter enrichment process.
[0091] After obtaining the above-mentioned multiple feature importance evaluation results, the feature importance results obtained by different methods are normalized to eliminate differences in numerical scale between the different evaluation methods. Subsequently, weights are assigned to different feature importance evaluation methods, and the normalized feature importance results are weighted and fused to obtain the comprehensive impact value of each income feature parameter.
[0092] Based on the comprehensive influence value, each input characteristic parameter is quantitatively ranked, and the ranking result is output as the quantitative characterization result of the main controlling factor of shale organic matter enrichment.
[0093] Step 6: Calculation of the comprehensive impact value of the main control factors, standardized output, and control mode identification
[0094] Based on the multiple feature importance evaluation results obtained in step 5, in order to integrate the advantages of different feature importance evaluation methods in terms of model interpretability and stability, a weighted fusion strategy is adopted to integrate the multiple feature importance evaluation results.
[0095] Specifically, based on the contribution of different feature importance assessment methods to the explanatory power of shale organic matter enrichment processes, the following weights are assigned: SHAP value analysis results have a weight of 30%, permutation feature importance analysis results have a weight of 25%, ANOVA variance analysis results have a weight of 20%, decision tree gain analysis results have a weight of 15%, and mutual information regression analysis results have a weight of 10%. The weights of different feature importance assessment methods can be set according to model prediction stability, validation set error, or consistency of evaluation results. The weights can be pre-set fixed values or adaptively adjusted based on model training results. This invention is not limited to the specific weight ratios mentioned above.
[0096] By weighting the importance results of each feature parameter obtained under different evaluation methods, the comprehensive influence value corresponding to each input feature parameter is obtained. Furthermore, the quantitative characterization results of the main controlling factors are stored and displayed in the form of data tables, graphical results, or electronic data for subsequent analysis and application.
[0097] Furthermore, based on the quantitative characterization results of the main controlling factors, the organic matter enrichment control patterns of shale in different layers or different samples in the study area are identified to distinguish between organic matter enrichment models dominated by a single environmental factor or controlled by multiple environmental factors in synergy.
[0098] The quantitative identification of control factors and determination of enrichment patterns in this embodiment are as follows:
[0099] 1. Overall characteristics of the 5th layer of the W1 Jingqiongzhusi Formation:
[0100] like Figure 4As shown, the enrichment of organic matter is the result of multi-factor coupling, among which redox conditions are the dominant factor, contributing 42.35%, while paleoproductivity and terrigenous input are secondary controlling factors, indicating that the well as a whole belongs to the enrichment mode dominated by preservation conditions.
[0101] W1 well section characteristics:
[0102] Section a: Redox conditions contributed 47.41% and productivity 31.84%;
[0103] Section b: Redox conditions contributed 31.10% and productivity 18.35%;
[0104] Section c: Productivity contribution rises to 49.81%, redox conditions account for 21.49%.
[0105] The results show that the main controlling factors differ significantly among different layers, reflecting the dynamic evolution characteristics of the sedimentary environment in the vertical direction.
[0106] 2. Overall characteristics of the 5th layer of the W2 Well Qiongzhu Temple Formation:
[0107] Organic matter enrichment was mainly dominated by paleoproductivity and terrestrial input, contributing 31.34% and 32.03% respectively. Redox conditions contributed less, only 10.97%, showing a typical productivity-source synergistic control mode.
[0108] W2 well segmentation characteristics:
[0109] Section a: Productivity-driven (47.19%)
[0110] Section b: Productivity and land-based inputs are jointly controlled;
[0111] Section c: Land-based inputs were dominant (51.17%).
[0112] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0113] In another embodiment, a machine learning-based quantitative characterization system for the main controlling factors of shale organic matter enrichment is provided. This system corresponds one-to-one with the machine learning-based quantitative characterization method for the main controlling factors of shale organic matter enrichment described in the above embodiment, such as... Figure 5 As shown, it includes:
[0114] The data acquisition module is used to acquire multi-source sedimentological parameters, geochemical parameters, and mineralogical parameters corresponding to shale samples;
[0115] The parameter filtering and preprocessing module is used to optimize the indicators of the acquired multi-source parameters.
[0116] The feature dimensionality reduction module is used to perform correlation analysis and feature dimensionality reduction on the optimized index parameters in order to construct an input feature parameter set for machine learning analysis.
[0117] The machine learning modeling module is used to construct a machine learning model of shale organic matter enrichment based on the input feature parameter set, and to train and validate the model.
[0118] The feature importance analysis module is used to calculate the influence of each input feature parameter on the enrichment degree of shale organic matter based on multiple feature importance assessment methods;
[0119] The fusion and output module is used to fuse and calculate the assessment results of the importance of different features, convert them into percentage form, and output the quantitative characterization results of the main controlling factors of shale organic matter enrichment.
[0120] Furthermore, the machine learning model used in the machine learning modeling module is a regression model based on the XGBoost algorithm, wherein the maximum depth of the tree is 3 to 10, the learning rate is 0.01 to 0.03, the number of weak learners is 50 to 500, and the sample subsampling ratio is 0.5 to 1.0.
[0121] Furthermore, in the fusion and output module, the weights for weighted fusion of the importance evaluation results of different features are set as follows: SHAP value analysis results have a weight of 30%, permutation feature importance analysis results have a weight of 25%, ANOVA variance analysis results have a weight of 20%, decision tree gain analysis results have a weight of 15%, and mutual information regression analysis results have a weight of 10%.
[0122] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0124] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for quantitatively characterizing the main controlling factors of shale organic matter enrichment based on machine learning, characterized in that, Includes the following steps: A parameter system for influencing factors of shale organic matter enrichment was constructed, including paleoclimate indicators, paleoproductivity indicators, redox indicators, terrigenous input intensity indicators, and sedimentation rate variation indicators. Correlation analysis and collinearity feature identification were performed on the parameter system of the influencing factors. For paleoproductivity and redox indices exhibiting multicollinearity, feature dimensionality reduction was performed. Principal component analysis was then used to reduce the dimensionality of the paleoproductivity and redox indices respectively, yielding principal component variables. Using total organic carbon content as a characterization parameter of shale organic matter enrichment, a gradient boosting tree-based machine learning model was constructed based on the principal component variables after dimensionality reduction and environmental indicators that did not participate in the dimensionality reduction process to model the nonlinear relationship between multi-source sedimentation, environmental factors and shale organic matter enrichment. A variety of feature importance assessment methods are used to systematically evaluate the role of each input feature parameter in the machine learning model. These methods include statistical analysis methods, information theory analysis methods, model structure-related methods, and model interpretability methods. The results obtained from different feature importance assessment methods are normalized and then weighted and fused according to preset weights to obtain the comprehensive influence value of each input feature parameter. The comprehensive influence value is converted into a percentage form, and the quantitative characterization results of the main controlling factors of shale organic matter enrichment are output.
2. The method for quantitative characterization of the main controlling factors of shale organic matter enrichment according to claim 1, characterized in that, In the steps of constructing the parameter system for influencing factors of shale organic matter enrichment, the preferred indicators include: paleoclimate indicators, paleoproductivity indicators, redox indicators, terrigenous input intensity indicators, and sedimentation rate variation indicators.
3. The method for quantitative characterization of the main controlling factors of shale organic matter enrichment according to claim 1, characterized in that, In the feature dimensionality reduction step, principal component analysis is used to reduce the dimensionality of the paleoproductivity index and the redox index respectively, and the top principal components that explain 70% to 80% of the cumulative variance are selected as comprehensive variables.
4. The method for quantitative characterization of the main controlling factors of shale organic matter enrichment according to claim 1, characterized in that, The machine learning model is a regression model based on the XGBoost algorithm, where the maximum tree depth is 3 to 10, the learning rate is 0.01 to 0.03, the number of weak learners is 50 to 500, and the sample subsampling ratio is 0.5 to 1.
0.
5. The method for quantitative characterization of the main controlling factors of shale organic matter enrichment according to claim 1, characterized in that, The methods for evaluating the importance of various features include: Statistical analysis methods: Pearson correlation coefficient analysis and ANOVA were used; Information theory analysis method: mutual information regression analysis is used; Model structure-related methods: Calculating decision tree gain values and permutation feature importance based on machine learning models; Model interpretation method: SHAP value analysis is used.
6. The method for quantitative characterization of the main controlling factors of shale organic matter enrichment according to claim 1, characterized in that, In the step of weighted fusion of the results obtained from different feature importance assessment methods, the SHAP value analysis result has a weight of 30%, the permutation feature importance analysis result has a weight of 25%, the ANOVA variance analysis result has a weight of 20%, the decision tree gain analysis result has a weight of 15%, and the mutual information regression analysis result has a weight of 10%.
7. The method for quantitative characterization of the main controlling factors of shale organic matter enrichment according to claim 1, characterized in that, In the step of outputting the quantitative characterization results of the main controlling factors of shale organic matter enrichment, the quantitative characterization results are stored and displayed in the form of data tables, graphical results, or electronic data.
8. A machine learning-based quantitative characterization system for the main controlling factors of shale organic matter enrichment, characterized in that, The method for quantitative characterizing the main controlling factors of shale organic matter enrichment according to any one of claims 1 to 7 includes: The data acquisition module is used to acquire multi-source sedimentological parameters, geochemical parameters, and mineralogical parameters corresponding to shale samples; The parameter filtering and preprocessing module is used to optimize the indicators of the acquired multi-source parameters. The feature dimensionality reduction module is used to perform correlation analysis and feature dimensionality reduction on the optimized index parameters in order to construct an input feature parameter set for machine learning analysis. The machine learning modeling module is used to construct a machine learning model of shale organic matter enrichment based on the input feature parameter set, and to train and validate the model. The feature importance analysis module is used to calculate the influence of each input feature parameter on the enrichment degree of shale organic matter based on multiple feature importance assessment methods; The fusion and output module is used to fuse and calculate the assessment results of the importance of different features, convert them into percentage form, and output the quantitative characterization results of the main controlling factors of shale organic matter enrichment.
9. The quantitative characterization system for the main controlling factors of shale organic matter enrichment according to claim 8, characterized in that, The machine learning model used in the machine learning modeling module is a regression model based on the XGBoost algorithm, where the maximum tree depth is 3 to 10, the learning rate is 0.01 to 0.03, the number of weak learners is 50 to 500, and the sample subsampling ratio is 0.5 to 1.
0.
10. The quantitative characterization system for the main controlling factors of shale organic matter enrichment according to claim 8, characterized in that, In the fusion and output module, the weights for weighted fusion of the importance assessment results of different features are set as follows: SHAP value analysis results have a weight of 30%, permutation feature importance analysis results have a weight of 25%, ANOVA variance analysis results have a weight of 20%, decision tree gain analysis results have a weight of 15%, and mutual information regression analysis results have a weight of 10%.