A shale core TOC prediction method and system based on adaptive double-path output
By constructing a multi-parameter matrix of reservoir properties using an adaptive dual-path output method for TOC prediction of shale cores and combining robust scaling and Huber loss function, the overfitting problem of existing TOC prediction methods in scenarios with small sample sizes or significant lithofacies differences is solved, achieving efficient and accurate TOC prediction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YANGTZE UNIVERSITY
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-21
AI Technical Summary
Existing TOC prediction methods are prone to overfitting or extrapolation distortion in scenarios with small sample size, high noise, or significant lithofacies differences. Furthermore, traditional chemical analysis methods are costly and time-consuming, and cannot quickly output continuous and comparable evaluation results.
An adaptive dual-path output method for predicting TOC in shale cores is adopted. A multi-parameter matrix of reservoir properties is constructed. Through adaptive binning and hierarchical verification and meta-learner training, combined with robust scaling and Huber loss function, local and global weights are constructed to achieve the fusion of the improved stacked integration path and the hierarchical adaptive performance-weighted integration path, and the output result with the higher comprehensive score is selected.
It improves the stability and adaptability of TOC prediction, reduces the subjectivity caused by manual weighting, enhances the ability to represent the changing patterns of TOC, reduces the impact of outliers and noise, and improves the consistency and accuracy of the prediction process.
Smart Images

Figure CN122432465A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unconventional reservoir evaluation technology, and in particular to a method and system for predicting TOC in shale cores based on adaptive dual-path output. Background Technology
[0002] Total organic carbon content, or TOC for short, can assess the hydrocarbon generation capacity of organic-rich rocks and is also the core basis for delineating the sweet spot range of shale reservoirs. It is the traditional way to obtain TOC data, which is mostly completed by chemical analysis and pyrolysis experiments. This type of method has high accuracy, high cost, and long experimental cycle. When the study needs to process a large number of samples or when the scale of the study area is large, the experiment cannot quickly output continuous and comparable evaluation results.
[0003] In existing TOC prediction methods, most approaches focus on simple fitting. These methods do not set physical consistency constraints for multiple physical property parameters. When a single ensemble strategy is applied to scenarios with small sample size, strong noise, or significant differences in lithofacies, it is prone to overfitting or distortion of extrapolation results. Summary of the Invention
[0004] This invention solves the aforementioned technical problems in the prior art by providing a method and system for predicting TOC in shale cores based on adaptive dual-path output.
[0005] This invention provides a method for predicting TOC in shale cores based on adaptive dual-path output, comprising: Construct a multi-parameter matrix of reservoir properties; Using the reservoir property multi-parameter matrix as the sample input matrix, and the corresponding TOC label vector of the sample... As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset; For the Each base model is used to train a sub-model on data that does not include the fold containing the current training sample. It also provides out-of-place predictions for the current training samples. A first-level prediction matrix is constructed from the out-of-range predictions of all base models. ; Combine the reservoir property multi-parameter matrix with the first-level prediction matrix Column concatenation is performed to obtain the meta-learner input matrix. ; Based on this, meta-learner The training objective is defined as ;in, For matrix The OK, Indicates a robust scaling operator. This represents the complexity constraint term of the meta-learner. The regularization coefficient is . For the first Measured TOC values of a training sample of shale cores Huber loss function; The validation samples are input into the base model to obtain the first-level prediction vector. ; The feature vector of the sample to be tested is compared with the first-level prediction vector. The vectors are concatenated to construct the meta-learner input vector for the test sample. ; The vector Input the meta-learner to obtain the output of the improved stacked integration path. ; For each base model and each TOC hierarchical interval Calculate the coefficient of determination of the model in this interval. With root mean square error And construct the original interval score, the calculation formula of which is: ;in, For the first The base model at the th The original score within each TOC stratification interval, subscript For base model index, subscript TOC hierarchical range index; For the first The base model at the th The coefficients of determination on a stratified interval validation set To avoid The score becomes invalid when the value is negative; For the first The base model at the th Root mean square error on a stratified interval validation set This is a local smoothing term; The original scores of all base models within the same TOC stratification interval are normalized to obtain the local weights corresponding to that stratification interval. The calculation formula is as follows: ;in, For the first The base model at the th Local weights within a TOC hierarchical interval, superscript Indicates the local weights of the hierarchy; For the first All within each stratified interval The sum of the original scores of each base model, with subscripts... Use the base model index for summation; Calculate the first value across the entire validation subset. The global original score of each base model is calculated using the following formula: ;in, For the first The global original score of each base model, superscript Indicates the global dimension, subscript Index for the base model; For the first The coefficients of determination of each basic model on the full set of shale core validation sets. To avoid The score becomes invalid when the value is negative; For the first The root mean square error of each basic model on the full shale core validation set; Through formula Normalization yields the global weights; where, For the first The global weights of each base model, superscript Indicates the global dimension, subscript Index for the base model; For all The sum of the global original scores of each base model, with subscripts... Summation is performed using the base model index; The local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ; For any test sample that enters the prediction phase First, determine the corresponding TOC stratification interval based on its primary prediction location. Subsequently, the hierarchical adaptive fusion weights corresponding to that interval are retrieved and invoked. The improved hierarchical adaptive performance-weighted ensemble path then affects the samples. The final output result is ;in, This indicates that the improved hierarchical adaptive performance-weighted ensemble path applies to any test sample. The final predicted TOC value is output. This represents the total number of base models; Index for the base model; Indicates sample The TOC hierarchical interval index according to preliminary estimation; Indicates the first call The base model at the th Hierarchical adaptive fusion weights on each hierarchical interval; Indicates the first The base model for this sample Independent predicted values; The comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path are compared, and the output of the path with the higher comprehensive score is selected as the final TOC prediction result.
[0006] Specifically, the construction of the reservoir property multi-parameter matrix includes: Electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters of shale core samples were obtained. The reservoir property multi-parameter matrix is constructed based on the electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters. ; The step of using the reservoir property multi-parameter matrix as the sample input matrix includes: For the reservoir property multi-parameter matrix The data in the dataset undergoes field organization, robust preprocessing, and enhanced feature construction to form a preprocessing matrix. ; Construct rock physical characteristics, mineral assemblage characteristics, and electrical-porosity-permeability-mineral coupling characteristics, and integrate them with the preprocessing matrix. Concatenate to form an enhanced feature matrix ; The enhanced feature matrix This serves as the sample input matrix.
[0007] Specifically, the structural rock physical characteristics, mineral assemblage characteristics, and electrical-porosity-permeability-mineral coupling characteristics are described, and are combined with the preprocessing matrix. Concatenate to form an enhanced feature matrix ,include: Through formula The wave velocity ratio characteristics were calculated. ;in, Indicates P-wave velocity. Indicates S-wave velocity; Through formula Calculated wave impedance characteristics ;in, Indicates rock density; Through formula Logarithmic characteristics of resistivity were calculated. ;in, Represents resistivity; Through formula The logarithmic characteristic of polarizability was calculated. ;in, Indicates polarizability; Through formula The dynamic elastic modulus characteristics were calculated. ; Through formula The interaction characteristics of resistivity and polarizability were calculated. ; Through formula The interaction characteristics between pyrite content and porosity were calculated. ;in, Indicates pyrite content, Indicates porosity; The wave velocity ratio feature Wave impedance characteristics Logarithmic characteristics of resistivity Logarithmic characteristics of polarizability Dynamic elastic modulus characteristics Interaction characteristics of resistivity and polarizability Interaction characteristics of pyrite content and porosity As a newly added feature column, it is used in conjunction with the preprocessing matrix. The enhanced feature matrix is formed by concatenating the columns. ; The reservoir property multi-parameter matrix is combined with the first-level prediction matrix. Perform column concatenation, including: The enhanced feature matrix With the first-level prediction matrix Perform column splicing.
[0008] Specifically, the local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ,include: Through formula Construct the hierarchical adaptive fusion weights ;in, For the first Local credibility coefficients for each TOC stratified interval.
[0009] Specifically, comparing the comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path, and selecting the output of the path with the higher comprehensive score as the final TOC prediction result, includes: Set path In the The local scores within each TOC interval are: ;in, For the first The path in the first Local scores within a TOC stratified interval, subscript For path index, subscript TOC hierarchical range index; For the first The path in the first The coefficient of determination on the TOC stratified intervals was validated. To avoid The score becomes invalid when the value is negative; For the first The path in the first Root mean square error on the TOC stratified interval validation samples; Based on the sample proportion of each interval The path is obtained by weighting and summing the local scores. Overall rating: ; Through formula The above final TOC prediction results are obtained. ;in, This represents the prediction result output by the improved stacked integration path; This represents the prediction result output by the improved hierarchical adaptive performance-weighted integration path; and These represent the combined scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path under the corresponding TOC hierarchical verification conditions.
[0010] This invention also provides a shale core TOC prediction system based on adaptive dual-path output, comprising: The reservoir property multi-parameter matrix construction module is used to construct the reservoir property multi-parameter matrix. The binning and layering module is used to take the reservoir physical property multi-parameter matrix as the sample input matrix and the corresponding TOC label vector of the sample. As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset; The first-level prediction matrix construction module is used for the first-level prediction matrix construction. Each base model is used to train a sub-model on data that does not include the fold containing the current training sample. It also provides out-of-place predictions for the current training samples. A first-level prediction matrix is constructed from the out-of-range predictions of all base models. ; The matrix splicing module is used to combine the reservoir property multi-parameter matrix with the first-level prediction matrix. Column concatenation is performed to obtain the meta-learner input matrix. ; The meta-learner definition module is used to define meta-learners based on this. The training objective is defined as ;in, For matrix The OK, Indicates a robust scaling operator. This represents the complexity constraint term of the meta-learner. The regularization coefficient is . For the first Measured TOC values of a training sample of shale cores Huber loss function; The first-level prediction vector acquisition module is used to input validation samples into the base model to obtain the first-level prediction vectors. ; The input vector construction module is used to combine the feature vector of the sample to be tested with the first-level prediction vector. The vectors are concatenated to construct the meta-learner input vector for the test sample. ; An improved stacked integrated path output module is used to output the vector Input the meta-learner to obtain the output of the improved stacked integration path. ; The interval raw score calculation module is used for each base model. and each TOC hierarchical interval Calculate the coefficient of determination of the model in this interval. With root mean square error And construct the original interval score, the calculation formula of which is: ;in, For the first The base model at the th The original score within each TOC stratification interval, subscript For base model index, subscript TOC hierarchical range index; For the first The base model at the th The coefficients of determination on a stratified interval validation set To avoid The score becomes invalid when the value is negative; For the first The base model at the th Root mean square error on a stratified interval validation set This is a local smoothing term; The local weight calculation module is used to normalize the original scores of all base models within the same TOC stratification interval to obtain the local weights corresponding to that stratification interval. The calculation formula is as follows: ;in, For the first The base model at the th Local weights within a TOC hierarchical interval, superscript Indicates the local weights of the hierarchy; For the first All within each stratified interval The sum of the original scores of each base model, with subscripts... Use the base model index for summation; The global raw score calculation module is used to calculate the score across the entire validation subset. The global original score of each base model is calculated using the following formula: ;in, For the first The global original score of each base model, superscript Indicates the global dimension, subscript Index for the base model; For the first The coefficients of determination of each basic model on the full set of shale core validation sets. To avoid The score becomes invalid when the value is negative; For the first The root mean square error of each basic model on the full shale core validation set; The global weight calculation module is used to calculate weights using formulas. Normalization yields the global weights; where, For the first The global weights of each base model, superscript Indicates the global dimension, subscript Index for the base model; For all The sum of the global original scores of each base model, with subscripts... Summation is performed using the base model index; The weighted convex combination module is used to combine the local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ; An improved hierarchical adaptive performance-weighted ensemble path output module is used for any test sample entering the prediction stage. First, determine the corresponding TOC stratification interval based on its primary prediction location. Subsequently, the hierarchical adaptive fusion weights corresponding to that interval are retrieved and invoked. The improved hierarchical adaptive performance-weighted ensemble path then affects the samples. The final output result is ;in, This indicates that the improved hierarchical adaptive performance-weighted ensemble path applies to any test sample. The final predicted TOC value is output. This represents the total number of base models; Index for the base model; Indicates sample The TOC hierarchical interval index according to preliminary estimation; Indicates the first call The base model at the th Hierarchical adaptive fusion weights on each hierarchical interval; Indicates the first The base model for this sample Independent predicted values; The final TOC prediction result output module is used to compare the comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path, and select the output result of the path with the higher comprehensive score as the final TOC prediction result.
[0011] Specifically, the reservoir property multi-parameter matrix construction module includes: The parameter acquisition submodule is used to acquire electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters of shale core samples. The reservoir property multi-parameter matrix construction submodule is used to construct the reservoir property multi-parameter matrix based on the electromagnetic parameters, density parameters, elastic parameters, porosity-permeability parameters, and mineral composition parameters. ; The compartmentation and layering module includes: The preprocessing matrix acquisition submodule is used to process the reservoir property multi-parameter matrix. The data in the dataset undergoes field organization, robust preprocessing, and enhanced feature construction to form a preprocessing matrix. ; The enhanced feature matrix acquisition submodule is used to construct rock physics derived features, mineral assemblage features, and electrical-porosity-permeability-mineral coupling features, and to integrate them with the preprocessed matrix. Concatenate to form an enhanced feature matrix ; The binning and layering execution submodule is used to process the enhanced feature matrix. As the sample input matrix, the corresponding TOC label vector of the sample As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset.
[0012] Specifically, the enhanced feature matrix acquisition submodule includes: The wave velocity ratio characteristic calculation unit is used to calculate the wave velocity ratio using the formula The wave velocity ratio characteristics were calculated. ;in, Indicates P-wave velocity. Indicates S-wave velocity; Wave impedance characteristic calculation unit, used to calculate the wave impedance characteristic using formula Calculated wave impedance characteristics ;in, Indicates rock density; The logarithmic resistivity characteristic calculation unit is used to calculate the logarithmic resistivity characteristic using the formula Logarithmic characteristics of resistivity were calculated. ;in, Represents resistivity; The polarizability logarithmic characteristic calculation unit is used to calculate the polarizability logarithmic characteristic using the formula The logarithmic characteristic of polarizability was calculated. ;in, Indicates polarizability; The dynamic elastic modulus characteristic calculation unit is used to calculate the characteristic of the formula. The dynamic elastic modulus characteristics were calculated. ; The unit for calculating the interaction characteristics of resistivity and polarizability is used to calculate the interaction characteristics using the formula... The interaction characteristics of resistivity and polarizability were calculated. ; The interaction characteristic calculation unit of pyrite content and porosity is used to calculate the interaction characteristic of pyrite content and porosity using the formula. The interaction characteristics between pyrite content and porosity were calculated. ;in, Indicates pyrite content, Indicates porosity; Enhanced feature matrix acquisition unit, used to obtain the wave velocity ratio feature Wave impedance characteristics Logarithmic characteristics of resistivity Logarithmic characteristics of polarizability Dynamic elastic modulus characteristics Interaction characteristics of resistivity and polarizability Interaction characteristics of pyrite content and porosity As a newly added feature column, it is used in conjunction with the preprocessing matrix. The enhanced feature matrix is formed by concatenating the columns. ; The matrix concatenation module is specifically used to concatenate the enhanced feature matrix. With the first-level prediction matrix Column concatenation is performed to obtain the meta-learner input matrix. .
[0013] Specifically, the weighted convex combination module is used to use the formula Construct the hierarchical adaptive fusion weights ;in, For the first Local credibility coefficients for each TOC stratified interval.
[0014] Specifically, the final TOC prediction result output module includes: The local score calculation submodule is used to set the path. In the The local scores within each TOC interval are: ;in, For the first The path in the first Local scores within a TOC stratified interval, subscript For path index, subscript TOC hierarchical range index; For the first The path in the first The coefficient of determination on the TOC stratified intervals was validated. To avoid The score becomes invalid when the value is negative; For the first The path in the first Root mean square error on the TOC stratified interval validation samples; The comprehensive score calculation submodule is used to calculate the sample proportion of each interval. The path is obtained by weighting and summing the local scores. Overall rating: ; The final TOC prediction result output submodule is used to output the formula. The above final TOC prediction results are obtained. ;in, This represents the prediction result output by the improved stacked integration path; This represents the prediction result output by the improved hierarchical adaptive performance-weighted integration path; and These represent the combined scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path under the corresponding TOC hierarchical verification conditions.
[0015] One or more technical solutions provided in this invention have at least the following technical effects or advantages: 1. This invention does not rely on lithological pre-grouping, but instead performs adaptive binning and stratification solely around the TOC target value. Under the same stratified validation conditions, it establishes an improved stacked ensemble path and an improved stratified adaptive performance-weighted ensemble path in parallel. The improved stacked ensemble path retains the original enhanced feature pass-through learner in the secondary fusion stage and employs robust loss to suppress anomalous perturbations. The improved stratified adaptive performance-weighted ensemble path simultaneously constructs local and global weights and performs convex combination using local confidence coefficients. Finally, it automatically outputs a superior prediction result based on a unified evaluation criterion. Furthermore, this invention does not rely on lithological pre-grouping during sample partitioning, but instead performs adaptive binning and stratification solely around the TOC target value, and sets dynamic backoff constraints during binning to ensure consistent distribution of the training and validation subsets across the TOC value range. This process reduces the target value distribution shift caused by random partitioning under small sample conditions, thereby improving the consistency between the training and validation processes. This invention retains the original enhanced feature pass-through learner in the stacked fusion path and introduces a second-level learner incorporating robust scaling and a Huber loss function. This allows the second-level fusion process to utilize the first-level prediction results while preserving the original physical information and mitigating the impact of outliers and noise disturbances on the fusion result. In the performance-weighted fusion path, instead of using a single globally pre-set weight, global weights are constructed based on the model fitting performance and error level during the validation phase. Local weights are further generated by combining the TOC (Theory of Consequences) stratification results. The local and global weights are then combined using local confidence coefficients, allowing samples in different TOC intervals to have different model contribution ratios. This process reduces the subjectivity of manual weighting and improves the adaptability of the weighted path to different target value intervals. Furthermore, the convex combination of local and global weights constructs a hierarchical adaptive fusion weight, avoiding excessive weight fluctuations in intervals with small sample sizes caused by using only local weights. This invention improves the stability and adaptability of the shale core TOC prediction process by running an improved stacked integration path and an improved hierarchical adaptive performance-weighted integration path in parallel under the same hierarchical verification conditions, and automatically selecting the final output path according to a unified evaluation criterion. This enables adaptive switching between the two fusion strategies for different data states.
[0016] 2. Based on multiple physical property parameters of shale cores, this invention constructs rock physics derived features and coupled interactive features, so that the input information not only includes the original physical property parameters, but also elastic characterization quantities, electrical parameters and porosity-permeability mineral coupling information related to TOC changes, thereby improving the ability of the input features to characterize the TOC change law and enhancing the physical consistency of the prediction process.
[0017] 3. To enhance the model's ability to characterize the complex nonlinear relationship between multiple physical parameters of shale cores and TOC, this invention further constructs rock physics derived features, mineral assemblage features, and electrical-porosity-permeability-mineral coupling features, and integrates them with the preprocessing matrix. Concatenate to form an enhanced feature matrix Specifically, rock physics derived quantities are constructed using elastic parameters. The P-wave to S-wave velocity ratio reflects the properties and compositional variations of the rock skeleton and has a certain ability to distinguish between organic-rich and non-organic-rich regions, thus constructing a wave velocity ratio characteristic. Wave impedance comprehensively considers rock density and wave velocity information, reflecting the impedance response of the medium to elastic wave propagation, thus constructing a wave impedance characteristic. By performing logarithmic transformations on resistivity and polarizability, logarithmic characteristics of resistivity and polarizability are obtained. This transformation can compress the data dimension span, thereby reducing the impact of outliers on the subsequent modeling process. To further characterize the rock mechanical properties, a dynamic elastic modulus characteristic is constructed. This characteristic comprehensively reflects the rock skeleton stiffness and elastic response characteristics, thereby enhancing the model's ability to identify differences in rock mechanical properties. Simultaneously, to characterize the combined effects between electrical parameters, an interaction characteristic of resistivity and polarizability is constructed. This characteristic is used to characterize the difference in electrical response under the combined action of low-resistivity organic matter and high-polarizability components. In addition, to characterize the coupling relationship between mineral composition and pore structure, an interactive feature of pyrite content and porosity was constructed. This feature is used to characterize the coupling effect between pyrite occurrence state and reservoir pore structure. Attached Figure Description
[0018] Figure 1 The flowchart shows the overall process of the shale core TOC prediction method based on adaptive dual-path output provided in this embodiment of the invention. Figure 2 This is a flowchart illustrating the construction of the improved stacked integrated path in the shale core TOC prediction method based on adaptive dual-path output provided in this embodiment of the invention. Figure 3 The flowchart shows the construction process of the improved hierarchical adaptive performance weighted integration path in the shale core TOC prediction method based on adaptive dual-path output provided in the embodiments of the present invention. Figure 4 A schematic diagram illustrating the principle of the shale core TOC prediction method based on adaptive dual-path output provided in this embodiment of the invention; Figure 5 This is a comparison chart of the TOC prediction results between the conventional multiple linear regression method and the method of this invention. Detailed Implementation
[0019] like Figure 1 As shown, the technical solution in this embodiment of the invention addresses the technical problems existing in the prior art, and the overall process is as follows: Step 1: Establish the reservoir property multi-parameter matrix ; Step 2: Original multi-parameter matrix After the data is established, the sample data undergoes field organization, robust preprocessing, and enhanced feature construction to form an enhanced feature matrix for subsequent modeling. .
[0020] Step 3: Complete data partitioning and stratified sampling. First, read the independent training set and independent test set. The independent test set is only used to judge the final model performance and does not participate in model training or parameter adjustment. Next, do not set grouping for lithology, but only perform adaptive binning of 2 to 5 bins based on the range of TOC percentage values in the training set. Use the binning results as the stratification basis to split the training set into two parts: a training subset and a validation subset. Each bin includes at least 2 samples. The proportion of the validation subset to the overall training set is determined by the current total sample size: when the total sample size exceeds 50, the proportion is 0.15; when the total sample size does not exceed 50, the proportion is 0.10.
[0021] Step 4: Train base models and adjust parameters. Build a base model library to store various models, including LightGBM regression, Random Forest regression, XGBoost regression, Support Vector Regression, Gradient Boosting regression, etc. Each base model is paired with RobustScaler to establish a processing flow, using RandomizedSearchCV to perform random searches on the training set and adjusting the corresponding parameters. The number of folds in cross-validation can be automatically adjusted according to the sample size, with a maximum setting of 5 folds. Select the folds during the cross-validation process. As the basis for parameter selection, the optimal estimator set corresponding to each base model is finally obtained.
[0022] Step 5: Complete parallel integration, extract the optimal base model set obtained in Step 4, and establish the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path respectively. The establishment process of the two types of paths is carried out simultaneously.
[0023] like Figure 2 As shown, the construction process of the improved stacked integration path is as follows: Step 1: Input the multi-property parameter matrix & Actual TOC label measurement ; Step 2: Organize the original sample data, complete robust preprocessing and enhance feature construction; Step 3: Divide the training set and validation set; Step 4: Input the training set into multiple base learners, complete the first-level learning for each base learner, and obtain the first-level prediction results for each base learner.
[0024] Step 5: Input the first-level prediction results of each base learner and the original enhanced features into the meta-learner, so that the second-level fusion stage can retain the original physical information while utilizing the first-level prediction information; the meta-learner preferably adopts a regressor that includes robust scaling and Huber loss function to reduce the impact of abnormal samples and noise perturbation on the second-level fusion results.
[0025] Step 6: The meta-learner outputs the TOC prediction results for the improved stacked integration path.
[0026] Specifically, let the first Each base model for samples The first-level prediction is The first-level prediction matrix is denoted as: Compare the first-level prediction matrix with the original enhanced feature matrix Concatenate the columns to obtain the meta-learner input matrix: The meta-learner outputs the stacked ensemble path prediction result as follows: The meta-learner preferably adopts a gradient boosting regressor that includes robust scaling and Huber loss function, so that the second-level fusion stage can utilize the first-level prediction information while maintaining a strong ability to suppress abnormal samples and noise disturbances.
[0027] like Figure 3 As shown, the construction process of the improved hierarchical adaptive performance-weighted integration path is as follows: Step 1: Input the multi-property parameter matrix & Actual TOC label measurement ; Step 2: Organize the original sample data, complete robust preprocessing and enhance feature construction; Step 3: Divide the training set and validation set; Step 4: Input the training set into each base learner separately and obtain the prediction results output by each base model independently.
[0028] Step 5: Calculate the local scores of each base learner in different TOC hierarchical intervals, and calculate the global scores within the overall validation set to obtain the local weights and global weights.
[0029] Step 6: Based on the number of samples or local confidence coefficients in each TOC stratification interval, perform a convex combination of local and global weights to construct stratified adaptive fusion weights, and output the corresponding weighted prediction results for the samples according to the interval where their predicted TOCs are located.
[0030] Let the first The base model at the th The prediction results within each TOC stratification interval are: The weighted integration path is in the interval The prediction results are as follows: in, The hierarchical adaptive fusion weights are calculated using the steps described above. This approach does not use manually preset weights, but instead directly embeds the TOC stratification results into the weight calculation process, allowing samples in different TOC intervals to have different model contribution ratios.
[0031] Step 6: Automatic Optimization Stage. This invention employs a dynamic optimization output strategy under unified hierarchical validation constraints. Under the same TOC hierarchical validation condition, the prediction performance of the improved stacked ensemble path and the improved hierarchical adaptive performance-weighted ensemble path are statistically analyzed in each TOC interval, and a comprehensive path score is formed by combining the sample proportion of each interval. Specifically, the comprehensive scores of the two paths are compared, and the path with the higher score is selected as the final output path. This process ensures that the final output result is constrained not only by the overall fitting performance but also by the local adaptation capability within different TOC intervals, thereby achieving dynamic path selection based on a unified hierarchical validation framework.
[0032] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0033] like Figure 4 As shown, the shale core TOC prediction method based on adaptive dual-path output provided in this embodiment of the invention includes: Step S1: Construct a multi-parameter matrix of reservoir properties; Step S2: Use the reservoir property multi-parameter matrix as the sample input matrix, and the corresponding TOC label vector of the sample. As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset; This prediction method is explained in detail, and a multi-parameter matrix of reservoir properties is constructed, including: Electromagnetic parameters, density parameters, elastic parameters, porosity-permeability parameters, and mineral composition parameters of shale core samples were obtained. Electromagnetic parameters included resistivity, polarizability, and magnetic susceptibility; density parameters included rock density; elastic parameters included P-wave velocity and S-wave velocity; porosity-permeability parameters included permeability and porosity; and mineral composition parameters included the content of quartz, potassium feldspar, plagioclase, calcite, dolomite, ferrodolithiasis, clay minerals, and pyrite. Specifically, the Autolab1000 laboratory instrument was used in the 1000 frequency band... -2 Hz~10 4Complex resistivity data of shale core samples at different saturations were tested within the Hz range to obtain the amplitude and phase information of complex resistivity. Based on the MGEMTIP model, induced polarization parameters (resistivity and polarizability) of the shale samples were extracted. Simultaneously, the P-wave and S-wave of the shale cores were measured using the Autolab1000 system. Then, using instruments such as a DH-1200 density meter, a KAPPAMETERKM-7 magnetic susceptibility meter, an FYHK-2A core overburden porosity and permeability testing system, and an SCMS-E high-temperature and high-pressure core multi-parameter measurement system, parameters such as density, magnetic susceptibility, permeability, and porosity of the actual rock were measured. Finally, X-ray diffraction analysis was used to obtain mineral composition parameters for quartz, potassium feldspar, plagioclase, calcite, dolomite, ferrodolithite, clay minerals, and pyrite.
[0034] A multi-parameter matrix of reservoir properties is constructed based on electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters. Specifically, the above parameters are organized according to the sample number and combined to form the original reservoir physical property multi-parameter matrix. .
[0035] In this case, the reservoir property multi-parameter matrix is used as the sample input matrix, including: Multi-parameter matrix of reservoir properties The data in the dataset undergoes field organization, robust preprocessing, and enhanced feature construction to form a preprocessing matrix. Original reservoir property multi-parameter matrix The input matrix consists of all samples and their original test parameters; the original features Representation matrix The feature column corresponds to a certain parameter, or represents the value of a sample in that feature dimension. For features... The truncation, scaling, and transformation are all performed on the original multi-parameter matrix. The process is performed column by column on each feature column. The specific processing procedure is as follows: First, for the original multi-parameter matrix All fields in the dataset were organized, and the correlation between each field and the TOC was analyzed. Numbered fields were deleted, and type validation and unit unification were completed for numerical fields. To mitigate the impact of a small number of outliers on the model training process, this invention employs a truncation strategy based on the interquartile range (IQR) to achieve robust outlier handling. Feature truncation upper and lower bounds are defined below the lower bound. and the Upper Realm The calculation is shown in equation (1), where... The cutoff factor is between 1.5 and 3, which can simultaneously satisfy the robustness and information retention requirements. Original features The lower quartile of the corresponding feature column, Original features The upper quartile of the corresponding feature column, The interquartile range represents the difference between the upper and lower quartiles. (1) For original features The clamping operation is completed, and the truncation feature is obtained. The corresponding processing method is shown in equation (2). This processing preserves the original shape of the main distribution and can reduce the bias effect of extreme values on the learner.
[0036] (2) In this embodiment, a cutoff factor is set for TOC(%). The cutoff coefficient is 2, and the other features correspond to the cutoff coefficients. For step 3, after completing the truncation step, use the truncation feature. To further implement robust scaling of the input, it is converted into a dimensionless form, as shown in equation (3): (3) In the formula, Indicates the first A truncated feature column, shows the median, Indicates the interquartile range. Indicates the number after robust scaling. Each feature column is scaled with the median as the center and the interquartile range as the scale, which has higher stability when dealing with non-Gaussian distributed data. After applying equation (3) to each feature column, all transformed feature columns are... Combine them according to the original column order to form a preprocessed matrix. ,Right now It is composed of the original multi-parameter matrix The resulting matrix after robust preprocessing.
[0037] Structural rock physical characteristics, mineral assemblage characteristics, and electrical-porosity-permeability-mineral coupling characteristics, and compared with the preprocessing matrix. Concatenate to form an enhanced feature matrix The enhanced feature matrix It will serve as a unified input for subsequent adaptive binning and hierarchical verification and dual-path integrated modeling.
[0038] Specifically, the rock physical characteristics, mineral assemblage characteristics, and electrical-porosity-permeability-mineral coupling characteristics are derived and compared with the preprocessing matrix. Concatenate to form an enhanced feature matrix ,include: Through formula The wave velocity ratio characteristics were calculated. ;in, Indicates P-wave velocity. Indicates S-wave velocity; Through formula Calculated wave impedance characteristics ;in, Indicates rock density; Through formula Logarithmic characteristics of resistivity were calculated. ;in, Represents resistivity; Through formula The logarithmic characteristic of polarizability was calculated. ;in, Indicates polarizability; Through formula The dynamic elastic modulus characteristics were calculated. ; Through formula The interaction characteristics of resistivity and polarizability were calculated. ; Through formula The interaction characteristics between pyrite content and porosity were calculated. ;in, Indicates pyrite content, Indicates porosity; Wave speed ratio characteristics Wave impedance characteristics Logarithmic characteristics of resistivity Logarithmic characteristics of polarizability Dynamic elastic modulus characteristics Interaction characteristics of resistivity and polarizability Interaction characteristics of pyrite content and porosity As a newly added feature column, it is used in conjunction with the preprocessing matrix. Concatenate columns to form an enhanced feature matrix. ; Enhance the feature matrix As a sample input matrix.
[0039] After completing data processing and feature enhancement construction, the enhanced feature matrix is used... As the sample input matrix, with the corresponding TOC label vector of the sample As a stratification criterion, adaptive binning stratified validation based on the TOC target value is performed. This process ensures that the sample inputs used for subsequent model training and validation maintain a unified enhanced feature representation, while guaranteeing the consistency of the distribution of the training subset and validation subset across the TOC value range.
[0040] To ensure consistency in the distribution of target values between the training and validation sets under small sample conditions, lithology is not pre-grouped; instead, adaptive quantile binning is performed solely based on the measured TOC value range. Let the number of bins be... The initial number of boxes is 2, gradually increasing to... For label vectors Dividing according to the quantile boundaries, we get: in, Indicates the first TOC bins. For any candidate bin number If the following conditions are met: If the binning scheme is deemed valid, it is considered valid; otherwise, it reverts to the previous binning scheme that satisfies the constraints, and that scheme is used as the final stratification basis. When a candidate binning scheme results in insufficient sample size in a local interval, it reverts to the previous binning scheme that satisfies the constraints to ensure statistical comparability in subsequent local weight calculations, dual-path comparisons, and dynamic optimization.
[0041] After determining the final binning scheme, the independent training set is further divided into a training subset and a validation subset. Preferably, when the total sample size is greater than 50, the validation subset accounts for 0.15%; when the total sample size is not greater than 50, the validation subset accounts for 0.10%.
[0042] Next, the base model is trained and its parameters are optimized. The specific process is as follows: Construct a set of base models on a subset of training models: The base model set includes tree models, boosting models, and kernel method models, and parameter optimization is performed through cross-validation-driven parameter search. Preferably, the base models include LightGBM regression models, random forest regression models, XGBoost regression models, support vector regression models, and gradient boosting regression models. Each base model is paired with a robust scaling pipeline for training, and parameter optimization is performed using cross-validation-driven parameter search. For the... The optimal parameters of a basis model are denoted as: in, For the first The parameter space to be searched for in each base model For cross-validation folds, a value no higher than 5 is preferred. The above process yields the optimal estimator set for each base model.
[0043] Step S31: For the first Each base model is used to train a sub-model on data that does not include the fold containing the current training sample. It also provides out-of-place predictions for the current training samples. A first-level prediction matrix is constructed from the out-of-range predictions of all base models. ; This step will be explained in detail. Assume that after robust preprocessing, feature enhancement, and hierarchical partitioning based on the TOC target value, the training subset is obtained as follows: in, Indicates the first Enhanced feature vectors of each sample, This represents the corresponding measured TOC value.
[0044] Let the set of base models after parameter optimization be... To avoid information leakage caused by directly using the fitting results of the first-level model on the same training samples in the second-level learning stage, a training subset is used. implement Cross-validation. For the first... Each base model is trained on data excluding the fold containing the current sample to obtain a sub-model. It also provides out-of-range predicted values for the current sample. The first-level prediction matrix is constructed from the out-of-range prediction results of all base models: .
[0045] Step S32: Combine the reservoir property multi-parameter matrix with the first-level prediction matrix Column concatenation is performed to obtain the meta-learner input matrix. ; This step will be explained in detail, involving the integration of the reservoir property multi-parameter matrix with the first-order prediction matrix. Perform column concatenation, including: Enhance the feature matrix With the first-level prediction matrix To concatenate columns, the expression is: ;in, This indicates a column concatenation operation.
[0046] Step S33: Based on this, the meta-learner The training objective is defined as ;in, For matrix The OK, Indicates a robust scaling operator. This represents the complexity constraint term of the meta-learner. The regularization coefficient is . For the first Measured TOC values of a training sample of shale cores Huber's loss function: ; in, For the first The prediction residual of a sample is defined as follows: ; Subscript For sample index, corresponding to the first sample in the training subset. One shale core sample; For the first Measured TOC values for each sample; For the first TOC predicted values for each sample; The threshold parameter for Huber loss is used to distinguish between small residual squared loss and large residual linear loss, and to suppress outliers and noise disturbances.
[0047] Step S34: Input the validation samples into the base model to obtain the first-level prediction vector. ; This step will be explained in detail after obtaining the meta-learner. Then, using the complete training subset to... Each model is retrained to obtain a set of retrained base models. Entering the prediction application phase, for any sample to be tested... (Sample to be tested) The test set data (after robust preprocessing and feature enhancement) is used to make predictions for each base model after retraining, obtaining the output results of each base model, and then constructing the first-level prediction vector for this single sample. .
[0048] Step S35: Combine the feature vector of the sample to be tested with the first-level prediction vector. The vectors are concatenated to construct the meta-learner input vector for the test samples. ; This step will be explained in detail, involving the enhancement feature vector of the sample to be tested. Its corresponding first-level prediction vector By concatenating the vectors, a meta-learner input vector for that sample can be constructed. Its satisfaction .
[0049] Step S36: Convert the vector Input the meta-learner to obtain the output of the improved stacked integration path. ; Step S41: Under the same training subset and validation subset partitioning conditions as the improved stacked integration path, let the validation subset be: In the performance-weighted integration path, the hierarchical set determined by adaptive binning based on the TOC target value is still used. ,in, The final effective number of boxes, Indicates the first The set of validation samples corresponding to each TOC interval.
[0050] For each base model and each TOC hierarchical interval Calculate the coefficient of determination of the model in this interval. With root mean square error And construct the original interval score, the calculation formula of which is: ;in, For the first The base model at the th The original score within each TOC stratification interval, subscript Base model index subscript For TOC hierarchical range index ; For the first The base model at the th The coefficients of determination on a stratified interval validation set To avoid The score becomes invalid when the value is negative; For the first The base model at the th Root mean square error on a stratified interval validation set For the minimum smoothing term (take...) ), used to avoid numerical anomalies; Step S42: Normalize the original scores of all base models within the same TOC stratification interval to obtain the local weights corresponding to that stratification interval. The calculation formula is as follows: ;in, For the first The base model at the th Local weights within a TOC hierarchical interval, superscript Indicates the local weights of the hierarchy; For the first All within each stratified interval The sum of the original scores of each base model, with subscripts... The base model index is used for summation; normalization is used to ensure that the sum of the weights of the base models in the same layer is 1, thereby achieving differentiated weighted fusion for layer adaptation and improving the accuracy and robustness of TOC prediction for shale cores.
[0051] Step S43: Calculate the first verdict within the entire validation subset. The global original score of each base model is calculated using the following formula: ;in, For the first The global original score of each base model, superscript Indicates the global dimension, subscript Base model index ; For the first The coefficients of determination of each basic model on the full set of shale core validation sets. To avoid The score becomes invalid when the value is negative; For the first The root mean square error of each basic model on the full shale core validation set; Step S44: Using the formula Normalization yields the global weights; where, For the first The global weights of each base model, superscript Indicates the global dimension, subscript Index for the base model; For all The sum of the global original scores of each base model, with subscripts... The summation is indexed by the base model; normalization ensures that the sum of the global weights is 1, which can be fused with the local weights to obtain the final fused weights.
[0052] Step S45: Adjust local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ; This step will be explained in detail, including the local weights. With global weight Perform convex combination to construct hierarchical adaptive fusion weights ,include: Through formula Constructing hierarchical adaptive fusion weights ;in, For the first Local credibility coefficients for each TOC stratified interval. With the number of samples in the interval The weight increases with the increase of the number of samples, so that intervals with more samples rely more on local weights, while intervals with fewer samples retain more global weight constraints.
[0053] Step S46: For any test sample entering the prediction phase (Sample to be tested) For the test set data after feature enhancement, the corresponding TOC stratification interval is first determined based on its primary prediction position (i.e., the preliminary predicted distribution of TOC values by each base model). Subsequently, the hierarchical adaptive fusion weights corresponding to that interval are retrieved and invoked. The improved hierarchical adaptive performance-weighted ensemble path then affects the samples. The final output result is ;in, This indicates that the improved hierarchical adaptive performance-weighted ensemble path applies to any test sample. The final predicted TOC value is output. This represents the total number of base models; Index for the base model; Indicates sample The TOC hierarchical interval index according to preliminary estimation; Indicates the first call The base model at the th Hierarchical adaptive fusion weights on each hierarchical interval; Indicates the first The base model for this sample Independent predicted values; If represented in the form of a sample set, then for the th All samples within each TOC stratification interval are: in, Indicates the weighted integration path at the th The overall prediction vector of all samples within each TOC stratification interval Indicates the first The base model for this first... The predicted vector of all samples within each stratified interval. Indicates the first The base model at the th The corresponding hierarchical adaptive fusion weight scalar on each hierarchical interval.
[0054] Step S5: Compare the comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path, and select the output of the path with the higher comprehensive score as the final TOC prediction result.
[0055] This step is explained in detail. The comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path are compared. The output of the path with the higher comprehensive score is selected as the final TOC prediction result, including: Set path In the The local scores within each TOC interval are: ;in, For the first The path in the first Local scores within a TOC stratified interval, subscript For path index, subscript TOC hierarchical range index; For the first The path in the first The coefficient of determination on the TOC stratified intervals was validated. To avoid The score becomes invalid when the value is negative; For the first The path in the first Root mean square error on the TOC stratified interval validation samples; Based on the sample proportion of each interval The path is obtained by weighting and summing the local scores. Overall rating: ; Through formula The above final TOC prediction results are obtained. ;in, This indicates the predicted output of the improved stacked integration path; This represents the prediction result output by the improved hierarchical adaptive performance-weighted ensemble path; and These represent the combined scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path under the corresponding TOC hierarchical verification conditions, respectively.
[0056] Furthermore, this invention also provides a shale core TOC prediction system based on adaptive dual-path output, comprising: The reservoir property multi-parameter matrix construction module is used to construct the reservoir property multi-parameter matrix. Specifically, the reservoir property multi-parameter matrix construction module includes: The parameter acquisition submodule is used to acquire electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters of shale core samples. Among them, electromagnetic parameters include resistivity, polarizability, and magnetic susceptibility; density parameters include rock density; elastic parameters include P-wave velocity and S-wave velocity; porosity and permeability parameters include permeability and porosity; and mineral composition parameters include the content of quartz, potassium feldspar, plagioclase, calcite, dolomite, ferrodolithite, clay minerals, and pyrite.
[0057] The reservoir property multi-parameter matrix construction submodule is used to construct a reservoir property multi-parameter matrix based on electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters. ; The binning and layering module is used to take the reservoir physical property multi-parameter matrix as the sample input matrix and the corresponding TOC label vector of the sample. As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset; Specifically, the compartmentalized and layered modules include: The preprocessing matrix acquisition submodule is used for processing multi-parameter matrices of reservoir properties. The data in the dataset undergoes field organization, robust preprocessing, and enhanced feature construction to form a preprocessing matrix. Original reservoir physical property multi-parameter matrix The input matrix consists of all samples and their original test parameters; the original features Representation matrix The feature column corresponds to a certain parameter, or represents the value of a sample in that feature dimension. For features... The truncation, scaling, and transformation are all performed on the original multi-parameter matrix. Implemented column by column on each feature column; The enhanced feature matrix acquisition submodule is used to construct rock physics derived features, mineral assemblage features, and electrical-porosity-permeability-mineral coupling features, and to integrate them with the preprocessed matrix. Concatenate to form an enhanced feature matrix ; The binning and layering execution submodule is used to enhance the feature matrix. As the sample input matrix, with the corresponding TOC label vector of the sample As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset.
[0058] The first-level prediction matrix construction module is used for the first-level prediction matrix construction. Each base model is used to train a sub-model on data that does not include the fold containing the current training sample. It also provides out-of-place predictions for the current training samples. A first-level prediction matrix is constructed from the out-of-range predictions of all base models. ; The matrix stitching module is used to combine the reservoir property multi-parameter matrix with the first-level prediction matrix. Column concatenation is performed to obtain the meta-learner input matrix. ; Specifically, the enhanced feature matrix acquisition submodule includes: The wave velocity ratio characteristic calculation unit is used to calculate the wave velocity ratio using the formula The wave velocity ratio characteristics were calculated. ;in, Indicates P-wave velocity. Indicates S-wave velocity; Wave impedance characteristic calculation unit, used to calculate the wave impedance characteristic using formula Calculated wave impedance characteristics ;in, Indicates rock density; The logarithmic resistivity characteristic calculation unit is used to calculate the logarithmic resistivity characteristic using the formula Logarithmic characteristics of resistivity were calculated. ;in, Represents resistivity; The polarizability logarithmic characteristic calculation unit is used to calculate the polarizability logarithmic characteristic using the formula The logarithmic characteristic of polarizability was calculated. ;in, Indicates polarizability; The dynamic elastic modulus characteristic calculation unit is used to calculate the characteristic of the formula. The dynamic elastic modulus characteristics were calculated. ; The unit for calculating the interaction characteristics of resistivity and polarizability is used to calculate the interaction characteristics using the formula... The interaction characteristics of resistivity and polarizability were calculated. ; The interaction characteristic calculation unit of pyrite content and porosity is used to calculate the interaction characteristic of pyrite content and porosity using the formula. The interaction characteristics between pyrite content and porosity were calculated. ;in, Indicates pyrite content, Indicates porosity; Enhanced feature matrix acquisition unit, used to obtain wave velocity ratio features Wave impedance characteristics Logarithmic characteristics of resistivity Logarithmic characteristics of polarizability Dynamic elastic modulus characteristics Interaction characteristics of resistivity and polarizability Interaction characteristics of pyrite content and porosity As a newly added feature column, it is used in conjunction with the preprocessing matrix. Concatenate columns to form an enhanced feature matrix. .
[0059] In this case, the matrix concatenation module is specifically used to combine the enhanced feature matrix. With the first-level prediction matrix Column concatenation is performed to obtain the meta-learner input matrix. The expression is ;in, This indicates a column concatenation operation.
[0060] The meta-learner definition module is used to define meta-learners based on this. The training objective is defined as ;in, For matrix The OK, Indicates a robust scaling operator. This represents the complexity constraint term of the meta-learner. The regularization coefficient is . For the first Measured TOC values of a training sample of shale cores Huber's loss function: ;in, For the first The prediction residual of a sample is defined as follows: ; Subscript For sample index, corresponding to the first sample in the training subset. One shale core sample; For the first Measured TOC values for each sample; For the first TOC predicted values for each sample; The threshold parameter for Huber loss is used to distinguish between small residual squared loss and large residual linear loss, and to suppress outliers and noise disturbances.
[0061] The first-level prediction vector acquisition module is used to input validation samples into the base model to obtain the first-level prediction vectors. ; Specifically, the first-level prediction vector acquisition module is used to utilize the complete training subset to obtain the vectors. Each model is retrained to obtain a set of retrained base models. Entering the prediction application phase, for any sample to be tested... (Sample to be tested) The test set data (after robust preprocessing and feature enhancement) is used to make predictions for each base model after retraining, obtaining the output results of each base model, and then constructing the first-level prediction vector for this single sample. .
[0062] The input vector construction module is used to combine the feature vector of the test sample with the first-level prediction vector. The vectors are concatenated to construct the meta-learner input vector for the test samples. ; Specifically, the input vector construction module is used to construct the enhanced feature vector of the sample to be tested. Its corresponding first-level prediction vector By concatenating the vectors, a meta-learner input vector for that sample can be constructed. Its satisfaction .
[0063] An improved stacked integrated path output module is used to convert vectors Input the meta-learner to obtain the output of the improved stacked integration path. ; The interval raw score calculation module is used for each base model. and each TOC hierarchical interval Calculate the coefficient of determination of the model in this interval. With root mean square error And construct the original interval score, the calculation formula of which is: ;in, For the first The base model at the th The original score within each TOC stratification interval, subscript Base model index subscript For TOC hierarchical range index ; For the first The base model at the th The coefficients of determination on a stratified interval validation set To avoid The score becomes invalid when the value is negative; For the first The base model at the th Root mean square error on a stratified interval validation set For the minimum smoothing term (take...) ), used to avoid numerical anomalies; The local weight calculation module is used to normalize the original scores of all base models within the same TOC stratification interval to obtain the local weights corresponding to that stratification interval. The calculation formula is as follows: ;in, For the first The base model at the th Local weights within a TOC hierarchical interval, superscript Indicates the local weights of the hierarchy; For the first All within each stratified interval The sum of the original scores of each base model, with subscripts... The base model index is used for summation; normalization is used to ensure that the sum of the weights of the base models in the same layer is 1, thereby achieving differentiated weighted fusion for layer adaptation and improving the accuracy and robustness of TOC prediction for shale cores.
[0064] The global raw score calculation module is used to calculate the score across the entire validation subset. The global original score of each base model is calculated using the following formula: ;in, For the first The global original score of each base model, superscript Indicates the global dimension, subscript Base model index ; For the first The coefficients of determination of each basic model on the full set of shale core validation sets. To avoid The score becomes invalid when the value is negative; For the first The root mean square error of each basic model on the full shale core validation set; The global weight calculation module is used to calculate weights using formulas. Normalization yields the global weights; where, For the first The global weights of each base model, superscript Indicates the global dimension, subscript Index for the base model; For all The sum of the global original scores of each base model, with subscripts... The summation is indexed by the base model; normalization ensures that the sum of the global weights is 1, which can be fused with the local weights to obtain the final fused weights.
[0065] The weighted convex combination module is used to combine local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ; Specifically, the weighted convex combination module is used to... (The sentence is incomplete and requires more context to translate accurately.) Constructing hierarchical adaptive fusion weights ;in, For the first Local credibility coefficients for each TOC stratified interval. With the number of samples in the interval The weight increases with the increase of the number of samples, so that intervals with more samples rely more on local weights, while intervals with fewer samples retain more global weight constraints.
[0066] An improved hierarchical adaptive performance-weighted ensemble path output module is used for any test sample entering the prediction stage. (Sample to be tested) For the test set data after feature enhancement, the corresponding TOC stratification interval is first determined based on its primary prediction position (i.e., the preliminary predicted distribution of TOC values by each base model). Subsequently, the hierarchical adaptive fusion weights corresponding to that interval are retrieved and invoked. The improved hierarchical adaptive performance-weighted ensemble path then affects the samples. The final output result is ;in, This indicates that the improved hierarchical adaptive performance-weighted ensemble path applies to any test sample. The final predicted TOC value is output. This represents the total number of base models; Index for the base model; Indicates sample The TOC hierarchical interval index according to preliminary estimation; Indicates the first call The base model at the th Hierarchical adaptive fusion weights on each hierarchical interval; Indicates the first The base model for this sample Independent predicted values; The final TOC prediction output module compares the comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path, and selects the output of the path with the higher comprehensive score as the final TOC prediction result.
[0067] Specifically, the final TOC prediction result output module includes: The local score calculation submodule is used to set the path. In the The local scores within each TOC interval are: ;in, For the first The path in the first Local scores within a TOC stratified interval, subscript For path index, subscript TOC hierarchical range index; For the first The path in the first The coefficient of determination on the TOC stratified intervals was validated. To avoid The score becomes invalid when the value is negative; For the first The path in the first Root mean square error on the TOC stratified interval validation samples; The comprehensive score calculation submodule is used to calculate the sample proportion of each interval. The path is obtained by weighting and summing the local scores. Overall rating: ; The final TOC prediction result output submodule is used to output the formula. The above final TOC prediction results are obtained. ;in, This indicates the predicted output of the improved stacked integration path; This represents the prediction result output by the improved hierarchical adaptive performance-weighted ensemble path; and These represent the combined scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path under the corresponding TOC hierarchical verification conditions, respectively.
[0068] The present invention will be described below with reference to an embodiment, but the scope of protection of the present invention is not limited to this embodiment.
[0069] Shale core samples were selected as the research object. Multiple physical property parameters of the core samples were obtained, with the measured total organic carbon (TOC) content used as a monitoring label. These multiple physical property parameters included rock density, P-wave velocity, S-wave velocity, resistivity, polarizability, porosity, permeability on a logarithmic scale, and mineral composition parameters, including quartz, potassium feldspar, plagioclase, calcite, dolomite, ferrodolithite, clay minerals, and pyrite. These parameters collectively constituted the original multi-parameter matrix. The label vector is formed by the measured TOC percentage of the corresponding samples. .
[0070] In the data preprocessing stage, the original sample table is first reorganized, numbered fields are deleted, numerical fields are validated and units are standardized, and the Spearman rank correlation coefficient matrix between TOC and each input feature is calculated to aid in subsequent feature enhancement construction and data verification. Subsequently, truncation based on interquartile range (IQR) is applied to both the original features and labels to reduce the impact of a small number of outliers on the model training process; for electrical parameters with significant skewness... Logarithmic transformation; robust scaling is applied to the truncated features to transform them into a dimensionless form suitable for subsequent modeling.
[0071] In the feature enhancement construction stage, based on the petrophysical mechanisms of shale reservoirs, derived features and coupled interactive features are constructed from the preprocessed multi-property parameters. Based on rock density, P-wave velocity, and S-wave velocity, P-wave to S-wave velocity ratio, wave impedance, elastic modulus, and strength characterization terms are constructed; based on resistivity, polarizability, porosity, permeability, and mineral composition, electrical-mineral coupling terms and porosity-permeability coupling terms are constructed. By concatenating these derived features and coupled interactive features with the preprocessed original features, an enhanced feature matrix is formed. .
[0072] During the sample partitioning phase, instead of pre-grouping based on lithology, adaptive binning and stratification are performed solely based on the TOC target value. Specifically, starting with 2 bins and gradually increasing the number of bins based on the TOC value range in the training set, the maximum number of bins is set to 5. When the number of samples in any bin of a candidate binning scheme falls below a preset threshold, the process reverts to the previous binning scheme that satisfies the constraints, and this scheme is used as the final stratification basis for the training and validation subsets. The proportion of the validation subset is determined based on the total sample size: 0.15 when the total sample size exceeds 50, and 0.10 when the total sample size does not exceed 50. This process ensures that the distribution of the training and validation subsets remains consistent across the TOC value range.
[0073] During the base model training phase, a set of base models is constructed, including tree models, boosting models, and kernel method models. Parameter optimization is performed through a cross-validation-driven parameter search method. Preferably, the base models include LightGBM regression models, random forest regression models, XGBoost regression models, support vector regression models, and gradient boosting regression models. Each base model is integrated with a robust scaling process to form a unified training pipeline, and parameter optimization is performed on a training subset to obtain the corresponding optimal estimator set.
[0074] In the improved stacked ensemble path, first-level prediction results are generated on the training subset using the optimal set of base models. To avoid information leakage caused by directly using the fitting results of the first-level models on the same training samples in the second-level learning stage, cross-validation is performed on the training subset to obtain the out-of-place prediction results of each base model, and the first-level prediction matrix is constructed from these results. Then, the first-level prediction matrix is concatenated column-wise with the original enhanced feature matrix as the input to the meta-learner. The meta-learner employs a regressor that includes robust scaling and a robust loss function, preferably a gradient boosting regressor with Huber loss, to reduce the impact of outlier samples and noise perturbations on the second-level fusion results. This path, by directly connecting the original enhanced features to the meta-learner, utilizes the complementary information of the first-level models while preserving the original physical information, ultimately outputting the prediction results of the improved stacked ensemble path.
[0075] In the improved hierarchical adaptive performance-weighted ensemble path, the goodness-of-fit index and error index of each base model are first calculated on the validation subset, and a global raw score and global normalized weights are generated accordingly. Then, combining the aforementioned TOC hierarchical results, the prediction performance of each base model within each TOC interval is statistically analyzed, further generating interval local scores and local normalized weights. To avoid excessive weight fluctuations due to insufficient local sample size, local weights are not directly used; instead, a convex combination of local and global weights is performed using local confidence coefficients to obtain the final hierarchical adaptive weights. TOC intervals with sufficient sample size rely more on local weights, while TOC intervals with fewer sample size retain more global weight constraints. Based on this, samples in different TOC intervals are weighted and fused using the corresponding hierarchical adaptive weights according to their respective intervals. This transforms the weighting path from a single global weighting into an interval-based weighting path that adaptively adjusts the model contribution ratio according to changes in the target value distribution.
[0076] In the dynamic optimization phase, under the same TOC hierarchical validation condition, the prediction performance of the improved stacked ensemble path and the improved hierarchical adaptive performance-weighted ensemble path are statistically analyzed in each TOC interval, and a comprehensive path score is formed by combining the sample proportion of each interval. The comprehensive scores of the two paths are directly compared, and the one with the higher score is selected as the final output path. Through this process, the final output result is constrained not only by the overall fitting performance but also by the local adaptation capability in different TOC intervals, thereby realizing dynamic path selection based on a unified hierarchical validation framework.
[0077] In the testing and results output phase, the independent test set is input into the final output model to obtain the TOC prediction results of shale core samples, and the coefficient of determination of the model on the test set is calculated. Evaluation metrics include root mean square error (RMSE). The system also outputs model performance comparison charts, TOC prediction scatter plots, and TOC prediction curves. Prediction results are stored in CSV format for comparative analysis of different base models, different integration paths, and the final output results, providing support for subsequent rock sample evaluation.
[0078] like Figure 5 As shown, when conventional multiple linear regression and the method of this invention are applied to TOC prediction, their output results differ. The figures indicate that the coefficient of determination (COD) of the ensemble learning method used in this invention is 0.8677, while that of the conventional multiple linear regression method is 0.8035. The former has a higher COD, indicating that the ensemble learning method fits the ideal prediction state better, better characterizes the nonlinear relationship between multiple physical parameters and TOC, and has higher prediction accuracy and stronger stability. This embodiment demonstrates that this invention can achieve adaptive prediction output for different TOC intervals under multiple physical parameter conditions in shale cores, and is applicable to quantitative TOC prediction of shale cores under conditions of small samples, uneven distribution of target values, and noise disturbance.
[0079] In summary, this invention provides a TOC prediction method for shale core samples based on TOC target hierarchical verification and an improved dual-path adaptive output. This method takes multiple physical property parameters of shale core samples as input and the measured total organic carbon (TOC) content as the supervision label. It sequentially completes the following steps: original multi-parameter matrix construction, robust preprocessing, construction of rock physics derived features and coupled interactive features, adaptive binning and hierarchical division based on TOC target values, base model training and parameter optimization, construction of an improved stacked ensemble path, construction of an improved hierarchical adaptive performance-weighted ensemble path, and dual-path parallel ensemble and dynamic optimization output. This invention focuses on rock physics mechanisms, processes stable data, and incorporates an adaptive ensemble learning method to achieve TOC prediction. This invention allows the prediction process to adapt to different lithofacies environments and parameter patterns while maintaining its stability, and also possesses interpretability and engineering usability.
[0080] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0081] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0082] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0084] Any aspects of this invention not described in detail in the embodiments are well-known techniques to those skilled in the art. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this invention and not to limit it. Although this invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this invention without departing from the spirit and scope of this invention, and all such modifications and substitutions should be covered within the scope of the claims of this invention.
Claims
1. A method for predicting TOC (Total Organic Carbon) in shale cores based on adaptive dual-path output, characterized in that, include: Construct a multi-parameter matrix of reservoir properties; Using the reservoir property multi-parameter matrix as the sample input matrix, and the corresponding TOC label vector of the sample... As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset; For the Each base model is used to train a sub-model on data that does not include the fold containing the current training sample. It also provides out-of-place predictions for the current training samples. A first-level prediction matrix is constructed from the out-of-range predictions of all base models. ; Combine the reservoir property multi-parameter matrix with the first-level prediction matrix Column concatenation is performed to obtain the meta-learner input matrix. ; Based on this, meta-learner The training objective is defined as ;in, For matrix The OK, Indicates a robust scaling operator. This represents the complexity constraint term of the meta-learner. The regularization coefficient is . For the first Measured TOC values of a training sample of shale cores Huber loss function; The validation samples are input into the base model to obtain the first-level prediction vector. ; The feature vector of the sample to be tested is compared with the first-level prediction vector. The vectors are concatenated to construct the meta-learner input vector for the test sample. ; The vector Input the meta-learner to obtain the output of the improved stacked integration path. ; For each base model and each TOC hierarchical interval Calculate the coefficient of determination of the model in this interval. With root mean square error And construct the original interval score, the calculation formula of which is: ;in, For the first The base model at the th The original score within each TOC stratification interval, subscript For base model index, subscript TOC hierarchical range index; For the first The base model at the th The coefficients of determination on a stratified interval validation set To avoid The score becomes invalid when the value is negative; For the first The base model at the th Root mean square error on a stratified interval validation set This is a local smoothing term; The original scores of all base models within the same TOC stratification interval are normalized to obtain the local weights corresponding to that stratification interval. The calculation formula is as follows: ;in, For the first The base model at the th Local weights within a TOC hierarchical interval, superscript Indicates the local weights of the hierarchy; For the first All within each stratified interval The sum of the original scores of each base model, with subscripts... Use the base model index for summation; Calculate the first value across the entire validation subset. The global original score of each base model is calculated using the following formula: ;in, For the first The global original score of each base model, superscript Indicates the global dimension, subscript Index for the base model; For the first The coefficients of determination of each basic model on the full set of shale core validation sets. To avoid The score becomes invalid when the value is negative; For the first The root mean square error of each basic model on the full shale core validation set; Through formula Normalization yields the global weights; where, For the first The global weights of each base model, superscript Indicates the global dimension, subscript Index for the base model; For all The sum of the global original scores of each base model, with subscripts... Summation is performed using the base model index; The local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ; For any test sample that enters the prediction phase First, determine the corresponding TOC stratification interval based on its primary prediction location. Subsequently, the hierarchical adaptive fusion weights corresponding to that interval are retrieved and invoked. The improved hierarchical adaptive performance-weighted ensemble path then affects the samples. The final output result is ;in, This indicates that the improved hierarchical adaptive performance-weighted ensemble path applies to any test sample. The final predicted TOC value is output. This represents the total number of base models; Index for the base model; Indicates sample The TOC hierarchical interval index according to preliminary estimation; Indicates the first call The base model at the th Hierarchical adaptive fusion weights on each hierarchical interval; Indicates the first The base model for this sample Independent predicted values; The comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path are compared, and the output of the path with the higher comprehensive score is selected as the final TOC prediction result.
2. The shale core TOC prediction method based on adaptive dual-path output as described in claim 1, characterized in that, The construction of the reservoir property multi-parameter matrix includes: Electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters of shale core samples were obtained. The reservoir property multi-parameter matrix is constructed based on the electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters. ; The step of using the reservoir property multi-parameter matrix as the sample input matrix includes: For the reservoir property multi-parameter matrix The data in the dataset undergoes field organization, robust preprocessing, and enhanced feature construction to form a preprocessing matrix. ; Construct rock physical characteristics, mineral assemblage characteristics, and electrical-porosity-permeability-mineral coupling characteristics, and integrate them with the preprocessing matrix. Concatenate to form an enhanced feature matrix ; The enhanced feature matrix This serves as the sample input matrix.
3. The shale core TOC prediction method based on adaptive dual-path output as described in claim 2, characterized in that, The structural rock physical characteristics, mineral assemblage characteristics, and electrical-porosity-mineral coupling characteristics are described, and are compared with the preprocessing matrix. Concatenate to form an enhanced feature matrix ,include: Through formula The wave velocity ratio characteristics were calculated. ;in, Indicates P-wave velocity. Indicates S-wave velocity; Through formula Calculated wave impedance characteristics ;in, Indicates rock density; Through formula Logarithmic characteristics of resistivity were calculated. ;in, Represents resistivity; Through formula The logarithmic characteristic of polarizability was calculated. ;in, Indicates polarizability; Through formula The dynamic elastic modulus characteristics were calculated. ; Through formula The interaction characteristics of resistivity and polarizability were calculated. ; Through formula The interaction characteristics between pyrite content and porosity were calculated. ;in, Indicates pyrite content, Indicates porosity; The wave velocity ratio feature Wave impedance characteristics Logarithmic characteristics of resistivity Logarithmic characteristics of polarizability Dynamic elastic modulus characteristics Interaction characteristics of resistivity and polarizability Interaction characteristics of pyrite content and porosity As a newly added feature column, it is used in conjunction with the preprocessing matrix. The enhanced feature matrix is formed by concatenating the columns. ; The reservoir property multi-parameter matrix is combined with the first-level prediction matrix. Perform column concatenation, including: The enhanced feature matrix With the first-level prediction matrix Perform column splicing.
4. The shale core TOC prediction method based on adaptive dual-path output as described in claim 1, characterized in that, The local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ,include: Through formula Construct the hierarchical adaptive fusion weights ;in, For the first Local credibility coefficients for each TOC stratified interval.
5. The shale core TOC prediction method based on adaptive dual-path output as described in any one of claims 1-4, characterized in that, The comparison of the comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path, and the selection of the path with the higher comprehensive score as the final TOC prediction result, includes: Set path In the The local scores within each TOC interval are: ;in, For the first The path in the first Local scores within a TOC stratified interval, subscript For path index, subscript TOC hierarchical range index; For the first The path in the first The coefficient of determination on the TOC stratified intervals was validated. To avoid The score becomes invalid when the value is negative; For the first The path in the first Root mean square error on the TOC stratified interval validation samples; Based on the sample proportion of each interval The path is obtained by weighting and summing the local scores. Overall rating: ; Through formula The above final TOC prediction results are obtained. ;in, This represents the prediction result output by the improved stacked integration path; This represents the prediction result output by the improved hierarchical adaptive performance-weighted integration path; and These represent the combined scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path under the corresponding TOC hierarchical verification conditions.
6. A shale core TOC prediction system based on adaptive dual-path output, characterized in that, include: The reservoir property multi-parameter matrix construction module is used to construct the reservoir property multi-parameter matrix. The binning and layering module is used to take the reservoir physical property multi-parameter matrix as the sample input matrix and the corresponding TOC label vector of the sample. As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset; The first-level prediction matrix construction module is used for the first-level prediction matrix construction. Each base model is used to train a sub-model on data that does not include the fold containing the current training sample. It also provides out-of-place predictions for the current training samples. A first-level prediction matrix is constructed from the out-of-range predictions of all base models. ; The matrix splicing module is used to combine the reservoir property multi-parameter matrix with the first-level prediction matrix. Column concatenation is performed to obtain the meta-learner input matrix. ; The meta-learner definition module is used to define meta-learners based on this. The training objective is defined as ;in, For matrix The OK, Indicates a robust scaling operator. This represents the complexity constraint term of the meta-learner. The regularization coefficient is . For the first Measured TOC values of a training sample of shale cores Huber loss function; The first-level prediction vector acquisition module is used to input validation samples into the base model to obtain the first-level prediction vectors. ; The input vector construction module is used to combine the feature vector of the sample to be tested with the first-level prediction vector. The vectors are concatenated to construct the meta-learner input vector for the test sample. ; An improved stacked integrated path output module is used to output the vector Input the meta-learner to obtain the output of the improved stacked integration path. ; The interval raw score calculation module is used for each base model. and each TOC hierarchical interval Calculate the coefficient of determination of the model in this interval. With root mean square error And construct the original interval score, the calculation formula of which is: ;in, For the first The base model at the th The original score within each TOC stratification interval, subscript For base model index, subscript TOC hierarchical range index; For the first The base model at the th The coefficients of determination on a stratified interval validation set To avoid The score becomes invalid when the value is negative; For the first The base model at the th Root mean square error on a stratified interval validation set This is a local smoothing term; The local weight calculation module is used to normalize the original scores of all base models within the same TOC stratification interval to obtain the local weights corresponding to that stratification interval. The calculation formula is as follows: ;in, For the first The base model at the th Local weights within a TOC hierarchical interval, superscript Indicates the local weights of the hierarchy; For the first All within each stratified interval The sum of the original scores of each base model, with subscripts... Use the base model index for summation; The global raw score calculation module is used to calculate the score across the entire validation subset. The global original score of each base model is calculated using the following formula: ;in, For the first The global original score of each base model, superscript Indicates the global dimension, subscript Index for the base model; For the first The coefficients of determination of each basic model on the full set of shale core validation sets. To avoid The score becomes invalid when the value is negative; For the first The root mean square error of each basic model on the full shale core validation set; The global weight calculation module is used to calculate weights using formulas. Normalization yields the global weights; where, For the first The global weights of each base model, superscript Indicates the global dimension, subscript Index for the base model; For all The sum of the global original scores of each base model, with subscripts... Summation is performed using the base model index; The weighted convex combination module is used to combine the local weights With global weight Perform convex combination to construct hierarchical adaptive fusion weights ; An improved hierarchical adaptive performance-weighted ensemble path output module is used for any test sample entering the prediction stage. First, determine the corresponding TOC stratification interval based on its primary prediction location. Subsequently, the hierarchical adaptive fusion weights corresponding to that interval are retrieved and invoked. The improved hierarchical adaptive performance-weighted ensemble path then affects the samples. The final output result is ;in, This indicates that the improved hierarchical adaptive performance-weighted ensemble path applies to any test sample. The final predicted TOC value is output. This represents the total number of base models; Index for the base model; Indicates sample The TOC hierarchical interval index according to preliminary estimation; Indicates the first call The base model at the th Hierarchical adaptive fusion weights on each hierarchical interval; Indicates the first The base model for this sample Independent predicted values; The final TOC prediction result output module is used to compare the comprehensive scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path, and select the output result of the path with the higher comprehensive score as the final TOC prediction result.
7. The shale core TOC prediction system based on adaptive dual-path output as described in claim 6, characterized in that, The reservoir property multi-parameter matrix construction module includes: The parameter acquisition submodule is used to acquire electromagnetic parameters, density parameters, elastic parameters, porosity and permeability parameters, and mineral composition parameters of shale core samples. The reservoir property multi-parameter matrix construction submodule is used to construct the reservoir property multi-parameter matrix based on the electromagnetic parameters, density parameters, elastic parameters, porosity-permeability parameters, and mineral composition parameters. ; The compartmentation and layering module includes: The preprocessing matrix acquisition submodule is used to process the reservoir property multi-parameter matrix. The data in the dataset undergoes field organization, robust preprocessing, and enhanced feature construction to form a preprocessing matrix. ; The enhanced feature matrix acquisition submodule is used to construct rock physics derived features, mineral assemblage features, and electrical-porosity-permeability-mineral coupling features, and to integrate them with the preprocessed matrix. Concatenate to form an enhanced feature matrix ; The binning and layering execution submodule is used to process the enhanced feature matrix. As the sample input matrix, the corresponding TOC label vector of the sample As a basis for stratification, adaptive binning stratification validation based on TOC target value is performed to determine the sample binning scheme and divide the training subset and validation subset.
8. The shale core TOC prediction system based on adaptive dual-path output as described in claim 7, characterized in that, The enhanced feature matrix acquisition submodule includes: The wave velocity ratio characteristic calculation unit is used to calculate the wave velocity ratio using the formula The wave velocity ratio characteristics were calculated. ;in, Indicates P-wave velocity. Indicates S-wave velocity; Wave impedance characteristic calculation unit, used to calculate the wave impedance characteristic using formula Calculated wave impedance characteristics ;in, Indicates rock density; The logarithmic resistivity characteristic calculation unit is used to calculate the logarithmic resistivity characteristic using the formula Logarithmic characteristics of resistivity were calculated. ;in, Represents resistivity; The polarizability logarithmic characteristic calculation unit is used to calculate the polarizability logarithmic characteristic using the formula The logarithmic characteristic of polarizability was calculated. ;in, Indicates polarizability; The dynamic elastic modulus characteristic calculation unit is used to calculate the characteristic of the formula. The dynamic elastic modulus characteristics were calculated. ; The unit for calculating the interaction characteristics of resistivity and polarizability is used to calculate the interaction characteristics using the formula... The interaction characteristics of resistivity and polarizability were calculated. ; The interaction characteristic calculation unit of pyrite content and porosity is used to calculate the interaction characteristic of pyrite content and porosity using the formula. The interaction characteristics between pyrite content and porosity were calculated. ;in, Indicates pyrite content, Indicates porosity; Enhanced feature matrix acquisition unit, used to obtain the wave velocity ratio feature Wave impedance characteristics Logarithmic characteristics of resistivity Logarithmic characteristics of polarizability Dynamic elastic modulus characteristics Interaction characteristics of resistivity and polarizability Interaction characteristics of pyrite content and porosity As a newly added feature column, it is used in conjunction with the preprocessing matrix. The enhanced feature matrix is formed by concatenating the columns. ; The matrix concatenation module is specifically used to concatenate the enhanced feature matrix. With the first-level prediction matrix Column concatenation is performed to obtain the meta-learner input matrix. .
9. The shale core TOC prediction system based on adaptive dual-path output as described in claim 6, characterized in that, The weighted convex combination module is specifically used to use the formula Construct the hierarchical adaptive fusion weights ;in, For the first Local credibility coefficients for each TOC stratified interval.
10. The shale core TOC prediction system based on adaptive dual-path output as described in any one of claims 6-9, characterized in that, The final TOC prediction result output module includes: The local score calculation submodule is used to set the path. In the The local scores within each TOC interval are: ;in, For the first The path in the first Local scores within a TOC stratified interval, subscript For path index, subscript TOC hierarchical range index; For the first The path in the first The coefficient of determination on the TOC stratified intervals was validated. To avoid The score becomes invalid when the value is negative; For the first The path in the first Root mean square error on the TOC stratified interval validation samples; The comprehensive score calculation submodule is used to calculate the sample proportion of each interval. The path is obtained by weighting and summing the local scores. Overall rating: ; The final TOC prediction result output submodule is used to output the formula. The above final TOC prediction results are obtained. ;in, This represents the prediction result output by the improved stacked integration path; This represents the prediction result output by the improved hierarchical adaptive performance-weighted integration path; and These represent the combined scores of the improved stacked integration path and the improved hierarchical adaptive performance-weighted integration path under the corresponding TOC hierarchical verification conditions.