Coal gangue-based geopolymer compressive strength high-precision prediction model based on Stacking ensemble learning and construction method thereof

By using Stacking ensemble learning and SHAP analysis, a predictive model for the compressive strength of coal gangue-based geopolymers was constructed. This model solves the problems of low prediction accuracy and poor interpretability in existing technologies, achieving high-precision prediction and interpretability, and promoting the industrial application of coal gangue resource utilization.

CN121237276APending Publication Date: 2025-12-30LIAONING TECHNICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511314239.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In existing technologies, the 28-day compressive strength prediction accuracy of coal gangue-based geopolymers is low, the integrated learning strategy does not incorporate the original material characteristics, resulting in poor generalization ability, and the model lacks interpretability, making it difficult to guide industrial proportion optimization.

Method used

A high-precision prediction model is constructed by using the Stacking ensemble learning method, which combines multi-base model optimization and feature space fusion. The contribution of features is quantified by SHapley additive interpretation (SHAP) analysis to form an 8-dimensional ensemble feature space, which is then combined with support vector machine (SVM) for final prediction.

Benefits of technology

It achieved high-precision prediction of 28-day compressive strength, improved the test set determination coefficient (R2) to 0.9251, reduced the mean absolute error (MAE) and root mean square error (RMSE), demonstrated stable model generalization ability and interpretability, guided industrial ratio adjustment, and promoted the high-value-added resource utilization of coal gangue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237276A_ABST
    Figure CN121237276A_ABST
Patent Text Reader

Abstract

The invention discloses a coal gangue-based geopolymer compressive strength prediction model construction method based on Stacking ensemble learning, and relates to the crossing field of coal gangue resource utilization, material performance prediction and machine learning. The method comprises the following steps: firstly, collecting the content of SiO2, Al2O3 and CaO, the mixing amount of an alkali activator, the water-solid ratio and 28-day compressive strength data, carrying out Z-score standardization treatment, and taking Zgt; 3, removing abnormal values; constructing RF, XGB and MLP base models, and optimizing hyper-parameters of each base model by combining grid search with nested 5-fold cross validation; then splicing a base model prediction result and original features into an 8-dimensional space, optimizing an SVM meta-model through nested cross validation, and constructing a Stacking integrated model; secondly, evaluating the performance of the model by adopting R2, MAE and RMSE; and finally, analyzing the contribution degree of each feature through an SHAP method. The method improves the prediction precision and generalization ability, can guide the proportion optimization, and promotes the high-value utilization of coal gangue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of coal gangue resource utilization and material performance prediction technology, specifically involving a prediction model for the compressive strength of coal gangue-based polymers based on Stacking ensemble learning and its construction method. This model and method are applicable to building material performance evaluation, coal gangue-based polymer formulation optimization, and industrial production parameter control. In particular, it can be directly applied to the prediction and proportioning design of 28-day compressive strength of coal gangue-based polymers in building structures, road engineering, and mine backfilling scenarios, providing technical support for the high-value-added resource utilization of coal gangue. Background Technology

[0002] Coal gangue is a major industrial solid waste generated during coal mining and washing, and its treatment and utilization have become a key issue restricting the green development of the coal industry. According to data from the "China Energy Statistical Yearbook 2023," as of 2023, the cumulative stockpile of coal gangue in my country exceeded 5 billion tons, with an annual increase of 150-200 million tons. The long-term stockpiling of this coal gangue not only occupies a large amount of land, but the heavy metals it contains may also leach into and pollute the soil and groundwater through rainwater. Dust pollution will also exacerbate regional air pollution, which is detrimental to the ecological environment and human health, necessitating efficient resource utilization methods.

[0003] Using coal gangue as the main raw material (mass percentage ≥ 50%) to prepare geopolymers is a core direction for their high-value-added utilization. Geopolymers are inorganic polymer materials formed by alkali activators activating aluminosilicate raw materials, mainly with an aluminosilicate network structure. Compared with traditional silicate cement, studies have shown that its carbon emissions can be reduced by 60%-80% (see the study published by McLellan B C et al. in the Journal of Cleaner Production 2011), and it is mechanically stable and environmentally friendly, with great potential for application in building structures, road engineering, mine backfilling and other fields.

[0004] The 28-day compressive strength is a core mechanical indicator for judging the engineering applicability of coal gangue-based geopolymers. According to GB / T17671-1999 "Test Method for Strength of Cement Mortar (ISO Method)," it is necessary to make it into 40mm×40mm×40mm test blocks and cure them for 28 days at (20±2)℃ and relative humidity ≥95%. The compressive strength needs to reach 15-40MPa to meet the requirements of different projects. Therefore, achieving high-precision prediction of this strength is a key prerequisite for promoting its industrial application.

[0005] While the industry is currently attempting to use machine learning to predict this intensity, three core shortcomings remain:

[0006] First, single models have low accuracy, mostly relying on single models such as random forests and gradient boosting trees, using only the original material features (SiO2 content, Al2O3 content, CaO content, alkali activator dosage, water-to-solid ratio) as input, without combining ensemble learning to explore nonlinear correlations. For example, the model test set R built by Wangwen H et al. in the Journal of Cleaner Production 2022 2 With a strength of only 0.81 and a MAE of 7.2 MPa, Zhang Daming (Northeastern University, 2017) also pointed out that when a single model is used to process the synergistic reaction of SiO2 and CaO, it is easy to lose feature information, resulting in a large prediction bias.

[0007] Second, the ensemble model has poor generalization. Existing Bagging and Boosting strategies simply weight or stack the prediction results of the base models without integrating the original material's physicochemical information. This makes them susceptible to the cumulative error of the base models, resulting in poor model R-value when the coal gangue production location changes. 2 Fluctuations ≥10% have been verified in the industrial-scale test conducted by Liao Yue (Liaoning University of Engineering and Technology, 2024).

[0008] Third, the models lack interpretability, mostly only outputting prediction results without quantitative analysis of the contribution of each feature to compressive strength, exhibiting "black box" characteristics. For example, the research by Liao Yue and Zhang Lin (Shenyang Jianzhu University, 2024) mentioned that mainstream models cannot clearly define "the impact of a 1% change in SiO2 content on compressive strength", making it difficult to guide industrial formulation.

[0009] In summary, existing technologies suffer from three bottlenecks: low accuracy of single models, lack of integration of original features in ensemble strategies, and lack of interpretability of models. This invention constructs a Stacking ensemble prediction model based on experimental data (including chemical composition, process parameters, and 28-day compressive strength) from 311 sets of publicly available academic literature, aiming to solve these defects and promote the industrialization of coal gangue resource utilization. Summary of the Invention

[0010] This invention aims to overcome three major shortcomings in existing technologies: "single models are insufficient to capture nonlinear features, resulting in low prediction accuracy; ensemble learning strategies do not integrate original material features, leading to poor generalization ability; and models lack interpretability, making them unable to guide proportion optimization." It provides a prediction model for the compressive strength of coal gangue-based geopolymers based on Stacking ensemble learning and its construction method. This model achieves high-precision prediction of 28-day compressive strength and quantifies the contribution of each input feature, providing a clear basis for the proportion design and industrial production parameter control of coal gangue-based geopolymers.

[0011] To achieve the above objectives, the present invention adopts the following technical solution. The present invention provides a prediction model for the compressive strength of coal gangue-based geological polymers based on Stacking ensemble learning and its construction method. The core of the method lies in the integrated technical solution of "multi-base model optimization + feature space fusion + interpretability analysis," specifically including the following steps:

[0012] First, data collection and preprocessing were performed. The system collected chemical composition, process parameters, and compressive strength values ​​of coal gangue-based geopolymers to form a sample set. The Z-score standardization method was used to remove outliers and standardize the data, eliminating the influence of dimensions and obtaining a high-quality standardized dataset.

[0013] Secondly, the base models were constructed and optimized. Three algorithms with complementary characteristics—Random Forest Regression (RF), Extreme Gradient Boosting (XGB), and Multilayer Perceptron (MLP)—were selected as base models. Using a grid search combined with a k-fold cross-validation strategy, and with the goal of minimizing the mean squared error, the optimal hyperparameter combination for each base model was determined to ensure optimal performance of each base model.

[0014] Next, a Stacking ensemble model is constructed. This invention innovatively employs a feature fusion strategy, using the optimized prediction results of each base model as new meta-features. These are then horizontally combined with the original five material features (SiO2 content, Al2O3 content, CaO content, alkali activator dosage, and water-to-solid ratio) to form an 8-dimensional ensemble feature space. A nested cross-validation framework is used to train a Support Vector Machine (SVM) as the meta-model, ultimately forming a high-precision, high-generalization Stacking ensemble prediction model.

[0015] Then, the coefficient of determination (R²) is used. 2 The performance of the ensemble model and the base model is evaluated using multiple metrics, including mean absolute error (MAE) and root mean square error (RMSE), to quantitatively verify the superiority of the invention.

[0016] Finally, the Shapley Additive Interpretation (SHAP) method was used to perform interpretability analysis on the integrated model, quantifying the contribution of each feature in the 8-dimensional integrated feature space to the final predicted compressive strength value, clarifying the key influencing factors and their effects, and transforming the "black box" model into a quantitative tool that can guide production.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] By integrating the advantages of three base models—RF, XGB, and MLP—through a stacking ensemble strategy and combining 8-dimensional feature fusion (original features + base model prediction results), the model can accurately capture nonlinear features such as the synergistic reaction between SiO2 and CaO, and the interaction between the alkali activator and the water-to-solid ratio. The test set R... 2 The accuracy reached 0.9251, a 5.65% improvement over the optimal base model (MLP); MAE = 0.2294, a 15.88% reduction compared to MLP; and RMSE = 0.3103, a 26.07% reduction compared to XGB, demonstrating a significant improvement in prediction accuracy. The nested cross-validation framework effectively avoided data leakage, and the 8-dimensional feature space fused with the original material's physicochemical information ensured that the model could adapt to different coal gangue production sites. 2 The fluctuation is ≤3%, far lower than that of existing integrated models (fluctuation ≥10%), and the generalization ability is stable. By quantifying the contribution of each feature through the SHAP method, the laws and specific quantitative effects of "CaO content promoting strength and water-to-solid ratio inhibiting strength" are clarified, solving the "black box" problem of existing models. It can directly guide industrial ratio adjustment, shorten the R&D cycle and reduce experimental costs. At the same time, it promotes the high-value-added resource utilization of coal gangue, reduces the environmental hazards of stockpiling, and the carbon emissions of coal gangue-based geopolymers are 60%-80% lower than those of traditional cement, which is in line with the national "dual carbon" strategy and solid waste comprehensive utilization policy, and has significant environmental and social benefits. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the entire process of model construction and prediction in an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of the nested cross-validation implementation process used in this embodiment of the invention;

[0021] Figure 3 This is a schematic diagram illustrating the construction process of the 8-dimensional integrated feature space in an embodiment of the present invention;

[0022] Figure 4 This is a schematic diagram of the SHAP feature contribution analysis path in an embodiment of the present invention;

[0023] Figure 5 This is a comparison chart of the error distribution of each model in the embodiments of the present invention;

[0024] Figure 6 This is a thermogram showing the correlation between the original features and compressive strength in an embodiment of the present invention;

[0025] Figure 7 This is a beehive diagram showing the characteristic SHAP values ​​in an embodiment of the present invention;

[0026] Figure 8 This is a scatter plot of the predicted values ​​and the true values ​​of the Stacking ensemble model in this embodiment of the invention. Detailed Implementation

[0027] The present invention will be described in detail below through specific embodiments. The implementation of each step can be understood with reference to the accompanying drawings.

[0028] I. Data Preparation and Preprocessing

[0029] This embodiment collected data from 311 coal gangue-based geopolymer samples, with the following characteristic values: SiO2 content 35%-55%, Al2O3 content 15%-25%, CaO content 5%-20%, alkali activator dosage 5%-15%, water-to-solid ratio 0.25-0.45, and 28-day compressive strength 10-45 MPa. Figure 1 The overall process, as shown, first involves preprocessing the raw data: Z-score standardization (with a threshold |Z|>3) is used to remove 27 outlier samples, retaining 285 valid data points. Subsequently, Z-score standardization is applied to the valid data to eliminate the influence of unit weights, resulting in a standardized dataset. The feature distribution of the preprocessed data is shown below. Figure 6 As shown in the correlation heatmap, there is a significant nonlinear relationship between each feature and compressive strength.

[0030] II. Base Model Training and Hyperparameter Optimization

[0031] Based on the Python platform and the Scikit-learn and XGBoost libraries, three base models—RF, XGB, and MLP—were built. The hyperparameters of each model were optimized using grid search combined with 3-fold cross-validation. The optimized hyperparameter combinations and performance are as follows: For the RF model, the optimal hyperparameters are n_estimators = 40, max_depth = 4, min_samples_leaf = 6, min_samples_split = 8, with a cross-validation RMSE of 0.3685; for the XGB model, the optimal hyperparameters are n_estimators = 50, learning_rate = 0.1, max_depth = 3, with a cross-validation RMSE of 0.3395; for the MLP model, the optimal hyperparameters are hidden_layer_sizes = (100, 50), activation = 'relu', learning_rate_init = 0.001, max_iter = 500, with a cross-validation RMSE of 0.4400. The error distributions of each base model are compared below. Figure 5 As shown.

[0032] III. Construction and Validation of the Stacking Integration Model

[0033] like Figure 2As shown in the nested cross-validation flowchart, a nested 5-fold cross-validation framework is used to construct the Stacking ensemble model. In the outer 5-fold loop, 228 samples are taken as the training set and 57 as the test set in each iteration. The inner layer further divides the training set into 182 training samples and 46 validation samples. After the base model is trained on the inner training set, it makes predictions on the validation set, and the prediction results are compared with the original 5 features. Figure 3 The illustrated process is used to concatenate the data to form an 8-dimensional ensemble feature space. Using SVM as the meta-model, the optimal hyperparameters were determined through grid search to be kernel = 'rbf', C = 10, and gamma = 0.1. The final ensemble model performed excellently on the independent test set (R² = 0.1). 2 =0.9251, MAE=0.2294MPa, RMSE=0.3103MPa), the scatter plot relationship between predicted and true values ​​is as follows: Figure 8 As shown, a high degree of linear correlation is observed.

[0034] IV. Feature Contribution Analysis and Application

[0035] The SHAP method is used to perform interpretability analysis on the model. The analysis path is as follows: Figure 4 As shown in the diagram. Figure 7 The SHAP swarm diagram clearly shows the importance ranking of each feature: CaO content > SiO2 content > water-to-solid ratio > Al2O3 content > alkali activator dosage. It also quantifies that for every 1% increase in CaO content, the compressive strength increases by 0.9 MPa, and for every 0.1 increase in the water-to-solid ratio, the compressive strength decreases by 1.2 MPa. Based on this pattern, the proportions of coal gangue raw materials (SiO2 = 48%, Al2O3 = 20%) from a certain region were optimized (CaO = 16%, water-to-solid ratio = 0.32, alkali activator dosage = 9%). The resulting geopolymer achieved a 28-day compressive strength of 38 MPa, meeting the requirements for building structural engineering.

[0036] The above embodiments describe in detail a prediction model for the compressive strength of coal gangue-based geological polymers based on Stacking ensemble learning and its construction method provided by the present invention. These embodiments are merely illustrative examples of the technical solution of the present invention and are not intended to limit the scope of protection of the present invention. For those skilled in the art, various modifications, substitutions, adjustments, and variations can be made to the data preprocessing method, base model type, hyperparameter combination, metamodel selection, and verification strategy without departing from the principles and essence of the present invention. All such changes and improvements should be considered to fall within the scope of protection defined by the claims of the present invention.

[0037] Background Technology Documents

[0038] [1] Energy Statistics Division, National Bureau of Statistics of China. China Energy Statistics Yearbook 2023 [M]. Beijing: China Statistics Press, 2023.

[0039] [2]MCLELLAnB C,WILLIAMS RP,LAY J,et al.Costs and carbonemissions forgeopolymer pastes incomparison to ordinary portland cement[J].JournalofCleaner Production,2011,19(9-10):1080-1090.

[0040] [3]WANGWEnH,YANG

[0041] [4] Zhang Daming. Preparation of coal gangue-based polymers and prediction of concrete strength [D]. Shenyang: Northeastern University, 2017.

[0042] [5] Liao Yue. Optimization of the proportion of coal gangue-slag-fly ash geopolymer grouting material and its rheological properties and durability study [D]. Fuxin: Liaoning University of Engineering and Technology, 2024.

[0043] [6] Zhang Lin. Study on hydration dynamic characteristics of alkali-activated self-combusting coal gangue-slag based all-solid waste cementitious materials [D]. Shenyang: Shenyang Jianzhu University, 2024.

Claims

1. A method for constructing a coal gangue-based geopolymer compressive strength prediction model based on stacking ensemble learning, characterized in that, Comprising the following steps: S1. Data collection and pretreatment: Collect the data related to coal gangue-based geopolymer from the published academic literature, including chemical composition SiO2 content, Al2O3 content, CaO content, alkali activator dosage, water-solid ratio, and 28-day compressive strength value; Adopt Z-score standardization method to detect and eliminate outliers, set threshold value as |Z|>3 (Z is the standard deviation distance between data point and mean value, calculation formula is x is the original data, μ is the characteristic mean value, and σ is the standard deviation); Adopt Z-score standardization formula to process the processed data, and obtain the standardized data set; S2. Base model construction and optimization: Constructing random forest regression model (RF), extreme gradient boosting regression model (XGB), and multilayer perception model (MLP) as base models; optimizing the hyperparameters of each base model by grid search, and the overall base model hyperparameter optimization process is based on a nested 5-fold cross-validation framework to realize model evaluation, wherein: the hyperparameters of RF include the total number of trees, the maximum depth of trees, the minimum leaf node sample size, and the minimum split sample size; the hyperparameters of XGB include the number of trees, the learning rate, and the maximum depth of trees; the hyperparameters of MLP include the number of hidden layer neurons, the activation function, the learning rate, and the maximum number of iterations; S3. Stacking ensemble model construction: using the "feature fusion + SVM ensemble" strategy, the prediction results of the base models optimized in step S2 are taken as 3 new features, which are horizontally spliced with the original 5 material features (SiO2 content, Al2O3 content, CaO content, alkali activator dosage, and water-solid ratio) to form an 8-dimensional integrated feature space; a nested cross-validation framework is used (5-fold cross-validation is used to divide the outer training set and test set, and 5-fold cross-validation is used to divide the outer training set into inner training set and test set), wherein the outer and inner cross-validation both use random division to ensure the consistency of the feature distribution of each fold; the base model is trained on the inner training set and outputs the prediction results of the inner test set, and only the prediction results of the inner test set are used as the training data of the SVM meta-model; the kernel function (RBF / linear), the regularization parameter (0.1-100), and the kernel coefficient (0.01-1) of the SVM meta-model are optimized by grid search, the SVM meta-model is trained using the 8-dimensional integrated features, and the Stacking ensemble model is constructed; S4. Model evaluation: The prediction performance of the Stacking ensemble model and each base model was evaluated using the coefficient of determination (R 2 ), mean absolute error (MAE), and root mean square error (RMSE), with the calculation formulas being: where y i is the actual value, y' is the predicted value, is the mean of the actual values, and n is the number of samples. S5. Feature contribution analysis: using the SHapley Additive Explanations (SHAP) method to analyze the contribution of each input feature in the 8-dimensional integrated feature space to the compressive strength of coal gangue-based geopolymer.

2. The method of claim 1, wherein, The parameter range of the grid search in step S2 is set according to the algorithm characteristics of each base model, the distribution of experimental data, and the results of pre-experiments and empirical rules in the field of machine learning, and the optimal hyperparameters are determined by the root mean square error minimization criterion of 5-fold cross-validation.

3. The method of claim 1, wherein, The outer training set and test set ratio of the nested cross-validation in step S3 is 4:1, and the inner training set and test set ratio is 4:

1. The ratio setting is based on sample size verification: when the ratio is <4:1, the test set sample size is insufficient, resulting in evaluation fluctuation (cross-validation error coefficient of variation CV value >10%); when the ratio is >4:1, the risk of overfitting of the training set increases (R 2 The difference >0.05); 4:1 is the optimal balanced ratio considering the evaluation stability and the risk of overfitting; and the prediction result of the base model is only used for the training process of the meta model and does not participate in the performance evaluation of the outer test set.

4. A coal gangue-based geopolymer compressive strength prediction model based on stacking ensemble learning, characterized by, The model comprises a base model layer and a meta-model layer; the base model layer includes random forest regression model (RF), extreme gradient boosting regression model (XGB), and multilayer perception model (MLP); the meta-model layer is a support vector machine model (SVM); the input feature space of the model is an 8-dimensional integrated feature space formed by horizontally splicing 5 original material features (SiO2 content, Al2O3 content, CaO content, alkali activator dosage, and water-solid ratio) and 3 base model prediction results (RF prediction result, XGB prediction result, and MLP prediction result) according to the feature dimension; the model is constructed by the method of claims 1-3.

5. The model of claim 4 is applied in the synergistic optimization of SiO2 and CaO ratio, interactive parameter adjustment of alkali activator dosage and water-solid ratio in coal gangue-based geopolymer, characterized in that, The contribution of each feature to the compressive strength is quantified by the SHapley Additive exPlanations (SHAP) method (e.g., the compressive strength is increased by 0.8 MPa on average when the SiO2 content is increased by 1%), which guides the precise control of the proportioning parameters to improve the 28-day compressive strength.