Geopolymer preparation and optimization method and system based on machine learning

By fusing multiple machine learning models and using an adaptive weighting mechanism, the problem of performance prediction under complex formulation conditions in geopolymer preparation was solved, achieving efficient and accurate performance optimization and material preparation, improving the strength, durability and permeability of geopolymers, and reducing production costs.

CN120977443AActive Publication Date: 2025-11-18GUIZHOU INST OF COAL SCI

Patent Information

Application Number
CN202511506146.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Traditional geopolymer preparation processes are complex, inefficient, and have inaccurate performance predictions. They are difficult to quickly optimize geopolymer performance under complex formulation conditions. Existing machine learning models have weak generalization ability and insufficient local feature capture when processing high-dimensional nonlinear data.

Method used

By fusing multiple machine learning models and introducing an adaptive weighting mechanism, an artificial neural network model using Rational-Minkowski kernel function support vector regression, piecewise Gaussian process regression, regularized neighborhood component analysis, and Levenberg-Marquardt algorithm is established to establish a nonlinear mapping relationship between raw material ratio and performance. Combined with particle swarm optimization algorithm, efficient parameter search under multi-objective constraints is achieved.

Benefits of technology

It significantly improves the accuracy and robustness of geopolymer performance prediction, reduces model generalization error, enhances the strength, durability, and permeability of geopolymer materials, reduces production costs, and broadens the resource utilization pathways for coal-based solid waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977443A_ABST
    Figure CN120977443A_ABST
Patent Text Reader

Abstract

The invention provides a geopolymer preparation and optimization method and system based on machine learning. The method is applied to the technical field of material science and machine learning. The method comprises the following steps: acquiring geopolymer preparation experimental data and preprocessing the data; performing nonlinear regression modeling on the geopolymer performance based on four machine learning regression algorithms, and constructing a geopolymer performance prediction model; calculating and distributing weights according to the mean square error of each machine learning model on the verification set, and performing weighted fusion to obtain a performance prediction result; receiving target performance parameters input by a user and an initial raw material ratio range, performing performance prediction by using the trained geopolymer performance prediction model, and reversely searching an optimal ratio combination meeting target performance constraints through an optimization algorithm; and preparing a geopolymer according to the optimal ratio combination to prepare the coal gangue-slag-fly ash geopolymer grouting material. According to the method, the prediction precision and the model generalization ability are effectively improved, and intelligent recommendation and accurate performance prediction of the raw material ratio are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of materials science and machine learning technology, and in particular to a method and system for the preparation and optimization of geopolymers based on machine learning. Background Technology

[0002] Geopolymers are a class of novel inorganic polymer materials formed through polymerization reactions of industrial solid wastes rich in silicon and aluminum components (such as fly ash, coal gangue, and slag) under alkaline activation conditions. Due to their excellent mechanical strength, good thermal stability, corrosion resistance, and environmentally friendly properties, geopolymers have broad application prospects in green building materials, environmental engineering, and transportation infrastructure.

[0003] However, coal-based solid wastes (especially coal gangue and fly ash) typically have low chemical activity and vary greatly in composition, leading to complex processes, numerous repetitive experiments, low efficiency, and large performance fluctuations in traditional geopolymer preparation methods. Especially in practical engineering, the key challenge hindering widespread adoption lies in rapidly predicting geopolymer performance and optimizing mixing parameters under complex formulation conditions.

[0004] Traditional methods have shortcomings: formulation design relies on empirical trial and error: hundreds of orthogonal experiments are required to obtain a feasible formulation, resulting in a long R&D cycle; performance prediction accuracy is insufficient: empirical formulas (such as Bolomey's formula) have large fitting errors for complex nonlinear relationships, making it difficult to meet engineering accuracy requirements; system optimization efficiency is low: multi-objective optimization (such as simultaneously pursuing strength and fluidity) requires thousands of iterative calculations, and existing methods (such as genetic algorithms) have slow convergence speeds, limiting their engineering applications.

[0005] In recent years, with the widespread application of machine learning in materials science, two major challenges remain in the field of geopolymers: weak generalization ability of single models: traditional support vector machines (SVM) are prone to overfitting when dealing with high-dimensional nonlinear data, while artificial neural networks (ANN) are sensitive to hyperparameters and require massive amounts of data; insufficient capture of local features: the performance of geopolymers is piecewise sensitive to the ratio parameters (e.g., the performance of water glass changes significantly in the range of 1.8-2.2), and existing global modeling methods are difficult to capture such local patterns.

[0006] Therefore, there is an urgent need for a new method and system for the preparation and optimization of geopolymers that can efficiently establish a complex nonlinear mapping relationship between proportions and performance, and realize intelligent optimization of the entire process from raw material proportions to performance prediction. Summary of the Invention

[0007] To address the problems of experience-dependent ratio design, inaccurate performance prediction, and low system optimization efficiency in the preparation of geopolymers, this invention proposes a machine learning-based method and system for geopolymer preparation and optimization. By integrating multiple machine learning models and introducing an adaptive weighting mechanism, the method effectively improves prediction accuracy and model generalization ability, enabling intelligent recommendation of raw material ratios and accurate performance prediction.

[0008] To achieve the above objectives, the following technical solution is adopted:

[0009] According to a first aspect of the present invention, a method for preparing and optimizing geopolymers based on machine learning is provided, comprising the following steps:

[0010] S1. Collect experimental data on the preparation of geopolymers, including parameters of solid raw materials, parameters of activators and performance output indicators, and perform missing value imputation, outlier removal, feature scaling, standardization, encoding and dataset partitioning on the experimental data;

[0011] S2. Based on four machine learning regression algorithms, nonlinear regression modeling of geopolymer performance is performed to construct a geopolymer performance prediction model integrating the four machine learning models, establishing a nonlinear mapping relationship between raw material ratio and performance, including:

[0012] Support vector regression model based on Rational-Minkowski kernel function;

[0013] The piecewise Gaussian process regression model divides the input space into sub-intervals, models each sub-interval independently, and then performs boundary weighted fusion.

[0014] The regularized neighborhood component analysis model learns feature weights by minimizing a loss function that includes an L2 regularization term, and defines a feature-weighted distance metric.

[0015] Artificial neural network model based on Levenberg-Marquardt algorithm;

[0016] S3. Calculate and assign weights based on the mean squared error of each machine learning model on the validation set in the dataset, and then perform weighted fusion to obtain the performance prediction result.

[0017] S4. Receive the target performance parameters and initial raw material ratio range input by the user, use the trained geopolymer performance prediction model to predict the performance, and use the optimization algorithm to search in reverse for the best ratio combination that meets the target performance constraints.

[0018] S5. Based on the optimal ratio combination, prepare a geopolymer to prepare a coal gangue-slag-fly ash geopolymer grouting material.

[0019] Furthermore, the construction of the support vector regression model based on the Rational-Minkowski kernel function includes:

[0020] Based on the Rational-Minkowski kernel function, an improved kernel function support vector regression model is constructed, where the expression for the Rational-Minkowski kernel function is:

[0021]

[0022] in, Let n be the sample input vector, and n be the number of samples. Minkowski distance The order norm, p∈[1,3], is used to adjust the norm order of the Minkowski distance; τ is the scaling parameter and τ>0, used to control the width of the kernel function; q is the exponential adjustment parameter and q∈[1,5], used to control the decay rate of the kernel function.

[0023] Furthermore, the construction of the piecewise Gaussian process regression model includes:

[0024] Input space partitioning: Based on principal component analysis and K-means clustering or expert knowledge rules, the geopolymer ratio input space is divided into M sub-intervals, each of which satisfies local data stability;

[0025] Local model training: Gaussian process regression sub-models are trained independently for each sub-interval, and kernel function type and hyperparameters are selected for them; the kernel function includes the squared exponential kernel function, Matérn3 / 2 kernel function or RationalQuadratic kernel function;

[0026] Segmented prediction ensemble: For a test sample, determine its sub-interval and use the corresponding sub-model for prediction; if it is located in a boundary region, the prediction results of adjacent sub-models are fused by weighted average.

[0027] Hyperparameter optimization: Estimate the hyperparameters of each sub-model by maximizing the log-marginal likelihood function.

[0028] Furthermore, if the location is in a boundary region, the prediction results of neighboring sub-models are fused by weighted averaging, including:

[0029]

[0030]

[0031] in, : Predicted mean, representing the mean given input Conditional expected values ​​of geopolymer performance indices under given conditions; : Prediction variance, representing the variance of the predicted mean A quantitative measure of uncertainty or confidence level; Input samples for testing; This represents the total number of sub-intervals or sub-models. Indicates the first Sub-model exist The predicted mean at the location; Indicates the first Sub-models in The predicted variance at the location; For the first The spatial weights of each sub-model satisfy the normalization constraint:

[0032]

[0033] Weight according to Calculation of the inverse proportional function of distance to the center of each subinterval:

[0034]

[0035] in, Indicates test input sample With the Center of each sub-interval Euclidean distance.

[0036] Furthermore, the estimation of hyperparameters of each sub-model by maximizing the log-marginal likelihood function includes:

[0037]

[0038] in, The training output vector, i.e., the target variable, includes compressive strength; The training input sample matrix, i.e., the geopolymer sizing parameters, has a total of There are samples, each sample has One input feature; The covariance matrix is ​​defined as follows: , by kernel function structure; : The variance of observation noise, used to characterize unexplainable noise in the data; : Identity matrix, dimension 1 , used to introduce noise.

[0039] Furthermore, the construction of the regularized neighborhood component analysis model includes the following steps:

[0040] Feature-weighted distance definition: Construct the weighted Manhattan distance metric function:

[0041]

[0042] in, : No. The input vector of each sample; : No. In the nth sample The values ​​of each feature; : No. Weight coefficients of each feature; distance metric It is a weighted Manhattan distance form;

[0043] Neighborhood probability calculation: Define sample choose The probability of being a valid neighbor:

[0044]

[0045] : Indicates a sample Will The higher the probability value, the better the probability of it being a "valid neighbor". The more likely to be The predictions have an effect; : No. The smoothing scale parameter for each feature is used to normalize the distance values ​​of different feature dimensions;

[0046] Model training: Feature weights are learned by minimizing a loss function that includes an L2 regularization term;

[0047] Feature selection: By compressing the influence of redundant features through optimized weights, high-dimensional feature selection and dimensionality reduction are achieved.

[0048] Furthermore, the artificial neural network model based on the Levenberg-Marquardt algorithm satisfies the following conditions:

[0049] Network structure: A three-layer feedforward neural network is adopted. The input layer receives the aggregate ratio parameters, including coal gangue content, fly ash content, slag content, water glass modulus, and Na2O content. The number of neurons in the hidden layer is 2-15. The output layer predicts the target performance indicators.

[0050] Weight update formula: The weights are dynamically adjusted using the Levenberg-Marquardt algorithm. The update formula is as follows:

[0051]

[0052] Let be the weight vector for the r-th iteration; This is the Jacobian matrix, containing the partial derivatives of the network output with respect to the weights; This is the error vector; To adjust the parameters; It is the identity matrix;

[0053] Training mechanism: When When →0, the algorithm approximates the Gauss-Newton method to accelerate convergence; when As the value increases, the algorithm degenerates into gradient descent to enhance stability;

[0054] Early stopping mechanism: During training, the validation set error is monitored, and training is terminated when the error does not decrease for a preset number of consecutive times to prevent overfitting.

[0055] Furthermore, in step S3, the weight calculation includes:

[0056] Calculate the mean squared error of each machine learning model on the validation set:

[0057]

[0058] in, For the first The mean squared error of a machine learning model; It is the first The predicted value of each model for the i-th sample; It is the true value of the i-th sample;

[0059] According to the mean square error Calculate the weight using the reciprocal :

[0060]

[0061] in, It is the first The weights of a machine learning model;

[0062] Based on the weights of the four machine learning models The prediction results of the four machine learning models are weighted and averaged to output the final performance prediction value.

[0063] Furthermore, in step S4: the user-defined target performance parameters include compressive strength > 80 MPa and permeability coefficient < 1 × 10⁻⁶. -7 cm / s, and durability indicators including sulfate corrosion resistance cycles and wet-dry cycle mass loss rate;

[0064] The initial raw material ratio range includes: coal gangue 60-80 wt%, slag 10-20 wt%, fly ash 10-20 wt%; water glass modulus 1.5-2.5, Na2O content 3-22 wt%; activator to solid raw material mass ratio 0.3-0.5;

[0065] Step S5, according to the optimal ratio combination, prepares a geopolymer for coal gangue-slag-fly ash geopolymer grouting material, including:

[0066] Coal gangue is crushed and ball-milled to a specific surface area of ​​400-600 m². 2 / kg, the particle size of slag and fly ash is controlled at 0.045-0.075 mm;

[0067] According to the recommended optimal ratio combination, coal gangue, slag, and fly ash are mixed, an activator is added, and the mixture is stirred until the fluidity of the slurry reaches 150-200 mm to form a coal gangue-slag-fly ash geopolymer grouting slurry.

[0068] Grouting and curing: After the coal gangue-slag-fly ash geopolymer grout is pre-cured at room temperature for 24 hours, it is transferred to an environment of 20±2℃ and humidity ≥95% for 28 days to form coal gangue-slag-fly ash geopolymer grouting material.

[0069] According to a second aspect of the present invention, a machine learning-based geopolymer preparation and optimization system is also provided, comprising the following steps:

[0070] The data acquisition module is used to collect experimental data on geopolymer preparation, including solid raw material parameters, activator parameters and performance output indicators, and to perform missing value imputation, outlier removal, feature scaling, standardization, encoding and dataset partitioning on the experimental data.

[0071] The model building module is used to perform nonlinear regression modeling of geopolymer performance based on four machine learning regression algorithms, construct a geopolymer performance prediction model that integrates the four machine learning models, and establish a nonlinear mapping relationship between raw material ratios and performance, including:

[0072] Support vector regression model based on Rational-Minkowski kernel function;

[0073] The piecewise Gaussian process regression model divides the input space into sub-intervals, models each sub-interval independently, and then performs boundary weighted fusion.

[0074] The regularized neighborhood component analysis model learns feature weights by minimizing a loss function that includes an L2 regularization term, and defines a feature-weighted distance metric.

[0075] Artificial neural network model based on Levenberg-Marquardt algorithm;

[0076] The prediction output module is used to calculate and assign weights based on the mean squared error of each machine learning model on the validation set in the dataset, and then perform weighted fusion to obtain the performance prediction result.

[0077] The ratio recommendation module is used to receive the target performance parameters and the initial raw material ratio range input by the user, use the trained geopolymer performance prediction model to predict the performance, and use the optimization algorithm to search for the best ratio combination that meets the target performance constraints.

[0078] The grouting preparation module is used to prepare geopolymer grouting materials for coal gangue-slag-fly ash according to the optimal ratio combination.

[0079] Compared with the prior art, the present invention achieves the following beneficial effects:

[0080] 1. Multi-model fusion improves prediction accuracy and robustness: By integrating support vector regression, piecewise Gaussian process regression, regularized neighborhood component analysis, and neural network models, a heterogeneous model fusion framework is constructed. Based on the adaptive weighting mechanism of validation set mean squared error (such as the complementary relationship between the sensitivity of the Rational-Minkowski kernel function to local features and the uncertainty quantification capability of Gaussian process regression), model bias and variance are effectively balanced, and the prediction error is effectively reduced compared to a single model.

[0081] 2. Enhanced Generalization Ability through Local Modeling and Feature Selection: Piecewise Gaussian process regression employs an input space partitioning strategy (e.g., dividing the proportioning space into 3-5 subdomains using K-means clustering), combined with feature weight learning through regularized neighborhood component analysis (feature dimension compression rate up to 60%), effectively addressing the performance mutation problem caused by differences in raw material composition. In cross-regional coal gangue sample testing, the model's generalization error fluctuation range is reduced, improving stability by 2.3 times compared to traditional global models.

[0082] 3. Inverse optimization drives intelligent ratio design: Based on the backpropagation mechanism of the Levenberg-Marquardt algorithm (convergence speed is 4-8 times faster than gradient descent), combined with the particle swarm optimization algorithm, efficient parameter search is achieved under multi-objective constraints.

[0083] 4. Enhanced Process Adaptability and Resource Conservation: This invention utilizes coal gangue, slag, and fly ash as main raw materials to prepare geopolymer grouting materials using an alkali activator. By optimizing parameters such as coal gangue content (greater than 60%), water glass modulus, and Na2O content, and integrating multiple process parameters such as activator modulus and particle size distribution through feature engineering, the model can adapt to the differences in activity of different solid waste raw materials. This results in the preparation of a grouting material with high strength, low permeability, and excellent durability, suitable for the construction of impermeable layers in coal-based solid waste ecological utilization sites. Experimental results show that this material has approximately 50% higher strength, approximately 40% higher durability, and approximately 15% higher wettability compared to traditional grouting materials. This technology fully utilizes the abundant coal gangue resources of Guizhou Province, combined with the synergistic effect of slag and fly ash, broadening the resource utilization pathways of coal-based solid waste while reducing production costs (by more than 30%) and environmental risks, providing new ideas for the research and development of geopolymer materials both domestically and internationally.

[0084] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0085] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0086] Figure 1 A schematic flowchart of a machine learning-based geopolymer preparation and optimization method according to an embodiment of the present invention is shown.

[0087] Figure 2 This is a schematic diagram of a machine learning-based geopolymer preparation and optimization system according to an embodiment of the present invention. Detailed Implementation

[0088] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0089] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0090] Figure 1 This diagram illustrates a flowchart of a machine learning-based geopolymer preparation and optimization method according to an embodiment of the present invention. Figure 1 and Figure 2 As shown, a machine learning-based method for geopolymer preparation and optimization includes the following steps:

[0091] S1. Collect experimental data on the preparation of geopolymers, including parameters of solid raw materials, parameters of activators and performance output indicators, and perform missing value imputation, outlier removal, feature scaling, standardization, encoding and dataset partitioning on the experimental data;

[0092] S11: Data Acquisition

[0093] This invention first conducts extensive experiments on polymer preparation, collecting experimental data under multiple sets of different formulation conditions to construct a complete input-output dataset, providing a reliable foundation for subsequent model training. The collected data covers the following aspects:

[0094] Raw material proportion parameters: This includes the proportion data of the main raw materials for coal-based solid waste, such as coal gangue (60%), fly ash (25%), and slag (15%). These raw materials each have their own characteristics, and their proportions have a significant impact on the structure and properties of the geopolymer.

[0095] Activator parameters: These mainly include alkaline activator parameters such as water glass modulus (SiO2 / Na2O ratio) and Na2O content. These activators determine the rate and extent of the geopolymerization reaction and are crucial to the final strength and durability.

[0096] Performance output indicators include compressive strength (MPa), permeability (e.g., permeability coefficient m / s), and durability (e.g., number of wet-dry cycles, resistance to sulfuric acid corrosion, etc.). These serve as prediction targets for the model, used to quantify the performance of geopolymer products in engineering applications.

[0097] S12: Data Preprocessing

[0098] After data collection is completed, comprehensive data preprocessing is performed, including the following steps:

[0099] Missing value imputation and outlier removal: For missing data or outlier samples with large errors that occur during the experiment, missing values ​​are imputed using the K-nearest neighbor algorithm (k=5), and outliers are removed using Tukey's rule (IQR coefficient is 1.5). Data cleaning is carried out by interpolation, mean substitution or direct removal to ensure the reliability and integrity of training data.

[0100] Feature scaling and standardization: To improve the convergence speed and prediction accuracy of machine learning models, input features are normalized or standardized according to a unified standard. Common methods include Z-score standardization and Min-Max normalization, which map feature data of different dimensions to the same numerical range, avoiding the dominance of any one feature in the learning results during model training.

[0101] Feature encoding and mapping: If there are non-numerical variables (such as raw material batch classification, activator type, etc.), they are converted into numerical forms that can be processed by the model using methods such as one-heat encoding or tag encoding.

[0102] Dataset partitioning: Divide the entire dataset into training, validation, and test sets, for example, in a ratio of 70%:15%:15% or 80%:10%:10%, to ensure that the model can learn patterns during the training phase, perform parameter tuning during the validation phase, and verify generalization ability during the testing phase.

[0103] S2. Based on four machine learning regression algorithms, nonlinear regression modeling of geopolymer performance is performed to construct a geopolymer performance prediction model integrating the four machine learning models, establishing a nonlinear mapping relationship between raw material ratio and performance, including:

[0104] Support vector regression model based on Rational-Minkowski kernel function;

[0105] The piecewise Gaussian process regression model divides the input space into sub-intervals, models each sub-interval independently, and then performs boundary weighted fusion.

[0106] The regularized neighborhood component analysis model learns feature weights by minimizing a loss function that includes an L2 regularization term, and defines a feature-weighted distance metric.

[0107] Artificial neural network model based on Levenberg-Marquardt algorithm;

[0108] More specifically, S2, based on four machine learning regression algorithms, performs nonlinear regression modeling of geopolymer performance, constructs a geopolymer performance prediction model integrating the four machine learning models, and establishes a nonlinear mapping relationship between raw material ratio and performance, including:

[0109] S21: Construct a support vector regression model based on the Rational-Minkowski kernel function:

[0110] Support Vector Regression (SVR) is a supervised learning method for solving regression problems, particularly suitable for small-sample, high-dimensional, and nonlinear problems. In modeling the relationship between aggregate proportions and target performance, SVR effectively captures the nonlinear mapping between input and output. SVR is a kernel-based machine learning algorithm that solves regression problems by constructing a hyperplane; its key lies in the kernel function. To further improve the model's ability to fit complex nonlinear relationships, this paper introduces the Rational-Minkowski kernel function to construct an improved kernel function support vector regression model, thereby enhancing prediction accuracy. The expression for the Rational-Minkowski kernel function is:

[0111]

[0112] : Sample input vector, where n is the number of samples; Minkowski distance The order norm, p∈[1,3], is used to adjust the norm order of the Minkowski distance; τ is the scaling parameter and τ>0, used to control the width of the kernel function; q is the exponential adjustment parameter and q∈[1,5], used to control the decay rate of the kernel function; when and At this point, the kernel function degenerates into the common RationalQuadratic kernel.

[0113] The Rational-Minkowski kernel function is characterized by a stronger penalty for distant samples; it maintains a larger kernel value among similar samples, which is beneficial for capturing local structures; and it has good adaptability to non-uniformly distributed data. By combining the p-order adjustable norm of the Minkowski distance (p∈[1,3]) with the exponential decay factor of the Rational function (q∈[1,5]), its generalization error is significantly reduced compared to the RBF kernel.

[0114] To ensure the generalization ability and accuracy of the support vector regression model in predicting geopolymer performance, a grid search combined with cross-validation technique was used to optimize the model. Parameters; and multiple divisions of the training / test set (e.g. Repeated experiments were conducted to evaluate the stability and robustness of the model.

[0115] The support vector regression model based on the Rational-Minkowski kernel function proposed in step S21 above introduces a composite kernel function, namely a combination of the Rational function and the Minkowski distance, to enhance the nonlinear modeling capability of the support vector regression model when processing high-dimensional sparse geopolymer ratio data. Compared with the traditional radial basis function, this kernel function has stronger robustness in fitting complex functional relationships and avoiding overfitting, significantly improving the model's adaptability and generalization ability in geopolymer performance prediction.

[0116] S22: Constructing a piecewise Gaussian process regression model:

[0117] Gaussian process regression is a non-parametric Bayesian regression method that can directly output predicted values ​​and their uncertainties during modeling. Addressing the highly nonlinear and locally fluctuating relationship between polymer proportions and performance, this invention further employs a piecewise modeling strategy to construct a piecewise Gaussian process regression model, thereby improving overall fitting accuracy and model interpretability. The specific process of modeling the piecewise Gaussian process regression model includes:

[0118] S221: Input space partitioning:

[0119] Based on principal component analysis and K-means clustering or expert knowledge rules, PCA was used to reduce the dimensionality to 3, followed by K-means clustering (M=4 was determined by the elbow rule). The sub-interval boundaries were delineated using Voronoi diagrams to define the geopolymer stoichiometry space. Divided into Sub-intervals The data within each subinterval satisfies local stability and structural consistency.

[0120] S222: Local model training:

[0121] For each subinterval Train independent Gaussian process regression models The optimal kernel function (such as the quadratic exponential kernel function, the Matérn 3 / 2 kernel function, or the Rational Quadratic kernel function) and hyperparameters can be selected individually for each sub-model.

[0122] S223: Segmented Predictive Ensemble:

[0123] For test samples First, determine which sub-interval it belongs to. Then use the corresponding sub-model Make a prediction; if In the boundary region, a weighted average method is used to fuse the prediction outputs of the two sub-models:

[0124]

[0125]

[0126] in, : Predicted mean, representing the mean given input Conditional expected values ​​of geopolymer performance indices under given conditions; : Prediction variance, representing the variance of the predicted mean A quantitative measure of uncertainty or confidence level; : Test input sample; : The total number of sub-intervals or sub-models (usually 2, when Located at the boundary of two sub-intervals); : No. Sub-model exist The predicted mean at the location; : No. Sub-models in The prediction variance at the location (representing uncertainty);

[0127] : No. The spatial weights of each sub-model satisfy the normalization constraint:

[0128]

[0129] Usually according to The inverse distance function to the center (or boundary) of each subinterval is defined, for example:

[0130]

[0131] in, express With the Center of each sub-interval Euclidean distance.

[0132] Sub-models Estimating hyperparameters by maximizing the logarithmic marginal likelihood function:

[0133]

[0134] in, Training output vector (i.e., target variable, such as compressive strength); The training input sample matrix (i.e., geopolymer ratio parameters) has a total of There are samples, each sample has One input feature; The covariance matrix is ​​defined as follows: , by kernel function structure; : The variance of observation noise, used to characterize unexplainable noise in the data (also known as white noise); : Identity matrix, dimension 1 , used to introduce noise; The determinant of the covariance matrix after adding noise is used to characterize the model complexity. Log-marginal likelihood is the objective function for hyperparameter optimization. The L-BFGS algorithm is used to maximize the log-marginal likelihood, with the iteration tolerance set to 1e-6. Maximizing it yields the optimal hyperparameters (such as the parameters in the kernel function). : Data fit term, measures the relationship between prediction and reality Consistency between them; Complexity penalty term, used to penalize overly complex models; : Normalization constant term, relative to the number of samples related.

[0135] Piecewise Gaussian process regression can provide the uncertainty at each prediction point. This helps with risk control and significantly improves model accuracy and computational efficiency when the sample size is large or local pattern differences are significant.

[0136] Based on the piecewise Gaussian process regression model proposed in step S22 above, a piecewise Gaussian process regression framework was designed to address the non-stationary regions and local abrupt changes in the performance curves of geopolymer materials. By fitting sub-models to different intervals and smoothly splicing the boundaries of each segment model, the model achieves both local accuracy and global consistency, making it particularly suitable for prediction tasks where performance indicators exhibit interval transitions or abrupt changes.

[0137] S23: Constructing a regularized neighborhood component analysis model:

[0138] Regularized neighborhood component analysis (NBI) is an improvement and extension of neighborhood component analysis (NBI) for regression problems, aiming to achieve high-dimensional feature selection through distance metric learning. Compared to traditional NBI, which is mainly used for classification tasks, regularized NBI, while retaining its nearest neighbor concept, introduces a regularization term, which not only effectively suppresses model overfitting but also improves predictive performance on continuous target variables (such as geopolymer compressive strength).

[0139] The core idea of ​​regularized neighborhood component analysis is to learn a set of feature weights among all predictor variables. A weighted distance metric is constructed to ensure that the output values ​​of similar sample pairs are similar in the new feature space, thereby achieving feature selection and dimensionality reduction. The specific construction process of the regularized neighborhood component analysis model includes the following steps:

[0140] S231: Definition of feature-weighted distance:

[0141] For the input sample set, construct a weighted Manhattan distance metric function. like:

[0142]

[0143] in, : No. The input vector of each sample; : No. In the nth sample The values ​​of each feature; : No. Weight coefficients of each feature (to be learned); distance metric It is a weighted Manhattan distance form (which can be extended to the Euclidean norm).

[0144] S232: Neighborhood probability calculation:

[0145] By learning appropriate Regularized neighborhood component analysis can compress the weights of irrelevant or redundant features, minimizing their impact on the final distance metric, thereby achieving feature selection. To construct the "proximity" relationship between samples based on this distance, samples are defined... choose The probability of being a valid neighbor:

[0146]

[0147] : Indicates a sample Will The higher the probability value, the better the probability of it being a "valid neighbor". The more likely to be The predictions have an effect; :sample In the Values ​​can be taken in each feature dimension; :sample In the Values ​​can be taken in each feature dimension; Total number of feature dimensions; : No. The learnable weight parameters of each feature are used to measure the importance of that feature in neighborhood determination. Ensure that the weights are non-negative, while enhancing the sensitivity to important features; : No. The smoothing scale parameter is used to normalize the distance values ​​of different feature dimensions, so that features with different dimensions or degrees of variability have a relatively consistent influence in the overall distance metric. Adaptive calculation based on characteristic variance: .

[0148] This mechanism ensures that the model can learn feature combinations that have better predictive performance within the "neighborhood", thereby improving the overall predictive ability.

[0149] S233: Model Training

[0150] To train the model, regularized neighborhood component analysis minimizes the following loss function with a regularization term:

[0151]

[0152] in, This is the actual output; This represents the predicted value obtained by weighting the values ​​based on the neighbor probabilities. λ is a regularization parameter used to penalize excessively large feature weights and prevent the model from becoming too complex. λ is selected from {0.01, 0.1, 1} through cross-validation. Mean absolute error; for Regularization term.

[0153] This objective function balances prediction accuracy and model simplicity. It learns feature weights by minimizing a loss function that includes an L2 regularization term, and automatically adjusts the feature weights accordingly. By ignoring unrepresentative variables, the effects of multicollinearity can be effectively mitigated.

[0154] S234: Feature Selection

[0155] By optimizing the weights to compress the influence of redundant features, high-dimensional feature selection and dimensionality reduction can be achieved.

[0156] In geopolymer composition optimization problems, features are often highly correlated, and direct modeling can easily lead to redundant interference. Regularized neighborhood composition analysis has the following significant advantages: it automatically determines the importance of variables through feature weight learning, without the need for pre-selection; it is insensitive to multicollinearity, and is particularly suitable for modeling indices related to physicochemical properties.

[0157] The regularized neighborhood component analysis model proposed in step S23 is used to further extract key feature dimensions between raw material ratios and performance response. This method avoids dimensional redundancy and overfitting while preserving local structure. The model significantly improves data separability and information retention after dimensionality reduction, providing a more effective feature input space for subsequent prediction models.

[0158] S24: Construct an artificial neural network model based on the Levenberg-Marquardt algorithm:

[0159] The Levenberg-Marquardt algorithm was used for training, which is highly efficient when dealing with networks containing hundreds of weights and biases.

[0160] Artificial neural network models based on the Levenberg-Marquardt algorithm satisfy the following conditions:

[0161] Network Structure: A typical three-layer feedforward neural network structure is adopted, including an input layer, hidden layers, and an output layer. The input layer receives multiple geopolymer composition parameters, including coal gangue content, fly ash content, slag content, water glass modulus, Na2O content, or NaOH solution concentration, such as: Coal gangue content (%) Fly ash content (%) Slag content (%) Water glass modulus; Dosage (%); the hidden layer uses an adjustable number of neurons (e.g., The hidden layer contains neurons (number of neurons) to uncover complex mapping relationships between features. Specifically, the number of neurons in the hidden layer can be determined through Bayesian optimization, with the objective function being the 5-fold cross-validation error and the search space being integers [4, 12]. The output layer outputs target performance parameters (such as compressive strength). ).

[0162] Weight Update: The weights are dynamically adjusted using the Levenberg-Marquardt algorithm. The update formula is as follows:

[0163]

[0164] in, Let be the weight vector for the r-th iteration; This is the Jacobian matrix, containing the partial derivatives of the network output with respect to the weights; This is the error vector; : Adjust the parameters (initially large, gradually decrease with iterations); : Identity matrix; Inverse operation; : Request the transpose operation.

[0165] Training mechanism: Levenberg-Marquardt combines two extreme cases:

[0166] when At that time, the algorithm is close to the Gauss-Newton method (fast convergence speed);

[0167] when When the value is large, the algorithm behaves as gradient descent (strong stability).

[0168] Continuously adjust during training This ensures that both convergence speed and stability are taken into account.

[0169] Early stopping mechanism: To prevent overfitting, an early stopping mechanism is introduced to ensure stability and reliability in practical applications. Specifically, during training, the validation set error is monitored, and training is terminated when the error fails to decrease for a preset number of consecutive iterations to prevent overfitting.

[0170] The artificial neural network weight update strategy based on the Levenberg-Marquardt algorithm, proposed in step S24, addresses the problems of slow convergence and susceptibility to local optima in traditional BP neural networks when learning complex stoichiometric data. The Levenberg-Marquardt optimization algorithm is employed to update network weights. This algorithm combines the advantages of gradient descent and Gauss-Newton's method, improving convergence speed while ensuring global stability. It significantly enhances the neural network model's ability to express complex nonlinear mapping relationships, making it suitable for predictive modeling tasks of multi-objective geopolymer performance.

[0171] The four machine learning models established in this invention operate collaboratively: the Rational-Minkowski kernel SVR captures local nonlinear relationships, the piecewise Gaussian process regression model handles the heterogeneity of the input space, the regularized neighborhood component analysis model suppresses high-dimensional noise, and the Levenberg-Marquardt algorithm accelerates ANN convergence. The four form a bias-variance balanced fusion framework.

[0172] S3. Calculate and assign weights based on the mean squared error of each machine learning model on the validation set in the dataset, and then perform weighted fusion to obtain the performance prediction result.

[0173] In practical regression problems, different regression models may exhibit varying predictive capabilities and error characteristics on different input features and datasets. To fully leverage the strengths of each model and avoid the limitations of a single model, a weighted fusion approach can be used to combine the prediction results of multiple models, thereby improving overall prediction performance, reducing overfitting, and enhancing model stability and robustness. In this invention, four different regression models are used in steps S21 to S24 for prediction. The core idea of ​​weighted fusion is to assign an appropriate weight to each regression model, thereby performing a weighted average of the model output results.

[0174] To ensure that the contribution of each model adapts to its predictive performance, the weight allocation strategy can be dynamically adjusted using error backpropagation. Specifically, the model weights should be adjusted based on each model's prediction error on the validation set, with models having smaller errors receiving larger weights.

[0175] Furthermore, in step S3, the weight calculation includes:

[0176] S31: Calculate the mean squared error of each machine learning model on the validation set:

[0177]

[0178] in, For the first The mean squared error of a machine learning model; It is the first The predicted value of each model for the i-th sample; It is the true value of the i-th sample;

[0179] S32: Based on mean square error Calculate the weight using the reciprocal :

[0180]

[0181] in, It is the first The weights of a machine learning model;

[0182] S33: Based on the weights of the four machine learning models The prediction results of the four machine learning models are weighted and averaged to output the final performance prediction value.

[0183] Through weighted fusion, the predictions of each model are effectively combined based on their performance on the validation set. If a model has a smaller error on the validation set, its prediction will have a significant impact on the final fused output. Weighted fusion of multiple models can improve prediction accuracy, especially when different models are complementary, where the fusion effect is particularly significant.

[0184] For example, in predicting the compressive strength of geopolymers, some regression models may have higher prediction accuracy in low-ratio ranges, while others perform better in high-ratio ranges. By using weighted fusion, the advantages of various models can be combined, avoiding performance fluctuations caused by changes in data characteristics of a single model.

[0185] S4. Receive the target performance parameters and initial raw material ratio range input by the user, use the trained geopolymer performance prediction model to predict the performance, and use the optimization algorithm to search in reverse for the best ratio combination that meets the target performance constraints.

[0186] Furthermore, in step S4: the user-defined target performance parameters include compressive strength > 80 MPa and permeability coefficient < 1 × 10⁻⁶. -7 cm / s, and durability indicators including sulfate corrosion resistance cycles and wet-dry cycle mass loss rate;

[0187] The initial raw material ratio range includes: coal gangue 60-80 wt%, slag 10-20 wt%, fly ash 10-20 wt%; water glass modulus 1.5-2.5, Na2O content 3-22 wt%; activator to solid raw material mass ratio 0.3-0.5;

[0188] The optimal combination of proportions that meets the target performance constraints is searched through back reasoning or optimization algorithms (such as particle swarm optimization or Bayesian optimization).

[0189] S5. Based on the optimal ratio combination, prepare a geopolymer to prepare a coal gangue-slag-fly ash geopolymer grouting material.

[0190] In existing technologies, coal gangue is difficult to use directly as a main raw material for geopolymers due to its uneven silicon and aluminum content and insufficient activity. This invention proposes a method for preparing geopolymer grouting materials using coal gangue as the main component, combined with slag and fly ash, to achieve efficient resource utilization of coal gangue and provide high-performance materials for seepage prevention layers in ecological sites. In some embodiments, step S5 specifically includes:

[0191] Raw material preparation: Typical coal gangue from Guizhou Province is selected and processed by crushing and ball milling to achieve a specific surface area of ​​400-600 m². 2 / kg; commercially available products are used for slag and fly ash, with particle size controlled between 0.045-0.075 mm.

[0192] Activator preparation: Mix water glass and Na2O or NaOH solution in a certain proportion according to the recommended optimal ratio to prepare a composite activator.

[0193] Solid raw material ratio: Prepare the amount of coal gangue, slag, and fly ash according to the recommended optimal ratio combination;

[0194] Mixing and stirring: Mix coal gangue, slag, and fly ash, add activator, and stir until the slurry fluidity reaches 150-200 mm to form a coal gangue-slag-fly ash geopolymer grouting slurry;

[0195] Grouting and curing: After the coal gangue-slag-fly ash geopolymer grout is pre-cured at room temperature for 24 hours, it is transferred to an environment of 20±2℃ and humidity ≥95% for 28 days to form coal gangue-slag-fly ash geopolymer grouting material.

[0196] Performance Verification and Feedback: Samples were prepared according to the recommended optimal formulation, and their compressive strength, permeability coefficient, and durability were tested to determine if they met the requirements of compressive strength > 80 MPa and permeability coefficient < 1 × 10⁻⁶. -7 Performance requirements such as cm / s; if the deviation between the measured value and the prediction exceeds 10%, the model parameters are triggered for iterative update.

[0197] According to the above embodiments of the present invention, by establishing a nonlinear regression model, the complex nonlinear relationship between different proportions (including coal gangue, fly ash, slag, water glass modulus, Na2O content, etc.) and geopolymer properties (compressive strength, permeability, durability) is modeled, and intelligent recommendation of raw material proportions is realized through model fusion, which significantly improves the comprehensive performance of geopolymers.

[0198] Figure 2 A schematic diagram of a machine learning-based geopolymer preparation and optimization system is shown. In some embodiments of the present invention, to achieve rapid application and intelligent proportion optimization of geopolymer materials in engineering practice, a machine learning-based geopolymer preparation and optimization system 200 is provided based on the aforementioned multi-model fusion prediction mechanism. This system 200 can automatically recommend corresponding raw material proportion combinations based on user-input target performance requirements and provide real-time feedback and performance prediction, significantly improving material design efficiency and intelligence. Figure 2 As shown, a machine learning-based geopolymer preparation and optimization system 200 includes the following steps:

[0199] The data acquisition module 210 is used to collect experimental data on the preparation of geopolymers, including solid raw material parameters, activator parameters and performance output indicators, and to perform missing value imputation, outlier removal, feature scaling, standardization, encoding and dataset partitioning on the experimental data.

[0200] The system receives two types of input through the user interface: target performance parameters (user input): compressive strength (MPa), corrosion resistance (optional indicators, such as salt spray test results, mass loss rate, etc.), impermeability, early strength, etc. (expandable support). Proportioning range constraints (user adjustable): upper / lower limits for material parameters such as fly ash content, slag ratio, activator type and concentration, and water-cement ratio.

[0201] Model building module 220 is used to perform nonlinear regression modeling of geopolymer performance based on four machine learning regression algorithms, construct a geopolymer performance prediction model integrating the four machine learning models, and establish a nonlinear mapping relationship between raw material ratio and performance, including:

[0202] Support vector regression model based on Rational-Minkowski kernel function;

[0203] The piecewise Gaussian process regression model divides the input space into sub-intervals, models each sub-interval independently, and then performs boundary weighted fusion.

[0204] The regularized neighborhood component analysis model learns feature weights by minimizing a loss function that includes an L2 regularization term, and defines a feature-weighted distance metric.

[0205] Artificial neural network model based on Levenberg-Marquardt algorithm;

[0206] The prediction output module 230 is used to calculate and assign weights based on the mean squared error of each machine learning model on the validation set in the dataset, and then perform weighted fusion to obtain the performance prediction result.

[0207] The ratio recommendation module 240 is used to receive the target performance parameters and the initial raw material ratio range input by the user, perform performance prediction using the trained geopolymer performance prediction model, and search for the best ratio combination that meets the target performance constraints through the optimization algorithm.

[0208] The proportion recommendation module 240 is performance-driven, mapping from "target performance" to "proportion recommendations" through the following steps: The user inputs a performance target, such as a desired compressive strength > 80 MPa. A fusion model is used to perform a reverse search for multiple initial proportion values ​​that meet this performance requirement (generated through optimization algorithms). The fusion model predicts the performance indicators of these proportion combinations, selecting the optimal solution or Pareto optimal solution (if multi-objective optimization is involved). The system outputs recommended raw material proportion combinations (including the proportion and amount of each material), predicted performance indicators and their confidence ranges, and multiple alternative schemes for the user to choose from or further adjust. The results are presented in charts and tables for the user to further filter or adjust the proportion parameters.

[0209] Grouting preparation module 250 is used to prepare geopolymer grouting material for coal gangue-slag-fly ash according to the optimal ratio combination.

[0210] Based on the aforementioned system, the user interface mainly includes the following functional sub-modules: Input area: used to input target performance values, acceptable error range, raw material types and proportion constraints, etc. Recommendation area: displays recommended formulation schemes in list format, supporting export and save functions. Performance prediction graph: displays the changing trends of various performance parameters under different formulations, supporting 2D / 3D visualization. Interactive adjustment area: users can drag sliders to adjust various material parameters, and the system automatically updates the prediction results. This interface can be implemented using web technologies (such as Flask + React) or desktop software (such as PyQt), possessing good scalability and compatibility, suitable for laboratory formulation design and engineering application scenarios. The system supports users manually modifying formulation parameters, and the system updates and re-predicts the corresponding performance indicators in real time, achieving a closed loop of human-computer interaction.

[0211] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0212] Furthermore, embodiments of this application also provide an electronic device, including: a processor, a memory, and a system bus; the processor and the memory are connected via the system bus; the memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the methods described above.

[0213] Furthermore, embodiments of this application also provide a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the methods described above.

[0214] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0215] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0216] It should also be noted that, in the embodiments of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0217] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in the embodiments of this application may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown in this application, but is to be accorded the widest scope consistent with the principles and novel features disclosed in the embodiments of this application.

Claims

1. A method for preparing and optimizing geopolymers based on machine learning, characterized in that, Includes the following steps: S1. Collect experimental data on the preparation of geopolymers, including parameters of solid raw materials, parameters of activators and performance output indicators, and perform missing value imputation, outlier removal, feature scaling, standardization, encoding and dataset partitioning on the experimental data; S2. Based on four machine learning regression algorithms, nonlinear regression modeling of geopolymer performance is performed to construct a geopolymer performance prediction model integrating the four machine learning models, establishing a nonlinear mapping relationship between raw material ratio and performance, including: Support vector regression model based on Rational-Minkowski kernel function; The piecewise Gaussian process regression model divides the input space into sub-intervals, models each sub-interval independently, and then performs boundary weighted fusion. The regularized neighborhood component analysis model learns feature weights by minimizing a loss function that includes an L2 regularization term, and defines a feature-weighted distance metric. Artificial neural network model based on Levenberg-Marquardt algorithm; S3. Calculate and assign weights based on the mean squared error of each machine learning model on the validation set in the dataset, and then perform weighted fusion to obtain the performance prediction result. S4. Receive the target performance parameters and initial raw material ratio range input by the user, use the trained geopolymer performance prediction model to predict the performance, and use the optimization algorithm to search in reverse for the best ratio combination that meets the target performance constraints. S5. Based on the optimal ratio combination, prepare a geopolymer to prepare a coal gangue-slag-fly ash geopolymer grouting material.

2. The method for preparing and optimizing geopolymers based on machine learning according to claim 1, characterized in that, in, The construction of the support vector regression model based on the Rational-Minkowski kernel function includes: Based on the Rational-Minkowski kernel function, an improved kernel function support vector regression model is constructed, where the expression for the Rational-Minkowski kernel function is: ; in, Let n be the sample input vector, and n be the number of samples. Minkowski distance The order norm, p∈[1,3], is used to adjust the norm order of the Minkowski distance; τ is the scaling parameter and τ>0, used to control the width of the kernel function; q is the exponential adjustment parameter and q∈[1,5], used to control the decay rate of the kernel function.

3. The method for preparing and optimizing geopolymers based on machine learning according to claim 2, characterized in that, in, The construction of the piecewise Gaussian process regression model includes: Input space partitioning: Based on principal component analysis and K-means clustering or expert knowledge rules, the geopolymer ratio input space is divided into M sub-intervals, each of which satisfies local data stability; Local model training: Gaussian process regression sub-models are trained independently for each sub-interval, and kernel function type and hyperparameters are selected for them; the kernel function includes the squared exponential kernel function, Matérn3 / 2 kernel function or RationalQuadratic kernel function; Segmented prediction ensemble: For a test sample, determine its sub-interval and use the corresponding sub-model for prediction; if it is located in a boundary region, the prediction results of adjacent sub-models are fused by weighted average. Hyperparameter optimization: Estimate the hyperparameters of each sub-model by maximizing the log-marginal likelihood function.

4. The method for preparing and optimizing geopolymers based on machine learning according to claim 3, characterized in that, in, If the location is in the boundary region, the prediction results of neighboring sub-models are fused by weighted average, including: ; ; in, : Predicted mean, representing the mean given input Conditional expected values ​​of geopolymer performance indices under given conditions; : Prediction variance, representing the variance of the predicted mean A quantitative measure of uncertainty or confidence level; Input samples for testing; This represents the total number of sub-intervals or sub-models. Indicates the first Sub-model exist The predicted mean at the location; Indicates the first Sub-models in The predicted variance at the location; For the first The spatial weights of each sub-model satisfy the normalization constraint: ; Weight according to Calculation of the inverse proportional function of distance to the center of each subinterval: ; in, Indicates test input sample With the Center of each sub-interval Euclidean distance.

5. The method for preparing and optimizing geopolymers based on machine learning according to claim 4, characterized in that, in, The estimation of hyperparameters for each sub-model by maximizing the log-marginal likelihood function includes: ; in, The training output vector, i.e., the target variable, includes compressive strength; The training input sample matrix, i.e., the geopolymer sizing parameters, has a total of There are samples, each sample has One input feature; The covariance matrix is ​​defined as follows: , by kernel function structure; : The variance of observation noise, used to characterize unexplainable noise in the data; : Identity matrix, dimension 1 , used to introduce noise.

6. The method for preparing and optimizing geopolymers based on machine learning according to claim 5, characterized in that, in, The construction of the regularized neighborhood component analysis model includes the following steps: Feature-weighted distance definition: Construct the weighted Manhattan distance metric function: ; in, : No. The input vector of each sample; : No. In the nth sample The values ​​of each feature; : No. Weight coefficients of each feature; distance metric It is a weighted Manhattan distance form; Neighborhood probability calculation: Define sample choose The probability of being a valid neighbor: ; : Indicates a sample Will The higher the probability value, the better the probability of it being a "valid neighbor". The more likely to be The predictions have an effect; : No. The smoothing scale parameter for each feature is used to normalize the distance values ​​of different feature dimensions; Model training: Feature weights are learned by minimizing a loss function that includes an L2 regularization term; Feature selection: By compressing the influence of redundant features through optimized weights, high-dimensional feature selection and dimensionality reduction are achieved.

7. The method for preparing and optimizing geopolymers based on machine learning according to claim 6, characterized in that, in, The artificial neural network model based on the Levenberg-Marquardt algorithm satisfies the following conditions: Network structure: A three-layer feedforward neural network is adopted. The input layer receives the aggregate ratio parameters, including coal gangue content, fly ash content, slag content, water glass modulus, and Na2O content. The number of neurons in the hidden layer is 2-15. The output layer predicts the target performance indicators. Weight update formula: The weights are dynamically adjusted using the Levenberg-Marquardt algorithm. The update formula is as follows: ; Let be the weight vector for the r-th iteration; This is the Jacobian matrix, containing the partial derivatives of the network output with respect to the weights; This is the error vector; To adjust the parameters; It is the identity matrix; Training mechanism: When When →0, the algorithm approximates the Gauss-Newton method to accelerate convergence; when As the value increases, the algorithm degenerates into gradient descent to enhance stability; Early stopping mechanism: During training, the validation set error is monitored, and training is terminated when the error does not decrease for a preset number of consecutive times to prevent overfitting.

8. The method for preparing and optimizing geopolymers based on machine learning according to claim 7, characterized in that, in, In step S3, the weight calculation includes: Calculate the mean squared error of each machine learning model on the validation set: ; in, For the first The mean squared error of a machine learning model; It is the first The predicted value of each model for the i-th sample; It is the true value of the i-th sample; According to the mean square error Calculate the weight using the reciprocal : ; in, It is the first The weights of a machine learning model; Based on the weights of the four machine learning models The prediction results of the four machine learning models are weighted and averaged to output the final performance prediction value.

9. The method for preparing and optimizing geopolymers based on machine learning according to claim 7, characterized in that, in, In step S4: the target performance parameters set by the user include compressive strength > 80 MPa and permeability coefficient < 1 × 10⁻⁶. -7 cm / s, and durability indicators including sulfate corrosion resistance cycles and wet-dry cycle mass loss rate; The initial raw material ratio range includes: coal gangue 60-80 wt%, slag 10-20 wt%, fly ash 10-20 wt%; water glass modulus 1.5-2.5, Na2O content 3-22 wt%; activator to solid raw material mass ratio 0.3-0.5; Step S5, according to the optimal ratio combination, prepares a geopolymer for coal gangue-slag-fly ash geopolymer grouting material, including: Coal gangue is crushed and ball-milled to a specific surface area of ​​400-600 m². 2 / kg, the particle size of slag and fly ash is controlled at 0.045-0.075 mm; According to the recommended optimal ratio combination, coal gangue, slag, and fly ash are mixed, an activator is added, and the mixture is stirred until the fluidity of the slurry reaches 150-200 mm to form a coal gangue-slag-fly ash geopolymer grouting slurry. Grouting and curing: After the coal gangue-slag-fly ash geopolymer grout is pre-cured at room temperature for 24 hours, it is transferred to an environment of 20±2℃ and humidity ≥95% for 28 days to form coal gangue-slag-fly ash geopolymer grouting material.

10. A machine learning-based system for preparing and optimizing geopolymers, characterized in that, Includes the following steps: The data acquisition module is used to collect experimental data on geopolymer preparation, including solid raw material parameters, activator parameters and performance output indicators, and to perform missing value imputation, outlier removal, feature scaling, standardization, encoding and dataset partitioning on the experimental data. The model building module is used to perform nonlinear regression modeling of geopolymer performance based on four machine learning regression algorithms, construct a geopolymer performance prediction model that integrates the four machine learning models, and establish a nonlinear mapping relationship between raw material ratios and performance, including: Support vector regression model based on Rational-Minkowski kernel function; The piecewise Gaussian process regression model divides the input space into sub-intervals, models each sub-interval independently, and then performs boundary weighted fusion. The regularized neighborhood component analysis model learns feature weights by minimizing a loss function that includes an L2 regularization term, and defines a feature-weighted distance metric. Artificial neural network model based on Levenberg-Marquardt algorithm; The prediction output module is used to calculate and assign weights based on the mean squared error of each machine learning model on the validation set in the dataset, and then perform weighted fusion to obtain the performance prediction result. The ratio recommendation module is used to receive the target performance parameters and the initial raw material ratio range input by the user, use the trained geopolymer performance prediction model to predict the performance, and use the optimization algorithm to search for the best ratio combination that meets the target performance constraints. The grouting preparation module is used to prepare geopolymer grouting materials for coal gangue-slag-fly ash according to the optimal ratio combination.

Citation Information

Patent Citations

  • Natural core polymer oil displacement recovery ratio prediction method based on machine learning

    CN117474158A

  • Perceptual concrete mix proportion optimization method based on interpretable machine learning

    CN119337706A

  • AU2020101453A4

Cited By

  • Pavement construction material ratio intelligent calculation and scheme generation system

    CN121215137A

  • An intelligent computing and scheme generating system for road construction material proportioning

    CN121215137B

  • Repair material performance prediction and formula optimization method based on machine learning algorithm

    CN121922286A

  • Oral absorption rate prediction and compound structure optimization method, device and equipment

    CN121938498A