An HPC compressive strength prediction method and system based on model fusion

Through model fusion-based methods and SHAP interpretability algorithms, the HPC compressive strength is accurately predicted, which solves the problems of relying on manual experience, large data demand and complex model training in the existing technology, and achieves high-precision, reliable and interpretable prediction results.

CN115497574BActive Publication Date: 2025-05-27CCCC SECOND HARBOR ENGINEERING CO LTD +1

Patent Information

Application Number
CN202211078389.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-05-27
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

The existing HPC compressive strength prediction methods rely on manual experience, have large data demand, complex model training, low accuracy of prediction results and lack interpretability, and are especially not suitable for high-risk and difficult data acquisition civil engineering scenarios.

Method used

Using a model fusion method, the first prediction model and the second prediction model are trained and the weighted average method is combined to obtain the fusion model of HPC compressive strength. At the same time, the model is interpreted and analyzed using the SHAP interpretability algorithm to measure the contribution value of each parameter to the compressive strength.

Benefits of technology

It improves the accuracy and reliability of compressive strength prediction, overcomes the problems of complexity and lack of interpretability in traditional model training, and provides more reliable and interpretable prediction results, suitable for civil engineering scenarios with high risk and difficulty in data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497574B_ABST
    Figure CN115497574B_ABST
Patent Text Reader

Abstract

The present invention discloses an HPC compressive strength prediction method and system based on model fusion, including collecting high-performance concrete related parameter data; exploratory data analysis and data cleaning; processing of concrete data outliers and data transformation; concrete data feature engineering; construction of a concrete compressive strength prediction model; parameter tuning of the concrete compressive strength prediction model; model fusion of the concrete predicted compressive strength prediction model; interpretability analysis of the concrete compressive strength prediction model based on SHAP; by using the method and system of the present invention, the defects that traditional neural network models are difficult to train and have a high demand for the amount of data are overcome; at the same time, the comprehensive average of the results of multiple models is more reliable compared with the prediction results using a single model; at the same time, the test period is short, the accuracy is high, and the test cost is low, and it has stronger engineering feasibility compared with the traditional empirical formula method and test method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large caisson construction, and particularly relates to a method and system for predicting the compressive strength of HPC based on model fusion. Background Technique

[0002] Due to its outstanding characteristics such as high strength and high durability, HPC (High-Performance Concrete) has been widely used in the construction of long-span bridges. As an important index for evaluating the quality of concrete, the compressive strength of concrete greatly reflects the safety performance of building structures. Therefore, studying an accurate prediction method for the compressive strength of high-performance concrete is of great significance for the accurate control of construction projects and the scientific evaluation of engineering projects.

[0003] At present, the methods for predicting compressive strength mainly include: empirical formula method, experimental method, and statistical machine learning-based methods.

[0004] Among them, the empirical formula method is based on artificial experience, and by establishing a complex mathematical model to fit various parameter indexes of concrete, a compressive strength calculation model is established. This type of method highly depends on artificial experience, the iterative calculation process is complex, the fitting accuracy is extremely limited, and it is not applicable to the calculation of the compressive strength of high-performance concrete with a large variety of mixed materials and extremely complex mixing ratio formula contents.

[0005] The experimental method is to monitor the structure of high-performance concrete materials before and after forming through various test instruments and equipment to obtain the compressive strength. This type of method has a high test accuracy and the test results have a high credibility; however, the test cycle of this type of method is long, the test cost is high, and it has a very high risk factor when testing in a complex construction environment on site. Therefore, it is more used for the compressive strength test in a laboratory environment.

[0006] The statistical machine learning-based method is based on machine learning theory. Through a data-driven method, without too many prerequisite assumptions, a compressive strength prediction model is directly established. It has low cost, short test cycle, and high test result accuracy, so it has extremely high research and application value.

[0007] So far, some studies have used machine learning methods to predict the compressive strength of concrete, such as methods based on the AdaBoost algorithm, random forest and intelligent algorithms, BP neural network or RBF neural network, support vector machine (SVM), linear regression (LR), deep learning (DL), etc. However, there are still some deficiencies in the above methods. For example, the above methods all predict the compressive strength based on the output results of a single model, lacking a certain degree of reliability. Methods based on artificial neural networks or deep learning especially rely on a large amount of experimental data and are not applicable to scenarios such as civil engineering with high risks and difficult data collection. At the same time, the model training of such methods is difficult, and it is easy to fall into local optima or overfitting. Methods based on SVM or LR are extremely vulnerable to outliers, resulting in a large difference between the prediction accuracy and the actual results, and it is difficult for the LR method to fit the complex linear relationship between the components of HPC. At the same time, the existing prediction methods are generally black-box models, lacking interpretability, and it is impossible to clearly know the specific impact of each data sample on the compressive strength, which is not conducive to the specific guidance of actual engineering projects. Summary of the Invention

[0008] Therefore, aiming at the problems of the current HPC (high-performance concrete) compressive strength prediction method, such as strong dependence on artificial experience, large data requirements, complex model training process, low accuracy of model prediction results, and lack of interpretability of model prediction results, the present invention proposes a method suitable for accurate prediction of the compressive strength of ordinary concrete, especially HPC.

[0009] A method for predicting the compressive strength of HPC based on model fusion to achieve one of the purposes of the present invention includes the following steps:

[0010] S1. Train the first prediction model and the second prediction model respectively according to the historical parameter data of the collected HPC to obtain the trained first prediction model and second prediction model for predicting the compressive strength of HPC; the HPC is high-performance concrete.

[0011] S2. Perform a composite operation on the trained first prediction model and the second prediction model to obtain a fusion model for predicting the compressive strength of HPC, and the fusion model for predicting the compressive strength of HPC outputs the final prediction result of the compressive strength of HPC.

[0012] A further technical solution includes that after step S2, there is also step S3:

[0013] S3. Use the first algorithm to interpret and analyze the fusion model for predicting the HPC compressive strength, and obtain the contribution value of the component content of each HPC parameter to the HPC compressive strength output by the fusion model for predicting the HPC compressive strength; the contribution value is used to measure whether the component content of HPC has the effect of enhancing the HPC compressive strength on the predicted value of the HPC compressive strength, and it is used to guide the actual concrete mix design.

[0014] A further technical solution includes: the first algorithm is a SHAP interpretability algorithm. The contribution value is represented by Shapley Value and is abbreviated as Ψ, and its definition is as follows:

[0015]

[0016] In the formula:

[0017] S is the subset of features input to the model to be explained, that is, the set of concrete parameters input to the fusion model for predicting the HPC compressive strength;

[0018] x j is the j-th feature variable of the sample to be explained; that is, the j-th concrete parameter;

[0019] p is the total number of features, that is, the total number of concrete parameters;

[0020] val x (S) represents the prediction result of the model for the sample x when S is the input feature, that is, the compressive strength result output by the model, where x is the sample, and the elements of x are denoted as x i ,x i is the value corresponding to the i-th feature variable;

[0021] The SHAP value represents the importance degree of the j-th feature to the model output result, that is, the marginal contribution, and the Shapley Value is the mean of each marginal contribution. The model interpretation result for the fusion model for predicting the HPC compressive strength is defined as the following formula:

[0022]

[0023] In the formula:

[0024] g is the compressive strength prediction model to be explained, that is, the fusion model for predicting the HPC compressive strength;

[0025] z' ∈ {0, 1} M is the combination vector, representing whether the feature z j (j ∈ [1, M]) exists, where z j is the j-th concrete parameter input to the fusion model for predicting the HPC compressive strength, and z' is used to identify z1 ~z M Exists in the parameter set input to the fusion model for HPC compressive strength prediction;

[0026] M is the number of combined features, that is, the number of concrete parameters input to the fusion model for HPC compressive strength prediction;

[0027] Is the Shapley Value of feature attribution for feature j, that is, the prediction result of the j-th parameter on the fusion model for HPC compressive strength prediction, that is, the contribution value to the compressive strength;

[0028] Ψ 0 Is the average prediction result of the fusion model for HPC compressive strength prediction, that is, the mean of the compressive strength prediction results.

[0029] The Shapley Value measures the contribution of the feature to the overall prediction result, Ψ j >0 indicates that the feature has a positive improvement effect on the predicted value of the compressive strength, that is, it has the effect of enhancing the HPC compressive strength.

[0030] The SHAP global feature importance is the average of the sum of the absolute values of the Shapley Value of each feature, that is

[0031] A further technical solution includes: in the step S2, the method for obtaining the fusion model for HPC compressive strength prediction includes performing a composite operation on the first prediction model and the second prediction model by using the weighted average method.

[0032] A further technical solution includes: the method for performing a composite operation on the first prediction model and the second prediction model by using the weighted average method includes:

[0033] minimiz e(Loss)s.t.w 1 +w 2 =1 and w 1 ≥0,w 2 ≥0

[0034] In the formula:

[0035] w 1 Represents the weight of the first prediction model;

[0036] w 2 Represents the weight of the second prediction model;

[0037] Loss is the loss function of the fusion model H(x) for HPC compressive strength prediction; its calculation method is as follows:

[0038]

[0039] In the formula:

[0040] N is the sample size, that is, the total number of concrete sample data collected;

[0041] is the actual compressive strength of the concrete corresponding to the i-th sample;

[0042] is the predicted value of the compressive strength of the concrete output by the fusion model H(x) for the compressive strength of the i-th sample in the prediction of HPC compressive strength;

[0043] The expression of H(x) is as follows:

[0044]

[0045] In the formula:

[0046] H(x): represents the finally predicted HPC compressive strength output by the fusion model for predicting HPC compressive strength;

[0047] w 1 and w 2 : respectively represent the weights of the first prediction model and the second prediction model;

[0048] h 1 and h 2 : respectively represent the compressive strengths of the concrete predicted by the first prediction model and the second prediction model.

[0049] A further technical solution includes: the first prediction model predicts the compressive strength of HPC based on the AdaBoost algorithm.

[0050] A further technical solution includes: the second prediction model predicts the compressive strength of HPC based on the CatBoost algorithm.

[0051] A further technical solution includes: when the first prediction model and the second prediction model are fused by the weighted average method, the weight of the second prediction model based on the CatBoost algorithm is greater than the weight of the first prediction model based on the AdaBoost algorithm.

[0052] A further technical solution includes: before the step S2, hyperparameter tuning is also performed on the first prediction model and the first prediction model by using the Bayesian optimization method; and cross-validation is performed on the model after parameter adjustment; the hyperparameters include the number of trees and depth, and the first prediction model and the second prediction model for predicting the compressive strength of HPC after hyperparameter tuning are obtained.

[0053] Both the AdaBoost model and the CatBoost model are tree-based ensemble models. Their base models are decision trees, that is, many base models (i.e., many trees) together constitute the ensemble models AdaBoost and CatBoost. AdaBoost and CatBoost as a whole are based on the Boosting ensemble learning framework. The number of trees in the AdaBoost model and the CatBoost model is the number of decision trees; the depth of the AdaBoost model and the depth of the CatBoost model are the number of layers of the AdaBoost model and the number of layers of the CatBoost model respectively.

[0054] The further technical solution includes: in step S1, new feature parameters are obtained from the collected historical parameter data of HPC through feature construction to expand the dataset and make the predicted compressive strength of HPC more accurate; the feature construction method is to combine engineering experience to perform mathematical calculations on different concrete parameters to obtain the ratio relationship of different concrete parameters.

[0055] A system for predicting the compressive strength of HPC based on model fusion to achieve the second object of the present invention includes a model training module and a model composite operation module;

[0056] The model training module is used to train the first model and the second model respectively according to the collected historical parameter data of HPC, and obtain the trained first prediction model and second prediction model for predicting the compressive strength of HPC respectively;

[0057] The model composite operation module is used to perform composite operations on the compressive strength of HPC output by the trained first prediction model and second prediction model for predicting the compressive strength of HPC to obtain a fusion model for predicting the compressive strength of HPC, and the fusion model for predicting the compressive strength of HPC outputs the final prediction result of the compressive strength of HPC.

[0058] Further, in the model composite operation module, the weighted average method is used to perform composite operations on the compressive strength of HPC output by the first prediction model and the second prediction model.

[0059] Further, it also includes a parameter tuning module, which is used to perform hyperparameter tuning on the trained first prediction model and second prediction model for predicting the compressive strength of HPC respectively to obtain the first prediction model and second prediction model with optimized parameters, and at the same time perform cross-validation on the tuned compressive strength prediction model; the cross-validation method includes five-fold cross-validation;

[0060] Further, it further includes a model interpretation and analysis module, which is used to interpret and analyze the fusion model for predicting the HPC compressive strength by using the first algorithm, and obtain the influence of the component content of each HPC parameter on the HPC compressive strength.

[0061] Further, it further includes an outlier processing module, which is used to detect outliers in the historical parameter data of the collected concrete; the method for outlier detection includes using an algorithm that combines K-Means++ clustering and isolation forest.

[0062] The steps of the K-Means++ algorithm are as follows:

[0063] a) Initialize an empty set M to store the initial cluster centers;

[0064] b) Randomly select the first cluster center μ (j) from the initial samples, and assign it to the set M;

[0065] c) For each sample x (i) (this sample does not belong to the set M), calculate its minimum squared distance d(x (i) , M) 2 ;

[0066] d) Randomly select the next centroid μ based on the weighted probability distribution (p) ;

[0067] e) Repeat the above steps b and c until K cluster centers are selected;

[0068] f) Based on the above set M, continue to use the classic K-Means algorithm;

[0069] g) Select the K-Means model with the best performance according to the SSE, that is, the sum of squared residuals, so as to obtain the best cluster centers;

[0070] Combined with the above K cluster centers, use the isolation forest algorithm to detect outliers in the data set.

[0071] The isolation forest algorithm is an outlier detection algorithm based on partitioning and ensemble learning. If the isolation forest algorithm is directly used for outlier detection without prior clustering analysis, problems such as large computational volume, long running cycle, and too strong artificiality in the partitioning process will be faced. Conducting data clustering analysis based on the K-Means++ algorithm can greatly improve the detection efficiency of the isolation forest algorithm.

[0072] Beneficial effects:

[0073] (1) The concrete compressive strength prediction model is established by means of ensemble learning, and the ensemble model is further fused again. This modeling method of group decision-making not only improves the prediction accuracy of the model, but also overcomes the defects that traditional neural network models are difficult to train and have high requirements for the amount of data. At the same time, compared with the prediction results of using a single model, the comprehensive average of multiple model results is more reliable.

[0074] (2) The HPC compressive strength prediction method based on statistical machine learning methods has a short test cycle, high accuracy, and low test cost at the same time, and has stronger engineering feasibility than traditional empirical formula methods, test methods, etc.

[0075] (3) The present invention combines model fusion with the SHAP interpretability algorithm for the prediction of concrete compressive strength, overcomes the unpredictability and black box nature of the traditional method in the modeling process, and more combines SHAP interpretability analysis to provide convenience for the development of concrete engineering, which is conducive to more accurately understanding the specific influence of each component on the compressive strength.

[0076] (4) In the process of handling outliers in concrete parameter data, the method of combining K-Means++ clustering and isolation forest is adopted, which overcomes the high computational complexity and artificial dependence on data partitioning faced by directly using the isolation forest method for processing.

[0077] (5) The present invention realizes the expansion of the dataset through feature engineering means, and newly constructs proportional features such as water-cement ratio and water-binder ratio, which is conducive to avoiding model overfitting. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 is the modeling flowchart of the HPC compressive strength prediction model of the present invention;

[0079] Figure 2 is the model fitting effect of the HPC compressive strength prediction model of the present invention Figure 1 ;

[0080] Figure 3 is the model fitting effect of the HPC compressive strength prediction model of the present invention Figure 2 ;

[0081] Figure 4 is the schematic diagram of the model fusion process of the HPC compressive strength prediction model of the present invention;

[0082] Figure 5 is the schematic diagram of the single-sample interpretation result of the SHAP model interpretability algorithm of the present invention Figure 1 ;

[0083] Figure 6 is the schematic diagram of the single-sample interpretation result of the SHAP model interpretability algorithm of the present inventionFigure 2 ;

[0084] Figure 7 This is a schematic diagram of the interpretation results of the SHAP model interpretability algorithm of the present invention on the entire data set. Detailed implementation manners

[0085] The following detailed implementation manners are used to interpret the technical solutions of the claims of the present invention, so that those skilled in the art can understand this claim book. The protection scope of the present invention is not limited to the following specific implementation structures. Those made by those skilled in the art that include the technical solutions of the claim book of the present invention and are different from the following detailed implementation manners are also within the protection scope of the present invention.

[0086] As Figure 1 shown, this embodiment includes the following steps:

[0087] Step 1: Collect relevant parameter data of high-performance concrete. Collect concrete-related parameter data at the site of concrete factories, mixing plants, etc. The concrete-related parameter data includes but is not limited to cement content, fly ash content, slag content, water reducer content, coarse / fine aggregate content, water content, curing period, temperature, slump, and form a sample set. Among them, the units of cement content, fly ash content, slag content, water reducer content, coarse / fine aggregate content, and water content are all (kg / m 3 ), that is, the mass of the corresponding components per cubic meter of concrete; the unit of the curing period is day (number of days), the unit of temperature is °C (degrees Celsius), and the unit of slump is mm (millimeters), all of which can be weighed or measured when configuring concrete;

[0088] At the same time, take the actual compressive strength corresponding to each sample as the target variable. Thus, an experimental data set is constructed. After the data construction is completed, store the data set on the local disk or in a relational database.

[0089] Step 2: Exploratory data analysis and data cleaning

[0090] Use means such as visualization and statistical analysis to conduct a preliminary exploration and understanding of the dataset. Combining the visualization analysis results, process the missing values, duplicate values, outliers, etc. in the collected data of high-performance concrete-related parameters, and at the same time understand the data distribution. For feature variables with fewer missing values (within 10%) and relatively small differences in corresponding feature extreme values, the missing values are filled with the mean; for feature variables with fewer missing values and relatively large differences in corresponding feature extreme values, the missing values are filled with the median; for feature variables with a missing value ratio reaching about 10%-50%, the missing values are predicted and filled using the decision tree algorithm; feature variables with a missing value ratio reaching more than 50% are excluded; duplicate values in the dataset are deleted;

[0091] Step 3. Handling outliers and data transformation of concrete data

[0092] Step 3.1. Data transformation

[0093] Normalize the feature variables with relatively large differences in feature extreme values. In this embodiment, the normalization process uses the Robust Scaler method that is robust to outliers. The steps of this method are as follows:

[0094] a) Calculate the quantiles of the data to be processed, where the quantile (i.e., the median) is removed, and then the corresponding quantiles are stored;

[0095] b) Calculate the IQR, which is defined as the difference between the quantile and the quantile;

[0096] c) Scale the feature variables using the IQR to achieve a unified scale;

[0097] According to the data visualization results, perform a logarithmic transformation on the dataset or feature variables that do not meet the current algorithm's inductive bias, that is, take the logarithm after adding 1 to the corresponding feature variable values, so that the distribution of each feature variable is closer to a normal distribution, and avoid the adverse impact of the skewness of the data distribution on the model prediction results;

[0098] Step 3.2. Outlier handling

[0099] Handle the outliers in the dataset using a combination of box plots and clustering + isolation forests. Among them, the box plot boxplot is used for the effect comparison before and after outlier handling and the confirmation of the handling benefit; the outlier detection adopted in this embodiment is based on the K-Means++ clustering algorithm + isolation forest. The K-Means++ algorithm generates more consistent results than the traditional K-Means by placing the initial centroids far from each other.

[0100] Step 4, Concrete Data Feature Engineering

[0101] Since the original dataset collected in Step 1 only contains the numerical contents of the components that make up HPC, without specific content ratio relationships, the dataset is further expanded through feature construction (i.e., mathematical calculations are performed on the numerical features of different components in combination with engineering experience). For example, the ratio relationship between the original feature "water content" and "cement content" is calculated to obtain the water-cement ratio; the ratio relationship between "water content" and "(cement content + slag content + fly ash content), abbreviated as gel" is calculated to obtain the water-binder ratio. Through feature engineering, the newly constructed features are shown in Table 1 below:

[0102]

[0103] Table 1

[0104] Step 5, Construction and Evaluation of Concrete Compressive Strength Prediction Model

[0105] Based on the dataset after the above preprocessing and feature engineering, the training set and test set are divided in a ratio of 8:2. Based on the Boosting framework, the first prediction model based on the AdaBoost algorithm and the second prediction model based on the CatBoost algorithm are established respectively. At the same time, combined with 5-fold cross-validation, the performance of the first prediction model and the second prediction model is evaluated using the regression model evaluation indicators. Among them, the evaluation indicators adopted in this embodiment are defined as follows:

[0106]

[0107]

[0108]

[0109]

[0110]

[0111] Where:

[0112] N is the total number of concrete sample data collected;

[0113] i is the sample number, that is, which sample;

[0114] is the predicted compressive strength value of the model for the i-th sample, that is, the predicted value;

[0115] is the actual compressive strength value corresponding to the i-th sample, that is, the observed value;

[0116] Step 6: Parameter Tuning of the Concrete Compressive Strength Prediction Model

[0117] Combined with the evaluation results of the above model performance, the hyperparameters of the first prediction model and the second prediction model are tuned using the Bayesian optimization method to further improve the model performance. Among them, some key hyperparameters to be adjusted include the number of base model trees and the depth of the ensemble model trees. The process of Bayesian optimization takes the model error as the objective function and finds the parameters corresponding to the minimum error through parameter combinations.

[0118] Step 7: Fusion of the First Prediction Model and the Second Prediction Model

[0119] As Figure 2 shown in the fitting effect diagram of the second prediction model based on the CatBoost algorithm on the test set, Figure 3 shown in the fitting effect diagram of the first prediction model based on the AdaBoost algorithm on the test set. It can be seen from the figure that the prediction effect of the second prediction model based on the CatBoost algorithm on the compressive strength is significantly better than that of the first prediction model based on the AdaBoost algorithm;

[0120] To improve the reliability of the model prediction results, the results predicted by multiple models are integrated using the method of group decision-making, and fusion is carried out at the model decision-making level to improve the accuracy of the model prediction results. As Figure 4 shown, in this embodiment, the weighted average method is used to fuse the CatBoost model and the AdaBoost model, and a greater weight is given to the CatBoost model during the fusion process.

[0121] The weighted average model fusion process is as follows:

[0122] Take the first prediction model based on the AdaBoost algorithm and the second prediction model based on the CatBoost algorithm as the base models h(x), and denote the fusion model for predicting the HPC compressive strength as H(x), which is expressed as follows:

[0123]

[0124] In the formula:

[0125] w i : represents the weight of the i-th base model. In this embodiment, w i ≥0 and satisfies

[0126] h i (x): represents the compressive strength of the HPC predicted by the i-th base model;

[0127] T: It represents the number of base models for model fusion. In this embodiment, the base models are AdaBoost and CatBoost, so T is equal to 2;

[0128] H(x): It represents the final prediction result of the compressive strength of concrete.

[0129] Among them, the fusion weight w of the base model i is determined as follows:

[0130] Combined with the definition of RMSE, the loss function of the fusion model is defined as follows:

[0131]

[0132] In the formula:

[0133] N is the sample size, that is, the total number of concrete sample data collected;

[0134] is the actual compressive strength of concrete corresponding to the i-th sample;

[0135] is the predicted value of the compressive strength of concrete of the i-th sample by the fusion model H(x) for the prediction of HPC compressive strength;

[0136] Therefore, the final optimization objective of the fusion model is defined as follows:

[0137] minimize (Loss) s.t. w 1 + w 2 = 1 and w 1 ≥ 0, w 2 ≥ 0 Equation (8)

[0138] In the formula:

[0139] Loss is the loss function of the ensemble model H(x);

[0140] s.t is the abbreviation of the constraint condition;

[0141] w 1 and w 2 are the fusion weights of the base models CatBoost and AdaBoost;

[0142] minimize means to minimize;

[0143] By solving the constrained minimum optimization problem shown in Equation (8), the fusion weight w of the base model is obtained.

[0144] Step 8. Interpretability analysis of the concrete compressive strength prediction model based on SHAP

[0145] Through the aforementioned model construction and evaluation, model parameter adjustment, and model fusion, a concrete compressive strength prediction model with good prediction ability is obtained. Further, in this embodiment, the SHAP model interpretability algorithm is combined to interpret and analyze the model prediction results, so as to better understand the influence of each eigenvalue on the concrete compressive strength predicted by the model. This influence is represented by the Shapley Value, briefly denoted as Ψ, and is defined as follows:

[0146]

[0147] In the formula:

[0148] S is the subset of features input to the model to be explained, that is, the set of concrete parameters input to the fusion model for HPC compressive strength prediction;

[0149] x j is the j-th feature variable of the sample to be explained; that is, the j-th concrete parameter;

[0150] p is the total number of features, that is, the total number of concrete parameters;

[0151] val x (S) represents the prediction result of the model to be explained for the sample x when S is the input feature, that is, the compressive strength result output by the model to be explained, where x is the sample, and the elements of x are denoted as x i ,x i is the value corresponding to the i-th feature variable;

[0152] The SHAP value is the degree of importance of the j-th feature for the output result of the model to be explained, that is, the marginal contribution, and the Shapley Value is the mean of each marginal contribution. The model interpretation result for the model to be explained is defined as follows:

[0153]

[0154] In the formula:

[0155] g is the model to be explained, which in this embodiment is the fusion model H(x) for HPC compressive strength prediction after the fusion of CatBoost and AdaBoost;

[0156] z' ∈ {0, 1} M is the combination vector, representing whether the feature z j (j ∈ [1, M]) exists, where z j is the j-th concrete parameter input to the fusion model for HPC compressive strength prediction in this embodiment, and z' is used to identify whether z 1 ~z M exists in the parameter set input to the fusion model for HPC compressive strength prediction;

[0157] M is the number of combined features, that is, the number of input parameters input to the model g to be explained;

[0158] is the Shapley Value of feature attribution for feature j, that is, the prediction result of the j-th parameter for the model g to be explained. In this embodiment, it is the contribution value to the compressive strength;

[0159] Ψ 0 is the average prediction result of the model g to be explained. In this embodiment, it is the mean of the compressive strength prediction results output by the fusion model for HPC compressive strength prediction.

[0160] The Shapley Value measures the contribution of a feature to the overall prediction result, Ψ j > 0 indicates that the feature has a positive improvement effect on the predicted value, that is, it has the effect of enhancing the compressive strength.

[0161] The SHAP global feature importance is the average of the sum of the absolute values of the Shapley Values of each feature, that is Among them, for the partial graph representation example of using the SHAP algorithm for model interpretability analysis in this embodiment, see Figures 5 to 7 ;

[0162] Figure 5 is the SHAP local interpretation of the first sample, Figure 6 is the SHAP local interpretation of the tenth sample; in the figure, the model average prediction result is 35.25 Mpa, the model's compressive strength prediction result for the first sample is 76.17 Mpa. The model's compressive strength prediction result for the 10th sample is 38.73 Mpa. Based on the model prediction results, actual observation results, and the value conditions of each parameter, it is possible to provide guidance for the design of concrete mix ratios.

[0163] Figure 7 is the SHAP global interpretation summary graph. The Y-axis arranges the feature factors affecting the compressive strength in descending order of contribution degree from top to bottom; among them, the curing period, water-cement ratio, and cement content are the top three factors that have a significant impact on the HPC compressive strength, followed by water, water-binder ratio, etc. The X-axis is the average influence value of each factor on the prediction result of the compressive strength prediction model. In the current experimental results, with the increase of the curing period, the compressive strength will increase by an average of 8 Mpa.

[0164] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0165] The embodiment of the present application also provides an embodiment of the system, including a model training module and a model composite operation module;

[0166] The model training module is used to train the first model and the second model respectively according to the collected historical parameter data of the HPC, and obtain the trained first prediction model and the second prediction model for predicting the compressive strength of the HPC respectively;

[0167] The model composite operation module is used to perform composite operation on the compressive strength of the HPC output by the trained first prediction model and the second prediction model for predicting the compressive strength of the HPC, obtain a fusion model for predicting the compressive strength of the HPC, and the fusion model for predicting the compressive strength of the HPC outputs the prediction result of the final compressive strength of the HPC.

[0168] In the model composite operation module, the weighted average method is used to perform composite operation on the compressive strength of the HPC output by the first prediction model and the second prediction model.

[0169] In another embodiment, it further includes a model interpretation and analysis module, which is used to interpret and analyze the fusion model for predicting the compressive strength of the HPC by using the first algorithm, and obtain the influence of the component content of each HPC parameter on the compressive strength of the HPC.

[0170] The content not detailed in this specification belongs to the prior art well known to those skilled in the art.

Claims

1. A method for predicting the compressive strength of HPC based on model fusion, characterized in that, it includes the following steps: S1. Train the first prediction model and the second prediction model respectively according to the historical parameter data of the collected HPC, and obtain the trained first prediction model and the second prediction model for predicting the compressive strength of HPC; S2. Perform a composite operation on the trained first prediction model and the second prediction model to obtain a fusion model for predicting the compressive strength of HPC, and the fusion model for predicting the compressive strength of HPC outputs the final prediction result of the compressive strength of HPC; In the step S2, the method for obtaining the fusion model for predicting the compressive strength of HPC includes performing a composite operation on the first prediction model and the second prediction model by using the weighted average method; The first prediction model predicts the compressive strength of HPC based on the AdaBoost algorithm; the second prediction model predicts the compressive strength of HPC based on the CatBoost algorithm; When performing a modulo composite operation on the first prediction model and the second prediction model by using the weighted average method, the weight of the second prediction model based on the CatBoost algorithm is greater than the weight of the first prediction model based on the AdaBoost algorithm; The calculation method of the weight of each prediction model in the weighted average method includes: by solving the constrained minimum optimization problem shown in the following formula, the weights of the first prediction model and the second prediction model are obtained: minimize(Loss)s.t.w 1 +w 2 =1 and w 1 ≥0,w 2 ≥0; In the formula: w 1 and w 2 correspond to the weights of the first prediction model and the second prediction model, respectively; Loss is the loss function for the following H ( x ) H ( x ) = (w 1 h 1 (x) + w 2 h 2 (x)); In the formula: H ( x ):Represents the final predicted HPC compressive strength output by the fusion model for HPC compressive strength prediction; w 1 、w 2 : represent the weights of the first prediction model and the second prediction model respectively; h 1 h(x) 2 h(x) represent the HPC compressive strengths predicted by the first prediction model and the second prediction model, respectively.

2. The method for predicting the compressive strength of HPC based on model fusion according to claim 1, characterized in that, after the step S2, it further includes a step S3: S3. Use the first algorithm to interpret and analyze the fusion model for predicting the compressive strength of HPC, and obtain the contribution value of the component content of each HPC parameter to the compressive strength of HPC output by the fusion model for predicting the compressive strength of HPC.

3. The method for predicting the compressive strength of HPC based on model fusion according to claim 2, characterized in that, the first algorithm is a SHAP interpretability algorithm.

4. A system for predicting the compressive strength of HPC based on model fusion using the method for predicting the compressive strength of HPC based on model fusion according to claim 1, characterized in that, it includes: a model training module and a model composite operation module; The model training module is used to train the first model and the second model respectively according to the historical parameter data of the collected HPC, and obtain the trained first prediction model and the second prediction model for predicting the compressive strength of HPC; The model composite operation module is used to perform a composite operation on the trained first prediction model and the second prediction model to obtain a fusion model for predicting the compressive strength of HPC, and the fusion model for predicting the compressive strength of HPC outputs the final prediction result of the compressive strength of HPC.

5. The system for predicting the compressive strength of HPC based on model fusion according to claim 4, characterized in that, It further includes a model interpretation and analysis module, which is used to interpret and analyze the fusion model for predicting the HPC compressive strength by using the first algorithm, and obtain the contribution value of the component content of each HPC parameter to the HPC compressive strength output by the fusion model for predicting the HPC compressive strength.

6. The HPC compressive strength prediction system based on model fusion according to claim 4, wherein, in the model composite operation module, the weighted average method is used to perform composite operation on the HPC compressive strengths output by the first prediction model and the second prediction model.

Citation Information

Patent Citations

  • Method of predicting compression strength of super high-early-strength concrete

    CN107133446A

  • Method of reinforced cementitious construction by high speed extrusion printing and apparatus for using same

    CN109923264A

Cited By

  • A method for constructing a strength prediction model of phosphogypsum lightweight aggregate concrete

    CN122778868A