Analysis method of the effect of ultra-high performance concrete components on dynamic compressive strength

By improving the data set and building a stacked Kraft model, combining CatBoost and random forest-based learners and linear regression element learners, the accuracy and robustness problems of UHPC dynamic compressive strength prediction are solved, and high-precision dynamic intensity prediction and feature impact analysis are achieved.

CN118866174BActive Publication Date: 2025-05-16DALIAN NATIONALITIES UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410879509.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2025-05-16
Estimated Expiration
2044-07-02

AI Technical Summary

Technical Problem

现有技术在预测超高性能混凝土(UHPC)的动态抗压强度时,缺乏对成分影响的全面考虑,导致预测精度和鲁棒性不足。

Method used

The isolated forest anomaly detection algorithm is used to improve the data set, a two-layer stacked Kraft model is built, combined with CatBoost and random forests as the basis learners, and a linear regression meta learner is used to optimize, and a high-precision prediction model is established through interpretable artificial intelligence technologies such as SHAP and PDP.

Benefits of technology

The prediction accuracy of UHPC dynamic compressive strength was significantly improved, and the RMSE was reduced from 24.07 to 18.53, and the feature impact relationship was accurately quantified, providing theoretical support for a single UHPC structural design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118866174B_ABST
    Figure CN118866174B_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of ultra-high performance concrete, and specifically discloses an analysis method for the influence of ultra-high performance concrete components on dynamic compressive strength, comprising: S1, obtaining test data, S2, anomaly detection based on isolation forest, S3, constructing a two-layer stacking-CARF model, S4, high-precision prediction of the model and S5, quantifying the influence of UHPC components on dynamic compressive strength; the isolation forest anomaly detection algorithm of the invention improves the data set by eliminating or reducing abnormal data, thereby significantly improving the prediction accuracy of the ML model, using CatBoost and RF as the base learner of the first layer, and LR as the meta learner of the second layer, the established stacking-CARF model has higher prediction accuracy and robustness compared with a single model and an EL model, LIME has stronger interpretability for specific samples, and can provide a theoretical reference for the design and performance optimization of a single UHPC structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of ultra-high performance concrete, and in particular relates to an analysis method for the influence of ultra-high performance concrete components on dynamic compressive strength. Background Art

[0002] Engineering structures may be subjected to extreme impact loads and blast loads during their service life, so the research on the dynamic compression performance of ultra-high performance concrete (UHPC) structures within a specific strain rate range has received increasing attention. However, traditional research is usually based on experiments and empirical formulas, which has certain limitations in terms of prediction accuracy and quantification of influencing factors.

[0003] Due to the challenges associated with traditional experimental studies, including the need for specialized instruments, technical complexity, high cost, and the fact that it can only reflect the dynamic compression performance of concrete under specific test conditions, some researchers have taken a different approach. They use existing test results as a benchmark to build verified numerical simulation models to describe the dynamic mechanical behavior of UHPC under impact. For example, Rong et al. used the finite element method (LS-DYNA) to simulate the entire impact process of UHPC. The Johnson_Holmquist_Concrete material constitutive model proposed by Johnson and Cook was verified to be applicable to the dynamic compression of UHPC by comparison with experimental results. Similarly, Xu et al. studied the dynamic material properties of steel fiber reinforced concrete (SFRC) through numerical simulation, combined fibers, aggregates, and cement mortar, established an axisymmetric mesoscale SFRC model, and conducted experimental verification.

[0004] On the other hand, researchers have proposed some analytical prediction models with various theoretical foundations based on the understanding of the dynamic mechanical properties of UHPC. For example, Ren et al. studied the effectiveness of traditional empirical formulas in describing the dynamic impact factors (DIFs) of UHPC. In view of the lack of universality of the traditional DIF model for UHPC materials, this study proposed a new empirical formula for calculating the DIFs of UHPC. Yu et al. systematically investigated the influence of different influencing factors on the dynamic compressive strength of UHPC. Based on the existing analysis, a dynamic strength prediction model considering strain rate and steel fiber volume was proposed. However, both numerical simulation and analytical prediction models still have limitations when studying the mechanical properties of UHPC. Due to the lack of consideration of all influencing factors, most models cannot accurately characterize the highly nonlinear relationship between UHPC composition and dynamic compressive strength.

[0005] In this regard, the inventors proposed an analysis method for the influence of ultra-high performance concrete components on dynamic compressive strength to solve the above problems. Summary of the invention

[0006] The purpose of the present invention is to provide an analysis method for the influence of ultra-high performance concrete components on dynamic compressive strength, so as to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The method for analyzing the effect of ultra-high performance concrete components on dynamic compressive strength includes the following steps:

[0009] S1. Acquire test data and collect ultra-high performance concrete (UHPC) dynamic compression data from the public literature on split-Hopkinson bar (SHPB) experiments as the initial data set;

[0010] S2, anomaly detection based on isolation forests, improves the dataset by eliminating or mitigating abnormal data, thereby avoiding poor prediction accuracy of the model due to data distribution defects and the presence of outliers;

[0011] S3. Construct a two-layer stacking-CARF model, use the first layer classification boosting, CatBoost and random forest, RF base learner method to make high-precision prediction of UHPC dynamic compressive strength, and use base learners for optimization and screening;

[0012] S4, high-precision prediction of the model, using the second-layer linear regression meta-learner to better learn the prediction results of the base learner to achieve high-precision prediction of UHPC dynamic strength;

[0013] S5. Quantify the influence of UHPC composition on dynamic compressive strength, and use explainable artificial intelligence (XAI) technology to assist stacked Kraft model prediction, including comparative analysis using Sharpless additive interpretation (SHAP) and partial dependence plot (PDP) to quantify the influence of single features and multiple interacting features on the global model. The local interpretable model (LIME) is used to supplement the local interpretation of the model prediction process, providing support for the design and performance optimization of individual ultra-high performance concrete (UHPC) structures.

[0014] Preferably, the input characteristic variables of the data set include basic material composition parameters of UHPC, namely cement, silica fume, fly ash, slag, sand and gravel, steel fiber, strain rate, high-efficiency water reducer and water content, and the output characteristic variable is the dynamic compressive strength of UHPC.

[0015] Preferably, the CatBoost combines multiple weak learners into a strong learner. The algorithm can adjust the sample weight of the next model learning data according to the learning results of the previous model, and construct the model in a step-by-step iterative manner. The weak learners constructed in each step of the iteration are to make up for the shortcomings of the existing model.

[0016] Preferably, the CatBoost divides the original dynamic compressive strength dataset into n training subsets and evenly distributes the corresponding sample weights Then, the model e1 is trained using the sample subset with sample weight d1. The error rate of the training model is calculated as follows:

[0017]

[0018] Where I is the accuracy of model classification, where 1 indicates correct classification and 0 indicates incorrect classification; each training model e i The weight coefficients are updated according to the error rate of the training model as follows:

[0019]

[0020] Next, each sample weight is adjusted according to the updated weight coefficient as follows:

[0021] d i =d i exp(2W1I(y i ≠e1(x i )))(i=1,2,...,n)

[0022] Finally, repeat the above process l times and set n training models e1, e2, ..., e n The weighted average of is taken as the final output of the model, and the output result is as follows:

[0023]

[0024] Preferably, the RF introduces a random subspace method, that is, randomly sampling features when constructing a regression decision tree, thereby alleviating the problem of overfitting and low precision of a single decision tree model. After selecting the partitioning features, an optimal eigenvalue cutting point is selected to partition the nodes.

[0025] Preferably, the CatBoost and RF models used as base learners are subjected to Bayesian automatic parameter optimization to obtain a model with the best hyperparameter value for training, the meta-features of the obtained training set and test set are merged as new training set and test set, the second layer is trained on the new training set as a linear regression model of the meta-learner, and the model is used to generate the final prediction result.

[0026] Preferably, the SHAP analysis interprets the predicted value of the model as the sum of the attributed values ​​of each feature, and the predicted value of the i-th sample is expressed as follows:

[0027]

[0028] Among them, y base is the predicted mean of all samples, p is the number of input features, is the SHAP value of the input feature, expressed as follows:

[0029]

[0030] Among them, S is a subset of features, x i is the feature vector of the sample, is the weight of subset S.

[0031] Preferably, PDP infers the relationship between features and predicted values ​​based on all samples in the data set. PDP is used to reveal how a single variable or the interaction of multiple variables affects the prediction results of the model. The marginal effect of the corresponding feature on the model output is expressed as follows:

[0032]

[0033] where x S is the feature that needs to be plotted, x C Except x S Other features besides is x C The true value of , k is the sample size in the data set. At the same time, the Monte Carlo method points out that the average marginal effect of the target feature on the prediction is estimated by calculating the average value in the training data. Overall, PDP can intuitively display the nonlinear relationship between a single feature and the output variable, helping to better understand the model behavior, the relationship between features, and the degree of influence of features on prediction.

[0034] Preferably, the LIME is used to explain the prediction of a single sample by a black box model, that is, firstly, a perturbation needs to be performed near a single sample point to generate a new data set around a specific sample, then a model is trained on the new data set, including a regression or decision tree model, and finally the model is used to replace the original model for local explanation, which is expressed as:

[0035]

[0036] in, is the original model to be explained, G is a set of simple models for explanation, g is the explanation model of instance x, Ω(g) is the complexity of the model, and the minimization loss L measures the difference between the explanation model g and the original model The closeness of the prediction, π x Represents the distance between instance x and the sample generated by the perturbation. The closer to the instance, the higher the weight.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) The quality of the dataset in the present invention is crucial to the prediction accuracy of the ML model. The isolation forest anomaly detection algorithm improves the dataset by eliminating or reducing abnormal data, thereby significantly improving the prediction accuracy of the ML model. The prediction results of the test set show that by removing 33% of the discrete abnormal data in the UHPC dynamic compression dataset, the RMSE value can be reduced from 24.07 to 18.53.

[0039] (2) The present invention uses CatBoost and RF as the base learners of the first layer, and LR as the meta-learner of the second layer. The stacking-CARF model established has higher prediction accuracy and robustness than the single model and EL model. Its prediction performance indicators can reach MAE=9.770, RMSE=15.865, R 2 =0.926, which can achieve high-precision prediction of the dynamic strength of UHPC.

[0040] (3) The comparative analysis of SHAP and PDP in the present invention can reveal relationships that are difficult to quantify by traditional experimental and theoretical analysis, such as feature importance ranking, marginal effects of features, and interactions. For example, in the low cement content range (<0.69), the effect of auxiliary cementitious materials on dynamic compressive strength plays a dominant role and can compensate for the loss of dynamic strength caused by low cement content to a certain extent. In addition, compared with the global explanation, LIME is more interpretable for specific samples and can provide a theoretical reference for the design and performance optimization of a single UHPC structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A flow chart of the method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength of the present invention;

[0042] Figure 2 It is the influence diagram of the pollution ratio on the prediction performance of the present invention;

[0043] Figure 3 It is the SF two-dimensional partial dependence grid map of the present invention;

[0044] Figure 4 It is the SF contour map of the present invention;

[0045] Figure 5 It is the FA two-dimensional partial dependence grid diagram of the present invention;

[0046] Figure 6 It is the FA contour map of the present invention;

[0047] Figure 7 is the S l ag two-dimensional partial dependence grid graph of the present invention;

[0048] Figure 8 is the S l ag contour map of the present invention;

[0049] Fig. 9 It is a fitting scatter plot of the experimental values ​​and the predicted values ​​of the single model of the present invention;

[0050] Fig.10 It is a fitting scatter plot of the experimental values ​​and predicted values ​​of the stacking-ML model of the present invention;

[0051] Fig.11 It is a distribution diagram of the experimental value, predicted value and error of the stacking-CARF model of the present invention on the test set. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] Embodiment 1:

[0054] See also Figure 1 to Figure 2 As shown in the figure, the analysis method of the influence of ultra-high performance concrete components on dynamic compressive strength is as follows:

[0055] Step S1, obtaining test data, after the 366 samples in the original data set are subjected to isolation forest anomaly detection, the remaining 245 groups of UHPC component-dynamic strength data sets are randomly divided into training sets and test sets. The division ratio of the data set is 8:2, and the random seed is set to 42; the input characteristic variables include the basic material composition parameters of UHPC, which are cement, silica fume, fly ash, slag, sand and gravel, steel fiber, strain rate, high-efficiency water reducer and water content, respectively. The output characteristic variable is the dynamic compressive strength of UHPC, which shows the simple statistical characteristics and data distribution of the input and output variables. Among them, except for the steel fiber content, which is measured by the fiber volume fraction, the mixing ratio of other components is the ratio of the mass of the corresponding material component to the mass of the bonding binder. In order to ensure the integrity and consistency of the data set, the following standards are formulated to define and organize the collected data: Considering that the type and distribution of steel fibers may affect the dynamic strength of UHPC, due to the limited experimental data on the influence of fiber distribution, the present invention only collects experimental data on random distribution of fibers. In order to collect as much experimental data as possible, there is no restriction on the type of steel fibers.

[0056] Step S2, anomaly detection based on isolation forest, uses the isolation forest algorithm to perform anomaly detection on the initial data set, and improves the data set by eliminating or reducing abnormal data, thereby avoiding poor prediction accuracy of the model due to data distribution defects and the existence of outliers.

[0057] Figure 2 The effect of the change in contamination rate on the RMSE value of the prediction performance evaluation index is shown. When the contamination rate changes from 0 to 50%, the RMSE value fluctuates significantly, which shows that the data processed by the isolation forest can significantly affect the prediction error. The data set after anomaly detection processing has a total of 245 samples, and its statistical characteristics are shown in Table 1 below:

[0058]

[0059]

[0060] Table 1

[0061] As can be seen from the above, when the contamination rate is 33%, the RMSE reaches the lowest 18.53, indicating that the prediction error obtained after removing 33% of the discrete anomaly data from the entire data set is as low as 18.53. It is worth noting that when no data is excluded, the prediction error of the model is 24.07. In other words, when anomaly detection is performed, the prediction error of dynamic intensity is reduced by nearly a quarter compared to when no processing is performed, which shows that the proposed anomaly detection method based on isolation forest is effective.

[0062] Embodiment 2:

[0063] See also Figures 3 to 11 As shown, step S3, construct a two-layer stacking-CARF model, use the first layer classification boosting (Categorical Boosting, CatBoost) and random forest (RF) base learner method to predict the dynamic compressive strength of UHPC with high precision, use the optimization and screening of the base learner, specifically: optimization and training of the base learner, the CatBoost and RF models as base learners are subjected to Bayesian automatic parameter optimization, and the models with the best hyperparameter values ​​are obtained for training, and the two base learners are trained using 5-fold cross validation to ensure that the model has a certain generalization ability. In the training process of the model, the training set is first divided into 5 equal subsets, and each fold uses 4 of the subsets to train the base learner, which means that 5 models will be trained for each base learner. Then, each base learner is used to predict the reserved subset, and the prediction results of the corresponding subset will be used as the meta-features of the training set. Finally, the model of each base learner on the five folds is used to predict the test set, and the prediction results of the five folds are averaged as the meta-features of the test set.

[0064] Step S4: high-precision prediction of the model, using the second-layer linear regression meta-learner to better learn the prediction results of the base learner, and achieve high-precision prediction of the dynamic strength of UHPC, specifically: training and prediction of the meta-learner. The meta-features of the training set and the test set obtained after the training of the above two base learners are combined as new training sets and test sets. The second-layer linear regression model as a meta-learner is trained on the new training set, and the model is used to generate the final prediction results;

[0065] Step S5, quantify the influence of UHPC composition on dynamic compressive strength, use explainable artificial intelligence (XAI) technology to assist stacked Kraft model prediction, including comparative analysis using Sharpless additive interpretation SHAP and partial dependence plot PDP is used to quantify the influence of single features and multiple interacting features on the global model, and the local interpretable model (LIME) is used to supplement the local interpretation of the model prediction process, providing support for the design and performance optimization of a single ultra-high performance concrete UHPC structure.

[0066] The single factor analysis results based on PDP show that the effect of low cement content on dynamic compressive strength is abnormal, because there is a certain interaction effect between cement and other auxiliary cementitious materials. Therefore, the interaction between different cementitious materials on dynamic compressive strength is analyzed, and the two-dimensional partial dependence grid diagram is used, such as Figure 3 , Figure 5 and Figure 7As shown and contour plots, Figure 4 , Figure 6 and Figure 8 As shown in Figure 2, the influence process is quantified and visualized. These two charts are very useful for identifying the complex interaction between two features and providing intuitive visual explanations for the prediction results of the model. Cement is the most important cementitious material, and silica fume, fly ash and blast furnace slag are selected to test the interaction of cement, as shown in Figure 2. Figures 3 to 8 shown.

[0067] When SF is greater than 0.21, FA is greater than 0.04, and S l ag is less than 0.1, the effects of the three supplementary cementitious materials on dynamic strength are significantly greater than the effects of low cement content (<0.69). In other words, no matter how the cement content changes between 0.4 and 0.69, when SF is greater than 0.21, FA is greater than 0.04, and S l ag is less than 0.1, the dynamic compressive strength will not change significantly. This result is Figure 4 , Figure 6 and Figure 8 This is particularly evident in the contour map. Figure 4 When the cement content is lower than 0.69 and SF is greater than 0.21, the contour line is almost perpendicular to the vertical axis, which means that the change of SF can significantly affect the dynamic compressive strength and the two are positively correlated. Similarly, in the low cement content range, FA greater than 0.04 and S l ag less than 0.1 are negatively correlated with the dynamic compressive strength. Therefore, low cement content does not significantly reduce the dynamic compressive strength of UHPC because of the cement content and the other three auxiliary cementitious materials, among which the auxiliary cementitious materials play a dominant role.

[0068] From the interaction of the three auxiliary cementitious materials on cement content, cement content is the main factor affecting the dynamic compressive strength of UHPC with high cement content (>0.69). Slight changes in cement content lead to large changes in dynamic compressive strength. At 0.74, the maximum single factor contribution of cement content to dynamic compressive strength is about 177.207MPa, but when combined with SF, FA and S lag, the contributions reach 181.918MPa, 178.071MPa and 186.116MPa, respectively. However, as the cement content decreases, auxiliary cementitious materials gradually become the dominant role. Especially when the cement content is lower than 0.69, the cement content loses its advantage in affecting the dynamic compressive strength and is replaced by the other three auxiliary cementitious materials.

[0069] Specifically, the CatBoost divides the original dynamic compressive strength dataset into n training subsets and evenly distributes the corresponding sample weights Then, the model e1 is trained using the sample subset with sample weight d1. The error rate of the training model is calculated as follows:

[0070]

[0071] Where I is the accuracy of model classification, where 1 indicates correct classification and 0 indicates incorrect classification; each training model e i The weight coefficients are updated according to the error rate of the training model as follows:

[0072]

[0073] Next, each sample weight is adjusted according to the updated weight coefficient as follows:

[0074] d i =d i exp(2W1I(y i ≠e1(x i )))(i=1,2,...,n)

[0075] Finally, repeat the above process l times and set n training models e1, e2, ..., e n The weighted average of is taken as the final output of the model, and the output result is as follows:

[0076]

[0077] Specifically, the RF introduces a random subspace method, that is, randomly sampling features when constructing a regression decision tree, thereby alleviating the problem of overfitting and low precision of a single decision tree model. After selecting the partitioning features, an optimal eigenvalue cutting point is selected to partition the nodes.

[0078] Specifically, the CatBoost and RF models used as base learners are subjected to Bayesian automatic parameter optimization to obtain a model with the best hyperparameter value for training. The meta-features of the obtained training set and test set are merged as new training set and test set. The second layer linear regression model used as the meta-learner is trained on the new training set, and the model is used to generate the final prediction result.

[0079] Specifically, the SHAP analysis interprets the predicted value of the model as the sum of the attributed values ​​of each feature, and the predicted value of the i-th sample is expressed as follows:

[0080]

[0081] Among them, y base is the predicted mean of all samples, p is the number of input features, is the SHAP value of the input feature, expressed as follows:

[0082]

[0083] Among them, S is a subset of features, x i is the feature vector of the sample, is the weight of subset S.

[0084] Specifically, PDP infers the relationship between features and predicted values ​​based on all samples in the data set. PDP can be used to reveal how a single variable or the interaction of multiple variables affects the prediction results of the model. The marginal effect of the corresponding feature on the model output is expressed as follows:

[0085]

[0086] where x S is the feature that needs to be plotted, x C Except x S Other features besides is x C The true value of , k is the sample size in the data set. At the same time, the Monte Carlo method points out that the average marginal effect of the target feature on the prediction can be estimated by calculating the average value in the training data. Overall, PDP can intuitively display the nonlinear relationship between a single feature and the output variable, helping to better understand the model behavior, the relationship between features, and the degree of influence of features on prediction.

[0087] Specifically, the LIME is used to explain the prediction of a single sample by a black box model. That is, first, a perturbation is performed near a single sample point to generate a new data set around a specific sample. Then, a simple model is trained on the new data set, usually a regression or decision tree model. Finally, the simple model is used to replace the original complex model for local explanation, which is expressed as:

[0088]

[0089] in, is the original model to be explained, G is a set of simple models for explanation, g is the explanation model of instance x, Ω(g) is the complexity of the model, and the minimization loss L measures the difference between the explanation model g and the original model The closeness of the prediction, π x It represents the distance between the instance x of interest and the sample generated by the perturbation. The closer to the instance, the higher the weight. LIME is used to understand the behavior of the model in a specific sample or local area, rather than a global explanation. The model independence of LIME is reflected in that it does not depend on the specific model structure. Specifically, a simple and easy-to-interpret local model is created based on the input and output of the model to explain the instance of interest, thereby better guiding decision-making.

[0090] The optimized RF, XGBoost and CatBoost were selected as candidate base models, and the simple LR model was selected as the meta-model. The three candidate base models were fused and compared in pairs using the stacking principle, and three stacking machine learning (stacking-ML) models with different base model combinations were established.

[0091] Table 2 compares the evaluation metrics of the stacking-ML model with those of a single ensemble learning model on the test set.

[0092]

[0093]

[0094] Table 2

[0095] The performance of the stacking-ML model is better than that of a single model, which indicates that the meta-model in the stacking model corrects the samples that are incorrectly predicted by the base learner to a certain extent, improving the prediction accuracy of the model. In addition, the stacking-CARF model obtained by combining CatBoost and RF as the base model has the best prediction performance. Its evaluation indicators on the test set are: MAE = 9.770, RMSE = 15.865, R 2 =0.926.

[0096] Fig. 9 and Fig.10 They are the distribution of predicted values ​​and true values ​​of a single ensemble learning model and a stacked model with different base learners on the test set, which are presented in the form of a fitted scatter plot. The diagonal line indicates that the predicted value and the true value are completely consistent, and the dotted line indicates that the error of the test set has an error range of ±20%, as shown in Fig. 9 and Fig.10 As shown in the figure, the distribution of scattered points on a single model is more discrete than that of the stacking-ML model, and the errors of some points are above or outside 20%. The performance of the stacking-CARF model is better than that of other stacking-ML models, and the prediction errors on the test set are all within 20%.

[0097] The comparison between the predicted values ​​and the true values ​​of the stacking-CARF model on the test set and the error distribution are shown in Figure 2. Fig.11 As shown in the figure, the predicted values ​​of the stacking-CARF model are consistent with the experimental values, and the errors on the test set are all below 20 MPa.

[0098] As can be seen from the above, the quality of the data set is crucial to the prediction accuracy of the ML model. The isolation forest anomaly detection algorithm improves the data set by eliminating or mitigating abnormal data, thereby significantly improving the prediction accuracy of the ML model. The prediction results of the test set show that by removing 33% of the discrete abnormal data in the UHPC dynamic compression data set, the RMSE value can be reduced from 24.07 to 18.53.

[0099] Using CatBoost and RF as the base learners of the first layer and LR as the meta-learner of the second layer, the stacked-CARF model established has higher prediction accuracy and robustness than the single model and EL model. Its prediction performance indicators can reach MAE=9.770, RMSE=15.865, R 2 =0.926, which can achieve high-precision prediction of the dynamic strength of UHPC.

[0100] Comparative analysis of SHAP and PDP can reveal relationships that are difficult to quantify by traditional experimental and theoretical analysis, such as feature importance ranking, marginal effects of features, and interactions. For example, in the low cement content range (<0.69), the effect of auxiliary cementitious materials on dynamic compressive strength plays a dominant role, which can compensate for the loss of dynamic strength caused by low cement content to a certain extent. In addition, compared with the above global explanation, LIME is more interpretable for specific samples, which can provide a theoretical reference for the design and performance optimization of a single UHPC structure.

[0101] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0102] In the drawings of the embodiments disclosed in the present invention, only the structures related to the embodiments disclosed in the present invention are involved, and other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other.

[0103] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength is characterized by: The following steps are involved: S1. Acquire test data and collect ultra-high performance concrete (UHPC) dynamic compression data from the public literature on split-Hopkinson bar (SHPB) experiments as the initial data set; S2, anomaly detection based on isolation forests, improves the dataset by eliminating or mitigating abnormal data, thereby avoiding poor prediction accuracy of the model due to data distribution defects and the presence of outliers; S3. Construct a two-layer stacking-CARF model, use the first layer classification boosting CategoricalBoosting, CatBoost and Random Forest Random Forest, RF base learner method to make high-precision prediction of UHPC dynamic compressive strength, and use base learners for optimization and screening; S4, high-precision prediction of the model, using the second-layer linear regression meta-learner to learn the prediction results of the base learner to achieve high-precision prediction of UHPC dynamic strength; S5. Quantify the influence of UHPC components on dynamic compressive strength, use explainable artificial intelligence (XAI) technology to assist stacked Kraft model prediction, including comparative analysis using Sharpless additive interpretation (SHAP) and partial dependence plot (PDP) to quantify the influence of single features and multiple interacting features on the global model, and use the local explainable model (LIME) to supplement the local explanation of the model prediction process, providing support for the design and performance optimization of a single ultra-high performance concrete (UHPC) structure.

2. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 1, characterized in that: The input characteristic variables of the data set include basic material composition parameters of UHPC, namely the content of cement, silica fume, fly ash, slag, sand and gravel, steel fiber, high-efficiency water reducer and water, and the output characteristic variable is the dynamic compressive strength of UHPC.

3. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 1, characterized in that: The CatBoost combines multiple weak learners into a strong learner. The algorithm can adjust the sample weight of the next model learning data according to the learning results of the previous model, and build the model in a step-by-step iterative manner. The weak learners built in each step of the iteration are to make up for the shortcomings of the existing model.

4. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 3, characterized in that: The CatBoost algorithm divides the original dynamic compressive strength dataset into n training subsets and evenly distributes the corresponding sample weights. i=1,2,...,n, then use the sample subset with sample weight d1 to train the model e1. The error rate of the training model is calculated as follows: Where I is the accuracy of model classification, where 1 indicates correct classification and 0 indicates incorrect classification; each training model e i The weight coefficients are updated according to the error rate of the training model as follows: Next, each sample weight is adjusted according to the updated weight coefficient as follows: d i =d i exp(2W1I(y i ≠e1(x i )))(i=1,2,...,n) Finally, repeat the above process l times and set n training models e1, e2, ..., e n The weighted average of is taken as the final output of the model, and the output result is as follows:

5. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 2, characterized in that: The RF introduces a random subspace method, that is, randomly sampling features when constructing a regression decision tree, thereby alleviating the problem of overfitting and low precision of a single decision tree model. After selecting the partitioning features, an optimal eigenvalue cutting point is selected to partition the nodes.

6. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 1, characterized in that: The CatBoost and RF models used as base learners are subjected to Bayesian automatic parameter optimization to obtain a model with the best hyperparameter value for training. The meta-features of the obtained training set and test set are merged as new training sets and test sets. The second layer of the linear regression model used as the meta-learner is trained on the new training set, and the model is used to generate the final prediction results.

7. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 1, characterized in that: The SHAP analysis interprets the model's predicted value as the sum of the attributed values ​​of each feature. The predicted value of the i-th sample is expressed as follows: Among them, y base is the predicted mean of all samples, p is the number of input features, is the SHAP value of the input feature, expressed as follows: Among them, S is a subset of features, x i is the feature vector of the sample, is the weight of subset S.

8. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 1, characterized in that: PDP infers the relationship between features and predicted values ​​based on all samples in the data set. PDP is used to reveal how a single variable or the interaction of multiple variables affects the prediction results of the model. The partial dependence function of the corresponding feature on the model output is expressed as follows: where x S is the feature that needs to be plotted, x C Except x S Other features besides P(x C ) is x C The probability measure of The partial dependence function It is estimated by calculating the average value in the training data, also known as the Monte Carlo method, and the corresponding partial dependence function is expressed as follows: in is x C The true value of , k is the number of samples in the data set. In general, PDP can intuitively display the nonlinear relationship between a single feature and the output variable, helping to understand the model behavior, the relationship between features, and the impact of features on prediction.

9. The method for analyzing the influence of ultra-high performance concrete components on dynamic compressive strength according to claim 1, characterized in that: The LIME is used to explain the prediction of a single sample by a black box model. That is, it is necessary to first perturb a single sample point to generate a new data set around the specific sample, then train a model on the new data set, including a regression or decision tree model, and finally use the model to replace the original model for local explanation, which is expressed as: in, is the original model to be explained, G is a set of simple models for explanation, g is the explanation model of instance x, Ω(g) is the complexity of the model, and the minimization loss L measures the difference between the explanation model g and the original model The closeness of the prediction, π x Represents the distance between instance x and the sample generated by the perturbation. The closer to the instance, the higher the weight.

Citation Information

Patent Citations

  • Method for predicting concrete compressive strength based on random forest and intelligent algorithm

    CN112069567A

  • Concrete compressive strength prediction model, construction method, medium and electronic equipment

    CN116434893A