Ultra-high performance concrete multi-performance prediction method based on machine learning

By constructing a UHPC performance dataset with multiple input variables, employing rigorous data preprocessing and feature selection, training the optimal prediction model using a variety of machine learning algorithms, and utilizing the SHAP method to interpret feature contributions, the problems of incomplete datasets and poor model interpretability in UHPC performance prediction are solved, achieving accurate and simultaneous prediction of multiple UHPC performance metrics.

CN121306353APending Publication Date: 2026-01-09XINJIANG BINGTUAN CONSTR ENG CO LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511400869.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing UHPC performance prediction methods suffer from incomplete datasets, insufficient data processing and feature engineering considerations, and poor model interpretability, making it difficult to accurately and simultaneously predict multiple properties of ultra-high performance concrete.

Method used

By constructing a UHPC performance dataset containing multiple input and output variables, employing rigorous data preprocessing and feature selection methods, combining various machine learning algorithms to train the optimal prediction model, and utilizing the SHAP method to interpret the contribution of features to the prediction results, the model is integrated into a graphical user interface software.

Benefits of technology

It achieves accurate and simultaneous prediction of various UHPC performance characteristics, improves the model's generalization ability and interpretability, and provides strong technical support for the research and development and application of UHPC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306353A_ABST
    Figure CN121306353A_ABST
Patent Text Reader

Abstract

The invention provides an ultra-high performance concrete multi-performance prediction method based on machine learning. The ultra-high performance concrete multi-performance prediction method comprises the following steps: Step 1, establishing a data set; step 2, data preprocessing is carried out; step 3, establishing an optimal prediction model: based on the feature subset, adopting a plurality of different machine learning algorithms for training, and selecting the machine learning algorithm with the best training effect as the optimal prediction model; step 4, selecting an optimal feature subset; step 5, explaining the influence of the features on model prediction: calculating the contribution degree of each feature to a prediction result based on the optimal prediction model and the optimal feature subset, and helping to understand the decision process of the model; and Step 6, performance prediction of the ultra-high performance concrete: inputting parameters of the to-be-predicted ultra-high performance concrete into the optimal prediction model to obtain a predicted value of the performance. The technical problems that an existing UHPC performance prediction method is incomplete in data set, insufficient in consideration of data processing and feature engineering and poor in model interpretation can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of materials science and engineering technology, and specifically to a method for predicting the multi-performance of ultra-high performance concrete based on machine learning. Background Technology

[0002] Ultra-high performance concrete (UHPC), as an advanced cement-based composite material, achieves superior properties such as compressive strength, flexural strength, toughness, and durability far exceeding those of traditional concrete through optimized gradation, reduced water-cement ratio, and the incorporation of active powders and fibers. It shows broad application prospects in bridges, high-rise buildings, and protective engineering. The performance characteristics of UHPC, including its mechanical properties (such as compressive and flexural strength), workability (such as flowability), and microstructural properties (such as porosity), are closely related to and exhibit a highly nonlinear dependence on numerous factors, including its complex material composition (such as cement, silica fume, slag, fly ash, quartz powder, aggregate type and particle size, fiber type and dosage, mix design, and curing regime (such as temperature and age). Therefore, accurately and efficiently predicting the various properties of UHPC is of great significance for guiding material design, optimizing mix proportions, controlling production quality, and promoting its engineering applications.

[0003] However, the complex composition of UHPC materials makes performance testing and evaluation a challenge. Traditional testing methods require extensive experiments and time, are costly and inefficient, and fail to meet the needs of engineering practice. Therefore, predicting the performance of UHPC has become a research hotspot in the field of materials engineering. Currently, UHPC performance prediction methods are mainly divided into two categories: 1) traditional empirical formulas and regression methods; 2) performance prediction models based on artificial intelligence technology. Traditional prediction methods rely on a large amount of experimental data and statistical analysis, making it difficult to accurately capture the complexity and nonlinear characteristics of materials. In recent years, with the rapid development of artificial intelligence and machine learning (ML) technologies, data-driven prediction methods have gradually become mainstream. Machine learning methods can more accurately capture complex nonlinear relationships and improve prediction accuracy. For example, artificial neural networks (ANN), support vector machines (SVR), random forests (RF), and gradient boosting (GB) have been applied to UHPC performance prediction. However, existing machine learning methods often ignore key influencing factors such as temperature, aggregate particle size, specimen size, etc. when constructing data sets for UHPC performance prediction, or rely on data sets with insufficient data volume. In addition, data preprocessing and feature engineering methods have a significant impact on the prediction accuracy and generalization ability of machine learning models, but existing research has not adequately considered these aspects. At the same time, machine learning models are often considered "black box" models, and the relationship between input and output is difficult to explain, so model interpretability remains an important research challenge. In summary, there is an urgent need in the current field to develop a new method or framework for UHPC performance prediction that can overcome the limitations of existing technology. This method should be based on a more comprehensive and high-quality data foundation, use effective feature engineering and advanced machine learning techniques, achieve accurate and simultaneous prediction of multiple key performances of UHPC, and have good model interpretability, thereby providing stronger technical support for the research and application of UHPC. SUMMARY

[0004] The technical problem to be solved by the present application is to provide a machine learning-based ultra-high performance concrete multi-performance prediction method, which solves the technical problems of existing UHPC performance prediction methods such as incomplete data sets, insufficient data processing and feature engineering, and poor model interpretability.

[0005] To solve the above technical problems, the technical solution adopted by the present application is: A machine learning-based ultra-high performance concrete multi-performance prediction method, comprising the following steps: Step 1, establishing a data set: constructing a UHPC performance data set containing multiple input variables and output variables by experiments and collecting data from domestic and foreign literature, the variables cover data information of material type, content, curing condition, aggregate particle size and specimen size; Step 2, data preprocessing: improving data quality through data normalization and outlier detection, and dividing different feature subsets through feature selection and processing; Step 3, establishing an optimal prediction model: based on the feature subsets, training multiple different machine learning algorithms, and selecting the machine learning algorithm with the best training effect as the optimal prediction model; Step 4, selecting the optimal feature subset: based on the optimal prediction model, evaluating the prediction effect of the model on different feature subsets, and selecting the optimal feature subset; Step 5, explaining the influence of features on model prediction: based on the optimal prediction model and the optimal feature subset, calculating the contribution of each feature to the prediction result, and helping to understand the decision-making process of the model; Step 6: Performance prediction of ultra-high performance concrete: Input the parameters of the ultra-high performance concrete to be predicted into the optimal prediction model to obtain the predicted value of the desired performance.

[0006] In Step 1 above, the input variables include: 1) Material composition related variables: cement type, cement dosage, silica fume dosage, slag dosage, fly ash dosage, quartz powder dosage, fine aggregate dosage, maximum fine aggregate particle size, coarse aggregate dosage, maximum coarse aggregate particle size, fiber type, fiber content, fiber length, fiber diameter, and water-reducing agent dosage. 2) Variables related to maintenance and environmental conditions: maintenance temperature and maintenance age; 3) Specimen information related variables: specimen size.

[0007] In Step 1 above, the output variables are the UHPC performance indicators that need to be predicted in Step 6, including one or more of the following: compressive strength, flexural strength, flowability, and porosity.

[0008] Step 2 above includes the following steps: Step 2.1: Use the Z-score normalization method to transform the data into a standard normal distribution with a mean of 0 and a standard deviation of 1, in order to eliminate dimensional differences between different features; Step 2.2: Use the Isolation Forest IF algorithm based on binary search tree (BST) to detect outliers and remove abnormal data. Step 2.3: The features are analyzed and screened using methods such as Pearson correlation analysis, Mantel test, one-way ANOVA, and mutual information (MI). Based on the MI score, a threshold is set to divide the features into multiple feature subsets with different dimensions.

[0009] In Step 3 above, the various machine learning algorithms specifically include at least two of the following: Ridge Regression (RR), Support Vector Regression (SVR), Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Extreme Gradient Boosting (XGBoost), and Lightweight Gradient Boosting Machine (LightGBM) ensemble learning or regression algorithms.

[0010] In Step 3 above, the Ridge Regression (RR) loss function can be expressed as follows: ; in, n It is the sample size. m It is the number of features. y i It is the first i The target value for each sample x ij It is the first i The first samplej 1 eigenvalue, It is the first j The weights of each feature λ This is the regularization parameter, which controls the strength of the regularization term; the formula for the regularization term is the latter half; without regularization... λ When the loss function is 0, ridge regression degenerates into ordinary least squares regression; the training process of the ridge regression model is to solve for the loss function. L ( θ Minimize the optimal parameters θ The process; For ridge regression, there exists a closed-form solution (analytical solution), and the optimal parameters can be calculated directly: ; in, X It is n × m A matrix, where each row represents a sample and each column represents a feature. y It is the target value vector. λ It is a regularization parameter. I It is an identity matrix.

[0011] In Step 3 above, Support Vector Regression (SVR) is an application of Support Vector Machine (SVM) to regression problems; for a given training dataset D = {( , ), ( , ), ..., ( , )},in It is the input feature vector. This corresponds to the true value; the goal of SVR is to find a linear function: ;in, w It is a weight vector. b This is the bias term; SVR allows a range of errors to be excluded from the loss calculation; this range is called the ε-insensitive band; the loss is calculated only when the absolute value of the difference between the predicted and the true values ​​is greater than ε; to handle data points outside the ε-insensitive band, slack variables are introduced. and ; This indicates the deviation of the sample point above the ε-insensitive band; This represents the bias of the sample points below the ε-insensitive band; the optimization objective of SVR can be expressed as the following original formula: Minimize: 0.5*|| w ||² + C ∑( + ); The constraints are: ; ; , ; Among them, || w ||² is a regularization term used to control the complexity of the model, making it smoother and preventing overfitting; C It is the penalty coefficient, a hyperparameter used to balance the regularization term and the training error term; C The larger the value, the greater the penalty for samples outside the ε-insensitive band, which may make the model more complex or even overfit. C The smaller the value, the stronger the model's fault tolerance, but this may lead to underfitting; the SVR training process involves inputting the configured parameters and prepared data into the algorithm; and obtaining the optimal weight vector by solving a convex quadratic programming problem. w and bias b .

[0012] In Step 3 above, Random Forest (RF) improves the accuracy and robustness of the model by constructing multiple decision trees and integrating their results. The training process of random forests introduces two types of randomness: random sampling and random feature selection; random sampling refers to selecting features from a collection of... N From the original training dataset, random sampling with replacement is performed to select samples. N This process is repeated to create a bootstrap sample set of the same size as the original dataset. T This time, thus obtaining T T different bootstrap sample sets are used to train T decision trees; random selection of features refers to randomly selecting features from these bootstrap samples when splitting at each node in the construction of each decision tree. M Select from features m These features increase the diversity between trees. In Step 3 above, Gradient Boosting Decision Tree (GBDT) iteratively trains a series of weak learners. Each new tree learns the residuals of the accumulated results of all previous trees, and finally, the predictions of all trees are summed to obtain the final prediction value. The final GBDT model is an additive model, and its mathematical form is as follows: Suppose we have M decision trees, for an input sample x The final predicted value F M ( x The sum of all tree predictions is: ; where, F M ( x ) is the final prediction model; y 0 is the initial prediction value of the model; is the prediction result of the i-th decision tree; this tree itself is used to fit the residual of the previous round model; m is the learning rate, which is a hyperparameter between 0 and 1; its role is to discount the prediction result of each new tree, prevent a single tree from contributing too much to the model, and thus reduce the risk of overfitting, making the model more robust. v

[0013] In Step 3 above, the extreme gradient boosting XGBoost includes the following: Regularization: built-in L 1 and L 2 regularization terms effectively control the complexity of the model and prevent overfitting; High-order gradient information: using the second-order Taylor expansion, using the first and second-order gradients makes the optimization process more accurate and converges faster; Efficient split point search algorithm: when searching for the best split point, a greedy strategy is adopted; when the data volume is huge and cannot be read into memory, XGBoost proposes an approximate algorithm based on quantiles, which greatly improves the efficiency; Sparsity-aware: can automatically handle sparse data and missing values; when searching for split points, it will calculate the gain of dividing missing values to the left and right child nodes respectively, and select the direction with greater gain as the default division direction of missing values; Parallel computing and cache optimization: parallel computing at the feature level; through clever data structure design, the gradient and other intermediate calculation results can better utilize CPU cache, reducing memory reading time.

[0014] ​In Step 3, the LightGBM algorithm improves the training efficiency and memory performance when dealing with large-scale data through histogram algorithm, GOSS (Gradient-based One-Side Sampling), and EFB (Efficient Feature Bundling). The histogram method is used to find the optimal split point of the tree, which discretizes continuous floating-point feature values into a fixed number of "bins" and constructs a histogram of feature values. During the process of finding the split point, EFB identifies mutually exclusive or approximately mutually exclusive features through algorithms such as graph coloring and bundles them into a single "super feature", thereby reducing the number of features. LightGBM adopts a tree growth strategy that grows leaves first. Each time, it finds the leaf with the largest split gain from all current leaves and splits it. This strategy may increase the risk of overfitting on small datasets, but it can converge faster to a high-precision model on large datasets.

[0015] In Step 3, decision trees in ensemble learning algorithms construct a tree structure model from data features by learning a series of simple "yes / no" rules, thereby predicting new samples. Starting from the root node, for each split child node, the optimal feature for splitting is repeatedly selected until the stopping conditions are met, such as the number of samples contained in the current node being below the pre-set threshold min_child_samples, the depth of the tree reaching the pre-set maximum value max_depth, the number of leaf nodes in a tree reaching the maximum number num_leaves, the learning rate learning_rate (also known as "step" or "shrinkage") in the ensemble model, and the coefficient for reducing the contribution of each tree. In Step 3, five-fold cross-validation technology is used to train and evaluate candidate machine learning algorithms.

[0016] The five-fold cross-validation technology specifically includes randomly dividing the training dataset into five equal parts, taking four parts as a sub-training set to train the model and the remaining one part as a validation set to evaluate the model performance, repeating five times, and calculating the average value of the performance indicators.

[0017] In Step 3, Bayesian Optimization is used to automatically optimize the hyperparameters of each candidate model. The Bayesian Optimization method uses existing experimental results, i.e., "prior knowledge", to guide the subsequent search direction. Its framework consists of two parts: Surrogate model: This is a probabilistic model constructed using Gaussian Process (GP). The role of the surrogate model is to fit an approximation of the "black box function" based on the existing combination of hyperparameters and their corresponding model performance data. Acquisition function: the acquisition function uses the information provided by the surrogate model to decide the next hyperparameter point to try; the goal is to strike a balance between "exploration" and "exploitation"; exploitation refers to searching near the area with the best known performance.

[0018] In Step 3, the performance evaluation indicators for evaluating and comparing models include one or more of the following: determination coefficient R 2 , root mean square error RMSE, mean absolute error MAE, and mean absolute percentage error MAPE; the optimal comprehensive indicator in cross-validation is selected as the optimal prediction model.

[0019] In Step 4, the correlation coefficient, standard deviation, and RMSE of the optimal prediction model on different feature subsets are evaluated and compared using visualization tools such as Taylor Diagram, and the optimal feature subset is intuitively selected.

[0020] In Step 5, the SHAP (SHapley Additive exPlanations) method is used to calculate the contribution of each input feature in the optimal feature subset to the prediction of each target performance. SHAP analysis includes: (1) Global explanation: analyze the average SHAP absolute value of each feature to obtain the feature importance ranking; analyze the relationship between feature value and SHAP value in the form of a bee swarm chart to understand the overall positive and negative impact of the feature on the prediction result. (2) Dependency analysis: draw a SHAP dependency chart for a single feature to show how changes in the feature value specifically affect the model output, and observe the interaction effects with other features.

[0021] In Step 5, the trained optimal model and preprocessing process are integrated into a graphical user interface (GUI) software.

[0022] The machine learning-based ultra-high performance concrete multi-performance prediction method mentioned in the application systematically combines experimental data and widely collected literature data through Step 1, and particularly pays attention to the key variables that are easily ignored in previous studies, such as aggregate particle size, curing temperature, specimen size, etc., to construct a large UHPC multi-performance data set with more comprehensive dimensions and stronger representation. Further, through the strict data preprocessing process in Step 2, especially the use of Isolation Forest IF algorithm for outlier detection and elimination, the quality of the data and the effectiveness of the training samples are significantly improved. Compared with the prior art which relies on limited, single source or insufficiently cleaned data, the present application is based on more solid and high-quality data, so that the machine learning model trained through Step 3 can capture more universal rules, significantly improving the accuracy, reliability and generalization ability of the prediction of multiple performances of UHPC with different sources and different proportions, and overcoming the limitations of traditional empirical formulas or modeling based on small samples / low-quality data. The framework of the present application can take compressive strength, flexural strength, fluidity, porosity and other key performances as prediction targets. Through Step 3, multiple advanced machine learning algorithms are systematically compared and selected, and means such as five-fold cross-validation and Bayesian hyperparameter optimization are used to ensure that the optimal prediction model selected has excellent performance. Unlike the prior art which mostly only focuses on a single performance such as only compressive strength prediction, the present application can accurately predict multiple performance indicators of UHPC based on the optimized high-performance model, providing a powerful tool for comprehensive performance evaluation and multi-objective optimization design of UHPC, and significantly outperforming the prediction ability and range of traditional methods.

[0023] The present application introduces a systematic feature analysis method in Step 2, and determines the optimal feature subset based on the actual performance of the optimal prediction model in Step 4. Through data-driven feature selection, the input dimension is effectively reduced, redundant or noise information is eliminated, the efficiency of model training is improved, and the prediction accuracy may be further improved. Strict cross-validation and hyperparameter optimization ensure the reliability of model performance evaluation, reduce the risk of overfitting, so that the finally selected model not only has high accuracy, but also has stable performance and strong generalization ability.

[0024] The application introduces an advanced SHAP (SHapley Additive exPlanations) model explanation method in Step 5, which is applied to the optimal model and the optimal feature set obtained through the foregoing steps. The "black box" problem commonly existing in high-performance machine learning models is overcome. The SHAP analysis can quantitatively and clearly reveal the specific contribution and influence mode (positive or negative) of each input feature material component, curing condition, etc. on each predicted performance (compressive strength, flexural strength, fluidity, and porosity). This not only enhances the user's trust in the prediction results, but more importantly, provides valuable, data-driven insights and basis for understanding the "component-process-structure-performance" relationship of UHPC materials, exploring performance influencing mechanisms, and performing mechanism-based material optimization design, which is difficult to achieve by traditional methods and ML models that only pursue prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0025] The application will be further described below in combination with the drawings and examples: Figure 1 The flowchart of the machine learning-based ultra-high performance concrete multi-performance prediction method in the application. DETAILED DESCRIPTION

[0026] The technical solutions of the application will be described in detail below in combination with the drawings and examples.

[0027] A machine learning-based ultra-high performance concrete multi-performance prediction method, comprising the following steps: Step 1, establishing a data set: through experiments and collecting data in domestic and foreign literature, a UHPC performance data set containing multiple input variables and output variables is constructed, and the variables cover data information of material type, content, curing condition, aggregate particle size, and specimen size; Step 2, data preprocessing: data quality is improved through data normalization and outlier detection, and different feature subsets are divided through feature selection and processing; Step 3, establishing an optimal prediction model: based on the feature subsets, multiple different machine learning algorithms are trained, and the machine learning algorithm with the best training effect is selected as the optimal prediction model; Step 4, selecting the optimal feature subset: based on the optimal prediction model, the prediction effect of the model on different feature subsets is evaluated, and the optimal feature subset is selected; Step 5, explaining the influence of features on model prediction: based on the optimal prediction model and the optimal feature subset, the contribution of each feature to the prediction result is calculated, which helps to understand the decision-making process of the model; Step 6, ultra-high performance concrete performance prediction: the parameters of the ultra-high performance concrete to be predicted are input into the optimal prediction model, and the predicted value of the performance is obtained.

[0028] In Step 1 above, the input variables include: 1) Material composition related variables: cement type (such as strength grade), cement content, silica fume content, slag content, fly ash content, quartz powder content, fine aggregate content, maximum fine aggregate particle size, coarse aggregate content (if any), maximum coarse aggregate particle size (if any), fiber type, fiber content, fiber length, fiber diameter, and water-reducing agent content, etc. 2) Variables related to maintenance and environmental conditions: maintenance temperature and maintenance age, etc.; 3) Specimen information related variables: specimen size, etc.

[0029] In Step 1 above, the output variables are the key performance indicators of UHPC that need to be predicted in Step 6, including one or more of the following: compressive strength, flexural strength, flowability, and porosity.

[0030] Step 2 above includes the following steps: Step 2.1: Use the Z-score normalization method to transform the data into a standard normal distribution with a mean of 0 and a standard deviation of 1, in order to eliminate dimensional differences between different features; Step 2.2: Use the Isolation Forest IF algorithm based on binary search tree (BST) to detect outliers, remove abnormal data, and make the data distribution more reasonable. Step 2.3: Analyze and screen features using methods such as Pearson correlation analysis, Mantel test, one-way ANOVA, and mutual information (MI); set thresholds based on MI scores or other importance indicators to divide features into multiple feature subsets with different dimensions.

[0031] In Step 3 above, the various machine learning algorithms specifically include at least two of the following ensemble learning or regression algorithms: Ridge Regression (RR), Support Vector Regression (SVR), Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Extreme Gradient Boosting (XGBoost), and Lightweight Gradient Boosting Machine (LightGBM).

[0032] In Step 3 above, Ridge Regression (RR) is an improved linear regression algorithm. By introducing an L2 regularization term into the loss function, it effectively prevents overfitting and enhances the model's stability when dealing with multicollinear data. Its loss function can be expressed as follows: ; in, n It is the sample size. m It is the number of features. y i It is the first i The target value for each samplex ij is the i th feature value of the j th sample, is the weight of the j th feature, λ is a regularization parameter that controls the strength of the regularization term; the formula of the regularization term is the latter part; in the case of no regularization, i.e. λ = 0, the ridge regression degenerates into the ordinary least squares regression; the training process of the ridge regression model is the process of solving the optimal parameters L that minimize the loss function θ ; θ ; For the ridge regression, there is a closed-form solution (analytical solution) that can directly calculate the optimal parameters: ; where X is a n × m matrix, where each row represents a sample and each column represents a feature. y is the target value vector, λ is the regularization parameter, I is the identity matrix.

[0033] In Step 3 above, the support vector regression (SVR) is the application of the support vector machine (SVM) to the regression problem; for a given training data set D = {( , ), ( , ),..., ( , )} where is the input feature vector, is the corresponding true value; the goal of SVR is to find a linear function: ; where w is the weight vector, b is the bias term; SVR allows errors within a certain range to not be counted in the loss; this range is called the ε-insensitive band; only when the absolute value of the difference between the predicted value and the true value is greater than ε, the loss is calculated; in order to handle those data points that are outside the ε-insensitive band, slack variables and are introduced; represents the deviation of the sample point above the ε-insensitive band; represents the deviation of the sample point below the ε-insensitive band; the optimization objective of SVR can be represented as the following original formula: minimize: 0.5*|| w ||² + C ∑( + ); The constraints are: ; ; , ; Among them, || w ||² is a regularization term used to control the complexity of the model, making it smoother and preventing overfitting; C It is the penalty coefficient, a hyperparameter used to balance the regularization term and the training error term; C The larger the value, the greater the penalty for samples outside the ε-insensitive band, which may make the model more complex or even overfit. C The smaller the value, the stronger the model's fault tolerance, but this may lead to underfitting; the SVR training process involves inputting the configured parameters and prepared data into the algorithm; and obtaining the optimal weight vector by solving a convex quadratic programming problem. w and bias b .

[0034] In Step 3 above, Random Forest (RF) is a powerful ensemble learning algorithm widely used in classification and regression tasks; it improves the accuracy and robustness of the model by constructing multiple decision trees and integrating their results. The training process of random forests introduces two types of randomness: random sampling and random feature selection; random sampling refers to selecting features from a collection of... N From the original training dataset, random sampling with replacement is performed to select samples. N This process is repeated to create a bootstrap sample set of the same size as the original dataset. T This time, thus obtaining T T decision trees are trained using T different bootstrap sample sets; random feature selection means that when splitting at each node in the construction of each decision tree, features are not selected from all available samples. M Instead of selecting the optimal feature from these features, it randomly selects one from these features. M Select from features m These features increase the diversity between trees. In this way, a large number of decision trees are trained, forming a forest. For regression problems, the model's final prediction is the average of the predictions from all the decision trees in the forest.

[0035] In Step 3 above, Gradient Boosting Decision Tree (GBDT) trains a series of weak learners, typically decision trees, by iteratively learning the residual error of the cumulative results of all previous trees. The final prediction value is obtained by adding the prediction results of all trees. The final model of GBDT is an additive model, which can be mathematically represented as follows: Assuming we have M decision trees, for an input sample x , the final prediction value F M ( x ) is the sum of the prediction results of all trees: ; where F M ( x ) is the final prediction model; y 0 is the initial prediction value of the model, usually the mean value of the training sample labels; is the prediction result of the m th decision tree; this tree itself is used to fit the residual error of the previous model; v is the learning rate, a hyperparameter between 0 and 1; its role is to discount the prediction result of each new tree to prevent a single tree from contributing too much to the model, thereby reducing the risk of overfitting and making the model more robust.

[0036] In Step 3 above, XGBoost is an efficient, flexible, and portable implementation of Gradient Boosting Decision Tree (GBDT). It has made several important improvements to GBDT, including the following: Regularization: It has built-in L 1 and L 2 regularization terms to effectively control model complexity and prevent overfitting; High-order gradient information: It uses a second-order Taylor expansion, utilizing both first and second-order gradients to make the optimization process more accurate and faster to converge; Efficient split point finding algorithm: When finding the best split point, it uses a greedy strategy; when the data volume is huge and cannot be read into memory, XGBoost proposes a quantile-based approximation algorithm, which greatly improves efficiency; Sparsity-aware: It can automatically handle sparse data and missing values; when finding the split point, it calculates the gain of dividing missing values to the left and right child nodes respectively, and selects the direction with greater gain as the default division direction for missing values; Parallel computing and cache optimization System Optimization: can be calculated in parallel at the feature level; through the ingenious design of data structure, the intermediate calculation results such as gradient can be better utilized CPU cache, and the memory reading time is reduced.

[0037] In summary, XGBoost has deeply optimized the theory and engineering implementation of GBDT.

[0038] In Step 3 above, LightGBM is a gradient boosting framework born for speed and efficiency. It greatly improves the training efficiency and memory performance of the model when processing large-scale data through three core technologies: histogram algorithm, one-sided gradient sampling GOSS, and mutually exclusive feature binding EFB. It uses the histogram method to find the optimal split point of the tree, that is, by discretizing the continuous floating-point feature values into a fixed number of "buckets" bins, and constructing a histogram of feature values. In the process of finding the split point, instead of traversing all data points, only the boundaries of these buckets are traversed as candidate split points. Since the data is binned, it is very efficient to calculate the gradient and the sum of the first and second order gradients of each bin. The calculation of split gain can be directly completed on the histogram. The core idea of one-sided gradient sampling GOSS is that the samples with large gradient, i.e. residual, are the ones that the model has made serious mistakes on, and their contribution to model optimization is also greater. Therefore, by retaining all samples with large gradient and randomly sampling a part of samples with small gradient and weighting, the number of samples can be effectively reduced to speed up training. EFB identifies mutually exclusive or nearly mutually exclusive features through graph coloring algorithms, and binds them into a single "super feature", thereby reducing the number of features. LightGBM adopts a tree growth strategy of growing leaves, and each time it finds the leaf with the largest split gain from all current leaves to split. This may increase the risk of overfitting on small datasets, but it can converge faster to a high-precision model on large datasets.

[0039] In Step 3, the decision tree in the ensemble learning algorithm is a basic supervised learning algorithm. It builds a tree structure model from data features by learning a series of simple "yes / no" rules to make predictions on new samples. The key to decision tree learning is how to select the optimal feature to split the node. The goal is to make the split child nodes as "pure" as possible, i.e., belonging to the same class or having small value differences. For regression tasks, the core idea of selecting a split node is to minimize the sum of the variances of the real values of the samples in the two child nodes after splitting. Starting from the root node, for each split child node, repeat the selection of the optimal feature for splitting until the stopping conditions are met, such as the number of samples in the current node is less than the pre-set threshold min_child_samples; the depth of the tree reaches the pre-set maximum value max_depth; the number of leaf nodes in a tree reaches the maximum number num_leaves; the learning rate learning_rate, also known as "step" or "shrinkage" shrinkage; a coefficient is used to reduce the contribution of each tree. A lower learning rate means more cautious iteration at each step, which usually requires more weak learner decision trees n_estimators to cooperate, but can make the model more robust and generalizable.

[0040] In Step 3, five-fold cross-validation technology is used to train and evaluate candidate machine learning algorithms.

[0041] The five-fold cross-validation technology specifically includes: randomly dividing the training dataset into five equal parts, taking four parts as a sub-training set to train the model, and the remaining one part as a validation set to evaluate the model performance, repeating five times, and calculating the average value of the performance indicators.

[0042] In Step 3, Bayesian optimization is used to automatically optimize the hyperparameters of each candidate model. The Bayesian optimization method uses existing experimental results, i.e., prior knowledge, to guide the search direction of subsequent experiments. Its framework mainly consists of two core parts: Surrogate model: This is a probabilistic model, usually built using Gaussian Process, GP. The role of the surrogate model is to fit an approximation of the "black box function" based on the existing hyperparameter combinations and their corresponding model performance data. Acquisition function: The acquisition function uses the information provided by the surrogate model to decide the next hyperparameter point to try; its goal is to strike a balance between 'exploration' and 'exploitation'; exploration refers to trying areas with higher uncertainty, as these areas may hide better results; exploitation refers to conducting a more detailed search near the areas with the best known performance.

[0043] In Step 3, the performance evaluation indicators for evaluating and selecting the model include one or more of the following: determination coefficient R 2 , root mean square error RMSE, mean absolute error MAE, and mean absolute percentage error MAPE; the model with the optimal comprehensive indicators such as the highest R 2 and the lowest RMSE / MAE / MAPE in cross-validation is selected as the optimal prediction model.

[0044] In Step 4, the correlation coefficient, standard deviation, and RMSE of the optimal prediction model on different feature subsets are evaluated and compared visually using tools such as Taylor Diagram to intuitively select the optimal feature subset.

[0045] In Step 5, the SHAP (SHapley Additive exPlanations) method is used to calculate the contribution of each input feature in the optimal feature subset to the prediction of each target performance. SHAP analysis includes: (1) Global explanation: analyze the average SHAP absolute value of each feature to obtain the feature importance ranking; analyze the relationship between feature value and SHAP value, such as the bee plot, to understand the overall positive and negative impact trend of the feature on the prediction result. (2) Dependency analysis: draw a SHAP dependency graph for a single feature to show how changes in the feature value specifically affect the model output, and observe the interaction effects with other features.

[0046] In Step 5, the trained optimal model and preprocessing process are integrated into a graphical user interface (GUI) software.

[0047] The present application can solve the following technical defects of traditional UHPC performance prediction methods: 1. Incomplete data set: existing methods often ignore key influencing factors such as temperature, aggregate particle size, and specimen size, or rely on data sets with insufficient data volume, resulting in limited prediction model accuracy. 2. Insufficient consideration of data processing and feature engineering: data preprocessing and feature engineering methods have a significant impact on the prediction accuracy and generalization ability of machine learning models, but existing research has insufficient consideration of these aspects. 3. Poor interpretability of models: Machine learning models are often considered "black box" models, and the relationship between inputs and outputs is difficult to interpret. The interpretability of models remains an important research challenge.

[0048] The main objectives of the present invention include: 1. Building a more comprehensive UHPC performance prediction model: Considering multiple influencing factors, a UHPC performance prediction model is established, including material composition, curing conditions, aggregate characteristics, etc., to improve prediction accuracy. 2. Optimizing data processing and feature engineering methods: Effective data preprocessing methods (such as outlier processing, missing value filling, etc.) and feature engineering methods (such as feature selection, feature extraction, etc.) are used to improve the generalization ability of the model. 3. Improving the interpretability of the model: Using SHAP and other model interpretation methods, the influence of input variables on the prediction results is analyzed, and the influencing mechanism of UHPC performance is revealed. 4. Developing UHPC performance prediction tools: The proposed prediction method is integrated into a visual software tool to provide technical support for the design and application of UHPC materials.

[0049] Example 1: A specific implementation process of an interpretable UHPC multi-performance prediction method based on machine learning is described, taking the prediction of UHPC compressive strength CS as an example. The same method is also applicable to the prediction of flexural strength, fluidity, porosity, and other properties.

[0050] S1: Building a UHPC compressive strength dataset This embodiment first builds a dataset for predicting the compressive strength of UHPC. The data sources include two parts: 1) self-conducted indoor test data, a total of 110 groups; 2) published data collected through literature research, a total of 814 groups. The initial dataset contains a total of 924 samples.

[0051] The dataset contains 20 input variables (features) and 1 output variable (target performance, i.e. compressive strength CS). The input variables cover material composition, curing conditions, aggregate characteristics, and specimen size, etc. key factors, and the specific variables and their statistical characteristics, such as range, mean, standard deviation, skewness, kurtosis, as shown in Table 1.

[0052] Table 1 Basic information of UHPC dataset

[0053] These variables include, for example, cement type CT, cement dosage C, slag S, silica fume SF, fine aggregate maximum particle size MSF, fiber dosage F, fiber diameter FD, curing temperature T, curing age Age, specimen size SS, etc.

[0054] S2: Data preprocessing and feature engineering S2.1 Outlier detection and processing: For the initial data set (924 groups), the Isolation Forest (IF) algorithm based on the principle of binary search tree (BST) was used for outlier detection. In this embodiment, 146 groups of abnormal data points were detected and removed, and 778 groups of valid data were retained for subsequent analysis.

[0055] S2.2 Data normalization: For the 778 groups of cleaned valid data, Z-score standardization method was used for normalization processing, so that all feature values obey the normal distribution with mean of 0 and standard deviation of 1. The processing formula is shown in equations (1)-(3).

[0056] (1) (2) (3) S2.3 Feature analysis and subset division: Feature analysis was performed on the normalized data. Pearson correlation coefficient was used to analyze the linear correlation between features. ANOVA F value test and mutual information (MI) method were used to analyze the linear and nonlinear relationship and importance between input features and output target CS. According to the MI score, different thresholds were set, such as MI>0, >0.1, >0.2, >0.3. The 18 input features (after preliminary screening) were divided into 4 different dimensional feature subsets: FS_18 (18 features), FS_14 (14 features), FS_10 (10 features), and FS_6 (6 features). The specific feature composition is shown in Table 2.

[0057] Table 2 Classification of feature sets based on MI method

[0058] S3: Training and selection of optimal machine learning prediction model S3.1 Candidate model: Six machine learning algorithms, including ridge regression (RR), support vector regression (SVR), random forest (RF), gradient boosting decision tree (GBDT), extreme gradient boosting (XGBoost), and light gradient boosting machine (LightGBM), were selected as candidate models.

[0059] S3.2 Data division and cross-validation: The 778 groups of valid data were divided into 80% (622 groups) as training set and 20% (156 groups) as test set. The 5-fold cross-validation method was used for model training and validation on the training set.

[0060] S3.3 Hyperparameter optimization: For each candidate model, the Bayesian optimization method is used to search for the optimal hyperparameter combination within its preset hyperparameter space (as shown in Table 3) during 5-fold cross-validation.

[0061] Table 3. Hyperparameter space of ML models

[0062] S3.4 Model performance evaluation and comparison: The coefficient of determination R2, root mean square error RMSE, mean absolute error MAE, and mean absolute percentage error MAPE are used as evaluation indicators. The average performance indicators of the 6 models on 4 different feature subsets (FS_18, FS_14, FS_10, FS_6) obtained through 5-fold cross-validation are compared.

[0063] S3.5 Optimal model determination: The analysis and comparison results show that the performance of RR and SVR models is poor (R2 is significantly lower than other models). Among the four ensemble learning models of RF, GBDT, XGBoost, and LightGBM, LightGBM shows higher prediction accuracy (such as R2>0.95) and better stability (DR2, DRMSE, DMAE, DMAPE, etc. are relatively small) on different feature subsets. Therefore, LightGBM is determined as the optimal machine learning model for predicting the compressive strength of UHPC. S4: Optimal feature subset determination The LightGBM model determined as the optimal model is trained and tested on the four feature subsets FS_18, FS_14, FS_10, and FS_6. The Taylor Diagram is used to comprehensively compare the correlation coefficient, standard deviation, and RMSE of the prediction results of the model on each subset with the actual observed values.

[0064] The results show that on the FS_10 feature subset, the prediction points of the LightGBM model are closest to the observation points on the Taylor Diagram, showing the highest correlation, the closest dispersion (standard deviation), and the smallest centralized root mean square error, and the Skill Score is close to 1.0. FS_14 and FS_18 may introduce interference due to too many features, and their performance is slightly worse; FS_6 has the worst performance due to insufficient information caused by too few features. Therefore, FS_10 (containing C, SF, W, FAG, MSF, FD, SP, T, Age, SS, a total of 10 features) is determined as the optimal feature subset. S5: Model interpretability analysis based on SHAP The optimal model LightGBM trained on the optimal feature subset FS_10 is subjected to interpretability analysis using the SHAP method.

[0065] S5.1 Global interpretation: Calculate and display the average SHAP absolute value (feature importance) and SHAP value distribution of each feature in FS_10 on CS prediction. The results show that, in this embodiment, the curing age Age has the greatest impact on CS, followed by the silica fume SF content and the fiber length FL (although FL is not in FS_10, it may refer to fiber-related features such as FD, fiber diameter FD, specimen size SS, cement dosage C, etc. The color in the figure distinguishes the high and low values of the features, and the position of the points represents the SHAP value (positive or negative), revealing the overall impact direction of each feature on CS prediction. For example, increasing Age generally results in a positive SHAP value (increasing CS), and increasing SF content within a certain range also increases CS, while excessive SP content may result in a negative SHAP value (decreasing CS).

[0066] S5.2 Dependency analysis: Draw SHAP dependency plots for key features such as Age, T, SP, W, C, and SF to show the relationship between changes in individual feature values and their SHAP values, further revealing their non-linear impact patterns on CS prediction. For example, the impact of Age on SHAP value presents a non-linear feature of rapid growth at the beginning and then tends to be flat. S6: Prediction application and performance verification Using the optimal LightGBM model trained on FS_10, predict the reserved 20% test set data.

[0067] On the test set of this embodiment, the performance indicators of the final model are: R2 = 0.9659, RMSE = 6.68 MPa, MAE = 4.67 MPa. Comparing this result with other existing researches, as shown in Table 4, it indicates that the method proposed in this invention has advantages in prediction accuracy and generalization ability.

[0068] Table 4

Claims

1. A machine learning based method for multi-property prediction of ultra-high performance concrete, characterized in that, The method comprises the following steps: Step 1, establishing a data set: through experiments and collecting data in domestic and foreign literature, a UHPC performance data set containing multiple input variables and output variables is constructed, and the variables cover data information of material type, content, curing condition, aggregate particle size and specimen size; Step 2, data preprocessing: the data quality is improved through data normalization and outlier detection methods, and different feature subsets are divided through feature selection and processing; Step 3, establishing an optimal prediction model: based on the feature subsets, multiple different machine learning algorithms are trained, and the machine learning algorithm with the best training effect is selected as the optimal prediction model; Step 4, selecting the optimal feature subset: based on the optimal prediction model, the prediction effect of the model on different feature subsets is evaluated, and the optimal feature subset is selected; Step 5, explaining the influence of features on model prediction: based on the optimal prediction model and the optimal feature subset, the contribution of each feature to the prediction result is calculated to help understand the decision-making process of the model; Step 6, predicting the performance of ultra-high performance concrete: the parameters of the ultra-high performance concrete to be predicted are input into the optimal prediction model to obtain the predicted value of the performance. 2.The method of claim 1, wherein, In Step 1, the input variables include: 1) material composition related variables: cement type, cement dosage, silica fume dosage, slag dosage, fly ash dosage, quartz powder dosage, fine aggregate dosage, maximum particle size of fine aggregate, coarse aggregate dosage, maximum particle size of coarse aggregate, fiber type, fiber dosage, fiber length, fiber diameter and water reducing agent dosage; 2) curing and environmental condition related variables: curing temperature and curing age; 3) specimen information related variables: specimen size.

3. The machine learning-based multi-performance prediction method of ultra-high performance concrete according to claim 2, characterized in that, Step 2 includes the following steps: Step 2.1, Z-score normalization method is used to convert the data into a standard normal distribution with a mean of 0 and a standard deviation of 1, so as to eliminate the dimension difference between different features; Step 2.2, the Isolation Forest (IF) algorithm based on Binary Search Tree (BST) is used to detect outliers and remove abnormal data; Step 2.3, Pearson correlation analysis, Mantel test, one-way ANOVA and mutual information (MI) are used to analyze and screen the features; based on the MI score threshold, the features are divided into multiple feature subsets with different dimensions.

4. The machine learning-based multi-performance prediction method of ultra-high performance concrete according to claim 3, characterized in that, In Step 3, the multiple different machine learning algorithms specifically include at least two of ridge regression (RR), support vector regression (SVR), random forest (RF), gradient boosting decision tree (GBDT), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM) and regression algorithm; The loss function of the ridge regression (RR) can be expressed as follows: ; wherein, n is the number of samples, m is the number of features, y i is the target value of the i th sample, x ij is the i th feature value of the j th sample, is the weight of the j th feature, λ is the regularization parameter, which controls the strength of the regularization term; the formula of the regularization term is the latter part; in the case of no regularization, i.e. λ = 0, the ridge regression degenerates into the ordinary least squares regression; the training process of the ridge regression model is the process of solving the optimal parameters L θ that minimize the loss function θ .​ For ridge regression, there is a closed-form solution, i.e. an analytical solution, which can directly calculate the optimal parameters: ; wherein, X is a matrix n × m wherein each row represents a sample and each column represents a feature; y is a target value vector, λ is a regularization parameter, I is an identity matrix; The support vector regression SVR is an application of support vector machine SVM in regression problems; for a given training data set D = {( , ), ( , ),..., ( , )} where is an input feature vector, is a corresponding real value; the goal of SVR is to find a linear function: ; wherein, w is a weight vector, b is a bias term; SVR allows for errors within a certain range to not be computed into the loss; this range is called the ε-insensitive band; the loss is only computed when the absolute value of the difference between the predicted value and the true value is greater than ε; to handle data points that are outside of the ε-insensitive band, slack variables and are introduced; represents the deviation of the sample point above the ε-insensitive band; represents the deviation of the sample point below the ε-insensitive band; the optimization objective of SVR can be represented as the following original formula: Minimize: 0.5*|| w ||² + C ∑( + ); with the constraint: ; ; , ; where, w ||² is the regularization term, which is used to control the complexity of the model, make it smoother, prevent overfitting; C is the penalty coefficient, which is a hyperparameter, used to balance the regularization term and the training error term; C The larger the value, the greater the penalty for samples outside the ε-insensitive band, and the model may become more complex and even overfit; C The smaller the value, the stronger the fault tolerance of the model, which may lead to underfitting; the SVR training process is to input the configured parameters and prepared data into the algorithm; by solving a convex quadratic programming problem, the optimal weight vector w and bias b are obtained; The random forest (RF) improves the accuracy and robustness of the model by constructing multiple decision trees and integrating their results. The training process of random forest introduces two kinds of randomness, random sampling and random feature selection; random sampling refers to randomly sampling with replacement from the original training dataset containing N samples to form a bootstrapped sample set with the same size as the original dataset for N times, and the process repeats for T times to get T different bootstrapped sample sets for training T decision trees; random feature selection refers to randomly selecting M features from the m features to split each node of each decision tree to increase the difference between trees; The gradient boosting decision tree GBDT trains a series of weak learners iteratively, each new tree learns the residual of the cumulative result of all previous trees, and finally adds the prediction results of all trees to obtain the final prediction value; the final model of GBDT is an additive model, and its mathematical form is as follows: Suppose we have M decision trees, and for an input sample x the final prediction F M ( x ) is the sum of all the tree predictions: ; in, F M ( x This is the final prediction model; y 0 represents the initial predicted value of the model; For the first m The prediction results of a decision tree; this tree itself is used to fit the residuals of the previous model; v The learning rate is a hyperparameter between 0 and 1; its function is to discount the prediction results of each new tree, preventing a single tree from contributing too much to the model, thereby reducing the risk of overfitting and making the model more robust. The extreme gradient boosting XGBoost includes the following contents: Regularization: built-in L 1 and L 2 regularization terms to effectively control model complexity and prevent overfitting; High-order gradient information: the second-order Taylor expansion is used, and the first-order and second-order gradients are used, so that the optimization process is more accurate and converges faster; Efficient split point searching algorithm: when searching for the best split point, a greedy strategy is adopted; when the data volume is huge and cannot be read into the memory, XGBoost proposes an approximate algorithm based on quantiles, which greatly improves the efficiency; Sparsity-aware: can automatically process sparse data and missing values; when searching for a split point, the gain of dividing missing values into the left and right child nodes is calculated respectively, and the direction with greater gain is selected as the default division direction of the missing values; Parallel computing and system optimization: parallel computing can be performed at the feature level; through the ingenious design of data structure, the gradient and other intermediate calculation results can be better utilized CPU cache, reducing memory reading time; The light gradient boosting machine LightGBM improves the training efficiency and memory performance of the model when processing large-scale data through histogram algorithm, one-side gradient sampling GOSS and exclusive feature binding EFB; the histogram method is used to find the optimal split point of the tree, that is, the continuous floating-point feature values are discretized into a fixed number of "barrels” bins, and a histogram of feature values is constructed; during the process of finding the split point, only the boundaries of these barrels are traversed as candidate split points; the exclusive feature binding EFB identifies mutually exclusive or approximately mutually exclusive features through algorithms such as graph coloring, and binds them into a single "super feature”, thereby reducing the number of features; LightGBM adopts a tree growth strategy of growing leaves, and each time it finds the leaf with the maximum split gain from all current leaves to split, which may increase the risk of overfitting on small data sets, but can converge to a high-precision model faster on large data sets; The decision tree in the ensemble learning algorithm constructs a tree structure model from data features by learning a series of simple "yes / no” rules, thereby predicting new samples; starting from the root node, for each split child node, the optimal feature is repeatedly selected for splitting until the stopping condition is reached, such as the number of samples contained in the current node being lower than the preset threshold min_child_samples; the depth of the tree reaches the preset maximum value max_depth; the number of leaf nodes in a tree reaches the maximum number num_leaves; the learning rate learning_rate in the ensemble model, also known as "step” or "shrinkage”; a coefficient is used to reduce the contribution of each tree.

5. The machine learning-based multi-performance prediction method of ultra-high performance concrete according to claim 4, characterized in that, In Step 3, the five-fold cross-validation technique is used to train and evaluate the candidate machine learning algorithms.

6. The machine learning-based ultra-high performance concrete multi-performance prediction method of claim 5, wherein, The five-fold cross-validation technique specifically includes: randomly dividing the training dataset into five equal parts, taking four parts as a sub-training set to train the model in turn, and taking the remaining one part as a validation set to evaluate the model performance, repeating five times, and calculating the average value of the performance indicators.

7. The machine learning-based ultra-high performance concrete multi-performance prediction method of claim 6, wherein, In Step 3, Bayesian Optimization is used to automatically optimize the hyperparameters of each candidate model. The Bayesian Optimization method uses existing experimental results, i.e., "prior knowledge", to guide the subsequent search direction; its framework consists of two parts: Surrogate model: This is a probabilistic model constructed using Gaussian Process (GP); the role of the surrogate model is to fit an approximation of the "black box function" based on the existing hyperparameter combinations and their corresponding model performance data; Acquisition function: The acquisition function uses the information provided by the surrogate model to decide the next hyperparameter point to try; the goal is to strike a balance between "exploration" and "exploitation"; exploitation refers to searching in the vicinity of the currently known best-performing area.

8. The machine learning-based ultra-high performance concrete multi-performance prediction method of claim 7, wherein, The performance evaluation index for evaluating and selecting the model in Step 3 specifically includes one or more of a determination coefficient R 2 , a root mean square error RMSE, a mean absolute error MAE, and a mean absolute percentage error MAPE; and the optimal prediction model is selected as the one with the optimal comprehensive index in cross validation.

9. The machine learning-based ultra-high performance concrete multi-performance prediction method of claim 8, wherein, In Step 4, Taylor Diagram and other visualization tools are used to comprehensively evaluate the correlation coefficient, standard deviation, and RMSE of the optimal prediction model on different feature subsets, and intuitively compare and select the optimal feature subset.

10. The machine learning-based ultra-high performance concrete multi-performance prediction method of claim 9, wherein, In Step 5, SHAP (SHapley Additive exPlanations) is used to calculate the contribution of each input feature in the optimal feature subset to the model's prediction of each target performance. SHAP analysis includes: (1) Global interpretation: analyze the average SHAP absolute value of each feature to obtain the feature importance ranking; analyze the relationship between feature value and SHAP value in the bee plot to understand the overall positive and negative impact trend of the feature on the prediction result; (2) Dependency analysis: draw the SHAP dependency graph of a single feature to show how changes in the feature value specifically affect the model output, and observe the interaction effects with other features; In Step 5, the trained optimal model and preprocessing process are integrated into the Graphical User Interface (GUI) software.

Citation Information

Cited By

  • Time domain electromagnetic transmitter performance evaluation method based on machine learning

    CN121786791A

  • Personalized acupoint physiotherapy wearing system and neck acupoint personalized positioning method

    CN121845934A

  • Assembly type prefabricated beam bridge construction quality intelligent evaluation method based on integrated tree model

    CN121960994A

  • Intelligent optimization design method for mix proportion of low-carbon concrete based on data-driven and mechanism constraint fusion

    CN122474226A