A machine learning based ultra-high performance concrete composition design method

CN122822154APending Publication Date: 2026-09-25NINGBO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610867578.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]为克服传统超高性能混凝土(UHPC)研发中“经验试错周期长、多组分协同效应难捕捉、性能预测精度低”的技术缺陷,本发明提供一种基于衍生特征嵌入与多尺度融合的机器学习UHPC性能预测及配合比设计方法

Benefits of technology

突破传统缺陷:首次通过衍生特征量化UHPC组分协同效应(如硅灰-钢纤维致密化协同、水胶比-密实度关联),解决了传统经验试错法“忽略组分交互作用、无法跨尺度融合微观属性与宏观参数”的问题,避免大量无效试配导致的资源浪费;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822154A_ABST
    Figure CN122822154A_ABST
Patent Text Reader

Abstract

The application provides a kind of ultra-high performance concrete component design method based on machine learning, related to ultra-high performance concrete performance prediction and machine learning engineering application field.The application first proposes a kind of component intelligent design method based on machine learning and multi-objective optimization specially for UHPC, constructs mechanism embedded feature engineering, and quantifies the densification contribution of silica fume and the bridging toughening effect of steel fiber through derived characteristic variables such as water-binder ratio, total amount of cementitious materials, mineral admixture proportion and steel fiber volume fraction;Meanwhile, the prediction results of different base learning machines are weighted and fused by using Stacking integrated learning framework, the trained performance prediction model is used as the fitness evaluation function of multi-objective optimization algorithm, and the UHPC candidate mix proportion that meets the strength constraint is searched by combining with NSGA-II algorithm and other algorithms, to realize the intelligent design process of "mix proportion input-performance prediction-Pareto solution screening-experimental verification".
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This study presents a machine learning-based method for predicting the mechanical properties and designing the composition of ultra-high performance concrete, which is based on feature engineering and ensemble learning. This method is applicable to the fields of ultra-high performance concrete performance prediction and machine learning engineering applications. Background Technology

[0002] In the research and development of ultra-high performance concrete (UHPC), traditional trial-and-error methods remain the mainstream experimental approach in the early stages. However, because UHPC involves a multi-component system including cement, silica fume, fly ash, quartz sand, steel fibers, and water-reducing agents, and the influence of each component on mechanical properties (compressive and flexural strength) is synergistic or competitive, the research and development process requires repeated adjustments to mix proportions, verification of curing processes, and performance testing, forming a cyclical iteration of "design-mixing-testing." This model not only prolongs the cycle from laboratory research to engineering application of UHPC but also results in a large amount of resource waste due to ineffective mix design. Furthermore, UHPC has high-dimensional component properties (such as more than 10 key variables including unit volume mass and curing parameters), making it difficult for traditional methods to efficiently process multi-source data and accurately capture core mechanisms such as silica fume densification and steel fiber interface strengthening. This approach is no longer sufficient to meet the rapid research and development needs of marine engineering scenarios.

[0003] With the development of artificial intelligence technology, machine learning methods have provided new pathways for materials research and development. In the field of ordinary concrete, existing technologies have attempted to utilize machine learning for mix design (e.g., CN202210762533) or durability prediction (e.g., CN202411298631). However, these methods have significant limitations when applied to ultra-high-performance concrete (UHPC): First, the model objectives are singular. Existing design methods mostly focus on basic indicators such as the strength and flowability of ordinary concrete, failing to encompass the synergistic optimization of multiple objectives pursued by UHPC, such as ultra-high compressive strength and high toughness. Second, the input feature representation is insufficient. Existing prediction models typically use the absolute mass of components or basic mix proportions as input, failing to quantify and integrate the intrinsic properties of core UHPC components (such as the pozzolanic activity of silica fume and the morphology of steel fibers) into feature engineering. The models are essentially learning "numerical proportions" rather than "physicochemical mechanisms," resulting in limited accuracy in predicting key UHPC performance and poor extrapolation. Third, the unique characteristics of the system are ignored. The low water-cement ratio, high cementitious material content, and fiber-reinforced system of UHPC cause its performance evolution to differ fundamentally from that of ordinary concrete. Existing models trained from ordinary concrete data, or models that do not consider the role of fibers, are difficult to directly transfer and accurately capture the unique behavior of UHPC.

[0004] To address the aforementioned gaps and deficiencies, this invention proposes for the first time a component-based intelligent design method specifically for UHPC (Ultra-High-Performance Compatibility), based on machine learning and multi-objective optimization. The core of this method lies in: constructing a feature engineering approach with embedded mechanisms, quantifying the densification contribution of silica fume and the bridging and toughening effect of steel fibers through derived features such as water-cement ratio, total amount of cementitious materials, proportion of mineral admixtures, and steel fiber volume fraction; simultaneously, utilizing a Stacking ensemble learning framework to weightedly fuse the prediction results of different base learners, where the weighting coefficients are automatically learned from the training set rather than being manually fixed. Furthermore, the trained performance prediction model is used as the fitness evaluation function of the multi-objective optimization algorithm, combined with algorithms such as NSGA-II to search for UHPC candidate mix proportions that satisfy strength constraints, thus realizing an intelligent design process from "mix proportion input—performance prediction—Pareto solution screening—experimental verification".

[0005] This invention effectively solves the three major technical pain points of traditional empirical methods (low efficiency) and existing machine learning methods for UHPC ("superficial feature representation, inaccurate model prediction, and single optimization objective"), providing a brand-new intelligent solution for the rapid, accurate, and high-performance development of UHPC. Summary of the Invention

[0006] To overcome the technical shortcomings of traditional ultra-high performance concrete (UHPC) R&D, such as "long trial-and-error cycles, difficulty in capturing multi-component synergistic effects, and low performance prediction accuracy," this invention provides a machine learning-based UHPC performance prediction and mix design method based on derived feature embedding and multi-scale fusion. This method replaces traditional simple component inputs by constructing a derived feature system with mechanism embedding, accurately quantifying core functions such as silica fume densification and steel fiber bridging; and utilizes an integrated learning model and optimization algorithm to achieve intelligent reverse design from multiple performance objectives to the optimal mix proportion.

[0007] The specific steps are as follows: Step 1: Construct a standardized UHPC raw dataset. The system collects UHPC formulation and performance data from published literature to construct an initial dataset covering raw material composition, process conditions, and key performance indicators.

[0008] Step 2: Data preprocessing and cleaning.

[0009] Step 3: Dataset partitioning and normalization.

[0010] Step 4: Construction of a basic prediction model based on original features and weighted fusion using Stacking. Using the unit volume mass and maintenance parameters of each component as input features, and compressive strength and flexural strength as output targets, decision trees, random forests, support vector machines, XGBoost, multilayer perceptrons, and Stacking ensemble models are constructed and trained based on the training set. The Stacking model adopts a two-layer structure. The first layer consists of multiple base learners outputting predicted values, and the second layer uses a linear regressor meta-learner to weightedly fuse the outputs of each base learner. Its expression is: In the formula, b0 is the intercept term, and b1 to b5 are weighting coefficients automatically learned from the training set. The performance of each model is evaluated using the test set, and indicators such as the coefficient of determination, root mean square error, and mean absolute error are compared comprehensively. The results show that the Stacking ensemble model has significantly better prediction accuracy than the single model and has been established as the foundational prediction model for subsequent optimization.

[0011] Step 5: Construct a mechanism-embedded derived feature system and optimize the model. To overcome the shortcomings of traditional models, such as "weak physical meaning of input features and difficulty in capturing nonlinear interactions between components," this invention constructs a core derived feature system for UHPC. Based on materials science principles, this system transforms the original component mass into derived features with clear physical meaning through proportional calculation, volume fraction conversion, and functional parameter reconstruction. These derived features, along with the original input, serve as input features for the Stacking model. It should be noted that the "weighted processing" in this invention mainly refers to the weighted fusion of prediction results from each base learner by the second-layer meta-learner of the Stacking model; the derived feature part mainly reflects the reconstruction of calculable proportions, volume fractions, and functional parameters, rather than manually assigning fixed weights to the raw material mass. In actual prediction, the candidate UHPC mix proportion X is first input; then the derived feature φ(X) is calculated according to the above formula; then X is normalized using the maximum and minimum values ​​saved during the training phase to obtain X*; and then X* is input into each base learner to obtain... f 1(X), f 2(X), ... f m (X); Finally, the second-layer meta-learner follows... Output predicted results such as compressive strength and flexural strength.

[0012] Step 6: Intelligent Design of UHPC Mix Proportion Based on Multi-Objective Optimization. The machine learning prediction model trained in Step 5 is embedded as a fitness evaluation function into the NSGA-II multi-objective optimization algorithm to achieve intelligent optimization of the UHPC mix proportion. The optimization variable is the unit volume usage of each raw material: q=[C,SF,SL,LP,QP,FA,W,Sa,Fi,SP] where C, SF, SL, LP, QP, FA, W, Sa, Fi, and SP represent the usage of cement, silica fume, slag, limestone powder, quartz powder, fly ash, water, sand, steel fiber, and high-efficiency water-reducing agent, respectively. The optimization objectives are to maximize the 28-day compressive strength Fc and the 28-day flexural strength Fs. Since NSGA-II is usually solved as a minimization problem, the objective function can be expressed as: During the optimization process, for each candidate mix proportion q generated by NSGA-II, it is first determined whether it meets the constraints such as the range of raw material usage, water-cement ratio, steel fiber volume fraction, and water-reducing agent dosage. If the constraints are met, its derived features are calculated according to step 5 and input into the trained machine learning model to obtain the corresponding compressive strength and flexural strength evaluation results. Subsequently, the algorithm performs non-dominated sorting, selection, crossover, and mutation based on these evaluation results, continuously updating the candidate mix proportion population, and finally obtaining a set of Pareto optimal mix proportion schemes.

[0013] Step 7: Experimental Verification of Optimized Mix Proportions. To verify the reliability of the machine learning prediction and multi-objective optimization method proposed in this invention, two representative mix proportion schemes were selected from the Pareto feasible solution set for experimental verification. Each scheme covered different optimization directions for high compressive strength and high flexural strength. At least three parallel specimens were prepared for each scheme, and 28-day compressive and flexural strength tests were conducted. See... Figure 11 .

[0014] Furthermore, to ensure data consistency, variables are strictly controlled: the composition of raw materials is accurate to the mass of each component per unit volume (kg / m³); performance indicators focus on mechanical properties (compressive strength, flexural strength).

[0015] Furthermore, the initial dataset undergoes systematic preprocessing to improve data quality. First, samples with missing data or significant record distortion are removed. Then, a boxplot method based on the interquartile range (IQR) is used to identify and remove statistical outliers from numerical features. For the limited number of non-numerical features, they are converted into discrete numerical values ​​through label encoding. These steps result in a well-structured and reliable initial dataset for UHPC modeling.

[0016] Furthermore, the cleaned dataset is randomly divided into training and test sets according to a preset ratio (e.g., 8:2). To eliminate the differences in scale between different features and accelerate model convergence, the max-min normalization method is used to standardize all numerical features, linearly mapping them to a unified interval.

[0017] Furthermore, the weighting coefficients in this invention b j The values ​​are not assigned by human experience, but are automatically obtained during the training process. Specifically, K-fold cross-validation is first used to obtain the out-of-fold prediction results of each base learner on the training samples, forming the meta-learning training matrix Z=[ f 1(X), f 2(X),…,f m (X)]; then, using the true performance value y as the supervision signal, the sum of squared residuals is minimized using a linear regressor learner. Thus, to obtain b 0 and b j After model training, the prediction performance is evaluated using a test set. Evaluation metrics include the coefficient of determination (R²), mean absolute error (MAE), mean squared error (MSE), mean absolute percentage error (MAPE), and prediction accuracy, calculated using the following formulas: ; ; ; ; .in, y i For the first i Measured values ​​of a sample i These are the model's predicted values. The average value is the measured value, and n is the number of test samples. The larger the R² and Accuracy, and the smaller the MAE, MSE, and MAPE, the better the model's predictive performance.

[0018] Advantages and technical effects of the present invention Breaking through traditional limitations: For the first time, the synergistic effect of UHPC components is quantified through derived features (such as the densification synergy between silica fume and steel fiber, and the correlation between water-binder ratio and density), which solves the problem of traditional empirical trial and error methods "ignoring component interactions and being unable to integrate microscopic properties and macroscopic parameters across scales", and avoids the waste of resources caused by a large number of ineffective trial formulations; Significantly improved prediction accuracy: Based on the weighted fusion of the second layer of the Stacking model, after introducing derived features with physical meaning, the model can comprehensively utilize the complementary information of base learners such as XGBoost, SVM, RF, DT and MLP. The prediction accuracy of compressive strength and flexural strength on the test set remains at a high level, meeting the high requirements of offshore UHPC for performance prediction. The above description is merely an overview of the technical solution of the present invention. To better understand the technical means, it can be implemented in accordance with the contents of the specification and the accompanying drawings. The details of the present invention will be further illustrated below through specific embodiments. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method of the present invention.

[0020] Figure 2 Box plots are processed to normalize the data.

[0021] Figure 3 This is a diagram showing the predicted results of the compressive strength model.

[0022] Figure 4 This is a diagram showing the predicted results of the flexural strength model.

[0023] Figure 5 The image shows the prediction results of the Stacking model before optimization.

[0024] Figure 6 The image shows the prediction results of the optimized Stacking model.

[0025] Figure 7 The Pareto optimal solution set diagram prioritizes compressive strength.

[0026] Figure 8 The mix proportion diagram is for the Pareto optimal solution set with priority given to compressive strength.

[0027] Figure 9 This is a Pareto optimal solution set diagram prioritizing flexural strength.

[0028] Figure 10 The mix proportion diagram is for the Pareto optimal solution set prioritizing flexural strength.

[0029] Figure 11 This is an error graph of the experimental results. Specific Implementation To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. In the following description, specific details such as specific configurations and components are provided merely to help fully understand the embodiments of this application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. In addition, for clarity and brevity, descriptions of known functions and structures are omitted in the embodiments.

[0031] Example 1 Step 1: Collect UHPC data. Mix proportion data includes the amounts of cement, fly ash, silica fume, blast furnace slag, quartz powder, water, and steel fiber. Predict the UHPC compressive strength and flexural strength parameters. The UHPC database contains 896 data sets, each pair including mix proportion data, environmental parameters, and prediction parameters. Basic information is shown in Table 1 below.

[0032] Table 1 shows some of the collected data. Step 2: Perform systematic preprocessing on the initial dataset. First, delete samples with missing data or obvious data distortion. Then, use a boxplot method based on the interquartile range (IQR) to identify and remove statistical outliers from numerical features; that is, for each feature, exclude extreme samples exceeding Q3 + 1.5 × IQR or falling below Q1 - 1.5 × IQR. For the limited number of non-numerical features, convert them into discrete numerical values ​​through label encoding. See... Figure 2 (a) Compressive strength before normalization; (b) Compressive strength after normalization; (c) Flexural strength before normalization; (d) Flexural strength after normalization.

[0033] Step 3: After label encoding, the initial datasets for compressive strength and flexural strength are divided into training set (R) and test set (T) in an 8:2 ratio, and the divided datasets are normalized; the training set and test set for compressive strength are R1 and T1, respectively, and the sets for flexural strength are R2 and T2. The formula for normalizing data is: in This represents the original data value, i.e., the data before normalization; and These represent the maximum and minimum values ​​in the original dataset, respectively. and These represent the maximum and minimum values ​​within the target interval, respectively. This represents the transformed data value; The six machine learning models constructed in step 4 are Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting (XGboost), and Multilayer Perceptron (MLP) ensemble model with Stacking. The evaluation metrics for the six machine learning models, such as the coefficient of determination (R²) and mean absolute error (MAE), are shown in the following tables. The prediction results of the five machine learning models are shown in Tables 2 and 3, where Table 2 shows the compressive strength results and Table 3 shows the flexural strength results. Table 2 shows the coefficients of determination, root mean square error, and mean absolute error of five machine learning models for predicting compressive strength. Table 3 shows the coefficients of determination, root mean square error, and mean absolute error of the five machine learning models for predicting flexural strength. As shown in the table, in compressive strength prediction, the Stacking model has an R² of 0.9054, which is higher than other single models, indicating that it has a better comprehensive predictive ability for compressive strength. In flexural strength prediction, the Stacking model has an R² of 0.9439, outperforming other models. Considering the unified prediction requirements of both compressive and flexural strength performance objectives, this invention preferably adopts the Stacking ensemble framework, which has model fusion capabilities and generalization stability, as a unified surrogate model for subsequent multi-objective optimization. This model can adaptively weight and fuse the prediction results of different base learners through a meta-learner, thereby improving the overall prediction stability and applicability.

[0034] Step 4: This study selected the unit volume doping content of each component of UHPC and the curing parameters as input features, and compressive strength and flexural strength as output targets. Six typical machine learning models were constructed and trained based on the training set. The first layer of the Stacking ensemble model includes base learners such as XGBoost, Support Vector Machine, Random Forest, Decision Tree, and Multilayer Perceptron. The second layer uses a linear regressor meta-learner. The model's prediction performance is as follows: Figure 3-4 As shown. Specifically, in the compressive strength-priority design of Example 1, the compressive strength fusion weight is preferentially used for prediction and evaluation, that is, the compressive strength weights of XGBoost, RF, SVM, DT, and MLP are 0.272, 0.203, 0.210, 0.131, and 0.184, respectively. Therefore, the compressive strength-priority optimization process gives more weight to the prediction results of XGBoost, while retaining the supplementary role of SVM, RF, and MLP in nonlinearity.

[0035] Step 5: Construct a mechanism-embedded derived feature system and optimize the model. To fundamentally overcome the shortcomings of traditional models, such as "weak physical meaning of input features and difficulty in capturing nonlinear interactions between components," this invention constructs a core derived feature system for UHPC. Based on materials science principles, this system transforms the original component mass into eight categories of derived features with clear physical meanings through proportional calculation, volume conversion, and functional parameter reconstruction. These features are then input into the Stacking model along with the original features. Here, "weighting" mainly refers to the weighted fusion of prediction results from different base learners by the Stacking meta-learner; the derived features themselves are calculated according to a defined formula. Key derived features include: ①. Water-cement ratio: Quantify the relationship between the total amount of water and cementitious materials to control workability and strength.

[0036] ②. Total amount of cementitious materials: characterizes the overall scale of the cementing phase.

[0037] ③. Total aggregate: Reflects the total amount of skeleton material used.

[0038] ④. Total amount of additives: Characterizes the total input of key functional components such as steel fibers and water-reducing agents.

[0039] ⑤. Cement to silica fume ratio: This describes the ratio of the main cementing material to the highly active micro powder, which is related to the densification of the microstructure.

[0040] ⑥. Mineral admixture ratio: reflects the degree of compositeness and greenness of the cementitious system.

[0041] ⑦. Water-reducing agent to cementitious material ratio: measures the effectiveness of the admixture relative to the cementitious system.

[0042] ⑧. Steel fiber volume fraction: The core parameter that directly determines the toughening effect of the fiber.

[0043] The Stacking ensemble model was retrained and optimized using this derived feature dataset. Validation results show that the optimized model maintains leading accuracy in predicting mechanical performance, demonstrating the effectiveness of the derived feature system in capturing multi-scale performance correlations in UHPC. The prediction results of the Stacking ensemble model before and after optimization are shown below. Figure 5-6 .

[0044] Step 6: Embed the trained high-precision performance prediction model as a fitness evaluation function into the NSGA-II multi-objective optimization algorithm. Example 1 prioritizes compressive strength as the recommended direction, with optimization variables being the amount of each raw material q=[C,SF,SL,LP,QP,FA,W,Sa,Fi,SP], and optimization objectives being to maximize F_c(q) and maximize F_s(q). In NSGA-II, maximizing the strength objective is transformed into minimizing -Fc(q) and -Fs(q). Each candidate mix ratio q requires first calculating derived features and inputting them into the trained Stacking model to obtain the predicted strength before participating in non-dominated ranking. After obtaining the Pareto solution set, a comprehensive score S(q) is used to select the compressive strength-priority scheme. The Pareto optimal solution set and Pareto optimal mix ratio for compressive strength priority are shown below. Figure 7-8 .

[0045] Example 2 Step 1: Collect UHPC data. Mix proportion data includes the amounts of cement, fly ash, silica fume, blast furnace slag, quartz powder, water, and steel fiber. Predict the UHPC compressive strength and flexural strength parameters. The UHPC database contains 896 data sets, each pair including mix proportion data, environmental parameters, and prediction parameters. Basic information is shown in Table 4 below.

[0046] Table 4 contains some of the collected data. Step 2: Perform systematic preprocessing on the initial dataset. First, delete samples with missing data or obvious data distortion. Then, use a boxplot method based on the interquartile range (IQR) to identify and remove statistical outliers in numerical features. That is, for each feature, extreme samples exceeding Q3 + 1.5 × IQR or falling below Q1 - 1.5 × IQR are excluded. For the limited number of non-numerical features, they are converted into discrete numerical values ​​through label encoding.

[0047] Step 3: After label encoding, the initial datasets for compressive strength and flexural strength are divided into training set (R) and test set (T) in an 8:2 ratio, and the divided datasets are normalized; the training set and test set for compressive strength are R1 and T1, respectively, and the sets for flexural strength are R2 and T2. The formula for normalizing data is: Where x represents the original data value, i.e., the data before normalization; x_max and x_min represent the maximum and minimum values ​​in the original dataset, respectively; y_max and y_min represent the maximum and minimum values ​​in the target interval, respectively; and y_k represents the transformed data value. The six machine learning models constructed in step 4 are Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting (XGboost), and Multilayer Perceptron (MLP) ensemble model with Stacking. The evaluation metrics for the six machine learning models, such as the coefficient of determination (R²) and mean absolute error (MAE), are shown in the following tables. The prediction results of the five machine learning models are shown in Tables 5 and 6, where Table 5 shows the compressive strength results and Table 6 shows the flexural strength results. Table 5 shows the coefficients of determination, root mean square error, and mean absolute error of five machine learning models for predicting compressive strength. Table 6 shows the coefficients of determination, root mean square error, and mean absolute error of five machine learning models for predicting flexural strength. As shown in the table, the Stacking model has an R² of 0.9054 in compressive strength prediction, demonstrating good overall predictive performance. In flexural strength prediction, the Stacking model still achieves an R² of 0.9439, exhibiting the highest prediction accuracy, indicating its equally good predictive ability for flexural strength. Example 2 prioritizes flexural strength, therefore, the analysis of flexural strength results focuses on the prediction results of the Stacking model. Furthermore, to ensure interface consistency among different performance prediction models in the multi-objective optimization process, this invention employs the Stacking integration framework as a unified proxy model, embedding the NSGA-II multi-objective optimization algorithm in a combination of multiple proxy models, including the "compressive strength Stacking prediction model" and the "flexural strength Stacking prediction model."

[0048] Step 4: This study selected the unit volume dosage of each component of UHPC and the curing parameters as input features, and compressive strength and flexural strength as output targets. Six typical machine learning models were constructed and trained based on the training set. Model evaluation used indicators such as R², MAE, MSE, RMSE, MAPE, and Accuracy. For the Stacking model, the first-layer base learner outputs predicted values, and the second-layer linear regression learner performs weighted fusion. Specifically, in the flexural strength-first design of Example 2, the flexural strength fusion weights were prioritized for prediction and evaluation. Specifically, the flexural strength weights of XGBoost, RF, SVM, DT, and MLP were 0.387, 0.142, 0.282, 0.080, and 0.109, respectively. Therefore, the flexural strength-first optimization process clearly favored the XGBoost prediction results, followed by the SVM prediction results, with RF and MLP providing auxiliary corrections, and DT having the lowest weight.

[0049] Step 5: Construct a mechanism-embedded derived feature system and optimize the model. Based on materials science principles, this system transforms the original component masses into eight categories of derived features with clear physical meanings through proportional calculations, volume conversions, and functional parameter reconstruction. These derived features, along with the original features, are input into the prediction model. The derived features include water-cement ratio, total amount of cementitious materials, total amount of skeleton materials, total amount of additives, cement to silica fume ratio, mineral admixture ratio, water-reducing agent to cementitious material ratio, and steel fiber volume fraction.

[0050] ①. Water-cement ratio: Quantify the relationship between the total amount of water and cementitious materials to control workability and strength.

[0051] ②. Total amount of cementitious materials: characterizes the overall scale of the cementing phase.

[0052] ③. Total aggregate: Reflects the total amount of skeleton material used.

[0053] ④. Total amount of additives: Characterizes the total input of key functional components such as steel fibers and water-reducing agents.

[0054] ⑤. Cement to silica fume ratio: This describes the ratio of the main cementing material to the highly active micro powder, which is related to the densification of the microstructure.

[0055] ⑥. Mineral admixture ratio: reflects the degree of compositeness and greenness of the cementitious system.

[0056] ⑦. Water-reducing agent to cementitious material ratio: measures the effectiveness of the admixture relative to the cementitious system.

[0057] ⑧. Steel fiber volume fraction: The core parameter that directly determines the toughening effect of the fiber.

[0058] The Stacking ensemble model was retrained and optimized using this derived feature dataset. Validation results show that the optimized model maintains leading accuracy in predicting mechanical performance, demonstrating the effectiveness of the derived feature system in capturing multi-scale performance correlations in UHPC. The prediction results of the Stacking ensemble model before and after optimization are shown below. Figure 5-6 After the candidate mix ratio q is input, intermediate variables are first calculated according to B=C+SF+SL+FA, P=C+SF+SL+LP+QP+FA, M=SF+SL+LP+QP+FA, A=Sa+QP, and Ad=Fi+SP. Then, derived features such as W / B, M / P, SP / B, and V_f are calculated. Subsequently, normalization is performed according to the maximum and minimum values ​​of the training set, and the results are input into the prediction model to obtain the performance prediction value under the priority target of flexural strength.

[0059] Step 6: Embed the trained high-precision performance prediction model as a fitness evaluation function into the NSGA-II multi-objective optimization algorithm. Example 2 prioritizes flexural strength as the recommended direction, with optimization variables being the amount of each raw material q=[C,SF,SL,LP,QP,FA,W,Sa,Fi,SP], and optimization objectives being to maximize F_c(q) and maximize F_s(q). In NSGA-II, maximizing strength is transformed into minimizing -Fc(q) and -Fs(q). Each candidate mix ratio q requires first calculating derived features and inputting them into the trained Stacking model to obtain predicted strength before participating in non-dominated ranking. After obtaining the Pareto solution set, a comprehensive score S(q) is used to select the flexural strength-priority scheme. The Pareto optimal solution set and Pareto optimal mix ratio for flexural strength priority are shown below. Figure 9-10 .

[0060] Step 7: Experimental Verification of Optimized Mix Proportions. To simultaneously verify the recommended mix proportions obtained in Examples 1 and 2, two representative schemes were selected from the Pareto feasible solution set for experimental verification. Each scheme covered different optimization directions for high compressive strength and high flexural strength. At least three parallel specimens were prepared for each scheme, and 28-day compressive strength and flexural strength tests were conducted. The results are shown in the table below. Figure 11 As shown.

[0061] The prediction error is calculated as Error = |Measured value - Predicted value| / Measured value × 100%. If the prediction errors of compressive strength and flexural strength of each representative scheme are within an acceptable range, it indicates that the closed-loop process of "machine learning prediction model - NSGA-II multi-objective optimization - experimental verification" established in this invention is reliable, as shown in Table 7.

[0062] Table 7 Comparison of Predicted and Actual Measurements The above description is merely a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any changes, modifications, substitutions, integrations, and parameter alterations to these embodiments within the spirit and principles of the present invention, achieved through conventional substitutions or by achieving the same function without departing from the principles and spirit of the present invention, fall within the scope of protection of the present invention.

Claims

1. A machine learning-based method for designing the composition of ultra-high performance concrete, characterized in that, Includes the following steps: Step 1: Construct the original dataset of ultra-high performance concrete, which includes the mix proportions and performance data of ultra-high performance concrete, and construct an initial dataset covering raw material composition, process conditions and key performance indicators. Step 2: Data preprocessing and cleaning; Step 3: Dataset partitioning and normalization; Step 4: Building a basic prediction model based on original features and using Stacking weighted fusion: Using the raw material composition in the original dataset of ultra-high performance concrete as input features and compressive strength and flexural strength as output targets, decision trees, random forests, support vector machines, XGBoost, multilayer perceptrons and stacking ensemble models are constructed and trained based on the training set. The Stacking model employs a two-layer structure. The first layer consists of multiple base learners that each outputs a predicted value. The second layer uses a linear regressive meta-learner to weight and fuse the outputs of each base learner. Its expression is as follows: In the formula, b0 is the intercept term, and b1 to b5 are weighting coefficients automatically learned from the training set; Step 5: Construct a derived feature system for mechanism embedding and optimize the model: The original component mass is transformed into derived features with clear physical meaning through proportional calculation, volume ratio conversion and functional parameter reconstruction. These features, together with the original input, serve as input features for the Stacking model. The Stacking ensemble model is then retrained and optimized using the derived feature dataset. Step 6: Intelligent design of UHPC mix proportions based on multi-objective optimization: The optimized model trained in step 5 is embedded as a fitness evaluation function into the NSGA-II multi-objective optimization algorithm to achieve intelligent optimization of the UHPC mix proportion. The optimization variable is the unit volume usage of each raw material: q=[C,SF,SL,LP,QP,FA,W,Sa,Fi,SP] where C, SF, SL, LP, QP, FA, W, Sa, Fi, and SP represent the usage of cement, silica fume, slag, limestone powder, quartz powder, fly ash, water, sand, steel fiber, and high-efficiency water-reducing agent, respectively. The objective function can be expressed as: During the optimization process, for each candidate mix proportion q generated by NSGA-II, it is first determined whether it meets the constraints of raw material dosage range, water-cement ratio, steel fiber volume fraction, and water-reducing agent dosage. After meeting the constraints, its derived features are calculated according to step 5 and input into the trained machine learning model to obtain the corresponding compressive strength and flexural strength evaluation results. Subsequently, the algorithm performs non-dominated sorting, selection, crossover, and mutation based on these evaluation results, continuously updating the candidate mix proportion population, and finally obtaining a set of Pareto optimal mix proportion schemes.

2. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 1, characterized in that, It also includes step 7: experimental verification of optimized mix proportions: To verify the reliability of the machine learning prediction and multi-objective optimization method proposed in this invention, two representative mix proportion schemes were selected from the Pareto optimal mix proportion schemes for experimental verification. Each scheme covers different optimization directions such as high compressive strength, high flexural strength, low electrical flux, and low cost. At least three parallel specimens were prepared for each scheme, and the compressive strength and flexural strength were measured after 28 days.

3. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 1, characterized in that, The composition of raw materials is standardized to the mass of each component per unit volume, in kg / m³. The performance indicators focus on mechanical properties, including compressive strength and flexural strength.

4. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 1, characterized in that, Data preprocessing and cleaning include: deleting samples with missing data or obvious data distortion, and using box plots based on interquartile range to identify and remove statistical outliers in numerical features.

5. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 4, characterized in that, For a limited number of non-numerical features, they are converted into discrete numerical values ​​through label encoding.

6. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 1, characterized in that, The cleaned dataset is randomly divided into training and test sets. To eliminate the differences in the units of different features and accelerate model convergence, the max-min normalization method is used to standardize all numerical features and linearly map them to a unified interval.

7. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 6, characterized in that, The training set and the test set are divided according to a preset ratio, which is 6-9:1-2.

8. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 1, characterized in that, The weighting coefficients b j The training process automatically yields the following results: First, K-fold cross-validation is used to obtain the out-of-fold prediction results of each base learner on the training samples, forming the meta-learning training matrix Z = [ f 1(X), f 2(X),…,f m (X)]; then, using the true performance value y as the supervision signal, the sum of squared residuals is minimized using a linear regressor learner. Thus, to obtain b 0 and b j After the model is trained, the prediction performance is evaluated using a test set.

9. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 8, characterized in that, The prediction performance is evaluated using a test set and the evaluation metrics include the coefficient of determination (R²), mean absolute error (MAE), mean squared error (MSE), mean absolute percentage error (MAPE), and prediction accuracy. The calculation formulas for these metrics are as follows: ; ; ; ; ,in, y i For the first i Measured values ​​of a sample i These are the model's predicted values. The average value is the measured value, and n is the number of test samples. The larger the R² and Accuracy, and the smaller the MAE, MSE and MAPE, the better the model's predictive performance.

10. The method for designing ultra-high performance concrete composition based on machine learning as described in claim 1, characterized in that, The derived features include multiple of the following: water-cement ratio, total amount of cementitious materials, total amount of aggregates, total amount of additives, ratio of cement to silica fume, ratio of mineral admixtures, ratio of water-reducing agent to cementitious materials, and volume fraction of steel fibers.

Citation Information

Patent Citations

  • Intelligent design method for concrete mix proportion based on machine learning

    CN115206463A

  • Concrete durability prediction method based on data expansion and machine learning

    CN119397415A