Concrete compressive strength prediction method based on GM model and machine learning model
By combining GM model and machine learning model, the input variables of concrete compressive strength prediction method are extended, and the problem of insufficient accuracy of concrete durability prediction in the prior art is solved, and higher prediction accuracy and model adaptability are achieved.
Patent Information
- Application Number
- CN202510208284.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to accurately predict the durability and service life of concrete when facing complex environments and variable material ratios, and machine learning models are prone to overfitting and insufficient accuracy when small samples or high uncertainties.
The concrete compressive strength prediction method based on GM model and machine learning model is adopted. By obtaining trusted literature corpus, the original data points are extracted, and the input variables of the data points are extended using the GM model prediction value and the GM residual model prediction value are used to construct a sample set to train the prediction model, and finally predict the compressive strength of the concrete to be measured based on the prediction model.
The accuracy of concrete compressive strength prediction is significantly improved, the model's learning ability of global and local characteristics of the data is enhanced, the ability to handle abnormal situations is improved, and the model's adaptability and generalization ability is improved.
Smart Images

Figure CN120144952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of non-destructive testing of concrete compressive strength, and particularly to a method for predicting concrete compressive strength based on a GM model and a machine learning model. Background Art
[0002] Concrete is an extremely important material in the construction field, and its durability is directly related to the long-term stability and safety of concrete structures. The durability of concrete has always received great attention in engineering practice and academic research. In practical engineering applications, concrete structures are often in complex and changeable environments and are faced with long-term tests of many adverse factors, such as wet-dry cycles, chemical erosion, etc. These complex environmental factors will gradually lead to the deterioration of the safety and functionality of concrete structures, thereby affecting their service life. The deterioration of concrete durability is a complex process of multi-factor interaction, involving material properties, environmental conditions, construction quality and other aspects. With the continuous deepening of research, scholars have found that there are limitations in single-factor research and have begun to pay attention to the influence of multi-factor coupling on durability. Therefore, accurately predicting the durability and service life of concrete under multi-factor coupling has become a core issue in engineering design and maintenance.
[0003] The prediction of traditional concrete durability is mainly based on experimental data and empirical formulas. Although it can reflect some performances, it is often difficult to achieve accurate prediction when facing complex environments and variable material ratios. On the one hand, empirical formulas are summarized based on experimental data under specific conditions. Due to the uncertainty of the test itself and the complexity of the actual use environment, it is difficult to adapt to diverse actual engineering scenarios. On the other hand, experiments have problems such as long cycle, high cost and low efficiency, which makes it difficult to quickly and accurately predict long-term durability performance. To solve these problems, some new evaluation methods and technologies have been proposed and applied in recent years. For example, data-driven evaluation methods use machine learning and data mining technologies to establish models through monitoring and detection data to evaluate the durability of concrete structures.
[0004] Machine learning technology uses algorithms and statistical models to enable computers to autonomously learn from data and continuously optimize, and then complete prediction and analysis tasks for complex systems. However, in the field of concrete durability testing, it is extremely difficult to obtain the sample size, and the number of real experimental data samples is too small. Machine learning models are prone to overfitting and insufficient accuracy in the case of small samples or high uncertainty. In this regard, some existing technologies attempt to achieve sample amplification through simulation and interpolation methods. However, the data quality obtained by the simulation method is low. When the number of real samples is too small, the high proportion of simulation samples is likely to have a negative impact on the training of the model, and the interpolation method still has a linear relationship with the original real samples and is difficult to solve the above problems of overfitting and insufficient accuracy. Summary of the Invention
[0005] The object of the present invention is to provide a method for predicting the compressive strength of concrete based on the GM model and the machine learning model.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] A method for predicting the compressive strength of concrete based on the GM model and the machine learning model, comprising:
[0008] Step S1: Obtain a credible literature corpus, and extract an original data set based on the credible literature corpus, wherein the original data set includes a plurality of original data groups, each data group consists of a plurality of original data points, and each original data point includes a plurality of input variables and an output variable;
[0009] Step S2: Process each original data group, and respectively calculate the GM model prediction value and the GM residual model prediction value of each original data point based on all the original data points in each original data group, and update the original data point with the GM model prediction value and the GM residual model prediction value as new input variables to obtain an amplified data point;
[0010] Step S3: Construct a sample set based on the amplified data points to train a prediction model;
[0011] Step S4: Predict the compressive strength of the concrete to be tested based on the prediction model.
[0012] The output variable of the original data point is the compressive strength.
[0013] The input variables of the original data point include water-binder ratio, sand ratio, RCA replacement rate, iron tailings sand replacement rate, fly ash content, silica fume content, PVA volume fraction, BF volume fraction, SF volume fraction, Na2SO4 mass fraction, NaCl mass fraction, immersion time, drying time, and number of wet-dry cycles.
[0014] The process of obtaining the GM model prediction value in step S2 specifically includes:
[0015] Step S2-1-1: Obtain an original sequence based on the input variables of all the original data points in the original data:
[0016]
[0017] Where: X 0 is the original sequence, is the initial value of the input variable of the first original data point, m is the length of the sequence, is the initial value of the input variable of the second original data point, is the initial value of the input variable for the m-th original data point;
[0018] Step S2-1-2: Perform a first-order accumulation on the original sequence to obtain a first-order accumulated sequence:
[0019]
[0020] where: X 1 is the first-order accumulated sequence, is the first-order accumulated value of the input variable of the first original data point, is the first-order accumulated value of the input variable of the second original data point, is the first-order accumulated value of the input variable of the m-th original data point, is the first-order accumulated value of the input variable of the k-th original data point, is the first-order accumulated value of the input variable of the i-th original data point;
[0021] Step S2-1-3: Calculate the mean of adjacent elements of the first-order accumulated sequence to obtain a mean sequence:
[0022]
[0023]
[0024] where: Z 1 is the mean sequence, is the sum of the first-order accumulated values of the input variables of the second original data point and the previous original data point, is the sum of the first-order accumulated values of the input variables of the third original data point and the previous original data point, is the sum of the first-order accumulated values of the input variables of the m-th original data point and the previous original data point, is the sum of the first-order accumulated values of the input variables of the (k + 1)-th original data point and the previous original data point, is the first-order accumulated value of the input variable of the (k + 1)-th original data point;
[0025] Step S2-1-4: Construct the first matrix B and the second matrix Y:
[0026]
[0027] Step S2-1-5: Based on the constructed first matrix and second matrix, calculate the first model parameter α and the second model parameter β:
[0028] [α β] T =(B T B) -1 B T Y
[0029] Step S2-1-6: Obtain the estimated values of the elements in the first-order accumulated sequence based on the obtained first model parameter α and second model parameter β:
[0030]
[0031] Where: is the estimated value of the first-order accumulated value of the input variable of the (k + 1)-th original data point;
[0032] Step S2-1-7: Based on the first model parameter α and second model parameter β, obtain the estimated values of the elements in the original sequence as the GM model prediction values:
[0033]
[0034] Where: is the estimated value of the initial value of the input variable of the (k + 1)-th original data point.
[0035] The process of obtaining the GM residual model prediction values in step S2 specifically includes:
[0036] Step S2-2-1: Determine the residual sequence:
[0037]
[0038] Where: ε 0 is the residual sequence, is the element in the residual sequence:
[0039] Step S2-2-2: Perform a first-order accumulation on the residual sequence to obtain a first-order residual accumulated sequence;
[0040] Step S2-2-3: Calculate the mean of adjacent elements of the first-order residual accumulated sequence to obtain a mean residual sequence;
[0041] Step S2-2-4: Construct a third matrix and a fourth matrix;
[0042] Step S2-2-5: Based on the constructed third matrix and fourth matrix, calculate to obtain the first residual model parameter α ε and the second residual model parameter β ε ;
[0043] Step S2-2-6: Based on the obtained first residual model parameter α ε and the second residual model parameter β ε obtain the estimated values of the elements in the first-order residual accumulated sequence:
[0044]
[0045] Where: is the estimated value of the k-th element in the first-order residual accumulation sequence;
[0046] Step S2-2-7: Based on the first residual model parameter α ε and the second residual model parameter β ε , obtain the estimated values of the elements in the residual sequence;
[0047] Step S2-2-8: Use the estimated values of the elements in the residual sequence to obtain the GM residual model prediction value:
[0048]
[0049] Where: is the GM residual model prediction value, is the estimated value of the element in the residual sequence.
[0050] The said step S3 includes:
[0051] Step S3-1: Based on a pre-configured ratio, divide the amplified data points into a training sample set and a test sample set;
[0052] Step S3-2: Use each amplified data point in the training sample set to train multiple candidate prediction models;
[0053] Step S3-3: Use each amplified data point in the test sample set to test each trained candidate prediction model;
[0054] Step S3-3: Select the candidate prediction model with the best test index as the selected prediction model.
[0055] The said test index includes goodness of fit, mean square error, root mean square error, relative root mean square error, rank sum ratio, mean absolute error, mean absolute percentage error, and variance ratio.
[0056] The said candidate prediction models are BP model, RBF model, and ELM model.
[0057] A concrete compressive strength prediction device based on GM model and machine learning model, including a memory, a processor, and a program stored in the said memory. When the processor executes the program, it implements the method as described above.
[0058] A storage medium, on which a program is stored. When the program is executed, it implements the method as described above.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] 1. By combining the GM model and the machine learning model, the input variables of the original data points are extended using the predicted values of the GM model and the predicted values of the GM residual model respectively. This enables the machine learning model to better learn the global and local features of the data, and thus have stronger generalization ability when facing new and unseen data. This multi-dimensional information input allows the machine learning model to more precisely adjust its prediction results, thereby significantly improving the prediction accuracy. In addition, the predicted values of the GM residual model can effectively capture the abnormal fluctuations in the data, further enhancing the model's ability to handle abnormal situations. By introducing the predicted values of the GM model and the predicted values of the GM residual model, the model can more flexibly adjust its prediction strategy to adapt to the production conditions and data characteristics of different manufacturers. This adaptability enables the model to more effectively handle various complex situations in practical applications.
[0061] 2. The input variables of the original data points include water-binder ratio, sand ratio, RCA replacement rate, iron tailings sand replacement rate, fly ash content, silica fume content, PVA volume fraction, BF volume fraction, SF volume fraction, Na2SO4 mass fraction, NaCl mass fraction, soaking time, drying time, and number of wet-dry cycles. The information in the credible literature corpus can be fully utilized to improve the prediction accuracy.
[0062] 3. In addition to the predicted values of the GM model, the predicted values of the GM residual model are also added to capture the local fluctuations and abnormal situations in the data. The predicted values of the GM model provide the macroscopic trend of the data, while the predicted values of the GM residual model supplement the microscopic details. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a schematic structural diagram of the present invention;
[0064] Figure 2 is a schematic diagram of the distribution of input variables and output variables in the embodiment, where (a) is X-1, (b) is X-2, (c) is X-3, (d) is X-4, (e) is X-5, (f) is X-6, (g) is X-7, (h) is X-8, (i) is X-9, (j) is X-10, (k) is X-11, (l) is X-12, (m) is X-13, (n) is X-14, (o) is X-15, (p) is X-16, (q) is Y;
[0065] Figure 3 is a schematic diagram of the relationship between input variables and output variables;
[0066] Figure 4 is a schematic diagram of the Pearson correlation coefficient between input variables and output variables;
[0067] Figure 5 is a schematic diagram of the grey correlation coefficient of input variables;
[0068] Figure 6 Bar charts of the GM(1,1) model and the GM(1,1) residual model;
[0069] Figure 7 Residual distributions of the GM(1,1) model and the GM(1,1) residual model, where (a) is the box plot of the residual distribution and (b) is the curve of the residual distribution;
[0070] Figure 8 Prediction performance metrics of three machine learning models under different data partition ratios, where (a) are the performance metrics of the BP model training set, (b) are the performance metrics of the BP model test set, (c) are the performance metrics of the RBF model training set, (d) are the performance metrics of the RBF model test set, (e) are the performance metrics of the ELM model training set, and (f) are the performance metrics of the ELM model test set;
[0071] Figure 9 Score rankings of three machine learning models under different data partition ratios, where (a) is the BP model, (b) is the RBF model, and (c) is the ELM model;
[0072] Figure 10 Prediction performance metrics of three machine learning models under the (85%:15%) data partition ratio, where (a) are the performance metrics of the training set and (b) are the performance metrics of the test set;
[0073] Figure 11 Score rankings of three machine learning models under the (85%:15%) data partition ratio;
[0074] Figure 12 Residual distributions of three machine learning models under the (85%:15%) data partition ratio, where (a) is the box plot of the residual distribution and (b) is the curve of the residual distribution;
[0075] Figure 13 Performance evaluation metrics of the fusion model, where (a) is R 2 , (b) is MSE, (c) is RMSE, (d) is RRMSE, (e) is RSR, (f) is MAE, (g) is MAPE, and (h) is VAF;
[0076] Figure 14 Score rankings of the fusion model;
[0077] Figure 15Distribution of the original values and predicted values of the fusion model, where (a) is the GM(1,1) model, (b) is the GM(1,1) residual model, (c) is the BP model, (d) is the BP-GM(1,1) model, (e) is the BP-GM(1,1) residual model, and (f) is the BP-GM(1,1)-GM(1,1) residual model;
[0078] Figure 16 Residual distribution of the fusion model, where (a) is the box plot of the residual distribution and (b) is the curve of the residual distribution;
[0079] Figure 17 Predicting the compressive strength of concrete in different literature data based on the fusion model. Detailed implementation manners
[0080] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0081] A method for predicting the compressive strength of concrete based on the GM model and the machine learning model, as Figure 1 shown, includes:
[0082] Step S1: Obtain a credible literature corpus, and extract an original data set based on the credible literature corpus, where the original data set includes multiple original data groups, each data group consists of multiple original data points, and each original data point includes multiple input variables and one output variable;
[0083] The output variable of the original data point is the compressive strength.
[0084] The input variables of the original data point include water-binder ratio, sand ratio, RCA replacement rate, iron tailings sand replacement rate, fly ash content, silica fume content, PVA volume fraction, BF volume fraction, SF volume fraction, Na2SO4 mass fraction, NaCl mass fraction, immersion time, drying time, and number of wet-dry cycles.
[0085] The credible literature corpus should specifically select real and effective literatures. In this embodiment, a total of 19 credible literature corpora are selected, and the specific literatures are as follows:
[0086] [1] Li W K, Xiang M S, Mi L, et al. Study on the resistance of mixed fiber recycled aggregate concrete to sulfate attack under dry-wet cycles[J]. Composites Science and Engineering, 2023, (10): 69-77.
[0087] [2] Lan L, Zhang K, He X, et al. Research on corrosion resistance of basalt fiber reinforced concrete under sulfate wet dry cycle[J]. Concrete, 2024, (6): 25-29.
[0088] [3] Zhong C H, Shi J N, Zhou J Z. Durability analysis of stain less steel fiber recycled concrete under sulfate-dry-wet cycle[J]. Water Resources and Hydropower Engineering, 2023, 54(4): 187-196.
[0089] [4] Liu P Q. Experimental study on the deterioration characteristics of conerete exposure to sulfate erosion under different drying-wetting conditions[J]. Zhengzhou University, 2021.
[0090] [5] Shao H J. Study on the influence of drying-wetting cycle on the physical and mechanical properties of concrete[J]. Northwest Agriculture and Forestry University, 2021.
[0091] [6] Zhang T H. Study on Deterioration of Concrete Performance under the Coupling of Sulfate attack and Dry-wet Cycle[J]. Xijing University, 2020.
[0092] [7] Shao L L. Research on the Transmission Law of SO42- in Concrete under High Water Pressure and Dry-Wet Cycling Mechanism[J]. China University of Mining and Technology, 2018.
[0093] [8] Tang Z Y. Research on the deterioration mechanism of basalt fiber recycled concrete under the action of dry wet cyclic corrosion of salt[J]. East China University of Technology, 2023.
[0094] [9] Liu D Q. Study on Durability of Recycled Concrete under the Effect of Sulfate Wet and Dry Cycle[J]. Xinjiang Agricultural University, 2018.
[0095]
[10] Liu J Y, Gong T Z, Li Q. Study on the corrosion performance of concrete under wetting-drying sulfate partial soaking attack[J]. New Building Materials, 2017, 44(4): 87-90.
[0096]
[11] Jia W L,Zhang Q,Li H.Effect of Sulfate Dry-Wet Cycles on thePerformance of Recycled Aggregate Concrete[J].Bulletin of the Chinese CeramicSociety,2016,35(12):3981-3986.
[0097]
[12] Gan L,Wu J,Shen Z Z,et al.Deterioration law of basalt fiberreinforced concrete under sulfate attack and dry-wet cycle[J].China CivilEngineering Journal,2021,54(11):37-46.
[0098]
[13] Zhou M R,Luo X B,Lu C G,et al.Experimental study of concretedurability under the action of sulfate and dry-wet circulation[J].Concrete,2017,(9):15-19.
[0099]
[14] Liu D Q.Durability and Life Prediction of Recycled Concrete underSulfate Attack and Wet-dry Cycle[J].Xinjiang Agricultural University,2018.
[0100]
[15] Wang L,Chen C,Liu R,et al.Chloride Corrosion Process of Concretewith Different Water–Binder Ratios under Variable Temperature Drying-WettingCycles[J].Materials,2024,17(10):2263.
[0101]
[16] Li L, Shi J, Kou J. Experimental study on mechanical properties of high-ductility concrete against combined sulfate attack and dry-wet cycles[J]. Materials, 2021, 14(14): 4035.
[0102]
[17] Wang H, Chen Y, Wang H. Under Sulfate Dry-Wet Cycling: Exploring the Symmetry of the Mechanical Performance Trend and Grey Prediction of Lightweight Aggregate Concrete with Silica Powder Content[J]. Symmetry, 2024, 16(3): 275.
[0103]
[18] Liu X, Xu H, Li B, et al. Investigation of the Mechanical Properties of Iron Tailings Concrete Subjected to Dry-Wet Cycle and Negative Temperature[J]. Materials, 2023, 16(13): 4602.
[0104]
[19] Liu S, Han F, Zheng S, et al. Study on Mechanical Properties and Erosion Resistance of Self-Compacting Concrete with Different Replacement Rates of Recycled Coarse Aggregates under Dry and Wet Cycles[J]. Applied Sciences, 2023, 13(19): 11101.
[0105] Specifically, for reference [1], reference [1] Figure 2 contains a total of 22 original data groups, and the length of each original data group is 7. Similarly, for other references, similar original data groups can be obtained.
[0106] Step S2: Process each original data group. Based on all the original data points in each original data group, calculate the GM model prediction value and the GM residual model prediction value for each original data point respectively, and use the GM model prediction value and the GM residual model prediction value as new input variables to update the original data point to an amplified data point;
[0107] In this step, the GM model prediction value is obtained from the GM(1,1) model. The specific process includes:
[0108] Step S2-1-1: Obtain the original sequence based on the input variables of all the original data points in the original data:
[0109]
[0110] where: X 0 is the original sequence, is the initial value of the input variable of the first original data point, m is the length of the sequence, is the initial value of the input variable of the second original data point, is the initial value of the input variable of the m-th original data point;
[0111] Step S2-1-2: Perform a first-order accumulation on the original sequence to obtain a first-order accumulated sequence:
[0112]
[0113] where: X 1 is the first-order accumulated sequence, is the first-order accumulated value of the input variable of the first original data point, is the first-order accumulated value of the input variable of the second original data point, is the first-order accumulated value of the input variable of the m-th original data point, is the first-order accumulated value of the input variable of the k-th original data point, is the first-order accumulated value of the input variable of the i-th original data point;
[0114] Step S2-1-3: Calculate the mean value of adjacent elements of the first-order accumulated sequence to obtain a mean value sequence:
[0115]
[0116] where: Z 1 is the mean value sequence, is the sum of the first-order accumulated values of the input variables of the second original data point and the previous original data point, is the sum of the first-order accumulated values of the input variables of the third original data point and the previous original data point, is the sum of the first-order cumulative value of the input variables of the m-th original data point and the previous original data point, is the sum of the first-order cumulative value of the input variables of the (k + 1)-th original data point and the previous original data point, is the first-order cumulative value of the input variable of the (k + 1)-th original data point;
[0117] Step S2-1-4: Construct the first matrix B and the second matrix Y:
[0118]
[0119] Step S2-1-5: Based on the constructed first matrix and second matrix, calculate the first model parameter α and the second model parameter β:
[0120] [αβ] T =(B T B) -1 B T Y
[0121] Step S2-1-6: Based on the obtained first model parameter α and second model parameter β, obtain the estimated values of the elements in the first-order cumulative sequence:
[0122]
[0123] where: is the estimated value of the first-order cumulative value of the input variable of the (k + 1)-th original data point;
[0124] Step S2-1-7: Based on the first model parameter α and the second model parameter β, obtain the estimated values of the elements in the original sequence as the GM model prediction values:
[0125]
[0126] where: is the estimated value of the initial value of the input variable of the (k + 1)-th original data point.
[0127] And, the GM residual model prediction value in this step is obtained from the GM(1,1) residual model, and the specific process includes:
[0128] Step S2-2-1: Determine the residual sequence:
[0129]
[0130] where: ε 0 is the residual sequence, is the element in the residual sequence:
[0131] Step S2-2-2: Perform a first-order accumulation on the residual sequence to obtain a first-order residual accumulation sequence;
[0132] Step S2-2-3: Calculate the mean of adjacent elements of the first-order residual cumulative sequence to obtain the mean residual sequence;
[0133] Step S2-2-4: Construct the third matrix and the fourth matrix;
[0134] Step S2-2-5: Based on the constructed third matrix and fourth matrix, calculate and obtain the first residual model parameter α ε and the second residual model parameter β ε ;
[0135] Step S2-2-6: Based on the obtained first residual model parameter α ε and the second residual model parameter β ε obtain the estimated values of each element in the first-order residual cumulative sequence:
[0136]
[0137] where: is the estimated value of the k-th element in the first-order residual cumulative sequence;
[0138] Step S2-2-7: Based on the first residual model parameter α ε and the second residual model parameter β ε , obtain the estimated values of the elements in the residual sequence;
[0139] Step S2-2-8: Use the estimated values of the elements in the residual sequence to obtain the GM residual model prediction value:
[0140]
[0141] where: is the GM residual model prediction value, is the estimated value of the element in the residual sequence.
[0142] The final variables are shown in Table 1:
[0143] Table 1
[0144]
[0145] Step S3: Construct a sample set based on the amplified data points to train and obtain a prediction model, including:
[0146] Step S3-1: Divide the amplified data points into a training sample set and a test sample set based on a pre-configured ratio. Specifically, in this embodiment, there are 3 pre-configured ratios, which are: (90%:10%), (85%:15%), (80%:20%)
[0147] Step S3-2: Train multiple candidate prediction models using each amplified data point in the training sample set;
[0148] Step S3-3: Test each trained candidate prediction model using each amplified data point in the test sample set;
[0149] Step S3-3: Select the candidate prediction model with the best test index as the selected prediction model.
[0150] Among them, in this embodiment, the test indexes include goodness of fit, mean square error, root mean square error, relative root mean square error, rank sum ratio, mean absolute error, mean absolute percentage error, and variance ratio, and the candidate prediction models are BP model, RBF model, and ELM model.
[0151] Step S4: Predict the compressive strength of the concrete to be measured based on the prediction model. Specifically, input the input parameters of the concrete to be measured into the selected prediction model to obtain the compressive strength output by the prediction model.
[0152] In this embodiment, a database is established by collecting a total of 982 amplified data points from 19 literatures. These data points cover 16 input factors and 1 output factor. Then, data visualization is performed, including data preprocessing and grey relational analysis, to display variable distribution and correlation. Subsequently, through model evaluation and data segmentation, 8 performance indexes are determined, and the database is randomly divided into a training set and a test set according to the ratios of (90%:10%), (85%:15%), and (80%:20%). After that, based on the further analysis of grey system theory, the data distribution and performance evaluation indexes are revealed, and GM models are established, including GM(1,1) model and GM(1,1) residual model. After that, the database is segmented according to 3 ratios, BP model, RBF model, and ELM model are established, and performance evaluation is carried out. Through database segmentation and performance evaluation, fusion models such as BP-GM(1,1) model, BP-GM(1,1) residual model, and BP-GM(1,1)-GM(1,1) residual model are established. Finally, by selecting the optimal model and collecting literature data, the strength and service life of the concrete are predicted. This framework provides a comprehensive and accurate method for concrete durability evaluation and service life prediction by systematically integrating various analysis methods.
[0153] The concept of machine learning is "the ability of a computer to improve its performance through experience accumulation". Machine learning is an important branch of artificial intelligence, and its core goal is to enable a computer to learn and improve through data without explicit programming. Since the 21st century, the rapid development of deep learning has greatly promoted the progress of machine learning, enabling it to achieve major breakthroughs in many fields such as image recognition, natural language processing, and autonomous driving. Currently, machine learning has become an important driving force for technological innovation, and its application fields continue to expand.
[0154] Using different evaluation metrics to measure the performance of a model is an important means to evaluate its predictive ability. The purpose of using multiple evaluation criteria is to overcome the inherent limitations of a single metric, so as to be able to evaluate the performance of the model more comprehensively. In the evaluation process of building a machine learning model, multiple evaluation metrics are usually used based on statistical methods to measure the accuracy of the machine learning model on the training set and the test set. The error metrics used in this study include goodness of fit (R 2 ), mean squared error (MSE), root mean squared error (RMSE), relative root mean squared error (RRMSE), rank sum ratio (RSR), mean absolute error (MAE), mean absolute percentage error (MAPE), and variance ratio (VAF). Among these metrics, the closer R 2 , VAF is to 1 and the smaller MSE, RMSE, RRMSE, RSR, MAE, MAPE are, the higher the degree of fitting of the model to the data and the better the predictive performance of the model.
[0155] The grey relational analysis method is used to calculate the significant influence of each input variable on the output variable. The calculation steps are as follows:
[0156] (1) Take the output variable as the system characteristic data sequence X 0 , and the input features as the related factor data sequences X i .
[0157] X 0 = {x 0 (1)x 0 (2)…x 0 (982)}
[0158] X i = {x i (1)x i (2)…x i (982)},(i = 1,2,…,16)
[0159] (2) To prevent the model from showing serious deviations, perform mean normalization on X 0 and X i .
[0160]
[0161] Among them, y 0 (k) and y i (k) are the results of averaging X 0 and X i respectively.
[0162] (3) Calculate the grey correlation coefficient δ 0 between X i and X i at point k.
[0163]
[0164] Δ i (k) = |y 0 (k) - y i (k)|, k = 1, 2, …, 982
[0165] Among them, φ is the resolution coefficient, and in this application, φ = 0.5 is taken.
[0166] (4) Calculate the grey correlation coefficient.
[0167]
[0168] When the compressive strength of the concrete is reduced to 75% of the undamaged state, it is considered that the concrete is damaged. When using the GM(1,1) residual model to predict the compressive strength and service life of the concrete, the Markov model needs to be used for symbol correction.
[0169] The Markov model statistically analyzes the positive and negative states of the residual sequence. It is defined that when the residual is positive, it is in state 1, and when the residual is negative, it is in state 2. The transition probability matrix is as follows.
[0170]
[0171] Among them, p ij is the probability that the residual transfers from state i to state j. e ij is the number of times the state transfers from state i to state j. e i is the number of times state i appears.
[0172] Name the predicted values obtained from the GM(1,1) model and the residual GM(1,1) model as X-15 and X-16. Based on 16 input variables and 1 output variable, establish the data volume for the evaluation of concrete durability and life prediction based on grey system theory and machine learning algorithms. The database collected a total of 982 data from 19 scientific papers. It can be seen from Table 2 the distribution range of the data in the 19 scientific papers. As can be seen in Table 2, the input variables are water-binder ratio, sand ratio, RCA replacement rate, iron tailings sand replacement rate, fly ash content, silica fume content, PVA volume fraction, BF volume fraction, SF volume fraction, Na2SO4 mass fraction, NaCl mass fraction, immersion time, drying time, wet-dry cycle times, GM(1,1) model predicted value, GM(1,1) residual model predicted value. The output variable is compressive strength.
[0173]
[0174] As can be seen in Table 2, there are significant differences in the variable ranges of the concrete input and output variables in different literatures. These differences may stem from various factors. Figure 2 Shows the distribution of each input variable (X-1 to X-16) and the output variable (Y). Figure 2 Each subgraph in is a histogram showing the frequency distribution of the data and superimposed with a fitted probability density function. From Figure 2 it can be seen that the distributions of the input and output variables are normal distributions. Figure 2 Shows that the distributions of X-1, X-2, X-14, X-15, X-16, and Y are relatively uniform, and X-1, X-2, X-14, X-15, X-16, and Y are evenly distributed in 0.3 - 0.5, 30% - 45%, 0 - 65times, 30 - 55MPa, 30 - 55MPa, 30 - 55MPa. The distributions of X-3, X-4, X-5, X-6, X-7, X-8, X-9, X-10, X-11, X-12, and X-13 are relatively concentrated, and X-3, X-4, X-5, X-6, X-7, X-8, X-9, X-10, X-11, X-12, and X-13 are roughly distributed in 0% - 30%, 0%, 0% - 30%, 0%, 0%, 0% - 0.2%, 0%, 0% - 6%, 11%, 10h - 20h, 0h - 10h.
[0175] Table 3 provides the minimum value (D min ), maximum value (D max ), first quartile (D 1 ), median (D 2 ), third quartile (D 3) Average and Standard Deviation (Std). Table 3 shows the numerical distribution characteristics of the input and output variables. Among them, D min and D max show the overall range of these variables. For example, the soaking time of X-12 varies from 9 to 72, indicating the diversity of experimental conditions. D 1 、D 2 、D 3 provide the ordered statistics of the data, explaining the central tendency and dispersion degree of the data. For example, the median of X-1 is 0.45, indicating that most data are concentrated at a lower water-cement ratio; Average reflects the central position of the variable. For example, the average value of X-2 is 40.817, indicating the general level of sand ratio in the study. Std measures the volatility of the data. The standard deviation of X-3 is 24.639, showing significant variability of the data.
[0176] Table 3
[0177]
[0178] Figure 3 shows the relationship between each input variable and the output variable, and superimposes the trend line. Figure 3 shows that as the water-cement ratio increases, the compressive strength shows a downward trend. The relationship between the sand ratio and the compressive strength is relatively complex. Generally, as the sand ratio increases, the compressive strength shows a downward trend. The RCA replacement rate, iron tailings replacement rate, fly ash content, silica fume content, PVA fiber content, Na 2 SO 4 mass fraction, number of dry-wet cycles and the compressive strength have an unclear trend, and the data points are relatively scattered. The BF fiber content shows a weak positive correlation trend with the compressive strength. The SF fiber content shows a weak negative correlation trend with the compressive strength. The NaCl mass fraction shows a weak negative correlation trend with the compressive strength. The relationship between the soaking time, drying time and the compressive strength generally shows a trend of increasing first and then decreasing. There is a strong positive correlation between the predicted values of the GM(1,1) model, the predicted values of the GM(1,1) residual model and the compressive strength.
[0179] Figure 4 shows the Pearson correlation coefficients between each input variable and the output variable. Figure 4 Uses colors and numbers to represent the strength and direction of the correlation. Among them, yellow indicates a positive correlation, blue indicates a negative correlation, and the darker the color, the stronger the correlation. A positive number indicates a positive correlation, a negative number indicates a negative correlation, and the closer the number is to 1, the stronger the correlation. Figure 4It shows that there is a significant negative correlation between X-1 and Y, indicating that a lower water-cement ratio helps to improve the compressive strength. There is a strong positive correlation between X-15 and X-16 and Y, showing that the GM(1,1) residual model and the GM(1,1) model have high accuracy in predicting the compressive strength. In addition, the strong positive correlation between X-4 and X-5 indicates the synergistic effect of these two materials in the concrete mix ratio.
[0180] The magnitude of the grey correlation coefficient reflects the similarity degree between the system characteristic data sequence and the related factor data sequence. When the correlation coefficient is large, it indicates a higher similarity degree and a stronger correlation degree between the system characteristic data sequence and the related factor data sequence, which means that this factor has a more significant impact on the target variable.
[0181] Calculate the significant influence of each input variable on the output variable based on the grey correlation analysis method, and the calculation results are shown in Figure 5 . From Figure 5 It can be seen that the grey correlation coefficients of X-4 and X-9 are the highest, 0.951 and 0.96 respectively. This indicates that the relationships between these two variables and the output variable are the closest, and their impacts on the output variable are the most significant. The grey correlation coefficients of X-11, X-8 and X-16 are also relatively high, 0.955, 0.931 and 0.88 respectively. The grey correlation coefficients of X-15, X-13, X-12, X-10, X-7 and X-6 are at a medium level, indicating that these variables have an impact on the output variable but not significantly. The grey correlation coefficients of X-5, X-3, X-2 and X-1 are relatively low, indicating that these variables have a weak impact on the output variable.
[0182] Randomly divide the database into a training set and a test set according to the ratios of (90%:10%), (85%:15%), and (80%:20%). The model development process will make full use of the data in the training set to construct and optimize the concrete strength prediction model. During the model training process, cross-validation techniques are used to evaluate the performance of different model configurations, thus avoiding overfitting and improving the generalization ability of the model. In the model validation stage, the developed prediction model is validated using the test set. By comparing the prediction results of the model on the test set with the actual observed values, various performance indicators can be calculated to quantify the prediction accuracy of the prediction model.
[0183] Calculate the performance of the GM(1,1) model and the GM(1,1) residual model on multiple evaluation indicators. Plot the results as Figure 6 . From Figure 6 It can be seen that among the 8 evaluation indicators, the GM(1,1) residual model only has worse prediction performance than the GM(1,1) model in terms of the RSR index, while its prediction performance is better than that of the GM(1,1) model in the remaining 7 indicators.
[0184] Figure 7 (a) of Figure 7 (a) shows the box plots of the residual distributions of the GM(1,1) model and the GM(1,1) residual model. It can be seen from
[0185] Figure 7 (a) that the residual distributions of the GM(1,1) model and the GM(1,1) residual model are normal distributions. The medians of the GM(1,1) model and the GM(1,1) residual model are close to 0, indicating that the deviations between the predicted values and the actual values of the two models are balanced on the whole. The interquartile range of 25%-75% shows the concentrated distribution range of the residuals, and most of the residuals of the GM(1,1) residual model are within this interval. The whisker lines of the GM(1,1) model and the GM(1,1) residual model do not seem to be exactly parallel to the zero line, indicating that there is a certain skewness in the residual distribution. The range covered by the whisker line of the GM(1,1) residual model is relatively narrow, suggesting that the residual distribution of this model is more concentrated. It reflects that the deviation between the predicted value and the actual observed value of this model is small, and the prediction stability and accuracy are high.
[0185] Figure 7 (b) shows the residual distribution curves of the GM(1,1) model and the GM(1,1) residual model. It can be seen from Figure 7 (b) that the residual distributions of the two models are normal distributions, and the residual distribution curve of the GM(1,1) residual model shows a certain skew to the upper right. The residual distribution of the GM(1,1) residual model is more concentrated and has a higher degree of fit with the normal distribution curve. The average value of the residuals of the GM(1,1) model is -0.002, close to 0, indicating that there is no systematic deviation in the GM(1,1) model during prediction. The average value of the residuals of the GM(1,1) model is 0.267, verifying that the GM(1,1) model shows a certain skew to the upper right during prediction. The standard deviation of the GM(1,1) model is 3.162, while the standard deviation of the GM(1,1) residual model is 1.968. The latter's residual distribution is more concentrated, indicating that the GM(1,1) residual model has higher prediction accuracy. The range of the GM(1,1) model is from -15.641 to 14.162, while the range of the GM(1,1) residual model is from -6.356 to 14.162. The latter's residual range is smaller, indicating that the GM(1,1) residual model has smaller prediction errors.
[0186] Overall, compared with the GM(1,1) model, the GM(1,1) residual model performs better in terms of error control, goodness of fit, and variance. By correcting the residuals of the original GM(1,1) model, the GM(1,1) residual model can capture the subtle changes and potential laws in the data more accurately, thus improving the accuracy and reliability of the prediction results. Therefore, it is more reasonable to choose the GM(1,1) residual model for prediction.
[0187] The database is randomly divided into a training set and a test set according to the ratios of (90%:10%), (85%:15%), and (80%:20%). The BP algorithm, RBF algorithm, and ELM algorithm are used to establish a concrete durability evaluation model. The model is tested multiple times, and the highest value is taken as the final calculation result. Figure 8 are the prediction performance indicators of three machine learning models (BP, RBF, ELM) under different data division ratios. Figure 8 shows that in all models, the R 2 value is generally higher in the training set than in the test set, indicating that the model has a better fitting effect in the training set but performs mediocrely in the test set. The indicators such as MSE, RMSE, RRMSE, RSR, MAE, and MAPE are generally higher under the data division ratio of (80%:20%), indicating that the model has poor performance under this data division. VAF varies greatly among different models and data division ratios, but generally performs better under the data division ratio of (90%:10%). The performance of the model in the training set is generally better than that in the test set, and the data division ratio has a significant impact on the model performance.
[0188] Figure 9 Shows the score rankings of three machine learning models (BP model, RBF model, and ELM model) under different data division ratios. From Figure 9 (a), it can be seen that the score rankings of the BP model are tied for first under the data division ratios of (90%:10%) and (85%:15%), and the score ranking is third under the data division ratio of (80%:20%). From Figure 9 (b), it can be seen that the score ranking of the RBF model is first under the data division ratio of (85%:15%), second under the data division ratio of (90%:10%), and third under the data division ratio of (80%:20%). From Figure 9 (c), it can be seen that the score ranking of the ELM model follows the same pattern as the BP model. For the BP model and the ELM model, as the data division ratio changes from (90%:10%) to (80%:20%), the score rankings first remain stable and then decline. This indicates that with more training data, the prediction performance of the model improves. For the RBF model, the score ranking first decreases and then increases. This indicates that the RBF model has the highest prediction performance under the data division ratio of (85%:15%).
[0189] Figure 10 Shows the prediction performance indicators of three machine learning models (BP model, RBF model, and ELM model) under the data division ratio of (85%:15%). Figure 11Shows the score rankings of three machine learning models under the data division ratio of (85%:15%). From Figure 10 and Figure 11 it can be seen that under the data division ratio of (85%:15%), the performance indicators of the BP model on the training set and the test set are generally better than those of the RBF model and the ELM model. Especially in terms of R 2 , MSE, RMSE, RSR, MAE and VAF, it performs the best. The BP model shows the best fitting and prediction ability under this data division ratio. The ELM model performs better in terms of MAE and RRMSE on the test set, indicating that it has certain advantages in terms of relative error and mean absolute error. The RBF model performs better in terms of MAPE, indicating that it has certain advantages in terms of mean absolute percentage error. Overall, the performance of the BP model is better, but in some specific indicators, the RBF model and the ELM model also have certain advantages.
[0190] Figure 12 Shows the residual distributions of three machine learning models under the data division ratio of (85%:15%). Figure 12 (a) is the box plot of the residual distribution. Figure 12 (b) is the curve of the residual distribution. Figure 12 (a) shows that the residual distributions of the three models are all concentrated around 0, but the residual distributions of the BP and RBF models are more concentrated, while the residual distribution of the ELM model is more dispersed. This indicates that the prediction error of the ELM model is larger. In addition, the residual distributions of the BP and RBF models are more symmetric, while the residual distribution of the ELM model is slightly skewed. Figure 12 (b) The bar chart shows the frequency distribution of the residuals of each model, and the dotted line represents the normal distribution curve. It can be seen that the residual distribution of the BP model is more concentrated and close to the normal distribution. This indicates that the prediction accuracy of the BP model is higher. The residual distributions of the RBF and ELM models are more dispersed and deviate from the normal distribution. Especially for the ELM model, the residual distribution is the widest, indicating that its prediction error is larger. The BP model performs the best in terms of mean, standard deviation and maximum value, indicating that its prediction error is the smallest and the most stable. Although the mean of the RBF model is close to 0, its standard deviation and maximum value are higher, indicating that its prediction error is larger and it is prone to larger prediction errors in some cases. The mean, standard deviation and minimum value of the ELM model all show that its prediction error is larger and unstable.
[0191] Establish a concrete durability evaluation fusion model based on the GM model and machine learning models according to the data division ratio of (85%:15%). Figure 13 Shows the performance of different models under multiple performance evaluation indicators, including R 2 , MSE, RMSE, RRMSE, RSR, MAE, MAPE and VAF. Figure 13Each sub - figure corresponds to an evaluation index, and the results of the training set (red) and the test set (blue) are represented by different three - dimensional graphs respectively. From Figure 13 it can be seen that the R 2 of all models on the training set is relatively high, close to 1, indicating that the models fit well on the training data. On the test set, the R 2 of the models is slightly insufficient, which reflects that their fitting effect on the data they have not encountered is poor. This is related to the over - fitting phenomenon of the models. The MSE, RMSE, RRMSE, RSR, MAE, and MAPE of the training set are generally lower than those of the test set, and the VAF of the training set is higher than that of the test set, which indicates that the models perform better on the training set than on the test set. The R 2 , MSE, RMSE, RRMSE, RSR, and VAF of the BP - GM(1,1) model on the two data sets all show good prediction performance. Generally speaking, when dealing with this data set, the BP - GM(1,1) model has better performance than using the BP or GM(1,1) model alone.
[0192] Figure 14 shows the score rankings of different fusion models. From Figure 14 it can be seen that the BP - GM(1,1) residual model performs best in terms of scores and rankings, with the lowest score (1.688) and the highest ranking (rank 1), indicating its optimal performance. The score of the BP - GM(1,1) - GM(1,1) residual model is the second lowest (2), and it ranks 2nd, with good performance. The score rankings of the GM(1,1) residual model, BP - GM(1,1) model, and GM(1,1) model are ranked 3rd, 4th, and 5th. The BP model has the highest score (6) and ranks 6th, with the worst performance. In the process of model fusion, combining the GM(1,1) residual model with the BP model can significantly improve the model performance. The BP - GM(1,1) residual model has advantages in terms of prediction accuracy, stability, or generalization ability. It shows that the performance of the GM model alone is average, but when its residual information is combined with other models, it can produce a complementary effect, thus improving the performance of the overall model.
[0193] Figure 15 shows the distributions of the original values and predicted values of different fusion models. Figure 15 Each sub - figure in it corresponds to a different model respectively. Figure 15 In it, the training data is represented by green dots and the test data is represented by blue dots. From Figure 15 it can be seen that the data points of all models are roughly distributed in the range of ±10% - ±20%, and their coefficient of determination r 2 is 0.846 - 0.983. The BP - GM(1,1) residual model and the BP - GM(1,1) - GM(1,1) residual model perform best in terms of goodness of fit, and their coefficient of determination r2 They reached 0.977 and 0.983 respectively, indicating that these two models can explain most of the data variability and have high prediction accuracy. Among them, the r of the BP model 2 is relatively low. The determination coefficients of the training set and test set of this model are 0.861 and 0.846 respectively, indicating that this model has certain limitations in capturing the complex relationships in the data. There are certain deviations between the predicted values and the original values of all models, but most data points are concentrated within the error range of ±10% - ±20%, which is usually an acceptable error range in engineering applications. The distribution trends of the training data and test data are basically the same, indicating that the model not only performs well on the training set, but also has good generalization ability on unseen data.
[0194] Figure 16 The residual distribution of the fusion model is shown. Figure 16 (a) shows the distribution of residuals of different models. Figure 16 (a) shows that the residual distributions of the GM(1,1) model and the GM(1,1) residual model are relatively symmetric, but there are a certain number of outliers. The residual distribution of the BP model is relatively symmetric, but there are more extreme values, indicating that the prediction of the BP model has a relatively large prediction error compared with other models. The residual distributions of the BP-GM(1,1) residual model and the BP-GM(1,1)-GM(1,1) residual model are relatively symmetric and concentrated, and there are fewer outliers, indicating that these two models perform well in dealing with residuals. Figure 16 (b) shows the distribution curves of residuals of different models. Figure 16 (b) shows that the averages of all models are close to 0, indicating that the deviations between the predicted values and the actual values of the models are generally balanced and there is no obvious systematic deviation. The standard deviation of the BP model is the largest (4.433), indicating that the volatility of its prediction results is the largest and the stability is relatively poor. The standard deviation of the BP-GM(1,1)-GM(1,1) residual model is the smallest (1.663), indicating that the volatility of its prediction results is the smallest and the stability is the best. The residual distributions of all models are roughly symmetric, and the peaks of most distribution curves are concentrated at the position where the residual is 0, indicating that the errors between the predicted values and the actual values of the models are close to 0 in many cases. The residual distributions of the BP-GM(1,1) residual and the BP-GM(1,1)-GM(1,1) residual model have a high degree of coincidence with the normal distribution curve. On the contrary, the residual distribution of the BP model is relatively dispersed and has a poor fitting degree with the normal distribution.
[0195] By sorting out the literature, the durability experimental values under different test conditions are obtained, as shown in Table 4. By inputting the experimental data into the fusion model, the compressive strength of concrete with different mix ratios under different erosion environments is quantitatively evaluated.
[0196] Figure 17 shows the change in the compressive strength of concrete after different numbers of wet-dry cycles in three different literature datasets (Li
[14] , Zhong
[16] and Shao
[20] ). Figure 17 (a) shows the change in the compressive strength of concrete with the number of wet-dry cycles in the dataset of Li
[14] . Figure 17 (a) shows that the original value gradually increases in the first 45 cycles, starts to decline after reaching the peak. The BP-GM(1,1) residual model is relatively close to the original value in the 45 - 90 cycles and has a larger gap with the original value in the 0 - 30 cycles. The BP-GM(1,1) residual model shows a good fitting effect throughout the wet-dry cycle process. Figure 17 (b) shows the change in the compressive strength of concrete with the number of wet-dry cycles in the dataset of Zhong
[16] . Figure 17 (b) shows that the peak of the curve of the original value changing with the number of wet-dry cycles appears at the 15th wet-dry cycle. Figure 17 (c) shows the change in the compressive strength of concrete with the number of wet-dry cycles in the dataset of Shao
[20] . Figure 17 (c) shows that the peak of the curve of the original value changing with the number of wet-dry cycles appears at the 20th wet-dry cycle. The BP-GM(1,1) residual model is relatively close to the original value throughout the wet-dry cycle process, showing a good fitting effect.
[0197] Table 4
[0198]
[0199] Based on the experimental data in Table 4 and Figure 17 , taking the reduction of the compressive strength of concrete to 75% as the failure index, the durability life prediction of concrete is carried out. The durability life prediction results of concrete are shown in Table 5. The prediction results of Li
[14]
[0200] show that the predicted values of the GM(1,1) residual model and the BP-GM(1,1) residual model are 105 times, and the predicted value of the BP-GM(1,1)-GM(1,1) residual model is less than 105 times. The prediction of the BP-GM(1,1)-GM(1,1) residual model is more conservative. The original value, the prediction results of the BP-GM(1,1) residual model, and the prediction results of the BP-GM(1,1)-GM(1,1) residual model of Zhong
[16] are less than 105 times. The original value, the prediction results of the BP-GM(1,1) residual model, and the prediction results of the BP-GM(1,1)-GM(1,1) residual model of Shao
[20] The prediction results show that the predicted value of the GM(1,1) residual model is less than 70 times, the predicted value of the BP-GM(1,1) residual model is greater than 70 times, and the predicted value of the BP-GM(1,1)-GM(1,1) residual model is approximately 80 times.
[0201] Table 5
[0202]
[0203] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A method for predicting concrete compressive strength based on GM model and machine learning model, characterized in that: include: Step S1: obtaining a credible document corpus, and extracting an original data set based on the credible document corpus, wherein the original data set includes a plurality of original data groups, each data group consists of a plurality of original data points, and each original data point includes a plurality of input variables and an output variable; Step S2: Processing each original data group, based on all the original data points in each original data group, respectively calculating the GM model prediction value and the GM residual model prediction value of each original data point, and using the GM model prediction value and the GM residual model prediction value as new input variables to update the original data point to an augmented data point; Step S3: constructing a sample set based on the amplified data points to train a prediction model; Step S4: predicting the compressive strength of the concrete to be tested based on the prediction model.
2. A method for predicting concrete compressive strength based on GM model and machine learning model according to claim 1, characterized in that: The output variable of the raw data points is compressive strength.
3. A method for predicting concrete compressive strength based on GM model and machine learning model according to claim 2, characterized in that: The input variables of the original data points include water-binder ratio, sand ratio, RCA replacement rate, iron tailings sand replacement rate, fly ash content, silica fume content, PVA volume fraction, BF volume fraction, SF volume fraction, Na2SO4 mass fraction, NaCl mass fraction, soaking time, drying time and number of dry-wet cycles.
4. The method for predicting concrete compressive strength based on GM model and machine learning model according to claim 2, characterized in that: The process of obtaining the GM model prediction value in step S2 specifically includes: Step S2-1-1: Get the original sequence based on the input variables of all the original data points in the original data: Where: X 0 is the original sequence, is the initial value of the input variable of the first original data point, m is the length of the sequence, is the initial value of the input variable for the second original data point, is the initial value of the input variable of the mth original data point; Step S2-1-2: Accumulate the original sequence once to obtain a first-order cumulative sequence: Where: X 1 is a first-order cumulative sequence, is the first-order cumulative value of the input variable of the first original data point, is the first-order accumulated value of the input variable of the second original data point, is the first-order cumulative value of the input variable of the mth original data point, is the first-order cumulative value of the input variable of the kth original data point, is the first-order cumulative value of the input variable of the i-th original data point; Step S2-1-3: Calculate the mean of adjacent elements of the first-order cumulative sequence to obtain the mean sequence: Where: Z 1 is the mean sequence, is the sum of the first-order accumulated values of the input variables of the second raw data point and the previous raw data point, is the sum of the first-order accumulated values of the input variables of the third raw data point and the previous raw data point, is the sum of the first-order accumulated values of the input variables of the mth original data point and the previous original data point, is the sum of the first-order accumulated values of the input variables of the k+1th original data point and the previous original data point, is the first-order cumulative value of the input variable of the k+1th original data point; Step S2-1-4: Construct the first matrix B and the second matrix Y: Step S2-1-5: Based on the constructed first matrix and second matrix, the first model parameter α and the second model parameter β are calculated: [αβ] T =(B T B) -1 B T Y Step S2-1-6: Obtain an estimated value of each element in the first-order cumulative sequence based on the obtained first model parameter α and second model parameter β: in: is the estimated value of the first-order cumulative value of the input variable of the k+1th original data point; Step S2-1-7: Based on the first model parameter α and the second model parameter β, obtain the estimated value of the element in the original sequence as the GM model prediction value: in: is the estimate of the initial value of the input variable for the k+1th original data point.
5. The method for predicting concrete compressive strength based on GM model and machine learning model according to claim 4, characterized in that: The process of obtaining the GM residual model prediction value in step S2 specifically includes: Step S2-2-1: Determine the residual sequence: Where: 0 is the residual sequence, are the elements in the residual sequence: Step S2-2-2: Accumulate the residual sequence once to obtain a first-order residual accumulation sequence; Step S2-2-3: Calculate the mean of adjacent elements of the first-order residual accumulation sequence to obtain a mean residual sequence; Step S2-2-4: construct a third matrix and a fourth matrix; Step S2-2-5: Based on the constructed third matrix and fourth matrix, calculate the first residual model parameter α ε and the second residual model parameter β ε ; Step S2-2-6: Based on the obtained first residual model parameter α ε and the second residual model parameter β ε Get the estimated value of each element in the first-order residual accumulation sequence: in: is the estimated value of the kth element in the first-order residual accumulation sequence; Step S2-2-7: Based on the first residual model parameter α ε and the second residual model parameter β ε , get the estimated value of the elements in the residual sequence; Step S2-2-8: Use the estimated values of the elements in the residual sequence to obtain the GM residual model prediction value: in: is the predicted value of the GM residual model, is the estimated value of the elements in the residual sequence.
6. A method for predicting concrete compressive strength based on GM model and machine learning model according to claim 5, characterized in that: The step S3 comprises: Step S3-1: Divide the amplified data points into a training sample set and a test sample set based on a preconfigured ratio; Step S3-2: training multiple candidate prediction models using each amplified data point in the training sample set; Step S3-3: testing each trained candidate prediction model using each amplified data point in the test sample set; Step S3-3: Select the candidate prediction model with the best test index as the selected prediction model.
7. The method for predicting concrete compressive strength based on GM model and machine learning model according to claim 6, characterized in that: The test indicators include goodness of fit, mean square error, root mean square error, relative root mean square error, rank sum ratio, mean absolute error, mean absolute percentage error and variance ratio.
8. The method for predicting concrete compressive strength based on GM model and machine learning model according to claim 1, characterized in that: The candidate prediction models include BP model, RBF model and ELM model.
9. A concrete compressive strength prediction device based on a GM model and a machine learning model, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.