Deep ensemble forest regression modeling method for measuring concrete compressive strength

The online soft measurement of the compressive strength of concrete through deep integrated forest regression method solves the problem of difficult real-time measurement in the prior art and improves the optimization control capability of the production process.

CN111931948BActive Publication Date: 2025-06-06BEIJING UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010263130.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-07
Publication Date
2025-06-06
Estimated Expiration
2040-04-07

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time online measurement of concrete compressive strength, which makes it difficult to achieve optimized control of the concrete production process.

Method used

Deep integrated forest regression (DEFR) method is used to pre-process the high-dimensional features through the dimension reduction module, the input layer forest module trains multiple sub-forest models, the intermediate layer forest module trains layer by layer, and the output layer forest module performs final prediction, realizing online soft measurement of concrete compressive strength.

Benefits of technology

The online soft measurement accuracy of concrete compressive strength is improved, real-time optimization control of the concrete production process is achieved, and the problem of low prediction accuracy in traditional methods is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111931948B_ABST
    Figure CN111931948B_ABST
Patent Text Reader

Abstract

The invention discloses a modeling method based on deep integrated forest regression for measuring the compressive strength of concrete, comprising: preprocessing the original high-dimensional features by adopting a dimensionality reduction strategy suitable for an industrial process to obtain a simplified feature vector; then, taking the simplified feature vector as input, training multiple sub-forest models, selecting the predicted values ​​of several sub-forests by a KNN nearest neighbor method to combine to obtain a layer regression vector, combining the layer regression vector with the simplified feature vector to obtain an enhanced layer regression vector, and then obtaining the output of the layer; secondly, taking the enhanced layer regression vector of the input layer as input to obtain the output of the second layer forest model, repeating the process in sequence until the output of the K-1th layer forest model is completed; finally, taking the output of the intermediate layer forest model of the K-1th layer as the input of the output layer forest model module, training multiple sub-forest models, and performing arithmetic averaging on the predicted outputs of the sub-forest models of the layer to obtain a final prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a deep integrated forest regression modeling method for measuring the compressive strength of concrete. Background Art

[0002] Due to the complex characteristics of complex physical / chemical production processes, such as unclear mechanisms, nonlinearity, and strong coupling, the key process parameters that characterize the product quality and environmental indicators of such processes are usually called difficult-to-measure parameters [1]. These parameters are obtained by first manually sampling at regular intervals and then analyzing them offline in the laboratory (such as concrete compressive strength, dioxin concentration in urban solid waste incineration pollution emissions, and grinding particle size that characterizes grinding quality) or by relying on excellent experts in the field to estimate them empirically at the production site (such as mill load that characterizes grinding efficiency). The above-mentioned inaccurate and long-delayed detection methods have become one of the main bottlenecks restricting the realization of operation optimization and feedback control of the production process [2]. Combining the production process mechanism and empirical knowledge, using offline easily detectable process variables to establish a soft measurement model for difficult-to-measure parameters is one of the effective ways to solve this problem [3].

[0003] As a major branch of machine learning, ensemble learning has been widely used in the field of soft measurement of difficult-to-measure parameters in industrial processes. Decision tree (DT), as a base learner of ensemble learning, can not only handle classification problems, but also regression problems. The most representative one is classification and regression tree (CART) [4]. The method of integrating DT is called forest algorithm (FM), among which the random forest (RF) algorithm proposed by Breiman [5] is the most representative.

[0004] Deep neural network learning algorithms [6] have made traditional machine learning methods lose their competitiveness in many fields. However, they are essentially "black box" models with many hyperparameters and difficult training. Zhou et al. [7] analyzed the internal reasons for the success of DNN and proposed a deep forest (DF) structure consisting of multi-granularity scanning and cascade forest. They conducted research on deep learning of non-neural network structures and preliminarily explored deep models composed of FM models. Kevin et al. [8] also drew inspiration from DNN and proposed a forward-looking deep random forest (FTDRF) by replacing neurons with DT. Although similar related studies are gradually increasing, their research fields are mainly focused on processing classification problems such as image recognition and natural language processing. Their main contribution is to use class distribution vectors as a feature representation method for transmission between layers. For continuous numerical data of industrial processes, reference [9] introduced a deep Boltzmann machine (DBM) to convert the original features into two-dimensional vectors before multi-granularity scanning, and then used the DF method to construct a classifier. The method was verified using industrial process fault diagnosis data. The experimental results show that the combination of DBM and DF effectively improves the recognition rate of fault diagnosis.

[0005] Concrete is an indispensable material in modern construction projects, and its compressive strength is the most important indicator of concrete. In concrete structure engineering, the strength of concrete is tested and evaluated by the results of the compressive strength test of concrete specimens. The inability to measure the real-time data of concrete compressive strength online makes it difficult to optimize and control the concrete production process. The compressive strength parameters of concrete usually require a long period of offline testing and analysis to obtain them. References [10, 11, 12] all proposed soft measurement modeling methods based on ensemble learning to achieve online soft measurement of concrete compressive strength. However, the structure of the soft measurement model of concrete compressive strength in the above research literature is complex, and the representation learning of features is not considered between modules. At the same time, there are problems such as low prediction accuracy of soft measurement values ​​of concrete compressive strength. Summary of the invention

[0006] The difficult-to-detect quality indicators or environmental protection parameters of complex industrial processes usually require long-term offline analysis to obtain. In order to achieve the optimal control of the operation of these processes, it is usually necessary to measure these difficult-to-measure parameters online in real time. The complexity of the mechanisms of industrial processes involving multiple physical and chemical principles makes it difficult to build an interpretable mapping model between high-dimensional input features and difficult-to-measure parameters.

[0007] In view of the above problems, the present invention proposes a modeling method based on deep ensemble forest regression (DEFR) for measuring the compressive strength of concrete, including: using a dimensionality reduction module to preprocess the original high-dimensional features by adopting a dimensionality reduction strategy suitable for industrial processes to obtain a reduced feature vector; using an input layer forest module to take the reduced feature vector as input, training multiple sub-forest models, selecting the predicted values ​​of several sub-forests by the KNN nearest neighbor method to combine to obtain a layer regression vector, combining it with the reduced feature vector to obtain an enhanced layer regression vector, and then obtaining the output of the layer; using an intermediate layer forest module including K-2 layers, which takes the enhanced layer regression vector of the input layer as input, and obtains the output of the second layer forest model in the same way as the input layer forest module, repeating in sequence until the output of the K-1th layer forest model is completed; using an output layer forest module to take the output of the intermediate layer forest model of the K-1th layer as the input of the output layer (Kth layer) forest model module, training multiple sub-forest models, and performing arithmetic averaging of the predicted outputs of the sub-forest models of this layer to obtain the final prediction result. The effectiveness of the proposed method was verified by simulating the concrete compressive strength data of the UCI platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 Flow chart of the present invention;

[0009] Figure 2 The tth sub-forest model F 1,tSchematic diagram of (·);

[0010] Figure 3 RMSE under different sample thresholds;

[0011] Figure 4 RMSE under different number of decision trees;

[0012] Figure 5 Prediction curve of concrete strength training set;

[0013] Figure 6 Prediction curves of the concrete strength validation set;

[0014] Figure 7 Prediction curve of concrete strength test data;

[0015] Figure 8 Correlation coefficient values ​​for different features;

[0016] Fig. 9 RMSE under different sample thresholds;

[0017] Fig.10 RMSE under different number of decision trees;

[0018] Fig.11 Prediction curve of concrete strength training set;

[0019] Fig.12 Prediction curves of the concrete strength validation set;

[0020] Fig.13 Prediction curves for the concrete strength test set. DETAILED DESCRIPTION

[0021] The present invention proposes a modeling method based on deep ensemble forest regression (DEFR) for measuring the compressive strength of concrete. DEFR modeling is realized by a dimensionality reduction module, an input layer forest module, an intermediate layer forest module and an output layer forest module, wherein the number of decision trees in each sub-forest model is J, such as Figure 1 shown.

[0022] Figure 1 In , x represents the original high-dimensional feature vector, which includes eight process measurement values: concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content and concrete placement days (the process measurement value is the feature of the data sample, which will be uniformly described as features below); x dimred represents the reduced feature vector after dimensionality reduction (the input feature vector of the input layer), that is, the dimensionality reduction of the eight features of concrete compressive strength is performed; F 1,t(·) represents the tth sub-forest model of the input layer forest model in the soft measurement of concrete compressive strength; Represents the tth sub-forest model F in the input layer forest model 1,t (·) concrete compressive strength prediction value vector generated by J decision trees; Represents the tth predicted value vector in the input layer forest model The predicted mean of Represents the tth predicted value vector from the input layer forest model using kNN Select the predicted mean Nearby k kNN A regression vector composed of the predicted values ​​of concrete compressive strength; Represents the layer regression vector composed of T regression vectors connected in series in the input layer forest model; Represents the reduced eigenvector x dimred The enhanced layer regression vector is composed of the layer regression vector in series with the input layer forest model, which is also the input feature vector of the middle layer (the second layer) in the soft sensor model of concrete compressive strength; Represents the input feature vector x dimred The enhanced layer regression vector composed of the layer regression vector of the k-1th layer forest model in series is the input feature vector of the kth layer forest model in the soft measurement model of concrete compressive strength; k = 1, 2, 2, K, K represents the number of layers (depth) of DEFR; F k,t (·) represents the tth sub-forest model in the kth layer forest model in the soft sensor model of concrete compressive strength; Represents the tth sub-forest model F in the kth layer forest model k,t (·) concrete compressive strength prediction value vector generated by J decision trees; Represents the tth predicted value vector in the kth layer forest model The predicted mean of Indicates the use of kNN from the predicted value vector Select the predicted mean Nearby kNN A regression vector composed of the predicted values ​​of concrete compressive strength; Represents the layer regression vector composed of T regression vectors connected in series in the kth layer forest model; Represents the input feature vector x dimred The enhanced layer regression vector composed of the layer regression vector of the kth layer forest model in series is the input feature vector of the k+1th layer forest model; Represents the layer regression vector composed of T regression vectors in series of the (K-1)th layer forest model; Represents the input feature vector xdimred The enhanced layer regression vector composed of the layer regression vector of the K-1th layer forest model in series is the input feature vector of the Kth layer forest model in the soft measurement model of concrete compressive strength; F K,t (·) represents the tth sub-forest model in the Kth layer forest model in the soft sensor model of concrete compressive strength; Represents the tth sub-forest model F in the Kth layer forest model K,t (·) concrete compressive strength prediction value vector generated by J decision trees; Represents the tth prediction value vector in the Kth layer forest The predicted mean of It represents the final concrete compressive strength prediction output value of DEFR.

[0023] The functions of the above modules are as follows:

[0024] (1) Dimensionality reduction module: The original high-dimensional feature vector in the concrete compressive strength data is preprocessed using the dimensionality reduction method to obtain the reduced feature vector;

[0025] (2) Input layer forest model module: The simplified feature vector is used as input, and T sub-forest models consisting of J decision trees are constructed to form the input layer forest model. In each sub-forest model, k are selected from the prediction value vector. kNN The predicted values ​​are combined into a layer regression vector, which is then combined with the simplified vector to form an enhanced layer regression vector, thereby obtaining the input of the middle layer forest model module;

[0026] (3) Intermediate forest model module: The enhanced layer regression vector obtained by the input forest model is used as input, and the K-2 layer forest model is trained in the same way as the input forest model.

[0027] (4) Output layer forest model module: The output of the K-1th layer forest model is used as the input of the output layer (Kth layer) forest model module to train the Kth layer forest model. Then, the T predicted true values ​​in the Kth layer forest model are arithmetic averaged to obtain the final concrete compressive strength prediction result.

[0028] The specific processing process of the dimensionality reduction module is as follows:

[0029] Complex physical / chemical production processes generally have strong coupling and nonlinear characteristics, which lead to many redundant features in process data, which easily lead to the curse of dimensionality in modeling

[13] . Consider using dimensionality reduction algorithms to reduce high-dimensional original feature vectors to finite dimensions before model training. Dimensionality reduction can be used to deal with the curse of dimensionality, improve algorithm efficiency and model interpretability, and visualize data. Since the output in the regression problem is a continuous real-valued variable, many simplification methods that work well in classification problems cannot achieve the optimal effect. The following lists linear and nonlinear dimensionality reduction methods for regression problems. When using the method proposed in this application, the corresponding dimensionality reduction method can be selected according to the characteristics of different data sets to obtain the dimensionality-reduced feature vector.

[0030] Among them, the linear dimensionality reduction methods include: (1) dimensionality reduction algorithms based on the first second-order moments: Sliced ​​Inverse Regression (SIR)

[14] , Sliced ​​Average Variance Estimation (SAE)

[15] , Principal Hessian Direction (pHd)

[16] , Directional Regression (DR)

[17] ; (2) Model-based dimensionality reduction algorithms: Principal Fitting Components

[18] ; (3) Mutual Information-Based Dimension Reduction Algorithms: Kernel Dimension Reduction (KDR)

[19] , Least-Squared Dimension Reduction (LSDR)

[20] , Mutual Information-Based Dimension Reduction (MI-BDI)

[21] , and Mutual Information-Based Dimension Reduction (MI-BDI)

[22] . Reduction, MIDR)

[21] ; (4) Dimensionality reduction algorithms based on dependency criteria: Hilbert-Schmidt Independence Criterion (HSIC)

[22] , Distance Covariance (DCOV)

[23] ; (5) Dimensionality reduction algorithms based on regression gradient: Gradient-Based Kernel Dimension Reduction (gKDR)

[24] , Least-Squares Gradients for Dimension Reduction (LSGDR)

[25] , etc.

[0031] The main nonlinear dimensionality reduction methods include: Covariance Operator Inverse Regression (COIR)

[26] , Kernel Sliced ​​Inverse Regression (KSIR)

[27] , etc.

[0032] The specific processing process of the input layer forest model module is:

[0033] The sub-forest in the DEFR structure can adopt various forms of regression forest models, such as random forest, completely random forest, etc. Bootstrap and random subspace method (RSM) are used to train the training set D = {(x i ,y i ),i=1,2,…N}∈R N×M Random sampling of samples and features is performed to increase the diversity of sub-forests.

[0034] First, the construction process of the input layer sub-forest is described.

[0035] Bootstrap and RSM are used to randomly sample the training set D containing eight features including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content and concrete placement days and the concrete compressive strength test value to input the J training subsets of the tth sub-forest model in the layer forest model. For example, its generation process can be expressed as:

[0036]

[0037] Where D represents the training set of concrete compressive strength test values ​​in the input layer forest model, which contains eight features including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content and concrete placement days in the soft measurement model of concrete compressive strength; J represents the number of Bootstrap times, and also represents the number of decision trees in each sub-forest model in the input layer forest model; represents the jth training subset of the tth sub-forest in the input layer forest model, where indicates that the jth training subset selects M from eight features including concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content and concrete placement days. jTraining samples of a feature, y j Represents the actual measured value of the concrete compressive strength; m = 1, …, M j , M j Represents the number of features selected from 8 features in the jth training set of the tth sub - forest in the input - layer forest model. Usually, there is M j <<M; t = 1, 2, …, T, where t represents the tth sub - forest model in the input - layer forest model.

[0038] Using the above J training subsets Construct J decision trees in the tth sub - forest of the concrete compressive strength soft - sensing model to obtain the tth sub - forest model F in the input layer 1,t (·), as shown Figure 2 as follows

[0039] The process of constructing the sub - forest model is detailed in reference

[28] . Repeat the above steps T times to obtain the set of input - layer forest models

[0040] Next, describe the generation process of the enhanced - layer regression vector of the input - layer forest model.

[0041] For the tth sub - forest model in the input - layer forest model, each decision - tree model will generate a predicted value of the concrete compressive strength for the concrete process measurement value sample Then obtain J predicted values of the concrete compressive strength Composing a predicted - value vector

[0042] Calculate the predicted mean of the tth sub - forest model in the input - layer forest model,

[0043]

[0044] Select the predicted mean through kNN Nearby k kNN Concrete compressive strength predicted values to form the regression vector of the tth sub - forest Repeat the above steps T times to obtain the layer regression vectors of T sub - forest models in the input - layer forest model

[0045] Next, concatenate and combine the reduced - dimension feature vector x after dimensionality reduction of the 8 features of the concrete compressive strength dimred with the layer regression vector to obtain the enhanced - layer regression vector as the output of the input - layer forest model which is the input of the middle - layer forest model (the second layer) of the concrete compressive strength soft - sensing model. Its generation process can be expressed as

[0046]

[0047] Among them, k kNN Indicates the number of predicted values ​​of concrete compressive strength near the selected predicted mean.

[0048] The specific processing process of the intermediate forest model module is:

[0049] Taking the kth layer forest model as an example, the construction process of the intermediate layer forest model module is introduced.

[0050] Training dataset for the kth layer forest model is the enhanced layer regression vector output by the k-1th layer forest model The combination of the concrete compressive strength test value, where the features include: concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, concrete placement days and other eight features and layer regression vector The representation process is:

[0051]

[0052] Where y represents the true value vector of concrete compressive strength in the training set D; N represents the number of samples in the training set D; The layer regression vector representing the k-1th layer of the forest model and the simplified feature vector x after the dimensionality reduction of the eight features of concrete compressive strength dimred The enhancement layer regression vector after concatenation; represents the input training set of the kth layer forest model, where x k,i It represents the ith layer regression vector containing eight features including concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, and concrete placement days. The training samples, y i represents the actual test value of the ith concrete compressive strength; M k =M+(k kNN ×T) indicates that the kth layer forest model contains eight features, including concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, concrete placement days, etc., and the layer regression vector and concrete compressive strength test value training data set D k The number of input features.

[0053] Then, Bootstrap and RSM were used to regress eight features, including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, and concrete placement days, layer regression vector and concrete compressive strength test value training data set D k The random sampling of samples and features can be expressed as follows:

[0054]

[0055] in, represents the jth training subset of the tth sub-forest model in the kth layer forest model, where represents the eight features and layer regression vectors in the jth training subset, including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, and concrete placement days. Select The training samples of features, y j Indicates the actual test value of concrete compressive strength; Represents the jth training set of the tth sub-forest model in the kth layer forest model from 8 features and layer regression vectors The number of features selected in

[0056] With the above J training subsets Construct J decision trees of the tth sub-forest model in the kth layer forest model in the soft measurement model of concrete compressive strength, and obtain the tth sub-forest model F in the kth layer forest model. k,t (·).

[0057] Repeat the above steps T times to get the set of kth layer forest models

[0058] Next, the process of generating the enhanced layer regression vector of the kth layer forest model is described.

[0059] The tth sub-forest model in the kth layer forest model, each decision tree model generates a concrete compressive strength prediction value for the input J predicted values ​​of concrete compressive strength can be obtained The predicted value vector

[0060] Calculate the predicted mean of the tth sub-forest model in the kth layer,

[0061]

[0062] Selecting the prediction mean through kNN Nearby k kNN The predicted values ​​of concrete compressive strength constitute the regression vector of the tth sub-forest model. Repeat the above steps T times to get the regression vectors of T sub-forest models, and combine them to get the layer regression vector of the kth layer forest model

[0063] Next, the simplified feature vector x after dimensionality reduction of the eight features of concrete compressive strength is dimred With layer regression vector Combine them in series to get the enhanced layer regression vector output by the kth layer forest model This is the input of the k+1th layer forest model. Its generation process can be expressed as:

[0064]

[0065] The specific processing process of the output layer forest model module is:

[0066] Training dataset for the Kth layer forest model is the enhanced layer regression vector output by the K-1th layer forest model The combination of the concrete compressive strength test value, where the features include: concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, concrete placement days and other eight features and layer regression vector The representation process is:

[0067]

[0068] in, represents the training set of the Kth layer forest model, where x K,i It represents the ith layer regression vector containing eight features including concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, and concrete placement days. The training samples, y i represents the actual test value of the ith concrete compressive strength; M K =M+(k kNN ×T) indicates that the Kth layer contains eight features, namely, concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, and concrete placement days, and the layer regression vector and concrete compressive strength test value training data set D K The number of features.

[0069] Then Bootstrap and RSM are used to regress eight features including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, concrete placement days, etc. and concrete compressive strength test value training data set D K By performing random sampling of samples and features, the process of generating J training subsets of the tth sub-forest model in the Kth layer forest model can be expressed as:

[0070]

[0071] in, represents the jth training subset of the tth sub-forest model in the Kth layer forest model, where represents the eight features and layer regression vectors in the jth training subset, including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, and concrete placement days. Select The training samples of features, y j Indicates the actual test value of concrete compressive strength; Represents the jth training set of the tth sub-forest model in the Kth layer forest model from 8 features and layer regression vectors The number of features selected in

[0072] Use the above J training subsets to construct J decision trees of the tth sub-forest model in the Kth layer, and obtain the tth sub-forest model F of the Kth layer. K,t (·). Repeat the above steps T times to obtain the model of the Kth layer forest module

[0073]

[0074] The tth sub-forest model in the Kth layer, each decision tree model will produce a prediction value of concrete compressive strength Then we get J predicted values ​​of concrete compressive strength The predicted value vector Calculate the predicted mean of the tth sub-forest model in the Kth layer,

[0075]

[0076] Repeat the above steps T times to obtain the prediction output set of T sub-forest models

[0077] Finally, the predicted values ​​of concrete compressive strength of T sub-forest models are arithmetic averaged.

[0078]

[0079] in, It represents the final concrete compressive strength prediction output of the DEFR model.

[0080] Simulation Verification of Embodiment

[0081] Experimental data description

[0082] The concrete compressive strength dataset [29,30] provided by the University of California Irvine (UCI) platform is used to verify the proposed method. The dataset contains 1030 samples, of which the first 8 columns are inputs, namely the content of concrete, blast furnace slag powder, fly ash, water, water reducer, coarse aggregate and fine aggregate in each cubic meter of concrete and the number of days the concrete is placed; the 9th column is the output, i.e. the concrete compressive strength. In this paper, 1 / 2 of the 1030 samples are used as training samples, 1 / 4 as validation samples, and 1 / 4 as test samples.

[0083] According to the characteristic attributes of the concrete compressive strength data set, the following experiments will be conducted on the dimensionless reduction module (hereinafter, for the sake of distinction, the model of the dimensionless reduction module is represented as DEFR-Nodimred) and the dimensional reduction module (the model of the dimensional reduction module is represented as DEFR-dimred). The initial parameters in the experiment are set as follows: the number of sub-forests in the forest layer of the concrete compressive strength soft measurement model is set to T = 8, including 4 random forests and 4 completely random forests, and the number of concrete compressive strength prediction values ​​selected by kNN is k. kNN =1.

[0084] Dimensionless reduction module

[0085] Experimental Results

[0086] The statistical results of the linear correlation coefficients between the eight characteristics of concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content and concrete placement days in the concrete compressive strength and the true value of concrete compressive strength are as follows: Figure 3 As shown,

[0087] The absolute values ​​of the correlation coefficients between the eight characteristics and the compressive strength of concrete are as follows: Figure 3As shown in the figure, the eight features are divided into two parts with 0.2 as the dividing line. Among them, the features greater than 0.2 are: concrete content, water content, water reducer content and concrete placement days. Therefore, the four features of concrete content, water content, water reducer content and concrete placement days are selected from the eight features of concrete compressive strength data as the training set through the dimensionality reduction module.

[0088] The average of 50 runs is taken as the final result, and the parameters are set to K=50, M j =4,k kNN =1,T=8,J=500, where M j =4 means that four features of concrete content, water content, water reducing agent content and concrete placement days in the concrete compressive strength data set are randomly selected from the layer regression vector. The training sample threshold θ of the test decision tree leaf node Forest The relationship between the RMSE of the soft sensor model DEFR-Nodimred for concrete compressive strength in the validation set is shown in the experimental results. Figure 4 shown.

[0089] Depend on Figure 4 It can be seen that when the training sample threshold θ of the leaf node DT =10, the RMSE (7.1736) value of the validation set reaches the minimum. DT When it increases, RMSE also increases. Therefore, the training sample threshold θ of the decision tree leaf node is selected. DT =10.

[0090] Then, the relationship between the number of decision trees J of the sub-forest model in the forest layer model and the RMSE of the concrete compressive strength soft sensor model DEFR-Nodimred in the validation set is tested, as shown in Figure 5 shown.

[0091] Depend on Figure 5 It can be seen that when the number of decision trees in the sub-forest of the forest layer in the concrete compressive strength soft sensor model DEFR-Nodimred is J=100, the RMSE (6.9979) value of the validation set reaches the minimum.

[0092] The parameters of the final concrete compressive strength soft sensor model DEFR-Nodimred are determined as: T = 8, k kNN =1,K=50,θ DT =10,M j =4,J=100.

[0093] Method comparison

[0094] The completely random forest method (CRF) and random forest (RF) are used to compare with the proposed method DEFR-Nodimred, where the CRF parameters are set to: θ DT =10,M j =4, J = 100, RF parameters are set to: θ DT =10,M j =4,J=100.

[0095] The prediction curves of different soft sensing methods are as follows: Figure 6 , 7 and 8.

[0096] Table 1 Comparison results of different methods

[0097]

[0098] Figure 6-8 The results in Table 1 show that: (1) CRF has the largest prediction error in predicting concrete compressive strength due to its inherent randomness, and the test set error is 9.3488; (2) RF uses the minimum average error rule to split the nodes of the decision tree, making its prediction performance of concrete compressive strength stronger than CRF, and the test set error is 7.5390; (3) The DEFR-Nodimred method proposed in this paper has the best prediction performance for predicting concrete compressive strength in the training set, validation set and test set, and the test set error is 7.2320, and its number of layers K = 3.

[0099] Dimensionality reduction module

[0100] Experimental Results

[0101] The average of 50 runs is taken as the final result, and the parameters are set to K=50, M j =4,k kNN =1,T=8,J=500, where M j =4 means randomly selecting 4 features from the 8 features and layer regression vectors in the concrete compressive strength dataset as input features. Training sample threshold θ for testing decision tree leaf nodes Forest The relationship between the RMSE of the concrete compressive strength soft sensor model DEFR-dimred in the validation set is shown in the experimental results. Fig. 9 shown.

[0102] Depend on Fig. 9 It can be seen that when the training sample threshold θ of the leaf node DT =10, the RMSE (7.4893) value of the validation set reaches the minimum. DT When it increases, RMSE also increases. Therefore, the training sample threshold θ of the decision tree leaf node is selected. DT=10.

[0103] Then, the relationship between the number of decision trees J of the sub-forest model in the forest layer model and the RMSE of the concrete compressive strength soft sensor model DEFR-dimred in the validation set is tested, as shown in Fig.10 shown.

[0104] Depend on Fig.10 It can be seen that when the number of decision trees in the sub-forest of the forest layer in the concrete compressive strength soft sensor model DEFR-dimred is J=200, the RMSE (7.4771) of the validation set reaches the minimum.

[0105] The parameters of the final concrete compressive strength soft measurement model DEFR-dimred are determined as: T = 8, k kNN =1,K=50,θ DT =10,M j =4,J=200.

[0106] Method comparison

[0107] The completely random forest method (CRF) and random forest (RF) are used to compare with the proposed method DEFR-dimred, where the CRF parameters are set to: θ DT =10,M j =4, J = 200, RF parameters are set to: θ DT =10,M j =4,J=200.

[0108] The prediction curves of different soft sensing methods are as follows: Fig.11 , 12 and 13.

[0109] The statistical results of different modeling methods are shown in Table 2.

[0110] Table 2 Comparison results of different methods

[0111]

[0112] Figure 11-13 The results in Tables 1 and 2 show that: (1) After omitting the dimensionality reduction module, the DEFR-dimred method proposed in this paper has the best prediction performance for the prediction of concrete compressive strength in the training set, validation set and test set. The error of the test set is 6.4018, and the number of layers K is 3; (2) Compared with DEFR-Nodimred without dimensionality reduction, DEFR-dimred is better than DEFR-Nodimred in predicting the compressive strength of concrete in both the validation set and the test set, which shows the effectiveness of the deep forest structure proposed in this paper.

[0113] Therefore, the deep integrated forest regression model first proposed by the proposed method in this paper has the best prediction performance in the soft measurement of concrete compressive strength.

[0114] Aiming at the soft sensor modeling of difficult-to-measure parameters in industrial processes, this paper proposes a modeling method based on deep ensemble forest regression. The main contribution is that it solves the feature representation method between levels in the deep ensemble forest regression problem for the first time, and realizes the application of the deep forest structure in the regression modeling problem for the first time. The effectiveness of the proposed method is verified by simulating the concrete compressive strength data of the UCI platform.

[0115] References

[0116] [1] Chai Tianyou. Operation optimization and feedback control of complex industrial processes. Acta Automatica Sinica, 2013, 39(11): 1744-1757.

[0117] [2] Tang Jian, Tian Fuqing, Jia Meiying, Li Dong. Soft measurement of rotating machinery load based on spectrum data drive[M]. National Defense Industry Press, June 2015, Beijing

[0118] [3]Kadlec P, Gabrys B, Strand S. Data-driven soft-sensors in the processindustry[J]. Computers and Chemical Engineering, 2009, 33(4):795-814.

[0119] [4] Breiman L, Friedman J, Stone C. Classification and Regression Trees. Wadsworth, 1984.

[0120] [5]Breiman, L. Random forests. Machine Learning, 2001, 45(1), 5-32.

[0121] [6] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, Cambridge, MA, 2016.

[0122] [7]ZhouZH,FengJ.Deepforest:Towards an alternative to deep neuralnetworks[J].eprintarXiv:1702.08835,2017.[8]K Miller,C Hettinger,et al.Forwardthinking:Building deep random forests.2017,arXiv:1705.07366.

[0123] [9]Hu G,Li H,Xia Y,Luo L.Adeep Boltzmann machine and multi-grainedscanning forest ensemble collaborative method and its application toindustrial fault diagnosis.Computers in Industry.2018.100,

[0124] 287-296.

[0125]

[10] Jian Tang,Jian Zhang,Zhiwei Wu,et al.Modeling collinear datausing double-layer GA-based selective ensemble kernel partial least squaresalgorithm.Neurocomputing.219(2017):248-262.

[0126]

[11] Jian Tang,Junfei Qiao,Jian Zhang,et al.Combinatorial optimizationof input features and learning parameters for decorrelated neural networkensemble-based soft measuring mosel.Neurocomputing.

[0127] 275(2018):1426-1440.

[0128]

[12] v Tang Jian, Qiao Junfei. Soft measurement of dioxin emission concentration in solid waste incineration process based on selective ensemble kernel learning algorithm[J]. Journal of Chemical Industry and Engineering, 2019, 70(02): 696-706.

[0129]

[13] Shan HM, Zhang J P.Real-valued multivariate dimension reduction:review[J].Journal of Automation, 2018,

[0130] 44(2):193-215.

[0131]

[14] Li K C.Sliced ​​inverse regression for dimension reduction.Journal of the American Statistical Associatio,

[0132] 1991,86(414):316-327.

[0133]

[15] Cook RD, Weisberg S. Sliced ​​inverse regression for dimensionreduction:comment. Journal of the American Statistical Association, 1991, 86 (414): 328-332.

[0134]

[16] Li K C.On principal Hessian directions for data visualization and dimension reduction:another application of Stein's lemma.Journal of the American Statistical Association,1992,87(420):1025-1039.

[0135]

[17] Li B,Wang S L.On directional regression for dimensionreduction.Journal of the American Statistical Association,2007,102(479):997-1008.

[0136]

[18] Cook R D.Fisher 1ecture:dimension reduction inregression.Statistical Science,2007,22(1):1-26

[0137]

[19] Fukumizu K,Bach F R,Jordan M I.Dimensionality reduction forsupervised learning with reproducing kernel Hilbert spaces.Journal of MachineLearning Research,2004,5:73—99.

[0138]

[20] Suzuki T,Sugiyama M.Suficient dimension reduction via squared—loss mutual information estimation.In Proceedings of InternationalConferenceon Artificial Intelligence and Statistics,Chia Laguna Resort,Sardinia,

[0139] Italy,2010,9:804-811.

[0140]

[21] Faivishevsky L,Goldberger J.Dimensionality reduction based onnon—parametric mutual information.

[0141] Neurocom—putting.2012.80:3l-37.

[0142]

[22] Gretton A,Bousquet O,Smola A,Schokopf B.Measuring Statisticaldependence with Hilbert—

[0143] Schmidtnorms.In:Proceedings of the 16th International ConferenceonAlgorithmic Learning Theory.Berlin,

[0144] Heidelberg:Springer-Verlag,2005.63-77.

[0145]

[23] Szekely G J,Rizzo M L,Bakirov N K.Measuring and testingdependence by correlation of distances.The Annals of Statistics,2007,35(6):2769-2794.

[0146]

[24] Fukumizu K,Leng C L.Gradient—based kernel dimension Reductionfor regression.Journal of the American Statistical Association,2014,109(505):359-370.

[0147]

[25] Sasaki H,Tangkaratt V,Sugiyama M.Suficient dimension reductionvia direct estimation of the gradients of logarithmic conditionaldensities.In:Proceedings of the 7th Asian Conferenceon MachineLearning.HongKong,

[0148] China:PMLR,2015.33-48.

[0149]

[26] Kim M,Pavlovic V,Central subspace dimensionality reduction usingcovariance operators.IEEE Transactions on Pattern Analysis and MachineIntelligence,20113,3(4):657—670.

[0150]

[27] Wu H M.Kernel sliced ​​inverse regression with applications to classification.Journal of Computational and Graphical Statistics,2012,17(3):590-610.

[0151]

[28] Tang Jian, Xia Heng, Qiao Junfei, Guo Zihao. , State Intellectual Property Office, application number: 202010083784.4, application date: February 10, 2020.

[0152]

[29] Yeh I C. Modeling of Strength of High Performance Concrete Using Artificial Neural Networks[J]. Cement and Concrete Research, 1998, 28(12): 1797-1808.

[0153]

[30] Tang J, Yu W, Chai TY, et al.On-line Principal Component Analysis with Application to Process Modeling[J].

[0154] Neurocomputing,2012,82(1):167-178.

Claims

1. A modeling method based on deep ensemble forest regression for measuring the compressive strength of concrete, It is characterized in that The following steps are involved: Step 1: Using a dimensionality reduction module to preprocess the original high-dimensional features by adopting a dimensionality reduction strategy suitable for industrial processes to obtain a reduced feature vector; Step 2: Use the input layer forest model module to take the simplified feature vector as input, train multiple sub-forest models, select the predicted values ​​of several sub-forests through the KNN nearest neighbor method to combine and obtain the layer regression vector, combine it with the simplified feature vector to obtain the enhanced layer regression vector, and then obtain the output of the layer; Step 3: The intermediate forest model module includes a K-2 layer, which takes the enhanced layer regression vector of the input layer as input, obtains the output of the second layer forest model in the same way as the input layer forest model module, and repeats the process until the output of the K-1 layer forest model is completed; Step 4: Use the output layer forest model module to take the output of the intermediate layer forest model module of the K-1th layer as the input of the output layer forest model module, train multiple sub-forest models, and perform arithmetic averaging of the predicted outputs of the sub-forest models of this layer to obtain the final prediction result; The original high-dimensional feature vector in step 1 includes: concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content and concrete placement days; The specific processing process of the input layer forest model module includes the following steps: Step 21, describe the construction process of the input layer subforest Bootstrap and RSM were used to randomly sample the training set D containing eight features, namely, concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, concrete placement days, and concrete compressive strength test values. Suppose the J training subsets of the tth sub-forest model in the input layer forest model module are With the above J training subsets Construct J decision trees in the tth sub-forest in the soft measurement model of concrete compressive strength, and obtain the tth sub-forest model F in the input layer. 1,t (·), Repeat the above steps T times to get the set of input layer forest model modules. Step 22, describing the process of generating the enhanced layer regression vector of the input layer forest model module For the tth sub-forest model in the input layer forest model module, each decision tree model will generate a concrete compressive strength prediction value for the concrete process measurement value sample Then we get J predicted values ​​of concrete compressive strength The predicted value vector Calculate the predicted mean of the tth sub-forest model in the input layer forest model module, Selecting the prediction mean through kNN Nearby k kNN The predicted values ​​of concrete compressive strength constitute the regression vector of the tth sub-forest Repeat the above steps T times to obtain the layer regression vector of T sub-forest models in the input layer forest model module Step 23: Reduce the dimensions of the eight features of concrete compressive strength to the reduced feature vector x dimred With layer regression vector Combine them in series to obtain the enhanced layer regression vector as the output of the input layer forest model module It is the input of the middle-layer forest model module of the soft-sensing model of concrete compressive strength; The specific processing process of the intermediate forest model module is: Suppose the training data set of the kth layer forest model is is the enhanced layer regression vector output by the k-1th layer forest model and the concrete compressive strength test value, where the features include: concrete content, blast furnace slag powder content, fly ash content, water content, water reducing agent content, coarse aggregate content, fine aggregate content, concrete placement days eight features and layer regression vector Bootstrap and RSM are used to regress eight features, including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, and concrete placement days. and concrete compressive strength test value training data set D k Perform random sampling of samples and features, Take the J training subsets of the tth sub-forest model in the Kth layer forest model Construct J decision trees of the tth sub-forest model in the kth layer forest model in the soft measurement model of concrete compressive strength, and obtain the tth sub-forest model F in the kth layer forest model. k,t (·), Repeat the above steps T times to get the set of kth layer forest models Next, the process of generating the enhanced layer regression vector of the kth forest model is described. The tth sub-forest model in the kth layer forest model, each decision tree model generates a concrete compressive strength prediction value for the input J predicted values ​​of concrete compressive strength can be obtained The predicted value vector Calculate the predicted mean of the tth sub-forest model in the kth layer, Selecting the prediction mean through kNN Nearby k kNN The predicted values ​​of concrete compressive strength constitute the regression vector of the tth sub-forest model. Repeat the above steps T times to get the regression vectors of T sub-forest models, and combine them to get the layer regression vector of the kth layer forest model Next, the simplified feature vector x after dimensionality reduction of the eight features of concrete compressive strength is dimred With layer regression vector Combine them in series to get the enhanced layer regression vector output by the kth layer forest model This is the input of the k+1th layer forest model; The specific processing process of the output layer forest model module is: Suppose the training data set of the Kth layer forest model is is the enhanced layer regression vector output by the K-1th layer forest model The combination of the concrete compressive strength test value includes eight features and layer regression vectors: concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, and concrete placement days. Bootstrap and RSM are used to regress eight features, including concrete content, blast furnace slag powder content, fly ash content, water content, water reducer content, coarse aggregate content, fine aggregate content, and concrete placement days. and concrete compressive strength test value training data set D K Perform random sampling of samples and features, Using the J training subsets of the tth sub-forest model in the Kth layer forest model, construct the J decision trees of the tth sub-forest model in the Kth layer to obtain the tth sub-forest model F of the Kth layer. K,t (·), repeat the above steps T times to get the model of the Kth layer forest module The tth sub-forest model in the Kth layer, each decision tree model will produce a prediction value of concrete compressive strength Then we get J predicted values ​​of concrete compressive strength The predicted value vector Calculate the predicted mean of the tth sub-forest model in the Kth layer, Repeat the above steps T times to obtain the prediction output set of T sub-forest models Finally, the predicted values ​​of concrete compressive strength of T sub-forest models are arithmetic averaged. in, It represents the final concrete compressive strength prediction output of the DEFR model.

Citation Information

Patent Citations

  • A method for predicting dioxin emission concentration

    CN111260149B

  • Non-linear optimization method for mix proportion of concrete

    CN104261742A

  • Method for predicting remaining service life of rolling bearing integrated with KELM

    CN109187025A