Soft Sensing Method for Dioxin Emission in MSWI Process Based on Missing Data Filling
Through expert experience and multi-model integration method of reducing feature, the missing data of the MSWI process is filled with, which solves the impact of missing data on soft measurement of DXN emission concentration, improves data quality and prediction accuracy, and realizes the optimization and pollution control of the MSWI process.
Patent Information
- Application Number
- CN202210606547.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-31
AI Technical Summary
During the MSWI process, due to sensor failure, process interference or human factors, there are different degrees of missing process data, which affects the soft measurement of DXN emission concentration and the optimization of MSWI process.
Expert experience and multi-model integration method for reducing feature were used to fill in the missing data of the MSWI process, identify the types of missing random distribution, time dimension and feature dimension, and fill them using linear interpolation and multi-model integration strategies respectively, and finally establish a DFR-clfc model for prediction of DXN emission concentration.
The quality of MSWI process data is improved, the accuracy of missing data filling and the prediction accuracy of DXN emission concentration are significantly improved, and the needs of MSWI process operation optimization and urban pollution control are met.
Smart Images

Figure CN114970353B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dioxin emission soft measurement, and particularly to a soft measurement method for dioxin emission in the MSWI process based on missing data filling. Background Art
[0002] Municipal Solid Waste Incineration (MSWI) has advantages such as harmlessness, reduction, and resource utilization compared to traditional methods such as landfilling and biological treatment. Dioxin (DXN) is a trace organic pollutant emitted during this process, and soft measurement of its emission concentration has always been a difficult and popular issue in the industry. Accurate and complete process data realization is the basis for DXN soft measurement, as well as for the operation optimization of the MSWI process and urban pollution control. Due to the long cycle and high cost characteristics of DXN detection, the corresponding historical process data is often stored on an hourly scale. Affected by sensor failures, process disturbances, or human factors, these process data often have varying degrees of missing values. Taking the recorded data of a certain incineration plant in Beijing from 2009 to 2020 as an example, it contains 2%-3% of missing data. Obviously, this has an adverse impact on mining the laws contained in the operation data and establishing a soft measurement model for DXN emission concentration. Therefore, it is crucial to fill the missing data in the MSWI process.
[0003] Currently, the methods for dealing with missing data are mainly divided into simple deletion method, weighting method, and filling method. The essence of the existing filling methods is to first infer multiple predicted values for the missing values, then comprehensively analyze multiple complete data sets generated based on these predicted values, and determine the final filling result according to a certain criterion, but ultimately, the estimated value of a single model is still selected for filling. For the MSWI process, different process variables have different distribution characteristics at different times. Usually, based on the physical meaning of the process variables and the changes in their distribution over time, domain experts can accurately infer discontinuous missing values. Summary of the Invention
[0004] The purpose of the present invention is to provide a soft measurement method for dioxin emission in the MSWI process based on missing data filling, which uses an expert experience and reduced feature multi-model integration filling method to fill the missing data in the MSWI process, improving the quality of the MSWI process data.
[0005] To achieve the above purpose, the present invention provides the following solutions:
[0006] A soft measurement method for dioxin emission in the MSWI process based on missing data filling specifically includes the following steps:
[0007] S1, Missing type identification: Classify according to different situations of missing data in the original MSWI process data, and identify the process data as three types: randomly distributed, time dimension, and feature dimension missing;
[0008] S2, Missing data filling: First, based on expert experience rules, fill in the missing data of randomly distributed and time dimension, and then use the method based on reduced feature multi-model integration to fill in the missing data of feature dimension to obtain the filled data set;
[0009] S3, Soft sensor modeling based on filled data: Establish a DFR-clfc model based on the filled data set to obtain the predicted value of DXN emission concentration.
[0010] Furthermore, in step S1, classify according to different situations of missing data in the original MSWI process data, and identify the process data as three types: randomly distributed, time dimension, and feature dimension missing, specifically including:
[0011] S101, Randomly distributed data missing situation:
[0012] The random occurrence of data missing has nothing to do with its own characteristics, the values of other characteristics, and time. The missing data is randomly distributed among different characteristics. The randomly distributed data missing situation can be expressed as follows:
[0013]
[0014] S102, Time dimension data missing situation:
[0015] At a certain moment, all variables have missing situations, that is, the missing row vector, which is expressed as follows:
[0016]
[0017] S103, Feature dimension data missing situation:
[0018] A certain feature value is missing in the process data, that is, the missing column vector, which is described as follows:
[0019]
[0020] Among them, NaN in the above formula represents missing data, N is the number of samples, and M is the number of feature dimensions.
[0021] Furthermore, in step S2, first, based on expert experience rules, fill in the missing data of randomly distributed and time dimension, and then use the method based on reduced feature multi-model integration to fill in the missing data of feature dimension to obtain the filled data set, specifically including:
[0022] S201. First, fill in the missing data in the random distribution and time dimension based on the expert experience rules. The results are as follows:
[0023]
[0024] Among them, is the data without missing values containing M - M′ features, (M′ << M) is the data to be filled;
[0025] S202. The filling of the missing data in the feature dimension takes as the input, and obtains the filling value by adopting the multi - model integration strategy based on the reduced features which includes three parts: reducing features based on mutual information MI, predicting the filling value based on multiple sub - models, and integrating the filling value based on BLR, to fill the missing feature Taking the filling of as an example, first obtain the correlation set of the feature to be filled after reducing the features. The processed dataset is as follows:
[0026]
[0027] Among them, is the feature to be filled, is the correlation set of;
[0028] Use it to construct sub - models based on the RF, GBDT, and BPNN algorithms, and record the corresponding predicted outputs as and Then, combine the above - mentioned predicted values to construct a fusion model based on BLR, that is:
[0029]
[0030] Among them, is composed of the predicted output and the feature to be filled ;
[0031] Finally, take its output as the final filling value of the m′ - th missing feature.
[0032] Furthermore, in step S3, establish a DFR - clfc model based on the filled dataset to obtain the DXN emission concentration prediction value, specifically including:
[0033] S301. For the input - layer forest model, adopt Bootstrap and RSM to randomly sample the training set D fill to construct I sub - forest models based on J DTs Accordingly, the predicted mean of the $i$-th sub-forest model is calculated by the following formula:
[0034]
[0035] where, is the predicted value of the $i$-th sub-forest model, is the predicted mean of the $J$ DTs in this sub-forest model;
[0036] S302. Then, $k$ kNN nearest neighbor values are reselected by the kNN method to form the regression vector of the $i$-th sub-forest model in the input layer After repeating the above steps $I$ times, the layer regression vector of the input layer forest model is obtained Accordingly, the input of the middle layer forest model is:
[0037]
[0038] where, is the enhanced regression vector of the input layer forest model, and $f$ FeaCom (·) is the combination function, is the layer regression vector of the input layer forest model, and $k$ kNN is the number of nearest neighbor values selected, $X$ is the process variable in $D$ fill in the training set; fill
[0039] The middle layer forest model contains $L - 2$ layers of sub-forest models, and the input of its $\lambda$ ($\lambda=2,3,\cdots,L - 1$) layer is expressed as:
[0040]
[0041] where, is the output feature vector of the $\lambda - 1$ layer, $y$ DXN is the emission concentration of $D_{X\times N}$, $N$ is the number of samples, and $M$ λ $=M+(k$ kNN $\times I)\times(\lambda - 1)$ is the number of features;
[0042] The enhanced regression vector of each layer is obtained by cross-layer fully connected method, and is expressed as:
[0043]
[0044] where, The acquisition method is the same as that of the input layer;
[0045] Therefore, the process of the $\lambda$ layer forest model generating the output feature vector is as follows:
[0046]
[0047] Among them, is the layer regression vector of the λ-th layer forest model;
[0048] Furthermore, the input of the L-th layer output layer forest model is expressed as:
[0049]
[0050] Among them, M L = M + (k kNN × I) × (L - 1) represents the number of features.
[0051] Based on D fill,L Construct I sub-forest models containing J DTs Among them, the predicted value vector of the i-th sub-forest model is is the predicted value obtained from the J DTs of the i-th sub-forest model in the L-th layer, and the average predicted value of the i-th sub-forest model is
[0052] Finally, the output of the DFR-clfc model is the predicted value of the DXN emission concentration
[0053]
[0054] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: The soft measurement method for dioxin emissions in the MSWI process based on missing data filling provided by the present invention adopts a filling method integrating expert experience and reduced features multi-model, and is compared with the traditional single-model filling method. The conclusions are as follows:
[0055] (1) After classifying the missing data in the MSWI process, first fill the randomly distributed and time-dimensional missing data based on expert experience, ensuring the accuracy of filling the missing values in the subsequent feature dimensions of the data;
[0056] (2) When using a learning model to fill the missing values in the feature dimensions, considering the redundancy and correlation among its input variables, first use the MI method for feature reduction and then construct a heterogeneous integrated model for missing value filling, effectively improving the filling accuracy compared with the traditional single-model method;
[0057] (3) Use the filled data to establish a DFR-clfc model to predict the DXN emission concentration, verifying that the filling effect of the proposed method is significantly better than the traditional single-model method, meeting the actual application needs while improving the data quality. Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 is the flowchart of the soft measurement method for dioxin emissions in the MSWI process based on missing data filling in the embodiment of the present invention;
[0060] Figure 2 is the missing data diagram of the MSWI process in the embodiment of the present invention;
[0061] Figure 3a is the schematic diagram of the filling result of the temperature at the top of the grate in the burnout section on the right in the embodiment of the present invention;
[0062] Figure 3b is the schematic diagram of the filling result of the water flow B of the mixer in the embodiment of the present invention;
[0063] Figure 4 is the group diagram of the data distribution of the flue gas oxygen concentration at the inlet of the NID system in the embodiment of the present invention;
[0064] Figure 5 is the data distribution diagram of the main steam flow rate in the embodiment of the present invention. Specific Embodiments
[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0066] The purpose of the present invention is to provide a soft measurement method for dioxin emissions in the MSWI process based on missing data filling, which uses an expert experience and a reduced feature multi-model integration filling method to fill the missing data in the MSWI process, improving the quality of the MSWI process data.
[0067] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0068] As Figure 1 shown, the soft measurement method for dioxin emissions in the MSWI process based on missing data filling provided by the present invention includes the following steps:
[0069] S1, Missing type identification: Classify according to different situations of missing data in the original MSWI process data, and identify the process data as three types: random distribution, time dimension, and feature dimension missing;
[0070] S2, Missing data filling: First, fill in the missing data of random distribution and time dimension based on expert experience rules, and then use the method based on multi-model integration of reduced features to fill in the missing data of feature dimension to obtain the filled data set;
[0071] S3, Soft sensor modeling based on filled data: Establish a DFR-clfc model based on the filled data set to obtain the predicted value of DXN emission concentration.
[0072] Figure 1 Among them, represents the original MSWI process data; the data after filling by expert experience is denoted as where X nomis represents the feature set without missing values, and X tbfil represents the feature to be filled (i.e., the missing feature) with dimension M'; taking the m'-th missing feature as an example for illustration, is the high-correlation set of and respectively represent the missing value prediction vectors of the m'-th RF, GBDT, and BPNN sub-models, represents the input of the m'-th BLR model, represents the finally obtained filled value; represents the missing data that has been filled;
[0073] represents the complete process data after filling; is the predicted value of DXN emission concentration.
[0074] In the research of this paper on DXN soft sensor, the MSWI process data is statistically recorded with an hourly sampling time, Figure 2 which respectively shows the data missing phenomena of random distribution, time dimension, and feature dimension in the data. Among them, in step S1, according to different situations of missing data in the original MSWI process data, the process data is identified as three types: random distribution, time dimension, and feature dimension missing, specifically including:
[0075] S101, Random distribution of data missing situation:
[0076] The random occurrence of data missing has nothing to do with its own characteristics, the values of other characteristics, and time. The missing data is randomly distributed among different features. The random distribution of data missing situation, such as Figure 2As shown on the left, it can be expressed as follows:
[0077]
[0078] S102, Data missing situation in the time dimension:
[0079] At a certain moment, all variables are missing, that is, the missing row vector, such as Figure 2 As shown in the middle, it is expressed as follows:
[0080]
[0081] S103, Data missing situation in the feature dimension:
[0082] A certain feature value is missing in the process data, that is, the missing column vector, such as Figure 2 As shown on the right, it is described as follows:
[0083]
[0084] Among them, NaN in the above formula represents missing data, N is the number of samples, and M is the number of feature dimensions.
[0085] Therefore, according to the MSWI process data record table, it should be judged whether there are missing data in the random distribution, time dimension and feature dimension, so as to facilitate filling according to different types. Such as Figure 2 As shown, the process data of MSWI is actually a mixed missing, so the process data should be processed for random and time dimension missing first, and then for feature dimension missing.
[0086] Among them, in step S2, first based on the expert experience rules, the missing data in the random distribution and time dimension are filled, and then a method based on the integration of reduced feature multi-models is used to fill the missing data in the feature dimension to obtain the filled data set, which specifically includes:
[0087] S201, First, based on the expert experience rules, the missing data in the random distribution and time dimension are filled, and the results are as follows:
[0088]
[0089] Among them, is the data without missing values containing M - M′ features, (M′ << M) is the data to be filled.
[0090] S202, The filling of the missing data in the feature dimension takes as the input, and the filling value is obtained by adopting the integration strategy of reduced feature multi-models It includes three parts: reducing features based on mutual information (MI), predicting and filling values based on multiple sub-models, and filling values based on BLR integration, to fill in the missing features Taking the filling of as an example, first obtain the correlation set of the feature to be filled after reducing features The dataset obtained after processing is as follows:
[0091]
[0092] Among them, is the feature to be filled, is 's correlation set.
[0093] Use it to construct sub-models based on the RF, GBDT, and BPNN algorithms, and record the corresponding predicted outputs as and Then, combine the above predicted values to construct a fusion model based on BLR, that is:
[0094]
[0095] Among them, is composed of the predicted output and the feature to be filled
[0096] Finally, take its output as the final filling value of the m'-th missing feature.
[0097] Among them, filling based on expert experience:
[0098] Expert experience is also known as prior knowledge, mechanism information, and domain knowledge, etc. In this paper, an expert experience filling method based on linear interpolation is mainly adopted.
[0099] In the MSWI process data, taking temperature and mixer water flow as examples, according to experience, it has linear characteristics within a certain range. For the missing data with random distribution and time dimension, use linear interpolation to fill it. Judge whether the data meets the condition that the time interval between the current missing value and the known upper and lower values is equal according to experience. If it is equal, take the average value as the filling value. If it is not equal, give the filling value based on experience.
[0100] Filling based on multi-model integration of reduced features:
[0101] First, based on data distribution and MI-reduced features:
[0102] For data missing in the feature dimension, feature selection is an important step in data filling preprocessing. The first problem to be solved is which known data to use for modeling to predict the filled values. The data distribution describes the frequency of occurrence of each value of the process variable. Due to the influence of many factors such as equipment parameters, operating conditions, and environmental interference, the data in different years of the MSWI process show different distributions. Therefore, it is necessary to select reasonable data according to the distribution similarity for predicting missing features. The present invention first uses histogram analysis to analyze the data distribution of each year, and then sets rules based on expert experience to select the data for establishing the missing value prediction model. Its general form can be written as:
[0103]
[0104] where N f is the year containing the missing column, N f ±n are the two years before and after the N f th year, and are the expected value and variance of the Gaussian distribution that the feature follows in the N f ±n year respectively.
[0105] The criterion of the present invention is to select the modeling data according to the trend presented by the change of the mean values in the previous and next two years. The specific steps are as follows: (1) If differs greatly from , the data in the N f ±2 year is not considered for modeling. Then, observe the similarity between and . If they are similar, select the data in the N f -1 year and N f +1 year for modeling; otherwise, select the data in the N f -1 year; (2) If does not differ much from , observe the change trends of to and to . If the trends are the same, that is, both increase or decrease, use the data of a total of 4 years before and after for modeling; if the trends are different, select the data in the N f -1 year and N f +1 year for modeling.
[0106] Since the high-dimensional feature space makes it difficult to mine the mapping rules between features and increases the difficulty of model training. Therefore, the present invention introduces the MI feature selection method, aiming to select the feature subset most relevant to the target filling feature. The following takes the missing feature A description will be given by taking as an example. MI is used to measure the degree of mutual dependence between two variables. MI feature selection ranks features by estimating the shared information content between features and the target variable, and selects features with high correlation. The steps are as follows:
[0107] Step(1): Set an empty set Calculate the MI value between feature and according to the following formula: and as follows:
[0108]
[0109] Step(2): Select the feature in X nomis that is most relevant to i.e., the feature with the largest MI value, as follows:
[0110]
[0111] At the same time, let and remove the feature in X nomis i.e.,
[0112] Step(3): Calculate the MI value between each feature in X nomis and and find the that satisfies the following conditions
[0113]
[0114] where β is a weight factor;
[0115] Step(4): Update and
[0116] Step(5): Repeat Step(3)-Step(4) until the MI value between each feature in X nomis and is obtained;
[0117] Step(6): Arrange in descending order according to the MI value with and obtain feature vectors in based on the set MI threshold. The dataset after feature reduction is as follows:
[0118]
[0119] Second, predict the filling value based on multiple sub-models
[0120] (1) Predict the filling value based on the RF sub-model
[0121] The RF uses a resampling method to generate a training subset to construct a parallel ensemble model based on a Decision Tree (DT), which has high tolerance for outliers and noise in the data.
[0122] First, the methods of bootstrap and RSM are used to perform J times of random sampling of samples and features. The process of generating the subset for the i-th time is as follows:
[0123]
[0124] where f RSM (·) is the subspace function, and f Bootstrap (·) is the random sampling function; M j represents the number of features selected for the sub-training set, usually
[0125] Next, divide the region where it is located to construct a DT. By traversing all samples and features, find the optimal variable number and the value of the splitting point The optimization is described as follows:
[0126]
[0127] where and represent a certain measurement value in the regions R 1 and R 2 ; and represent the average value of the measurement values in the regions R 1 and R 2 ; θ is the threshold of the number of samples included in the leaf node;
[0128] Use the following criterion to divide the region and determine the corresponding output value:
[0129]
[0130] Then, repeat the above steps for R 1 and R 2 until the number of samples in the leaf node is less than the set threshold θ. Accordingly, the input space is divided into R V regions, and the obtained DT model is denoted as the following formula:
[0131]
[0132]
[0133] where Rv The region contains samples ; Denote as R v the nth true value in the jth subset within the region; when it exists, I(·) = 1, otherwise I(·) = 0.
[0134] Finally, repeat the above process J times on to obtain the RF sub - model based on the following formula
[0135]
[0136] Correspondingly, the predicted value of the RF sub - model is as follows
[0137]
[0138] (2) Predict the filling value based on the GBDT sub - model
[0139] Different from the parallel integration mode of RF, GBDT adopts a serial integration mode with DT as the base learner. Its idea is to fit the residual of the previous round of base learner through the negative gradient of the loss function
[0140] First, after constructing the DT model by using as described above, set the initial constant model Let the value of the negative gradient of the loss function at the jth DT model be r jn , as follows
[0141]
[0142] where L(·) is the loss function is the (j - 1)th DT model
[0143]
[0144] Then, learn a new DT model jn by fitting the residual r and update the jth DT model as follows
[0145]
[0146] Finally, the GBDT sub - model is obtained by stacking the DTs obtained J times, as follows
[0147]
[0148] Correspondingly, the predicted value of the GBDT sub - model is as follows
[0149]
[0150] (3) Predicting the filling value based on the BPNN sub-model
[0151] Different from the above two DT-based modeling algorithms, BPNN is a multi-layer feedforward network trained according to the error backpropagation algorithm. A typical BPNN includes an input layer, a hidden layer, and an output layer. Its learning starts from the input layer, and the error corrects the weights and thresholds of each layer by the gradient descent method. Through the repeated forward propagation of information and backpropagation of errors, the weights and thresholds of each layer are continuously adjusted. When the output error is reduced to the desired degree or reaches the preset number of iterations, the network completes learning.
[0152] Here, the input vector is The output of the hidden layer is
[0153]
[0154] The output layer uses a linear activation function:
[0155]
[0156] Among them, ω 1 and ω 2 are the weights of the hidden layer and the output layer respectively, b 1 and b 2 are the biases of the hidden layer and the output layer respectively, is the activation function of the hidden layer, F BPNN (·) is the activation function of the output layer
[0157] Third, integrating the filling value based on BLR
[0158] As can be seen from the above, BPNN is a learning machine in the neural network mode, and RF and GBDT are parallel and serial ensemble learning machines based on the non-neural network mode, with differences from each other. BLR regards the model parameters as random variables and calculates their posterior probabilities through the prior probabilities of these parameters (weight coefficients), and has good reasoning ability, which is suitable for fusing the prediction results of these different sub-models in this paper.
[0159] Denote the input of the BLR model as The design matrix of is as follows:
[0160]
[0161] Furthermore, denote the prediction model of BLR as:
[0162]
[0163] Among them, w is the weight matrix.
[0164] Solve the posterior distribution probability of w as follows:
[0165]
[0166] Among them, the prior probability P(w) = N(0, σ P ) follows a Gaussian distribution (σ P is set to 1).
[0167] Maximum likelihood Calculate as follows:
[0168]
[0169] That is
[0170] Furthermore, the posterior distribution of w can be obtained as follows:
[0171]
[0172] According to the linear term and quadratic term of w in the above formula, the mean and variance in w ~ N(μ w , σ w ) are as follows:
[0173]
[0174] Among them,
[0175]
[0176] Finally, the calculation of the integrated filling value is as follows:
[0177]
[0178] Perform the above filling process on all M' missing values to obtain all filling values
[0179] Among them, in step S3, establishing a DFR-clfc model based on the filled dataset to obtain the DXN emission concentration prediction value specifically includes:
[0180] S301, for the input layer forest model, use Bootstrap and RSM to randomly sample the training set D fill , and construct I sub-forest models based on J DTs Correspondingly, the predicted mean of the i-th sub-forest model is calculated by the following formula:
[0181]
[0182] Among them, is the predicted value of the i-th sub-forest model, is the predicted mean of J DTs in the sub-forest model.
[0183] S302, then, reselect k kNN number of nearest neighbor values to form the regression vector of the i-th sub-forest model in the input layer After repeating the above steps I times, the layer regression vector of the input layer forest model is obtained Correspondingly, the input of the middle layer forest model is:
[0184]
[0185] where, is the enhanced regression vector of the input layer forest model, f FeaCom (·) is the combination function, is the layer regression vector of the input layer forest model, k kNN is the number of nearest neighbor values selected, X fill is the process variable in D fill in the training set;
[0186] The middle layer forest model contains L - 2 layers of sub-forest models, and the input of its λ (λ = 2, 3, …, L - 1) layer is expressed as:
[0187]
[0188] where, is the output feature vector of the λ - 1 layer, y DXN is the emission concentration of DXN, N is the number of samples, M λ = M + (k kNN × I) × (λ - 1) is the number of features;
[0189] The enhanced regression vector of each layer is obtained by cross-layer fully connected method, and is expressed as:
[0190]
[0191] where, The acquisition method is the same as that of the input layer;
[0192] Therefore, the process of generating the output feature vector by the λ layer forest model is as follows:
[0193]
[0194] where, is the layer regression vector of the λ layer forest model;
[0195] Furthermore, the input of the L-th layer output layer forest model is expressed as:
[0196]
[0197] Among them, M L = M + (k kNN × I) × (L - 1) represents the number of features;
[0198] Based on D fill,L Construct I sub - forest models containing J DTs Among them, the predicted value vector of the i - th sub - forest model is is the predicted value obtained from the J DTs of the i - th sub - forest model in the L - th layer, and the average predicted value of the i - th sub - forest model is
[0199] Finally, the predicted output of the DFR - clfc model is:
[0200]
[0201] Experimental verification based on the industrial MSWI dataset:
[0202] 1. Data description
[0203] This experiment aims to construct a DXN soft - sensing model, mainly studying the process data of the 2# furnace of a certain MSWI power plant in Beijing. The dataset processed covers the range from 2009 to 2020, with a time scale of hours and contains 116 features. Among them, there are missing data in both random distribution and time dimension in each year, mainly including individual speed variables, temperature variables, lime silo feed rate, activated carbon silo feed rate, and mixer water flow rate; there are missing values of the flue gas oxygen concentration at the NID system inlet and the main steam flow rate between 2014 and 2015.
[0204] In this paper, the modeling dataset after data analysis based on expert experience is divided into a training set and a validation set, and the performance of the proposed reduced - feature multi - model integration filling model is evaluated using the validation set; at the same time, the dataset to be filled is used as a test set to test the performance of the DXN soft - sensing model constructed based on the filled data.
[0205] 2. Evaluation metrics
[0206] Using the dataset to evaluate the model performance, N var is the number of samples of, the root mean square error (RMSE), the mean absolute error (MAE), and the goodness of fit (R 2 ), and the calculation formulas are as follows:
[0207]
[0208]
[0209]
[0210] Among them, mod represents the RF, GBDT, and BPNN sub-models. is the predicted value of the sub-model mod on the nth var sample.
[0211] 3. Verification Results
[0212] Taking the filling of the flue gas oxygen concentration value and the main steam flow value at the inlet of the NID system as examples, the data filling results based on expert experience, the filling results based on the integration of reduced feature multi-models, and the soft sensor modeling results are respectively shown.
[0213] Data filling results based on expert experience:
[0214] The data in 2017 has random distribution and missing time dimensions. Linear interpolation is used to fill the missing values. Figure 3a and 3b are the filling results of the temperature on the right at the top of the grate in the burnout section and the water flow B of the mixer.
[0215] From Figure 3a and 3b it can be seen that in the records of the temperature on the right at the top of the grate in the burnout section on a certain day, the recorded values at the 5th and 18th hours are missing, while the recorded values of the water flow B of the mixer are missing at the 7th and 21st hours. According to linear interpolation, the average value of the data before and after is used as the filling value.
[0216] Filling results based on the integration of reduced feature multi-models:
[0217] The data collected in 2014 and 2015 has the situation of missing flue gas oxygen concentration values at the inlet of the NID system. Appropriate data modeling needs to be selected to predict the missing values. Figure 4 is the group diagram of the distribution of the flue gas oxygen concentration data at the inlet of the NID system before and after the missing years.
[0218] From Figure 4The main concentration ranges of oxygen concentration for each year can be obtained, such as: the oxygen concentration in 2012 and 2017 is concentrated between [7.65, 8.2], the main one in 2013 is between [5.88, 6.65], and the most in 2015 and 2016 are between [6.4, 7.5] and [7.5, 8.6]. According to expert experience, the change trend of oxygen concentration from 2012 to 2017 shows a trend of first decreasing and then increasing. The missing data in 2014 and 2015 years and the data in 2013 have a high similarity in distribution. Therefore, it is reasonable to select the data in 2013 for modeling.
[0219] At the same time, there are missing values of main steam flow rate in 2014 and 2015. Modeling data is selected by analyzing the data distribution of adjacent years. Figure 5 It is the distribution of main steam flow data before and after the missing years.
[0220] From Figure 5 It can be seen that the main steam flow concentration ranges of 2013, 2016 and 2017 are similar, all being [75.65 - 78.25], which is higher than the main concentration value range of 2012 [72.5 - 74.25]. According to expert experience, the change trend of the missing main steam flow rate from 2014 to 2015 is similar to that of 2013, 2016 and 2017. It is reasonable to use the MSWI process data of the above three years for modeling.
[0221] According to the above description, a feature set with high correlation with the oxygen concentration and main steam flow rate at the inlet of the NID system is selected. The features with MI value greater than the threshold of 0.7 are used as the high - correlation feature set of the oxygen concentration value at the inlet of the NID system, and the feature information is shown in Table 1.
[0222] Table 1 High - correlation feature set of oxygen concentration value at the inlet of the NID system
[0223]
[0224] The features with MI value greater than the threshold of 0.75 are used as the high - correlation feature set of the main steam flow rate, and the feature information is shown in Table 2:
[0225] Table 2 High - correlation feature set of main steam flow rate
[0226]
[0227]
[0228] The filling value prediction of the sub - model is carried out according to the method described above. Table 3 and Table 4 respectively record the multi - model evaluation results of the oxygen concentration and main steam flow rate at the inlet of the NID system.
[0229] Table 3 Multi-model evaluation results of the oxygen concentration in the flue gas at the inlet of the NID system
[0230]
[0231] As can be seen from Table 3, among multiple sub-models, the GBDT has the highest prediction accuracy, and the MAE value on the validation dataset is small, indicating that the generalization performance of this model is good; the RF can obtain a better RMSE value on the validation set; the BPNN has the best RMSE and R on the training set 2 which is the best, indicating that its fitting effect on the training set is the best, but its MAE value is poor. Based on the results of the three groups, the final predicted value is obtained by using the method in this paper. The results show that all indicators on the training set and the validation set are the best, indicating that it is effective to fill the characteristics of the oxygen concentration in the flue gas at the inlet of the NID system by using this method.
[0232] Table 4 Multi-model evaluation results of the main steam flow rate
[0233]
[0234] As can be seen from Table 4, among multiple sub-models for predicting the main steam flow rate, all evaluation indicators of the BPNN are the best; while the GBDT has a slightly better MAE value than the RF on the training set and the validation set, but its RMSE is not as good as that of the RF. The results of using the method in this paper based on the three groups of results show that all indicators on the training set and the validation set are the best.
[0235] Results of the soft sensor modeling module based on the filled data:
[0236] To verify whether the filled MSWI process data can effectively predict the DXN concentration, a DFR-clfc model was established. Table 5 shows the results of the modeling experiment. The results show that the model established based on the filled MSWI process data can effectively predict the DXN concentration. The parameters involved in this model are as follows: the minimum number of samples is 9, 11 features are selected, and the decision tree is set to 500.
[0237] Table 5 Modeling experiment results based on the filled data
[0238]
[0239]
[0240] As can be seen from the above results, the Prec performance of the modeling results based on the data filled by the method of the present invention is not as good as that of the sub-models on the validation set and the test set, but they all have the best RMSE, MAE and R 2 indicator values, indicating that the data filled by the method of the present invention can effectively predict the DXN emission concentration and can meet the monitoring requirements of pollutant emissions.
[0241] The soft measurement method for dioxin emissions in the MSWI process based on missing data filling provided by the present invention first identifies process data as three types according to the missing situation, namely, randomly distributed, missing in the time dimension, and missing in the feature dimension. Then, after filling the first two types based on expert experience rules, a filling model based on reduced feature multi-model integration is constructed for the missing data in the feature dimension. The data distribution analysis and the Mutual Information (MI) algorithm are used to select relevant process variables as inputs for the missing features, and sub-models of Random Forest (RF), Gradient Boosting Decision Tree (GBDT), and Backpropagation neural network (BPNN) with complementary characteristics are established to preliminarily predict the missing values. Then, the Bayesian Linear Regression (BLR) is used to integrate the above predicted values to obtain the final filling value. Finally, a soft measurement model for dioxin (DXN) emission concentration based on Deep forest regression based on cross-layer full connection (DFR-clfc) is established using the filled MSWI data to verify the filling effect.
[0242] In this article, specific examples are used to elaborate on the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A soft measurement method for dioxin emissions in the MSWI process based on missing data filling, characterized in that, it specifically includes the following steps: S1, Missing type identification: Classify according to different situations of missing data in the original MSWI process data, and identify the process data as three types: random distribution, time dimension, and feature dimension missing; S2, Missing data filling: First, based on expert experience rules, fill in the missing data of random distribution and time dimension, and then use the method based on reduced feature multi-model integration to fill in the missing data of feature dimension to obtain a filled data set; S3, Soft measurement modeling based on filled data: Establish a DFR-clfc model based on the filled data set to obtain the predicted value of DXN emission concentration; The specific steps of S2 include: S201, First, based on expert experience rules, fill in the missing data of random distribution and time dimension, and the results are as follows: Among them, is data without missing values containing M - M' features, (M' << M) is the data to be filled; S202, the filling of the missing data in the feature dimension is based on as the input, and the filling value is obtained by adopting a multi-model integration strategy based on the reduced features It includes three parts: reducing features based on mutual information MI, predicting filling values based on multiple sub-models, and integrating filling values based on BLR, so as to take the filling of the missing features as an example. First obtain the correlation set of the features to be filled through feature reduction The data set obtained after processing is as follows: Among them, is the feature to be filled, is the relevance set of It is used to construct sub-models based on RF, GBDT, and BPNN algorithms, and the corresponding predicted outputs are denoted as and Next, the above predicted values are combined to construct a fusion model based on BLR, that is: Among them, is composed of a predicted output and a feature to be filled ; Finally, output it as the final filling value for the m'-th missing feature.
2. The soft measurement method for dioxin emissions in the MSWI process based on missing data filling according to claim 1, characterized in that, In step S1, classify according to different situations of missing data in the original MSWI process data, and identify the process data as three types: random distribution, time dimension, and feature dimension missing. Specifically, it includes: S101, The situation of missing data in random distribution: The random occurrence of missing data has nothing to do with its own characteristics, the values of other characteristics, and time. The missing data is randomly distributed among different characteristics. The situation of missing data in random distribution can be expressed as follows: S102, The situation of missing data in time dimension: At a certain moment, all variables are missing, that is, a missing row vector, which is expressed as follows: S103, The situation of missing data in feature dimension: A certain feature value in the process data is missing, that is, a missing column vector, which is described as follows: Among them, NaN in the above formula represents missing data, N is the number of samples, and M is the number of feature dimensions.
3. The soft measurement method for dioxin emissions in the MSWI process based on missing data filling according to claim 1, characterized in that, In step S3, establish a DFR-clfc model based on the filled data set to obtain the predicted value of DXN emission concentration. Specifically, it includes: S301. For the input layer forest model, use Bootstrap and RSM to randomly sample the training set D fill , and construct I sub-forest models based on J DTs Correspondingly, the predicted mean of the i-th sub-forest model is calculated by the following formula: Among them, is the predicted value of the i-th sub-forest model, is the predicted mean of J DTs in this sub-forest model; S302, reselect k through the kNN method kNN ones of the nearest neighbor values to form the regression vector of the i-th sub-forest model in the input layer Repeat the operation of reselecting k through the kNN method for I times kNN ones of the nearest neighbor values to form the regression vector of the i-th sub-forest model in the input layer After that, the layer regression vector of the input layer forest model is obtained Correspondingly, the input of the middle layer forest model is as follows: Among them, is the enhanced regression vector of the input layer forest model, and f FeaCom (·) is the combination function, is the layer regression vector of the input layer forest model, and k kNN is the number of nearest neighbor values selected, X fill is the process variable in the training set D fill ; The intermediate layer forest model contains L - 2 layer sub-forest models, and the input of its λ (λ = 2, 3,..., L - 1) layer is expressed as: Among them, is the output feature vector of the (λ - 1)-th layer, y DXN is the emission concentration of DXN, N is the number of samples, M λ = M + (k kNN × I) × (λ - 1) is the number of features; Enhanced regression vector for each layer Obtained by using a cross-layer fully connected method, expressed as: Among them, The acquisition method is the same as that of the input layer; Therefore, the process of generating the output feature vector by the λ layer forest model is as follows: Among them, is the layer regression vector of the λ-th layer forest model; Furthermore, the input of the L layer output layer forest model is expressed as: Among them, M L = M + (k kNN × I) × (L - 1) represents the number of features; Based on D fill,L Construct I sub - forest models containing J DTs Among them, the prediction value vector of the i - th sub - forest model is is the prediction value obtained from the J DTs of the i - th sub - forest model in the L - th layer. The average prediction value of the i - th sub - forest model is Finally, the output of the DFR-clfc model is the predicted value of DXN emission concentration