On-line soft measurement method for dioxin emission concentration in MSWI process
By constructing a typical sample pool using the K-means weighted algorithm and principal component analysis, and combining it with the FTBL model, the problem of real-time monitoring of dioxin emission concentration during MSWI was solved, achieving efficient and accurate prediction of dioxin emission concentration and rapid model updates.
Patent Information
- Application Number
- CN202211651114.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-12-21
AI Technical Summary
Existing technologies make it difficult to achieve real-time monitoring of dioxin emission concentrations during MSWI, and traditional detection methods are time-consuming and costly, making it difficult for models to accurately reflect changes in the incinerator's operating status.
A typical sample pool was constructed using the K-means weighted algorithm and principal component analysis, and online and offline predictions were performed using the FTBL model. Through a combination of feature mapping, enhancement, and incremental layers, soft measurement of dioxin emission concentration was achieved.
It improves the accuracy of dioxin emission concentration prediction, reduces detection costs, and enables real-time monitoring and rapid model updates for MSWI processes.
Smart Images

Figure CN116110506B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of pollutant monitoring, in particular to an on-line soft measurement method for dioxin emission concentration in MSWI process. BACKGROUND
[0002] Dioxin (DXN) is a persistent organic pollutant generated in the process of municipal solid waste incineration (MSWI), which is an important environmental indicator that needs to be controlled to achieve the lowest emission. However, it is difficult to monitor in real time due to factors such as detection technology, economy and labor cost.
[0003] Municipal solid waste incineration is one of the main technologies for waste power generation and has been widely used worldwide. As a typical industrial process for the harmless, reduction and resource utilization of municipal solid waste (MSW), although the advantages of MSWI technology outweigh the disadvantages, the toxic and harmful substances contained in the exhaust gas have always been the focus of public attention. Dioxin is currently the most toxic persistent organic pollutant known to man and the environment, and has become one of the factors restricting the development of MSWI technology. Due to the complex types of DXN compounds, the complex analysis and testing steps, the long measurement delay and the high price, it is difficult to measure the DXN emission concentration in real time. DXN is an important environmental indicator for intelligent optimization control of MSWI process. Therefore, the application of soft measurement technology to solve the high economic and labor cost of current DXN offline detection is an effective means to achieve ultra-low emission of DXN.
[0004] Current soft measurement technology mainly includes mechanism model and data-driven model strategies. Since the generation, adsorption and emission mechanism of DXN is not clear, mechanism model has not yet appeared. Data-driven models based on easily measured process variables are widely used due to their high efficiency, low cost and easy implementation. Therefore, the present embodiment studies a soft measurement method for DXN emission based on MSWI process data.
[0005] Due to the complexity of incineration process mechanism, the fluctuation of MSW raw material composition and the randomness of artificial control by field experts, the working condition drift phenomenon of MSWI process frequently occurs. In addition, it is difficult to establish a reliable soft measurement model of DXN emission concentration for actual engineering application. Therefore, the online soft measurement of DXN emission needs to solve the following problems: in the MSWI process, the fluctuation of sensor data such as temperature, pressure and flow measured by process variables is the main basis for representing the change of running condition, and how to identify the drift change of running state according to the process data to assist the realization of accurate detection and model updating is one of the challenges in the online application stage of soft measurement model; how to maintain the continuous and rapid learning ability of new data (drift data) in the process of online measurement of DXN emission concentration is a key problem that needs to be solved in the practical engineering application of soft measurement technology; making full use of historical data for offline modeling is the first step to realize online soft measurement, and how to select historical data to build an offline model with low search cost and maintain its best performance is the primary problem that needs to be solved in DXN soft measurement modeling. SUMMARY
[0006] In order to overcome the shortcomings of the prior art, the purpose of the present application is to provide an online soft measurement method for dioxin emission concentration of MSWI process.
[0007] In order to achieve the above purpose, the present application provides the following scheme:
[0008] An online soft measurement method for dioxin emission concentration of MSWI process, comprising:
[0009] Based on the K-means weighted algorithm, the process data of the typical sample pool is determined according to the historical process data set of MSWI;
[0010] The principal component analysis is carried out according to the process data of the typical sample pool, and the drift index control limit reflecting whether the MSWI process changes is obtained;
[0011] An offline model based on FTBL is constructed, and the process data of the typical sample pool and the historical DXN true value data of MSWI are input into the offline model for prediction calculation to obtain offline calculation results; the offline model comprises a feature mapping layer, an enhancement layer and an incremental layer;
[0012] According to the obtained online data, principal component analysis is performed, and whether the online data is drift data or normal data is judged according to the drift index control limit, if it is the normal data, jump to the step of 'constructing an offline model based on FTBL, and inputting the process data of the typical sample pool and the historical DXN true value data of MSWI into the offline model for prediction calculation to obtain the calculation result'; if it is the drift data, constructing an online model based on FTBL, and inputting the process data of the typical sample pool, the drift data and the output data of the incremental layer of the offline model into the online model for prediction calculation to obtain the online calculation result; the online model comprises an online incremental layer;
[0013] According to the offline calculation result and the online calculation result, the DXN emission concentration prediction value is determined.
[0014] Preferably, according to the historical process data set of MSWI, the process data of the typical sample pool is determined based on the K-means weighted algorithm, which comprises:
[0015] Obtaining a historical process data set X of MSWI His ;
[0016] According to the historical process data set X His , historical data is obtained , wherein x n is the nth sample, y n is the prediction value corresponding to the nth sample; N is the sample number of the historical data set of MSWI, and M is the feature number of the historical data set of MSWI;
[0017] Randomly selecting I instances as initial centroids
[0018] According to the weighted Euclidean distance between the samples and the centroids, all samples are classified into I classes:
[0019]
[0020] , wherein C i represents the ith class; represents the weight vector of the process variable, wherein, , wherein H(·) represents the information entropy of a random variable, x m is the mth feature vector, y represents the DXN concentration, x n,m represents the mth feature value of the nth sample, p(x n,m )p(y n ) represents the marginal probability distribution, and p(x n,m , y n ) is the joint probability distribution;
[0021] Update centroid C using inter-class samples i :
[0022] in, This represents the number of samples in the i-th cluster;
[0023] The centroids are updated cyclically, and all centroids are obtained through preset conditions, which are expressed as follows: Where, δ TS For the evaluation index R iter The threshold, iter, represents the number of iterations, and the formula for calculating the metric is:
[0024] The TSP (True Strength Point) is established by minimizing cluster similarity, and the formula is as follows:
[0025] Among them, R DB This is a cluster similarity measure index; where, Among them, S i M represents the sum of distances for the i-th class. ij This represents the Minkowski metric criterion.
[0026] Preferably, principal component analysis is performed based on the process data from the typical sample pool to obtain drift index control limits reflecting whether the MSWI process has changed, including:
[0027] The correlation coefficient matrix of TSP data Represented as Where, N TSP For TSP data D TSP The number of elements; R is the correlation coefficient matrix of the TSP data;
[0028] Perform singular value decomposition on R and calculate its eigenvalues; the formula is R = U. M×M Σ M×M [V M×M ] T Among them, U M×M and V M×M Denotes an orthogonal matrix, Σ M×M It is an M-dimensional diagonal matrix;
[0029] Using the cumulative contribution rate η and the PCA contribution threshold δ PCA Dimensional reduction:
[0030] Among them, P PCA P represents the number of selected principal components. PCA Less than M;
[0031] Rewrite the calculation formula as follows: in, is the loading matrix;
[0032] According to the score matrix T and the loading matrix X TPS is represented as:
[0033] where, X TSP is represented as: TSP is the projection of X on the residual space; the satisfies the orthogonal relationship;
[0034] In addition, satisfies the orthogonal relationship, which is proved as follows:
[0035]
[0036] The drift index control limit is represented as:
[0037]
[0038]
[0039] where, is the control limit of Hotelling's T 2 ; SPE CL is the control limit of SPE, P PCA is the number of selected principal components F α (P PCA , N TSP -P PCA ) represents the F distribution with degrees of freedom P PCA and (N TSP -P PCA ); c α represents the normal deviation not exceeding (1-α); and h0are calculated as shown below:
[0040]
[0041]
[0042] where, h0、 and are intermediate variables for the calculation of the SPE control limit, σ m is the eigenvalue of the singular value decomposition.
[0043] Preferably, the construction of an offline model based on FTBL, and the input of the process data of the typical sample pool and the historical DXN ground truth data of MSWI into the offline model for prediction calculation, to obtain the offline calculation results, includes:
[0044] For a given TSP data In D TSP In the process, a node splitting function is defined by randomly selecting an eigenvalue x. It is represented as: n∈(1,N TSP )and m∈(1,M); where, The sign function is `rand(·)`, and the random number generator function is `rand(·)`, where `n` and `m` do not take maximum or minimum values.
[0045] For TS fuzzy inference, K fuzzy rules are determined, and the k-th rule can be expressed as: in,
[0046] Among them, R k For the k-th fuzzy rule, c k,m and σ k,m Representing Gaussian functions respectively Center and width, t leaf Indicates the t-th leaf leaf nodes; It is a Gaussian function;
[0047] Based on the above K fuzzy rules, the result of FDT can be described as follows:
[0048] in, Where f(·) is the FDT model, and g k (·) indicates the antecedent and consequent parts of TS fuzzy inference. Represents the weight of the consequent features;
[0049] The gradient descent method is applied to update the parameters of the FDT model f(·) during training. These parameters include the center c. k Width σ k and weight The output of the feature mapping layer is represented as follows:
[0050] in, in, It is the nth FM An FDT model is input X TSP The output.
[0051] Enhancement layer with The output of the enhancement layer is represented as input
[0052] The output of the feature mapping layer and the enhancement layer is
[0053] The weight between the predicted output and the input is calculated by using a ridge regression learning algorithm As follows:
[0054] Wherein, is a pseudo-inverse matrix, λ is a regularization coefficient, and I is a unit matrix;
[0055] The FDT model is added in the incremental layer, and the pseudo-inverse matrix is dynamically updated, so that as input, as output; the pseudo-inverse matrix updating process of the incremental process is as follows:
[0056] Wherein, D, H k+1 , B T , C are intermediate variables of the pseudo-inverse matrix updating process;
[0057] The new weight matrix is represented as:
[0058]
[0059] The prediction calculation process of the offline model FTBL is as follows:
[0060]
[0061] Preferably, the principal component analysis is performed according to the obtained online data, and whether the online data is drift data or normal data is judged according to the drift index control limit, if the normal data is obtained, the step of “constructing an offline model based on FTBL, and inputting the process data of the typical sample pool and the historical DXN true value data of MSWI into the offline model for prediction calculation to obtain the calculation result” is jumped to; if the drift data is obtained, an online model based on FTBL is constructed, and the process data of the typical sample pool, the drift data and the output data of the incremental layer of the offline model are input into the online model for prediction calculation to obtain online calculation results, including:
[0062] The drift value in the new window is calculated based on the following formula:
[0063]
[0064] Wherein, and is the (N TSP +1)th process data statistical index of the statistical index of the (N represents a new load matrix, represents a new diagonal matrix;
[0065] determine whether the sample is a drift sample or a normal sample by a judgment formula; the judgment formula is:
[0066]
[0067] For normal samples, reuse the offline model of the FTBL to perform soft measurement of the DXN concentration, which can be expressed as:
[0068]
[0069] For drift samples, the soft measurement value calculation method is:
[0070] wherein, ε Offset is an offset value of the offline FTBL prediction output, and is as follows:
[0071] wherein, N t represents the total amount of data arriving at t, represents all prediction values arriving at t, represents the mathematical expectation of the vector , and x IncTem represents the temperature of the incinerator at t.
[0072] When the detection true value exists, input the TSP data, the drift data and the output of the increment layer into the online model; the prediction value of the online model is:
[0073] wherein, represents a weight matrix, is the output matrix of the N FM +N En +N In +N OI FDT.
[0074] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0075] The application provides an online soft measurement method for dioxin emission concentration of MSWI process, and has the characteristics that: based on K-means weighted algorithm, process data of a typical sample pool is determined according to historical process data set of MSWI; principal component analysis is performed according to process data of the typical sample pool, and a drift index control limit reflecting whether the MSWI process changes is obtained; an offline model based on FTBL is constructed, and process data of the typical sample pool and historical DXN true value data of MSWI are input into the offline model for predictive calculation, and offline calculation results are obtained; the offline model comprises a feature mapping layer, an enhancement layer and an incremental layer; principal component analysis is performed according to obtained online data, and whether the online data is drift data or normal data is judged according to the drift index control limit; if it is the normal data, it is jumped to the step of constructing the offline model based on FTBL, and the process data of the typical sample pool and the historical DXN true value data of MSWI are input into the offline model for predictive calculation, and calculation results are obtained; if it is the drift data, an online model based on FTBL is constructed, and the process data of the typical sample pool, the drift data and output data of the incremental layer of the offline model are input into the online model for predictive calculation, and online calculation results are obtained; the online model comprises an online incremental layer; and DXN emission concentration prediction value is determined according to the offline calculation results and the online calculation results. The application can effectively improve the accuracy of the DXN emission concentration prediction value. BRIEF DESCRIPTION OF DRAWINGS
[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0077] Figure 1 MSWI process and DXN emission diagram provided for the embodiments of the present application;
[0078] Figure 2 DXN concentration measurement process diagram provided for the embodiments of the present application;
[0079] Figure 3 Method flowchart provided for the embodiments of the present application;
[0080] Figure 4 DXN concentration soft measurement strategy diagram provided for the embodiments of the present application;
[0081] Figure 5 Fuzzy decision tree diagram provided for the embodiments of the present application;
[0082] Figure 6 A first schematic diagram of DXN data provided for an embodiment of the present application;
[0083] Figure 7 A second schematic diagram of DXN data provided for an embodiment of the present application;
[0084] Figure 8 A schematic diagram of training data fitting curves of different methods provided for an embodiment of the present application;
[0085] Figure 9 A schematic diagram of test data fitting curves of different methods provided for an embodiment of the present application;
[0086] Figure 10 A schematic diagram of three-dimensional curves of a typical sample provided for an embodiment of the present application;
[0087] Figure 11 A schematic diagram of two-dimensional curves of a typical sample provided for an embodiment of the present application;
[0088] Figure 12 A schematic diagram of three-dimensional curves of a typical sample provided for an embodiment of the present application after deleting redundant samples;
[0089] Figure 13 A schematic diagram of T 2 curves provided for an embodiment of the present application;
[0090] Figure 14 A schematic diagram of SPE curves provided for an embodiment of the present application;
[0091] Figure 15 A schematic diagram of prediction results in an offline stage provided for an embodiment of the present application
[0092] Figure 16 A schematic diagram of prediction results in an online stage provided for an embodiment of the present application
[0093] Figure 17 A schematic diagram of laboratory simulation application test provided for an embodiment of the present application;
[0094] Figure 18 A schematic diagram of industrial field online application test provided for an embodiment of the present application. DETAILED DESCRIPTION
[0095] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0096] Reference to "an embodiment" in this disclosure means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that any of the embodiments described in this disclosure can be incorporated in a combination of embodiments.
[0097] The terms "first", "second", "third", and "fourth" and the like in the description and in the claims of the application and the accompanying drawings, are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. Moreover, the terms "include", and "have", and any variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, product, or apparatus that comprises a list of steps, processes, methods, or components does not necessarily comprise only those listed steps, processes, methods, or components but can optionally include additional steps, processes, methods, or components not expressly listed or inherent to such process, method, product, or apparatus.
[0098] The purpose of the present application is to provide an online soft measurement method for MSWI process dioxin emission concentration, which can effectively improve the accuracy of DXN emission concentration prediction value.
[0099] In order to make the above-mentioned purpose, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0100] As shown in Figure 1 , the present embodiment divides the MSWI process into five stages, in turn:
[0101] (1) MSW is sent into the hopper by mechanical grab bucket, and then is transported to the incinerator by moving grate;
[0102] (2) In the high temperature environment above 850℃, MSW goes through drying, pyrolysis and gasification stages on the grate to form ash;
[0103] (3) The heat energy of high temperature flue gas is converted into high temperature and high pressure steam through heat exchange devices such as water cooled wall, superheater, evaporator, coal economizer and steam drum, and then the steam is used to drive steam turbine generator to realize waste-to-energy;
[0104] (4) Selective non-catalytic reduction (SNCR) system, activated carbon adsorption, semi-dry deacidification, bag type dust collector and other devices are used to purify the flue gas generated in the MSWI process, such as CO, HCl, SO2, NOx, particulate matter and DXN;
[0105] (5) The flue gas meeting the pollutant control emission standard (GB18485-2014) is discharged into the atmosphere through a 80-meter-high chimney.
[0106] Studies have shown that MSW contains trace amounts of DXN. DXN goes through high-temperature decomposition, low-temperature de novo synthesis, activated carbon adsorption, and bag dust removal in turn during the combustion and purification process of MSW, and the remaining trace DXN is discharged into the atmosphere. In addition, due to the accumulation of fly ash in MSWI equipment, DXN has a significant "memory effect". Therefore, it is a challenging task to detect the DXN emission concentration by using soft measurement method. Generally, the detection position of the DXN emission concentration of the MSWI process is located at the entrance of the chimney (i.e. Figure 1 the position marked by the dot where the flue gas meets the chimney).
[0107] So far, the detection of DXN emission concentration mainly relies on manual field sampling and laboratory analysis. The detection process is as shown in Figure 2 .
[0108] In Figure 2 , the first stage is manual field sampling by an on-site engineer, which requires using a constant speed sampling device to continuously collect real-time flue gas in the pipeline for two hours to obtain a quartz filter (solid phase), a resin cylinder (gas phase), and a condensate (liquid phase), which are then packaged and sent to the laboratory for analysis; the second stage is sample analysis, which requires pretreatment, extraction, and Soxhlet extraction of the samples, and then placing them in designated reagent bottles, and then using high-resolution gas chromatography-high-resolution mass spectrometry (HRGC / HRMS) equipment to analyze the samples; the third stage is to analyze and calculate the HRGC / HRMS data to obtain a DXN concentration report containing 17 compounds.
[0109] The detection of DXN concentration has the disadvantages of long time consumption, high labor and material costs, and cannot continuously and timely reflect the DXN emission of the incinerator, which also leads to the characteristics of small sample size and high dimension of the modeling data based on the above experiments. In addition, due to unclear DXN mechanism and many human interference factors in the detection link, the soft measurement technology based on actual DXN data is difficult to achieve satisfactory detection accuracy. Therefore, it is a great challenge to research a DXN emission concentration soft measurement model with stable and high performance.
[0110] Figure 3 The MSWI process dioxin emission concentration online soft measurement method provided in this embodiment comprises:
[0111] Step 100: determining process data of a typical sample pool based on K-means weighted algorithm according to historical process data set of MSWI;
[0112] Step 200: performing principal component analysis according to the process data of the typical sample pool to obtain a drift index control limit reflecting whether the MSWI process changes;
[0113] Step 300: Construct an offline model based on FTBL, and input the process data of the typical sample pool and the historical DXN ground truth data of MSWI into the offline model for prediction calculation to obtain the offline calculation result; the offline model includes a feature mapping layer, an enhancement layer and an incremental layer;
[0114] Step 400: Perform principal component analysis on the acquired online data, and determine whether the online data is drifting data or normal data based on the drift index control limit. If it is normal data, proceed to step "Construct an offline model based on FTBL, and input the process data of the typical sample pool and the historical DXN ground truth data of MSWI into the offline model for prediction calculation to obtain the calculation result"; if it is drifting data, construct an online model based on FTBL, and input the process data of the typical sample pool, the drifting data, and the output data of the incremental layer of the offline model into the online model for prediction calculation to obtain the online calculation result; the online model includes an online incremental layer;
[0115] Step 500: Determine the predicted value of DXN emission concentration based on the offline calculation results and the online calculation results.
[0116] The soft measurement structure for DXN emission concentration proposed in this embodiment is as follows: Figure 4 As shown, the process includes two phases: offline and online. In the offline phase, TSP, FTBL algorithms, and PCA analysis are used to obtain an offline model and drift index control limits based on historical process data. In the online phase, a sliding window recursive PCA adaptive monitoring process is used to achieve drift detection, online measurement, and FTBL dynamic learning. Figure 4 The meanings of different symbols are shown in Table 1 below.
[0117] Table 1
[0118]
[0119]
[0120]
[0121]
[0122] This embodiment first constructs a typical sample pool acquisition module based on k-means. Historical data collected from complex industrial processes typically exhibits high dimensionality, strong correlation, and high redundancy, leading to an asymmetry between modeling data and valuable information. Therefore, this embodiment proposes a weighted k-means method for constructing a typical sample pool (TSP). Theoretically, TSP can achieve the same (or higher) modeling performance as the original dataset.
[0123] Based on historical data First, randomly select I instances as the initial centroids. Then, based on the weighted Euclidean distance between the sample and the centroid, all samples are classified into class I:
[0124]
[0125] Among them, C i Indicates the i-th class; x n Represents the nth sample; The weight vector representing the process variable is determined by the information values between the process variable and the DXN concentration, as shown below:
[0126]
[0127] Where H(·) represents the information entropy of the random variable, x m Let y be the m-th eigenvector, and x be the concentration of DXN. n,m Let p(x) represent the m-th feature value of the n-th sample. n,m )p(y n p(x) represents the marginal probability distribution. n,m, y n ) represents the joint probability distribution.
[0128] Update centroid C using inter-class samples i :
[0129]
[0130] in, This represents the number of samples in the i-th cluster.
[0131] By iteratively updating the centroids using equations (1)-(3), all centroids are finally obtained under the following conditions, which can be expressed as:
[0132]
[0133] Where, δ TS For the evaluation index R iter The threshold, iter, represents the number of iterations, and the metric is calculated as follows:
[0134]
[0135] As can be seen from the above formula, it represents the sum of Euclidean distances between clustered samples and centroids.
[0136] Then, by minimizing the cluster similarity, the TSP can be established as follows:
[0137]
[0138] where R DB is a Cluster Similarity Measure (CSM) index.
[0139]
[0140] where S i is the sum of distances of the ith cluster, M ij is the Minkowski metric criterion.
[0141] Secondly, the embodiment constructs a drift index calculation module based on PCA, and for high-dimensional process variables, the embodiment uses a principal component analysis (PCA) model to calculate a drift index control limit for determining whether the MSWI process has changed.
[0142] Firstly, the correlation coefficient matrix R of the TSP data is calculated as follows:
[0143]
[0144] where N TSP is the number of TSP data D TSP .
[0145] In order to obtain latent variables, i.e., principal components, which can represent the original high-dimensional process variables, the embodiment performs singular value decomposition (SVD) on R and calculates eigenvalues:
[0146] R = U M×M Σ M×M V M×M T (10)
[0147] where U M×M and V M×M are orthogonal matrices (V M×M = [U M×M ] T ), Σ M×M is an M-dimensional diagonal matrix (i.e., diag (σ1, σ2, …, σ M )), and σ m is an eigenvalue on the diagonal.
[0148] The embodiment uses the eigenvalue cumulative contribution rate η and the PCA contribution threshold δ PCA to reduce the dimension:
[0149]
[0150] where P PCA is the number of selected principal components, and P PCA is less than M.
[0151] Accordingly, by P PCA A drift control limit for monitoring whether the working condition has changed is determined.
[0152] Further, formula (10) is rewritten as:
[0153]
[0154] According to the score matrix T and the load matrix X TPS is expressed as:
[0155]
[0156] wherein, X TSP The projection on the principal component space, X TSP The projection on the residual space.
[0157] In addition, and satisfy the orthogonal relationship, which is proved as follows:
[0158]
[0159] Therefore, using and has the advantage that the two parts are independent of each other (i.e., the statistical information does not interfere with each other).
[0160] This embodiment uses the Hotelling’s T2 and SPE statistical indicators calculated according to the TSP as control limits for identifying working condition drift. The two statistical indicators can be expressed as:
[0161]
[0162]
[0163] wherein, F α (P PCA ,N TSP -P PCA ) represents the F distribution with degrees of freedom P PCA and (N TSP -P PCA ); c α represents the normal deviation not exceeding (1-α); and h0are calculated as shown below:
[0164]
[0165]
[0166] Again, the embodiment is constructed with an FTBL-based offline model construction module, and the FTBL offline modeling method is composed of a feature mapping layer, an enhancement layer, and an incremental layer. Figure 5 Compared with the traditional DT, the basic unit of each layer is replaced by FDT instead of neurons, and the structure of FDT is shown in Figure 5 FDT is a kind of binary tree, and the structure includes non-leaf nodes and leaf nodes, wherein: the former is used for feature selection, and the latter is used for TS fuzzy inference system.
[0167] 1) Feature mapping layer
[0168] For a given TSP data The embodiment randomly selects a feature value x in D TSP to define the node splitting function , which can be expressed as:
[0169]
[0170] Wherein, is a symbol function, and rand(·) is a random number generation function (n and m do not take the maximum and minimum values).
[0171] The embodiment can obtain T / 2-1 non-leaf nodes (i.e. ) and T / 2 leaf nodes constructed by formula (20). Since the paths from the root node to the leaf nodes are different, the input data of each leaf node is different. Therefore, the TS fuzzy inference process based on D Leaf is as follows.
[0172] K fuzzy rules are determined for the TS fuzzy inference, and the kth rule can be expressed as:
[0173]
[0174]
[0175] Wherein,
[0176]
[0177] Wherein, c k,m and σ k,m represent the center and width of the Gaussian function , and t leaf represents the t leaf th leaf node.
[0178] According to the above K fuzzy rules, the result of FDT can be described as:
[0179]
[0180] where,
[0181]
[0182]
[0183] where, f(·) is the FDT model, and g k (·) represents the antecedent and consequent parts of the TS fuzzy inference, represents the weight of the consequent part feature.
[0184] The present embodiment applies gradient descent method to update the parameters in the training process of the FDT model f(·), including the center c k , width σ k and weight
[0185] Therefore, the output of the feature mapping layer is represented as follows:
[0186]
[0187] where,
[0188]
[0189] where, is the output of the nth FM FDT model through the input X TSP
[0190] 2) Enhancement layer
[0191] The enhancement layer takes as input, and its output can be represented as:
[0192]
[0193] Therefore, the output of the feature mapping layer and the enhancement layer is The ridge regression learning algorithm is used to calculate the weight between the predicted output and the output of the feature mapping layer and the enhancement layer, as follows:
[0194]
[0195] where, is the pseudo-inverse matrix, λ is the regularization coefficient, and I is the identity matrix.
[0196] 3) Incremental layer
[0197] To obtain sufficiently good performance, the FDT model is further increased in the increment layer, and the pseudo-inverse matrix is dynamically updated.
[0198] As input, As output, the pseudo-inverse matrix updating process of the increment process is as follows:
[0199]
[0200] Wherein,
[0201]
[0202] Therefore, the new weight matrix Can be expressed as:
[0203]
[0204] The prediction calculation process of the offline model FTBL is as follows:
[0205]
[0206] Again, the embodiment also constructs an online drift identification and soft measurement module, which is based on the PCA model and drift index control limit established above to monitor the MSWI process in real time, and based on the drift identification result to output the corresponding soft measurement. The embodiment uses a fixed window size to obtain process data, and when the data fills the window, the drift identification and soft measurement are started.
[0207] The drift value in the new window is calculated based on the following formula:
[0208]
[0209]
[0210] Wherein, And The statistical index of the N TSP +1)th process data , represents the new load matrix, Represents the new diagonal matrix (i.e. ).
[0211] The samples in the window are defined as drift and normal samples, through the following conditions:
[0212]
[0213] For normal samples, reuse the FTBL offline model for DXN concentration soft measurement, which can be expressed as:
[0214]
[0215] For the drift sample, considering that the true value detection of DXN emission in actual engineering is time-consuming and it is difficult to obtain the true value data of DXN in real time, in addition to prompting the need for detection, the following soft measurement value calculation method is given:
[0216]
[0217] Wherein, ε Offset is the offset value of offline FTBL prediction output, as follows:
[0218]
[0219] Wherein, N t represents the total amount of data arriving at t time, represents all the prediction values arriving at t time, represents the mathematical expectation of the vector , x IncTem represents the incinerator temperature at t time.
[0220] In addition, when the true value exists, the following method is used to update the FTBL.
[0221] Finally, the embodiment also constructs an FTBL model online dynamic updating module, for the drift data, the embodiment increases an online incremental layer to quickly learn the characteristics of the drift data. In the online incremental process, the offline FTBL is described as an independent model, and the concept of expanding the width of the original model is used to enhance the online measurement accuracy of the model. At this time, the input data includes TSP data (X TSP ), drift data (X Dri ) and the output of the incremental layer . The prediction value of the new FTBL model is denoted as:
[0222]
[0223] Wherein, represents the weight matrix, is N FM +N En +N In +N OI FDT output matrix, and its dynamic updating process is consistent with equations (29)-(31).
[0224] After completing the online learning and prediction, it is also necessary to update the TSP and and SPE CL control limit to adapt to the online measurement of the next window.
[0225] AsFigure 6 and Figure 7 The DXN data of a large MSWI plant in Beijing from 2009 to 2020 was used in this example. Figure 6 and Figure 7 The left side is the actual concentration value of DXN.
[0226] It can be seen from Figure 6 and Figure 7 that the DXN concentration of the MSWI plant has not exceeded 0.1 TEQ ng / m 3 since it was put into operation, meeting the requirements of the pollutant control emission standard (GB18485-2014). Considering the difficulty in obtaining the true value of DXN, only an offline model can be established for soft measurement based on historical data. Therefore, the true value samples from 2009 to 2016 were taken as the historical data, and the true value samples from 2016 to 2020 were taken as the test data. The true value of DXN concentration was taken as the average emission concentration of the MSWI process within 2 hours. At the same time, in the discrete control system (DCS), process variables such as temperature, pressure, and flow are generated within seconds, and the samples corresponding to the true value of DXN obtained within the average sampling time are required. The division of DXN data in this example is shown in Table 2.
[0227] Table 2 Division of DXN data
[0228]
[0229] It can be seen from Figure 7 that there is a significant difference between the distribution of historical data and test data, which will lead to difficulty in accurately detecting the DXN emission concentration based on classical modeling methods.
[0230] Random forest (RF), back propagation neural network (BPNN), deep forest regression (DFR), support vector regression (SVR), fuzzy neural network (FNN), and width learning system (BLS) were used in this example. These are classical and popular methods, and were compared with the FTBL method proposed in this example.
[0231] The indicators used were root mean square error (RMSE) and explainable variance (EV), which were calculated as follows:
[0232]
[0233]
[0234] The experimental data and fitting curves are shown in Table 3 and Figure 8 and Figure 9 .
[0235] Table 3 Comparison results of different methods
[0236]
[0237] From Table 3 and Figure 8 and Figure 9 It can be seen that the modeling performance of the above methods is good, specifically:
[0238] (1) In the training data, the RMSE and EV of BPNN, FNN and FTBL are better than RF, DFR, SVR and BLS, which shows that BPNN, FNN and FTBL have better fitting performance on the training data;
[0239] (2) In the test data, due to the difference in data distribution, there is a significant difference in prediction performance, which leads to the difficulty of all methods to effectively fit the distribution of test data;
[0240] (3) The RMSE of SVR is the lowest (2.0233E-02), and the EV of RF is the lowest (-1.9795E-01), and the prediction curves of SVR, RF and DFR methods tend to be a straight line, which means that the response ability of RF, SVR and DFR methods to test data is poor; In addition, for the remaining methods such as BP, FNN, BLS and FTBL, their response to online data is more sensitive. This shows that the prediction trend of these methods fluctuates with the change of online process. The FTBL method proposed in this embodiment has higher advantages in modeling accuracy and model sensitivity.
[0241] The above experimental results show that the proposed offline modeling method has higher modeling accuracy and sensitivity than RF and DFR methods; In addition, compared with BP, FNN, SVR and BLS methods, FTBL has the same modeling accuracy and more stable sensitivity.
[0242] In this embodiment, the weighted k-means clustering is used to train the data, and the CSM index is used to construct the TSP data D TSP . The clustering results are taken as an example of the first three dimensions x1, x2, x3 of the historical data, and the results are shown in Figure 10 and Figure 11 .
[0243] Figure 10 and Figure 11 show the effectiveness of weighted k-means clustering in three-dimensional space, and there is a clear spatial distance between clusters. In the iteration process, the total distance R iter between clusters gradually decreases until convergence Figure 11 . Then, according to the CSM index, the redundant samples are deleted. The results are shown in Figure 12 and Table 4. Comparing the results of Figure 10 and Figure 12 , it can be seen that the samples within the class are close to the periphery of the centroid, and the distance between the classes is enlarged. In addition, from the statistical results of the CSM index, the clustering similarity of the typical samples is lower.
[0244] CSM of typical samples
[0245]
[0246] The principal components are determined by the contribution threshold δ PCA = 0.9, and 19 principal components are selected for statistical analysis. According to the equations (15), (16) and the 19 principal components, two and SPE CL control limits are further obtained, which are 13.2270 and 19.3475 respectively, and the confidence level is 95%.
[0247] The moving window size of the fixed online monitoring is 5, and the online drift identification and soft measurement results based on the offline model established in the last section are shown in Figure 13 , Figure 14 and Table 5.
[0248] Table 5 Results of online measurement process
[0249]
[0250] The experimental results show that there is a significant difference between the process data in the online stage and the historical data. According to the data statistical index control limit, all process data are drift samples. The number of FDT in the online incremental update layer is set to 5, which effectively controls the time cost of soft measurement. The results in Table 5 show that the FTBL online updating process has high fitting accuracy and short updating time, and the corresponding experimental results are shown in Figure 15 and Figure 16 .
[0251] As shown in Table 6, the RMSE and EV of the historical data are 9.9065E-03 and 8.9321E-01 respectively, and the RMSE and EV of the test data are 2.1595E-02 and 9.5089E-01 respectively. From the following aspects: 1) the offline FTBL model uses TSP, although the modeling accuracy is slightly decreased, the modeling time cost and the calculation cost of online updating process are reduced; 2) compared with the offline modeling part, the test monitoring stage realizes good fitting and high modeling accuracy of the data. The results show that the offline modeling and online measurement strategy proposed in this embodiment is sufficient.
[0252] Table 6 Prediction index statistics of offline and online stages
[0253]
[0254] For the working condition drift, the proposed strategy is realized in combination with actual engineering application. In this embodiment, the DXN concentration emission online detection system is developed based on C# programming language, and experimental tests are carried out on the laboratory semi-physical simulation platform and MSWI industrial site respectively. The corresponding hardware structure and software interface are shown in Figure 17 and Figure 18 .
[0255] After the software interface displays the online predicted DXN emission concentration in real time after starting running through the "start" button, the software interface displays the online predicted DXN emission concentration in real time. The hardware structure of the semi-physical simulation platform is displayed above the software interface. The effectiveness of the software system is verified by the test results of the soft measurement system on the semi-physical simulation platform, which shows that the method proposed in this embodiment can be applied to engineering and provide strong support for actual engineering.
[0256] The beneficial effects of the present application are as follows:
[0257] For the soft measurement problem of DXN emission concentration in MSWI process, this embodiment proposes a soft measurement strategy based on fuzzy tree width learning, the main contributions are as follows: a new FTBL algorithm is proposed to construct an offline DXN emission model, an online working condition drift identification and corresponding soft measurement strategy are proposed, an online dynamic updating method of FTBL model is proposed, and the effectiveness of the proposed strategy is verified on the laboratory simulation experiment platform based on actual process data, and the proposed method is verified in actual industrial process.
[0258] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be referred to each other.
[0259] In this embodiment, specific examples are applied to illustrate the principles and implementation modes of the present application. The above embodiment description is only used to help understand the method and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In view of the above, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for on-line soft-sensing of MSWI process dioxin emission concentration, characterized in that, The method comprises the following steps: determining process data of a typical sample pool according to historical process data sets of the MSWI based on a K-means weighted algorithm; performing principal component analysis on the process data of the typical sample pool to obtain a drift index control limit reflecting whether the MSWI process has changed; constructing an offline model based on FTBL, and inputting the process data of the typical sample pool and historical DXN true value data of the MSWI into the offline model to perform predictive calculation to obtain an offline calculation result; the offline model comprises a feature mapping layer, an enhancement layer and an incremental layer; performing principal component analysis on the obtained online data, and judging whether the online data is drift data or normal data according to the drift index control limit; if the online data is the normal data, jumping to the step of constructing the offline model based on FTBL and inputting the process data of the typical sample pool and the historical DXN true value data of the MSWI into the offline model to perform predictive calculation to obtain a calculation result; if the online data is the drift data, constructing an online model based on FTBL, and inputting the process data of the typical sample pool, the drift data and output data of the incremental layer of the offline model into the online model to perform predictive calculation to obtain an online calculation result; the online model comprises an online incremental layer; determining a DXN emission concentration prediction value according to the offline calculation result and the online calculation result.
2. The MSWI process dioxin emission concentration on-line soft-sensing method according to claim 1, characterized in that, The method for determining process data of a typical sample pool according to historical process data sets of the MSWI based on a K-means weighted algorithm comprises the following steps: Acquiring a historical process dataset X of MSWI His ; According to the historical process data set X His Obtaining historical data Wherein, x n is the nth sample, y n is the prediction value corresponding to the nth sample; N is the sample number of the historical data set of MSWI, and M is the feature number of the historical data set of MSWI. randomly select I instances as initial centroids arranging all samples into Class I according to weighted Euclidean distances between the samples and the centroids: where C i represents the i-th class; represents a weight vector of process variables, where, where H(·) represents the information entropy of a random variable, x m is the m-th feature vector, y represents the DXN concentration, x n,m represents the m-th feature value of the n-th sample, p(x n,m ) p(y n ) represents the marginal probability distribution, p(x n,m , y n ) is the joint probability distribution; updating the centroid C with the inter-class samples i : wherein, denotes the number of samples in the i-th cluster; The centroid is cyclically updated, and all centroids are obtained through preset conditions, which are expressed as: Wherein, δ TS is the threshold of evaluation index R iter , iter represents the iteration number, and the calculation formula of the measurement index is: establishing a TSP by minimizing clustering similarity, and establishing a formula as follows: where R DB is a cluster similarity measure index; where, M ij = {∑|C i -C j | b} 1 / b ; where S i denotes the sum of distances of the ith cluster, M ij denotes the Minkowski metric criterion.
3. The MSWI process dioxin emission concentration on-line soft-sensing method according to claim 2, characterized in that, performing principal component analysis on the process data of the typical sample pool to obtain a drift index control limit reflecting whether the MSWI process has changed, which comprises the following steps: The correlation coefficient matrix of the TSP data is represented as where N TSP is the number of TSP data D TSP ; R is the correlation coefficient matrix of the TSP data R is singular value decomposed and eigenvalues are calculated; the calculation formula is R = U M×M Σ M×M [V M×M ] T ; wherein, U M×M and V M×M represent orthogonal matrices, and Σ M×M is an M-dimensional diagonal matrix; performing dimensionality reduction by using a feature cumulative contribution rate η and a PCA contribution threshold value δ PCA: where P PCA is the number of selected principal components, P PCA is less than M; Rewrite the calculation formula as: wherein, is the load matrix; According to the score matrix T and the load matrix X is represented as: TPS is represented as: wherein represents X TSP projection onto the principal component space, represents X TSP projection onto the residual space; said with satisfies an orthogonal relationship; Furthermore, With The orthogonality is satisfied, which is proved as follows: expressing the drift index control limit as follows: wherein, Hotelling's T 2 control limit, SPE CL control limit of SPE, P PCA number of selected principal components, F α (P PCA , N TSP -P PCA ) represents F distribution with degrees of freedom of P PCA and (N TSP -P PCA ); c α represents normal deviation not more than (1 - a); Θ1, Θ2, and h0are calculated as shown below: where h0, Θ1, and Θ2 are all intermediate variables for the SPE control limit calculation, σ m are eigenvalues of singular value decomposition.
4. The MSWI process dioxin emission concentration on-line soft-sensing method according to claim 3, characterized in that, constructing an offline model based on FTBL, and inputting the process data of the typical sample pool and historical DXN true value data of the MSWI into the offline model to perform predictive calculation to obtain an offline calculation result, which comprises the following steps: For a given TSP data In D TSP A feature value x is randomly selected in D to define the node split function It is expressed as: n∈(1,N TSP )and m∈(1,M);where, is a sign function, rand(·) is a random number generation function, n and m do not take the maximum and minimum values; K fuzzy rules are determined for the TS fuzzy inference, and the kth rule can be expressed as: wherein, wherein, R k is the kth fuzzy rule, c k,m and σ k,m represent the center and width of the Gaussian function , respectively, and t leaf represents the tth leaf node; leaf is the Gaussian function; the result of the FDT is described as follows according to the K fuzzy rules: wherein, wherein f(·) is the FDT model, and g k (·) represents the antecedent and consequent parts of the TS fuzzy inference, representing the weight of the consequent part feature; The parameters of the FDT model f(·) during the training process are updated using a gradient descent method, including the center c k , the width σ k , and the weight The feature mapping layer outputs are represented as follows: wherein, wherein, is the nth FM output of the FDT model by inputting X TSP . The enhancement layer is to As input, the output of the enhancement layer is represented as: The output of the feature mapping layer and the augmentation layer is ridge regression learning algorithm weights between the input and the predicted output as follows: wherein is the pseudo-inverse matrix, λ is the regularization coefficient, and I is the identity matrix. The FDT model is added in the incremental layer, and the pseudo-inverse matrix is dynamically updated, so that As input, As output; the pseudo-inverse matrix update process of the incremental process is as follows: wherein, D, H k+1 , B T , C are intermediate variables of the pseudo-inverse matrix update process; new weight matrix is represented as: the predictive calculation process of the offline model FTBL is as follows:
5. The MSWI process dioxin emission concentration on-line soft-sensing method according to claim 4, characterized in that, performing principal component analysis on the obtained online data, and judging whether the online data is drift data or normal data according to the drift index control limit; if the online data is the normal data, jumping to the step of constructing the offline model based on FTBL and inputting the process data of the typical sample pool and the historical DXN true value data of the MSWI into the offline model to perform predictive calculation to obtain a calculation result; if the online data is the drift data, constructing an online model based on FTBL, and inputting the process data of the typical sample pool, the drift data and output data of the incremental layer of the offline model into the online model to perform predictive calculation to obtain an online calculation result, which comprises the following steps: calculating a drift value in a new window based on the following formula: wherein, and is a statistical indicator of the (N TSP +1)th process data , denotes a new load matrix, denotes a new diagonal matrix; The sample is determined as a drift sample or a normal sample by a judgment formula; the judgment formula is: For normal samples, reuse the offline model of the FTBL to perform soft measurement of the DXN concentration, which can be expressed as: For drift samples, the soft measurement value is calculated as: where ε Offset is the offset value for offline FTBL prediction output as follows: where N t denotes the total amount of data arriving at time t, denotes all the predicted values arriving at time t, denotes the mathematical expectation of the vector x IncTem denotes the incinerator temperature at time t. When the detection true value exists, input the TSP data, drift data, and the output of the incremental layer into the online model; the prediction value of the online model is: wherein, denotes a weight matrix, is N FM +N En +N In +N OI the output matrix of the FDT.
Citation Information
Patent Citations
Dioxin emission concentration transfer learning prediction method based on random forest
CN111461355A
MSWI process dioxin emission prediction method based on multi-window concept drift detection
CN114330845A