A Construction Method for a Combustion State Recognition Model of MSWI Process Flame Images
By formulating classification standards and using ViT depth feature extraction module to process flame images and building a combustion state recognition model, the problems of lack of flame image classification standards and insufficient model accuracy in the MSWI process in the prior art are solved, and more efficient combustion state recognition and cost reduction are achieved.
Patent Information
- Application Number
- CN202310144167.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-02-21
AI Technical Summary
In the prior art, the classification standards for flame image of MSWI process are relatively scarce, the model accuracy is insufficient, the data redundancy is used, and the cost is high.
By formulating classification standards for combustion states, the flame images collected on-site are classified, and the original flame images are processed, screened and spliced, and the combustion state recognition model is constructed.
This improves the accuracy of model judgment, reduces the time and cost of processing data, and achieves more efficient combustion state recognition.
Smart Images

Figure CN116152557B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of combustion state recognition, and in particular to a method for constructing a combustion state recognition model of a flame image in the MSWI process. Background Art
[0002] The continuous expansion of the urban scale has led to a gradual increase in the generation rate of municipal solid waste (MSW). Among them, the incineration treatment method that can achieve resource recovery has become the primary choice for many countries. The energy generated by MSW incineration (MSWI) can be used in the form of steam for waste heat power generation and community heating. In fact, the physical and chemical properties of MSW as a fuel have large fluctuations. Among them: the instability of the calorific value will cause fluctuations in the combustion calorific value, the fluctuation of the water content will cause uncertainty in the burnout time, and the difference in particle size will lead to changes in the weight flow of the grate bed feeding. These uncertain factors have become the main reasons hindering the stable combustion in the MSWI process, and thus also make it difficult to control the pollutant emission concentration. With the continuous deterioration of the global climate environment, countries around the world have formulated more stringent emission limit standards, which require the adoption of more efficient combustion control strategies to reduce the pollution emissions in the MSWI process. Obviously, the primary goal of combustion control is to perceive the process state, that is, to master the combustion state of MSW.
[0003] The combustion state of the MSWI process can be predicted through the visual characteristics of the flame image. Compared with traditional monitoring instruments such as temperature sensors and flame detectors, the image sequence containing space and time captured by a CCD camera can provide richer information such as flame temperature, emissivity, and flue gas concentration. In the research results, the accuracy of the model constructed by relying on feature engineering is weaker than that of the end-to-end recognition model. The reason is that the former relies too much on expert experience and needs to debug the adaptability of the model based on feature selection; however, the latter also has problems such as the need to process big data and long training time. In the actual industrial field, insufficient data will lead to problems such as low model accuracy or poor generalization ability, but data collection and annotation require a large amount of human and time costs, especially in the MSWI process. In addition, the current research is mostly based on the combustion line for state recognition, which cannot reflect the global state of the flame image in the MSWI process. To sum up, the existing technologies have problems such as a lack of classification standards for the flame image in the MSWI process, insufficient model accuracy, redundant data processing, and high cost. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for constructing a combustion state recognition model of a flame image in the MSWI process. The present invention solves the problems of the prior art such as a lack of classification standards for the flame image in the MSWI process, insufficient model accuracy, redundant data processing, and high cost.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A method for constructing a combustion state recognition model of a MSWI process flame image, comprising:
[0007] Formulating classification criteria for combustion states, classifying the flame images collected on-site, and obtaining original flame images;
[0008] Processing the original flame images by using a ViT deep feature extraction module to obtain first flame image features;
[0009] Screening the first flame image features to obtain second flame image features;
[0010] Stitching the second flame image features and the original flame images to obtain third flame image features;
[0011] Constructing a combustion state recognition model according to the third flame image features.
[0012] Preferably, the processing the original flame images by using a ViT deep feature extraction module to obtain first flame image features includes:
[0013] Using the embedding layer of the ViT deep feature extraction module to convert the original flame images into 2D matrices, and taking the 2D matrices as the inputs of the encoding layer of the ViT deep feature extraction module; the encoding layer is composed of layer normalization, multi-head attention mechanism, MLP, and residual;
[0014] Processing the 2D matrices by using the encoding layer to obtain the outputs of the encoding layer, and taking the outputs of the encoding layer as the inputs of the multi-layer perceptron layer of the ViT deep feature extraction module;
[0015] Processing the outputs by using the multi-layer perceptron layer of the ViT deep feature extraction module to obtain a set of prediction results;
[0016] Evaluating the set of prediction results, adjusting the network parameters when the prediction results do not meet the index requirements until the accuracy requirement is met and then stopping the training, and finally taking the output of the encoding layer as the first flame image features.
[0017] Preferably, the using the embedding layer of the ViT deep feature extraction module to convert the original flame images into 2D matrices includes:
[0018] Linearly embedding each image into a sequence containing multiple image patches;
[0019] Projecting each image patch to D dimensions through the embedding matrix, and concatenating trainable parameters for classification in front of the D-dimensional vectors;
[0020] Obtain the position information of the image patch according to the randomly initialized position encoding parameters;
[0021] Obtain the 2D matrix according to the image patch sequence, the trainable parameters for classification, and the position information of the image patch.
[0022] Preferably, processing the 2D matrix by using the encoding layer to obtain the output of the encoding layer, including:
[0023] Perform normalization processing on the 2D matrix by using the layer normalization to obtain a processed 2D matrix;
[0024] Use the multi-head attention mechanism to endow the model with the ability to notice information in different subspaces according to the processed 2D matrix;
[0025] Use the output of the embedding layer and the output of the multi-head attention mechanism as the input of the residual structure to obtain the input of the layer normalization;
[0026] Use the output result of the layer normalization as the input of the MLP, and the MLP consists of two linear mapping layers and a GELU activation function layer;
[0027] Process the output obtained by the layer normalization by using the MLP to obtain the output of the MLP;
[0028] Use the output of the MLP and the output of the residual structure as the input of the next residual structure to obtain the output of the encoding layer.
[0029] Preferably, processing the output by using the multi-layer perceptron layer of the ViT deep feature extraction module to obtain a set of prediction results, including:
[0030] The input of the multi-layer perceptron layer is the output of the last encoding layer. Extract the sequence position where the trainable parameters for classification in the output of the encoding layer are located as the corresponding feature representation of the image and use it as the input of the MLP to obtain the output of the MLP;
[0031] Perform Softmax processing on the output of the MLP to obtain the final classification result;
[0032] Based on the ViT deep feature extraction module, obtain a set of prediction results according to the final classification result.
[0033] Preferably, evaluate the set of prediction results. When the prediction results do not meet the index requirements, adjust the network parameters until the accuracy requirement is met and then stop training. Finally, take the output of the encoding layer as the first flame image feature, including:
[0034] Pre-train based on the ImageNet dataset and iteratively optimize the ViT model. When the performance of the network meets the requirements, save the weight and bias parameters in the model;
[0035] Based on the extraction of transfer features from the image dataset, input the original flame image dataset into the ViT model to obtain the first flame image features.
[0036] Preferably, feature selection is performed on the first flame image features to obtain second flame image features, including:
[0037] Based on expert experience, select the features extracted from the last three layers of the first flame image features as the second flame image features.
[0038] Preferably, the process of splicing the second flame image features and the original flame image to obtain third flame image features includes:
[0039] Reduce the original image to 0.1 times to obtain the reduced original image;
[0040] Splice the second flame image with the reduced original image to obtain the third flame image features.
[0041] Preferably, the process of constructing a combustion state recognition model based on the third flame image features includes:
[0042] Input the third flame image features into the first CF layer of the CF module to obtain a class distribution vector;
[0043] Concatenate the class distribution vector and the third flame image features to obtain a feature vector, and use the feature vector as the input to the next CF layer of the CF module to construct the combustion state recognition model.
[0044] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0045] The present invention provides a method for constructing a combustion state recognition model for the flame image in the MSWI process. By extracting and screening features from the original flame image and inputting the screened image into the CF module, a combustion state recognition model is constructed to improve the accuracy of model determination and reduce the time and cost of processing data. Brief Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 Flowchart of a method for constructing a combustion state recognition model of a flame image in the MSWI process provided by an embodiment of the present invention;
[0048] Figure 2 Schematic diagram of video acquisition and on-site monitoring provided by an embodiment of the present invention;
[0049] Figure 3 Structural diagram of the ViT model for flame image feature extraction; where (a) is the structural diagram of the ViT model; (b) is the structural diagram of the Transformer encoding layer;
[0050] Figure 4 Structural diagrams of the MSA and SA provided by an embodiment of the present invention; where (a) is the structural diagram of the MSA, and (b) is the structural diagram of the SA;
[0051] Figure 5 Improved deep forest recognition model provided by an embodiment of the present invention; Detailed implementation manners
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0053] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0054] In the description, claims, and drawings of this application, terms such as "first", "second", "third", and "fourth" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, including a series of steps, processes, methods, etc. is not limited to the listed steps, but optionally also includes steps not listed, or optionally also includes other step elements inherent to these processes, methods, products, or devices.
[0055] The object of the present invention is to provide a method for constructing a combustion state recognition model for MSWI process flame images. The present invention solves the problems in the prior art such as the lack of classification criteria for MSWI process flame images, insufficient model accuracy, redundant data processing, and high cost.
[0056] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0057] As Figure 1 shown, the present invention provides a method for constructing a combustion state recognition model for MSWI process flame images, including:
[0058] Step 100: Formulate classification criteria for combustion states, and classify the flame images collected on-site to obtain an original flame image dataset;
[0059] Step 200: Process the original flame images using a ViT deep feature extraction module to obtain first flame image features;
[0060] Step 300: Screen the first flame image features to obtain second flame image features;
[0061] Step 400: Stitch the second flame image features and the original flame images to obtain third flame image features;
[0062] Step 500: Construct a combustion state recognition model based on the third flame image features.
[0063] Specifically: Construct an MSWI flame image combustion state dataset;
[0064] Use the pre-trained ViT network Transformer encoding layer to extract multi-layer visual transformation features of the flame images, and select the deep features based on expert experience.
[0065] Combine the selected ViT visual transformation features and the original flame images as the input of the cascade forest to construct an improved deep forest recognition (IDFC) model.
[0066] The constructed flame image dataset is used to verify the model.
[0067] The results show that the recognition accuracies of the proposed method for the left and right grate flame images reach 98.15% and 96.43% respectively, achieving a low error acceptable in industry.
[0068] This implementation also discloses the MSWI process flow of the grate furnace type. The MSWI process mainly includes six process stages: storage and fermentation, solid waste combustion, waste heat exchange, steam power generation, flue gas treatment, and flue gas emission. Among them, solid waste combustion is the core of the entire MSWI process. The combustion state of MSW plays a decisive role in the costs and efficiencies of waste heat exchange, steam power generation, flue gas treatment and other links.
[0069] Figure 2 Shown is the schematic diagram of MSWI process video acquisition and on-site monitoring. Figure 2 As can be seen, in order to monitor the combustion state of MSW inside the furnace in real time, industrial cameras are installed at the positions obliquely above the ends of the left and right grates respectively to obtain real-time video streams. After being transmitted through coaxial cables, they are stored in the monitoring machine by a video capture card and displayed in real time. Experienced on-site experts make judgments on the combustion state of MSW by observing the flame images, and then formulate corresponding combustion control strategies. This shows that the recognition of the combustion state of MSW can provide important feedback information for the intelligent control of the MSWI process. The inside of the furnace includes: a feeder, a drying grate, a combustion grate 1, a combustion grate 2, and an afterburning grate.
[0070] The main defect of the combustion state division standard based on the MSWI combustion line position is that the actual combustion state of MSW is complex and changeable. Although the standard for dividing the combustion line state based only on the combustion line position has physical meaning and interpretability, there is a problem of lack of information representation, that is, it is only based on one line; essentially, various complex information such as flames, combustion line positions, and bed layer distributions jointly constitute the visual characteristics of the incinerator combustion state. Therefore, the combustion state division standard based on a single combustion line position has the problem of overgeneralization. To sum up, the feedback information provided by the combustion state recognition model constructed based on the existing combustion line division standard is very limited, and it is necessary to construct a typical combustion state dataset.
[0071] This implementation discloses the specific meanings of the parameters used in the following instructions:
[0072]
[0073]
[0074]
[0075]
[0076] In the modeling strategy proposed in this paper:
[0077] (1) Typical combustion state dataset construction module: Based on expert experience and literature research in related fields, a classification standard including four combustion states of normal, partial burning, flash burning, and smoldering is formulated, and the flame images collected on-site are classified.
[0078] (2) ViT deep feature extraction module: First, the ViT network parameters pre-trained on the ImageNet dataset are migrated to the task of extracting features from MSWI process flame images by loading the weight file; then the flame images are input into the embedding layer and the Transformer encoding layer to obtain 12-layer flame image features with a global attention interaction mechanism; finally, the last 3 layers are selected as effective features based on expert experience.
[0079] (3) Improved deep forest recognition model construction module: The visual features of the flame images extracted based on the ViT model are mapped and spliced with the original flame images, and finally input into the improved CF to construct a combustion state recognition model to obtain the combustion state recognition result.
[0080] Due to the characteristics of MSW such as low calorific value, large range of moisture change, complex ash content and composition, the combustion process is complex and variable. For the convenience of research, this paper temporarily does not consider the flame segments whose combustion states are difficult to define even relying on existing expert experience, and only summarizes and organizes the combustion states with clear features. Based on the actual combustion situation on-site in the MSWI process, by combining on-site expert experience and the research on combustion phenomena in the combustion process of biomass grate furnaces, this paper proposes a classification criterion based on four combustion states of normal, partial burning, flash burning, and smoldering. The features and causes of each criterion are described in detail as follows:
[0081] (1) Normal: The combustion line is linearly distributed within the pixel band corresponding to the combustion section, the flame is stable, bright and concentrated, and the MSW combustion state is good at this time;
[0082] (2) Partial burning: The combustion line runs through the drying section, the combustion section and the burnout section and is curved, the flame is uneven in height, bright but scattered, and its generation reason is the uneven distribution of the MSW material layer;
[0083] (3) Flash burning: The combustion line is scattered within the pixel band corresponding to the drying section and the combustion section, the flame has local flashovers, is locally bright and is in a divergent jet shape, and its generation reason is short-term lack of fuel;
[0084] (4) Smoldering: The combustion line is star-shaped distributed on the three-section grate, there is a large area of lack of fire inside the furnace, and a large area of black block area appears, and the MSW combustion state is poor.
[0085] Based on the above classification standard, a typical combustion state dataset is constructed by manual annotation.
[0086] The basic model structure of ViT is as shown in Figure 3 (a), which mainly consists of an embedding layer, a Transformer encoding layer, and a multi-layer perceptron.
[0087] Furthermore, using the ViT deep feature extraction module to process the original flame image to obtain the first flame image feature includes:
[0088] Using the embedding layer of the ViT deep feature extraction module to convert the original flame image into a 2D matrix, and taking the 2D matrix as the input of the encoding layer of the ViT deep feature extraction module; the encoding layer consists of layer normalization, multi-head attention mechanism, MLP, and residual;
[0089] Using the encoding layer to process the 2D matrix to obtain the output of the encoding layer, and taking the output of the encoding layer as the input of the multi-layer perceptron layer of the ViT deep feature extraction module;
[0090] Using the multi-layer perceptron layer of the ViT deep feature extraction module to process the output to obtain a set of prediction results;
[0091] Evaluating the set of prediction results. When the prediction results do not meet the index requirements, the network parameters need to be adjusted until the accuracy requirement is met and the training stops. Finally, take the output of the encoding layer as the first flame image feature.
[0092] Furthermore, using the embedding layer of the ViT deep feature extraction module to convert the original flame image into a 2D matrix includes:
[0093] Linearly embedding each image into a sequence containing multiple image patches;
[0094] Projecting each image patch to D dimensions through the embedding matrix, and concatenating trainable parameters for classification in front of the D-dimensional vector;
[0095] According to the randomly initialized position encoding parameters, obtaining the position information of the image patches;
[0096] Obtaining the 2D matrix according to the image patch sequence, the trainable parameters for classification, and the position information of the image patches.
[0097] For the image dataset (where X n ∈R H×W×C represents the nth image, H, W, and C respectively represent its height, width, and number of channels, y n represents the label, and N represents the size of the dataset sample), the following steps are used to convert it into a 2D matrix:
[0098] Step (1): Linearly embed each image X n into a sequence (B 1 , B 2 , …, B m , …, B M ) n containing M image patches, where for each image patch p represents the width and height of each image patch, and M = HW / p 2 ;
[0099] Step (2): Project each image patch into D dimensions through an embedding matrix and concatenate trainable parameters dedicated for classification in front of the D-dimensional vectors
[0100] Step (3): Record the position information of the image patches through randomly initialized position encoding parameters ;
[0101] Step (4): Obtain a 2D serialized matrix as the input to the encoding layer
[0102] In summary, the processing procedure of the embedding layer can be expressed as follows:
[0103]
[0104] Furthermore, use the encoding layer to process the 2D matrix to obtain the output of the encoding layer, including:
[0105] Normalize the 2D matrix using the layer normalization to obtain a processed 2D matrix;
[0106] Use the multi-head attention mechanism to endow the model with the ability to notice information in different subspaces according to the processed 2D matrix;
[0107] Use the output of the embedding layer and the output of the multi-head attention mechanism as the input to the residual structure to obtain the input to the layer normalization;
[0108] Process the input obtained by the layer normalization to obtain an output result; the output result is the input to the MLP, and the MLP consists of two linear mapping layers and a GELU activation function layer;
[0109] Process the output obtained by the layer normalization using the MLP to obtain the output of the MLP;
[0110] Use the output of the MLP and the output of the residual structure as the input to the next-layer residual structure to obtain the output of the encoding layer
[0111] The Transformer encoding layer mainly consists of layer normalization (LayerNorm, LN), multi-head self-attention mechanism (Multi-head Self-Attention, MSA), MLP, and residual (ADD) structure, as shown in Figure 3 (b). Here, the first encoding layer is taken as an example for description.
[0112] (1) LN1: Normalize the samples, which can ensure the stability of the data feature distribution as the network depth increases, thus accelerating the model convergence speed. Its input is When, the output is
[0113] As shown in Figure 4 (a), (2) MSA: It is the core component of the encoding layer. The self-attention mechanism (Self-Attention, SA) is the core of MSA. The schematic diagram of the SA principle is shown in Figure 4 (b).
[0114] The steps of SA calculation are as follows:
[0115] Step (1): Through the transformation matrices and transform the input vector into the query matrix the key matrix and the value matrix three different matrices, as shown in the following formula:
[0116]
[0117] Among them, and The dimension d k is equal to The dimension d v and there is d k = d v = D / U;
[0118] Step (2): Obtain the attention factor of each through the dot product of and ;
[0119] Step (3): Scale the attention factor by times;
[0120] Step (4): Perform Softmax processing on the attention factor of each to obtain the weight;
[0121] Step (5): Multiply the weight with its corresponding value to obtain the final attention value.
[0122] In summary, the SA calculation process can be expressed as the following formula:
[0123]
[0124] where, denotes the Softmax process, and its calculation formula is as follows:
[0125]
[0126] where, G represents the dimension of the original vector, e represents the natural logarithm, and c represents the calculation result of SA, and its element composition is (c 1 , c 2 , …, c G ). This function compresses the G-dimensional vector c of any real number into the vector such that the range of each of its elements is within (0, 1).
[0127] The MSA structure diagram is shown in Figure 4(a). Similar to the idea of using multiple convolutional kernels in the convolutional layer to learn image features from different perspectives, the MSA structure enables the model to notice information in different subspaces. The steps are as follows:
[0128] Step (1): Map the input vector to U spaces respectively. Taking the u-th space as an example, it is as follows:
[0129]
[0130] Step (2): Perform SA processing on the input mapped vector in each space to obtain as follows:
[0131]
[0132] Step (3): Concatenate the SA output results to obtain the MSA result, as shown below:
[0133]
[0134] where, ζ(·) represents the MSA operation, is the transformation matrix, and Concat(·) represents the stacking operation in the column vector direction.
[0135] (3) ADD1: This structure can alleviate the vanishing gradient phenomenon, thus facilitating the construction of deep networks. Its input is the output of the embedding layer and the MSA output The output is The processing process is as follows:
[0136]
[0137] (4) LN2: The input is the output of ADD When, the output is
[0138] (5) MLP: It consists of two linear mapping layers and a GELU activation function layer, and its input is The output is Its processing process is expressed as the following formula:
[0139]
[0140] where k 1 , k 2 respectively represent the coefficients of the first and second linear layers, and GELU(·) represents the GELU function, and its specific expression is as follows:
[0141]
[0142] where, * represents the function input, σ represents the standard deviation, and μ represents the mean.
[0143] (6) ADD2: Its input is and The output is
[0144] Correspondingly, the processing process of the l-th layer encoder for the image patch sequence can be expressed as formulas (11)-(12):
[0145]
[0146]
[0147] where LN(·) represents the LN operation, MLP(·) represents the MLP operation, and L = 12 represents the number of Transformer encoding layers.
[0148] Finally, after being processed by the L-th layer encoding layer, we can obtain
[0149] Furthermore, the multi-layer perceptron layer of the ViT depth feature extraction module is used to process the output to obtain a set of prediction results, including:
[0150] The input of the multi-layer perceptron layer is the output of the last encoding layer, and the sequence position of the output of the encoding layer is extracted as the corresponding feature representation of the image and used as the input of the MLP to obtain the output of the MLP;
[0151] Perform Softmax processing on the output of the MLP to obtain the final classification result;
[0152] Based on the ViT deep feature extraction module, obtain a set of prediction results according to the final classification result.
[0153] The structure of the multi-layer perceptron adopted here is the same as above. The difference is that its input is the output of the L-th encoding layer, and the predicted category of its output is
[0154] For the output of the last encoder layer Extract The sequence position where it is located As the corresponding feature representation of the image and used as the input of the MLP, its processing process is as follows:
[0155]
[0156] Take the output of the MLP Perform Softmax processing to obtain the final classification result
[0157]
[0158] Then the image set After passing through ViT, obtain the corresponding set of prediction results
[0159] The training process of the deep network is generally based on the gradient descent method, and its minimum value is found along the direction of the loss function decreasing. The learning process of the Vit model specifically includes calculating the gradient by the backpropagation algorithm and updating the weight parameters by the Adam algorithm.
[0160] (1) Calculate the gradient by the backpropagation algorithm
[0161] Here, ViT(X n ) represents the prediction result of the ViT network for the input image X n That is y n Represents the true category corresponding to the image, then the specific expression of the loss function is as follows:
[0162] Loss(ViT(X n ),y n );
[0163] The gradient calculation formula of the loss function for the input X n Is:
[0164]
[0165] The specific process of the backpropagation algorithm for deriving the gradients of network nodes from back to front is as follows:
[0166] First, calculate the error of the j-th layer:
[0167]
[0168] Then, calculate the backpropagation error of the j-th layer:
[0169]
[0170] Among them, ω represents the weight matrix of the network.
[0171] Thus, the gradient values of ω and the bias matrix β in the network at the j-th layer are obtained and
[0172]
[0173] (2) Update the weight parameters by the Adam algorithm
[0174] The network parameters are updated using the Adam gradient descent algorithm,
[0175]
[0176] In the formula, θ t represents the network parameters at the t-th iteration; α is the learning rate; T represents the total number of network training iterations; γ is a very small positive real number to prevent the denominator from being zero; and represent the first-order momentum and second-order momentum of the network at the t-th update, as follows:
[0177]
[0178]
[0179] In the formula, and The default values of are 0.9 and 0.999.
[0180] For the t-th iteration, the update steps of the network parameters are as follows:
[0181] Step(1): Calculate the gradients of the parameters of each layer currently according to formulas (17) - (20), and then obtain the gradients of the entire network parameters
[0182] Step(2): Calculate the first-order momentum and the second-order momentum
[0183] Step (3): Calculate the descent gradient η at the t-th time t , as follows:
[0184]
[0185] Step (4): Use η t to update θ t to obtain the network parameter θ at the (t + 1)-th time t+1 , as follows:
[0186] θ t+1 = θ t - η t ;
[0187] Because and Therefore, the first-order momentum and the second-order momentum are close to 0 at the initial stage of updating the parameters. Therefore, equations (20) and (21) need to be corrected by the following deviation:
[0188]
[0189]
[0190] The weight and bias parameters are iteratively updated by the Adam algorithm. When the Loss(Vit(X n ), y n ) stops decreasing, the training effect of the model reaches the optimum, and the model stops training.
[0191] Furthermore, evaluate the prediction result set. When the prediction result does not meet the index requirements, the network parameters need to be adjusted until the accuracy requirement is met and then the training stops. Finally, take the output of the encoding layer as the first flame image feature, including:
[0192] Pre-train based on the ImageNet dataset and iteratively optimize the ViT model. When the performance of the network meets the requirements, save the weight and bias parameters in the model;
[0193] Extract the transfer features based on the image dataset. Input the original flame image dataset into the ViT model to obtain the first flame image feature.
[0194] Furthermore, perform feature selection on the first flame image feature to obtain the second flame image feature, including:
[0195] Based on expert experience, select the features extracted from the last three layers of the first flame image feature as the second flame image feature.
[0196] Aiming at the problems of difficult acquisition of effective labeled samples and high manual labor cost in the MSWI process flame image dataset, the image feature knowledge on the ImageNet dataset is migrated by TL and applied to the combustion state recognition of the MSWI process. The specific steps are as follows:
[0197] Step(1): Based on the pre-training of the ImageNet dataset, the ViT model is iteratively optimized. When the performance of the network meets the requirements, the weight and bias parameters in the model are saved.
[0198] Step(2): Migration feature extraction based on the image dataset. The flame image dataset is input into the ViT model to obtain the feature sequence after the global attention interaction of each encoding layer of the Transformer on the flame image, as follows:
[0199]
[0200] where f ViT (·) represents the ViT deep feature extraction model composed of the embedding layer and the encoding layer.
[0201] Step(3): Migration feature selection. Based on expert experience, only the feature extraction results of the last three layers are selected as the deep features of the flame image.
[0202] Further, the splicing of the second flame image feature and the original flame image to obtain the third flame image feature includes:
[0203] The original image is reduced to 0.1 times to obtain the reduced original image;
[0204] The second flame image is spliced with the reduced original image to obtain the third flame image feature.
[0205] Further, the construction of the combustion state recognition model according to the third flame image feature includes:
[0206] The third flame image feature is input into the first CF layer of the CF module to obtain the class distribution vector;
[0207] The class distribution vector and the third flame image feature are concatenated to obtain a feature vector, and the feature vector is used as the input of the next CF layer of the CF module to construct the combustion state recognition model.
[0208] Since the flame image is an RGB three-channel image with a large size, using the multi-granularity scanning of the original DFC will bring huge computational consumption. Therefore, in this paper, a pre-trained ViT model is adopted as the primary feature extractor. In addition, to preserve the original features, referring to the mechanism of the residual learning model, the ViT output is concatenated with the original flame image to obtain multi-scale features. Furthermore, to reduce the computational consumption, is reduced by a factor of 0.1 to obtain The feature concatenation is as follows:
[0209]
[0210] where flatten(·) represents the image flattening operation.
[0211] The base learners adopted by each CF layer of the CF module are RF and CRF. Among them, the number of decision trees Tree_Number and the minimum number of samples Mini_Samples in the leaf nodes of the CF layer forest need to be determined. The deep features of the flame image extracted based on the ViT transfer model, the original reduced image, and CF constitute the IDFC model, and its structure is as Figure 5 shown.
[0212] In the IDFC model, each layer of CF contains 2 RFs and 2 CRFs for cascade learning, and the Stack idea is adopted to construct the CF layer model. When is input into CF to construct the recognition model, except that the first CF layer directly takes as the input feature of each forest learner, the subsequent CF layers need to take the class distribution vector output by the previous layer and the concatenated vector as the input of this CF layer to effectively prevent the overfitting phenomenon of the Stack strategy. The number of CF layers is adaptively adjusted by cross-validation.
[0213] When is input into the IDFC model, the recognition prediction result of the combustion state of the flame image is obtained:
[0214]
[0215] where f DFC (·) represents the IDFC model.
[0216] The flame image data adopted in this embodiment is sourced from a certain MSWI power plant in Beijing. To monitor the combustion state in real time, an industrial camera is installed at the end of the grate. After the MSWI flame video is transmitted through a coaxial cable, it is stored in the monitoring machine using a video capture card.
[0217] For the flame video with a frame rate of 25 frames / s collected on-site, flame images are obtained by first screening video segments of typical combustion states and then sampling the selected videos at a sampling rate of 1 frame / min. Among them, the video segments need to be manually labeled, and the sampling is based on a Matlab program. Finally, the total number of typical combustion state images of the left and right grate bars is 3,289 and 2,685 respectively. The more specific distribution of each combustion state of the left and right grate bars is shown in Table 1.
[0218] Table 1 is the flame image dataset, as shown below:
[0219] Table 1 Flame Image Dataset
[0220]
[0221] As can be seen from Table 1, the combustion states of the left and right grate bars are inconsistent because the internal space of the furnace is large. Therefore, combustion state recognition models should be established for the left and right grate bars respectively.
[0222] To verify the accuracy of this model, this implementation discloses the introduction of the model recognition results:
[0223] To evaluate the performance of the model, the confusion matrix, accuracy, precision, and recall are used as evaluation indicators. The confusion matrix of the classification results is shown in Table 2.
[0224] Table 2 is the confusion matrix of the classification results, as shown below:
[0225] Table 2 Confusion Matrix of Classification Results
[0226]
[0227] In Table 2, the column direction of the confusion matrix represents the predicted category, and the row direction represents the true category. It can be seen from the confusion matrix which parts the model is confused about during prediction.
[0228] Accuracy:
[0229] Precision:
[0230] Recall:
[0231] The hardware configuration environment used to build the model is CPU: Intel(R) Core(TM) i9-11900K, RAM: 32G, GPU: NVIDIA GeForce RTX3060Ti. The integrated development environment used is PyCharm Community Edition-2022.2, and the model is built based on the Pytorch deep learning framework. The training set and test set are divided according to the ratio of 0.7:0.3 of the total number of samples.
[0232] To clarify the extraction of flame image features by the Transformer encoding layer during model training, taking the right grate flame image as an example, the Gradient-weighted Class Activation Mapping (Grad-CAM) is used to visualize the output feature map of the Transformer encoding layer. In the above feature maps, the brighter the color area, the higher the attention of the model to it. For the four typical flame images, the extraction effect of features in the first layer is weak. The second and third layers respectively learn some features of the flame area and the material burnout area. The fourth to ninth layers focus on the comparison and balance of features between the flame area and the material burnout area. The tenth to twelfth layers further extract the features of each area of the flame image. Therefore, the process of feature extraction of the flame image by each layer of the Transforer encoding layer has different effects.
[0233] Regarding the finally extracted features of the twelfth layer: for the flame images in the normal and partial burning states, the focus of feature extraction is on the material burnout area, which is consistent with the experience of on-site experts judging based on the shape of the combustion line; the focus of feature extraction of the flame image in the flashover state is on the sparsity of the flame area, indicating that the model can notice the jet-divergent shape presented by the flame during flashover; the focus of feature extraction of the flame image in the smoldering state is on the entire flame image. At this time, the overall image shows a darker characteristic, so the focus of the model is no longer limited to a certain local area.
[0234] In addition, to make the best use of the features of flame images at different scales to improve the recognition effect of the IDF model, the features extracted from the tenth to twelfth layers of the Transformer encoding layer are used as the CF input here, so as to build a recognition model.
[0235] The parameter settings and recognition results of the IDFC model are shown in Table 3.
[0236] Table 3 Recognition Results of the IDFC Model
[0237]
[0238] As can be seen from Table 3, from the perspective of recognition accuracy, the IDFC recognition models constructed based on the left and right grate flame images can reach 98.15% and 96.43%, respectively, which can meet the recognition accuracy requirements of the actual industrial site. The performance such as the accuracy rate, precision rate, and recall rate of the model constructed based on the left grate is about 2 percentage points higher than that of the right grate, which may be related to the difference in the quality of the left and right grate flame images.
[0239] The beneficial effects of the present invention are as follows:
[0240] The present invention extracts and screens features from the original flame images, inputs the screened images into the CF module, and constructs a combustion state recognition model to improve the determination accuracy of the model and reduce the time and cost of processing data. Each embodiment in this specification is described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0241] Specific examples are used in this article to elaborate on the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for constructing a combustion state recognition model of a MSWI process flame image, characterized in that, it includes: Formulating classification criteria for combustion states, classifying the flame images collected on-site, and obtaining an original flame image dataset with markings completed; Using the ViT deep feature extraction module to process the original flame image to obtain the first flame image feature; Screening the first flame image feature to obtain the second flame image feature; Stitching the second flame image feature and the original flame image to obtain the third flame image feature; Constructing a combustion state recognition model according to the third flame image feature; The using the ViT deep feature extraction module to process the original flame image to obtain the first flame image feature includes: Using the embedding layer of the ViT deep feature extraction module to convert the original flame image into a 2D matrix, and taking the 2D matrix as the input of the encoding layer of the ViT deep feature extraction module; the encoding layer consists of layer normalization, multi-head attention mechanism, MLP and residual; Using the encoding layer to process the 2D matrix to obtain the output of the encoding layer, and taking the output of the encoding layer as the input of the multi-layer perceptron layer of the ViT deep feature extraction module; Using the multi-layer perceptron layer of the ViT deep feature extraction module to process the output to obtain a set of prediction results; Evaluating the set of prediction results, adjusting the network parameters when the prediction results do not meet the index requirements until the accuracy requirement is met and then stopping training, and finally taking the output of the encoding layer as the first flame image feature; Performing feature selection on the first flame image feature to obtain the second flame image feature, including: Based on expert experience, selecting the last three layers of features of the first flame image feature as the second flame image feature.
2. The method for constructing a combustion state recognition model of a MSWI process flame image according to claim 1, characterized in that, The using the embedding layer of the ViT deep feature extraction module to convert the original flame image into a 2D matrix includes: Linearly embedding each image into a sequence containing multiple image patches; Projecting each image patch to D dimensions through the embedding matrix, and concatenating trainable parameters for classification in front of the D-dimensional vector; According to the randomly initialized position encoding parameters, obtaining the position information of the image patches; Obtaining the 2D matrix according to the image patch sequence, the trainable parameters for classification and the position information of the image patches.
3. The method for constructing a combustion state recognition model of a MSWI process flame image according to claim 1, characterized in that, The using the encoding layer to process the 2D matrix to obtain the output of the encoding layer includes: Using the layer normalization to perform normalization processing on the 2D matrix to obtain the processed 2D matrix; Using the multi-head attention mechanism to endow the model with the ability to notice information in different subspaces according to the processed 2D matrix; Using the output of the embedding layer and the output of the multi-head attention mechanism as the input of the residual structure to obtain the input of the layer normalization; Use the output result obtained by the layer normalization as the input of the MLP, where the MLP consists of two linear mapping layers and a GELU activation function layer; Process the output obtained by the layer normalization using the MLP to obtain the output of the MLP; Use the output of the MLP and the output of the residual structure as the input of the next residual structure to obtain the output of the encoding layer.
4. A method for constructing a combustion state recognition model of a MSWI process flame image according to claim 1, characterized in that, The multi-layer perceptron of the ViT deep feature extraction module is used to process the output to obtain a set of prediction results, including: The input of the multi-layer perceptron is the output of the last encoding layer. Extract the sequence positions of the trainable parameters for classification in the output of the encoding layer as the corresponding feature representation of the image, and use it as the input of the MLP to obtain the output of the MLP; Perform Softmax processing on the output of the MLP to obtain the final classification result; Based on the ViT deep feature extraction module, obtain a set of prediction results according to the final classification result.
5. A method for constructing a combustion state recognition model of a MSWI process flame image according to claim 1, characterized in that, Evaluate the set of prediction results. When the prediction results do not meet the index requirements, the network parameters need to be adjusted until the accuracy requirement is met and the training is stopped. Finally, take the output of the encoding layer as the first flame image feature, including: Based on pre-training on the ImageNet dataset, iteratively optimize the ViT model. When the performance of the network meets the requirements, save the weight and bias parameters in the model; Based on the transfer feature extraction of the flame image dataset, input the original flame image dataset into the ViT model, and take the output of each Transformer encoder in the encoding layer as the first flame image feature.
6. A method for constructing a combustion state recognition model of a MSWI process flame image according to claim 1, characterized in that, The splicing of the second flame image feature and the original flame image to obtain the third flame image feature includes: Reduce the original image to 0.1 times to obtain the reduced original image; Splice the second flame image with the reduced original image to obtain the third flame image feature.
7. A method for constructing a combustion state recognition model of a MSWI process flame image according to claim 1, characterized in that, Constructing a combustion state recognition model according to the third flame image feature includes: Input the third flame image feature into the first CF layer of the CF module to obtain a class distribution vector; Concatenate the class distribution vector and the third flame image feature to obtain a feature vector, and use the feature vector as the input of the next CF layer of the CF module to construct the combustion state recognition model.
Citation Information
Patent Citations
MSWI process combustion state identification method based on multi-feature fusion and improved cascade forest
CN114882391A