Carbon monitoring data anomaly detection method and device and storage medium
By employing multi-scale feature extraction and a cascaded encoder-decoder structure, combined with dilated convolution and multi-head attention mechanisms, the accuracy problem of anomaly identification in CEMS carbon monitoring data was solved, improving the accuracy and reliability of carbon monitoring data from thermal power plants, reducing the risk of overfitting, and enhancing the model's generalization ability.
Patent Information
- Application Number
- CN202510810112.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, CEMS carbon monitoring data often exhibits anomalies, leading to a decrease in the accuracy and reliability of the monitoring data. Furthermore, manual intervention is inefficient and prone to errors, making it difficult to accurately identify abnormal data.
Employing a multi-scale feature extraction and cascaded encoder-decoder structure, combined with dilated convolution, multi-head attention mechanism, and Gaussian modulation, outliers in carbon monitoring data are accurately identified through multi-level feature extraction and reconstruction.
It improves the accuracy and reliability of carbon monitoring data, reduces the risk of model overfitting, enhances the ability to understand complex data patterns, and can better adapt to the complexity and diversity of thermal power unit operation data, accurately identifying abnormal data.
Smart Images

Figure CN120953628A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial carbon emission monitoring technology, specifically to a method, device, and storage medium for detecting anomalies in carbon monitoring data. Background Technology
[0002] With increasing global emphasis on environmental protection and carbon emission control, the accuracy and reliability of emission monitoring data from the thermal power industry, a major source of carbon dioxide emissions, have become particularly crucial. CEMS (Continuous Emission Monitoring System), a type of continuous emission monitoring system, is widely used in monitoring exhaust emissions from thermal power plants. However, in actual operation, anomalies frequently appear in CEMS monitoring data. These anomalies not only affect the accuracy and reliability of the monitoring data but may also mislead environmental protection authorities' decision-making.
[0003] The CEMS carbon monitoring system mainly consists of four parts: a particulate matter monitoring unit, a gaseous pollutant monitoring unit, a flue gas parameter monitoring unit, and a data acquisition and processing unit. This system can track and measure the concentration of pollutants in flue gas emitted from boilers, industrial furnaces, incinerators, etc., in real time, and upload the data to local environmental management departments, providing a scientific basis for environmental supervision. However, anomalies in CEMS carbon monitoring data are relatively common in actual operation. These anomalies may be caused by various reasons, including equipment failure, improper calibration, and environmental interference.
[0004] Traditional anomaly detection methods rely heavily on manual intervention, which is not only inefficient but also prone to errors. Furthermore, when processing large volumes of data, manual methods may miss some anomalies due to the sheer volume, often failing to accurately identify abnormal data and impacting the accuracy and reliability of monitoring data. With increasingly stringent environmental standards, the accuracy and stability of CEMS carbon monitoring have become crucial for ensuring emission compliance. Therefore, developing a method and system capable of automatically identifying and processing abnormal CEMS carbon monitoring data is of great significance for improving the accuracy and efficiency of environmental monitoring in the thermal power industry.
[0005] In publicly available literature and materials, some studies have made relevant attempts to automatically identify anomalies in carbon monitoring data. For example, in related technologies, patent application document CN118171771A proposes to extract the temporal content variation characteristics of gas monitoring parameters and the temporal variation characteristics of gas emission flow based on deep learning technology, and decode the carbon emission amount based on the fusion characteristics of the two, so that timely warnings can be issued when carbon emissions exceed a threshold. However, this method requires setting a threshold, and the warning effect largely depends on the selection of the threshold. The paper "Research on Time Series Prediction Model Based on Deep Spatiotemporal Characteristics, Xiao Yuteng, Doctoral Dissertation" proposes four time series prediction models to improve the accuracy of time series prediction. The models have been validated on public datasets such as the meteorological dataset SML2010 and the stock dataset Nasdaq100, as well as time series datasets in the real-world application scenario of underground coal gasification. However, the multivariate time series prediction model based on convolutional-long short-term memory modules with time-dependent features proposed in this paper mainly expands the network's perception range by introducing dilated causal convolution and residual networks, but it does not fully consider the features of different time scales. It may have the problem of relying on features of a single scale, thereby increasing the risk of model overfitting. The paper "Design of a Real-Time Methane Concentration Detection Device Based on One-Dimensional Dentual Convolution, Experimental Technology and Management, Vol. 42, No. 1" designs a methane concentration detection device based on dilated convolution. It uses one-dimensional dilated convolution to extract second harmonic signal features, uses an inverse residual module to improve the model convergence speed and reduce training time, and adds a wide convolution kernel receptive field module to expand the receptive field and reduce model parameters while ensuring detection accuracy. However, this study also has the problem of not fully considering the characteristics of different time scales. Summary of the Invention
[0006] The technical problem to be solved by this invention is how to accurately locate outliers in carbon monitoring data.
[0007] The present invention solves the above-mentioned technical problems through the following technical means:
[0008] A method for detecting anomalies in carbon monitoring data is proposed, including:
[0009] Acquire real-time operating data of thermal power units, perform multi-scale feature extraction on the real-time operating data, and obtain preliminary features at multiple levels;
[0010] Multiple levels of preliminary features are input into the encoders of the corresponding levels, so that the encoders encode the received preliminary features and the encoded features output by the previous level encoder to obtain the encoded features of the encoder.
[0011] The encoded features corresponding to each level of encoder are input to the corresponding level of decoder, so that the decoder can decode the received encoded features and the decoded features output by the previous level decoder to obtain the decoded features of the decoder.
[0012] The decoding features output from each level of decoder are concatenated to obtain a feature map, and the predicted carbon dioxide concentration is calculated based on the feature map.
[0013] By comparing the predicted carbon dioxide concentration with the measured carbon dioxide concentration, outliers in the carbon monitoring data can be identified.
[0014] Furthermore, the operating data adopts multi-dimensional operating data, which includes at least active power, oxygen content, flue gas flow rate, and sulfur dioxide concentration. After acquiring the real-time operating data of the thermal power unit, the method further includes:
[0015] The multidimensional time series composed of the multidimensional operational data is divided into several time series blocks according to the time length T.
[0016] Furthermore, the multi-scale feature extraction of real-time operational data yields preliminary features at multiple levels, including:
[0017] Wavelet transform is used to extract multi-scale features from real-time running data, resulting in preliminary features at multiple levels.
[0018] Further, the step of inputting preliminary features at multiple levels into the encoder at the corresponding level, so that the encoder encodes the received preliminary features with the encoded features output by the previous level encoder to obtain the encoder's encoded features, includes:
[0019] Multiple levels of preliminary features are input into the encoders of the corresponding levels, so that the encoders fuse the received preliminary features with the encoded features output by the previous level encoder to obtain the first fused feature.
[0020] Add location information to the first fused feature to obtain a location-guided first fused feature;
[0021] The scaling dot product attention mechanism is used to process the first fusion feature guided by position, and the obtained first multi-head fusion feature is merged with the preliminary feature of the corresponding level as the encoding feature of the encoder.
[0022] Furthermore, the formula for the first fusion feature is expressed as:
[0023]
[0024] In the formula, The first fused feature corresponding to the k-th level encoder. The k-th level preliminary feature received by the k-th level encoder. Let K be the encoded features output by the (k-1)th level encoder. Concat() represents a concatenation operation along the feature map dimension, diConv() represents dilated convolution, and K is the total number of encoders.
[0025] Furthermore, the formula for the encoding features of the encoder is expressed as:
[0026]
[0027] In the formula, For the coding features of the k-th level encoder, The output result obtained by passing the first fused feature guided by the position corresponding to the k-th level encoder through the h-th attention head. For the initial features of the k-th level received by the k-th level encoder, Concat() represents the connection operation along the feature map dimension, and FNN() represents a fully connected network.
[0028] Further, the step of inputting the encoded features corresponding to each level of encoder to the corresponding level of decoder, so that the decoder performs decoding processing on the received encoded features and the decoded features output by the previous level decoder to obtain the decoded features of the decoder, includes:
[0029] The encoded features corresponding to each level of encoder are input to the corresponding level of decoder, so that the decoder fuses the received encoded features with the decoded features output by the previous level decoder to obtain the second fused feature;
[0030] Add location information to the second fused feature to obtain a location-guided second fused feature;
[0031] The position-guided second fusion feature is processed using a scaled dot product attention mechanism, and the resulting second multi-head fusion feature is merged with the encoded feature output by the corresponding level encoder as the decoder's decoding feature.
[0032] Furthermore, the formula for the second fusion feature is expressed as:
[0033]
[0034] In the formula, The second fused feature corresponding to the k-th level decoder. The encoded features output by the k-th level encoder. Let K be the decoded features output by the (k+1)th level decoder. Concat() represents the concatenation operation along the feature map dimension, diConv() represents dilated convolution, and K is the total number of decoders.
[0035] Furthermore, the formula for the decoding features output by the decoder is expressed as:
[0036]
[0037] In the formula, The decoding features output by the k-th level decoder. The output result obtained by passing the second fused feature guided by the position corresponding to the k-th level decoder through the h-th attention head. For the encoded features output by the k-th level encoder, Concat() represents the concatenation operation along the feature map dimension, and FNN() represents a fully connected network.
[0038] Further, the step of concatenating the decoding features output from each level of decoder to obtain a feature map, and calculating the predicted carbon dioxide concentration based on the feature map, includes:
[0039] Will Connect and perform convolution to obtain feature map H. s ,in, Let K be the decoding feature output by the k-th level decoder, where K is the total number of decoders.
[0040] For feature map H s Gaussian modulation is performed to obtain the modulated feature map H. g The formula is expressed as:
[0041]
[0042] In the formula, For indicator functions, τ is a fixed threshold. This represents Gaussian modulation, where μ and σ are H, respectively. s The mean and variance; ⊙ indicates element-wise multiplication;
[0043] The predicted carbon dioxide concentration is calculated based on the modulated feature map.
[0044] Furthermore, the comparison between the predicted and measured carbon dioxide concentrations to locate outliers in the carbon monitoring data includes:
[0045] Calculate the difference between the predicted carbon dioxide concentration and the measured carbon dioxide concentration, and identify the measured carbon dioxide concentration that exceeds a set threshold as an outlier.
[0046] In addition, the present invention also proposes a carbon monitoring data anomaly detection device, including a data monitoring module, a concentration prediction module and an anomaly detection module connected in sequence. The concentration prediction module includes a multi-scale feature extraction unit, a number of encoders connected in sequence and a number of decoders connected in sequence, and a modulation unit. The multi-level outputs of the multi-scale feature extraction unit are respectively connected to the encoders of the corresponding levels, and the outputs of each encoder are respectively connected to the decoders of the corresponding levels.
[0047] The data monitoring module is used to acquire real-time operating data of the thermal power unit;
[0048] The multi-scale feature extraction unit is used to extract multi-scale features from real-time running data, and the resulting preliminary features at multiple levels are input into the encoders at the corresponding levels.
[0049] The encoder is used to encode the received preliminary features and the encoded features output by the previous level encoder to obtain the encoded features of the encoder, which are then input to the corresponding level decoder.
[0050] The decoder is used to decode the received encoded features and the decoded features output by the previous decoder to obtain the decoded features of the decoder.
[0051] The modulation unit is used to connect the decoding features output from each stage of the decoder to obtain a feature map, and calculate the predicted value of carbon dioxide concentration based on the feature map;
[0052] The anomaly detection module is used to compare the predicted carbon dioxide concentration with the measured carbon dioxide concentration to locate abnormal values in the carbon monitoring data.
[0053] Furthermore, the present invention also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the carbon monitoring data anomaly detection method as described above.
[0054] The advantages of this invention are:
[0055] (1) Since the operating data of thermal power units typically includes multiple dimensions, and this operating data is generated continuously over time, it has obvious time series characteristics. There is a time dependency between data at different time points. Since carbon emissions are closely related to the operating status of thermal power units, and with the periodic changes in electricity demand, the carbon emissions of thermal power units also have periodic characteristics, which are reflected in the monitoring data as multi-scale features. Moreover, the operating status and carbon emissions of thermal power units are affected by a variety of factors, and the data may contain complex patterns and relationships. Therefore, this invention first extracts the multi-scale preliminary features of the real-time operating data of thermal power units. Through multi-scale feature extraction, the data can be decomposed into multiple levels of preliminary features, thereby better capturing changes at different time scales. Then, the multi-scale preliminary features are input into the encoders of the corresponding levels, which can enhance the... The model's ability to capture data changes across different time scales is enhanced by the cascaded structure of the encoder and decoder, which encodes and decodes features at different levels. Each encoder level focuses on processing features at a specific scale and fuses them with the output of the previous level, gradually building richer feature representations. This structure effectively captures feature changes from local to global perspectives, enhancing the model's understanding of complex data patterns. Through multi-level feature extraction and reconstruction, it avoids over-reliance on features at a single scale, reducing the risk of overfitting and effectively capturing temporal and spatial features in the input sequence. The encoder progressively extracts and refines features through multi-level encoding, while the decoder progressively reconstructs feature maps through multi-level decoding. This progressive refinement process effectively reduces the impact of noise and interference, improving the model's accuracy in detecting outliers. Therefore, this structure better adapts to the complexity and diversity of thermal power unit operating data, improving the model's generalization ability and thus more accurately locating outliers.
[0056] (2) By employing advanced techniques such as dilated convolution, multi-head attention mechanism and position embedding in the encoder, the temporal dependence and spatial features in the input sequence can be effectively captured. This is because the encoder and decoder can progressively pass and process the temporal dependence in the time series. By introducing positional information (such as position embedding) in the encoder and decoder, the model can better capture the temporal order and time interval information of the data. Furthermore, by using techniques such as multi-head attention mechanism to model the spatial features between different dimensions, it is possible to better understand the interaction between different operating parameters, thereby more accurately identifying abnormal data.
[0057] (3) The present invention uses Gaussian modulation to modulate the feature map output by the model, and identifies uncertain regions by setting an indicator function and a fixed threshold, thereby locating outliers more accurately.
[0058] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0059] Figure 1 This is a flowchart illustrating a method for detecting anomalies in carbon monitoring data proposed in one embodiment of the present invention.
[0060] Figure 2 This is a schematic diagram of the structure of a carbon monitoring data anomaly detection device proposed in one embodiment of the present invention. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] like Figure 1 As shown, the first embodiment of the present invention proposes a method for detecting anomalies in carbon monitoring data, the method comprising the following steps:
[0063] S10. Obtain real-time operating data of thermal power units, perform multi-scale feature extraction on the real-time operating data, and obtain preliminary features at multiple levels.
[0064] It should be noted that this embodiment extracts multi-scale features from the real-time operating data of thermal power units and inputs the extracted preliminary features at multiple levels into the encoders at the corresponding levels to enhance the model's ability to capture data changes at different time scales.
[0065] It should be noted that, since the acquired real-time operating data of thermal power units is relatively long and covers a wide time range, this embodiment divides the real-time operating data into several sequence blocks with a time length of T before inputting them into the model for processing.
[0066] S20. Input the preliminary features of multiple levels into the encoder of the corresponding level, so that the encoder encodes the received preliminary features and the encoded features output by the previous level encoder to obtain the encoded features of the encoder.
[0067] S30. Input the encoding features corresponding to each level of encoder to the corresponding level of decoder, so that the decoder can decode the received encoding features and the decoding features output by the previous level decoder to obtain the decoder's decoding features.
[0068] S40. Connect the decoding features output from each level of decoder to obtain a feature map, and calculate the predicted carbon dioxide concentration based on the feature map;
[0069] S50. Compare the predicted carbon dioxide concentration with the measured carbon dioxide concentration to locate outliers in the carbon monitoring data.
[0070] It should be noted that thermal power unit operation data and its carbon emission data often have the following characteristics: (1) Multi-dimensional: Thermal power unit operation data usually includes multiple dimensions, such as active power, oxygen content, flue gas flow rate, sulfur dioxide concentration, etc. These data together reflect the operating status and carbon emission of thermal power units. (2) Time series: These data are generated continuously over time and have obvious time series characteristics. There is a time dependence between data at different time points. (3) Periodicity: Since carbon emission is closely related to the operating status of thermal power units, and with the periodic changes in electricity demand, the carbon emission of thermal power units also has periodic characteristics, such as seasonal changes, annual changes and long-term trend changes, which are reflected in the monitoring data as multi-scale characteristics. (4) Complexity: The operating status and carbon emission of thermal power units are affected by a variety of factors, and the data may contain complex patterns and relationships. (5) Sparsity of abnormal data: Under normal operating conditions, there are relatively few abnormal data, but once an anomaly occurs, it may have a serious impact on the environment and equipment.
[0071] Therefore, thermal power unit operating data contains features at different time scales, such as short-term fluctuations and long-term trends. This embodiment utilizes multi-scale feature extraction (such as wavelet transform) to decompose the data into multiple levels of preliminary features, thereby better capturing changes at different time scales. The cascaded structure of the encoder and decoder can encode and decode features at different levels separately. Each encoder level can focus on processing features at a specific scale and fuse them with the output of the previous encoder level, thus gradually constructing a richer feature representation. This structure effectively captures feature changes from local to global, enhancing the model's ability to understand complex data patterns. The encoder gradually extracts and refines features through multi-level encoding, while the decoder gradually reconstructs the feature map through multi-level decoding. This gradual refinement process effectively reduces the impact of noise and interference, improving the model's accuracy in detecting abnormal data.
[0072] Therefore, the cascaded encoder and decoder structure can avoid over-reliance on single-scale features through multi-level feature extraction and reconstruction, thereby reducing the risk of model overfitting. This structure can better adapt to the complexity and diversity of thermal power unit operation data and improve the model's generalization ability.
[0073] As a further preferred technical solution, the operating data adopts multi-dimensional operating data, which includes at least active power, oxygen content, flue gas flow rate, and sulfur dioxide concentration; correspondingly, after acquiring the real-time operating data of the thermal power unit, the method further includes the following steps:
[0074] The multidimensional time series composed of the multidimensional operational data is divided into several time series blocks according to the time length T.
[0075] It should be noted that this embodiment uses a modeling method combining multi-dimensional operating data of thermal power units and carbon emission concentration monitoring data to identify outliers, effectively avoiding the problem of misjudging normal mutation values caused by identifying outliers from a single carbon monitoring data.
[0076] As a further preferred technical solution, in step S10, multi-scale feature extraction is performed on the real-time running data to obtain preliminary features at multiple levels, specifically including:
[0077] Wavelet transform is used to extract multi-scale features from real-time running data, resulting in preliminary features at multiple levels.
[0078] Specifically, taking the time series block X of the running data as an example. s Taking multi-scale feature extraction as an example, extract X s The preliminary multi-scale features are obtained by wavelet transform of a certain wavelet basis, as follows:
[0079]
[0080] in, This represents the result of a wavelet transform with a scale of 2 raised to the power of k. With X s Same dimensions This is referred to as the preliminary feature of the k-th level wavelet transform.
[0081] As a further preferred technical solution, step S20: inputting the preliminary features of multiple levels into the encoder of the corresponding level, so that the encoder encodes the received preliminary features with the encoded features output by the previous level encoder to obtain the encoded features of the encoder, specifically including the following steps:
[0082] S21. Input the preliminary features of multiple levels into the encoder of the corresponding level, so that the encoder fuses the received preliminary features with the encoded features output by the previous level encoder to obtain the first fused feature;
[0083] It should be noted that each of the K encoders receives the preliminary features of its corresponding level as input; that is, the k-th level encoder receives the preliminary features. As input; wherein, the first-level encoder receives preliminary features. As input, the 2nd to Kth encoders simultaneously receive the encoded features output by their previous level encoder and fuse the received preliminary features with the encoded features output by the previous level encoder. For example, the kth level encoder simultaneously receives... The encoded features output by the (k-1)th level encoder To integrate.
[0084] It should be noted that each encoder level can focus on processing features at a specific scale. By fusing the initial features and the encoded features of the previous encoder through concatenation and dilated convolution along the feature map dimension, the first fused features are obtained. This process gradually builds richer feature representations. This structure can effectively capture feature changes from local to global and enhance the model's ability to understand complex data patterns.
[0085] S22. Add location information to the first fusion feature to obtain the location-guided first fusion feature;
[0086] Specifically, this embodiment uses learnable location information Added to the first fusion feature In the process, the first fusion feature guided by location is obtained.
[0087]
[0088] This embodiment can effectively capture the temporal dependence in the input sequence by adding location information to the first fusion feature.
[0089] It should be noted that the encoder and decoder can progressively pass on and process the temporal dependencies in the time series. By introducing positional information (such as positional embedding) into the encoder and decoder, the model can better capture the temporal order and time interval information of the data.
[0090] It should be noted that for the first-stage encoder, the initial features it receives... This is considered the first fusion feature.
[0091] It should be noted here that if it is a first-stage encoder, there are no output features from the previous stage encoder; the received preliminary features are directly used. This is considered the first fusion feature.
[0092] S23. The scaling dot product attention mechanism is used to process the first fusion feature guided by position, and the obtained first multi-head fusion feature is merged with the preliminary feature of the corresponding level as the encoding feature of the encoder.
[0093] It should be noted that multi-head attention mechanisms and other techniques are used to model spatial features across different dimensions. This structure allows for a better understanding of the interactions between different operating parameters, thereby enabling more accurate identification of anomalous data.
[0094] As a further preferred technical solution, the formula for the first fusion feature obtained in step S21 is expressed as:
[0095]
[0096] In the formula, The first fused feature corresponding to the k-th level encoder. The k-th level preliminary feature received by the k-th level encoder. Let K be the encoded features output by the (k-1)th level encoder. Concat() represents a concatenation operation along the feature map dimension, diConv() represents dilated convolution, and K is the total number of encoders.
[0097] It should be understood that in this embodiment, a splicing layer and a dilated convolutional layer can be set in the encoder. The splicing layer connects the preliminary features and the encoded features output by the previous level encoder along the feature dimension. The connected features are then input into the dilated convolution to obtain the first fused feature corresponding to the k-th level encoder.
[0098] As a further preferred technical solution, the dilated convolution in this embodiment adopts standard convolution or depthwise separable convolution. Among them, standard convolution is the most common type of convolution. It extracts features by sliding a fixed-size filter (kernel) on the input feature map. It has wide applications in image recognition, speech processing and other fields.
[0099] Depthwise separable convolution breaks down standard convolution into two steps: depthwise convolution and pointwise convolution. Depthwise convolution applies a convolution kernel independently to each input channel, while pointwise convolution applies a 1×1 convolution kernel to the result of depthwise convolution to combine channel information.
[0100] Specifically, the combined channel information consists of different channels of the input feature map. Depthwise convolution performs convolution operations on each channel of the input feature map separately, without crossing channels; while pointwise convolution uses a 1×1 convolution kernel to linearly combine the results of depthwise convolution to obtain the output that fuses the channel information.
[0101] As a further preferred technical solution, step S23: processing the position-guided first fusion feature using a scaled dot product attention mechanism, and merging the obtained first multi-head fusion feature with the corresponding level of preliminary features as the encoding feature of the encoder, specifically including:
[0102] The first fusion feature guided by location The input to the Scaled Dot-ProductAttention (SDA) module is as follows:
[0103]
[0104] Among them, Q,K e V represents Query, Key, and Value; Linear represents the linear transformation layer; Softmax represents the normalized exponential function; Attention represents the attention operation; d k For K e The dimension is , and the superscript T indicates transpose.
[0105] The scaling dot product attention module employs a multi-head attention mechanism, which divides the input features into multiple heads. Each head independently undergoes the scaling dot product attention operation. Finally, the outputs of all heads are concatenated, and a linear transformation is applied to obtain the final output. Therefore:
[0106] Q h ,K h V h =Linear(Q,K) e ,V) h
[0107]
[0108] in, The first fusion feature guided by the position in the k-th level encoder The output obtained by inputting the h-th attention head.
[0109] Multiple The first multi-head fusion feature and the primary feature obtained by fusion By merging, we obtain the encoder's encoding features:
[0110]
[0111] in, This represents the encoded features output by the k-th level encoder, where FNN is a fully connected network.
[0112] As a further preferred technical solution, step S30: inputting the encoded features corresponding to each level of encoder to the corresponding level of decoder, so that the decoder performs decoding processing on the received encoded features and the decoded features output by the previous level decoder to obtain the decoded features of the decoder, specifically including the following steps:
[0113] S31. Input the encoding features corresponding to each level of encoder into the decoder of the corresponding level, so that the decoder fuses the received encoding features with the decoded features output by the previous-level decoder to obtain a second fused feature;
[0114] It should be noted that the number and position of the decoders set in this embodiment correspond to the encoders, and there are also levels 1 to K. When 1 ≤ k < K, the k-th level decoder receives the encoding features output by the k-th level encoder and the output decoded features of the previous-level decoder and fuses them to obtain a second fused feature When k = K, the K-th level decoder only receives the encoding features output by the K-th level encoder The encoding features can be used as the second fused feature of the K-th level decoder.
[0115] S32. Add position information to the second fused feature to obtain a position-guided second fused feature;
[0116] Specifically, similar to the encoder, in this embodiment, position information of different levels is embedded into the second fused feature to obtain a position-guided second fused feature
[0117] S33. Process the position-guided second fused feature using the scaled dot-product attention mechanism, and merge the obtained second multi-head fused feature with the encoding features output by the encoder of the corresponding level as the decoded feature of this decoder.
[0118] As a further preferred technical solution, in the step S31, the formula representation of the obtained second fused feature is:
[0119]
[0120] In the formula, is the second fused feature corresponding to the k-th level decoder, is the encoding feature output by the k-th level encoder, is the decoded feature output by the (k + 1)-th level decoder, Concat() represents the concatenation operation along the feature map dimension, diConv() represents the dilated convolution, and K is the total number of decoders.
[0121] It should be understood that the dilated convolution set in the decoder in this embodiment can use standard convolution or depthwise separable convolution.
[0122] As a further preferred technical solution, step S33: The second fusion feature guided by position is processed using a scaled dot product attention mechanism, and the obtained second multi-head fusion feature is merged with the encoded feature output by the corresponding level encoder as the decoding feature of the decoder, specifically including:
[0123] The second fusion feature guided by location Input scaling dot product attention mechanism, in the k-th level decoder The output obtained by inputting the h-th attention head is Then the output results of multiple attention heads The second multi-head fusion feature obtained by fusion is combined with the encoded feature output by the k-th level encoder. By merging, we obtain the decoding features of the k-th level decoder:
[0124]
[0125] As a further preferred technical solution, step S40: concatenating the decoding features output by each level of decoder to obtain a feature map, and calculating the predicted carbon dioxide concentration value based on the feature map, specifically includes the following steps:
[0126] S41, will Connect and perform convolution to obtain feature map H. s for:
[0127]
[0128] in, Let K be the decoded feature output by the k-th level decoder, where K is the total number of decoders, and Conv represents ordinary convolution.
[0129] S42, Regarding feature map H s Gaussian modulation is performed to obtain the modulated feature map H. g The formula is expressed as:
[0130]
[0131] In the formula, For indicator functions, τ is a fixed threshold. This represents Gaussian modulation, activated by partitioning through an indicator function; μ and σ represent H, respectively. s The mean and variance; ⊙ indicates element-wise multiplication;
[0132] It should be noted that τ = 0.1 can be chosen, which means that activation values in the range [0.1, 0.9] are considered indeterminate.
[0133] This embodiment uses Gaussian modulation to modulate the feature map output by the model, which can reduce the impact of noise and effectively smooth local fluctuations in the feature map, making the features more stable. This smoothing effect helps the model to better capture long-term trends rather than be disturbed by short-term random fluctuations, which can significantly improve the performance and stability of the model. By setting an indicator function and a fixed threshold to identify uncertain regions, outliers can be located more accurately.
[0134] S43. Calculate the predicted carbon dioxide concentration based on the modulated feature map.
[0135] Specifically, in this embodiment, the modulated feature map H g By using extreme learning machine regression, H can be obtained. s Corresponding predicted carbon dioxide concentration
[0136] As a further preferred technical solution, step S50: comparing the predicted carbon dioxide concentration with the measured carbon dioxide concentration to locate outliers in the carbon monitoring data, specifically includes:
[0137] Calculate the difference between the predicted and measured carbon dioxide concentrations. and the difference The measured value y of carbon dioxide concentration that exceeds the set threshold Δ is determined to be an outlier.
[0138] It should be noted that, since the running data has a time duration, this embodiment can calculate the difference sequence. Carbon monitoring values in the sequence δ that exceed the manually set threshold Δ are identified as outliers.
[0139] It should be understood that the 3-sigma criterion can also be used to determine outliers in this embodiment. This embodiment does not specifically limit the method of determining outliers.
[0140] As a further preferred technical solution, before predicting carbon dioxide concentration based on thermal power unit operating data, this embodiment also needs to train the constructed carbon dioxide concentration prediction model. The training process is as follows:
[0141] Dataset Construction: Obtain historical multidimensional operating data of thermal power units, including at least four parameters: active power, oxygen content, flue gas flow rate, and sulfur dioxide concentration. Construct a multidimensional time series, and divide the multidimensional time series into blocks according to a certain time length T to obtain a series of time series blocks, denoted by X. s X represents a specific time series block. s The width is T, and the height is the number of unit operating parameters. The carbon dioxide concentration value y for the corresponding time period is obtained from the carbon monitoring system; the length of y is T, and a large number of (X) parameters are used.s The dataset consists of y) data.
[0142] The model is trained using a dataset, where the model consists of K encoders and decoders, and the model is formalized as follows:
[0143]
[0144] The input to the model is X. s The model output is the carbon dioxide concentration, i.e., it uses... This represents the predicted carbon dioxide concentration time series.
[0145] During model training, the general MSE loss function is used to measure the prediction results. The difference between y and y. Observe the change of MSE loss during training, and stop training when the following two situations occur: (1) the training loss no longer changes and fluctuates within a small range; (2) the loss value undergoes a significant sudden change.
[0146] In addition, such as Figure 2 As shown, the second embodiment of the present invention proposes a carbon monitoring data anomaly detection device, including a data monitoring module 10, a concentration prediction module 20 and an anomaly detection module 30 connected in sequence. The concentration prediction module 20 includes a multi-scale feature extraction unit, a plurality of encoders connected in sequence, a plurality of decoders connected in sequence and a modulation unit. The multi-level outputs of the multi-scale feature extraction unit are respectively connected to the encoders of the corresponding levels, and the outputs of each encoder are respectively connected to the decoders of the corresponding levels.
[0147] The data monitoring module 10 is used to acquire real-time operating data of the thermal power unit;
[0148] The multi-scale feature extraction unit is used to extract multi-scale features from real-time running data, and the resulting preliminary features at multiple levels are input into the encoders at the corresponding levels.
[0149] The encoder is used to encode the received preliminary features and the encoded features output by the previous level encoder to obtain the encoded features of the encoder, which are then input to the corresponding level decoder.
[0150] The decoder is used to decode the received encoded features and the decoded features output by the previous decoder to obtain the decoded features of the decoder.
[0151] The modulation unit is used to connect the decoding features output from each stage of the decoder to obtain a feature map, and calculate the predicted value of carbon dioxide concentration based on the feature map;
[0152] The anomaly detection module 30 is used to compare the predicted value of carbon dioxide concentration with the measured value of carbon dioxide concentration to locate abnormal values in the carbon monitoring data.
[0153] As a further preferred technical solution, the data monitoring module 10 can be implemented using a CEMS carbon system, which collects thermal power unit operation data through particulate matter monitoring units, gaseous pollutant monitoring units, and flue gas parameter monitoring units in the system.
[0154] As a further preferred technical solution, the multi-scale feature extraction unit is specifically used to perform multi-scale feature extraction on real-time running data using wavelet transform to obtain preliminary features at multiple levels.
[0155] As a further preferred technical solution, the encoded features output by the encoder are:
[0156]
[0157] in, The first fused feature corresponding to the k-th level encoder. The k-th level preliminary feature received by the k-th level encoder. The encoded features are the output of the (k-1)th level encoder.
[0158] As a further preferred technical solution, the decoding features output by the decoder are:
[0159]
[0160] Where, in the formula, The second fused feature corresponding to the k-th level decoder. The encoded features output by the k-th level encoder. This represents the decoding feature output by the (k+1)th level decoder.
[0161] As a further preferred technical solution, the modulation unit is specifically used for:
[0162] Will Connect and perform convolution to obtain feature map H. s for:
[0163]
[0164] in, Let K be the decoded feature output by the k-th level decoder, where K is the total number of decoders, and Conv represents ordinary convolution.
[0165] For feature map H s Gaussian modulation is performed to obtain the modulated feature map H. g The formula is expressed as:
[0166]
[0167] In the formula, For indicator functions, τ is a fixed threshold. This represents Gaussian modulation, activated by partitioning through an indicator function; μ and σ represent H, respectively. s The mean and variance; ⊙ indicates element-wise multiplication;
[0168] The predicted carbon dioxide concentration is calculated based on the modulated feature map.
[0169] As a further preferred technical solution, the detection module 30 is specifically used to calculate the difference between the predicted value of carbon dioxide concentration and the measured value of carbon dioxide concentration. and the difference The measured value y of carbon dioxide concentration that exceeds the set threshold Δ is determined to be an outlier.
[0170] It should be noted that other embodiments or specific implementation methods of the carbon monitoring data anomaly detection device described in this invention can refer to the above-mentioned method embodiments, and will not be repeated here.
[0171] Furthermore, the third embodiment of the present invention also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the carbon monitoring data anomaly detection method as described in the first embodiment above.
[0172] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0173] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0174] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0175] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0176] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for detecting anomalies in carbon monitoring data, characterized in that, include: Acquire real-time operating data of thermal power units, perform multi-scale feature extraction on the real-time operating data, and obtain preliminary features at multiple levels; Multiple levels of preliminary features are input into the encoders of the corresponding levels, so that the encoders encode the received preliminary features and the encoded features output by the previous level encoder to obtain the encoded features of the encoder. The encoded features corresponding to each level of encoder are input to the corresponding level of decoder, so that the decoder can decode the received encoded features and the decoded features output by the previous level decoder to obtain the decoded features of the decoder. The decoding features output from each level of decoder are concatenated to obtain a feature map, and the predicted carbon dioxide concentration is calculated based on the feature map. By comparing the predicted carbon dioxide concentration with the measured carbon dioxide concentration, outliers in the carbon monitoring data can be identified.
2. The carbon monitoring data anomaly detection method as described in claim 1, characterized in that, The operating data is multidimensional operating data, which includes at least active power, oxygen content, flue gas flow rate, and sulfur dioxide concentration. After acquiring the real-time operating data of the thermal power unit, the method further includes: The multidimensional time series composed of the multidimensional operational data is divided into several time series blocks according to the time length T.
3. The carbon monitoring data anomaly detection method as described in claim 1, characterized in that, The process involves multi-scale feature extraction from real-time operational data to obtain preliminary features at multiple levels, including: Wavelet transform is used to extract multi-scale features from real-time running data, resulting in preliminary features at multiple levels.
4. The carbon monitoring data anomaly detection method as described in claim 1, characterized in that, The step of inputting preliminary features at multiple levels into the encoder at the corresponding level, so that the encoder encodes the received preliminary features with the encoded features output by the previous level encoder to obtain the encoder's encoded features, includes: Multiple levels of preliminary features are input into the encoders of the corresponding levels, so that the encoders fuse the received preliminary features with the encoded features output by the previous level encoder to obtain the first fused feature. Add location information to the first fused feature to obtain a location-guided first fused feature; The scaling dot product attention mechanism is used to process the first fusion feature guided by position, and the obtained first multi-head fusion feature is merged with the preliminary feature of the corresponding level as the encoding feature of the encoder.
5. The carbon monitoring data anomaly detection method as described in claim 4, characterized in that, The formula for the first fusion feature is expressed as: In the formula, The first fused feature corresponding to the k-th level encoder. The k-th level preliminary feature received by the k-th level encoder. Let K be the encoded features output by the (k-1)th level encoder. Concat() represents a concatenation operation along the feature map dimension, diConv() represents dilated convolution, and K is the total number of encoders.
6. The carbon monitoring data anomaly detection method as described in claim 4, characterized in that, The formula for the encoding features of the encoder is expressed as follows: In the formula, For the coding features of the k-th level encoder, The output result obtained by passing the first fused feature guided by the position corresponding to the k-th level encoder through the h-th attention head. For the initial features of the k-th level received by the k-th level encoder, Concat() represents the connection operation along the feature map dimension, and FNN() represents a fully connected network.
7. The carbon monitoring data anomaly detection method as described in claim 1, characterized in that, The step of inputting the encoded features corresponding to each level of encoder to the corresponding level of decoder, so that the decoder performs decoding processing on the received encoded features and the decoded features output by the previous level decoder to obtain the decoded features of the decoder, includes: The encoded features corresponding to each level of encoder are input to the corresponding level of decoder, so that the decoder fuses the received encoded features with the decoded features output by the previous level decoder to obtain the second fused feature; Add location information to the second fused feature to obtain a location-guided second fused feature; The position-guided second fusion feature is processed using a scaled dot product attention mechanism, and the resulting second multi-head fusion feature is merged with the encoded feature output by the corresponding level encoder as the decoder's decoding feature.
8. The carbon monitoring data anomaly detection method as described in claim 7, characterized in that, The formula for the second fusion feature is expressed as follows: In the formula, The second fused feature corresponding to the k-th level decoder. The encoded features output by the k-th level encoder. Let K be the decoded features output by the (k+1)th level decoder. Concat() represents the concatenation operation along the feature map dimension, diConv() represents dilated convolution, and K is the total number of decoders.
9. The carbon monitoring data anomaly detection method as described in claim 7, characterized in that, The formula for the decoding features output by the decoder is expressed as follows: In the formula, The decoding features output by the k-th level decoder. The output result obtained by passing the second fused feature guided by the position corresponding to the k-th level decoder through the h-th attention head. For the encoded features output by the h-th level encoder, Concat() represents the concatenation operation along the feature map dimension, and FNN() represents a fully connected network.
10. The carbon monitoring data anomaly detection method as described in claim 1, characterized in that, The step of concatenating the decoding features output from each level of decoder to obtain a feature map, and calculating the predicted carbon dioxide concentration based on the feature map, includes: Will Connect and perform convolution to obtain feature map H. s ,in, Let K be the decoding feature output by the k-th level decoder, where K is the total number of decoders. For feature map H s Gaussian modulation is performed to obtain the modulated feature map H. g The formula is expressed as: In the formula, For indicator functions, τ is a fixed threshold. This represents Gaussian modulation, where μ and σ are H, respectively. s The mean and variance; ⊙ indicates element-wise multiplication; The predicted carbon dioxide concentration is calculated based on the modulated feature map.
11. The carbon monitoring data anomaly detection method as described in claim 10, characterized in that, The step of comparing the predicted carbon dioxide concentration with the measured carbon dioxide concentration to locate outliers in the carbon monitoring data includes: Calculate the difference between the predicted carbon dioxide concentration and the measured carbon dioxide concentration, and identify the measured carbon dioxide concentration that exceeds a set threshold as an outlier.
12. A carbon monitoring data anomaly detection device, characterized in that, It includes a data monitoring module, a concentration prediction module and an anomaly detection module connected in sequence. The concentration prediction module includes a multi-scale feature extraction unit, a number of encoders connected in sequence and a number of decoders connected in sequence. The multi-level outputs of the multi-scale feature extraction unit are respectively connected to the encoders at the corresponding levels, and the outputs of each encoder are respectively connected to the decoders at the corresponding levels. The data monitoring module is used to acquire real-time operating data of the thermal power unit; The multi-scale feature extraction unit is used to extract multi-scale features from real-time running data, and the resulting preliminary features at multiple levels are input into the encoders at the corresponding levels. The encoder is used to encode the received preliminary features and the encoded features output by the previous level encoder to obtain the encoded features of the encoder, which are then input to the corresponding level decoder. The decoder is used to decode the received encoded features and the decoded features output by the previous decoder to obtain the decoded features of the decoder. The modulation unit connects the decoding features output from each stage of the decoder to obtain a feature map, and calculates the predicted value of carbon dioxide concentration based on the feature map; The anomaly detection module is used to compare the predicted carbon dioxide concentration with the measured carbon dioxide concentration to locate abnormal values in the carbon monitoring data.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the carbon monitoring data anomaly detection method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Carbon emission detection and early warning method and system
CN118171771A