A volatile organic compound treatment facility operation state monitoring method and system
By preprocessing data and identifying multidimensional feature vectors for volatile organic compound (VOC) treatment facilities, combined with target process thresholds, unsupervised anomaly detection, and time series decomposition, the problem of difficulty in identifying facility operating status in existing technologies has been solved, achieving more accurate and timely status monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA NAT ENVIRONMENTAL MONITORING CENT
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies are insufficient to comprehensively and accurately identify the operating status of volatile organic compound (VOC) treatment facilities, resulting in delayed identification of equipment malfunctions and failure to address them in a timely manner, thus posing a risk of environmental pollution.
By acquiring emission monitoring data and equipment operating parameters from volatile organic compound (VOC) treatment facilities, a multi-dimensional feature vector is generated after preprocessing. Then, a triple identification process is performed using the target process operating threshold, unsupervised anomaly detection mode, and time series decomposition mode. Finally, the fusion results determine the facility's operating status.
It enables a multi-dimensional comprehensive assessment of the operational status of volatile organic compound (VOC) treatment facilities, improving the accuracy and timeliness of identification and reducing the risk of environmental pollution.
Smart Images

Figure CN122489907A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method and system for monitoring the operational status of volatile organic compound (VOC) treatment facilities. Background Technology
[0002] Volatile organic compounds (VOCs) are the main components that form fine particulate matter (PM2.5). 2.5 VOCs are key precursors to both oxygen (VOCs) and ozone (O3), posing a serious threat to air quality and human health. With the successive introduction of a series of environmental regulations and management policies by the state and relevant departments, the regulatory requirements for corporate VOCs emissions are becoming increasingly stringent. Ensuring the stable and efficient operation of treatment facilities has become a basic requirement for corporate compliance.
[0003] Currently, enterprises primarily rely on two methods for monitoring the operation of VOCs treatment facilities: manual inspections and single online emission concentration monitoring. Manual inspections are inefficient and slow to respond, making it difficult to promptly detect and address sudden equipment malfunctions. Single online monitoring devices only collect a few indicators such as VOCs emission concentrations, failing to comprehensively capture the operational status of the treatment facilities and hindering the establishment of correlation analysis between equipment operating parameters and emission data. When potential anomalies or minor malfunctions occur in the treatment facilities, existing monitoring methods struggle to identify them in a timely manner, only discovering problems after excessive emissions have occurred, thus posing potential environmental pollution risks. Summary of the Invention
[0004] In view of the above problems, this application provides a method and system for monitoring the operational status of volatile organic compound (VOC) treatment facilities, aiming to improve the accuracy and timeliness of identifying the operational status of VOC treatment facilities. The specific solution is as follows:
[0005] The first aspect of this application provides a method for monitoring the operational status of a volatile organic compound (VOC) treatment facility, comprising:
[0006] Obtain emission monitoring data and equipment operating parameters from volatile organic compound (VOC) treatment facilities;
[0007] The emission outlet monitoring data and the equipment operating parameters are preprocessed to obtain a multidimensional feature vector;
[0008] The multidimensional feature vector is subjected to a first identification process based on the target process operation threshold to obtain a first result;
[0009] The multidimensional feature vector is subjected to a second identification process based on an unsupervised anomaly detection mode to obtain a second result; the unsupervised anomaly detection mode represents a processing method that identifies abnormal states by calculating sample anomaly scores.
[0010] The multidimensional feature vector is subjected to a third identification process based on the time series decomposition model to obtain a third result; the time series decomposition model represents the decomposition of key parameters into trend terms, seasonal terms and residual terms, and the handling method of abnormal fluctuations is identified based on the residual terms analysis.
[0011] Based on the first result, the second result, and the third result, the operating status of the volatile organic compound treatment facility is determined.
[0012] In one possible implementation, acquiring emission monitoring data and equipment operating parameters of the volatile organic compound (VOC) treatment facility includes:
[0013] In response to the volatile organic compound treatment facility employing activated carbon adsorption technology, the collected inlet and outlet pressure difference, temperature, relative humidity, and flue gas velocity of the activated carbon adsorption process are determined as equipment operating parameters.
[0014] In response to the volatile organic compound treatment facility employing catalytic combustion technology, the collected combustion chamber temperature, catalyst bed temperature, and combustion aid dosage of the catalytic combustion process are determined as the equipment operating parameters;
[0015] In response to the volatile organic compound treatment facility adopting a regenerative thermal combustion process, the collected combustion chamber temperature and switching valve operating frequency of the regenerative thermal combustion process are determined as the equipment operating parameters;
[0016] The volatile organic compound concentration, flue gas flow rate, and flue gas temperature collected at the emission outlet are determined as the emission outlet monitoring data.
[0017] In one possible implementation, the preprocessing of the emission outlet monitoring data and the equipment operating parameters to obtain a multidimensional feature vector includes:
[0018] Outlier removal processing is performed on the raw data of the emission outlet monitoring data and the equipment operating parameters;
[0019] Smooth the data after removing outliers;
[0020] Feature construction is performed on the smoothed data to obtain basic features and derived features. The basic features represent the values of the smoothed emission outlet monitoring data and equipment operating parameters after standardization. The derived features represent the parameter change rate, parameter correlation and process efficiency calculated based on the basic features.
[0021] The basic features and the derived features are combined into a multidimensional feature vector.
[0022] In one possible implementation, the outlier removal processing of the raw data of the emission outlet monitoring data and the equipment operating parameters includes:
[0023] Based on the time-series characteristics of the emission outlet monitoring data and the equipment operating parameters, a target processing mode is adopted to remove outliers from the raw data. The target processing mode is characterized by calculating the mean and standard deviation of the raw data of the emission outlet monitoring data and the equipment operating parameters, and identifying outliers as values whose absolute value of the difference from the mean is greater than three times the standard deviation and removing them.
[0024] The smoothing process for the data after removing outliers includes:
[0025] To address the fluctuation characteristics of the operating parameters of the volatile organic compound (VOC) treatment facility, a moving average method is used to smooth the data after removing outliers. The size of the moving window in the smoothing process is dynamically set based on the parameter sampling frequency and the process response time.
[0026] In one possible implementation, the unsupervised anomaly detection mode includes an anomaly detection model constructed using the isolated forest algorithm, wherein the second identification processing of the multidimensional feature vector based on the unsupervised anomaly detection mode to obtain a second result includes:
[0027] The multidimensional feature vector is input into the anomaly detection model. The path length of each sample vector in the multidimensional feature vector is calculated in the decision tree constructed based on randomly selected features and split points in the anomaly detection model.
[0028] The anomaly score for each sample vector is determined based on the path length.
[0029] The second result is determined based on the comparison between the anomaly score and the preset anomaly threshold.
[0030] In one possible implementation, the third identification processing of the multidimensional feature vector based on the time series decomposition mode to obtain a third result includes:
[0031] The key parameters in the multidimensional feature vector are decomposed using the STL decomposition method to obtain the trend term, seasonal term, and residual term;
[0032] Calculate the mean and standard deviation of the residual terms;
[0033] If the difference between the residual term and the mean is greater than a preset multiple of the standard deviation, it is determined that there is abnormal fluctuation.
[0034] The third result is determined based on the abnormal fluctuation judgment results.
[0035] In one possible implementation, determining the operating status of the volatile organic compound (VOC) treatment facility based on the first result, the second result, and the third result includes:
[0036] Based on the pre-defined correspondence between the operating status and score of the volatile organic compound treatment facility, the first result, the second result, and the third result are quantitatively assigned values to obtain the first assigned value result, the second assigned value result, and the third assigned value result.
[0037] Based on the preset weight coefficients corresponding to each recognition process, the weight values corresponding to the first result, the second result, and the third result are determined, and the first assignment result, the second assignment result, and the third assignment result are weighted and summed based on the weight values to obtain the fusion score;
[0038] The operating status of the volatile organic compound (VOC) treatment facility is determined based on the comparison between the fusion score and the target fusion threshold.
[0039] In one possible implementation, the method further includes:
[0040] Based on the operating status and target warning level, determine the abnormal parameter information;
[0041] Based on the abnormal parameter information, fault diagnosis push information is determined so that the fault diagnosis push information is sent to the destination.
[0042] In one possible implementation, determining the fault diagnosis push information based on the abnormal parameter information includes:
[0043] The abnormal parameter information is input into a pre-built fault diagnosis model. The fault diagnosis model is a multi-classification model built based on the random forest algorithm. The fault diagnosis model uses historical abnormal parameters and corresponding fault causes as training samples and fault cause categories as output labels. By constructing multiple decision trees and integrating the voting results of each decision tree, a mapping relationship between abnormal parameters and fault causes is formed.
[0044] In response to the fault diagnosis model, the abnormal parameter information is processed to obtain the fault cause corresponding to the abnormal parameter information and the confidence level corresponding to each fault cause;
[0045] Based on the cause of the fault and the confidence level, a corresponding handling suggestion is matched from a preset handling suggestion library;
[0046] The cause of the fault, the confidence level, and the handling suggestions are combined into a fault diagnosis push message.
[0047] The second aspect of this application provides a monitoring system for the operational status of a volatile organic compound (VOC) treatment facility, comprising:
[0048] The data acquisition module is used to acquire emission monitoring data and equipment operating parameters of volatile organic compound (VOC) treatment facilities.
[0049] The data preprocessing module is used to preprocess the emission outlet monitoring data and the equipment operating parameters to obtain a multi-dimensional feature vector;
[0050] The first identification module is used to perform a first identification process on the multidimensional feature vector based on the target process operation threshold to obtain a first result;
[0051] The second identification module is used to perform a second identification process on the multidimensional feature vector according to the unsupervised anomaly detection mode to obtain a second result; the unsupervised anomaly detection mode represents the processing method of identifying abnormal states by calculating sample anomaly scores.
[0052] The third identification module is used to perform third identification processing on the multidimensional feature vector based on the time series decomposition mode to obtain a third result; the time series decomposition mode represents the decomposition of key parameters into trend terms, seasonal terms and residual terms, and the processing method for identifying abnormal fluctuations based on residual term analysis.
[0053] The fusion module is used to determine the operating status of the volatile organic compound treatment facility based on the first result, the second result, and the third result.
[0054] By employing the above technical solution, this application provides a method and system for monitoring the operational status of volatile organic compound (VOC) treatment facilities. This method obtains a multi-dimensional feature vector by preprocessing emission outlet monitoring data and equipment operating parameters. The multi-dimensional feature vector is then subjected to triple identification processing based on the target process operating threshold, unsupervised anomaly detection mode, and time series decomposition mode. Finally, the three identification results are fused to determine the operational status of the treatment facility. This achieves a multi-dimensional comprehensive assessment of the operational status of the treatment facility, overcoming the technical problem that a single monitoring indicator or single identification method cannot comprehensively and accurately reflect the true operating condition of the equipment. This improves the accuracy and timeliness of identifying the operational status of VOC treatment facilities. Attached Figure Description
[0055] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0056] Figure 1A flowchart illustrating a method for monitoring the operational status of a volatile organic compound (VOC) treatment facility provided in this application;
[0057] Figure 2 This is a schematic diagram of the structure of a volatile organic compound (VOC) treatment facility operation status monitoring system provided in this application. Detailed Implementation
[0058] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0059] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0060] The terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0061] This application provides a method for monitoring the operational status of volatile organic compound (VOC) treatment facilities. This method can be applied to online monitoring platforms or environmental supervision systems for VOC treatment facilities. It is used for real-time monitoring and intelligent determination of the operational status of treatment facilities using different processes such as activated carbon adsorption, catalytic combustion, or regenerative thermal combustion. For example, it can be used for real-time monitoring and anomaly identification of treatment facilities in industries with high VOC emissions, such as chemical, printing, and coating industries.
[0062] In the embodiments of this application, volatile organic compounds (VOCs) may include various volatile organic compounds emitted during industrial production processes, such as benzene compounds, hydrocarbons, esters, aldehydes, ketones, etc. Corresponding VOC treatment facilities may include at least one of activated carbon adsorption facilities, catalytic combustion facilities, regenerative combustion facilities, zeolite rotor adsorption and concentration facilities, biotrickling filtration facilities, low-temperature plasma facilities, and photocatalytic oxidation facilities.
[0063] See Figure 1The illustration shows a flowchart of a method for monitoring the operational status of a volatile organic compound (VOC) treatment facility according to an embodiment of this application. The method may include the following steps:
[0064] S101. Obtain emission monitoring data and equipment operating parameters of volatile organic compound (VOC) treatment facilities.
[0065] Volatile organic compound (VOC) treatment facilities generate two main types of data during operation: one is the emission monitoring data from the treatment facility's outlets, which reflects the emissions of waste gas after treatment; the other is the operating parameters of the equipment itself, which reflect the working condition of the treatment facility.
[0066] Emission outlet monitoring data mainly includes parameters such as volatile organic compound concentration, flue gas flow rate, and flue gas temperature. These parameters are usually automatically collected by an online monitoring system installed on the emission pipeline at a fixed frequency (e.g., once per hour). Equipment operating parameters vary depending on the specific process used in the treatment facility, requiring the collection of differentiated key parameters tailored to the characteristics of different processes.
[0067] As a preferred approach, the volatile organic compound (VOC) treatment facility in this embodiment employs activated carbon adsorption technology. Activated carbon adsorption utilizes the porous structure of activated carbon to adsorb VOCs from waste gas, and its adsorption efficiency is influenced by various factors. To accurately assess the operating status of this process, it is necessary to collect key parameters reflecting the adsorption capacity of activated carbon. Therefore, in this step, in response to the adoption of activated carbon adsorption technology in the treatment facility, the inlet and outlet pressure difference, temperature, relative humidity, and flue gas velocity of the activated carbon adsorption process are determined as the equipment operating parameters. Specifically, the inlet and outlet pressure difference reflects the degree of blockage in the activated carbon layer, temperature affects adsorption efficiency, relative humidity affects the water absorption performance of activated carbon, and flue gas velocity determines the contact time between the waste gas and the activated carbon. These parameters are collected in real time by various sensors installed in the inlet and outlet pipes of the activated carbon adsorption box, inside the box, and before and after the fan.
[0068] In response to the adoption of catalytic combustion technology in volatile organic compound (VOC) treatment facilities, the collected data on combustion chamber temperature, catalyst bed temperature, and combustion aid dosage are determined as equipment operating parameters. Catalytic combustion oxidizes and decomposes VOCs at relatively low temperatures under the action of a catalyst. Combustion chamber temperature directly affects the reaction rate, catalyst bed temperature reflects the catalyst's activity state, and combustion aid dosage affects combustion efficiency. By collecting these parameters, the operating conditions of the catalytic combustion process can be comprehensively understood.
[0069] In response to the adoption of regenerative thermal oxidizer (RTO) technology in volatile organic compound (VOC) treatment facilities, the collected data on the combustion chamber temperature and switching valve operating frequency are determined as equipment operating parameters. RTO recovers heat generated during combustion through a heat storage medium, achieving highly efficient and energy-saving waste gas treatment. The combustion chamber temperature determines the oxidation and decomposition effect, while the switching valve operating frequency affects heat recovery efficiency and system stability. By collecting these parameters, the operating status of the RTO can be effectively monitored.
[0070] It should be noted that the above three process types and corresponding parameter acquisition methods are merely preferred examples of embodiments of this application and are not intended to limit the scope of protection of this application. When the treatment facility adopts other processes (such as zeolite rotor adsorption concentration, bio-trickling filtration, low-temperature plasma, photocatalysis, etc.), those skilled in the art can determine the corresponding key operating parameters as equipment operating parameters for acquisition based on the characteristics of the specific process.
[0071] Through step S101, this embodiment of the application realizes the collection of data across the entire chain of governance facilities, providing a data foundation for subsequent multi-dimensional analysis.
[0072] S102. Preprocess the emission outlet monitoring data and the equipment operating parameters to obtain a multidimensional feature vector.
[0073] The raw emission outlet monitoring data and equipment operating parameters are often isolated numerical sequences and may contain noise, outliers, or missing values, making them unsuitable for direct comprehensive judgment. Therefore, step S102 can be used to preprocess the raw data, converting it into a multi-dimensional feature vector that can comprehensively reflect the operating status of the treatment facilities.
[0074] The preprocessing process typically includes three stages: first, outlier removal, eliminating obviously erroneous data caused by sensor malfunctions or communication anomalies; second, smoothing, eliminating random noise interference; and third, feature construction, extracting characteristic indicators that can characterize the operational status from the raw data. Through these stages, the raw time-series data is transformed into a multi-dimensional feature vector containing basic and derived features. Basic features are the standardized values of the original parameters, eliminating the influence of dimensions; derived features are parameter change rates, parameter correlations, and process efficiency calculated based on the basic features, reflecting the dynamic relationships between parameters and the treatment effect. These features together constitute a multi-dimensional representation of the operational status of the treatment facility, providing a unified input format for subsequent identification and processing.
[0075] S103. Perform a first identification process on the multidimensional feature vector based on the target process operation threshold to obtain a first result.
[0076] Under normal operating conditions, various parameters of different treatment processes typically fall within specific threshold ranges. For example, activated carbon adsorption processes require temperatures below 40°C, relative humidity below 80%, and pressure differentials between 500-1500 Pa. When a parameter deviates from its normal threshold, it often indicates that the equipment has experienced some type of malfunction.
[0077] Based on these principles, a first identification process is performed on the multidimensional feature vector using preset target process operation thresholds. Specifically, each parameter in the multidimensional feature vector is compared one by one with the preset threshold of the corresponding process. Based on whether a parameter exceeds the threshold and the extent of the exceedance, the parameter deviation is marked as a slight deviation or a severe deviation. For example, if a parameter exceeds the threshold but by less than 10%, it can be marked as a slight deviation; if the exceedance reaches or exceeds 10%, it can be marked as a severe deviation. A first result is generated based on the deviation markings of all parameters. This result reflects whether there are obvious, known types of parameter anomalies in the treatment facility, serving as the basis for determining the operational status.
[0078] S104. Perform a second identification process on the multidimensional feature vector according to the unsupervised anomaly detection mode to obtain a second result.
[0079] Unsupervised anomaly detection model characterization identifies abnormal states by calculating sample anomaly scores. Fixed threshold determination can effectively identify explicit anomalies where parameters exceed limits, but it can be biased when all parameters are within the threshold range but the overall combination pattern is abnormal. For example, in activated carbon adsorption processes, temperature and humidity may be within normal ranges, but their combined relationship may deviate from normal operating patterns, indicating impending equipment failure. These latent anomalies require more intelligent methods for identification.
[0080] This step employs an unsupervised anomaly detection approach to address this issue. Unsupervised anomaly detection does not require pre-labeling of normal or anomalous samples; instead, it identifies outliers that deviate from the behavior of the majority of samples by analyzing the distribution characteristics of the data itself. Specifically, an unsupervised anomaly detection model is first constructed. Then, each sample vector from the multidimensional feature vector is input into the model, which calculates an anomaly score for each sample vector. A higher anomaly score indicates a greater deviation from the normal pattern, and a higher probability of an anomaly. Based on the comparison between the anomaly score and a preset anomaly threshold, a second result is generated. This result reflects whether the governance facility exhibits potential, non-obvious anomaly patterns, effectively complementing the first identification process.
[0081] S105. Perform a third identification process on the multidimensional feature vector based on the time series decomposition mode to obtain a third result.
[0082] Time series decomposition models characterize the decomposition of key parameters into trend, seasonal, and residual terms, and identify handling methods for abnormal fluctuations based on residual term analysis. The operating parameters of pollution control facilities often exhibit time-series characteristics; parameter changes are not only affected by the equipment's own condition but may also be influenced by the periodicity of external factors such as fluctuations in production conditions and changes in environmental factors. For example, increased output in certain production processes during specific time periods leads to higher exhaust gas concentrations, which in turn causes corresponding changes in the parameters of the pollution control facilities. Such periodic fluctuations are normal phenomena and should not be misjudged as equipment malfunctions.
[0083] These types of problems can be addressed using time series decomposition. Time series decomposition breaks down the original parameter sequence into the following components: a trend term, reflecting the long-term trend of the parameter, such as the slow decline in efficiency due to equipment aging; a seasonal term, reflecting the periodic fluctuations of the parameter, such as regular changes caused by day / night cycles, seasons, or production cycles; and a residual term, reflecting random fluctuations after removing trend and seasonal factors, which should normally follow a normal distribution. By analyzing the residual term, if a significant deviation is found (e.g., the difference from the mean exceeds two standard deviations), it indicates abnormal fluctuations in the parameter. These fluctuations cannot be explained by long-term trends or periodic patterns and are likely a warning sign of equipment failure. Based on the abnormal fluctuation judgment results of the residual analysis, a third result is generated. This result can capture gradual, trend-based anomalies, enabling early warning of equipment failure.
[0084] S106. Based on the first result, the second result, and the third result, determine the operating status of the volatile organic compound treatment facility.
[0085] The operational status of the governance facilities was identified from three different dimensions: the first result reflects explicit parameter anomalies, the second result reflects implicit pattern anomalies, and the third result reflects trend fluctuation anomalies. These three results each have their own emphasis and complement each other, but a single dimension's result may lead to misjudgment or omission. For example, a momentary sensor malfunction may cause a false alarm in the first result, which can be corrected by the second and third results; some abnormal patterns may initially only manifest as residual fluctuations (identified by the third result), without triggering threshold alarms (not identified by the first result) or increased anomaly scores (not identified by the second result). In such cases, combining the three results can achieve earlier early warning.
[0086] The specific integration method can employ a weighted voting mechanism, quantifying and assigning values to the three results (e.g., 1 point for good operation, 2 points for low risk, and 3 points for high risk), and then summing them according to preset weights (e.g., 0.4 for the first result, 0.35 for the second result, and 0.25 for the third result) to obtain an integration score. Then, by comparing the integration score with preset integration thresholds (e.g., 1.5 and 2.5), the final operational status is determined as good operation, low risk, or high risk. This multi-dimensional integration judgment allows for a more comprehensive and accurate assessment of the actual operational status of the governance facilities, avoiding the limitations of single-dimensional judgments.
[0087] This application provides a method for monitoring the operational status of volatile organic compound (VOC) treatment facilities. This method obtains a multi-dimensional feature vector by preprocessing emission outlet monitoring data and equipment operating parameters. The multi-dimensional feature vector is then subjected to triple identification processing based on the target process operating threshold, an unsupervised anomaly detection mode, and a time series decomposition mode. Finally, the three identification results are fused to determine the operational status of the treatment facility. This achieves a multi-dimensional comprehensive assessment of the facility's operational status, overcoming the technical problem that a single monitoring indicator or identification method cannot comprehensively and accurately reflect the true operating condition of the equipment. This improves the accuracy and timeliness of identifying the operational status of VOC treatment facilities.
[0088] In one possible implementation, the emission outlet monitoring data and equipment operating parameters are preprocessed to obtain a multi-dimensional feature vector, including:
[0089] S201. Perform outlier removal processing on the raw data of emission outlet monitoring data and equipment operating parameters.
[0090] In actual monitoring, due to occasional sensor malfunctions, signal transmission interruptions, or momentary external interference, the raw data collected may contain some outliers that significantly deviate from the normal range. Using these outliers directly in subsequent analysis without processing may lead to misjudgments. Therefore, it is necessary to remove outliers from the raw data.
[0091] In one implementation of this application, a target processing mode is used to remove outliers from the raw data, taking into account the time series characteristics of the emission outlet monitoring data and the equipment operating parameters. The target processing mode is based on the 3σ principle, and its processing logic is as follows: for the currently processed raw data sequence, the mean and standard deviation of the sequence are calculated; values whose absolute value of the difference from the mean is greater than three times the standard deviation are identified as outliers and removed.
[0092] The theoretical basis of the 3σ principle is that for data following a normal distribution, approximately 99.7% of the values fall within the range of the mean ± three standard deviations. Values outside this range are considered low-probability events and can be regarded as outliers. Applying the 3σ principle in the preprocessing of operational data from pollution control facilities can effectively identify and remove obviously erroneous data caused by sensor malfunctions or communication anomalies, while retaining normal fluctuation data reflecting the true operating conditions of the equipment. For example, a temperature sensor in an activated carbon adsorption facility collected 60 temperature data points in one hour (sampling frequency of 1 time / minute), and the mean of this set of data was calculated to be 35℃, with a standard deviation of 2℃. According to the 3σ principle, the normal temperature range should be 35±6℃, i.e., 29℃-41℃. If a certain temperature data point is 45℃, its difference from the mean is 10℃, which is greater than three standard deviations (6℃), then this data point is determined to be an outlier and is removed.
[0093] S202. Smooth the data after removing outliers.
[0094] Even after outlier removal, the data may still contain random fluctuations caused by measurement noise or minor variations in operating conditions. These random fluctuations may mask the true trend of the data and affect the accuracy of subsequent feature construction and anomaly identification. Therefore, it is necessary to smooth the data after outlier removal to eliminate random noise interference and highlight the original characteristics of the data.
[0095] In one implementation, to address the fluctuation characteristics of operating parameters of volatile organic compound (VOC) treatment facilities, a moving average method is used to smooth the data after removing outliers. The moving average method involves taking the average of several adjacent points as the new value for each data point in the time series, thereby reducing the impact of random fluctuations. The size of the moving window determines the degree of smoothing; a larger window results in a more pronounced smoothing effect, but may also lose details of the data's variation.
[0096] In this embodiment, the sliding window size is dynamically set based on the parameter sampling frequency and the process response time. The sampling frequency determines the temporal resolution of the data, while the process response time reflects how quickly parameter changes affect the process state. For example, for parameters that change relatively slowly, such as temperature, a larger sliding window can be set to effectively eliminate noise; for parameters that change rapidly, such as differential pressure, a smaller sliding window is needed to retain real-time data changes. Assuming a parameter is sampled at a frequency of 1 time per hour, and the process response time of this parameter is approximately 2-3 hours, the sliding window size can be set to 5, meaning the average of the current data point and the two data points before and after it is used as the smoothed new value. For boundary data points (such as the beginning or end of a sequence), a symmetrical filling method can be used, i.e., filling with data from symmetrical positions in the sequence to ensure that all data points are smoothed.
[0097] S203. Perform feature construction on the smoothed data to obtain basic features and derived features.
[0098] Even after outlier removal and smoothing, the data remains a raw numerical sequence and has not yet formed a feature system that comprehensively reflects the operational status of the treatment facilities. Therefore, it is necessary to construct features from this processed data and extract key feature indicators that can characterize the equipment's operational status.
[0099] The basic characteristic characterization represents the standardized values of emission outlet monitoring data and equipment operating parameters after smoothing. The purpose of standardization is to eliminate the dimensional influence between different parameters, enabling comparison and analysis of parameters with different dimensions on the same scale. The standardization calculation method is as follows: for each parameter, the current value is mapped to the 0-1 interval based on its historical minimum and maximum values within its normal operating range. Specifically, the standardized value = (current value - historical minimum) / (historical maximum - historical minimum). Through standardization, parameters with different dimensions such as temperature, pressure, and concentration are converted into dimensionless values, facilitating subsequent fusion analysis.
[0100] Derived features characterize parameter change rates, parameter correlations, and process efficiency calculated from basic features. These features reflect the dynamic relationships between parameters and the treatment effect, representing a deepening and expansion of the basic features. Parameter change rates reflect the trend of parameter changes over time, capturing dynamic information. For example, the hourly temperature change rate can be expressed as the difference between the current temperature and the temperature one hour ago divided by the time interval, used to monitor sudden temperature changes; the daily pressure difference change rate reflects the clogging rate of the activated carbon layer. Parameter correlation reflects the degree of association between different parameters, revealing their intrinsic connections. For example, in catalytic combustion processes, there should be a positive correlation between combustion chamber temperature and VOC concentration; as the combustion chamber temperature increases, VOC treatment efficiency improves, and emission concentration decreases accordingly. By calculating the Pearson correlation coefficient between combustion chamber temperature and VOC concentration, the normality of the correlation can be assessed. If the correlation coefficient deviates significantly from the normal range, it may indicate a decrease in catalyst activity or incomplete combustion. Process efficiency reflects the actual treatment effect of the treatment facility and is one of the core indicators for evaluating the operational status of the facility. For example, in activated carbon adsorption processes, the adsorption efficiency can be calculated by dividing the difference between the VOCs concentration at the inlet and the VOCs concentration at the outlet by the VOCs concentration at the inlet. The level of adsorption efficiency directly reflects the adsorption capacity of activated carbon. In catalytic combustion processes, the removal efficiency can be characterized by the rate of change of VOCs concentration at the inlet and outlet.
[0101] S204. Combine basic features and derived features into a multidimensional feature vector.
[0102] After constructing the basic and derived features, these two types of features need to be combined to form a multi-dimensional feature vector that can comprehensively characterize the operational status of the governance facilities. This multi-dimensional feature vector is a data structure containing multiple feature dimensions. Each dimension corresponds to a basic or derived feature. The dimensions are independent of each other but complementary to each other, together constituting a multi-dimensional representation of the operational status of the governance facilities.
[0103] As an example, for activated carbon adsorption processes, a multidimensional feature vector can include the following dimensions: standardized inlet and outlet pressure difference, standardized temperature, standardized relative humidity, standardized flue gas velocity, pressure difference change rate, temperature change rate, humidity change rate, flow rate change rate, correlation between pressure difference and flow rate, and adsorption efficiency. These dimensions reflect the operating status of the activated carbon adsorption process from different perspectives, providing a unified input format for subsequent three-level identification processing.
[0104] Through the preprocessing steps S201 to S204 described above, the original time series data is transformed into a multidimensional feature vector containing basic and derived features, realizing the mapping from the original data to the feature space, and providing a data foundation for subsequent threshold determination, anomaly detection and residual analysis.
[0105] In one possible implementation, unsupervised anomaly detection includes an anomaly detection model constructed using the Isolation Forest algorithm. Specifically, a second identification process is performed on the multidimensional feature vector according to the unsupervised anomaly detection pattern to obtain a second result, including:
[0106] S301. Input the multidimensional feature vector into the anomaly detection model. Calculate the path length of each sample vector in the multidimensional feature vector within the decision tree constructed based on randomly selected features and split points in the anomaly detection model.
[0107] The Isolation Forest algorithm is an ensemble-based unsupervised anomaly detection method that "isolates" anomalous samples by randomly partitioning the feature space. Unlike traditional distance- or density-based anomaly detection methods, Isolation Forest does not calculate the distance or density between samples. Instead, it directly utilizes the characteristic of anomalous samples being "few but distinct" to make them easier to isolate through random partitioning. The construction process of the Isolation Forest model includes: randomly selecting multiple subsets of samples from the training dataset, with each subset used to build an isolation tree. The size of the subsets is typically set to a small constant (e.g., 256) to ensure good computational efficiency and noise resistance. For each isolation tree, a feature dimension is randomly selected from the multidimensional feature vector, and a split point is randomly selected within the range of values for that feature dimension. Then, based on this split point, the sample space of the current node is divided into left and right subspaces. Samples smaller than the split point are assigned to the left subtree, and samples larger than or equal to the split point are assigned to the right subtree. This process is recursively repeated in the left and right subspaces until a stopping condition is met (e.g., the tree height reaches a preset limit, only one sample remains in the node, or all samples have the same feature value). Through this random partitioning method, each isolated tree progressively divides the training samples into different leaf nodes. Because outlier samples are "few and different," they are more easily isolated earlier during the random partitioning process, meaning the path length from the root node to the leaf node is shorter; while normal samples require more partitioning to be isolated, resulting in longer path lengths. This characteristic forms the basis for the isolated forest's ability to identify outlier samples.
[0108] After constructing the isolation forest model, the resulting multidimensional feature vectors are input into the model for anomaly detection. The multidimensional feature vectors contain multiple sample vectors, each corresponding to the operational status of the governance facility at a given sampling time. For each sample vector, it is input into each isolated tree in the isolation forest. Starting from the root node, the model traverses downwards based on the segmentation features and values of each node until a leaf node is reached. The number of edges traversed from the root node to the leaf node is recorded, representing the path length of the sample vector in the current isolated tree. Since the construction process of each isolated tree is random, the path length of the same sample vector may differ across different isolated trees. The average path length of the sample vector across all isolated trees is obtained by averaging the calculated path lengths across the entire isolation forest. The average path length reflects the ease with which a sample vector is isolated; a shorter path length indicates a more easily isolated sample vector, suggesting a higher probability of it being an anomaly, while a longer path length indicates a more difficult sample vector to isolate, suggesting a higher probability of it being a normal sample.
[0109] S302. Determine the anomaly score for each sample vector based on the path length.
[0110] After obtaining the average path length of each sample vector, it needs to be converted into a standardized anomaly score for unified comparison and threshold setting. In this embodiment, the anomaly score is calculated by comparing the average path length of the sample vector with the theoretical average path length of all samples, and then normalizing it to obtain an anomaly score between 0 and 1. The closer the anomaly score is to 1, the shorter the average path length of the sample vector, and the higher its probability of being anomaly; the closer the anomaly score is to 0, the longer the average path length of the sample vector, and the higher its probability of being a normal sample. It quantifies the degree of deviation of the sample vector from the normal pattern. When the feature combination of the sample vector is consistent with most historical samples, its path length is longer and the anomaly score is lower; when the feature combination of the sample vector shows an abnormal pattern, it is more likely to be isolated during random partitioning, its path length is shorter, and its anomaly score is higher.
[0111] S303. Determine the second result based on the comparison between the anomaly score and the preset anomaly threshold.
[0112] After obtaining the anomaly scores for each sample vector, the anomaly scores need to be graded according to a preset anomaly threshold to determine the second result corresponding to each sample vector. In this embodiment, the preset anomaly threshold may include a first preset threshold and a second preset threshold, wherein the first preset threshold is less than the second preset threshold. Based on the comparison between the anomaly score and these two thresholds, the second result is determined to be one of the following states:
[0113] If the anomaly score is less than the first preset threshold, it indicates that the feature pattern of the sample vector is highly consistent with historical normal samples, and there are no signs of anomaly. The second result is then determined to be "running well." If the anomaly score is greater than or equal to the first preset threshold but less than the second preset threshold, it indicates that the feature pattern of the sample vector deviates to some extent from historical normal samples, but the degree of deviation has not reached the level of significant anomaly. There may be potential anomaly risk, and the second result is determined to be "low risk." If the anomaly score is greater than or equal to the second preset threshold, it indicates that the feature pattern of the sample vector deviates significantly from historical normal samples, and there are obvious signs of anomaly. The second result is then determined to be "high risk."
[0114] In this embodiment, the first preset threshold is set to 0.3, and the second preset threshold is set to 0.7. For a given sample vector, if the calculated anomaly score is 0.2, it is determined to be operating well; if the anomaly score is 0.5, it is determined to be low risk; and if the anomaly score is 0.8, it is determined to be high risk. It should be noted that the above threshold values are only illustrative examples of this embodiment. In practical applications, they can be dynamically adjusted and optimized based on the historical operating data of the governance facility and the requirements for anomaly detection.
[0115] In one possible implementation, a third identification process is performed on the multidimensional feature vector based on the time series decomposition pattern to obtain a third result, including:
[0116] S401. The key parameters in the multidimensional feature vector are decomposed using the STL decomposition method to obtain the trend term, seasonal term and residual term.
[0117] STL decomposition is a time series decomposition method based on locally weighted regression, which can robustly decompose a time series into three components: a trend term, a seasonal term, and a residual term. It can handle any type of seasonal and trend changes, has good robustness to outliers, and allows for control over the smoothing of the trend and seasonal terms through parameter adjustment.
[0118] In this embodiment, key parameters reflecting the core operating status of the treatment facility are selected from the multidimensional feature vector and decomposed using STL. The selection of key parameters is determined according to the characteristics of the treatment process. For example, for activated carbon adsorption processes, the inlet and outlet pressure difference can be selected as a key parameter because changes in pressure difference can directly reflect the degree of blockage in the activated carbon layer; for catalytic combustion processes, the catalyst bed temperature can be selected as a key parameter because changes in catalyst bed temperature can reflect the decay of catalyst activity; for regenerative thermal combustion processes, the temperature difference of the regenerator can be selected as a key parameter because changes in temperature difference can reflect changes in the efficiency of the regenerator.
[0119] The STL decomposition process involves two nested loops: the inner loop updates the trend and seasonal terms, while the outer loop adjusts robust weights to reduce the impact of outliers. Through iterative calculations, the original time series is ultimately decomposed into the following three components:
[0120] Trend terms reflect the long-term changing trends of parameters, characterizing the gradual changes in equipment operating status. For example, as activated carbon is used for longer periods, its adsorption capacity gradually decreases, leading to a slow increase in the inlet and outlet pressure difference; the catalyst's activity gradually declines during use, resulting in a slow decrease in the catalyst bed temperature. Extracting trend terms can reflect the equipment's deterioration process and aging trends, providing a basis for preventative maintenance.
[0121] Seasonal terms reflect the periodic fluctuations of parameters, characterizing regular changes caused by external factors. For example, influenced by periodic adjustments in production processes, the intake load of treatment facilities may exhibit regular changes with a daily, weekly, or monthly cycle; affected by diurnal temperature variations, activated carbon adsorption efficiency may exhibit fluctuations with a 24-hour cycle. Extracting seasonal terms can separate these normal fluctuations from the original sequence, avoiding misjudgments as equipment malfunctions.
[0122] The residual term reflects the random fluctuations after removing trend and seasonal terms, characterizing instantaneous changes caused by random factors. Under normal circumstances, the residual term should follow a normal distribution with a mean of zero, and its fluctuation range should remain within a certain range. When equipment experiences sudden failures or abnormal events, the residual term will deviate significantly, manifesting as sudden spikes or continuous shifts.
[0123] For example, the catalyst bed temperature of a catalytic combustion facility was collected at a frequency of once per hour for 30 consecutive days (a total of 720 data points). Using the STL decomposition method, this time series was decomposed. The trend term might show a slow decreasing trend in temperature (e.g., gradually decreasing from 380℃ to 365℃), reflecting the gradual decline in catalyst activity; the seasonal term might show temperature fluctuations with a 24-hour cycle (e.g., temperature increases during high production loads during the day and decreases at low production loads at night), reflecting the regularity of the production cycle; the residual term is the random fluctuation after removing trend and seasonal factors, which should normally be randomly distributed around zero.
[0124] S402. Calculate the mean and standard deviation of the residual terms.
[0125] After obtaining the residual term sequence, its statistical characteristics need to be analyzed to determine whether there are any abnormal fluctuations. Under normal circumstances, the residual terms should follow a normal distribution with a mean of zero, and their mean should be close to zero. The standard deviation reflects the fluctuation range of the residual terms. This step calculates the mean and standard deviation of the residual terms as a benchmark for subsequent abnormal fluctuation judgment. The formula for calculating the mean is: the sum of all values in the residual term sequence divided by the sequence length; the formula for calculating the standard deviation is: the square root of the sum of the squares of the differences between each value in the residual term sequence and the mean divided by the sequence length.
[0126] It should be noted that the mean and standard deviation can be calculated using a sliding window approach, which means that only residual data from the most recent period is used for calculation, in order to adapt to the dynamic changes in the equipment's operating status. For example, a 7-day time window can be set, and the calculated results of the mean and standard deviation can be updated daily, allowing the benchmark values to be adjusted along with the normal aging process of the equipment.
[0127] S403. If the difference between the residual term and the mean is greater than the preset multiple of the standard deviation, it is determined that there is abnormal fluctuation.
[0128] After obtaining the mean and standard deviation of the residuals, the residual values at each time point are examined to determine whether they exceed the normal fluctuation range.
[0129] Specifically, the absolute value of the difference between the current residual value and the mean is calculated, and this absolute value is compared with a preset multiple of the standard deviation. If the absolute value is greater than the preset multiple of the standard deviation, it is determined that there is an abnormal fluctuation at this time point; otherwise, it is determined to be a normal fluctuation. The selection of the preset multiple needs to balance the detection sensitivity and the false alarm rate. The smaller the multiple, the higher the detection sensitivity, but the higher the false alarm rate; the larger the multiple, the lower the false alarm rate, but some small-amplitude anomalies may be missed. In this embodiment, the preset multiple can be set to 2, that is, 2 times the standard deviation is used as the judgment threshold. According to the normal distribution theory, under normal circumstances, the probability that the residual value falls within the range of mean ± 2 times the standard deviation is about 95.5%, and the probability of exceeding this range is about 4.5%, which is a low-probability event and can be regarded as an abnormal fluctuation. For example, suppose that the mean of a certain residual term is 0 and the standard deviation is 2. The residual value at the current time point is 5, and the absolute value of its difference from the mean is 5, which is greater than 2 times the standard deviation (4), so it is determined that there is an abnormal fluctuation at this time point. If the residual value is 3, and the absolute value of the difference between it and the mean is 3, which is less than 2 times the standard deviation (4), then it is judged as normal fluctuation.
[0130] It should be noted that the preset multiplier value can be adjusted according to the actual application scenario. For scenarios with high requirements for anomaly detection sensitivity, a smaller multiplier (such as 1.5) can be set; for scenarios with high requirements for false alarm rate, a larger multiplier (such as 2.5 or 3) can be set. In addition, differentiated multipliers can be set according to the characteristics of different parameters. For example, a larger multiplier can be set for parameters with large fluctuations, and a smaller multiplier can be set for parameters with small fluctuations.
[0131] S404. Determine the third result based on the abnormal fluctuation judgment result.
[0132] After determining the abnormal fluctuations at each point in time, a third overall result needs to be determined based on the results. The third result should reflect whether there are any trend-based or gradual anomalies in the treatment facilities, as well as the degree of the anomaly.
[0133] In this embodiment, the method for determining the third result can be flexibly set according to application requirements. As a preferred method, if abnormal fluctuations are determined to exist at multiple consecutive time points (e.g., three consecutive time points), or if the frequency of abnormal fluctuations within a unit of time exceeds a preset threshold, it indicates that the parameter has experienced continuous abnormal fluctuations, which may be a precursor to equipment failure, and the third result is determined to be high-risk. If a single time point is determined to have abnormal fluctuations, but no continuous trend is formed, it indicates that the parameter has experienced occasional instantaneous anomalies, which may be caused by external interference or sensor noise, and the third result is determined to be low-risk. If no abnormal fluctuations are determined at any time point, it indicates that the parameter is operating smoothly without any abnormal signs, and the third result is determined to be operating well.
[0134] For example, if the inlet and outlet pressure difference of an activated carbon adsorption facility, after STL decomposition, shows a residual term exceeding twice the standard deviation for three consecutive hours, it indicates a continuous abnormal fluctuation in the pressure difference, which may foreshadow impending blockage of the activated carbon layer. In this case, the third result is classified as high risk. If the abnormal fluctuation only occurs at a single point in time, while the time points before and after are normal, it may be due to transient interference from the sensor, and the third result is classified as low risk.
[0135] The above processing enables time series decomposition and residual analysis of multidimensional feature vectors, which can capture gradual and trend-based anomalies and achieve early warning of equipment failures.
[0136] In one possible implementation, the operational status of the volatile organic compound (VOC) treatment facility is determined based on the first, second, and third results, including:
[0137] S501. Based on the pre-defined correspondence between the operating status and score of the volatile organic compound treatment facility, the first result, the second result, and the third result are quantitatively assigned values to obtain the first assigned value and the second assigned value.
[0138] Because the three identification results have different output formats (e.g., the first result might output labels such as "slight deviation" or "severe deviation," the second result might output categories such as "operating well," "low risk," or "high risk," and the third result might output judgments such as "abnormal fluctuations exist" or "no abnormal fluctuations"), they need to be uniformly quantified into comparable scores to facilitate subsequent fusion calculations. This step quantifies and assigns scores based on a preset correspondence between operating status and scores. As an example, the following correspondence is set: 1 point for "operating well," 2 points for "low risk," and 3 points for "high risk."
[0139] Specifically, the first result is assigned a value based on the judgment result of step S103: if all parameters are within the normal threshold range, the first result corresponds to "operating well" and is assigned 1 point; if there is a slight deviation, it is assigned 2 points; if there is a severe deviation, it is assigned 3 points. The second result is assigned a value based on the judgment result of step S104: if it is operating well, it is assigned 1 point; if it is low risk, it is assigned 2 points; if it is high risk, it is assigned 3 points. The third result is assigned a value based on the judgment result of step S105: if it is operating well (no abnormal fluctuations), it is assigned 1 point; if it is low risk (occasional abnormal fluctuations), it is assigned 2 points; if it is high risk (persistent abnormal fluctuations), it is assigned 3 points.
[0140] The assignment results obtained through the above assignment process can be labeled as: first assignment result Score1, second assignment result Score2, and third assignment result Score3.
[0141] S502. Based on the preset weight coefficients corresponding to each recognition process, determine the weight values corresponding to the first result, the second result, and the third result, and perform a weighted summation on the first assignment result, the second assignment result, and the third assignment result based on the weight values to obtain the fusion score.
[0142] The three identification results have different levels of importance in determining the final running status, and therefore require different weighting coefficients. As an example, the following weighting coefficients are set: weight w1 = 0.4 for the first identification process (threshold determination), weight w2 = 0.35 for the second identification process (anomaly detection), and weight w3 = 0.25 for the third identification process (residual analysis).
[0143] The formula for calculating the fusion score is: Score = w1 × Score1 + w2 × Score2 + w3 × Score3, with a score range of 1 to 3.
[0144] For example, suppose the recognition results at a certain sampling time are: the first result is assigned 1 point (good performance), the second result is assigned 2 points (low risk), and the third result is assigned 2 points (low risk). Calculated according to the above weights, the fusion score is: Score = 0.4 × 1 + 0.35 × 2 + 0.25 × 2 = 0.4 + 0.7 + 0.5 = 1.6 points.
[0145] S503. Based on the comparison results between the fusion score and the target fusion threshold, determine the operating status of the volatile organic compound treatment facility.
[0146] After obtaining the fusion score, the continuous score values need to be mapped to discrete running state categories according to preset target fusion thresholds. As an example, the following target fusion thresholds are set: First target fusion threshold TH1 = 1.5, second target fusion threshold TH2 = 2.5. Based on the comparison between the fusion score and these two thresholds, the final running state is determined as follows:
[0147] If the fusion score is less than the first target fusion threshold (Score < 1.5), it indicates that all three recognition results tend to be normal, and the final operating status is determined to be good. If the fusion score is greater than or equal to the first target fusion threshold and less than the second target fusion threshold (1.5 ≤ Score < 2.5), it indicates that there are some signs of anomaly, and the final operating status is determined to be low risk. If the fusion score is greater than or equal to the second target fusion threshold (Score ≥ 2.5), it indicates that there are obvious signs of anomaly, and the final operating status is determined to be high risk.
[0148] To illustrate with the previous example: the fusion score is 1.6, which is greater than or equal to 1.5 and less than 2.5, therefore the final operating status is determined to be low risk. Although the first result shows good operation, the second and third results both indicate low risk. The comprehensive assessment suggests that there is a potential anomaly that requires attention.
[0149] Through the above-mentioned fusion processing, a comprehensive judgment of the three identification results is achieved, which can more comprehensively and accurately assess the actual operating status of the governance facilities and avoid the limitations of single-dimensional judgment.
[0150] In one possible implementation, the method further includes:
[0151] Based on the operating status and target warning level, determine abnormal parameter information; based on the abnormal parameter information, determine fault diagnosis push information so that the fault diagnosis push information is sent to the destination.
[0152] After determining the operational status of the treatment facilities, corresponding early warning and response measures need to be taken based on different operational statuses so that operation and maintenance personnel can understand the equipment status in a timely manner and take appropriate actions. This step triggers tiered early warnings based on the determined operational status and pushes relevant abnormal parameter information and fault diagnosis suggestions.
[0153] As a preferred approach, this application embodiment sets differentiated warning levels based on different operating states: when the operating state is low-risk, a yellow warning is triggered, indicating the need for attention; when the operating state is high-risk, a red warning is triggered, indicating the need for immediate action. Warning information is pushed through multiple channels, including monitoring platform pop-ups, SMS notifications, and mobile push notifications, ensuring that relevant personnel can receive it in a timely manner. The warning push information includes the following: operating state (low-risk or high-risk), abnormal parameter information (such as which parameters are abnormal, abnormal scores, residual fluctuations, etc.), and fault diagnosis suggestions (possible causes of the fault and suggested solutions).
[0154] In one possible implementation, the specific process of determining the fault diagnosis push information may include:
[0155] Input abnormal parameter information into a pre-built fault diagnosis model; respond to the fault diagnosis model to process the abnormal parameter information, obtain the fault cause corresponding to the abnormal parameter information and the confidence level corresponding to each fault cause; match the corresponding handling suggestions from the pre-set handling suggestion library according to the fault cause and confidence level; combine the fault cause, confidence level and handling suggestions into fault diagnosis push information.
[0156] The fault diagnosis model is a multi-class classification model built on the random forest algorithm. It uses historical abnormal parameters and their corresponding causes as training samples, and the fault cause category as the output label. By constructing multiple decision trees and integrating the voting results of each tree, a mapping relationship between abnormal parameters and fault causes is formed. When currently monitored abnormal parameter information is input, the model can automatically identify the most likely fault cause and match the corresponding treatment suggestion from a pre-set treatment suggestion library. For example, for activated carbon adsorption processes, when a continuous increase in the inlet and outlet pressure difference is detected and the abnormal score is high, the fault diagnosis model may output the following fault cause and treatment suggestion: the fault cause is "activated carbon layer blockage," and the treatment suggestion is "check whether the activated carbon layer is over-compacted or has pulverized particles; replace or regenerate if necessary."
[0157] The aforementioned fault causes and handling suggestions are combined with operational status and abnormal parameter information to form a complete fault diagnosis push message, which is sent to the target end (such as the mobile phone of maintenance personnel or the monitoring platform) through an early warning channel, providing on-site personnel with clear guidance for fault investigation and handling. Through the above processing, this embodiment of the application not only realizes the monitoring and early warning of the operating status of the treatment facility, but also provides targeted fault diagnosis suggestions to help maintenance personnel quickly locate problems and take effective measures, thereby improving the operation and maintenance efficiency and reliability of the treatment facility.
[0158] The following describes the method for monitoring the operational status of volatile organic compound (VOC) treatment facilities according to the corresponding application scenarios. It should be noted that the example values in this scenario embodiment are only for the purpose of explanation, and the specific values need to be determined in conjunction with the actual application.
[0159] A chemical industrial park has a volatile organic compound (VOC) treatment facility using activated carbon adsorption technology to treat benzene-containing waste gas generated during production. The facility includes an activated carbon adsorption box (e.g., with an inner diameter of D), a fan, and an exhaust stack. An online VOC monitoring system is installed at the facility's outlet. Differential pressure sensors, temperature sensors, humidity sensors, and flow rate sensors are installed on the inlet and outlet pipes of the activated carbon adsorption box, inside the box, and before and after the fan. All sensors and online monitoring equipment transmit the collected data to a monitoring platform at a sampling frequency of once per hour.
[0160] For example, the monitoring platform receives the following raw data:
[0161] Monitoring data from emission outlets includes: volatile organic compound (VOC) concentration C. out = 25 mg / m³; flue gas flow rate Q g = 8000 m³ / h; flue gas temperature T gas = 38℃.
[0162] Equipment operating parameters (activated carbon adsorption process): Inlet and outlet pressure difference ΔP = 1200 Pa; temperature T = 32℃; relative humidity RH = 65%; flue gas velocity v = 0.5 m / s. In addition, a VOCs concentration monitor is installed at the inlet to measure the VOCs concentration C at the inlet. in = 150 mg / m³.
[0163] The monitoring platform performs outlier detection on the temperature data sequence over the past 24 hours. Let the temperature data sequence be X = {x1, x2, ..., x...} 24 The mean μ was calculated to be 31.5℃ and the standard deviation σ was 1.2℃.
[0164] The mathematical expression for the 3σ principle is: if |x i If -μ| > 3σ, then determine x i These are outliers and will be removed.
[0165] For the current temperature value x = 32℃, we calculate |x - μ| = |32 - 31.5| = 0.5℃. Since 0.5 < 3 × 1.2 = 3.6, the current temperature value is within the normal range and is retained. Suppose that at a certain moment the collected temperature value is 36℃, we calculate |36 - 31.5| = 4.5℃. Since 4.5 > 3.6, this value is determined to be an outlier and is removed.
[0166] The data after outlier removal is smoothed, with a sliding window size of 5. Boundary data is processed using symmetrical fill. The formula for the moving average method is as follows (taking the i-th data point as an example):
[0167] .
[0168] For the first point in the sequence (e.g., i=1), x i-2 and x i-1 If it does not exist, use the symmetrical filling method, that is, use x i+2 and x i+1 Fill with symmetrical values.
[0169] Feature construction is performed on the smoothed data to obtain basic features and derived features.
[0170] Basic characteristics: Standardization of all parameters. Taking temperature as an example, the minimum value T of the historical normal operating range. min = 20℃, maximum value T max = 40℃, the standardized formula is: T std =(TT min ) / (T max -T minSimilarly, other parameters are standardized.
[0171] Derived features: Calculate parameter change rate, parameter correlation and process efficiency based on basic features.
[0172] For example, the hourly rate of temperature change (the difference between the current temperature and the temperature one hour ago divided by the time interval): r T =(T t -T t-60 ) / 60; Assuming the temperature T 1 hour ago t-60 =31℃, then r T Substituting this into the above formula yields approximately 0.017℃ / min.
[0173] Process efficiency characteristics (activated carbon adsorption efficiency): C in This refers to the VOCs concentration at the air inlet.
[0174] Then, a three-level recognition process is performed:
[0175] The first identification process involves fixed threshold judgment. For example, for activated carbon adsorption processes, the following normal operating thresholds are preset: temperature threshold: T < 40℃; relative humidity threshold: RH < 80%; flue gas velocity threshold: v < 0.6 m / s for granular activated carbon; pressure difference threshold: ΔP = 500 - 1500 Pa. Then, based on the inner diameter of the activated carbon box, such as D = 2m, the cross-sectional area S is calculated: S = (πD) / (πD) = 2m / ( ... / (πD) = 2m / (πD) / (πD) / (πD) = 2m / (πD) / (πD) / (πD) = 2m / (πD) / (πD) / (πD) / (πD) = 2m / (πD) / (πD) / (πD) / (πD) = 2m / (πD) / (πD) / (πD) / (πD) / (πD) = 2m / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD) / (πD 2 ) / 4=3.14m 2 .
[0176] flue gas velocity v = Q g / S = 8000 / 3.14 ≈ 2548 m / h ≈ 0.71 m / s (unit conversion is required; this is just an example).
[0177] Correction formula for cumulative operating time of activated carbon (based on flue gas velocity v = 2 m / s): t run,actual =t run,record ×2 / v actual .
[0178] Assume the running time t is recorded. run,record = 400h, actual flow rate v actual = 0.5 m / s, then the actual equivalent running time is: t run,actual =1600h, compared with the threshold of 500h, it has exceeded the replacement cycle.
[0179] Compare the current parameters with the threshold:
[0180] Temperature 32℃ < 40℃, normal; relative humidity 65% < 80%, normal; flue gas velocity 0.5 m / s < 0.6 m / s (granular activated carbon), normal; pressure difference 1200 Pa within the range of 500-1500 Pa, normal; equivalent operating time 1600h > 500h, exceeding the threshold; due to the equivalent operating time exceeding the threshold, it is marked as a slight deviation (assuming an excess of less than 10% is considered a slight deviation, the excess here is relatively large, and it may actually be marked as a severe deviation). The first result is determined to be low risk (assigned 2 points).
[0181] Then, unsupervised anomaly detection is performed as the second identification process. An anomaly detection model is constructed using the Isolation Forest algorithm. For sample X, the anomaly score is calculated using the following formula:
[0182] S(X)=2 -E(h(X)) / c(n) Where E(h(X)) is the average path length of sample X across all decision trees; c(n) is the average path length (correction factor) when the number of samples is n; the judgment rule is: S(X) < 0.3 indicates good performance, 0.3 ≤ S(X) < 0.7 indicates low risk, and S(X) ≥ 0.7 indicates high risk. For example, since S(X) = 0.81 ≥ 0.7, the second result is determined to be high risk (assigned a score of 3).
[0183] The third identification process, based on time series decomposition (STL decomposition method), is then performed. The STL decomposition method is used to perform time series decomposition on the key parameter of inlet and outlet pressure difference. The mathematical model for STL decomposition is as follows:
[0184] Y t =T t +S t +R t .
[0185] Among them, T t For the trend term (reflecting long-term trends), S t For seasonal terms (reflecting periodic fluctuations), R t For the residual term (reflecting random fluctuations, which should normally follow a normal distribution), residual analysis is used to capture abnormal fluctuations. If |R t -μ R |>2σ R If the result is abnormal, it is determined that there is an abnormal fluctuation, and the result is corrected based on the second result.
[0186] For example, performing STL decomposition on the differential pressure data from the past 30 days yields the residual term sequence R. t Calculate the mean μ of the residual term. R = 0 Pa, standard deviation σ R = 15 Pa. The residual value R at the current time. t= 42 Pa. The condition for determining whether abnormal fluctuations exist is: |R t -μ R |>2σ R If the calculated value is 42 > 30, then an abnormal fluctuation is determined. The preset judgment rule is that abnormal fluctuations occurring at multiple consecutive time points are considered high-risk, while occasional abnormal fluctuations are considered low-risk. Upon inspection, residuals exceeded limits consecutively within the past 3 hours; therefore, the third result is determined to be high-risk (assigned a score of 3).
[0187] Based on the preset correspondence between operating status and score (good operation = 1 point, low risk = 2 points, high risk = 3 points), the three identification results are assigned values: First result: low risk, assigned 2 points; Second result: high risk, assigned 3 points;
[0188] Third Result: High Risk, assigned a score of 3. Preset weighting coefficients: First identification processing weight w1 = 0.4, Second identification processing weight w2 = 0.35, Third identification processing weight w3 = 0.25. Fusion score calculation formula: score = w1×score1 + w2×score2 + w3×score3, Substituting the above values = 0.4×2 + 0.35×3 + 0.25×3 = 0.8 + 1.05 + 0.75 = 2.6.
[0189] Preset target fusion thresholds: First target fusion threshold TH1 = 1.5, second target fusion threshold TH2 = 2.5.
[0190] Judgment rules: If score < 1.5, it is judged as running well; if 1.5 ≤ score < 2.5, it is judged as low risk; if score ≥ 2.5, it is judged as high risk; the current fusion score = 2.6 ≥ 2.5, so the final running status is determined to be high risk.
[0191] Finally, a red alert is triggered based on the high-risk status, and the push notification includes abnormal parameter information, such as equivalent operating time of 1600 hours, exceeding the replacement cycle (500 hours); anomaly score of 0.81, exceeding the high-risk threshold of 0.7; and differential pressure residual exceeding the limit for 3 consecutive hours. The abnormal parameter information is then input into a fault diagnosis model built based on the random forest algorithm. This model uses historical abnormal parameters and corresponding fault causes as training samples, constructs multiple decision trees, and integrates the voting results of each decision tree to form a mapping relationship between abnormal parameters and fault causes. The model outputs: the fault cause is "activated carbon adsorption capacity decay, reaching the replacement cycle," with a confidence level of 92%. A corresponding disposal suggestion is matched from a preset disposal suggestion library: "It is recommended to immediately arrange for activated carbon replacement, and appropriately reduce the production load before replacement to avoid exceeding emission standards." Furthermore, the operating status (high risk), abnormal parameter information, fault cause, and disposal suggestion are combined into a complete fault diagnosis push notification, which is sent to maintenance personnel via monitoring platform pop-ups and SMS.
[0192] In this embodiment, the system supports iterative model updates (e.g., monthly updates) to continuously optimize the accuracy of analysis based on closed-loop management data of early warning events. After each early warning event is processed, the actual cause of the fault and the processing result are fed back to the system as new training samples to update the isolated forest anomaly detection model, STL decomposition parameters, fault diagnosis model, and various thresholds and weight coefficients. This allows the system to continuously adapt to the operating characteristics of the governance facility, improving the accuracy and targeting of subsequent monitoring.
[0193] This application also provides a monitoring system for the operational status of volatile organic compound (VOC) treatment facilities, see [link / reference]. Figure 2 The system includes:
[0194] Data acquisition module 21 is used to acquire emission monitoring data and equipment operating parameters of volatile organic compound treatment facilities;
[0195] Data preprocessing module 22 is used to preprocess the emission outlet monitoring data and the equipment operating parameters to obtain a multidimensional feature vector;
[0196] The first identification module 23 is used to perform a first identification process on the multidimensional feature vector based on the target process operation threshold to obtain a first result;
[0197] The second identification module 24 is used to perform a second identification process on the multidimensional feature vector according to the unsupervised anomaly detection mode to obtain a second result; the unsupervised anomaly detection mode represents the processing method of identifying abnormal states by calculating sample anomaly scores.
[0198] The third identification module 25 is used to perform third identification processing on the multidimensional feature vector based on the time series decomposition mode to obtain a third result; the time series decomposition mode represents the decomposition of key parameters into trend terms, seasonal terms and residual terms, and the processing method for identifying abnormal fluctuations based on residual term analysis.
[0199] The fusion module 26 is used to determine the operating status of the volatile organic compound treatment facility based on the first result, the second result, and the third result.
[0200] Optionally, the data acquisition module includes:
[0201] The first data acquisition submodule is used to determine the inlet and outlet pressure difference, temperature, relative humidity, and flue gas velocity of the activated carbon adsorption process as equipment operating parameters in response to the activated carbon adsorption process being used in the volatile organic compound treatment facility.
[0202] The second data acquisition submodule is used to determine the collected combustion chamber temperature, catalyst bed temperature, and combustion aid dosage of the catalytic combustion process as the equipment operating parameters in response to the volatile organic compound treatment facility adopting the catalytic combustion process.
[0203] The third data acquisition submodule is used to determine the collected combustion chamber temperature and switching valve operating frequency of the regenerative combustion process as the equipment operating parameters in response to the regenerative combustion process adopted by the volatile organic compound treatment facility.
[0204] The fourth data acquisition submodule is used to determine the collected volatile organic compound concentration, flue gas flow rate, and flue gas temperature at the emission outlet as emission outlet monitoring data.
[0205] Optionally, the data preprocessing module includes:
[0206] The outlier removal submodule is used to remove outliers from the raw data of the emission outlet monitoring data and the equipment operating parameters.
[0207] The smoothing submodule is used to smooth the data after removing outliers;
[0208] The feature construction submodule is used to construct features from the smoothed data to obtain basic features and derived features. The basic features represent the values after standardizing the smoothed emission outlet monitoring data and equipment operating parameters, and the derived features represent the parameter change rate, parameter correlation and process efficiency calculated based on the basic features.
[0209] The combination submodule is used to combine the basic features and the derived features into a multidimensional feature vector.
[0210] Optionally, the removal submodule is configured as follows:
[0211] Based on the time-series characteristics of the emission outlet monitoring data and the equipment operating parameters, a target processing mode is adopted to remove outliers from the raw data. The target processing mode is characterized by calculating the mean and standard deviation of the raw data of the emission outlet monitoring data and the equipment operating parameters, and identifying outliers as values whose absolute value of the difference from the mean is greater than three times the standard deviation and removing them.
[0212] The smoothing submodule is configured as follows:
[0213] To address the fluctuation characteristics of the operating parameters of the volatile organic compound (VOC) treatment facility, a moving average method is used to smooth the data after removing outliers. The size of the moving window in the smoothing process is dynamically set based on the parameter sampling frequency and the process response time.
[0214] Optionally, the unsupervised anomaly detection mode includes an anomaly detection model constructed using the Isolation Forest algorithm, and a second identification module:
[0215] The input submodule is used to input the multidimensional feature vector into the anomaly detection model, and calculate the path length of each sample vector in the multidimensional feature vector in the decision tree through the decision tree constructed based on randomly selected features and split points in the anomaly detection model.
[0216] The first determining submodule is used to determine the anomaly score of each sample vector based on the path length;
[0217] The second determining submodule is used to determine the second result based on the comparison result between the anomaly score and the preset anomaly threshold.
[0218] Optionally, the third identification module is configured as follows:
[0219] The key parameters in the multidimensional feature vector are decomposed using the STL decomposition method to obtain the trend term, seasonal term, and residual term;
[0220] Calculate the mean and standard deviation of the residual terms;
[0221] If the difference between the residual term and the mean is greater than a preset multiple of the standard deviation, it is determined that there is abnormal fluctuation.
[0222] The third result is determined based on the abnormal fluctuation judgment results.
[0223] Optionally, the fusion module is configured as follows:
[0224] Based on the pre-defined correspondence between the operating status and score of the volatile organic compound treatment facility, the first result, the second result, and the third result are quantitatively assigned values to obtain the first assigned value result, the second assigned value result, and the third assigned value result.
[0225] Based on the preset weight coefficients corresponding to each recognition process, the weight values corresponding to the first result, the second result, and the third result are determined, and the first assignment result, the second assignment result, and the third assignment result are weighted and summed based on the weight values to obtain the fusion score;
[0226] The operating status of the volatile organic compound (VOC) treatment facility is determined based on the comparison between the fusion score and the target fusion threshold.
[0227] Optionally, it also includes:
[0228] The parameter determination module is used to determine abnormal parameter information based on the operating status and the target warning level;
[0229] The push information determination module is used to determine fault diagnosis push information based on the abnormal parameter information, so as to send the fault diagnosis push information to the destination.
[0230] Optionally, the push information determination module is configured as follows:
[0231] The abnormal parameter information is input into a pre-built fault diagnosis model. The fault diagnosis model is a multi-classification model built based on the random forest algorithm. The fault diagnosis model uses historical abnormal parameters and corresponding fault causes as training samples and fault cause categories as output labels. By constructing multiple decision trees and integrating the voting results of each decision tree, a mapping relationship between abnormal parameters and fault causes is formed.
[0232] In response to the fault diagnosis model, the abnormal parameter information is processed to obtain the fault cause corresponding to the abnormal parameter information and the confidence level corresponding to each fault cause;
[0233] Based on the cause of the fault and the confidence level, a corresponding handling suggestion is matched from a preset handling suggestion library;
[0234] The cause of the fault, the confidence level, and the handling suggestions are combined into a fault diagnosis push message.
[0235] It should be noted that the specific implementation of each module and sub-module in this embodiment can refer to the corresponding content of the volatile organic compound treatment facility operation status monitoring method mentioned above, and will not be described in detail here.
[0236] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the volatile organic compound (VOC) treatment facility operation status monitoring methods provided in this application.
[0237] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the volatile organic compound (VOC) treatment facility operation status monitoring methods provided in this application.
[0238] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0239] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0240] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0241] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for monitoring the operational status of a volatile organic compound (VOC) treatment facility, characterized in that, include: Obtain emission monitoring data and equipment operating parameters from volatile organic compound (VOC) treatment facilities; The emission outlet monitoring data and the equipment operating parameters are preprocessed to obtain a multidimensional feature vector; The multidimensional feature vector is subjected to a first identification process based on the target process operation threshold to obtain a first result; The multidimensional feature vector is subjected to a second identification process based on an unsupervised anomaly detection mode to obtain a second result; the unsupervised anomaly detection mode represents a processing method that identifies abnormal states by calculating sample anomaly scores. The multidimensional feature vector is subjected to a third identification process based on the time series decomposition model to obtain a third result; the time series decomposition model represents the decomposition of key parameters into trend terms, seasonal terms and residual terms, and the handling method of abnormal fluctuations is identified based on the residual terms analysis. Based on the first result, the second result, and the third result, the operating status of the volatile organic compound treatment facility is determined.
2. The method according to claim 1, characterized in that, The acquisition of emission outlet monitoring data and equipment operating parameters of the volatile organic compound (VOC) treatment facility includes: In response to the volatile organic compound treatment facility employing activated carbon adsorption technology, the collected inlet and outlet pressure difference, temperature, relative humidity, and flue gas velocity of the activated carbon adsorption process are determined as equipment operating parameters. In response to the volatile organic compound treatment facility employing catalytic combustion technology, the collected combustion chamber temperature, catalyst bed temperature, and combustion aid dosage of the catalytic combustion process are determined as the equipment operating parameters; In response to the volatile organic compound treatment facility adopting a regenerative thermal combustion process, the collected combustion chamber temperature and switching valve operating frequency of the regenerative thermal combustion process are determined as the equipment operating parameters; The volatile organic compound concentration, flue gas flow rate, and flue gas temperature collected at the emission outlet are determined as the emission outlet monitoring data.
3. The method according to claim 1, characterized in that, The preprocessing of the emission outlet monitoring data and the equipment operating parameters to obtain a multi-dimensional feature vector includes: Outlier removal processing is performed on the raw data of the emission outlet monitoring data and the equipment operating parameters; Smooth the data after removing outliers; Feature construction is performed on the smoothed data to obtain basic features and derived features. The basic features represent the values of the smoothed emission outlet monitoring data and equipment operating parameters after standardization. The derived features represent the parameter change rate, parameter correlation and process efficiency calculated based on the basic features. The basic features and the derived features are combined into a multidimensional feature vector.
4. The method according to claim 3, characterized in that, The outlier removal process for the raw data of the emission outlet monitoring data and the equipment operating parameters includes: Based on the time-series characteristics of the emission outlet monitoring data and the equipment operating parameters, a target processing mode is adopted to remove outliers from the raw data. The target processing mode is characterized by calculating the mean and standard deviation of the raw data of the emission outlet monitoring data and the equipment operating parameters, and identifying outliers as values whose absolute value of the difference from the mean is greater than three times the standard deviation and removing them. The smoothing process for the data after removing outliers includes: To address the fluctuation characteristics of the operating parameters of the volatile organic compound (VOC) treatment facility, a moving average method is used to smooth the data after removing outliers. The size of the moving window in the smoothing process is dynamically set based on the parameter sampling frequency and the process response time.
5. The method according to claim 1, characterized in that, The unsupervised anomaly detection mode includes an anomaly detection model constructed using the isolated forest algorithm, wherein the second identification processing of the multidimensional feature vector according to the unsupervised anomaly detection mode to obtain a second result includes: The multidimensional feature vector is input into the anomaly detection model. The path length of each sample vector in the multidimensional feature vector is calculated in the decision tree constructed based on randomly selected features and split points in the anomaly detection model. The anomaly score for each sample vector is determined based on the path length. The second result is determined based on the comparison between the anomaly score and the preset anomaly threshold.
6. The method according to claim 1, characterized in that, The third identification process based on the time series decomposition mode to obtain the third result includes: The key parameters in the multidimensional feature vector are decomposed using the STL decomposition method to obtain the trend term, seasonal term, and residual term; Calculate the mean and standard deviation of the residual terms; If the difference between the residual term and the mean is greater than a preset multiple of the standard deviation, it is determined that there is abnormal fluctuation. The third result is determined based on the abnormal fluctuation judgment results.
7. The method according to claim 1, characterized in that, Determining the operating status of the volatile organic compound (VOC) treatment facility based on the first result, the second result, and the third result includes: Based on the pre-defined correspondence between the operating status and score of the volatile organic compound treatment facility, the first result, the second result, and the third result are quantitatively assigned values to obtain the first assigned value result, the second assigned value result, and the third assigned value result. Based on the preset weight coefficients corresponding to each recognition process, the weight values corresponding to the first result, the second result, and the third result are determined, and the first assignment result, the second assignment result, and the third assignment result are weighted and summed based on the weight values to obtain the fusion score; The operating status of the volatile organic compound (VOC) treatment facility is determined based on the comparison between the fusion score and the target fusion threshold.
8. The method according to claim 1, characterized in that, The method further includes: Based on the operating status and target warning level, determine the abnormal parameter information; Based on the abnormal parameter information, fault diagnosis push information is determined so that the fault diagnosis push information is sent to the destination.
9. The method according to claim 8, characterized in that, The step of determining fault diagnosis push information based on the abnormal parameter information includes: The abnormal parameter information is input into a pre-built fault diagnosis model. The fault diagnosis model is a multi-classification model built based on the random forest algorithm. The fault diagnosis model uses historical abnormal parameters and corresponding fault causes as training samples and fault cause categories as output labels. By constructing multiple decision trees and integrating the voting results of each decision tree, a mapping relationship between abnormal parameters and fault causes is formed. In response to the fault diagnosis model, the abnormal parameter information is processed to obtain the fault cause corresponding to the abnormal parameter information and the confidence level corresponding to each fault cause; Based on the cause of the fault and the confidence level, a corresponding handling suggestion is matched from a preset handling suggestion library; The cause of the fault, the confidence level, and the handling suggestions are combined into a fault diagnosis push message.
10. A monitoring system for the operational status of a volatile organic compound (VOC) treatment facility, characterized in that, include: The data acquisition module is used to acquire emission monitoring data and equipment operating parameters of volatile organic compound (VOC) treatment facilities. The data preprocessing module is used to preprocess the emission outlet monitoring data and the equipment operating parameters to obtain a multi-dimensional feature vector; The first identification module is used to perform a first identification process on the multidimensional feature vector based on the target process operation threshold to obtain a first result; The second identification module is used to perform a second identification process on the multidimensional feature vector according to the unsupervised anomaly detection mode to obtain a second result; the unsupervised anomaly detection mode represents the processing method of identifying abnormal states by calculating sample anomaly scores. The third identification module is used to perform third identification processing on the multidimensional feature vector based on the time series decomposition mode to obtain a third result; the time series decomposition mode represents the decomposition of key parameters into trend terms, seasonal terms and residual terms, and the processing method for identifying abnormal fluctuations based on residual term analysis. The fusion module is used to determine the operating status of the volatile organic compound treatment facility based on the first result, the second result, and the third result.