Abnormality analysis method and system based on pollution source data and storage medium
Through deep learning models and multi-dimensional data fusion technology, enterprises' pollutant discharge abnormalities are identified, and the problems of insufficient accuracy and availability in traditional methods are solved, and efficient and accurate abnormal clue generation and off-site supervision are achieved.
Patent Information
- Application Number
- CN202510949297.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the existing technology to accurately identify whether the company has fraudulent pollution discharge behavior. It does not combine the emission laws of the company's industry and the company itself. Traditional methods find that the usability is not high, and the intrinsic connection between multiple monitoring factors is not established, so the accuracy is difficult to improve.
The pollutant discharge characteristics of enterprises and industries are extracted through deep learning models, combined with online monitoring data for matching analysis, data abnormal information is identified, and abnormal clues are generated through multi-dimensional data correlation analysis, including data validity judgment and multi-dimensional data fusion scoring model.
It improves the efficiency and accuracy of abnormal clues of pollutant discharge from pollution sources, improves the accuracy and availability of identification, and realizes the effectiveness of off-site supervision.
Smart Images

Figure CN120449065A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an automatic monitoring system for pollution sources, and in particular to an abnormality analysis method, system and storage medium based on pollution source data. Background Art
[0002] With the continuous development of industrialization and the acceleration of urbanization, environmental pollution has become a widespread concern. To effectively monitor and manage pollution sources, various regions have implemented automated pollution source monitoring systems. Traditional monitoring methods rely on setting thresholds to determine whether indicators are abnormal. However, with the deepening of law enforcement, there are cases of companies engaging in fraud. Monitoring data may meet standards, but during on-site inspections, companies are actually engaging in illegal emissions, making traditional methods difficult to detect.
[0003] Neural network + data mining.
[0004] By combining the existing pollution source enterprise information, matching the pollution source enterprises with the pollution source enterprise characteristic indicators, assigning precise characteristic indicators to the pollution sources, combining the pollution discharge characteristics and online monitoring data for matching analysis, identifying data anomaly information, and then judging the validity of the data anomaly information, eliminating invalid data, and finally conducting multi-dimensional data correlation analysis to generate anomaly clues.
[0005] A similar invention can be found in "A Method for Identifying Abnormal Data of Pollution Sources Based on Deep Learning Algorithms", publication number CN109711547A, which is described as follows: The present invention discloses a method for identifying abnormal data of pollution sources based on a deep learning algorithm, which specifically includes the following steps: S1. First, the system processing module controls the data acquisition module to collect detection data of external pollution sources, and transmits the collected data to the data classification module for classification. The present invention relates to the field of data processing technology. The method for identifying abnormal data of pollution sources based on a deep learning algorithm can realize the use of original data as input for the feature learning algorithm, and the learning process adopts an unsupervised feature learning process, which can greatly enhance the accuracy of identifying abnormal data of pollution sources, shorten the time for identifying abnormal data, and realize the use of deep learning model algorithms to characterize rich information of data and improve classification performance, achieving the purpose of no manual extraction in the entire process of feature learning and data anomalies, and well solving the problem of difficulty in obtaining labeled abnormal data.
[0006] Disadvantages of existing technology: 1. Traditional discoveries are not very usable and difficult to interpret and apply; 2. If the data does not exceed the standard, it is impossible to determine whether the company has engaged in fraudulent activities; 3. Failure to consider the emission patterns of the industry to which the enterprise belongs and the enterprise itself; 4. The intrinsic connection between multiple monitoring factors has not been established; 5. Without considering the characteristics of the pollution source industry and the emission patterns of the enterprises, it is difficult to further improve the accuracy. Summary of the Invention
[0007] In order to solve the problems in the prior art, the present invention provides an anomaly analysis method based on pollution source data, comprising the following steps: Step 1: Extract pollution characteristics of enterprises and industries through deep learning models; Step 2: Combine the pollution discharge characteristics with the online monitoring data for matching analysis to identify data anomalies; Step 3: Determine the validity of data anomaly information and eliminate invalid data; Step 4: Multi-dimensional data association analysis to generate abnormal clues.
[0008] As a further improvement of the present invention, in step 1, the following is further included: Step S1: Collect historical pollution data, including production data, emission data and environmental monitoring data of pollution-discharging units; Step S2: Build a deep neural network model, train it on historical pollution data, and learn the pollution emission characteristic patterns of different enterprises and industries; Step S3: The deep neural network model constructed in step S2 outputs a pollution emission feature vector, including pollutant composition characteristics and emission intensity characteristics.
[0009] As a further improvement of the present invention, the step 2 is specifically as follows: Collect online monitoring data from pollutant discharge units in real time, match the feature vector parameters of the real-time monitoring data with the pollution discharge characteristics output by the deep learning model, and mark data that does not meet the pollution discharge feature vector parameter range as data anomaly information.
[0010] As a further improvement of the present invention, in step 3, the following is further included: Data tag validity judgment step: Verify whether the abnormal data is caused by equipment failure or calibration, and eliminate data abnormal information related to equipment or calibration; Steps for judging the validity of the operating condition mark: Analyze the correlation between the operating status of the production facilities and the emission data, and eliminate abnormal data information in abnormal operating conditions; Steps for judging the validity of operation and maintenance records: Verify the correlation between equipment maintenance records and abnormal time points, and eliminate abnormal data information during equipment maintenance; Steps for judging the effectiveness of video surveillance: Compare the time series consistency between the video device status and abnormal data, and eliminate data anomaly information caused by abnormal video device status.
[0011] As a further improvement of the present invention, the step 4 is specifically as follows: Integrate video alarm data, penalty data, and complaint data, build a multi-dimensional correlation scoring model, perform weighted scoring on abnormal events, generate abnormal clue reports for abnormal events with scores exceeding the threshold, and output abnormal clue reports. The abnormal clue reports contain abnormal description, abnormal type, confidence level, and processing priority information.
[0012] As a further improvement of the present invention, step 1 further includes a sewage feature extraction step, which is specifically as follows: Step 1: Collect pollution source data through environmental monitoring equipment, including pollutant concentration, emission flow, ambient temperature, and humidity; The second step: preprocessing the collected data, including data cleaning, removing outliers, duplicate values and missing values; The third step: using the long short-term memory network model to explore the relationship and abnormal patterns between multiple factors; Step 4: Save the model parameters and feature extraction layer output as the pollution feature vector, including pollutant composition characteristics and emission intensity characteristics.
[0013] As a further improvement of the present invention, the third step specifically includes: Step a1: Arrange the multiple factor data pre-processed in the second step according to time series to form a multidimensional data sample; Step a2: Build an LSTM network model, determine the number of layers, the number of hidden layer neurons, the input dimension, and the output dimension. The input dimension is determined by the number of factors. Step a3: Divide the dataset into training, validation, and test sets, and use the training set to train the LSTM model. During the training process, adjust the model parameters by minimizing the loss function. Step a4: Use the validation set to evaluate the model during training, select the model with the best performance on the validation set, and conduct a comprehensive evaluation of the final model on the test set to determine the accuracy and reliability of the model; Step a5: When new multi-factor time series data is input into the trained model, the model outputs the prediction result of the data status to determine whether there is an anomaly. If there is an anomaly, the coordinated change relationship between multiple factors is further analyzed to explore the direct pollution discharge rules of multiple factors.
[0014] As a further improvement of the present invention, in step 4, the calculation method of the multidimensional data association score includes: Step B1: Calculate the score based on the following formula: , Among them, the weight of each type of data is , the sum of all weights , The standardized value of each abnormal data information is ,in: Video alarm data standard value ,According to the video alarm type T and the number of video alarms N in the corresponding time period of the data anomaly information, a standardized conversion is performed. The specific formula is: , Where k represents the attenuation coefficient, which is set to 1.61 based on experience. Represents the number of alarms corresponding to video alarm type T, Represents the score weight corresponding to the video alarm type T; Law enforcement data standards ,According to the time T between the law enforcement data and the data anomaly information and the number of law enforcement data N, a standardized conversion is performed. The specific formula is: , Where k represents the attenuation coefficient, which is set to 1.61 based on experience. Represents the number of alarms corresponding to the time interval T, Represents the score weight corresponding to the time interval T; Complaint Data Standards ,According to the time T between the complaint data and the data anomaly information and the number of law enforcement N, a standardized conversion is performed. The specific formula is: , Among them, k represents the attenuation coefficient, which is set to 1.61 according to the empirical value. Represents the number of alarms corresponding to the time interval T, Represents the score weight corresponding to the time interval T; Step B2: Based on the calculated comprehensive score S, different score intervals are divided to interpret the correlation degree and abnormal tendency of multidimensional data: When S>= 80, it indicates that the multidimensional data are closely correlated and there is a high risk of abnormality, which requires special attention; When 60<= S<80, it indicates that the multidimensional data are correlated to a certain extent, there is a potential anomaly, and continuous monitoring is required; When S<60, it means that the correlation degree of multidimensional data is low and is in a relatively normal state.
[0015] The present invention also discloses an anomaly analysis system based on pollution source data, comprising: a memory, a processor, and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the anomaly analysis method of the present invention when called by the processor.
[0016] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the abnormality analysis method described in the claims of the present invention when called by a processor.
[0017] The beneficial effects of the present invention are: 1. The present invention improves the efficiency and accuracy of discovering abnormal clues of pollution discharge from pollution sources; 2. The present invention improves the availability and interpretability of abnormal clues of pollution discharge from pollution sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the abnormality analysis method based on pollution source data of the present invention. DETAILED DESCRIPTION
[0019] The present invention combines the pollution emission characteristics of pollution source industries with the emission patterns of enterprises to establish a complete abnormal clue analysis process, which greatly improves the recognition rate and accuracy of enterprise pollution emission anomalies and effectively realizes off-site supervision.
[0020] like Figure 1 As shown, the present invention discloses an abnormality analysis method based on pollution source data, comprising the following steps: Step 1: Extract pollution characteristics of enterprises and industries through deep learning models; Step 2: Combine the pollution discharge characteristics with the online monitoring data for matching analysis to identify data anomalies; Step 3: Determine the validity of data anomaly information and eliminate invalid data; Step 4: Multi-dimensional data association analysis to generate abnormal clues.
[0021] In step 1, it also includes: Step S1: Collect historical pollution data, including production data, emission data and environmental monitoring data of pollution-discharging units; Step S2: Build a deep neural network model, train it on historical pollution data, and learn the pollution emission characteristic patterns of different enterprises and industries; Step S3: The deep neural network model constructed in step S2 outputs the pollution feature vector, including pollutant composition characteristics and emission intensity characteristics.
[0022] Step 2 is as follows: Collect online monitoring data from pollutant discharge units in real time, match the feature vector parameters of the real-time monitoring data with the pollution discharge characteristics output by the deep learning model, and mark data that does not meet the pollution discharge feature vector parameter range as data anomaly information.
[0023] Step 3 also includes: Data tag validity judgment step: Verify whether the abnormal data is caused by equipment failure or calibration, and eliminate data abnormal information related to equipment or calibration; Steps for judging the validity of the operating condition mark: Analyze the correlation between the operating status of the production facilities and the emission data, and eliminate abnormal data information in abnormal operating conditions; Steps for judging the validity of operation and maintenance records: Verify the correlation between equipment maintenance records and abnormal time points, and eliminate abnormal data information during equipment maintenance; Steps for judging the effectiveness of video surveillance: Compare the time series consistency between the video device status and abnormal data, and eliminate data anomaly information caused by abnormal video device status.
[0024] Step 4 is as follows: Integrate video alarm data, penalty data, and complaint data, build a multi-dimensional correlation scoring model, perform weighted scoring on abnormal events, generate abnormal clue reports for abnormal events with scores exceeding the threshold, and output abnormal clue reports. The abnormal clue reports contain abnormal description, abnormal type, confidence level, and processing priority information.
[0025] In step 1, a sewage feature extraction step is also included, which is as follows: Step 1: Collect extensive data on pollution sources through various environmental monitoring devices, such as sensors and monitoring stations. This data includes but is not limited to pollutant concentrations, emission flow rates, ambient temperature, humidity, and other dimensions. The second step: preprocessing the collected data, including data cleaning, removing outliers, duplicate values and missing values; Step 3: Use models such as the Long Short-Term Memory (LSTM) network to explore the complex relationships and abnormal patterns between multiple factors. The specific process is as follows: Step a1: Arrange the multiple factor data after preprocessing in the second step in time series to form a multidimensional data sample; for example, the data of multiple factors such as pollutant concentration, emission flow, ambient temperature, etc. at the same time point are taken as a sample.
[0026] Step a2: Build an LSTM network model, determine the number of layers, the number of hidden layer neurons, the input dimension, and the output dimension. The input dimension is determined by the number of factors. Step a3: Divide the dataset into training, validation, and test sets, and use the training set to train the LSTM model. During the training process, the model parameters are adjusted by minimizing the loss function (cross entropy loss function). Step a4: Use the validation set to evaluate the model during training, select the model with the best performance on the validation set, and conduct a comprehensive evaluation of the final model on the test set to determine the accuracy and reliability of the model; Step a5: When new multi-factor time series data is input into the trained model, the model outputs the prediction results of the data status to determine whether there are anomalies. If anomalies exist, the model further analyzes the coordinated change relationship between multiple factors, such as which factor changes have a greater impact on the occurrence of anomalies, and the causal relationship between these factors, so as to explore the direct pollution discharge patterns of multiple factors.
[0027] Step 4: Save the model parameters and feature extraction layer output as the pollution feature vector, including pollutant composition characteristics and emission intensity characteristics.
[0028] In step 4, the calculation method of the multidimensional data association score includes: Step B1: Calculate the score based on the following formula: , Among them, the weight of each type of data is , for example, the weight of video alarm data is , the law enforcement data weight is , the complaint data weight is , the sum of all weights , the standardized value of each abnormal data information is ,in: Video alarm data standard value ,According to the video alarm type T (low, medium, high) and the number of video alarms N in the corresponding time period of the data anomaly information, a standardized conversion is performed. The specific formula is: , Where k represents the attenuation coefficient, which is set to 1.61 based on experience. Represents the number of alarms corresponding to the video alarm type T (low, medium, high), Represents the score weight (0.1, 0.2, 0.7) corresponding to the video alarm type T (low, medium, high); Law enforcement data standards , based on the time T (far, medium, near) between the law enforcement data and the data anomaly information and the number of law enforcement data N, the standardized conversion is performed. The specific formula is: , Complaint Data Standards , a standardized conversion is performed based on the time T (far, medium, near) between the complaint data and the data anomaly information and the number of law enforcement N. The specific formula is: , Among them, k represents the attenuation coefficient, which is set to 1.61 according to the empirical value. Represents the number of alarms corresponding to the time interval T (far, medium, near), Represents the score weight corresponding to the time interval T (far, medium, near) (0.1, 0.2, 0.7); Step B2: Based on the calculated comprehensive score S, different score intervals are divided to interpret the correlation degree and abnormal tendency of multidimensional data: When S>= 80, it indicates that the multidimensional data are closely correlated and there is a high risk of abnormality, which requires special attention; When 60<= S<80, it indicates that the multidimensional data are correlated to a certain extent, there is a potential anomaly, and continuous monitoring is required; When S<60, it means that the correlation degree of multidimensional data is low and is in a relatively normal state.
[0029] The present invention also discloses an anomaly analysis system based on pollution source data, comprising: a memory, a processor, and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the anomaly analysis method of the present invention when called by the processor.
[0030] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the abnormality analysis method of the present invention when called by a processor.
[0031] Advantages of the present invention: 1. Combine the industry to which the enterprise belongs and the emission patterns of the enterprise itself to improve the recognition accuracy; 2. Establish the internal connection between multiple pollution source monitoring factors to expand the scope and recognition rate of abnormal clues of pollution source discharge; 3. Multi-dimensional data fusion: Perform correlation analysis on video surveillance data, online monitoring data, etc. to improve the accuracy of anomaly detection.
[0032] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for abnormal analysis based on pollution source data, characterized in that: The following steps are involved: Step 1: Extract pollution characteristics of enterprises and industries through deep learning models; Step 2: Combine the pollution discharge characteristics with the online monitoring data for matching analysis to identify data anomalies; Step 3: Determine the validity of data anomaly information and eliminate invalid data; Step 4: Multi-dimensional data association analysis to generate abnormal clues.
2. The abnormality analysis method according to claim 1, characterized in that: In step 1, the method further includes: Step S1: Collect historical pollution data, including production data, emission data and environmental monitoring data of pollution-discharging units; Step S2: Build a deep neural network model, train it on historical pollution data, and learn the pollution emission characteristic patterns of different enterprises and industries; Step S3: The deep neural network model constructed in step S2 outputs a pollution emission feature vector, including pollutant composition characteristics and emission intensity characteristics.
3. The abnormality analysis method according to claim 1, characterized in that: The step 2 is specifically as follows: Collect online monitoring data from pollutant discharge units in real time, match the feature vector parameters of the real-time monitoring data with the pollution discharge characteristics output by the deep learning model, and mark data that does not meet the pollution discharge feature vector parameter range as data anomaly information.
4. The abnormality analysis method according to claim 1, characterized in that: In step 3, the method further includes: Data tag validity judgment step: Verify whether the abnormal data is caused by equipment failure or calibration, and eliminate data abnormal information related to equipment or calibration; Steps for judging the validity of the operating condition mark: Analyze the correlation between the operating status of the production facilities and the emission data, and eliminate abnormal data information in abnormal operating conditions; Steps for judging the validity of operation and maintenance records: Verify the correlation between equipment maintenance records and abnormal time points, and eliminate abnormal data information during equipment maintenance; Steps for judging the effectiveness of video surveillance: Compare the time series consistency between the video device status and abnormal data, and eliminate data anomaly information caused by abnormal video device status.
5. The abnormality analysis method according to claim 1, characterized in that: The step 4 is specifically as follows: Integrate video alarm data, penalty data, and complaint data, build a multi-dimensional correlation scoring model, perform weighted scoring on abnormal events, generate abnormal clue reports for abnormal events with scores exceeding the threshold, and output abnormal clue reports. The abnormal clue reports contain abnormal description, abnormal type, confidence level, and processing priority information.
6. The abnormality analysis method according to claim 2, characterized in that: In step 1, a sewage discharge feature extraction step is also included, which is specifically as follows: Step 1: Collect pollution source data through environmental monitoring equipment, including pollutant concentration, emission flow, ambient temperature, and humidity; The second step: preprocessing the collected data, including data cleaning, removing outliers, duplicate values and missing values; The third step: using the long short-term memory network model to explore the relationship and abnormal patterns between multiple factors; Step 4: Save the model parameters and feature extraction layer output as the pollution feature vector, including pollutant composition characteristics and emission intensity characteristics.
7. The abnormality analysis method according to claim 6, characterized in that: The third step specifically includes: Step a1: Arrange the multiple factor data pre-processed in the second step according to time series to form a multidimensional data sample; Step a2: Build an LSTM network model, determine the number of layers, the number of hidden layer neurons, the input dimension, and the output dimension. The input dimension is determined by the number of factors. Step a3: Divide the dataset into training, validation, and test sets, and use the training set to train the LSTM model. During the training process, adjust the model parameters by minimizing the loss function. Step a4: Use the validation set to evaluate the model during training, select the model with the best performance on the validation set, and conduct a comprehensive evaluation of the final model on the test set to determine the accuracy and reliability of the model; Step a5: When new multi-factor time series data is input into the trained model, the model outputs the prediction result of the data status to determine whether there is an anomaly. If there is an anomaly, the coordinated change relationship between multiple factors is further analyzed to explore the direct pollution discharge rules of multiple factors.
8. The abnormality analysis method according to claim 5, characterized in that: In step 4, the calculation method of the multidimensional data association score includes: Step B1: Calculate the score based on the following formula: , Among them, the weight of each type of data is , the sum of all weights , The standardized value of each abnormal data information is ,in: Video alarm data standard value ,According to the video alarm type T and the number of video alarms N in the corresponding time period of the data anomaly information, a standardized conversion is performed. The specific formula is: , Where k represents the attenuation coefficient, which is set to 1.61 based on experience. Represents the number of alarms corresponding to video alarm type T, Represents the score weight corresponding to the video alarm type T; Law enforcement data standards ,According to the time T between the law enforcement data and the data anomaly information and the number of law enforcement data N, a standardized conversion is performed. The specific formula is: , Where k represents the attenuation coefficient, which is set to 1.61 based on experience. Represents the number of alarms corresponding to the time interval T, Represents the score weight corresponding to the time interval T; Complaint Data Standards ,According to the time T between the complaint data and the data anomaly information and the number of law enforcement N, a standardized conversion is performed. The specific formula is: , Among them, k represents the attenuation coefficient, which is set to 1.61 according to the empirical value. Represents the number of alarms corresponding to the time interval T, Represents the score weight corresponding to the time interval T; Step B2: Based on the calculated comprehensive score S, different score intervals are divided to interpret the correlation degree and abnormal tendency of multidimensional data: When S >= 80, it indicates that the multidimensional data are closely correlated and there is a high risk of abnormality, which requires special attention; When 60<= S <80, it indicates that the multidimensional data are correlated to a certain extent, there is a potential abnormality, and continuous monitoring is required; When S < 60, it means that the multidimensional data has a low correlation and is in a relatively normal state.
9. An abnormality analysis system based on pollution source data, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the abnormality analysis method according to any one of claims 1 to 8 when called by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the abnormality analysis method according to any one of claims 1 to 8 when called by a processor.
Citation Information
Patent Citations
A pollution source abnormal data identification method based on a deep learning algorithm
CN109711547A
Method of assisting in judging pollution source monitoring data validity by utilizing neural network
CN104063609A
Big data identification method for industry enterprise data abnormal behaviors
CN110990393A
Environment data processing method and device
CN111310803A
Pollution discharge enterprise abnormity judgment method based on electric power big data
CN113962518A