Industrial pollution source emission data auditing and statistical analysis system
By providing a pollution source emission data audit and statistical analysis system that includes multiple modules, the lack of automation and real-time nature in the existing technology is solved, efficient audit and real-time analysis of pollution source emission data is achieved, and scientific environmental governance and policy formulation basis is provided.
Patent Information
- Application Number
- CN202510439858.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The lack of automated intelligent means in the existing technology, which is difficult to fully reflect the complexity of pollution source emissions, and the analysis cycle is long, making it difficult to meet the needs of real-time supervision.
It provides an industry pollution source emission data review and statistical analysis system, including survey unit screening module, data review module, pollution emission index calculation module, regional or industry comprehensive index determination module and analysis verification module. Data screening and analysis are carried out through multi-index data sets, K-means clustering and random forest algorithms, and pollution emission index is calculated dynamically, and real-time updates are supported.
It realizes automated audit and real-time analysis of pollution source emission data, improves data quality and analysis efficiency, meets real-time regulatory needs, and provides scientific basis for environmental governance and policy formulation.
Smart Images

Figure CN119963059A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of environmental monitoring and data analysis technology, and specifically relates to an industry pollution source emission data review and statistical analysis system. Background Technology
[0002] Traditional pollution emission statistical accounting methods are usually based on fixed formulas and weights, which are difficult to adapt to industry development and policy changes. Taking a single pollutant or indicator as the accounting object, it is difficult to provide a comprehensive basis for environmental management. The cycle from data collection to analysis is long, which is difficult to meet the needs of real-time supervision.
[0003] Therefore, how to solve the above key issues in the audit of industry pollution source emission data is a topic worthy of study in this field. SUMMARY OF THE INVENTION
[0004] In view of the above analysis, the embodiment of the present invention aims to provide an industry pollution source emission data review and statistical analysis system to solve the problems in the prior art of lack of automated intelligent means, difficulty in fully reflecting the complexity of pollution source emissions, and long analysis cycle.
[0005] The first aspect of the present application provides an industry pollution source emission data review and statistical analysis system, including: The survey unit screening module is configured to receive enterprise scale, production process, emission characteristics and policy change data, and use the preset sample enterprise screening rules to screen the enterprises involved in the survey industry to obtain sample enterprises; The data audit module is configured to dynamically screen the pollution emission data of sample enterprises for abnormal values and eliminate abnormal data by establishing a multi-indicator data set, combining K-means clustering and random forest algorithms; The pollution emission index calculation module is configured to obtain the pollutant indicators and preset weights included in the index calculation of the industry to be investigated, receive the pollution emission data of the sample enterprises that have been screened for abnormal values, calculate the sub-index of each pollutant based on the ratio of the current emission of each pollutant indicator to the base period emission; perform weighted calculation on each pollutant sub-index according to the preset weights to calculate the pollution emission index of the sample enterprise; and dynamically synchronize the calculation results to the regional or industry index determination module; The regional or industry comprehensive index determination module is configured to summarize the pollution emission index of each sample enterprise in different regional or industry dimensions, and calculate the pollution emission index of the selected region or selected industry using a dynamic weighted average method; The analysis and verification module is configured to verify the accuracy of each pollution emission index through correlation analysis and trend analysis, and generate an analysis report based on the change trend of the pollution emission index.
[0006] Optionally, also include: The tuning module is configured to dynamically set the tuning coefficient according to policy objectives, regional needs or industry characteristics, automatically apply the tuning coefficient to adjust the relevant pollution emission index, and limit the calculated pollution emission index to a specified value range.
[0007] Optionally, the data audit module is configured to dynamically screen the pollution emission data of the sample enterprises by establishing a multi-indicator data set, combining K-means clustering and random forest algorithms, and eliminating abnormal data including: The data audit module performs data standardization on the pollution emission data and constructs a multi-index feature matrix of the enterprise as input data: data matrix = {X 1 ,X 2, ...,X n};X i is the pollution emission data corresponding to the i-th pollutant of the enterprise; Use K-means algorithm to cluster multi-index data sets, compare the differences between different enterprises in multiple indicators, and identify enterprises whose data points are more than the preset distance threshold from their clusters as the first abnormal enterprises; Train the random forest model, use the pollution emission data of the enterprise as input features, and train the model to determine whether the enterprise meets the emission pattern of normal enterprises in the industry; if the emission pattern of the enterprise does not meet the normal standard, it will be marked as the second abnormal enterprise; Perform intersection processing on the first abnormal enterprise and the second abnormal enterprise to determine the final abnormal enterprise, and feed back the abnormal enterprise to the investigation unit screening module.
[0008] Optionally, the survey unit screening module is configured to receive enterprise scale, production process, emission characteristics and policy change data, and screen the enterprises involved in the industry to be surveyed using preset sample enterprise screening rules, and the sample enterprises include: The survey unit screening module is configured to use the preset sample enterprise screening rules to eliminate enterprises that do not meet the conditions and obtain sample enterprises; The K-means clustering method is used to group enterprises according to enterprise scale, emission characteristics, and production process, and representative enterprises are selected as sample enterprises based on the clustering results.
[0009] Optionally, the analysis and verification module is configured to verify the accuracy of each pollution emission index through correlation analysis and trend analysis, and generate an analysis report based on the change trend of the pollution emission index, including: The analysis and verification module is used to calculate the correlation between the pollution emission index and external macro data, and mark the index with low correlation as abnormal data; Calculate the month-on-month change rate and year-on-year change rate of the pollution emission index, and mark it as abnormal fluctuation data when the month-on-month change rate and year-on-year change rate are greater than the preset change rate threshold.
[0010] Optionally, the analysis and verification module is configured to use the Pearson correlation coefficient to calculate the correlation r between the output data of the main products of the industry and the calculated pollution emission index of the industry: Wherein, X i is the output data of main products of relevant industries in external macro data, Y i is the pollution emission index of the industry, 、 are the means of X and Y respectively; When the correlation r is lower than the preset correlation threshold, it is determined that the corresponding pollution emission index is abnormal.
[0011] Optionally, it also includes: a pollutant index determination module, which is configured to summarize the pollution emission index of each sample enterprise in different pollutant dimensions, and calculate the pollution emission index of the set pollutant using a dynamic weighted average method.
[0012] Optionally, it also includes: a query and calculation module, which is configured to receive query conditions input by the user, generate query statements, locate the target partition from the data storage partition storing the specific dimension, and query the target data set from the target partition; slice the target data set according to industry, region or pollutant type, and use parallel computing technology to independently calculate each slice; summarize the calculation results of each slice to generate the final calculation result.
[0013] Optionally, it also includes: a data export module, which is configured to receive the query conditions input by the user, output the pollution emission index of the selected area or industry or pollutant, and generate a trend visualization diagram of the corresponding pollution emission index.
[0014] Optionally, the pollution emission index calculation module includes: a rule solidification submodule, which is configured to embed a preset statistical caliber and calculation method into an automated calculation tool so as to be called during the calculation process.
[0015] The industry pollution source emission data review and statistical analysis system provided in this application, the survey unit screening module uses screening rules to screen out sample enterprises according to enterprise scale, production process, emission characteristics and policy change data, eliminates small, irrelevant or low-emission enterprises, and focuses on analyzing the main emission sources. Adjust the screening rules in time according to policy changes to ensure the timeliness and accuracy of the analysis object. The data review module detects outliers and eliminates abnormal data on the pollution emission data of sample enterprises through the establishment of a multi-indicator data set, combined with K-means clustering and random forest algorithms, reduces the deviation caused by erroneous data, and improves the quality of the data. The pollution emission index calculation module simplifies complex multi-pollutant emissions into a single value, which is convenient for intuitive evaluation of the emission status of enterprises. Supports real-time updating of calculation results to meet the dynamic accounting needs of regional or industry indexes. The regional or industry comprehensive index determination module aggregates enterprise-level indexes into regional or industry-level indexes to provide overall emission status. Optimize weight distribution according to policy requirements or regional characteristics to improve the decision-making support capabilities of the results. The analysis and verification module compares historical data with external data (such as statistics bureau data) to ensure the rationality of the index, and generates analysis reports based on the index change trend to provide a scientific basis for environmental governance and policy formulation. This application has significantly improved data quality, dynamic adaptability, comprehensive evaluation, automation and efficiency through the collaborative work between modules. Brief Description of the Figures In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0016] Figure 1 This is a structural diagram of a specific implementation of the industry pollution source emission data review and statistical analysis system provided in this application; Figure 2 This is a schematic diagram of the changing trend of the pollution emission index and the added value of the secondary industry. Specific implementation method
[0017] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. It should be noted that the embodiments and features in the embodiments of the present disclosure can be combined, separated, interchanged and / or rearranged with each other in the absence of conflict. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0018] The terms used herein are for the purpose of describing specific embodiments and are not intended to be limiting. As used herein, the singular forms "one", "the" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, when the terms "comprise" and / or "include" and their variations are used in this specification, it is indicated that the stated features, wholes, steps, operations, parts, components and / or their groups exist, but the existence or addition of one or more other features, wholes, steps, operations, parts, components and / or their groups is not excluded. It should also be noted that, as used herein, the terms "substantially", "approximately" and other similar terms are used as approximate terms rather than terms of degree, and as such, they are used to explain the inherent deviations in measurements, calculations and / or provided values that will be recognized by those of ordinary skill in the art.
[0019] The structural block diagram of a specific implementation of the industry pollution source emission data review and statistical analysis system provided in this application is as follows Figure 1 As shown in the figure, the system specifically includes: The survey unit screening module 100 is configured to receive enterprise scale, production process, emission characteristics and policy change data, and use preset sample enterprise screening rules to screen enterprises involved in the industry to be surveyed to obtain sample enterprises.
[0020] Collect data from multiple sources, including production scale, emission characteristics, process characteristics and policy requirements. Define screening conditions for different industries: scale (such as output value, output), emission characteristics (such as major pollutant types and emissions), and process (such as ironmaking and steelmaking processes in the steel industry). Output a list of qualified companies based on the screening rules as the basis for subsequent data processing.
[0021] The data audit module 200 is configured to dynamically screen the pollution emission data of sample enterprises for abnormal values and eliminate abnormal data by establishing a multi-indicator data set and combining K-means clustering and random forest algorithms.
[0022] As a specific implementation method, the pollution emission data of the sample enterprises in this application adopts quarterly statistical data. Quarterly statistics are between monthly and annual statistics, which can better balance the timeliness and stability of data and are suitable for capturing medium-term changes and cyclical trends. Compared with other span data, it can meet certain timeliness requirements and provide sufficient stability and representativeness.
[0023] Establish a data set containing multiple indicators, such as the emission of various pollutants, and can further include energy consumption data and product output data.
[0024] Use K-means clustering detection to perform unsupervised clustering of the pollution emission data of sample enterprises and identify abnormal points that deviate from the cluster center. At the same time, build a random forest model based on pollutant emission data to identify outliers. Combine the detection results of the two algorithms to eliminate data marked as abnormal. Output the audited data set for use in subsequent calculation modules.
[0025] The pollution emission index calculation module 300 is configured to obtain the pollutant indicators and preset weights included in the index calculation of the industry to be investigated, receive the pollution emission data of the sample enterprises that have been screened for abnormal values, calculate each pollutant sub-index based on the ratio of the current emission amount of each pollutant indicator to the base period emission amount; perform weighted calculation on each pollutant sub-index according to the preset weight, calculate the pollution emission index of the sample enterprise; and dynamically synchronize the calculation results to the regional or industry index determination module.
[0026] This application proposes to quantify the changes in pollution emissions by an indexation method with a fixed base period as a reference to form trend-based and intuitive results. As a specific implementation method, the base period can be the first quarter of 2022.
[0027] Key pollutants included in the calculation of pollution emission index can be determined based on characteristic pollutants of different industries. For example, traditional industries such as thermal power, steel and cement focus on sulfur dioxide (SO2), nitrogen oxides (NOx) and particulate matter, papermaking and cotton printing and dyeing focus on chemical oxygen demand (COD) and ammonia nitrogen, and other industries such as petroleum refining and coking focus on volatile organic compounds (VOCs).
[0028] A specific correspondence between pollutant indicators included in the index calculation for different industries is shown in Table 1 below.
[0029] Table 1 Industry Included in the index calculation of pollutant emissions Thermal power SO2, NOx, particulate matter Steel SO2, NOx, particulate matter Petroleum refining VOCs Cement SO2, NOx, particulate matter Papermaking COD, ammonia nitrogen Cotton printing and dyeing finishing COD, ammonia nitrogen Coking VOCs Organic chemical raw materials manufacturing VOCs Nitrogen fertilizer manufacturing VOCs Primary plastics and synthetic resin manufacturing VOCs Chemical API manufacturing VOCs For the objects selected by key industry enterprises, the fixed base ratio is calculated according to the main pollutant emissions of different industries, and the calculation is completed and written into the database table, which can also be directly output as an Excel spreadsheet file.
[0030] The regional or industry comprehensive index determination module 400 is configured to summarize the pollution emission index of each sample enterprise in different regional or industry dimensions, and calculate the pollution emission index of the selected region or selected industry using a dynamic weighted average method.
[0031] Classification by region or industry can group enterprises according to their region or industry. Example: Cement, steel, thermal power are divided by industry. Weighted average of the pollution index of enterprises in each group. Output pollution index by region or industry for use by analysis module.
[0032] The industry index determination module refers to the full-caliber industry index calculation for all pollutants in a certain industry, which can realize the total value of the pollution emission index of all pollutants in a certain industry at the national, provincial, key regional, eastern and central regions, and prefecture-level cities, as well as the analysis of the corresponding industry index trend. The regional index determination module refers to the full-caliber pollution emission index calculation for all industries and all pollutants at the national, provincial, key regional, eastern and central regions, and prefecture-level cities, including the total value of the corresponding pollutant emission index of each industry in a certain region, forming the trend analysis results for more than a dozen quarters from 2022 to 2024.
[0033] The analysis and verification module 500 is configured to verify the accuracy of each pollution emission index through correlation analysis and trend analysis, and generate an analysis report based on the change trend of the pollution emission index.
[0034] By checking the correlation between the pollution emission index and external macro data (such as the added value of the secondary industry), analyzing the year-on-year and month-on-month change rates of the pollution emission index, a report is automatically generated, including charts, trend analysis and policy recommendations.
[0035] The system provided in this application realizes the automation of the entire process from survey unit screening to data review, index calculation, regional aggregation and result verification through modular design. Each module ensures the accuracy and timeliness of the pollution emission index through automated algorithms and dynamic tuning mechanisms, providing scientific support for environmental governance and decision-making.
[0036] Based on the above embodiments, the industry pollution source emission data review and statistical analysis system provided by this application also includes: a tuning module, which is configured to dynamically set the tuning coefficient according to policy objectives, regional needs or industry characteristics, and automatically apply the tuning coefficient to adjust the relevant pollution emission index, and limit the calculated pollution emission index to a specified value range.
[0037] When the index value of the index needs to be limited to a certain value range, it can be adjusted through the tuning coefficient. This application develops a tuning module, and can set the value of the corresponding tuning coefficient according to the tuning requirements, so as to achieve recalculation of all data and improve analysis efficiency.
[0038] Among them, the data audit module is configured to dynamically screen the pollution emission data of sample enterprises by establishing a multi-indicator data set, combining K-means clustering and random forest algorithms, and eliminating abnormal data including: The data audit module performs data standardization on the pollution emission data and constructs a multi-index feature matrix of the enterprise as input data: data matrix = {X 1 ,X 2, ...,X n};X i is the pollution emission data corresponding to the i-th pollutant of the enterprise.
[0039] Use K-means algorithm to cluster multi-index data sets, compare the differences between different companies in multiple indicators, and identify companies whose data points are more than the preset distance threshold from their clusters as the first abnormal companies.
[0040] Train the random forest model, use the pollution emission data of the enterprise as input features, and train the model to determine whether the enterprise meets the emission pattern of normal enterprises in the industry; if the emission pattern of the enterprise does not meet the normal standard, it will be marked as the second abnormal enterprise.
[0041] Perform intersection processing on the first abnormal enterprise and the second abnormal enterprise to determine the final abnormal enterprise, and feed back the abnormal enterprise to the investigation unit screening module.
[0042] Thermal power industry uses thermal power generation, SO 2 Emissions, NOx emissions and particulate matter emissions are used as key analysis indicators, while the crude oil refining industry uses crude oil processing volume and VOCs emissions as core indicators. For the key enterprises included in the analysis, by analyzing the internal logical relationship between these indicators and combining the results of the two abnormal screening methods, the data of enterprises that may have abnormalities can be determined, and further verification or rectification suggestions can be made. Through such settings, key indicators and abnormal data can be automatically identified, and verification results can be output.
[0043] Among them, the survey unit screening module is configured to use the preset sample enterprise screening rules to eliminate enterprises that do not meet the conditions and obtain sample enterprises; it can also group enterprises according to enterprise scale, emission characteristics, and production process through the K-means clustering method, and select representative enterprises as sample enterprises based on the clustering results. This ensures that the selected units can accurately reflect the overall characteristics of the industry.
[0044] Specifically, apply screening conditions according to the characteristics of different industries to ensure that the screened companies are representative and data is valid. Sample company screening rules may include: (1) Screening by enterprise size: For industries with a large number of small and medium-sized enterprises and high quality control difficulties (such as cement, papermaking, and cotton printing and dyeing finishing), select large-scale enterprises.
[0045] Screening logic: Cement industry: retain enterprises with production line capacity greater than 600,000 tons / year or quarterly clinker output greater than 100,000 tons. Chemical API manufacturing industry: exclude small enterprises with output value less than 10 million yuan and VOCs emissions less than 1 ton.
[0046] (2) Screening by operating status: Screen out unstable or newly established enterprises (such as abnormal reported data, incomplete operating indicators).
[0047] Filtering logic: Only enterprises that operate normally in each quarter and have no missing items in the main indicators are retained. Example: In the thermal power industry, enterprises with gas boilers with a rated output of more than 100,000 kilowatts and complete quarterly data are retained.
[0048] During the audit process, it was found that the data of some survey objects had obvious irregularities. After verification, these problems were mainly caused by the following reasons: abnormal operating conditions of individual enterprises, replacement of coefficient method and monitoring method, and instability of newly built enterprises. These erroneous information may cause deviations in the judgment of macro data trends. Therefore, by screening the operating status and pollution emission status of enterprises, enterprises with long-term stable production and continuous pollution emissions are retained as key focus objects.
[0049] (3) Screening by key processes: Focus on the key processes of the industry and eliminate small enterprises that are not involved in the main emission links.
[0050] Screening logic: Iron and steel industry: retain iron and steel making enterprises, and exclude ferroalloy smelting and steel rolling processing enterprises. Petroleum refining industry: exclude enterprises whose crude oil processing volume and VOCs emissions account for 1%-2% in total.
[0051] The focus of key industries is on specific key process links. For example, in the steel industry, more attention is paid to the ironmaking and steelmaking processes, while small enterprises in processes such as ferroalloy smelting and steel rolling are not the focus of supervision. By screening key processes, key regulatory targets can be determined, such as screening enterprises with stable crude oil processing volume in the oil refining industry, screening enterprises involved in ironmaking and steelmaking processes in the steel industry, and screening enterprises engaged in clinker production in the cement industry.
[0052] The implementation of the screening function module can be completed through programming tools (such as Python), and the data can be automatically processed using preset screening rules. According to the characteristics of different industries, the corresponding condition sets are defined, and logical operators are used to filter the enterprise data that meets the conditions. At the same time, in order to maintain the timeliness of the screening results, the screening rules can be dynamically updated according to changes in industry policies and emission standards.
[0053] Specifically, based on the results of statistical analysis, we design screening conditions according to the characteristics of different industries, and use programming tools (such as Python) to automate the screening rules. The specific screening conditions are as follows: (1) Thermal power industry: Screen by scale and retain gas boilers with a rated output of more than 100,000 kilowatts, gas turbines and coal-fired boilers with a rated output of more than 200,000 kilowatts. The screening results must undergo data review to ensure that the company is operating normally in each quarter and that the main indicator data is complete and without missing items.
[0054] (2) Iron and steel industry: Focus on steelmaking and ironmaking process enterprises, and delete steel rolling processing enterprises and ferroalloy smelting enterprises that do not involve ironmaking process according to the output of pig iron and crude steel products.
[0055] (3) Petroleum refining industry: Screen by the combined proportion of crude oil processing volume and VOCs emissions, eliminate enterprises with a combined proportion of 1%-2%, and only retain key enterprises with higher emissions.
[0056] (4) Cement industry: Based on the scale screening, enterprises with production line capacity greater than 600,000 tons / year or quarterly clinker output greater than 100,000 tons will be retained.
[0057] (5) Papermaking industry: Delete enterprises with no wastewater discharge or uncalculated wastewater pollutants such as COD and ammonia nitrogen, or zero discharge, and only retain enterprises with actual wastewater discharge.
[0058] (6) Cotton printing and dyeing finishing industry: Eliminate enterprises with wastewater discharge less than 100 tons, COD generation less than 10 tons, or COD emission less than 0.1 tons.
[0059] (7) Coking industry: Screen by production process, exclude enterprises whose coking process is heat recovery process and semi-coke production, and only keep enterprises with traditional coking process.
[0060] (8) Organic chemical raw material manufacturing industry: Small and micro enterprises are eliminated based on their scale, and enterprises with an output value of less than 20 million yuan, a main product output of less than 1,000 tons, or VOCs emissions of less than 1 ton are deleted.
[0061] (9) Primary plastics and synthetic resin manufacturing industry: Small and micro enterprises with an output value of less than 10 million yuan or VOCs emissions of less than 1 ton will be eliminated.
[0062] (10) Chemical API manufacturing industry: Delete small and micro enterprises with an output value of less than 10 million yuan or VOCs emissions of less than 1 ton.
[0063] Through the above screening rules, the enterprise data is comprehensively screened and analyzed, and the industrial enterprises included in the statistical survey and analysis are finally determined to ensure the representativeness and validity of the data.
[0064] Based on the survey unit screening conditions in the definition phase, select the fixed statistical survey objects that meet the requirements, and ensure that these objects maintain the data integrity of normal operation in each quarter. In each analysis, the data is screened and analyzed based on the statistical survey list that has been formed. Through the setting of the survey unit screening module, the problem of insufficient representativeness of the survey units is effectively solved, ensuring the scientificity and reliability of the statistical analysis results.
[0065] The analysis and verification module is used to calculate the correlation between the pollution emission index and external macro data, and mark the index with low correlation as abnormal data.
[0066] Among them, the Pearson correlation coefficient can be used to calculate the correlation r between the output data of the main products of the industry and the calculated pollution emission index of the industry: Wherein, X i is the output data of main products of relevant industries in external macro data, Y i is the pollution emission index of the industry, 、 are the means of X and Y respectively; When the correlation r is lower than the preset correlation threshold, it is determined that the pollution emission index is abnormal.
[0067] Since the number of key survey subjects in the coking, chemical medicine, plastic manufacturing, nitrogen fertilizer and other industries is small, or the changes in some abnormal values between quarters have a significant impact on the statistical results, the data is processed in the following ways: the coefficient of variation is calculated based on the quarterly summary data to screen the discrete data, and the correlation analysis is conducted with the output data of major products in key industries released by government departments, and the rationality of the data is verified industry by industry. The data of enterprises with obvious abnormalities are verified and corrected or deleted as appropriate. This process is mainly based on cross-validation after statistical data analysis to ensure the accuracy and credibility of the data.
[0068] Calculate the month-on-month change rate and year-on-year change rate of the pollution emission index, and mark it as abnormal fluctuation data when the month-on-month change rate and year-on-year change rate are greater than the preset change rate threshold.
[0069] Focus on the development trends of regions and industries, especially the year-on-year and month-on-month changes in statistical data, as well as data analysis of pollution emission index. Through the index calculation method, the first quarter of 2022 is used as the benchmark value to calculate the ratio of data in different quarters to the benchmark quarter. While calculating the parameters required for the index, the ratio of changes between quarters is output, and enterprises with an increase or decrease in the change range of more than 20 times are verified, and the data is corrected or eliminated based on the verification results, so as to realize the automatic identification and output of key indicators and abnormal data.
[0070] Figure 2 Shows a schematic diagram of the change trend of the pollution emission index and the added value of the secondary industry. The left vertical axis in the figure represents the pollution emission index, which is a dimensionless value ranging from 0 to 4000. The right vertical axis represents the added value of the secondary industry, which is 100 million yuan and ranges from 80,000 to 160,000. Figure 2 It can be seen that the trend of the pollution emission index and the added value of the secondary industry is basically consistent, especially at the peak and trough of the index, the two change in the same direction. This shows that pollution emissions may be positively correlated with the intensity of economic activities in the secondary industry.
[0071] Based on any of the above embodiments, it can further include: a pollutant index determination module, which is configured to summarize the pollution emission index of each sample enterprise in different pollutant dimensions, and calculate the pollution emission index of the set pollutant using a dynamic weighted average method. The index is calculated for a certain pollutant according to the relevant industries included in the index analysis, and can achieve the total value of the emission index of a certain pollutant in the whole industry at the national, provincial, key regional, eastern and central regions, and prefecture-level cities, as well as the analysis of the trend of the corresponding pollutant index.
[0072] The query and calculation module is configured to receive the query conditions input by the user, generate query statements, locate the target partition from the data storage partition storing the specific dimension, query the target data set from the target partition; slice the target data set according to industry, region or pollutant type, and use parallel computing technology to independently calculate each slice; summarize the calculation results of each slice to generate the final calculation result.
[0073] Users enter query conditions (such as industry, region, pollutant type, time range, etc.) through the interface. Automatically build query statements based on the conditions entered by the user. Data is partitioned and stored according to dimensions (such as time, region, industry, pollutant type) to improve query efficiency. According to the query statement, quickly find matching partitions from the storage path. Load the data set from the located target partition into the memory or computing environment.
[0074] Slice the target dataset by industry, region or pollutant type. Perform calculation tasks on each slice data independently, using a parallel computing framework. Collect the calculation results of each slice and merge them into the final result. Present the final calculation results in the form of charts or reports. It can support export to multiple formats (such as Excel, PDF). Among them, the parallel computing framework can be carried out in multi-threaded, multi-process or distributed computing.
[0075] Through partition storage and target partition positioning, the query scope is significantly reduced and the query response speed is improved. Parallel computing technology is adopted, and multi-core CPU or distributed computing resources are used to process multiple data shards at the same time to accelerate the calculation process. In addition, this module supports flexible user input conditions, and can dynamically adjust the query and calculation logic according to real-time needs to enhance the applicability of the system. Through fast query and efficient calculation, users are provided with real-time analysis results to support decisions such as pollution control and policy adjustments.
[0076] It can be seen that in addition to being able to generate the calculation structure of various indexes with one click, this application also has a query and calculation module. Through user-friendly query condition input, efficient query positioning based on partitions, parallel computing technology to accelerate processing, and result aggregation and output, a flexible and efficient analysis platform has been built. For the index changes of a certain industry, a certain pollutant, or a certain region, fast query and calculation can be achieved without the need to recalculate and search all.
[0077] In addition, the system provided by this application may also be provided with a data export module, which is configured to receive the query conditions input by the user, output the pollution emission index of the selected area or industry or pollutant, and generate a trend visualization diagram of the corresponding pollution emission index. This module can export the main pollutant index value, industry index, pollutant index, comprehensive index, year-on-year and month-on-month changes, data review issues and other data of the specific survey object, and provide a trend visualization diagram of the relevant data so that users can understand the trend changes more intuitively.
[0078] In addition, the system provided by this application may also include: a statistical analysis and verification module, which conducts statistical analysis and verification based on the index data results, mainly including data collection and correlation coefficients of key industry government departments, which significantly improves the accuracy and decision-making support capabilities of the system.
[0079] Wherein, the pollution emission index calculation module may specifically include: a rule solidification submodule, which is configured to embed the preset statistical caliber and calculation method into the automated calculation tool so as to be called during the calculation process.
[0080] Statistical caliber is used to clarify the definition, dimension and calculation scope of indicators. For example, emissions are measured in tons, including direct emissions and indirect emissions. The calculation method is the specific method for calculating the pollutant sub-indices and various comprehensive indices mentioned above. The statistical caliber and calculation method are embedded in the automated calculation tool. During the calculation process, the applicable statistical caliber and calculation formula are called. Automated calculations are performed according to the rules and input data, and the calculation results are output.
[0081] By solidifying the unified statistical caliber and calculation method, we ensure that all calculations follow a unified standard, without relying on manual judgment, avoiding data incomparability caused by method adjustments and avoiding calculation errors. At the same time, automated tools can quickly read rules and perform calculations, greatly reducing the complexity of manual operations. All calculation processes are based on solidified rules, which facilitates traceability and verification and enhances the credibility of results. Modular design and automated implementation enable the system to handle large-scale, multi-industry data.
[0082] In summary, the industry pollution source emission data review and statistical analysis system provided by this application, the survey unit screening module uses screening rules to screen out sample enterprises according to enterprise scale, production process, emission characteristics and policy change data, eliminates small, irrelevant or low-emission enterprises, and focuses on analyzing the main emission sources. Adjust the screening rules in time according to policy changes to ensure the timeliness and accuracy of the analysis object. The data review module detects outliers and eliminates abnormal data on the pollution emission data of sample enterprises through the establishment of a multi-indicator data set, combined with K-means clustering and random forest algorithms, reduces the deviation caused by erroneous data, and improves the quality of the data. The pollution emission index calculation module simplifies complex multi-pollutant emissions into a single value, which is convenient for intuitive evaluation of the emission status of enterprises. Supports real-time updating of calculation results to meet the dynamic accounting needs of regional or industry indexes. The regional or industry comprehensive index determination module aggregates enterprise-level indexes into regional or industry-level indexes to provide overall emission status. Optimize weight distribution according to policy requirements or regional characteristics to improve the decision-making support capabilities of the results. The analysis and verification module compares historical data with external data (such as statistics bureau data) to ensure the rationality of the index, and generates analysis reports based on the index change trend to provide a scientific basis for environmental governance and policy formulation. This application has significantly improved data quality, dynamic adaptability, comprehensive evaluation, automation and efficiency through the collaborative work between modules.
[0083] Professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0084] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, software modules executed by a processor, or a combination of the two. The software modules may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.
[0085] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An industry pollution source emission data review and statistical analysis system, characterized in that: include: The survey unit screening module is configured to receive enterprise scale, production process, emission characteristics and policy change data, and screen the enterprises involved in the industry to be surveyed using preset sample enterprise screening rules to obtain sample enterprises; The data audit module is configured to dynamically screen the pollution emission data of sample enterprises for abnormal values and eliminate abnormal data by establishing a multi-indicator data set and combining K-means clustering and random forest algorithms; The pollution emission index calculation module is configured to obtain the pollutant indicators and preset weights included in the index calculation of the industry to be investigated, receive the pollution emission data of the sample enterprises that have been screened for abnormal values, and calculate the sub-index of each pollutant based on the ratio of the current emission amount of each pollutant indicator to the base period emission amount; The pollutant sub-index is weighted according to the preset weights to calculate the pollution emission index of the sample enterprises; And dynamically synchronize the calculation results to the regional or industry index determination module; The regional or industry comprehensive index determination module is configured to summarize the pollution emission index of each sample enterprise in different regional or industry dimensions, and calculate the pollution emission index of the selected region or selected industry by using a dynamic weighted average method; The analysis and verification module is configured to verify the accuracy of each pollution emission index through correlation analysis and trend analysis, and generate an analysis report based on the changing trend of the pollution emission index.
2. The industry pollution source emission data review and statistical analysis system according to claim 1 is characterized in that: Also includes: The tuning module is configured to dynamically set the tuning coefficient according to policy objectives, regional needs or industry characteristics, automatically apply the tuning coefficient to adjust the relevant pollution emission index, and limit the calculated pollution emission index to a specified value range.
3. The industry pollution source emission data review and statistical analysis system according to claim 1 is characterized in that: The data audit module is configured to dynamically screen the pollution emission data of sample enterprises for abnormal values by establishing a multi-indicator data set, combining K-means clustering and random forest algorithms, and eliminating abnormal data including: The data audit module performs data standardization on the pollution emission data and constructs a multi-index feature matrix of the enterprise as input data: data matrix = {X1,X 2, ...,X n };X i is the pollution emission data corresponding to the i-th pollutant of the enterprise; Use K-means algorithm to cluster multi-index data sets, compare the differences between different enterprises in multiple indicators, and identify enterprises whose data points are more than the preset distance threshold from their clusters as the first abnormal enterprises; Train the random forest model, use the pollution emission data of the enterprise as input features, and train the model to determine whether the enterprise meets the emission pattern of normal enterprises in the industry; if the emission pattern of the enterprise does not meet the normal standard, it will be marked as the second abnormal enterprise; The first abnormal enterprise and the second abnormal enterprise are subjected to intersection processing to determine the enterprise with final abnormal performance, and the abnormal enterprise is fed back to the investigation unit screening module.
4. The industry pollution source emission data review and statistical analysis system according to claim 1 is characterized in that: The survey unit screening module is configured to receive enterprise scale, production process, emission characteristics and policy change data, and screen the enterprises involved in the industry to be surveyed using the preset sample enterprise screening rules, and the sample enterprises include: The survey unit screening module is configured to eliminate enterprises that do not meet the conditions by using preset sample enterprise screening rules to obtain sample enterprises; The K-means clustering method is used to group enterprises according to their scale, emission characteristics, and production process, and representative enterprises are selected as sample enterprises based on the clustering results.
5. The industry pollution source emission data review and statistical analysis system according to any one of claims 1 to 4, characterized in that: The analysis and verification module is configured to verify the accuracy of each pollution emission index through correlation analysis and trend analysis, and generate an analysis report according to the change trend of the pollution emission index, including: The analysis and verification module is used to calculate the correlation between the pollution emission index and external macro data, and mark the index with low correlation as abnormal data; The month-on-month change rate and year-on-year change rate of the pollution emission index are calculated, and when the month-on-month change rate and year-on-year change rate are greater than the preset change rate threshold, they are marked as abnormal fluctuation data.
6. The industry pollution source emission data review and statistical analysis system according to claim 5 is characterized in that: The analysis and verification module is configured to use the Pearson correlation coefficient to calculate the correlation r between the output data of the main products of the industry and the calculated pollution emission index of the industry: Among them, X i is the output data of main products of relevant industries in external macro data, Y i is the pollution emission index of the industry, , are the means of X and Y respectively; When the correlation r is lower than a preset correlation threshold, it is determined that the corresponding pollution emission index is abnormal.
7. The industry pollution source emission data review and statistical analysis system according to any one of claims 1 to 4, characterized in that: It also includes: a pollutant index determination module, which is configured to summarize the pollution emission index of each sample enterprise in different pollutant dimensions, and use a dynamic weighted average method to calculate the pollution emission index of the set pollutant.
8. The industry pollution source emission data review and statistical analysis system according to claim 7 is characterized in that: Also includes: A query and calculation module is configured to receive a query condition input by a user, generate a query statement, locate a target partition from a data storage partition storing a specific dimension, and query a target data set from the target partition; The target data set is divided into slices according to industry, region or pollutant type, and each slice is calculated independently using parallel computing technology; Summarize the calculation results of each shard to generate the final calculation result.
9. The industry pollution source emission data review and statistical analysis system according to claim 8 is characterized in that: Also includes: The data export module is configured to receive query conditions input by the user, output the pollution emission index of the selected area or industry or pollutant, and generate a trend visualization diagram of the corresponding pollution emission index.
10. The industry pollution source emission data review and statistical analysis system according to claim 9 is characterized in that: The pollution emission index calculation module includes: a rule solidification submodule, which is configured to embed a preset statistical caliber and calculation method into an automated calculation tool so as to be called during the calculation process.
Citation Information
Patent Citations
Method for monitoring pollution discharge behavior according to power utilization of enterprise based on random forest algorithm
CN114580494A
Fixed point source pollutant and carbon emission list dynamic accounting system and method
CN116976562A
Enterprise pollution discharge management method, device and equipment based on characteristic factors
CN118350535A
Recommendation method and device for troubleshooting and classified supervision of soil pollution risk enterprises
CN118627883A
Device and method for verifying deep learning-based carbon accounting data and carbon emissions
KR102690855B1
Cited By
Emission source statistical data comparative analysis and visualization system based on localization rule
CN120763373A