Emission source statistical data comparison and visualization system based on localization rules
By designing a localized rules-based emission source statistical data comparison, analysis, and visualization system, the problem of time-consuming and labor-intensive manual processing in existing technologies has been solved. This system enables automated verification and multi-dimensional data analysis, adapts to the environmental management needs of different regions, and improves data processing efficiency and decision support capabilities.
Patent Information
- Application Number
- CN202511280169.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-09
AI Technical Summary
The current statistical data processing of emission sources relies on manual methods, which has the problems of large amount of calculation, time-consuming and labor-intensive, and prone to errors. In addition, the data is not consistent between different environmental management systems, and the timeliness is insufficient, making it difficult to provide multi-dimensional information services and decision support.
Design a system for comparative analysis and visualization of emission source statistics based on localized rules, including modules for data access, abnormal data identification, comparative analysis, visualization, and output. Support automated verification and visualization, and improve data processing efficiency and accuracy through custom rules and logical verification.
It enables automated verification and comparative analysis of emission source statistics, significantly improving work efficiency, reducing manual operation time, providing intuitive multi-dimensional data analysis and decision support, and adapting to the environmental management needs of different regions.
Smart Images

Figure CN120763373B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing and visualization technology, and in particular to a system for comparative analysis and visualization of emission source statistical data based on localized rules. Background Technology
[0002] Ecological and environmental statistics are an important component of socio-economic statistics and a fundamental task in environmental management. They can promptly reflect environmental conditions and pollutant emission levels, provide a scientific basis for the formulation of environmental development plans and the strengthening of environmental management, and reflect the scientific nature and effectiveness of environmental policies.
[0003] However, the current emission source statistics overlap with multiple environmental management systems. The investigation methods and standards of different systems are inconsistent, the timeliness is insufficient, and the application of information technology is lagging behind. There is an urgent need to improve the efficiency of ecological and environmental statistics and pollution source supervision, and to provide multi-field and multi-dimensional information services and decision support for environmental management, situation analysis, etc.
[0004] The current methods for identifying and comparing abnormal data in emission source statistics, creating charts and writing reports mainly rely on manual methods. The steps include extracting the basic data that needs to be identified and compared, conducting anomaly identification and comparison, calculation and analysis, creating relevant charts, and writing technical analysis reports. This method has obvious shortcomings: it involves many forms and review requirements, the amount of calculation is large and the time is tight, any change to the report will lead to changes in the data in the charts and reports, requiring multiple rounds of verification, and manual processing is time-consuming, labor-intensive and prone to errors. Summary of the Invention
[0005] The present invention proposes a system for comparative analysis and visualization of emission source statistical data based on localized rules, in order to solve the problems mentioned in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A system for comparing, analyzing, and visualizing emission source statistics based on localized rules, comprising the following modules:
[0008] Data Access Module: This module supports importing data forms from the National Ecological and Environmental Statistics Business System. All tables are imported in Excel format. The module has a built-in configurable target area parameter library, which includes longitude range thresholds, latitude range thresholds, and enterprise name keywords corresponding to the target area. Users can edit and save the parameters in the target area parameter library according to actual application scenarios. After data import, the module calls the currently effective target area parameters and performs preliminary verification of the regional validity of the data by combining conditions such as whether the enterprise name contains keywords from the parameter library, whether the longitude of the enterprise center is within the longitude range defined by the parameter library, and whether the latitude of the enterprise center is within the latitude range defined by the parameter library.
[0009] Anomaly data identification module: Constructs localized anomaly data identification rules, supports adding, deleting and versioning rules, records the update time and modifier information of each rule, and associates each rule in the rule base with specific forms and fields;
[0010] Comparison and Analysis Module: This module enables cross-year data comparison, statistically analyzing relevant annual data from five dimensions: citywide, region, industry, key enterprises, and key indicators.
[0011] Visualization module: This module presents comparative analysis results in the form of scatter plots, heat maps, trend lines, bar charts, etc. It is used to show the correlation between indicator data, the differences in emission intensity distribution in different regions, the annual change trend of indicators, and the numerical differences between different indicators or different years. The module supports filtering by city, region, industry, and key indicators to display visual charts and technical analysis reports.
[0012] Output module: This module can export verification results, various charts and technical analysis reports. The output formats supported are EXCEL, WORD and PDF. The exported images support mainstream image formats such as JPG, PNG and SVG. The module supports editing custom report templates. By editing different chart and report templates, various charts and technical analysis reports that meet individual needs can be generated.
[0013] Furthermore, it also includes a logical verification submodule, which works in conjunction with the anomaly data identification module and the comparison analysis module. The anomaly data identification module calls logical rules from the rule base, and the submodule follows the formula... Calculate, where C logic V is the logical verification index. current V represents the indicator value for that year. last The value is the indicator value from the previous year. The minimum value is represented by I(·), which is an indicator when C is calculated. logic =0 and V current and Vlast If all values are greater than 0, the submodule marks the data as a logical anomaly and pushes the anomaly information to the comparison and analysis module.
[0014] Furthermore, it also includes a rationality verification submodule, which works in conjunction with the data access module and the abnormal data identification module. The data access module extracts relevant data from the imported data table, and the abnormal data identification module calls the rationality rules in the rule base. The data access module is also associated with a configurable industry threshold library, which can be entered and updated by the user based on industry statistical data of the target region, including the highest and lowest thresholds of key indicators for each industry. The submodule uses formulas to calculate key indicators, and when the enterprise's calculation result exceeds the threshold range of the corresponding industry in the currently effective industry threshold library, the data is marked as unreasonable.
[0015] Furthermore, the comparison analysis module uses a formula when calculating the rate of change. Where K is the corrected rate of change, and V current For the indicator data of that year, V last The data is based on the previous year's indicators. W is the minimum value. season This is a seasonal correction factor. When the absolute value of the calculated rate of change |K| > 80%, the module highlights the data.
[0016] Furthermore, the abnormal data identification module supports user-defined rule input, allowing users to enter rules using formulas. Configure the rule logic, where Condition is the validation condition that conforms to the form field specifications, Result is the exception message, and Score is the exception severity score; W rule The rule weight is defined as follows: Normal indicates the normal state; user-defined rules are bound to specified database forms and fields, and are activated in response to the system administrator's approval instruction for the submitted rules; the activated rules are stored in the version management library, and the rule's version number, effective timestamp, and operator identification information are recorded.
[0017] Furthermore, the visualization module uses a formula when generating the regional heat map. Mapping color depth, where H is the region's overall thermal value, V region P represents the index value for this region. pop V represents the regional population density weight. total A represents the total value of the corresponding indicator for the target region. region S represents the area of the region. weightThe sensitivity coefficient, with a thermal value ranging from 0 to 1, uses color depth to represent the differences in emission intensity across different regions. This module also uses scatter plots to show the correlation between indicator data, trend lines to show annual changes in indicators, and bar charts to compare different indicators or different years. It supports filtering of these charts by region, industry, and key indicators.
[0018] Furthermore, the data access module uses a formula when verifying the amount of abnormal indicators generated. ,in This represents the amount of the corrected indicator. The amount generated by the original indicator. This is the production scale coefficient. This is the production process efficiency coefficient. The seasonal correction factor is used to calculate the corrected output of the indicator. The module then checks the ratio of the corrected output to the benchmark value per unit of output. If the ratio is not within the range, the module marks the data as abnormal.
[0019] Furthermore, it also includes an integrity verification submodule, which works in conjunction with the data access module and the output module. The data access module counts the total number N of required fields in each form. required At the same time, count the number N fields that have actually been filled in. filled And distinguish between key fields and ordinary fields, the submodules follow the formula Calculate the data integrity rate, where C comp For the weighted completeness ratio, N key_filled N represents the number of key fields that have been filled. norm_filled N represents the number of filled ordinary fields. key_required N represents the total number of required key fields. norm_required W represents the total number of required general fields. key W is the weight of the key field. norm For ordinary field weights, when the calculated completeness rate C comp When the percentage is less than 100%, the submodule will send the missing field names and their corresponding row numbers to the output module.
[0020] Furthermore, the rectification tracking function of the output module adopts a formula. Calculate the weighted rectification rate, where T is the weighted rectification rate and D is the weighted rectification rate. fixed,i Let D be the number of data entries that have been rectified in the i-th type of problem. total,i Let W be the total number of data points for the i-th type of problem. i The weight is denoted as i, and n is the total number of problem types. The module supports filtering data by form type and rectification status, generating a rectification rate trend chart. The trend chart is updated weekly to provide progress reference for managers.
[0021] Furthermore, the comparison and analysis module uses formulas when comparing industry data. , where I diff To correct the difference in industry indicators, V current,i V represents the indicator data for the i-th industry in that year. last,i For the indicator data of industry i in the previous year, S current,i S represents the output value of industry i in that year. last,i For the output value of industry i in the previous year, A avg,i This represents the average output value of the i-th industry over the past three years.
[0022] Compared with existing technologies, the beneficial effects of this invention are:
[0023] In terms of data processing efficiency, the system automates the verification and comparative analysis of emission source statistics, replacing the tedious process of manually filtering and calculating each item in Excel. It can quickly perform multi-dimensional verification of 25 base tables and 14 comprehensive tables, considering completeness, logic, rationality, mutability, and standardization. This significantly reduces manual operation time, avoids repetitive work, frees staff from mechanical calculations, and allows them to focus on data rectification, analysis, and decision-making, thus significantly improving work efficiency.
[0024] In terms of localization adaptability, the system has a built-in configurable target area database (containing data such as a list of enterprises in the target area and pollutant concentration thresholds). Users can input or update the enterprise list (such as a list of power plants in a province or a list of chemical industrial parks in a city) and pollutant concentration thresholds in the database according to the statistical needs of the actual application area (such as provincial, municipal, or district-level administrative regions). The system verifies the validity of the data area through a combination of logic: 'keyword matching of enterprise name (keywords can be customized) + latitude and longitude range verification (latitude and longitude ranges can be edited)', ensuring that only data that meets the requirements of the target area is processed, thus adapting to the statistical characteristics of emission sources in different regions. At the same time, the system's rule base supports the customization or adjustment of rules based on the local management needs of the target area, which can better meet the actual needs of environmental management and situation analysis in different regions.
[0025] In terms of ease of application of results, the visualization module presents comparison results intuitively through scatter plots, heat maps, trend lines, and bar charts, supporting multi-dimensional filtering and allowing users to quickly grasp data distribution and trends. The results output module can export charts and technical analysis reports in various formats based on custom editing templates, providing clear and reliable data support for environmental management decisions and promoting the development of emission source statistics towards refinement and intelligence. Attached Figure Description
[0026] Figure 1 This is a schematic block diagram of the emission source statistical data comparison, analysis and visualization system proposed in this invention;
[0027] Figure 2 Line chart comparing the efficiency of different verification methods;
[0028] Figure 3 A bar chart comparing rule coverage and anomaly detection rate under different verification dimensions;
[0029] Figure 4 A bar chart comparing the adjusted change rates of emissions across various industries. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0032] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.
[0033] Reference Figures 1 to 4 A system for comparing, analyzing, and visualizing emission source statistics based on localized rules, comprising the following modules:
[0034] Data Access Module: This module supports importing 25 emission source statistical base tables and 14 comparative analysis summary tables. All tables are imported in Excel format. The 25 base tables are: Base Table 101, Base Table 101 - Hazardous Waste, Base Table 102, Base Table 102 - Details, Base Table 104 - Details, Base Table 105 - Details, Base Table 106 - Details, Base Table 106, Base Table 107, Base Table 108, Base Table 109 - Details, Base Table 110, Base Table 111 - Conventional, Base Table 111 - Heavy Metals, Base Table 112, Base Table 401, Base Table 401 - Details, Base Table 402, and Base Table 402 - Wastewater Monitoring. The 14 comprehensive tables are: Comprehensive Table 402, Basic Table 403, Basic Table 403 - Incinerator Details, Basic Table 403 - Wastewater Monitoring, Basic Table 108 - Flare Details, and Basic Table 105; and Comprehensive Table 101, Comprehensive Table 101 - Industry Classification, Comprehensive Table 101 - General Solid Waste and Hazardous Waste, Comprehensive Table 107, Comprehensive Table 201, Comprehensive Table 202, Comprehensive Table 203, Comprehensive Table 301, Comprehensive Table 401, Comprehensive Table 403, Comprehensive Table 404, Comprehensive Table 501, Comprehensive Table 502, and Comprehensive Table 601. The module has a built-in configurable target area parameter library and industry threshold library: users can enter a list of enterprises in the target area (such as a list of power plants in a province, a list of chemical industrial parks in a city), and threshold data such as wastewater treatment plant outlet concentration; configuration is supported through 'template import' (e.g., selecting 'Shanghai parameter template' and entering longitude 120°-121°E, latitude 30°-31°N) or 'custom entry' (manually entering the keyword 'XX district' for the enterprise name of a certain district). After the data is imported, the module calls the currently effective parameters and verifies the regional validity of the data based on the condition 'enterprise name contains parameter library keyword + longitude and latitude within the parameter library range'. Only data that meets the target area requirements is further processed (e.g., when importing Shanghai data, Beijing enterprise data is rejected), and only data that meets the regional requirements is further processed.
[0035] Anomaly Data Identification Module: This module constructs a localized rule base for anomaly data verification. Rules are categorized into five dimensions: completeness, logicality, rationality, mutability, and standardization. The rule dictionary contains three fields: rule code, rule name, and remarks. Rule types cover various categories such as change rate, comparison with the previous year, comparison between two columns, NOT NULL, IF condition judgment, and rules that cannot have both elements present or absent. The module supports adding, deleting, and modifying rules and has version control functionality, recording the update time and modifier information for each rule to adapt to the annual update requirements of national rules. Each rule in the rule base is associated with specific forms and fields, such as logical rules for wastewater-related indicators in Table 101 and rationality rules for sludge production in Table 401.
[0036] Comparison and Analysis Module: This module enables cross-year data comparison, analyzing data from the perspectives of the entire city, regions, industries, key enterprises, and key indicators. Specifically, it compares data across the entire city, districts, enterprises, key indicators, industries, solid waste and hazardous waste categories, power plant online data, and all indicators. The comparison of data from each district is linked to Tables 101, 401, 402, and 403, statistically comparing relevant indicator data for the current year and the previous year for each district, as well as the change rate and percentage of the same indicators over the two years; the comparison of enterprise data is linked to Tables 101, 401, 402, 403, 501, Table 101 (Solid Waste Details), and Table 101 (Hazardous Waste Details), statistically analyzing relevant indicator data for enterprises for the current year and the previous year, as well as the difference, change rate, and sum of some indicators over the two years; the comparison of key indicators is linked to Tables 101, 401, 403, 404, 301, 201, 202, 203, and 501, statistically analyzing the total data for the current year and the previous year for key indicators, as well as the change rate of the same indicators over the two years; the comparison of industry data is linked to Table 102... Table 102 – General Solid Waste and Hazardous Waste – provides statistics on major air emission factors for the current and previous years, categorized by industry, along with the differences and rates of change between the two years. Solid waste and hazardous waste categories are compared against Tables 101 – Solid Waste Details and 101 – Hazardous Waste Details, with the generation volume for the current and previous years calculated by solid waste and hazardous waste name, and the differences and rates of change between the two years. Online data from power plants is compared against Table 102 – Details, with the emission volume of major air emission factors calculated by power plant enterprise. All indicators are compared against Tables 101, 107, 201, 202, 203, 301, 401, 403, 404, 501, 502, and 601, categorized by source items such as industrial, agricultural, domestic, sewage treatment plants, and mobile sources, with some key indicators for the current and previous years, and the rates of change between the two years are compared.
[0037] The visualization module presents comparative analysis results using scatter plots, heatmaps, trend lines, and bar charts. It supports filtering by city, region, industry, key enterprise, and key indicator, displaying various charts and technical analysis reports. Scatter plots show the correlation between indicator data; heatmaps show regional distribution, with different regions displaying different color depths based on indicator values; trend lines show the annual changes in indicators, clearly reflecting their increasing or decreasing trends; and bar charts are used for comparing different indicators or different years. The module supports filtering by region (e.g., central urban area, suburbs), indicator (e.g., ammonia nitrogen emissions, chemical oxygen demand generation), and industry (e.g., chemical, power, metallurgy). The output visualization report includes problem data location markers, highlighting outliers in red. The data labels in the report simultaneously display specific values and deviation rates, facilitating quick identification of problematic data.
[0038] Output Module: This module can export verification results, various charts, and technical analysis reports. Output formats supported include EXCEL, WORD, and PDF. Image exports support mainstream image formats such as JPG, PNG, and SVG. The anomaly identification results table includes the row number of the erroneous data, the corresponding verification rules, problem description, and error message. Charts present the comparative analysis results, including scatter plots, heatmaps, trend lines, and bar charts. The technical analysis report summarizes the key results and main conclusions of the comparative analysis. The module supports custom report template editing, allowing users to generate various charts and technical analysis reports by editing different chart and report templates.
[0039] This invention also includes a logical verification submodule, which works in conjunction with the abnormal data identification module and the comparison analysis module. The abnormal data identification module calls logical rules from the rule base, such as those in Table 101: "If the wastewater ammonia nitrogen production (tons) is greater than 0 in the previous year or the current year, then the current year's value should be different from the previous year's value," and "If the wastewater chemical oxygen demand production (tons) is greater than 0 in the previous year or the current year's value, then the current year's value should be different from the previous year's value," etc. The submodule follows the formula... Perform calculations, where C logic V is the logic check index (ranging from 0 to 1, with a value of 0 indicating an anomaly). current V represents the annual target value (such as ammonia nitrogen production). last The index value is from the previous year. The minimum value is 0.001 (to avoid a denominator of 0), and I(·) is an indicator function (1 if the condition is met, 0 otherwise). When C is calculated... logic =0 and V current and V lastWhen all values are greater than 0, the submodule marks the data as a logical anomaly and pushes the anomaly information to the comparison and analysis module. The comparison and analysis module will incorporate the anomaly information when processing the data. Finally, in the visualization module, the anomaly data is marked in orange to distinguish it from other normal data.
[0040] This invention also includes a rationality verification submodule, which works in conjunction with the data access module and the abnormal data identification module. The data access module extracts relevant data such as the operating cost of waste gas treatment facilities and the total industrial output value from the imported base table. The abnormal data identification module calls rationality rules from the rule base, such as "operating cost of waste gas treatment facilities (yuan) / 10000 < total industrial output value" and "base 401 - sludge production (tons) [dry sludge] / actual wastewater treatment volume (ten thousand tons) within the threshold (0.2,4) tons". The submodule uses a formula to verify the rationality of the operating cost of waste gas treatment facilities and the total industrial output value. The calculation is performed, where R is the ratio of facility costs to output (adjusted for scale), and V... facility V represents the operating cost of waste gas treatment facilities (in yuan). output S represents the total industrial output value (in ten thousand yuan). scale This is the industry scale coefficient (e.g., set based on the average output value of various industries in Shanghai, with a value range of 0.8-1.2). When the calculated R value exceeds the threshold range of the corresponding region in the currently effective industry threshold library (e.g., 1.2 for the chemical industry and 0.8 for the power industry), the submodule marks the data as unreasonable and notes the problem description "abnormal ratio of facility costs to output value" in the verification result table generated by the output module, so that users can conduct targeted data verification.
[0041] In this invention, the comparison analysis module uses a formula to calculate the rate of change. Where K is the corrected rate of change (in %), and V current For the indicator data of the year (such as the emissions of cadmium and its compounds in waste gas in Table 101, and the ammonia nitrogen production in wastewater in Table 101), V last The data is based on the previous year's indicators. For the minimum value (0.001), W season This is a seasonal correction factor (set according to the seasonality of industry production, with a value range of 0.9-1.1). When the absolute value of the calculated rate of change |K| > 80%, the module marks the data as an indicator mutation. In the trend chart generated by the visualization module, the mutated data is marked with a dashed box. At the same time, it is associated with the "mutation audit" rule in the abnormal data identification module to generate an abnormal description of "indicator change rate exceeds 80%, belonging to mutation anomaly", which is included in the comparative analysis results.
[0042] In this invention, the abnormal data identification module supports user-defined rule input, and the user can use formulas... Configure the rule logic, where Condition is the validation condition that conforms to the form field specifications, such as "Base 402 - Wastewater (including leachate) discharge (tons) cannot be empty" and "If the hazardous waste utilization and disposal method is comprehensive utilization (code 1), then the actual utilization (tons) cannot be empty", etc.; Result is the exception prompt information; Score is the exception severity score (1-5 points, the higher the score, the more severe); W rule This is the rule weight (set according to the rule's importance, 0.5-1.5); Normal is the normal status indicator (e.g., "Passed"). User-defined rules need to be associated with specific forms and fields. After submission, they must be approved by the system administrator to take effect and be included in the rule library for management. The rule's version number, effective timestamp, and operator identification information are recorded.
[0043] In this invention, the visualization module uses a formula when generating a regional heat map. Mapping color depth, where H is the region's overall thermal value, V region P represents the indicator value for this region (such as the chemical oxygen demand emissions of a certain region). pop V represents the regional population density weight (10,000 people / square kilometer). total A represents the total value of this indicator for the entire city. region S represents the area of the region (square kilometers). weight The sensitivity coefficient is set according to the regional function: 1.2 for industrial areas and 0.8 for residential areas. The heat value ranges from 0 to 1, and the corresponding color gradually transitions from light blue (0) to dark red (1). The color depth intuitively shows the difference in emission intensity in different areas, helping users quickly grasp the regional emission distribution characteristics. At the same time, this module shows the correlation relationship of indicator data through scatter plots, shows the annual changes of indicators through trend lines, and compares different indicators or different years through bar charts. It supports filtering the above charts by region, industry, and key indicators.
[0044] In this invention, the data access module uses a formula when verifying the amount of abnormal indicators generated. ,in This represents the amount of the corrected indicator. This refers to the original indicator generation amount (such as the generation amount of a certain pollutant recorded in the base table, in tons). This is the production scale coefficient (set based on the ratio of the company's actual production capacity to its designed production capacity, with a value range of 0.8-1.2). This is the production process efficiency coefficient (set according to the industry's technological level, with a value range of 0.9-1.0). The seasonal adjustment factor (set according to the seasonal differences in industry production, such as 1.1 in peak season and 0.9 in off-season) is used to calculate the adjusted output of the indicator. The ratio of the adjusted output to the benchmark output per unit (i.e., ...) is then used. ,in Whether the company's total industrial output value (in ten thousand yuan) falls within the industry threshold range. Internal (e.g., thresholds are set based on local Shanghai industry statistics, such as the chemical oxygen demand generation threshold in the chemical industry). tons / ten thousand yuan (tons / ten thousand yuan). If the data is not within this range, the module will mark it as abnormal.
[0045] This invention also includes an integrity verification submodule, which works in conjunction with the data access module and the output module. The data access module counts the total number N of required fields in each form. required For example, in Table 101, fields such as company code, company name, and contact information are required fields. The table also counts the number of fields (N) that have actually been filled in. filled And distinguish between key fields (such as company code) and ordinary fields (such as remarks). Submodules follow the formula Calculate the data integrity rate, where C comp N represents the weighted completeness rate (in %). key_filled N represents the number of key fields that have been filled. norm_filled N represents the number of filled ordinary fields. key_required N represents the total number of required key fields. norm_required W represents the total number of required general fields. key For the key field weight (1.0), W norm The weight is set to 0.5 for a normal field. When the calculated completeness rate C... comp When the accuracy is less than 100%, the submodule sends the missing field names and row numbers to the output module. The output module lists this information in the generated verification result table, such as "The 'Contact Information Mobile Phone' field in row 15 of table 101 is not filled in" and "The 'Wastewater Discharge' field in row 8 of table 402 is not filled in", to help users supplement and improve the data.
[0046] In this invention, the rectification tracking function of the results output module adopts a formula. Calculate the weighted rectification rate, where T is the weighted rectification rate (in %), and D... fixed,i Let D be the number of data entries that have been rectified in the i-th type of problem. total,i Let W be the total number of data points for the i-th type of problem. iThe weight for the i-th type of problem is set according to severity, with scores of 1-5 corresponding to weights of 1.0-0.2, and n is the total number of problem types. The module supports filtering data by form type (such as base 101 form, base 402 form, etc.) and rectification status (such as "pending rectification", "rectified", etc.), generating a rectification rate trend chart. The trend chart is updated weekly, clearly showing the progress of rectification work, providing managers with an intuitive progress reference, and promoting the timely rectification of problem data.
[0047] In this invention, the comparison and analysis module uses a formula when performing industry data comparison. , where I diff To correct the difference in industry indicators, V current,i V represents the indicator data for the i-th industry in that year (such as sulfur dioxide emissions from the chemical industry). last,i For the indicator data of industry i in the previous year, S current,i S represents the output value (in ten thousand yuan) of industry i in that year. last,i For the output value (in ten thousand yuan) of industry i in the previous year, A avg,i This represents the average output value of the i-th industry over the past three years.
[0048] The following two examples further illustrate the specific implementation of this system:
[0049] Example 1: Verification and Multi-dimensional Analysis of City-level Total Emission Source Statistical Data
[0050] This embodiment is applied to the Shanghai municipal-level ecological and environmental statistics department, covering the full data verification and cross-dimensional analysis of 25 base tables (base tables 101 to 403 and various detailed tables) and 14 comprehensive tables (comprehensive tables 101 to 601) for the entire city. The data access module supports batch import of Excel format forms, and automatically performs regional validity verification during import: the proportion of enterprises containing the keyword "Shanghai" in their names must be ≥60%, and the proportion of enterprises with a central longitude (120°-121°) and a central latitude (30°-31°) must both be ≥90%. If these conditions are not met, a "non-Shanghai region data" prompt is triggered and the data is rejected. The built-in localized lists (such as power plant enterprise lists and chemical zone enterprise lists) are associated with the base table fields. For example, the "enterprise name" in the details of base table 102 must match the power plant list; otherwise, it is marked as "non-power plant enterprise incorrectly filled".
[0051] The abnormal data identification module loads over 600 localized rules, categorized and called by dimension: integrity rule verification is based on required fields such as "wastewater discharge" in table 402, and is performed using formulas. Calculate the weighted completeness rate (key fields such as enterprise code have a weight of 1.0, and ordinary fields such as remarks have a weight of 0.5). If the completeness rate is less than 100%, list the missing fields and row numbers; logical rules, such as the "ammonia nitrogen production" in table 101, must meet the following requirements. (V)current >0∨V last >0), C logic If the value is 0 and all values are greater than 0, it is marked as "abnormal with consistent values for two consecutive years"; in the reasonableness rules, "cost of waste gas treatment facilities" is calculated as follows: Calculation (Chemical Industry S) scale =1.2, if it exceeds the industry threshold, it will be marked as "abnormal cost-to-output ratio".
[0052] The comparison and analysis module performs in-depth calculations across five dimensions: comparing and correlating data from each region with basal table 101 and basal table 401, through... (Rainy Season W) season =1.1) Calculate the corrected rate of change; when |K| > 80%, mark the abrupt change with a dashed box in the trend line; industry data comparison uses... After adjusting for the impact of output value scale, red / green bars are used to distinguish between increases and decreases in indicators. The visualization module generates a city-wide heat map, categorized by... (Industrial Zone S) weight =1.2) Mapping colors, with red highlighting high-emission areas. The output module exports an Excel-formatted verification table (including abnormal row numbers, rule codes, and calculation process) and a PDF analysis report, supporting filtering by dimensions such as "Pudong / Puxi" and "Chemical / Power Industry".
[0053] Table 1: Comparison of Efficiency and Functionality between Manual Processing and System Processing
[0054]
[0055] This table (Table 1) verifies the system's efficiency and accuracy advantages: manual processing is extremely inefficient due to incomplete rule memorization and cumbersome calculations; the system achieves full form validation through automated rule invocation, with 100% rule coverage, and anomaly location is accurate to specific row numbers and calculation logic (e.g., "In Table 101, row 23, the ammonia nitrogen production is 50 tons for two consecutive years, C..."). logic =0”, providing reliable data support for city-level overall management.
[0056] Example 2: District-level Key Form Validation and Enterprise-level Rectification Tracking
[0057] This embodiment targets the ecological and environmental department of a district in Shanghai, focusing on six core forms, including Basic Form 101 and Basic Form 401, emphasizing enterprise-level data verification and closed-loop management of rectification. The data access module is simplified to import only the specified six forms, automatically associating them with a list of district-owned enterprises (such as enterprises in the Lingang New Area), verifying that the longitude of the enterprises is concentrated at 121°±0.5° and the latitude at 30.5°±0.5°, ensuring data relevance.
[0058] The abnormal data identification module has enabled district-level special rules: the verification of abnormal indicator generation in Table 401 has passed. Calculate (production scale coefficient) This is the ratio of a company's actual production capacity to its designed production capacity, ranging from 0.8 to 1.2; seasonal adjustment factor. The efficiency coefficient is 1.1 in peak season and 0.9 in off-season; (Value range: 0.9-1.0) ( (The enterprise's total industrial output value in the same period, in RMB 10,000) is not within the industry threshold range. Internal (e.g., the threshold for chemical oxygen demand generation in the chemical industry) tons / ten thousand yuan (tons / ten thousand yuan), generating detailed anomaly information (e.g., "Corrected indicator generation = Original indicator generation × Production scale coefficient × Seasonal correction coefficient × Production process efficiency coefficient = 500 × 1.0 × 1.1 × 0.95 = 522.5 tons, 522.5 / 1000 = 0.5225 tons / ten thousand yuan > 0.5"); The normative rule verification "Contact information (telephone and mobile phone) cannot be empty" is performed using a custom formula. The flag is incorrect.
[0059] The comparison and analysis module only enables enterprise data comparison, by... The corrected difference is calculated, and changes in emission indicators of key enterprises (such as those in chemical industrial parks) within the jurisdiction are closely monitored. The rectification tracking function is implemented through... Calculate the weighted rectification rate (severe anomaly W) i =1.0), generating weekly trend charts, such as "The rectification rate of a chemical company's sudden change in waste gas emissions has increased from 30% to 90%".
[0060] The visualization module is simplified into an enterprise-level bar chart comparison, with red indicating unrectified anomalies and green indicating rectified items. The output module generates an "Enterprise Rectification List," which includes fields such as "Problem Description - Responsible Unit - Rectification Deadline - Current Status," and supports Excel export and progress updates.
[0061] Table 2: Efficiency Comparison of Traditional Manual Methods and This System in Various Aspects of Environmental Management
[0062]
[0063] Table 2 demonstrates the adaptability of the district-level application: the system optimizes the process of abnormal data identification, chart processing and report writing, the system directly calls relevant data, which greatly shortens the processing time, the system's lightweight design fits the "small but sophisticated" requirement, and at the same time fully covers core functions such as abnormal data identification and custom templates, so as to realize the accuracy and efficiency of grassroots statistical work.
[0064] Reference Figure 2 , Figure 2A line graph comparing the efficiency of different verification methods is used, with Base 101, Base 401, Base 402, Comprehensive 101, and Comprehensive 401 as the horizontal axis and verification time (minutes) as the vertical axis. By comparing the time taken for manual verification and system verification, the line graph shows that the time taken for system verification of a single table is significantly reduced, far lower than that taken for manual verification, which intuitively demonstrates the significant advantage of the system in data verification efficiency.
[0065] Reference Figure 3 , Figure 3 A bar chart comparing rule coverage and anomaly detection rate under different verification dimensions is presented. The horizontal axis represents completeness, logic, rationality, mutation, and standardization, while the vertical axis represents percentage. The bar chart shows that the system significantly improves rule coverage and anomaly detection rate in all dimensions, which is better than the rule coverage and anomaly detection rate of manual verification, proving the completeness and accuracy of the system's verification dimensions.
[0066] Reference Figure 4 , Figure 4 A bar chart comparing the corrected change rates of emissions across various industries is provided. The horizontal axis represents chemical, power, metallurgy, light industry, and other industries, while the vertical axis represents the corrected change rate (%). The bar chart presents the changes in emission indicators across different industries over the years. Positive numbers indicate an increase in emissions, and negative numbers indicate a decrease in emissions. This verifies the system's industry data comparison function and data correction effect, and adapts to the emission statistics needs of different industries.
[0067] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A system for comparative analysis and visualization of emission source statistical data based on localized rules, characterized in that, Includes the following modules: Data Access Module: This module supports importing data forms from the National Ecological and Environmental Statistics Business System. All tables are imported in Excel format. The module has a built-in configurable target area parameter library, which includes longitude range thresholds, latitude range thresholds, and enterprise name keywords corresponding to the target area. Users can edit and save the parameters in the target area parameter library according to actual application scenarios. After data import, the module calls the currently effective target area parameters and performs preliminary verification of the regional validity of the data by combining conditions such as whether the enterprise name contains keywords from the parameter library, whether the longitude of the enterprise center is within the longitude range defined by the parameter library, and whether the latitude of the enterprise center is within the latitude range defined by the parameter library. Anomaly data identification module: Constructs localized anomaly data identification rules, supports adding, deleting and versioning rules, records the update time and modifier information of each rule, and associates each rule in the rule base with specific forms and fields; Comparison and Analysis Module: This module enables cross-year data comparison, statistically analyzing relevant annual data from five dimensions: citywide, region, industry, key enterprises, and key indicators. Visualization module: This module presents comparative analysis results in the form of scatter plots, heat maps, trend lines, and bar charts. It is used to show the correlation between indicator data, the differences in emission intensity distribution in different regions, the annual change trend of indicators, and the numerical differences between different indicators or different years. The module supports filtering by region, industry, and key indicators, and displays charts and technical analysis reports. Output module: This module can export verification results, various charts and technical analysis reports. The output formats supported are EXCEL, WORD and PDF. It also supports editing custom report templates. By editing different charts and report templates, various charts and technical analysis reports that meet individual needs can be generated. The logical verification submodule works in conjunction with the anomaly data identification module and the comparison analysis module. The anomaly data identification module calls logical rules from the rule base, and the submodule follows the formula... Calculate, where C logic V is the logical verification index. current V represents the indicator value for that year. last The value is the indicator value from the previous year. The minimum value is represented by I(·), which is an indicator when C is calculated. logic =0 and V current and V last When all values are greater than 0, the submodule marks the data as a logical anomaly and pushes the anomaly information to the comparison and analysis module; The rationality verification submodule works in conjunction with the data access module and the abnormal data identification module. The data access module extracts relevant data from the imported data table, and the abnormal data identification module calls the rationality rules in the rule base. The data access module is also associated with a configurable industry threshold library, which can be entered and updated by the user based on industry statistical data of the target region, specifying the highest and lowest thresholds for key industry indicators. The submodule uses formulas to calculate key indicators, and when the enterprise's calculation result exceeds the threshold range of the corresponding industry in the currently effective industry threshold library, the data is marked as unreasonable.
2. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, The comparison analysis module uses a formula when calculating the rate of change. Where K is the corrected rate of change, and V current For the indicator data of that year, V last The data is based on the previous year's indicators. For the minimum value, W season This is a seasonal correction factor. When the absolute value of the calculated rate of change |K| > 80%, the module highlights the data.
3. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, The abnormal data identification module supports user-defined rule input, and users can use formulas... Configure the rule logic, where Condition is the validation condition that conforms to the form field specifications, Result is the exception message, and Score is the exception severity score; W rule The rule weight is defined as follows: Normal indicates the normal state; user-defined rules are bound to specified database forms and fields, and are activated in response to the system administrator's approval instruction for the submitted rules; the activated rules are stored in the version management library, and the rule's version number, effective timestamp, and operator identification information are recorded.
4. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, The visualization module uses a formula when generating regional heat maps. Mapping color depth, where H is the region's overall thermal value, V region P represents the index value for this region. pop V represents the regional population density weight. total A represents the total value of the corresponding indicator for the target region. region S represents the area of the region. weight The sensitivity coefficient, with a thermal value ranging from 0 to 1, uses color depth to represent the differences in emission intensity across different regions. This module also uses scatter plots to show the correlation between indicator data, trend lines to show annual changes in indicators, and bar charts to compare different indicators or different years. It supports filtering of these charts by region, industry, and key indicators.
5. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, The data access module uses a formula when verifying the amount of abnormal indicators generated. ,in This represents the amount of the corrected indicator. The amount generated by the original indicators. This is the production scale coefficient. This is the production process efficiency coefficient. The seasonal correction factor is used to calculate the corrected output of the indicator. The module then checks the ratio of the corrected output to the benchmark value per unit of output. If the ratio is not within the range, the module marks the data as abnormal.
6. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, It also includes an integrity verification submodule, which works in conjunction with the data access module and the output module. The data access module counts the total number N of required fields in each form. required At the same time, count the number N fields that have actually been filled in. filled And distinguish between key fields and ordinary fields, the submodules follow the formula Calculate the data integrity rate, where C comp For the weighted completeness ratio, N key_filled N represents the number of key fields that have been filled. norm_filled N represents the number of filled ordinary fields. key_required N represents the total number of required key fields. norm_required W represents the total number of required general fields. key W is the weight of the key field. norm For ordinary field weights, when the calculated completeness rate C comp When the percentage is less than 100%, the submodule will send the missing field names and their corresponding row numbers to the output module.
7. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, The rectification tracking function of the output module uses a formula. Calculate the weighted rectification rate, where T is the weighted rectification rate and D is the weighted rectification rate. fixed,i Let D be the number of data entries that have been rectified in the i-th type of problem. total,i Let W be the total number of data points for the i-th type of problem. i The weight is denoted as i, and n is the total number of problem types. The module supports filtering data by form type and rectification status, generating a rectification rate trend chart. The trend chart is updated weekly to provide progress reference for managers.
8. The emission source statistical data comparison, analysis, and visualization system based on localized rules according to claim 1, characterized in that, The comparison and analysis module uses a formula when performing industry data comparison. Where Idiff is the corrected difference in industry indicators, Vcurrent,i is the indicator data of industry i in the current year, Vlast,i is the indicator data of industry i in the previous year, Scurrent,i is the output value of industry i in the current year, Slast,i is the output value of industry i in the previous year, and Aavg,i is the average output value of industry i over the past three years.
Citation Information
Patent Citations
Recycling full-period monitoring method and system for waste paper products
CN119850196A
Carbon emission accounting system
CN119962825A