Pollution source general survey data summarization and analysis system

By designing a contamination source census data summary and analysis system, the shortcomings of the existing system in data summary and intelligent audit are solved, efficient data summary and intelligent audit are achieved, and data quality and processing efficiency are improved.

CN120012000AInactive Publication Date: 2025-05-16CHINA NAT ENVIRONMENTAL MONITORING CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510439855.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing pollution source data management system lacks complex summary and intelligent audit functions, resulting in low data aggregation efficiency and serious data quality problems, which cannot meet the analysis needs of large-scale pollution source data.

Method used

A pollution source census data summary and analysis system is designed, including heterogeneous data standardization module, multi-level multi-dimensional summary engine, intelligent audit cluster module, distributed computing framework and data push module, supporting structured data conversion, multi-dimensional summary, intelligent audit and parallel computing of massive data.

Benefits of technology

The provincial data aggregation time has been shortened to 1 minute and the national abnormality review time has been shortened to 30 minutes, improving data quality and processing efficiency, and standardizing calculations and data reviews for more than 3 million companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012000A_ABST
    Figure CN120012000A_ABST
Patent Text Reader

Abstract

The invention relates to a pollution source census data summarization and analysis system, and belongs to the technical field of environment big data processing. The scheme mainly comprises: a heterogeneous data standardization module, which is used for performing structured conversion and unified format processing on industrial source, agricultural source, centralized pollution control facility, mobile source, living source and non-point source data; the multi-level multi-dimensional summarization engine supports dynamic combination summarization according to administrative division levels, industry classification and pollution source types; the intelligent auditing cluster module comprises a K-means clustering model, an isolated forest anomaly detection model and a decision tree classification model; the distributed computing framework is used for realizing mass data parallel computing based on a multi-thread fragmentation processing mechanism; and the data pushing module is used for realizing provincial compression of abnormal data and automatic mail pushing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of environmental big data processing, and in particular relates to a pollution source census data summary and analysis system. Background Art

[0002] In the field of pollution source census, such as the Second National Pollution Source Census, referred to as the "Second Pollution Census", the invention and development of corresponding data analysis and summary tools are very limited and in very small quantities, unable to meet work needs.

[0003] The existing pollution source data management system has the following problems: the pollution source management system only supports basic data collection and lacks complex aggregation and intelligent review functions; efficiency bottleneck, the traditional single-threaded processing mode causes provincial data aggregation to take more than 8 hours; data quality problems, the manual review method has a missed detection rate of up to 42%, and is unable to detect deep-seated data logic errors.

[0004] The actual situation we face is the following: First, the census business has been stopped, and the corresponding business system functions can no longer be maintained and used normally, but the value of the data still needs to be further explored, and the data review and aggregation functions need to continue to be used. There are currently no tools that can achieve this function; second, the business system software invention during the census work is mainly based on collection and process review, and does not have the complex aggregation function of statistical data, and cannot export data in real time; third, there is a lack of data review modules based on big data.

[0005] In summary, existing technologies lack effective analysis tools for large-scale pollution source data. Summary of the invention

[0006] In view of the above analysis, in order to solve the above problems, an embodiment of the present invention provides a pollution source census data summary and analysis system, including: Heterogeneous data standardization module, used for structural conversion and unified format processing of industrial source, agricultural source, centralized pollution control facility, mobile source, domestic source and non-point source data; Multi-level and multi-dimensional aggregation engine supports dynamic combination aggregation by administrative division level, industry classification, and pollution source type; Intelligent audit cluster module, including K-means clustering model, isolation forest anomaly detection model and decision tree classification model; Distributed computing framework, based on multi-threaded sharding processing mechanism to achieve massive data parallel computing; The data push module realizes the compression of abnormal data by province and automatic email push.

[0007] In some embodiments, the heterogeneous data normalization module includes: Industrial source special table parsing unit, which parses G101-G104 series tables through the association mechanism between the main index table and the secondary index table; Energy conversion unit, with built-in standard coal conversion coefficient matrix for 36 types of energy, supports energy consumption comparability analysis; Text normalization unit, using SIMPLIFIED CHINESE_CHINA.UTF8 encoding to process minority characters.

[0008] In some embodiments, the multi-level multi-dimensional aggregation engine includes: Dynamic administrative division verification submodule verifies the validity of administrative division codes based on the Luhn algorithm; Clustering units are divided into different industries, and the silhouette coefficient method is used to determine the optimal industry classification scheme; Hierarchical inheritance unit, which defines district, county, prefecture, province, and national level summary functions through Python class inheritance mechanism.

[0009] In some embodiments, the intelligent audit cluster module includes: Outlier detection matrix, constructing multi-dimensional anomaly determination rules through Z-score, MIN-MAX standardization and K-means algorithm; Missing field screening unit generates a mandatory field association rule base based on a decision tree model; The problem data tracing module uses index codes to reversely trace the original reported data table.

[0010] In some embodiments, the distributed computing framework includes: Elastic thread pool management unit, dynamically allocates computing resources based on ThreadPoolExecutor; The time-sharing transfer module automatically starts writing in batches when the amount of data processed at a time exceeds 10,000; The load balancing unit uses multi-threaded processing for the industrial source G101 table and single-threaded processing for the G104 special table.

[0011] In some embodiments, the industrial source table parsing unit includes: The primary and secondary index association tool associates the sub-table data of wastewater, waste gas, solid waste, etc. through the unified social credit code of the enterprise; Pollutant accounting unit, with built-in cross-validation function for pollution generation and emission coefficient and monitoring method dual models; The production process analysis module identifies the optimal processing technology level based on the process code library.

[0012] In some embodiments, the data push module includes: The provincial screening unit automatically splits abnormal data across the country based on administrative division codes; Compression and packaging tool, using ZipFile library to generate encrypted compressed packages according to problem categories; Mail cluster management unit supports automatic SMTP protocol retry and sending status log recording.

[0013] In some embodiments, it also includes: a standard coal conversion function module, which realizes unified accounting of multiple energy sources through the following formula: ;as well as The index code generation module assigns a 15-digit unique identification code to each calculation result. The encoding rules are as follows: The first 6-digit administrative division code + 3-digit industry code + 3-digit pollution source type code + 3-digit serial number.

[0014] In some embodiments, the interaction between the system and the Oracle database includes: Dynamic splicing of SQL statements is achieved through cx_Oracle module; Use SQLAlchemy ORM framework to define database table mapping relationships; Develop a batch writing tool and set the single write data volume threshold to 1,000-10,000 records.

[0015] In some embodiments, the multi-threaded processing performance of the system satisfies: The aggregation time for provincial administrative regions (≥100,000 survey units) is ≤1 minute; The abnormal review time nationwide (3 million survey units) is ≤30 minutes.

[0016] The present invention can perform standardized calculations for more than 3 million enterprises and realize the construction of structured databases for different survey units; secondly, data audit was carried out based on the big data model during the census, tens of thousands of problems were fed back and rectification was carried out, and the provincial screening, archiving and push functions of data audit problems were also developed separately in the invention; thirdly, in the initial development process of the census business system, the census summary function was first possessed, and the data processing results corrected the formal census data collection system; fourthly, it is the only tool with the most complete existing data analysis function for the second national pollution source census data and the ability to develop national data statistical analysis, and it is an invention of a second national pollution source census data audit, summary and analysis tool that can operate normally. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0018] Figure 1 A schematic flow chart of an ozone anomaly detection method based on multi-dimensional time series analysis provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments disclosed in this disclosure can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0020] The terms used here are for the purpose of describing specific embodiments, and are not intended to be restrictive. As used here, unless the context clearly indicates otherwise, the singular forms "one (kind, person)" and "said (the)" are also intended to include plural forms. In addition, when the terms "comprise" and / or "include" and their variations are used in this specification, it is explained that there are stated features, integral bodies, steps, operations, parts, assemblies and / or their groups, but it is not excluded that there are or add one or more other features, integral bodies, steps, operations, parts, assemblies and / or their groups. It should also be noted that, as used here, the terms "substantially", "approximately" and other similar terms are used as approximate terms and not as degree terms, so that they are used to explain the inherent deviations of the measured values, calculated values ​​and / or the values ​​provided that will be recognized by those of ordinary skill in the art.

[0021] The present disclosure is described below through several specific embodiments. In order to keep the following description of the embodiments of the present invention clear and concise, the present invention omits the detailed description of known functions and known components.

[0022] Example 1 See also Figure 1 The embodiment of the present invention provides a pollution source census data summary and analysis system, including: Heterogeneous data standardization module, used for structural conversion and unified format processing of industrial source, agricultural source, centralized pollution control facility, mobile source, domestic source and non-point source data; Multi-level and multi-dimensional aggregation engine supports dynamic combination aggregation by administrative division level, industry classification, and pollution source type; Intelligent audit cluster module, including K-means clustering model, isolation forest anomaly detection model and decision tree classification model; Distributed computing framework, based on multi-threaded sharding processing mechanism to achieve massive data parallel computing; The data push module realizes the compression of abnormal data by province and automatic email push.

[0023] In some embodiments, the heterogeneous data normalization module includes: Industrial source special table parsing unit, which parses G101-G104 series tables through the association mechanism between the main index table and the secondary index table; Energy conversion unit, with built-in standard coal conversion coefficient matrix for 36 types of energy, supports energy consumption comparability analysis; Text normalization unit, using SIMPLIFIED CHINESE_CHINA.UTF8 encoding to process minority characters.

[0024] In some embodiments, the multi-level multi-dimensional aggregation engine includes: Dynamic administrative division verification submodule verifies the validity of administrative division codes based on the Luhn algorithm; Clustering units are divided into different industries, and the silhouette coefficient method is used to determine the optimal industry classification scheme; Hierarchical inheritance unit, which defines district, county, prefecture, province, and national level summary functions through Python class inheritance mechanism.

[0025] In some embodiments, the intelligent audit cluster module includes: Outlier detection matrix, constructing multi-dimensional anomaly determination rules through Z-score, MIN-MAX standardization and K-means algorithm; Missing field screening unit generates a mandatory field association rule base based on a decision tree model; The problem data tracing module uses index codes to reversely trace the original reported data table.

[0026] In some embodiments, the distributed computing framework includes: Elastic thread pool management unit, dynamically allocates computing resources based on ThreadPoolExecutor; The time-sharing transfer module automatically starts writing in batches when the amount of data processed at a time exceeds 10,000; The load balancing unit uses multi-threaded processing for the industrial source G101 table and single-threaded processing for the G104 special table.

[0027] In some embodiments, the industrial source table parsing unit includes: The primary and secondary index association tool associates the sub-table data of wastewater, waste gas, solid waste, etc. through the unified social credit code of the enterprise; Pollutant accounting unit, with built-in cross-validation function for pollution generation and emission coefficient and monitoring method dual models; The production process analysis module identifies the optimal processing technology level based on the process code library.

[0028] In some embodiments, the data push module includes: The provincial screening unit automatically splits abnormal data across the country based on administrative division codes; Compression and packaging tool, using ZipFile library to generate encrypted compressed packages according to problem categories; Mail cluster management unit supports automatic SMTP protocol retry and sending status log recording.

[0029] In some embodiments, it also includes: a standard coal conversion function module, which realizes unified accounting of multiple energy sources through the following formula: ;as well as The index code generation module assigns a 15-digit unique identification code to each calculation result. The encoding rules are as follows: The first 6-digit administrative division code + 3-digit industry code + 3-digit pollution source type code + 3-digit serial number.

[0030] In some embodiments, the interaction between the system and the Oracle database includes: Dynamic splicing of SQL statements is achieved through cx_Oracle module; Use SQLAlchemy ORM framework to define database table mapping relationships; Develop a batch writing tool and set the single write data volume threshold to 1,000-10,000 records.

[0031] In some embodiments, the multi-threaded processing performance of the system satisfies: The aggregation time for provincial administrative regions (≥100,000 survey units) is ≤1 minute; The abnormal review time nationwide (3 million survey units) is ≤30 minutes.

[0032] Example 2 In one embodiment, the present invention provides a pollution source census data summary and analysis system, including: 1. Basic code call module, mainly designed based on the second national pollution source census data code, administrative division code, industry classification code, etc. need to be called by the system module; the main functions include: ① The Python function reference modules required by the tool include mathematical calculations, Oracle database reading and writing, seaborn drawing, threading multi-threaded execution, sklearn big data analysis scientific computing, etc. The function modules required for data aggregation and analysis can greatly reduce the development difficulty and improve the development efficiency by referencing the tool; ② The text environment and database environment setting module mainly adopts SIMPLIFIED CHINESE_CHINA.UTF8 text environment, so that the minority characters in the census database can be read; and the oracle environment of the system is configured, and the database is read through the sqlalchemy module. In order to read and write the database better and faster, tools such as reading all of a certain table, reading some fields, reading a certain row and column data, and filtering and reading according to multiple conditions are developed. At the same time, a flexible definition of the database writing format is developed to store the data results in the database table, and a time-sharing queue transfer tool is developed when facing massive data to avoid system crashes. Among them, the function of reading the database is mainly based on the cx_Oracle function to spell the sql statement. For example, the get_specific function is developed to determine the sql statement by determining the variables such as the read data, indicators and screening targets, and execute the database to obtain the required data results; for the storage database, the mapping_types function is mainly developed to set the storage format and length based on text and numerical data judgment. When storing a large amount of data, the simultaneous writing speed of 1000-10000 is set. This function can read and write more than 3 million survey objects and thousands of indicators.

[0033] ③ The code module marks the institutions and developers. According to the source table classification of the second national pollution source census system, the table names that need to be read in the database of various base tables of industrial sources, agricultural sources, domestic sources, centralized pollution control facilities and mobile sources are set, such as industrial sources 'T_BAS_G101_1', 'T_BAS_G101_2', etc. Industrial sources also need to read the index tables in different separate tables for the survey objects; the national economic industry classification code, including major industries, medium industries, and minor industries; the administrative division codes above the county level in the country, which are subdivided into provinces, cities, and districts and counties; in response to the problem of large amounts of industrial source data, 2.47 million survey units have set up a reading method for the library table index by province, so as to achieve efficient reading of national data of a province, or to query data within the data of a province, thereby improving the retrieval time; at the same time, an index code function tool has been developed, which can implement specific 15-bit index codes for different calculation results and realize step-by-step calculation; 2. Basic calculation module, which is mainly designed for the calculation functions required for the summary calculation of survey units, various industries, and administrative levels, such as energy accounting, KMeans algorithm, artificial neural network algorithm, decision tree algorithm, data text format conversion, text format to data conversion and other basic work; ① The basic calculation module can realize statistical description calculations such as maximum value, minimum value, average value, standard deviation, etc. on the read data; perform percentage calculations; 0-1 judgment, whether it is a logical judgment; judge whether a value is in the column; perform data addition, non-empty calculation, e standard value, etc. The development of calculation functions is not a simple reference to numpy statements, but a detailed calculation tool for data standardization, which requires numerical or text judgment, and requires additional filtering and conversion settings, etc.

[0034] ② Special algorithm module, for 36 energy sources such as coal, natural gas, biomass energy, etc. that need to be focused on in environmental management, they are uniformly recorded as standard coal, and the conversion coefficients, energy types, and numerical reporting are standardized, so that the energy consumption of different survey units can be compared. At the same time, for enterprises that fill in multiple energy sources, their main energy sources can be indicated according to the standard coal consumption; for detailed data of multiple products, processes and raw materials, the most filled-in content in the corresponding column is calculated; for the census involving more processes and facilities, the most available processes and facilities can be calculated and determined through process codes, so as to achieve comparability of the optimal technical levels of different survey units; it can also output the proportion of different attribute indicators in the same column of data.

[0035] ③Outlier review algorithm module, developed Z-scoe, MIN-MAX, Kmeans, iforest and other algorithms, and developed the outliers_circle tool based on the results of algorithm screening and calculation, which can judge the outliers identified by the algorithm and form a predetermined report.

[0036] 3. The one-click classification, compression and push function module of data issues is mainly designed to split the unified issues by province after screening the data issues, summarize and compress them according to different issues, and push them directly to the receiving email address through the input of provincial problem feedback.

[0037] ① The abnormal problem splitting module mainly splits the data through the table index. For example, by province, the abnormal problem collection data is conditionally read according to different provinces, and the reading results are stored in a new library table or output as an Excel table; ②Data sorting, merging and packaging functions can identify data from different problems according to indexes such as provinces. First, different folders will be created in the hard disk partition, and different problem categories will be saved in the corresponding folders. The folders will then be compressed and renamed using zipfile.

[0038] ③ Email push function, develop a simple email sending tool based on tools such as smtplib, which needs to synchronize the email address to be pushed, form the target sending and email address into a dictionary item, and can also edit the title and body when sending. It also supports reminders for sending failures and the function of resending at time intervals.

[0039] 4. Industrial source standard normalization module, including the calculation of data in table G101-102; ① The basic situation description module of industrial source enterprises queries the situation of enterprises from various forms, such as industry code, production time, wastewater generation, waste gas generation, use of organic solvents, etc.; if there are required items that are not filled in, they need to be returned and deleted except for errors; the calculation of a single industrial source enterprise is realized, such as wastewater generation per unit product, energy consumption per unit product, products, production processes, etc.; the important contents, main matters, energy, products, raw materials, etc. of different industrial enterprises are summarized to form a data description of the fixed source of a single enterprise.

[0040] ② Library table index query tool module. The tables of industrial sources are very complex, including wastewater, industrial boilers, industrial furnaces, coking, sintering, ironmaking, clinker production, petrochemical industry, organic liquid storage tanks, use of raw and auxiliary materials containing volatile organic compounds, industrial solid waste, other waste gas, industrial solid waste and hazardous waste, as well as production and emission coefficient accounting and monitoring method accounting tables. Therefore, in the design of the census database, different contents are stored in different tables, and different tables have corresponding index codes, that is, through an index table, primary index and secondary index. The primary index is the basic information of the enterprise, and the secondary index corresponds to the content in other tables of different enterprises. The designed query module can realize the reading and analysis of all data by reading the primary and secondary indexes.

[0041] ③ The normalized design and calculation modules of various special tables include the basic information table of the enterprise, covering the unified social credit code, enterprise name, industry code, province, city and county, enterprise scale, receiving water body and other information; main raw materials, products, processes, energy consumption; enterprise wastewater generation, treatment and emission, covering enterprise water intake, treatment process, number of wastewater outlets, chemical oxygen demand, ammonia nitrogen, total nitrogen, total phosphorus, petroleum, heavy metals and other pollutants generation and emission; industrial boilers, kilns installed capacity, energy consumption, treatment process, sulfur dioxide, nitrogen oxides, particulate matter and volatile organic compounds, heavy metals and other pollutants generation and emission, etc., as well as other coking, sintering pelletizing, ironmaking, clinker production, petrochemical industry, organic liquid storage tanks, use of raw and auxiliary materials containing volatile organic compounds, industrial solid waste, other waste gas, industrial solid waste and hazardous waste, as well as the corresponding indicators and calculation results of the pollution generation and emission coefficient accounting and monitoring method accounting table.

[0042] ④Special table separate calculation and overall summary module. The industrial source survey table is divided into different special tables. Through the Industry_get_info tool, the calculation of all library tables and sub-tables can be realized, and step-by-step calculation can be realized, which supports the summary of pollutant generation and emission in the sub-tables.

[0043] ⑤ Rapid data analysis function module. Since the screening of enterprise information such as enterprise name, unified credit code, organizational structure code, etc. often requires further reading of the index table, in order to read and calculate data faster, we have developed a corresponding rapid calculation through ID index code.

[0044] 5. Industry data analysis function module, which determines the pollutants generated, the products, production processes and raw and auxiliary materials involved in the industry, and treatment technologies from the census database and industrial source coefficient form; ① Industry situation description module, for example, the proportion of wastewater generated, the proportion of boilers used, the proportion of certain treatment processes adopted, the proportion of energy used, the proportion of various types of enterprises, etc.; ②Industry analysis module, classify the situations of different industries and output the statistical contents to the database. It is better to cluster them according to the situations of industry subcategories to find out the industries with similar management and the industries that need key control. ③ The statistical characteristic accounting module clusters different industries, finds out the emission characteristics of the industry, and calculates the emission volume based on the emission characteristics; analyzes the coefficients of the industry, and reviews the generation and emission of wastewater and waste gas that need to be reported.

[0045] ④ The industry summary tool module is similar to the summary statistical function of industrial sources. It can realize the screening of conditions such as industry code and administrative area code, and screen out the corresponding data of a certain province, city, county or key area, a certain major, medium and minor industry, including wastewater, industrial boilers, industrial furnaces, coking, sintering pellets, ironmaking, clinker production, petrochemical industry, organic liquid storage tanks, use of raw and auxiliary materials containing volatile organic compounds, industrial solid waste, other waste gas, industrial solid waste and hazardous waste, as well as the calculation results of production and emission coefficient accounting and monitoring method accounting table, and can also realize functions such as rapid summary and sub-table summary.

[0046] 6. Standard normalization module for livestock and poultry farms and centralized pollution control facilities, including data accounting for urban sewage treatment plants, domestic garbage plants, and centralized hazardous waste treatment plants; the main module settings are similar to those for industrial sources, and functional modules such as basic information description, special table accounting, and statistical feature analysis and summary have been developed separately for different types of survey objects.

[0047] ① The standard normalization module for large livestock and poultry farms is designed to read the index information of the agricultural source N101-1 and N101-2 tables, the corresponding basic information reports, the generation and emission of pollutants, and can realize differentiated summary analysis based on the farm type, livestock and poultry breeding type, and breeding scale.

[0048] ② The normalization module of centralized pollution control facilities is designed according to the storage format of centralized pollution control facilities in the census database. The index information of J101-1, J101-2 and J101-3 tables of sewage treatment facilities is calculated to realize the calculation of basic information, sewage treatment methods, optimal treatment processes, sewage design treatment capacity, actual treatment volume, pollutant removal volume, etc. It also realizes differentiated summary analysis of the presence or absence of recycled water treatment processes, sludge stabilization treatment methods, recycled water scale, sewage treatment facility types and design treatment capacity scale.

[0049] ③ The standard normalization module for domestic waste treatment plants is designed according to the storage format of domestic waste sites in the census database. The index information of J102-1, J102-2 and J104 tables of centralized waste treatment plants is designed to realize the calculation of different treatment methods of domestic waste sites, the filled capacity of landfills, the landfill volume, the operation load of waste incineration plants, the membrane concentrate treatment method, the production and emission of major pollutants and other related indicators, and can realize differentiated summary analysis of landfill, incineration, anaerobic fermentation, biological decomposition, other treatment methods and different treatment capacity scales.

[0050] ④ Hazardous waste disposal plants. According to the storage format of hazardous waste disposal facilities in the census database, the index information of J103-1, J103-2 and J104 tables of centralized hazardous waste disposal sites is designed to realize the calculation of different treatment methods for centralized hazardous waste disposal, the calculation of different types of disposal volume of industrial hazardous waste and medical waste, the calculation of different indicators such as operation mode of different treatment methods, proportion of energy consumption, and production and emission of major pollutants, and can also realize differentiated summary analysis of treatment methods such as burial and incineration, and different treatment capacity scales.

[0051] 7. Standard normalization module for other point source and surface source data, including data on towns, administrative villages, planting, livestock and poultry breeding, aquaculture, motor vehicle ownership, construction machinery ownership, etc. Since the table is relatively simple, some data are extracted and some indicators are also normalized. For example, in the oil storage, transportation and sales table, whether there is accounting for volatile organic compound emissions from oil and gas recovery devices, the number of heavy diesel trains is calculated in the motor vehicle ownership, as well as the corresponding pollutant emissions, etc.

[0052] 8. The administrative region analysis module can realize flexible real-time summary of different industries, different administrative levels, different regional combinations, and different report selections according to different needs.

[0053] ① The administrative region code processing module sets the administrative division codes of provinces, cities and counties in the 34 provincial-level administrative regions across the country, and develops a screening and identification tool for administrative division levels based on coding rules; an administrative division jurisdiction analysis module is developed, which can identify all administrative regions under its jurisdiction based on the input administrative division code, so as to ensure that any administrative division input can be judged whether it is wrong, which administrative level it belongs to, and the areas under its jurisdiction, etc.

[0054] ② The list screening module of the survey units in the region. In order to realize the summary, it is necessary to obtain the list of various survey units in the administrative area, or to obtain the list screening of the operating status of a certain industry, a certain city or county, or a certain enterprise according to the conditions.

[0055] ③ The summary analysis function can be a county, a prefecture-level city, a province, or the whole country. It summarizes and calculates the situation of enterprises in the province, such as the total number of enterprises and the number of enterprises in different industries. It is mainly based on the results of the census data summary, and can also be based on the data reported in the base table; it realizes the data of industrial enterprises in the region, defines the list of enterprises in the region, calculates the management level, and conducts a detailed analysis of the enterprises; it realizes the collection and analysis of data on other contents in the region, the situation of river codes, and sorts out the situation of rivers in the region; it collects the situation of centralized governance facilities, summarizes the number, process, and discharge outlets of centralized facilities; it collects the situation of domestic boilers, summarizes the number of domestic boilers and sewage outlets into rivers, and the networking situation; it collects data related to domestic sources, including data from urban construction departments such as population; it collects data related to agricultural sources, including various types of livestock and poultry breeding data in the region; it collects data related to mobile sources, including data on the basic situation of motor vehicles in the region, and the situation of pollutant generation and emissions in different regions; it clusters regional conditions, classifies by city, and conducts precise measurements; the developed data summary and analysis function is synchronized with the summary analysis and calculation of the survey unit.

[0056] ④ Define modules step by step. Due to the different levels of administrative regions, there are great differences in the data results obtained from the census work, and there are also great differences in the indicators that can be summarized and analyzed. Through the inheritance function of Python, rural, district, county, city, province and national functions are defined. The functions can be called directly according to the level that needs to be analyzed. The functions include basic calculation functions, administrative regions, tables involved in analysis and summary analysis functions, which can directly query and develop extended calculation functions.

[0057] 9. Multi-threaded fast summary module, through the threading program, can automatically optimize the distribution of thread calculations during the process of large-scale data processing, thereby greatly optimizing the process speed. For a data volume of about 100,000 households, the calculation can be completed within 1 minute.

[0058] ① Rapid summary definition module. Administrative regions with a small number of survey units can be directly summarized. However, if there are many survey units, such as provinces such as Jiangsu, Guangdong, and Shandong, the reading and data processing speed of the national database will be slowed down. Different reading and processing speeds are set for different types of summary indicators. For example, table G101-1 is set for threaded processing, while table G104 and other special tables for cement, steel, etc. use single-threaded processing. At the same time, an integration function for various summary modules of industrial sources, agricultural sources, centralized pollution control facilities, domestic sources, and mobile sources is set, which can generate summary data of various sources of the second national pollution source census in administrative regions with one click.

[0059] ② The multi-thread processing module mainly uses the thread module function to automatically split the read data content according to the total amount for multi-thread processing, which can reduce the processing time by 30%. For the summary module, it is processed through the processing module.

[0060] ③ Automatic save function. Due to the large amount of data reading and calculation, the server is prone to overload, resulting in data loss. A step-by-step save method is designed. According to the amount of data processing or the processing time, when a certain time limit or a certain load is reached, the automatic database write function is started, which can realize the saving and calling of process data, release the memory after saving, and ensure the stability of code operation.

[0061] ④ Index quick query module: Based on the index query of the survey unit, a quick query tool that can correspond to the index and sub-index is specially designed. It mainly reads the index, splits it according to the amount of data, and then distributes it, gradually queries the sub-index, and alleviates the code load.

[0062] ⑤ Administrative division code verification function: the administrative divisions filled in for some survey data are incorrect, which will affect the summary results. By comparing and matching in the administrative division code list, the data that has not been matched and is unqualified will be screened and reminded.

[0063] 10. National-level audit function module: Based on the audit functions that need to be carried out during the second national pollution source census, a functional module for statistical data has been specially developed, including missing and gap filling based on coefficient algorithms, etc.

[0064] ① Data screening output and save function: In order to save data results and conduct more detailed analysis, a one-click transfer function of the library table content in the system is developed.

[0065] ② The outlier review function module can use algorithms such as Kmeans, zscore, and isolation forest to conduct one-click review of outliers in the numerical data of all survey units in the system according to industry, city, etc., and output a template for data return and rectification.

[0066] ③ Perform statistical classification through decision tree models, perform data underfill and underreporting screening based on statistical laws, and develop various audit tools as needed.

[0067] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0068] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0069] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A pollution source census data summary and analysis system, characterized in that: include: Heterogeneous data standardization module, used for structural conversion and unified format processing of industrial source, agricultural source, centralized pollution control facility, mobile source, domestic source and non-point source data; Multi-level and multi-dimensional aggregation engine supports dynamic combination aggregation by administrative division level, industry classification, and pollution source type; Intelligent audit cluster module, including K-means clustering model, isolation forest anomaly detection model and decision tree classification model; Distributed computing framework, based on multi-threaded sharding processing mechanism to achieve massive data parallel computing; The data push module realizes the compression of abnormal data by province and automatic email push.

2. The pollution source census data summary and analysis system according to claim 1 is characterized by: The heterogeneous data standardization module includes: Industrial source special table parsing unit, which parses G101-G104 series tables through the association mechanism between the main index table and the secondary index table; Energy conversion unit, with built-in standard coal conversion coefficient matrix for 36 types of energy, supports energy consumption comparability analysis; Text normalization unit, using SIMPLIFIED CHINESE_CHINA.UTF8 encoding to process minority characters.

3. The pollution source census data summary and analysis system according to claim 1 is characterized by: The multi-level multi-dimensional aggregation engine includes: Dynamic administrative division verification submodule verifies the validity of administrative division codes based on the Luhn algorithm; Clustering units are divided into different industries, and the silhouette coefficient method is used to determine the optimal industry classification scheme; Hierarchical inheritance unit, which defines district, county, prefecture, province, and national level summary functions through Python class inheritance mechanism.

4. The pollution source census data summary and analysis system according to claim 1 is characterized by: The intelligent audit cluster module includes: Outlier detection matrix, constructing multi-dimensional anomaly determination rules through Z-score, MIN-MAX standardization and K-means algorithm; Missing field screening unit generates a mandatory field association rule base based on a decision tree model; The problem data tracing module uses index codes to reversely trace the original reported data table.

5. The pollution source census data summary and analysis system according to claim 1 is characterized by: The distributed computing framework includes: Elastic thread pool management unit, dynamically allocates computing resources based on ThreadPoolExecutor; The time-sharing transfer module automatically starts writing in batches when the amount of data processed at a time exceeds 10,000; The load balancing unit uses multi-threaded processing for the industrial source G101 table and single-threaded processing for the G104 special table.

6. The pollution source census data summary and analysis system according to claim 2 is characterized by: The industrial source special table parsing unit includes: The primary and secondary index association tool associates the wastewater, waste gas, and solid waste sub-table data through the enterprise unified social credit code; Pollutant accounting unit, with built-in cross-validation function for pollution generation and emission coefficient and monitoring method dual models; The production process analysis module identifies the optimal processing technology level based on the process code library.

7. The pollution source census data summary and analysis system according to claim 1 is characterized by: The data push module includes: The provincial screening unit automatically splits abnormal data across the country based on administrative division codes; Compression and packaging tool, using ZipFile library to generate encrypted compressed packages according to problem categories; Mail cluster management unit supports automatic SMTP protocol retry and sending status logging.

8. The pollution source census data summary and analysis system according to claim 7, characterized in that: Also includes: The standard coal conversion function module realizes unified accounting of multiple energy sources through the following formula: ; as well as The index code generation module assigns a 15-digit unique identification code to each calculation result. The encoding rules are as follows: The first 6-digit administrative division code + 3-digit industry code + 3-digit pollution source type code + 3-digit serial number.

9. The pollution source census data summary and analysis system according to claim 1 is characterized by: The interaction between the system and the Oracle database includes: Dynamic splicing of SQL statements is achieved through cx_Oracle module; Use SQLAlchemy ORM framework to define database table mapping relationships; Develop a batch writing tool and set the single write data volume threshold to 1,000-10,000 records.

10. The pollution source census data summary and analysis system according to claim 1, characterized in that: The multi-thread processing performance of the system meets the following requirements: The aggregation time for provincial administrative regions is ≤ 1 minute; The national abnormal review time is ≤30 minutes.

Citation Information

Patent Citations

  • Method for recognizing heavy metal pollution source in soil

    CN105631203A

  • Method, apparatus and device for locating pollution source on basis of big data, and storage medium

    WO2021174751A1