A multi-modal clustering-based high-value pollution diagnosis method and system for state-controlled sites
The high-value pollution diagnosis method for national monitoring stations using multimodal clustering solves the problems of high implementation threshold and low response efficiency of high-value pollution diagnosis methods. It achieves second-level identification of high-value pollution events, provides real-time technical support, is applicable to high-value pollution diagnosis in different regions, and reduces data acquisition and system deployment costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京首创大气环境科技股份有限公司
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-21
AI Technical Summary
Existing high-value pollution diagnostic methods have high implementation thresholds, low response efficiency, and weak regional adaptability, making it difficult to meet the needs of routine and large-scale pollution prevention and control.
A high-value pollution diagnosis method for national monitoring stations using multimodal clustering is proposed, which includes multi-source data standardization preprocessing, construction of a high-value automatic identification rule base, high-value process identification and multimodal feature construction, intelligent classification of high-value process pollution patterns, classification of pollution pattern causes, and model self-learning and continuous optimization. By constructing integrated data preprocessing and pollution pattern cause classification, the entire process of automated diagnosis is achieved.
It achieves high-value pollution event identification in seconds, provides real-time technical support, reduces reliance on manual intervention, is applicable to high-value pollution diagnosis in different regions, reduces data acquisition and system deployment costs, and has broad application scenario adaptability.
Smart Images

Figure CN122432875A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of air pollution control technology, and in particular to a method and system for diagnosing high-value pollution at national monitoring stations based on multimodal clustering. Background Technology
[0002] High-value pollution levels at national monitoring stations are key signals for identifying local pollution sources, and their rapid diagnosis is crucial for precise atmospheric pollution control. Currently, high-value analysis mainly relies on human experience, physical models, or dense monitoring networks. In practical engineering applications, this has revealed prominent problems such as high implementation thresholds, low response efficiency, and weak regional adaptability, making it difficult to meet the needs of routine, large-scale pollution control.
[0003] Specifically, physical model-based methods rely on high-precision meteorological fields or pollution source inventories, and their effectiveness is easily limited in areas with relatively insufficient basic data support (e.g., patent publications CN109583743A and CN118782169A). While integrated air-space-ground monitoring can provide rich information, its widespread application faces challenges due to equipment and deployment costs (e.g., patent publication CN111121862A). Traditional manual analysis methods also show some pressure when dealing with large-scale, real-time data, while existing automated methods lack closed-loop learning mechanisms, making continuous optimization difficult (e.g., patent publication CN113919533A). Furthermore, some technical approaches require high monitoring network density, and their identification and diagnostic capabilities may not be fully realized in counties in the central and western regions where monitoring network coverage is insufficient (e.g., patent publications CN113763221A and CN112131739A).
[0004] Against this backdrop, the industry urgently needs a high-value pollution intelligent diagnostic solution that combines high efficiency, universal applicability, low barriers to entry, full-process automation, and accurate diagnosis. This solution aims to overcome existing technological bottlenecks and achieve second-level identification, intelligent classification, and accurate cause determination of high-value pollution, thereby providing efficient and reliable technical support for precise prevention and control of air pollution in different regions and scenarios. Summary of the Invention
[0005] The purpose of this invention is to address the technical problems in existing high-value pollution diagnosis, such as high data requirements, strong model dependence leading to high implementation thresholds, insufficient efficiency of manual analysis resulting in delayed response, and poor adaptability of static methods leading to difficulties in regional promotion. This invention provides a high-value pollution diagnosis method and system based on multimodal clustering for national control stations that combines low threshold, high timeliness, and strong adaptability.
[0006] To achieve the above objectives, this invention provides a method for diagnosing high-value pollution at national monitoring stations based on multimodal clustering. The method includes the following steps: standardized preprocessing of multi-source data; construction of a rule base for automatic high-value identification; identification of high-value processes and construction of multimodal features; intelligent classification of pollution patterns in high-value processes; classification of pollution pattern causes; intelligent diagnosis of new high-value events; and model self-learning and continuous optimization to complete the intelligent diagnosis of high-value pollution at national monitoring stations.
[0007] Optionally, the multi-source data standardization preprocessing of the target site includes: stable acquisition of multi-source data; missing data repair and removal; abnormal observation value identification and filtering; and multi-scale data standardization. This invention proposes an integrated closed-loop processing technology solution of stable acquisition-repair-removal-standardization. By constructing a data type differentiation repair mechanism (using linear interpolation for pollutant concentration data and forward filling for meteorological parameters), it effectively solves the technical challenges of inconsistent formats, redundant invalid values, skewed distribution, and interference from abnormal data in multi-source heterogeneous data. Through precise noise reduction and unified normalization mechanisms for multi-source data, it effectively overcomes the limitations of traditional data processing, significantly improves data integrity and standardization, and fully preserves the core features and temporal continuity of the data. This provides standardized, high-quality time-series data support for the subsequent application of the high-value identification rule base, ensuring the accuracy of subsequent rule matching and high-value determination. It guarantees the reliability and stability of the entire diagnostic process from the data source, laying a solid data foundation for intelligent diagnosis of high-value pollution.
[0008] Optionally, the construction of the high-value automatic identification rule base includes: formal definition of high-value judgment rules; setting of high-value judgment combination logic; and adaptive strategy for high-value judgment thresholds. This invention formalizes high-value judgment logic into three core quantitative conditions: trigger, duration, and termination, creating a multi-threshold collaborative combination (dynamic threshold, static threshold, minimum threshold, and increasing threshold) rule system. This effectively solves the technical problems of poor scenario adaptability, ambiguous rule logic, and high rates of false positives and false negatives associated with traditional single-threshold rules. This multi-threshold collaborative and adaptive rule system effectively overcomes the scenario limitations of traditional fixed rules, significantly enhancing the accuracy and regional adaptability of rule execution. It provides a scientific and standardized basis for the accurate identification of subsequent discrete high-value moments, significantly improving the targeting and regional universality of high-value identification, providing accurate high-value moment data support for the aggregation of complete pollution processes, and laying a technical foundation for the complete characterization of high-value pollution processes.
[0009] Optionally, the high-value process identification and multimodal feature construction includes automatic identification and marking of discrete high-value moments; aggregation of continuous high-value pollution processes; construction of multimodal diagnostic feature vectors; and establishment of a pollution pattern feature matrix. This invention proposes a discrete high-value aggregation technology logic that couples common pollutant inspection with continuous verification, designs a process boundary correction mechanism with a 24-hour backtracking window, and creates a multimodal feature extraction system encompassing four dimensions: pollution intensity, temporal evolution, meteorological correlation, and chemical composition. This effectively solves the technical problems of discrete high values lacking complete process descriptions and the one-sidedness of single feature analysis leading to insufficient basis for subsequent causal classification. The complete pollution process identification and multimodal feature extraction mechanism effectively overcomes the limitations of traditional feature analysis, significantly improving the comprehensiveness and quantification of pollution event characterization. It achieves complete identification from discrete high-value moments to the pollution process, and the constructed multimodal feature matrix comprehensively covers the core attributes of pollution events, accurately quantifying the essential characteristics of pollution. It provides comprehensive and standardized feature inputs for subsequent intelligent classification of pollution patterns, ensuring the scientific rigor and reliability of cluster analysis, and constructing a quantitative feature correlation system between pollution phenomena and causes, providing key technical support for pollution cause classification.
[0010] Optionally, the intelligent classification of high-value process pollution patterns includes: feature matrix reconstruction and enhancement; intelligent data imputation and feature standardization; SOM clustering model execution; post-processing and quality assessment of clustering results; business-oriented interpretation of pollution patterns; and summarization of pollution pattern clustering results. This invention creates a feature optimization system with enhanced key feature weights, proposes a PCA-based SOM network weight initialization method, and designs a small cluster merging optimization and a three-dimensional quality assessment mechanism for cluster density, feature concentration, and quantization error, effectively solving the technical bottlenecks of strong subjectivity and low clustering resolution in manual classification. This key weight-enhanced SOM clustering model effectively overcomes the blindness and low resolution problems of traditional clustering, significantly improving the objectivity and interpretability of the classification results. It provides a clear basis for pollution pattern classification for subsequent pollution cause determination, clarifies the initial direction of pollution causes, and provides technical support for accurate determination of pollution causes.
[0011] Optionally, the pollution pattern causation classification includes: constructing a pollution causation classification rule base; setting causation determination priorities; and summarizing pollution causation classification results. This invention creates a modular quantitative rule base covering 6 core pollution sources, 3 special pollution sources, and 2 data anomalies, designs a four-level priority determination mechanism (data anomaly determination > special pollution determination > conventional pollution source determination > fallback handling), and establishes a multi-source coexistence labeling and judgment basis full traceability technical solution. This effectively solves the technical problems of chaotic pollution causation determination logic, difficulty in distinguishing multi-source composite pollution, and lack of scientific basis for causation results. This modular rule base and four-level priority determination mechanism effectively overcome the ambiguity and inefficiency of traditional causation determination logic, significantly improve the accuracy and traceability of pollution causation results, and can effectively distinguish the core characteristics of different pollution causes. It provides standardized and reusable pollution causation classification rules and logical support for subsequent intelligent diagnosis of new high-value events, achieving precise location of pollution causes.
[0012] Optionally, the intelligent diagnosis of new high-value events includes: input and preprocessing of new high-value event data; identification and feature extraction of high-value processes; pollution pattern matching and confidence assessment; and classification and determination of pollution pattern causes. This invention creates a fully automated closed-loop technical solution encompassing preprocessing, high-value identification, feature extraction, pattern matching, and pollution cause determination. It proposes a rapid pollution pattern matching algorithm based on weighted Euclidean distance and establishes a two-dimensional confidence assessment system (number of triggering rules × pattern matching degree), effectively solving the technical challenges of low efficiency, strong reliance on manual intervention, and lack of quantitative assurance of result reliability in traditional high-value event diagnosis. This fully automated diagnosis and confidence assessment mechanism effectively overcomes the lag and subjectivity of traditional manual diagnosis, significantly improving diagnostic efficiency and result reliability. Furthermore, a closed-loop mechanism triggering manual review for low-confidence events provides effective feedback data (low-quality diagnostic cases) in real-world scenarios for subsequent model self-learning and continuous optimization, enabling rapid response and accurate diagnosis of new high-value events and saving crucial time for pollution emergency response.
[0013] Optionally, the model self-learning and continuous optimization includes: expert feedback data collection and standardization; incremental learning triggering and execution; full retraining mechanism and execution; and model version management and deployment. This invention creates a closed-loop evolutionary technology system of expert feedback-data standardization-incremental learning-full retraining-version management-canary release. It employs a feedback sample weighted training strategy, a dynamic triggering mechanism for incremental and full retraining, and a version traceability and canary release guarantee scheme. This effectively solves the technical bottlenecks of traditional models, such as model fixation after one training iteration, inability to adapt to dynamic changes in pollution characteristics, and poor system stability due to chaotic version management. This closed-loop self-learning and version management mechanism effectively overcomes the static limitations of traditional models, significantly enhancing the system's dynamic adaptability and long-term stability. The accuracy of pollution pattern clustering and pollution cause determination continuously improves with application scenarios, and the version management mechanism ensures the traceability of system operation. This provides continuous optimization technical support for the long-term application of the system, ensuring that diagnostic accuracy continuously improves with application scenarios, and significantly enhancing the system's long-term applicability and core competitiveness.
[0014] Secondly, the present invention provides a high-value pollution diagnosis system for national monitoring stations based on multimodal clustering. The system is used to execute the high-value pollution diagnosis method for national monitoring stations based on multimodal clustering described in the first aspect. The system includes a data acquisition layer module, a data preprocessing module, a high-value identification module, a feature engineering module, a clustering and typing module, a causal diagnosis module, a new high-value intelligent diagnosis module, and a feedback and learning module.
[0015] Based on the above technical solution, the advantages of the present invention are: (1) Data-driven high-efficiency response: Based on the core technology path of data-driven and machine learning integration, the efficiency of the diagnostic process is greatly improved, and the diagnostic results are output in seconds. It can quickly identify sudden pollution events, provide immediate technical support for pollution emergency response, and effectively solve the pain points of slow traditional diagnosis and delayed emergency response. (2) Self-optimization and strong adaptability: Relying on closed-loop feedback and automatic optimization mechanism, the system can dynamically adapt to the changes in pollution characteristics in different times and spaces, thus being suitable for high-value pollution diagnosis needs at different regional scales without frequent manual adjustments; at the same time, the fully automated analysis and design process greatly reduces the reliance on manual intervention, lowers labor costs and operating thresholds, and balances adaptability and ease of use. (3) Low threshold and easy to implement: It supports full-link diagnosis based on monitoring data from ordinary national control stations, without the need for special data acquisition equipment, which significantly reduces the cost of data acquisition and system deployment; it adopts a modular architecture and configurable design, supports offline and cloud dual-mode deployment, and can be quickly implemented and promoted without complex hardware and professional algorithm knowledge, and has a wide range of application scenarios.
[0016] This invention constructs a full-link intelligent diagnostic system for high-value processes, encompassing data preprocessing, feature extraction, pollution pattern classification, pollution cause determination, and automatic model optimization, through close collaboration and deep functional integration of various modules. It systematically solves the core problems of existing methods, such as high implementation thresholds, slow response, and poor adaptability, from a technical perspective. Real-time information interaction and deep functional coupling between modules form a closed-loop synergistic effect, providing strong system support for precise prevention and scientific governance of air pollution. It possesses significant technological innovation value and broad engineering application prospects. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 A flowchart of a method for diagnosing high-value pollution at national monitoring stations based on multimodal clustering; Figure 2 This is an example of air quality monitoring data from national air quality control stations in an embodiment of the present invention; Figure 3 This is an example of meteorological data in an embodiment of the present invention; Figure 4 This is an example of the historical high value identification result in an embodiment of the present invention; Figure 5 This is an example of the high-value process identification result in an embodiment of the present invention; Figure 6 This is an example of a pollution mode feature matrix in an embodiment of the present invention; Figure 7 This is an example of cluster optimization results in an embodiment of the present invention; Figure 8 This is an example of clustering results in an embodiment of the present invention; Figure 9 This is an example of a historical case library in the embodiments of the present invention; Figure 10 This is an example of new high-value air quality data in an embodiment of the present invention; Figure 11 This is an example of new high-value meteorological data in an embodiment of the present invention; Figure 12 This is an example of new high value identification and its features in an embodiment of the present invention; Figure 13 This is the pollution cause classification result of the XGBoost model for the new high-value event in the embodiments of the present invention; Figure 14 This is an example of the intelligent diagnostic results for new high-value events in an embodiment of the present invention; Figure 15 This is an example of expert feedback and its incremental training results in an embodiment of the present invention; Figure 16 This is a diagram of a high-value pollution diagnostic system for national monitoring stations based on multimodal clustering. Detailed Implementation
[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0019] This invention overcomes the limitations of inconsistent heterogeneous data formats through an integrated preprocessing mechanism for multi-source data, providing high-quality data support for diagnosis. It solves the problem of poor adaptability in traditional fixed-threshold scenarios through a dynamically adaptable high-value identification rule system. Based on multimodal features and enhanced clustering algorithms, it achieves objective classification of pollution patterns, eliminating the subjectivity of manual analysis. Relying on a modular rule base and priority determination mechanism, it achieves accurate determination of pollution causes. Through a self-learning optimization mechanism, it enables dynamic model evolution, ensuring adaptability for long-term application. Finally, through fully automated design, it significantly improves diagnostic timeliness, providing effective technical support for scientific decision-making and refined management of air pollution.
[0020] One embodiment of the present invention is based on a high-value pollution diagnosis method for national monitoring stations based on multimodal clustering. This method collects routine monitoring data from seven national monitoring stations in Jiujiang City (Xiyuan, Shili, Maoshantou, Petrochemical Plant, Foreign Language School, No. 8 Lianhua Avenue, and Zonghe Industrial Park) from January 1, 2022 to May 1, 2025, including six pollutants (PM2.5, PM2.5, PM10, PM2.5 ... 2.5 PM 10 The system collected approximately 1.3 million data points, including four meteorological parameters (wind speed, temperature, etc.) with a time resolution of one hour.
[0021] like Figure 1 As shown, the high-value pollution diagnosis method for national monitoring stations specifically includes the following steps: Step S1: Preprocessing of monitoring data from national monitoring stations; This step establishes an integrated preprocessing pipeline of "stable acquisition - cleaning and repair" to standardize air quality and meteorological monitoring data from national control stations, ensuring that the data quality meets the needs of subsequent high-value identification, feature extraction and other analysis, and fundamentally solves the problems of inconsistent data formats, many invalid values and poor stability.
[0022] S1.1: Stable acquisition of multi-source data from national monitoring stations; Data Content: This data includes station-level data and corresponding city-level data from national monitoring stations under the jurisdiction of Jiujiang City (including Xiyuan, Shili, Maoshantou, Petrochemical Plant, Foreign Language School, No. 8 Lianhua Avenue, and Comprehensive Industrial Park), covering SO2, NO2, O3, CO, and PM2.5. 10 PM 2.5 The data includes concentrations of six core pollutants, as well as six key meteorological parameters: temperature, humidity, air pressure, wind speed, wind direction, and precipitation. Figure 2 , Figure 3 As shown.
[0023] Data integrity assurance: For approximately 1.3 million monitoring data points from national monitoring stations between January 1, 2022 and January 1, 2025, with a time resolution of 1 hour, to ensure the continuity and integrity of the data sequence, missing or abnormal data records are identified and logged, and corresponding notifications are triggered.
[0024] Data filtering: Parametric filtering is performed by time range (accurate to the hour), city code (such as the administrative division code corresponding to Jiujiang City), and site type (focusing on national control sites in this study) to avoid invalid data redundancy.
[0025] S1.2: Missing data repair and removal; This step is the core of the data preprocessing for monitoring data from national control stations. Its purpose is to establish a standardized repair and removal mechanism to address data loss caused by equipment malfunctions or transmission interruptions during data acquisition. Its function is to filter valid data and supplement recoverable data, providing a complete and reliable data source for subsequent analyses such as high-value process identification and pollution cause determination, thus avoiding analytical biases caused by missing data. The specific steps are as follows: (1) Repair threshold setting: Define the continuous missing duration threshold as 2 hours, that is, when the continuous missing duration of a certain parameter is ≤2 hours, it is judged as short-term missing; Repair method: For short-term missing data, the repair method is adaptively selected according to the data type: (1) The linear interpolation method is used for pollutant concentration data; (2) The forward filling method is used for meteorological data (such as wind direction and wind speed) to ensure the continuity of data time sequence.
[0026] (2) Elimination rules: When the continuous missing duration is >2 hours, the data in that period is determined to be invalid and the complete data sequence corresponding to that period is directly eliminated. If a high-value process has been initially identified, the associated high-value process is eliminated simultaneously to avoid invalid data interfering with the analysis results.
[0027] S1.3: Identification and filtering of outlier observations; (1) Invalid value identification strategy: The types of invalid values identified include negative numbers and invalid flag values (such as 999, NAN, etc.). (2) Filtering rules: Replace invalid values with invalid values and then supplement them according to the missing data processing logic to ensure the validity of the data sequence.
[0028] Step S2: Construction of a rule base for automatic high-value recognition; This step is one of the core innovative steps of this invention. It breaks through the limitation of traditional high-value identification relying on a single fixed threshold and constructs a high-value identification rule base that can adapt to regional features and supports flexible configuration. Through multi-dimensional threshold collaboration and dynamic logic combination, it achieves accurate capture of discrete high-value moments, fundamentally solving the technical difficulties of traditional methods, such as poor adaptability in different regions and monitoring scenarios, and susceptibility to misjudgment and missed judgment due to environmental background differences. The specific construction steps are as follows: S2.1: Definition of high value determination rules; High-value identification employs a multi-dimensional threshold collaborative judgment mechanism, setting four core threshold categories to replace the traditional single-threshold judgment logic. High-value marking is triggered only if the pollutant concentration meets any combination of the following conditions: (1) Basic combination: simultaneously exceeding the dynamic threshold, static threshold and minimum threshold; (2) Sudden combination: simultaneously exceeding the static threshold, the minimum threshold, and the growth threshold; (3) Strong high value rule: simultaneously exceeding the dynamic threshold, static threshold, minimum threshold, and growth threshold; The thresholds above are defined as follows: Dynamic threshold: A relative threshold derived from historical data of pollutants, reflecting the relatively high level of pollutants under current environmental conditions; adapting to changes in pollution concentration over different time periods and regions; providing a dynamic benchmark and avoiding the limitations of fixed thresholds. The calculation formula is as follows: Dynamic threshold = Average concentration of all stations at the same time point × Threshold dynamic factor; Based on the air pollution prevention and control requirements of Jiujiang City, the threshold dynamic multiple of each pollutant is uniformly set to 1.2 to reflect the relatively high value characteristics of pollutants based on the current overall pollution level in the region. The method of the present invention can be flexibly adjusted according to the requirements of different cities and different periods.
[0029] Static threshold: An incremental threshold based on the average pollutant concentration at all monitoring stations at the current moment. Its purpose is to: (1) provide a stable concentration gradient requirement based on the dynamic average; (2) ensure that identified high-value points have sufficient concentration increases; and (3) provide a dual-protection mechanism as a supplement to the dynamic threshold. The calculation formula is as follows: Static threshold = average concentration of all stations at the same time point + threshold increment; This embodiment is based on the air pollution prevention and control requirements of Jiujiang City, and the threshold increment is set as follows: SO2: threshold increment = 5 μg / m³ 3 NO2: Threshold increment = 10 μg / m 3 O3: Threshold increment = 10 μg / m 3 CO: Threshold increment = 0.5 mg / m³ 3 PM 10Threshold increment = 20 μg / m 3 PM 2.5 Threshold increment = 10 μg / m 3 The method of this invention can be flexibly adjusted according to the requirements of different cities and different periods.
[0030] Minimum threshold: The lowest trigger concentration set to avoid misjudgment of low concentration fluctuations, based on the background concentration characteristics of each pollutant. It can be used to: (1) filter the background concentration levels of each pollutant to avoid misjudgment of normal low concentration fluctuations; (2) ensure that the identified high values have practical environmental significance in terms of concentration magnitude; (3) eliminate the influence of monitoring equipment noise and small fluctuations on the identification results. The specific values are as follows: This embodiment is based on the air pollution prevention and control requirements of Jiujiang City, where the minimum threshold for SO2 is 20 μg / m³. 3 NO2: Minimum threshold = 20 μg / m 3 O3: Minimum threshold = 150 μg / m 3 CO: Minimum threshold = 1 mg / m³ 3 PM 10 Minimum threshold = 30 μg / m3, PM 2.5 Minimum threshold = 20 μg / m 3 The method of this invention can be flexibly adjusted according to the requirements of different cities and different periods.
[0031] Growth threshold: A threshold based on the relative increase in pollutant concentration. It can be used to: (1) identify sudden increases and rapid accumulation processes of pollutants; and (2) promptly detect potential environmental pollution events or abnormal emissions. The calculation formula is as follows: Growth rate = current concentration / previous concentration; This embodiment is based on the air pollution prevention and control requirements of Jiujiang City, and the concentration growth rate threshold is defined as 1.5. The method of the present invention can be flexibly adjusted according to the requirements of different cities and different periods.
[0032] In this embodiment, the calculation parameters for the above pollutant dynamic threshold, static threshold, minimum threshold, and growth threshold are summarized in the following table:
[0033] S2.2: High-value judgment combination logic settings; A combinatorial judgment process based on Boolean algebra is designed to achieve accurate adaptation to complex high-value scenarios: S2.2.1: Construction of composite rules; By combining a single condition using the logical AND (&) and OR (|) operators, targeted compound rules can be formed: (1) Conventional high value rule: The concentration threshold conditions simultaneously satisfy {dynamic threshold & static threshold & minimum threshold}; (2) Sudden high value rule: The concentration threshold conditions simultaneously satisfy {static threshold & minimum threshold & growth threshold}; (3) Strong high value rule: The concentration threshold conditions simultaneously satisfy {dynamic threshold & static threshold & minimum threshold & growth threshold}; Meeting any of the above rules is sufficient to determine that the corresponding pollutant has a high value.
[0034] S2.2.2: Dynamic expansion rule settings; This step serves as a flexible supplementary module to the multi-dimensional high-value threshold identification system. Its core function is to address the rule adaptation issues caused by differences in pollutant types across different regions and seasonal fluctuations in pollution characteristics. It supports the dynamic addition of custom extended rules based on actual monitoring needs without affecting the stability of the original judgment logic. All rule expressions can be adjusted by combining regional pollution cause classification results and seasonal meteorological change patterns, breaking the limitations of traditional fixed rule bases and improving the rule base's scenario adaptability. This embodiment, based on the high-value characteristics of Jiujiang City, adds the following extended rules: (1) Extended rules for determining high SO2 values: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 20 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 5 μg / m³) 3 (Time period: November to February of the following year | 10 PM to 6 AM the next day) (Meteorological conditions: Wind speed ≤ 2 m / s); (2) Extended rules for judging high NO2 values: (static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 20 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 10 μg / m³) 3 (Time period conditions: morning peak 7-9 am | evening peak 17-19 pm) (Meteorological conditions: calm and stable weather, air pressure ≥1010 hPa); (3) Extended rules for judging high O3 values: (static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 150 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 10 μg / m³) 3 (Time period: 2 PM - 8 PM) (Meteorological conditions: Temperature ≥ 25℃ & Humidity ≤ 60%) (4) Extended rules for judging high CO values: (static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 1 mg / m³) 3 (Growth threshold: proportion ≥ 1 & absolute growth ≥ 0.5 mg / m³) 3 (Time period conditions: Winter heating season | Morning and evening peak hours) (Coordination conditions: PM)2.5 Concentration synchronous increase ≥10μg / m 3 ); (5) PM 10 High-value determination extended rule: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 30 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 20 μg / m³) 3 (Meteorological conditions: wind speed ≥ 3 m / s | within 24 hours after precipitation) (Synchronous conditions: PM) 2.5 / PM 10 ≤0.6); (6) PM 2.5 High-value determination extended rules: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 20 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 10 μg / m³) 3 (Meteorological conditions: wind speed ≤ 1.5 m / s & humidity ≥ 60%) (Time period: 8 PM to 8 AM the next day | calm and stable weather period); The above is a high-value extended regulation set based on the air pollution prevention and control requirements of Jiujiang City in this embodiment. The method of this invention can be flexibly adjusted according to the requirements of different cities and different periods when applied to other regions.
[0035] Step S3: High-value process identification and multimodal feature construction; This step aggregates discrete high-value moments into a complete pollution process, extracts quantitative features from multiple dimensions, provides support for subsequent pollution pattern clustering and pollution cause classification, and solves the problems of lack of unified description and insufficient analytical basis for pollution processes.
[0036] S3.1: Automatic identification and marking of discrete high-value moments; S3.1.1: Evaluation of batch rules for time series data; (1) Parallel processing: for SO2, NO2, O3, CO, PM 10 PM 2.5 The concentration data of six pollutants are used in batches to apply the rule base constructed in step S2 to calculate the satisfaction status of four types of concentration threshold conditions at each time step. (2) Composite rule determination: For each pollutant, determine whether the current moment is a high value moment for the pollutant based on the composite rule expression configured in step S2; (3) Tag aggregation: Aggregate the names of all pollutants that trigger high values at the same time to generate a comprehensive high value tag. If the tag is not empty, mark the time as a high value time.
[0037] S3.1.2: High-value event metadata records; Record complete metadata for each discrete high-value moment: (1) Triggering information: Record the specific pollutant name that triggered the high value and the corresponding rule combination number; (2) Environmental snapshot: Simultaneously save meteorological data (temperature, humidity, wind speed, wind direction, etc.), background average of pollutants (average of the last 24 hours), concentration value of the previous moment, etc., to preserve the complete environmental background for subsequent analysis.
[0038] S3.2: Aggregation of continuous high-value processes; By merging discrete high-value moments that are spatially and temporally close into a complete contamination process, the integrity and accuracy of the process are ensured. S3.2.1: Initial seed generation for high-value processes; Starting from the first discrete high-value moment, an initial contamination process is created: (1) Record the starting point of the process: Set this moment as the initial start time; (2) Record core pollutants: The list of core pollutants during the initialization process consists of high-value labels at that moment; (3) Record peak information: Record the initial peak concentration of each pollutant at this moment and the corresponding peak time.
[0039] S3.2.2: Process merging determination based on pollutant continuity; (1) Common pollutant inspection: Calculate the intersection of the current high value time and the pollutant set in the subsequent process. If there is no intersection, create a new process. (2) High-value process update: If the common pollutant inspection is met, the current time is merged into the subsequent ongoing process, and the process end time and core pollutant list are updated; if the pollutant concentration at the current time exceeds the existing peak value of the process, the peak concentration and peak time are updated. S3.2.3: Intelligent correction and trend marking of process boundaries; (1) Set a backtracking window: For each post-aggregation pollution process, set a 24-hour backtracking window before the initial start time to find the true starting point where the pollutant concentration increases significantly; (2) Sliding window analysis: A sliding window of variable size of 2-6 hours is used to move within the backtracking window. Linear fitting is performed on the concentration data in each window, and the fitting slope is calculated. (3) Determination of actual starting point: When the fitting slope is ≥0.1 (μg / (m 3 •h) ), cumulative increase in concentration ≥5μg / m 3 When the initial concentration is lower than the initial starting time concentration but higher than the background average, the starting time of this window is determined as the actual starting point of the process. (4) High value process trend marking: If the actual starting point is identified, the process is marked as having a significant upward trend; otherwise, the initial starting time is used and marked as "no significant upward trend".
[0040] S3.3: Construction of the feature matrix; For each complete pollution process, quantitative features are extracted from four dimensions: pollution intensity characteristics, temporal evolution characteristics, meteorological correlation characteristics, and specific pollutant ratio characteristics, to construct a multimodal feature matrix: S3.3.1: Pollution intensity characteristics; (1) High concentration of pollutants: Record the maximum concentration of pollutants during the high-value process, including SO2, NO2, and PM2.5. 10 PM 2.5 O3 is the core pollutant, while CO is another pollutant; (2) High value marking characteristics: Mark whether the pollutant is the main pollutant at the time of high value (0 / 1); (3) Characteristics of pollutant importance: the importance of pollutants in high-value processes; (4) Correlation characteristics of pollutants: The correlation coefficient between pollutants was calculated using the Pearson correlation coefficient method.
[0041] S3.3.2: Characteristics of temporal evolution; (1) Growth characteristics: Calculate the average growth rate of the concentration of the six pollutants from the actual starting point to the peak time (peak concentration - initial concentration) / (actual starting point at the peak time); (2) Time characteristics: Duration: The exact duration from the actual start point to the end of the process (unit: hours); Duration type: The process is divided into short-duration processes (<2 hours, marked as 1) and long-duration processes (≥2 hours, marked as 2) according to the duration; High value moment: The moment when the high value event occurs (0-23).
[0042] S3.3.3: Meteorological correlation characteristics; (1) Meteorological characteristics: temperature, humidity, and wind speed at high values; (2) Wind direction characteristics: the dominant wind direction during the statistical process (wind direction with a frequency of ≥50%), and converted into an eight-way code (N=0, NE=1, E=2, SE=3, S=4, SW=5, W=6, NW=7). (3) Pollutant-meteorological correlation characteristics: The correlation coefficient between pollutant concentration and each meteorological factor was calculated using the Pearson correlation coefficient method.
[0043] S3.3.4: Characteristics of ratios of specific pollutants; Calculating the concentration ratio of specific pollutants mainly involves two ratios: (1) PM2.5 / PM 10 Ratio: Characterizes the proportion of fine particulate matter and reflects the source of pollution (a high ratio indicates a bias towards secondary or mobile sources, while a low ratio indicates a bias towards dust sources). (2) SO2 / NO2 ratio: characterizes the relative contribution of coal-fired sources and motor vehicle emissions (a high ratio is biased towards coal-fired sources, and a low ratio is biased towards mobile sources).
[0044] S3.4: Establishment of the pollution pattern feature matrix; Including all the features of the above four categories, constructing such as Figure 6 The pollution pattern feature matrix shown provides quantitative input for subsequent pollution pattern typing.
[0045] Step S4: Intelligent typing of high-value process contamination patterns; This step is one of the core innovative steps of this invention. It uses the self-organizing map network (SOM) unsupervised clustering algorithm to avoid the limitation of traditional clustering requiring a preset number of clusters. It improves the clustering accuracy of high-dimensional data by relying on its topology preservation characteristics, and solves the problem of fractal bias that is easy to occur in high-dimensional data processing by traditional techniques.
[0046] Furthermore, by amplifying key features and configuring differentiated weights, the core information for pollution pattern identification is strengthened, ultimately achieving objective and accurate clustering of pollution patterns. This determination provides a reliable pattern classification basis for subsequent classification and analysis of pollution causes, establishing a key technological link for pollution cause identification. Simultaneously, the trained SOM model can quickly match new high-value events, improving the response efficiency and technical support capabilities of air pollution control. Furthermore, the trained SOM model can achieve rapid matching and classification of new high-value pollution events, significantly shortening the pollution pattern identification cycle and substantially improving the response efficiency of emergency air pollution control.
[0047] S4.1: Feature Matrix Reconstruction and Enhancement; This step aims to improve the accuracy of cluster analysis of high-value pollution processes. It addresses common technical issues in the basic feature matrix, such as key feature overload, dimensional redundancy, and insufficient discriminative power. By setting key features to reconstruct the feature matrix and weighting the core features, it achieves targeted enhancement and optimizes the feature matrix, providing high-quality feature input for subsequent pollution process classification and pollution cause determination.
[0048] S4.1.1: Feature matrix reconstruction; High-value features, core features, pollutant-related features, and growth features are set as key features. The feature matrix structure is reconstructed and used as the input for the subsequent SOM model.
[0049] S4.1.2: Core Feature Enhancement; The core features include the concentration of core pollutants, meteorological correlation features, temporal features, and the ratio of characteristic pollutants; the core features are enhanced through optimization and weighted strengthening.
[0050] S4.1.2.1: Core Feature Optimization Rules: (1) Optimization of meteorological correlation features: The wind direction (1-8) is mapped to eight directions and uniquely encoded to generate eight binary features (such as north wind = 1, others = 0). At the same time, the wind direction angle is transformed into continuous features through sine and cosine transformations to retain the periodic information of wind direction. (2) Time feature optimization: The high-value moments (0-23) are transformed into continuous features through sine and cosine transformations to express the time periodicity (such as the similarity between early morning and night).
[0051] S4.1.2.2: Core Feature Enhancement Rules; (1) Core features include: duration type, duration; hourly sine and cosine; prevailing wind direction, wind direction sine and cosine; PM 2.5 / PM 10 Ratio, SO2 / NO2 ratio; SO2, NO2, PM 10 PM 2.5 The average concentration of O3 during the process; (2) Enhancement of core features: Duration type and duration magnification factor is 1.8; hourly sine and cosine magnification factor is 1.5; dominant wind direction magnification factor is 1.1, and wind direction sine and cosine magnification factor is 1.7; PM 2.5 / PM 10 The amplification factor for the ratio and the SO2 / NO2 ratio is 2.0; SO2, NO2, PM 10 PM 2.5 The peak concentration amplification factor for the high-value O3 process is 1.2.
[0052] S4.2: Intelligent data completion and feature standardization; S4.2.1: Differentiation-based filling strategy; For different types of missing features, a precise imputation method is used: (1) Binary features (such as growth and decline markers): The mode of this feature is used to fill in the gaps and ensure the consistency of the classification features; (2) Wind direction characteristics: Grouped by the month in which the process occurs, and the dominant wind direction of that month is used to fill in the gaps, in accordance with the seasonal wind direction pattern; (3) Duration characteristics: Grouped by high value label type (0 / 1), the median within each group is used to fill in the gaps and preserve the duration distribution characteristics; (4) General numerical characteristics: When the missing rate is <30%, the global median is used for imputation; when the missing rate is >30%, the data is grouped by duration type (short-term / long-term) and imputed by the mean within the group to balance data integrity and accuracy.
[0053] S4.2.2: Feature Standardization and Weight Configuration (1) Numerical feature standardization: For all continuous numerical features except binary features, the Z-score standardization method is used to eliminate dimensional differences; (2) Feature weighting system: Construct a differentiated feature weight vector to strengthen the influence of core features: wind direction feature weight = 30; special pollutant ratio feature (PM 2.5 / PM 10 Weight of ratio (SO2 / NO2 ratio) = 28; weight of time feature = 20; weight of core pollutant concentration feature = 20; weight of other pollutant concentration and meteorological feature = 15; weight of high value marker feature = 10; weight of other features = 1.
[0054] S4.3: SOM clustering model execution; S4.3.1: Adaptive mesh parameter configuration; The SOM network structure is dynamically adjusted based on the number of samples to ensure that the clustering resolution is adapted to the amount of data: Small samples (number of samples < 20): 3×3 grid; Medium samples (20 ≤ number of samples < 50): 4×4 grid; Large samples (number of samples ≥ 50): 7×7 grid; Default configuration: 5×5 grid, which balances versatility and discriminative power, and is used in this embodiment. The SOM model grid parameters are shown in the table below:
[0055] S4.3.2: Network initialization and training optimization; (1) Weight initialization: The SOM network weights are initialized based on the PCA dimensionality reduction results to avoid training fluctuations caused by random initialization and improve convergence stability; (2) Training parameter settings: Learning rate strategy: The initial learning rate is set to 0.08 and updated using an inverse time decay function, so that it decreases adaptively with the increase of training rounds, ensuring that the network quickly approximates the macro structure of the data in the early stage and finely adjusts the micro weights in the later stage to achieve a balance between exploration and convergence; Neighborhood function design: The initial neighborhood radius covers most of the network area. The influence strength of the neuron is defined according to the Gaussian function. Its width parameter sigma is set to 0.4 and decays exponentially with training, so that the weight update range gradually focuses from global collaboration to the local of the winning neuron, accurately depicting the topological features of the data; Training control and convergence criteria: The maximum number of training rounds is set to 2000, and an early stopping mechanism based on the change norm of the weight matrix is provided (convergence threshold is 1e-5). When the network state tends to be stable, the iteration is automatically terminated to avoid overfitting and improve computational efficiency; Core feature enhancement: A core feature enhancement mechanism is constructed. During the weight update stage, a specified gain factor is applied to the learning rate of the specified important features, so that the model can learn and separate the differences in key business dimensions while maintaining the overall structure. Based on the air pollution prevention and control requirements of Jiujiang, this embodiment sets the core feature learning rate gain factor to 10, and the SOM clustering model parameters are set as shown in the table below:
[0056] S4.3.3: Key Feature Enhancement Clustering; In the SOM distance calculation process, the feature weights set in step S4.2.2 are incorporated into the distance formula, and the influence of key features is amplified according to the key feature weight enhancement rule in step S4.1.3, so as to ensure that the clustering results are highly consistent with the business needs of the pollution pattern.
[0057] S4.4: Post-processing and quality assessment of clustering results; S4.4.1: Cluster optimization merging; (1) Small cluster identification: Clusters with less than 5 samples are identified as small clusters and need to be merged and optimized; (2) Merging rules: Calculate the weighted Euclidean distance between the small cluster and other clusters. In this embodiment, the magnification factor in the weighted Euclidean distance calculation of the core features is 8. If the distance is <0.1 and the variance of the key features increases by ≤1.5 times after merging, the small cluster will be merged into the nearest large cluster to ensure the concentration of features within the cluster.
[0058] S4.4.2: Quantitative assessment of cluster quality; Three types of indicators were used to comprehensively evaluate the clustering effect: (1) Cluster density: This measures the concentration of samples within each cluster on key features. It is evaluated by calculating the average variance of key features within the cluster, with a value ≤0.5 being considered excellent. The calculation formula is as follows: For each cluster ( ): Density ( )= ; It is a cluster The number of samples; It is a weighted Euclidean distance, calculated using the following formula: ; To further emphasize the concentration of core features, the magnification factor in the weighted Euclidean distance calculation of core features in this embodiment is 8. Therefore, the formula for calculating the weighted Euclidean distance of core features is: ; (2) Cluster concentration: Calculate the ratio of the intra-cluster variance to the global variance for each key feature. A value <0.7 is considered well-concentrated, indicating that the feature has high consistency within the cluster. The calculation formula is as follows: For each key characteristic (f): Concentration ( )= ; It is the number of clusters; feature In cluster The variance in; yes Variance across the entire dataset; (3) Quantization error: This measures how well the SOM network fits the original data, i.e., the average distance from all samples to the weight vector of their respective best-fit unit (BMU). A value <0.3 is considered excellent. The calculation formula is as follows: Quantization error = ; It is the total number of samples; It is exactly the same as the weighted distance function used in S-cluster compactness.
[0059] In this embodiment, the overall performance results of the SOM model are shown in the table below:
[0060] Furthermore, the concentration results of the core features of the SOM model in this embodiment are shown in the following table:
[0061] S4.5: Summary of clustering results; (1) Pollution process cluster label table: includes the cluster number and cluster quality score of each pollution process; (2) The trained SOM model contains information such as network weights, cluster center features, and quantization error, which are used for matching new high-value events in the future.
[0062] Based on the above steps, the cluster optimization result in this embodiment is as follows: Figure 7 As shown, the clustering results are as follows Figure 8As shown.
[0063] Step S5: Diagnosis of the cause of contamination patterns; This step, as a key follow-up to step S4, takes the pollution pattern clustering results output from step S4 as input and constructs a modular pollution cause classification rule base covering 6 types of core pollution sources, 2 types of special pollution, and 3 types of data anomalies to achieve full coverage of pollution scenarios. Furthermore, the rule base adopts a modular design, which can flexibly add, delete, and adjust the judgment logic of various pollution sources according to the industrial structure, geographical features, and pollution prevention and control priorities of different regions, and has strong regional adaptability.
[0064] S5.1: Construction of a rule base for classifying pollution causes; A modular rule base for judgment was constructed, covering 6 types of core pollution sources, 2 types of special pollution, and 3 types of data anomalies. Each rule clearly defines the quantified threshold and logical conditions. S5.1.1: Rules for determining mobile sources; NO2 is the main pollutant, with an SO2 / NO2 ratio ≤ 0.7, and meeting any of the following conditions: the process occurs during the morning peak (7-9 am) or evening peak (17-19 pm); the correlation coefficient between NO2 and CO is ≥ 0.6; and the CO growth rate is > 0.8.
[0065] S5.1.2: Rules for determining the source of fireworks displays; (1) Time period determination: The process occurs between January 1st and 10th (Spring Festival) or January 15th (Lantern Festival), and the proportion of cases in this period is >0.4% of the total cases; (2) SO2-dominant type: SO2 is the main pollutant, the SO2 / NO2 ratio is ≥0.8, and the time period is determined. At the same time, the meteorological conditions are humidity ≥60% and wind speed ≤2m / s. (3) PM 2.5 Dominant type: PM 2.5 The main pollutant is SO2, with an SO2 / NO2 ratio ≥ 0.8. SO2 is a high-growth pollutant, and the time period and the above meteorological conditions are met.
[0066] S5.1.3: Rules for determining dust sources; (1) PM 10 Dominant type: PM 10 It is the primary pollutant and meets all of the following conditions: PM 2.5 / PM 10 Ratio ≤ 0.7; PM 10 Concentration ≥115μg / m 3 Wind speed ≥ 3m / s; (2) Meteorological conditions: humidity is humid or high humidity; wind speed is light or calm.
[0067] S5.1.4: Rules for determining coal sources; (1) SO2 is the main pollutant, and the SO2 / NO2 ratio is ≥0.75 and the SO2 concentration is ≥15 μg / m³. 3 Or SO2, PM 2.5 Both are major pollutants and their correlation coefficient is ≥0.6; (2) Auxiliary conditions for determining coal combustion sources in winter: The process occurs in December or January; SO2 and PM 2.5 Highly correlated; SO2 concentration ≥ 20 μg / m³ 3 .
[0068] S5.1.5: Rules for determining industrial sources; NO2 is the primary pollutant, the SO2 / NO2 ratio is <0.7, and any of the following conditions are met: it is not during morning or evening peak hours; there is no high correlation between NO2 and CO (<0.6); CO is not a high-growth pollutant; PM2.5 is the primary pollutant. 2.5 As the main pollutant, PM 2.5 It is a high-growth pollutant.
[0069] S5.1.6: Rules for determining secondary sources; (1) PM 2.5 Dominant type: PM 2.5 PM2.5 is the main pollutant. 2.5 / PM 10 The ratio is ≤1.2 and ≥0.7, and the humidity is ≥60%, while the wind speed is ≤1m / s; (2) Auxiliary conditions: O3 is a high-growth pollutant; the process occurs at night or in the early morning; PM 2.5 Correlation with O3 > 0.5.
[0070] S5.1.7: Rules for determining special pollution sources; (1) Biomass combustion source: The process occurs from March to May or from September to November (crop harvest burning season), and at the same time CO or PM 2.5 It must be a primary pollutant. And it must meet the following supplementary conditions: PM2.5 2.5 / PM 10 Ratio > 0.4, CO and PM 2.5 High correlation; ②PM 10 Not a major pollutant; (2) Sources of cooking fumes: The process occurs at noon (11:00-13:00) or in the evening (18:00-20:00), PM 2.5 The main pollutants are SO2 and NO2, which are not high-growth pollutants, and there is no SO2-PM. 2.5 NO2-PM 2.5 High correlation.
[0071] It should be noted that in this embodiment, the above parameters are based on the requirements for preventing and controlling air pollution in Jiujiang City, and special pollution sources are set. When the method of the present invention is applied to other regions, the special pollution sources can be modified and supplemented based on the prevention and control requirements of different regions.
[0072] S5.1.8: Data anomaly determination rules; (1) Equipment failure: It is determined as equipment failure if any of the following conditions is met: ① CO is the main pollutant, CO > 4 mg / m 3 and the average growth rate > 2; ② SO2 is the main pollutant, SO2 > 300 μg / m 3 and the average growth rate > 2; ③ NO2 is the main pollutant, NO2 > 300 μg / m 3 and the average growth rate > 2; ④ PM 10 is the main pollutant and PM 10 > 800 μg / m 3 ; ⑤ PM 2.5 is the main pollutant and PM 2.5 > 800 μg / m 3 ; (2) PM 2.5 inversion: PM 2.5 is the main pollutant and the ratio of PM 2.5 / PM 10 > 1.25, which is determined as PM 2.5 inversion; (3) Ozone pollution (meeting any of the following conditions): O3 is the main pollutant, and 150 < O3 concentration ≤ 160 μg / m 3 (ozone pollution); O3 concentration > 160 μg / m 3 (ozone is on the high side).
[0073] In this embodiment, the above parameters are based on the requirements for preventing and controlling air pollution in Jiujiang City, and data anomaly determination rules are set. When the method of the present invention is applied to other regions, the data anomaly determination rules can be modified and supplemented based on the prevention and control requirements of different regions.
[0074] S5.2: Setting the priority of cause determination; Set four levels of determination priorities and determine in the order of priorities: (1) The 1st priority: Data anomaly determination (such as equipment failure, PM 2.5 inversion, ozone pollution), if triggered, directly output the result and terminate the subsequent determination; (2) The 2nd priority: Special pollution determination (such as biomass combustion source, catering oil fume source); (3) Third priority: core pollution source determination (dust source, fireworks source, mobile source, industrial source, coal source, secondary source), supporting multiple sources coexisting (if dust source and secondary source are triggered at the same time, both are recorded); (4) Fourth priority: fallback processing. If no rule is triggered, mark it as "unrecognized" and record the event characteristics for subsequent rule optimization.
[0075] S5.3: Summary of pollution cause classification results; (1) Pollution source determination result table: includes the pollution source label (single / multiple pollution sources) and the specific rule number that triggered each pollution process; (2) Pollution cause analysis report: detail the key basis for determining the pollution source (such as SO2 / NO2=0.9≥0.8, which meets the coal combustion source rule) and the reasons for excluding other pollution sources; (3) Abnormal event list: includes equipment failure and unidentified events, for manual review and rule optimization.
[0076] S5.4: Construct a historical case library of pollution causes; like Figure 9 As shown, the historical pollution cause case library is based on previous feature data, clustering results, and a rule base determined by expert experience to identify pollution causes. This facilitates the subsequent XGBoost classification model's classification of pollution causes based on the historical pollution cause case library constructed in this step, thereby achieving intelligent diagnosis of pollution causes at new high-value times. The historical pollution cause case data includes: (1) Feature data: Extraction of feature matrix of high-value process in historical data; (2) Clustering data: Cluster label column after SOM model clustering; (3) Causes of pollution: The list of causes of pollution corresponding to the high-value processes in historical data.
[0077] Step S6: Intelligent diagnosis of new high-value events; This step utilizes the XGBoost model to perform fully automated diagnosis of newly added high-value events, quickly outputting the pollution cause determination results. Specifically, it includes the following steps: S6.1: Input and preprocessing of monitoring data from new national monitoring stations; Data integrity assurance: for the monitoring data of the aforementioned sites; (1) Acquisition of monitoring data from national control stations: In this embodiment, the station-level data and corresponding city-level data of national control stations under the jurisdiction of Jiujiang City (including Xiyuan, Shili, Maoshantou, Petrochemical Plant, Foreign Language School, No. 8 Lianhua Avenue, and Comprehensive Industrial Park) are acquired; the time range is from January 1, 2025 to November 20, 2025, and the time resolution is 1 hour for national control stations; the monitoring data includes station code, monitoring time (accurate to the minute), concentration of six pollutants, and six meteorological parameters; (2) Standardized preprocessing of monitoring data from national control stations: Based on the preprocessing pipeline in step S1, complete missing data repair and invalid value filtering (including negative numbers, invalid marker values, etc.), and output preprocessed time series data to ensure that the data format is consistent with the training data.
[0078] S6.2: High-value process identification and feature extraction; where: (1) Discrete high value identification: Based on the high value identification rule base in step S2, the preprocessed data is scanned time by time to identify discrete high value moments; (2) Aggregation of continuous high-value processes: According to the logic of step S3.2, aggregate discrete high-value moments into complete pollution processes, correct process boundaries and mark trends; (3) Construction of high-value process feature matrix: According to the multimodal feature system in step S3.3, extract the pollution intensity features, time evolution features, meteorological correlation features, and special pollutant ratio features of the event to generate a feature matrix; S6.3: Automated diagnosis of the causes of new high-value events caused by pollution patterns; This step uses a classification model to automatically diagnose the causes of pollution patterns in new high-value events. The XGBoost model is selected, which leverages its strong predictive classification capabilities to automatically classify the causes of pollution in new high-value events based on a historical pollution cause case library.
[0079] S6.3.1: Training optimization; (1) Initialization parameter configuration: ① Objective function: multi:softprob (multi-class probability output); ② Maximum tree depth: 6; ③ Learning rate: 0.1; ④ Subsample ratio: 0.8; ⑤ Feature sampling ratio: 0.8; ⑥ Random seed: 42 (to ensure reproducibility of results); ⑦ Number of trees: 100; ⑧ Incremental training threshold: default 5, in this embodiment, it is set to ≥20 feedback data points according to the requirements of Jiujiang air pollution prevention and control, and incremental retraining begins; ⑨ Full retraining threshold: default 5 times, in this embodiment, it is set to start full retraining after ≥30 incremental training times; (2) Model initialization training: The model is initially trained using historical pollution cause case data. This includes the following: ① Training data sources: Historical pollution cause and pollution case data files; monitoring data from the Jiujiang City National Monitoring Station (six pollutants and six meteorological parameters). ② Feature processing: Missing values in numerical feature columns are filled with 0 to ensure feature consistency; ③ Tag Encoding: Converting text-based pollution cause tags into numerical codes; ④ Model saving: After training is completed, the model, feature names, label encoder, historical data and other components are serialized and saved, supporting direct loading and use later; ⑤ Historical data maintenance: During initial training, the complete dataset is saved as a historical pollution cause case library, which serves as the basis for subsequent incremental learning and full retraining; S6.3.2: Intelligent diagnosis of the causes of pollution in new high-value events; This step utilizes the XGBoost contamination pattern classification model to achieve a rapid response based on data-driven cause determination. S6.3.2.1 Input of new high value event data; The feature matrix of this high-value process is extracted, including: (1) Numerical feature data: Select numerical columns as input model features; (2) Feature consistency processing: Ensure that the feature dimensions of the new data are consistent with those during training, and fill missing features with 0; (3) Data preprocessing: Fill in missing values to ensure data integrity; S6.3.2.2 XGBoost model execution; The model execution process includes: (1) Feature matching: Align the input data with the feature names used during model training to ensure dimensionality consistency; (2) Prediction execution: Use the trained XGBoost model to predict the classification of pollution causes; (3) Probability output: Provides the probability distribution of each pollution cause to support uncertainty assessment; (4) Label decoding: Convert the numerical prediction results back into the original text label form of pollution causes.
[0080] S6.3.2.3 Post-processing and quality assessment of classification results; (1) Result analysis: The numerical classification results output by the model are mapped into readable descriptions of pollution causes; (2) Confidence assessment: assessing the reliability of the prediction probability results; (3) Historical records: The prediction results and related feature data are recorded in the historical cause case library for continuous model optimization.
[0081] S6.4: Classification and determination of pollution pattern causes; Using the above method, the pollution cause classification results of the XGBoost model for new high-value events are as follows: Figure 13 As shown, the intelligent diagnostic results for new high-value events are as follows: Figure 14 As shown.
[0082] Step S7: Model self-learning and continuous optimization; This step is the core of the long-term guarantee of the technical solution system of this invention. Based on expert feedback data from the intelligent diagnosis of new high-value events, a full-process model self-optimization mechanism is constructed, encompassing expert feedback, model updates, effect evaluation, and version management, to achieve dynamic iteration of the XGBoost pollution cause classification model. This mechanism drives continuous optimization of the model's core performance indicators through feedback from domain experts, breaking through the technical bottlenecks of traditional static models where parameters are fixed in a single training iteration and the model's generalization ability drifts and decays with data distribution. This provides support for the long-term reliable operation of the pollution source tracing system.
[0083] S7.1: Expert feedback data collection and standardization; S7.1.1: Feedback data source and type; (1) Expert manual review and feedback: Experts manually review the pollution cause results output in step S6 and correct the pollution cause classification results; (2) Experts actively label new high-value cases to supplement the training samples; (3) Expert suggestions for rule modification: Based on their business experience, experts propose suggestions for threshold adjustment, rule addition and modification.
[0084] S7.1.2: Feedback data screening and validity determination; (1) Feedback data screening: Based on the site code, process time and core pollutant combination, records that are duplicates of historical feedback data are removed and the feedback data is screened out; (2) Labeling quality verification: Verify that the classification results of the expert labels are in the existing cluster list (or are new labels) and that the pollution source labels conform to the rule base definition (or are new types). If the labeling is not standardized, it will be returned for supplementation.
[0085] S7.1.3: Standardized processing of feedback data; (1) Feature data alignment: Complete all quantitative features of the event data annotated by experts according to the feature system in step S3.3 to ensure that it is consistent with the training data format; (2) Label standardization: unify the naming conventions for type labels and pollution source labels, and name the new labels according to the combination of type + serial number; (3) Data format conversion: Convert the standardized feedback data into the feature matrix and label vector required for model training. The dimension of the feature matrix is the same as that in step S4.
[0086] S7.2: Incremental learning triggering and execution; S7.2.1: Incremental learning triggering conditions; (1) Data volume trigger: When the amount of newly added valid data is greater than or equal to the set threshold (the set threshold in this embodiment is 20) and does not involve new tags, incremental learning is automatically triggered; (2) Time trigger: If the data volume threshold is not met, but the amount of newly added valid data is greater than or equal to the time trigger threshold (the threshold is set to 10 in this embodiment), incremental learning will be automatically triggered every 180 days to avoid long-term lack of updates.
[0087] S7.2.2 Incremental learning execution process; (1) Model loading: Load the latest version of the XGBoost classification model (including feature weights, class weights, and historical training data). (2) Feature weight adaptation: Use the feature weights of the trained XGBoost classification model; if experts suggest adjusting the feature weights, update the feature weight vector; (3) Incremental training parameters: The trained XGBoost classification model is used, but the number of decision trees is adjusted. In this embodiment, it is set to increase by 10 (only the model parameters are updated, and the model structure is not changed).
[0088] S7.2.3: Evaluation of incremental training results; The performance of the updated model was evaluated using a test set (20% of historical data). The core metrics included: ① Classification accuracy: the number of correctly classified events / the total number of events, which must be ≥ the original accuracy + 2%; ② Contamination cause determination accuracy: the number of events with the correct contamination cause / the total number of events, which must be ≥ the original accuracy + 2%; ③ New sample fit: the prediction accuracy of the newly added feedback samples ≥ 90%.
[0089] S7.3: Full retraining mechanism and execution; S7.3.1: Full retraining trigger condition; (1) Cumulative triggering of incremental learning: The cumulative execution of incremental learning reaches a set threshold (the threshold is set to 30 times in this embodiment), which automatically triggers full retraining to ensure the overall stability of the model; (2) New label trigger: After a new label is detected and the expert completes the rule base update in the format of step S5.1, the number of new labels in the XGBoost model is automatically adjusted, triggering full retraining; (3) Time trigger: If the trigger threshold is not met, but the amount of newly added valid data is greater than or equal to the time trigger threshold (the threshold is set to 15 times in this embodiment), a full retraining will be automatically triggered every 360 days to avoid long-term lack of updates.
[0090] S7.3.2: Preparation of Full Retraining Data (1) Data integration: Integrate historical training data (feature matrix and labels from steps S3-S4) with all valid feedback samples, and remove duplicate data; (2) Data partitioning: Divide the data into training set, validation set and test set in a ratio of 7:2:1 to ensure that the data is evenly distributed.
[0091] S7.3.3: Full retraining execution process; (1) Feature engineering reconstruction: Re-execute steps S4.1-S4.2 to reconstruct, fill, and standardize the feature matrix to ensure feature quality; (2) Initialization of XGBoost model retraining parameters: (3) XGBoost model retraining: Retrain the XGBoost pollution cause classification model according to step S6. (4) Rule base synchronous update: If the threshold of some rules is found to be unreasonable during the retraining process (such as the trigger rate of a certain rule <1%), the system will automatically propose adjustment suggestions, which will be confirmed by experts and the rule base will be updated.
[0092] S7.3.4: Quality assessment of retrained models; The performance of the retrained model is comprehensively evaluated using multiple dimensions of indicators, and all indicators must meet the following requirements: (1) Pollution cause assessment: The pollution cause classification accuracy of the XGBoost model is ≥85%; the pollution cause classification accuracy of the XGBoost model after retraining is greater than that of the original model. (2) Generalization ability assessment: The accuracy of the XGBoost model in classifying pollution causes decreased by ≤5% compared with the training set, ensuring that the model does not overfit; (3) Evaluation report: Generate a model retraining evaluation report, including comparison of various indicators, sample distribution, and error case analysis, for expert review.
[0093] S7.4: Model version management and deployment; S7.4.1: Version Management Rules; (1) Version number: Use the format "V[full retraining times].[incremental learning times]", such as V2.1 (2 full retraining times, 1 incremental learning time); (2) Version storage: Each version of the model must store complete files, including XGBoost model parameters, feature weights, training data, and evaluation reports, with a storage period of ≥3 years; (3) Version rollback: Supports rollback to any historical version. When the performance of the new version model degrades (e.g., accuracy drops by ≥5%), it can quickly switch to the best historical version.
[0094] S7.4.2: Model Deployment and Switching; (1) Deployment method: Supports offline deployment (local server) and cloud deployment (cloud server), and the model response delay after deployment is ≤3 seconds; (2) Gray release: After the new version model is deployed, the new version is used to diagnose 10% of new events, while the old version is used for the rest, and monitoring continues for 24 hours. (3) Full release switch: If the accuracy of the new version’s classification of subtypes / contamination causes during the gray release period is ≥ the old version + 3%, then the full release will be automatically switched; if the accuracy decreases, the rollback will be performed, the problem will be recorded and optimized.
[0095] S7.4.3: Optimization effect monitoring and feedback; (1) Real-time monitoring: After deployment, the core indicators of the real-time monitoring model (accuracy of pollution cause classification, confidence distribution, generalization ability) are monitored once per hour; (2) Abnormal alarm: When the indicator fails to meet the requirements for 3 consecutive hours, an alarm will be automatically triggered and the administrator will be notified; (3) Regular evaluation: A model optimization effect report is generated monthly to compare the performance changes of different versions, the adoption of expert feedback, and the effect of business applications (such as early warning accuracy) to provide direction for subsequent optimization.
[0096] To efficiently implement the high-value pollution diagnosis method for national monitoring stations based on multimodal clustering provided in this invention, see [link to relevant documentation]. Figure 16 This invention provides a high-value pollution diagnosis system for national monitoring stations based on multimodal clustering. Specifically, it includes a data acquisition layer module, a data preprocessing module, a high-value identification module, a feature engineering module, a pollution pattern clustering and classification module, a pollution pattern cause diagnosis module, a new high-value intelligent diagnosis module, a feedback and learning module, and a display and scheduling layer module. This system can quickly review data from national monitoring stations, automatically identify high-value pollution processes, and achieve rapid response in intelligent diagnosis of pollution causes. Through continuous self-optimization and adaptation to multi-regional characteristics, it provides precise decision support for air pollution prevention and control and emergency response.
[0097] 1: Data acquisition layer module; The data acquisition layer module is used to realize the automated, incremental acquisition and storage of multi-source air quality and meteorological data. It is characterized by including a data source interface adaptation unit, an acquisition scheduling unit, a raw data storage unit, and an acquisition integrity verification unit. Each unit is developed based on the Python language and relies on toolkits such as Celery, Redis, InfluxDB-Python, and SQLAlchemy. The specific implementation is as follows: 1.1 Data source interface adaptation unit; 1.1.1 National-level monitoring station data interface encapsulation subunit; Based on the HTTP / HTTPS protocol communication principle, a standardized API request encapsulation logic is constructed, and a request parameter template containing site code, hourly timestamp range, and parameter type enumeration (six pollutants and six meteorological parameters) is solidified. Relying on the requests library, the open API interface of the ecological environment monitoring platform is called to obtain hourly monitoring data of national control stations in the target area in batches. A 30-second interface timeout threshold is configured, and a fault tolerance mechanism with a maximum of 3 retries and a 5-second retry interval is set to ensure the reliability and integrity of data collection.
[0098] 1.1.2 City-level data synchronization interface subunit; Based on the paramiko library, FTP / SFTP protocol client communication logic is implemented to build a daily incremental synchronization mechanism for a city-level air quality summary database. The synchronized data covers SO2, NO2, O3, CO, and PM2.5. 10 PM 2.5 The system includes statistical values for six pollutant concentrations and six meteorological parameters: temperature, humidity, air pressure, wind speed, wind direction, and precipitation. A synchronized log recording mechanism is designed, including filename, data volume, synchronization status (success / failure), and error messages. When synchronization fails, an email alert is triggered via the smtplib library, synchronizing the data to a designated administrator's email address.
[0099] 1.2 Data Acquisition and Scheduling Unit; 1.2.1 Scheduled data collection task configuration subunit; Based on the Celery distributed task scheduling framework, a scheduled data collection task scheduling logic was designed. The cron standard format for hourly data collection tasks is configured as "0 * * * *", which triggers real-time data retrieval at the site level. The cron expression for daily data completion tasks is configured as "0 1 * * *", which completes city-level data completion for the previous day. The task execution status (running / successful / failed) is stored in a Redis cache using the redis-py library, providing a real-time query interface for the system monitoring module.
[0100] 1.2.2 Incremental Acquisition Strategy Implementation Sub-unit; A timestamp indexing mechanism is constructed. Before data collection, the latest data timestamp stored in the MySQL metadata table is called, and incremental time period data is only requested from the data source to improve collection efficiency. For network interruption scenarios, breakpoint resume logic is designed. The received data fragments are cached through temporary files. After the connection is restored, the integrity of the fragments is verified based on the MD5 checksum calculated by the hashlib library, and the data splicing and integration are completed.
[0101] 1.3 Raw data storage unit; 1.3.1 Time-series database storage sub-unit; InfluxDB time-series database is used as the original data storage medium. Data writing code is written through the InfluxDB-Python client. The data model is designed as follows: the measurement field is defined as "sation_monitor", the tag field includes site code, parameter code, and parameter type, the field field includes monitoring value (data type is float64) and data quality flag (0=valid, 1=suspicious, 2=invalid), and the timestamp field is accurate to the millisecond level; the data retention strategy is configured, the data retention period is set to 3 years, and an automatic sharding storage mechanism is adopted on a daily basis.
[0102] 1.3.2 Metadata Record Subunit; A metadata table is created in the MySQL database. The metadata writing logic is designed based on the SQLAlchemyORM framework. The metadata table contains fields such as a unique data identifier generated by UUID, a data source identifier (API / FTP), collection time, interface version, data byte count, and data verification code. A storage mechanism is implemented to associate metadata with raw data to ensure that metadata updates are executed synchronously with the collection task and to guarantee data traceability.
[0103] 1.4 Data Acquisition Integrity Verification Unit; 1.4.1 Data Count Verification Subunit; Design a theoretical data count calculation model based on the formula: number of stations × total number of parameters × time span (hours) (total number of parameters = 6 pollutants + 6 meteorological parameters = 12) to calculate the theoretical data count; construct a deviation rate judgment logic based on the formula: |actual number of data entries - theoretical number of data entries| / theoretical number of data entries to calculate the deviation rate. When the deviation rate > 5%, mark the data status as "data missing" and write the missing time period and missing station information into the anomaly log table.
[0104] 1.4.2 Data format verification subunit; Based on the principles of regular expressions and data type judgment, a data format validation logic is constructed to verify the completeness of required fields (site code, parameter code, monitoring value, timestamp), the type validity of numeric parameters (must be int / float), and the validity of timestamps (valid Unix timestamps). Data that fails validation is written to an abnormal data temporary table through code, which includes the original data, error type (missing field, incorrect type, invalid timestamp), and collection time, for manual review.
[0105] 2: Data preprocessing module; The data preprocessing module is used to clean, repair, and filter anomalies in multi-source data, constructing a configurable, automated, and traceable preprocessing pipeline. Its features include a preprocessing pipeline initialization unit, a missing data processing unit, and an invalid value detection and filtering unit. It primarily relies on toolkits such as pandas, numpy, scipy, and scikit-learn. The specific implementation is as follows: 2.1 Preprocessing initialization unit; 2.1.1 Loading subunits using configuration files; The JSON configuration file parsing logic is designed based on the json library, and the core preprocessing parameters are loaded: the missing file repair threshold is set to 2 hours, the Z-score anomaly detection threshold is set to 3, and the isolated forest anomaly score threshold is set to 0.8. The code implements a hot update mechanism for the configuration file, which loads the latest configuration without restarting the service by modifying the timestamp of the configuration file.
[0106] 2.1.2 Parameter initialization subunit; A preprocessed data cache pool is built based on Redis, with a batch size of 1000 records. Core algorithm instances are pre-initialized, including a linear interpolation function based on scipy.interpolate.interp1d, a forward fill function based on pandas.DataFrame.fillna, and an isolated forest model (100 decision trees and a maximum sample size of 256) pre-trained based on the IsolationForest class in the scikit-learn library, to reduce the time consumption caused by repeated initialization.
[0107] 2.2 Missing data processing unit; 2.2.1 Short-term missing repair subunit; Based on the principle of parameter type-differentiated repair, an adaptive repair logic is constructed, with input data sequence and parameter type (pollutant parameters and meteorological parameters). (1) When the parameter type is pollutant, the missing value repair is achieved by using the linear interpolation method (interp1d linear interpolation); (2) When the parameter type is a meteorological parameter such as wind direction or wind speed, the forward filling method is used to achieve the repair; (3) After repair, the data is marked with a repair identifier and stored in the preprocessing intermediate table to ensure that the data processing is traceable.
[0108] 2.2.2 Long-term missing sub-unit removal; Design a sliding window missing data detection logic with a window size of 1 hour, traverse the data sequence to calculate the duration of consecutive missing data; when the duration of consecutive missing data is greater than 2 hours, execute an SQL delete statement to remove all parameter data for that period; construct a high-value process association processing mechanism, by associating with high-value pre-marked fields, if an associated high-value process is detected, synchronously update the status of the corresponding record in the high-value process table to "invalid".
[0109] 2.3 Invalid value detection and filtering unit; A unified invalid value replacement mechanism is designed to replace detected invalid values (including negative numbers and invalid marker values) with NaN missing value identifiers, automatically triggering the missing data processing logic of unit S2.2 in this module; a preprocessing log recording mechanism is constructed to store the original invalid value, detection method (Z-score and isolated forest), replacement time, and processing result in the preprocessing log table to ensure traceability of the entire data processing process; 2.4 Data Output Unit; Design the data reorganization logic of site code, time, and preprocessed data dimensions. Based on the PyArrow library, convert the reorganized data into Parquet files in snappy compressed format (compression ratio 3:1) and store them in the HDFS distributed file system. Build a data status update mechanism to mark the preprocessed data as "preprocessing complete" and synchronize it to the data status table to provide data call interfaces for subsequent modules. 3: High-value identification module; The high-value recognition module is used to build an adaptive and configurable high-value recognition rule base to achieve accurate recognition of high-value moments. It is characterized by including a rule base management unit, a threshold calculation unit, a high-value judgment engine unit, and a rule priority execution unit. Its core dependencies include toolkits such as pandas, numpy, and Celery. The specific implementation is as follows: 3.1 Rule base management unit; 3.1.1 Rule configuration file parsing subunit; Based on the xml.etree.ElementTree library, we designed the XML rule configuration file parsing logic, extracted the quantitative parameters of triggering conditions, persistence conditions, and termination conditions, and converted them into Python dictionary-structured rule objects containing a unique rule identifier, rule type (including strong high value, sudden high value, and regular high value), combination of triggering conditions, persistence conditions, and termination conditions. We support dynamic adjustment of rules through adding, deleting, and modifying XML nodes to adapt to the high value recognition needs in different scenarios.
[0110] 3.1.2 Registering sub-units using custom rules; A RESTful API interface (POST / api / rule / register) is built based on the Flask framework. The interface processing logic supports user-uploaded custom rules in JSON format, including fields such as rule name, Boolean logical expression, target pollutant, and priority. A rule syntax validation mechanism is constructed to verify the legality of logical operators (&, |, and NOT) and the rationality of threshold parameters. After the validation is passed, the rule is written to the rule base table and assigned a unique rule identifier.
[0111] 3.2 Threshold rule setting unit; 3.2.1 Dynamic threshold setting subunit; The dynamic threshold calculation logic is designed based on the pandas library. For each site and each pollutant, the dynamic threshold is calculated as: (average concentration of all sites at the same time point) × dynamic threshold factor. Based on the air pollution prevention and control requirements of Jiujiang City, the threshold dynamic multiple of each pollutant is uniformly set to 1.2 to reflect the relatively high value characteristics of pollutants based on the pollution level in Jiujiang City; the specific parameters of this invention can be flexibly adjusted according to the requirements of different regions and cities and different periods.
[0112] 3.2.2 Static threshold setting subunit; The static threshold calculation logic is designed based on the pandas library: Static threshold = average concentration of all stations at the same time point + threshold increment; This embodiment is based on the air pollution prevention and control requirements of Jiujiang City, and the threshold increment is set as follows: SO2: threshold increment = 5 μg / m³ 3 NO2: Threshold increment = 10 μg / m 3 O3: Threshold increment = 10 μg / m 3 CO: Threshold increment = 0.5 mg / m³ 3 PM 10 Threshold increment = 20 μg / m 3 PM 2.5 Threshold increment = 10 μg / m 3 The specific parameters of this invention can be flexibly adjusted to suit the requirements of different regions, cities, and periods.
[0113] 3.2.3 The minimum threshold setting sub-unit long threshold verification should be flexibly adjusted according to the city's and different period's policy requirements; To avoid misjudgment of low concentration fluctuations, a minimum trigger concentration (minimum threshold) is set based on the background concentration characteristics of each pollutant. Based on the air pollution prevention and control requirements of Jiujiang City, this embodiment sets the minimum threshold for six pollutants as follows: SO2 is 20 μg / m³. 3 NO2 was 20 μg / m3 O3 was 150 μg / m 3 CO is 1 mg / m³ 3 PM 10 30 μg / m 3 PM 2.5 20 μg / m 3 The specific parameters of this invention can be flexibly adjusted to suit the requirements of different regions, cities, and periods.
[0114] 3.2.4 Growth threshold setting subunit; The growth threshold is the threshold for determining the growth rate, where the growth rate = current concentration / previous concentration; This embodiment is based on the air pollution prevention and control requirements of Jiujiang City, and sets the growth threshold at 1.2. The specific parameters of this invention can be flexibly adjusted according to the requirements of different regions and cities and different periods.
[0115] 3.3 High-value determination engine unit; 3.3.1 High-value determination basic triggering condition sub-unit; Design a multi-combination basic high-value determination trigger logic: (1) Conventional high value judgment trigger combination: the concentration threshold conditions simultaneously satisfy {dynamic threshold & static threshold & minimum threshold}; (2) Trigger combination for sudden high value determination: The concentration threshold conditions simultaneously satisfy {static threshold & minimum threshold & growth threshold}; (3) Trigger combination for high value determination: The concentration threshold conditions are simultaneously satisfied {dynamic threshold & static threshold & minimum threshold & growth threshold}.
[0116] 3.3.2 High-value determination dynamic expansion rule triggering condition sub-unit; (1) Triggering conditions for the extended rule for determining high SO2 values: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 20 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 5 μg / m³) 3 (Time period: November to February of the following year | 10 PM to 6 AM the next day) (Meteorological conditions: Wind speed ≤ 2 m / s); (2) Triggering conditions for the extended rule for determining high NO2 values: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 20 μg / m 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 10 μg / m³) 3 (Time period conditions: morning peak 7-9 am | evening peak 17-19 pm) (Meteorological conditions: calm and stable weather, air pressure ≥1010 hPa); (3) Triggering conditions for the extended rule for determining high O3 values: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 150 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 10 μg / m³) 3 (Time period: 2 PM - 8 PM) (Meteorological conditions: Temperature ≥ 25℃ & Humidity ≤ 60%) (4) Triggering conditions for the extended rule for judging high CO values: (static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 1 mg / m³) 3 (Growth threshold: proportion ≥ 1 & absolute growth ≥ 0.5 mg / m³) 3 (Time period conditions: Winter heating season | Morning and evening peak hours) (Coordination conditions: PM) 2.5 Concentration synchronous increase ≥10μg / m 3 ); (5) PM 10 High-value determination extended rule triggering conditions: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 30 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 20 μg / m³) 3 (Meteorological conditions: wind speed ≥ 3 m / s | within 24 hours after precipitation) (Synchronous conditions: PM) 2.5 / PM 10 ≤0.6); (6) PM 2.5 High-value determination extended rule triggering conditions: (Static threshold ≥ average concentration × 1.2 & minimum threshold ≥ 20 μg / m³) 3 (Growth threshold: ratio ≥ 1.5 & absolute growth ≥ 10 μg / m³) 3 (Meteorological conditions: wind speed ≤ 1.5m / s & humidity ≥ 60%) (Time period: 20:00-08:00 the next day | calm and stable weather period).
[0117] The above is an extended regulation for high-value determination based on the air pollution prevention and control requirements of Jiujiang City in this embodiment. The specific parameters of this invention can be flexibly adjusted according to the requirements of different regions and cities and different periods.
[0118] 3.3.3 High-value tag generation subunit; Meeting either rule 3.1 or 3.2 will trigger the determination of high values of the corresponding pollutants. The Boolean expression of the rule is parsed by the expression parsing function, and the matching result is stored in the trigger identifier field (1=triggered / 0=not triggered), so as to achieve accurate determination of high value triggering. Construct a high-value label generation logic to generate the following for time periods that meet the trigger conditions: ① High-value process pollutant label: "{pollutant}_high" format label (1 = is the main pollutant at the high-value time / 0 = is not); ② Association rule identifier: such as "dynamic threshold & static threshold & minimum threshold"; store in the high-value event table.
[0119] 3.4 High-value process aggregation unit; 3.4.1 Continuous high-value polymer subunit; Constructing discrete high-value aggregation logic based on common pollutants: (1) Create the initial pollution process: Starting from the first discrete high value moment, create the initial pollution process: ① Initial start time; ② Core pollutant; ③ Peak information: peak concentration and peak time; (2) Common pollutant inspection: Calculate the intersection of the current high value time and the pollutant set in the subsequent process. If there is no intersection, create a new process. (3) High-value process update: If the common pollutant inspection is met, the current time is merged into the subsequent ongoing process, and the process end time and core pollutant list are updated; if the pollutant concentration at the current time exceeds the existing peak value of the process, the peak concentration and peak time are updated.
[0120] 3.4.2 High-value process boundary correction sub-unit; The design process involves boundary correction logic, setting a backtracking window 24 hours before the initial start time, and using a variable sliding window of 2-6 hours to traverse the backtracking data; linear fitting is performed using scipy.stats.polyfit (fit times=1) to calculate the slope, and when the slope is ≥0.1μg / (m 3 (h) and the cumulative increase is ≥5 μg / m 3 When updating the window start time to the actual start time, record the start time before and after the correction, and mark the trend indicator (1 = obvious upward trend, 0 = none).
[0121] 3.4.3 High-value process metadata encapsulation subunit; The process metadata encapsulation logic is constructed, which encapsulates a metadata structure containing process identifier, actual start time, end time, core pollutant list, peak concentration, peak time, and trigger rule identifier; it is stored in the process metadata table in JSON format for subsequent step S4 feature matrix construction.
[0122] 4: Feature Engineering Module; The feature engineering module extracts the quantitative features of the identified high-value processes, providing input data for subsequent contamination pattern clustering and typing. It is characterized by including a high-value process multimodal feature extraction unit, a feature optimization processing unit, and a feature matrix generation unit. Its core dependencies include toolkits such as pandas, numpy, scipy, and scikit-learn. The specific implementation is as follows: 4.1 Multimodal feature extraction unit; 4.1.1 Pollution Intensity Feature Extraction Subunit; Logic for extracting pollution intensity features based on the NumPy library: (1) High-value concentration of pollutants: peak concentration of pollutants during the high-value extraction process; (2) High-value label: pollutant label (0 / 1) during the high-value extraction process; (3) Pollutant correlation characteristics: the correlation coefficient between pollutants was calculated using the Pearson correlation coefficient method.
[0123] 4.1.2 Sub-unit for extracting temporal evolution features; Logic for extracting time evolution features based on the NumPy library: (1) Growth characteristics: Calculate the rate of increase ((peak concentration - initial concentration) / (peak time - actual start time)) (2) Time characteristics: ① Duration: The exact duration (in hours) from the actual start point to the end time of the process is calculated based on the datetime module; ② Duration type: The process is divided into short-duration process (<2 hours, marked as 1) and long-duration process (≥2 hours, marked as 2) according to the duration; ③ High value moment: The occurrence time of high value event (0-23) is extracted, and the sine (2π×hour / 24) and cosine (2π×hour / 24) of the high value moment are calculated.
[0124] 4.1.3 Meteorological correlation feature extraction subunit; Meteorological correlation feature extraction logic based on NumPy library: (1) Meteorological characteristics: extract the temperature, humidity and wind speed at the time of high values; (2) Wind direction characteristics: ① Extract the wind direction that occurs more than 50% of the time during the high-value process as the dominant wind direction and convert it into an eight-directional code (N=0, NE=1, E=2, SE=3, S=4, SW=5, W=6, NW=7); ② Calculate the sine and cosine of the wind direction angle; (3) Pollutant-meteorological correlation characteristics: The correlation coefficient between pollutant concentration and each meteorological factor was calculated using the Pearson correlation coefficient method; the effective range [0,1] was set. This implementation uses the NumPy library to calculate the sine and cosine periodic characteristics of wind direction angles, while retaining time periodic information.
[0125] 4.1.4 Sub-unit for calculating the ratio characteristics of special pollutants; The logic for calculating ratios of special pollutants is constructed, including two types of special ratios: (1) PM 2.5 / PM 10 Ratio: PM at high values 2.5 With PM 10 Concentration ratio, with a denominator of 0 set as a missing value; set the valid range of the ratio to [0,3]; (2) SO2 / NO2 ratio; SO2 / NO2 concentration ratio at high values, with a missing value when the denominator is 0; set the effective range of the ratio to [0,3]; 4.2 Feature Matrix Reconstruction and Core Feature Enhancement Processing Unit; 4.2.1 Feature matrix reconstruction; The feature matrix reconstruction logic is as follows: high-value features, core features, pollutant correlation features, and growth features are set as key features. The feature matrix structure is reconstructed using these key features, and this feature matrix is used as the input for the subsequent SOM model. Among them, the core features include core pollutant concentration, meteorological correlation features, time features, and characteristic pollutant ratio features. The wind direction encoding logic is constructed based on the OneHotEncoder class of the scikit-learn library. One-hot encoding is performed on the eight-directional encoding to generate 8-dimensional binary features. The time period feature transformation logic is designed and implemented by the NumPy library: the hourly periodic features of sine (2π×hour / 24) and cosine (2π×hour / 24) are calculated, and the time dimension information is preserved.
[0126] 4.2.2 Core Feature Enhancement Subunit; ① Duration type and duration magnification factor are 1.8; ② Hourly sine and cosine magnification factors are 1.5; ③ Dominant wind direction magnification factor is 1.1, and wind direction sine and cosine magnification factors are 1.7; ④ PM 2.5 / PM 10 The ratio and the SO2 / NO2 ratio amplification factor are 2.0; ⑤ SO2, NO2, PM 10 PM 2.5 The peak concentration amplification factor for the high-value O3 process is 1.2.
[0127] A core feature weight enhancement model is constructed, using NumPy arrays to build a weight matrix, amplifying the core features according to configuration: ① Duration type and duration (1.8x); ② Hourly sine and cosine (1.5x); ③ Dominant wind direction (1.1x), wind direction sine / cosine (1.7x); ④ Ratio of specific pollutants (2.0x); ⑤ SO2, NO2, PM2.5 10 PM 2.5The peak concentration of O3 during high-value processes (1.2 times); weight amplification is achieved through matrix multiplication to enhance the influence of core features in clustering.
[0128] 4.2.3 Missing feature imputation subunit; The missing value processing logic is constructed based on the differentiated imputation strategy: ① Binary features are imputed using the mode; ② Wind direction features are imputed by month (taking the dominant wind direction of the month); ③ Duration is imputed by the median when grouped by high value label type; ④ General numerical features are imputed by missing rate (<30% uses global median, >30% uses group mean); ⑤ Imputation identifier and method are recorded to ensure feature traceability.
[0129] 4.3 Feature Matrix Generation Unit; 4.3.1 Feature dimension alignment of sub-units; The design unifies the logic of feature dimensions and standardizes all pollution process features; the dictionary mapping ensures consistent feature order, and missing dimensions are filled with default values (0 for binary features and global mean for numerical features) to ensure the uniformity of input dimensions.
[0130] 4.3.2 Matrix Format Conversion Subunit; We construct a feature matrix format conversion logic to convert the aligned features into a float32 type NumPy array and store it as a .npy format file using the numpy.save() function, thereby optimizing model loading efficiency.
[0131] 4.3.3 Feature metadata record subunit; The design logic generates feature metadata, producing a metadata file containing feature name, dimension index, calculation method, data source, and missing rate. The metadata is linked to the feature matrix through timestamps, providing a basis for feature analysis for model training.
[0132] 5: Pollution pattern clustering and classification module; The clustering and typing module uses an enhanced self-organizing map (SOM) network to achieve objective typing of contamination patterns. It is characterized by including a clustering model initialization unit, a feature preprocessing unit, an SOM clustering execution unit, and a clustering result post-processing unit. Its core dependencies include toolkits such as numpy, scipy, and scikit-learn. The specific implementation is as follows: 5.1 Clustering model initialization unit; 5.1.1 SOM Network Parameter Configuration Subunit; Design the dynamic configuration logic for the SOM network mesh, dynamically adapting the mesh size based on the number of samples: (1) Sample size < 20: 3×3 grid; (2) 20 ≤ sample size < 50: 4×4 grid; (3) Sample size ≥ 50: 7×7 grid; The default configuration is a 5×5 grid; the network input dimension equals the feature matrix dimension, and the output dimension equals the number of grid nodes (5×5=25), ensuring the model's input and output adaptability.
[0133] 5.1.2 Weight initialization of sub-units; Based on the PCA dimensionality reduction initialization strategy, the SOM weight initialization logic is constructed: the feature matrix is reduced to 2 dimensions using the PCA class of the scikit-learn library, the dimensionality-reduced data is mapped to SOM grid nodes, and the number of grid nodes × the feature matrix dimension weight matrix is initialized; this avoids training fluctuations caused by random initialization and improves the model convergence stability.
[0134] 5.1.3 Loading Training Parameters Sub-unit; The design incorporates training parameter configuration loading logic, reading core parameters from a JSON configuration file: initial learning rate = 0.08, learning rate decay coefficient = 0.4, maximum number of training epochs = 2000, convergence threshold = 1e-5 (early termination occurs when weight change is less than the threshold), and core feature learning rate gain factor = 10. Dynamic parameter adjustment is supported, allowing adaptation to different training needs without recompilation.
[0135] 5.2 Feature Preprocessing Unit; 5.2.1 Feature-standardized subunit; The feature standardization logic is constructed to perform Z-score standardization ((feature value - mean) / standard deviation) on non-binary features, where the mean and standard deviation are calculated based on the training set to avoid data leakage; the standardized features are stored in Redis cache through the redis-py library to provide an efficient data interface for model training.
[0136] 5.2.2 Feature weight configuration sub-unit; Design feature weight vector construction logic: Based on the air pollution prevention and control requirements of Jiujiang City, this embodiment allocates the following weights: ① Wind direction feature weight = 30; ② Special pollutant ratio feature (PM 2.5 / PM 10 The weights are as follows: ① Ratio, SO2 / NO2 ratio (weight = 28); ② Time feature weight = 20; ③ Core pollutant concentration feature weight = 20; ④ Other pollutant concentration and meteorological feature weight = 15; ⑤ High value marker feature = 10; ⑥ Other feature weight = 1. The weight vector is stored as a NumPy array to provide the weight basis for distance calculation.
[0137] 5.2.3 Data format verification subunit; Construct data format validation logic to verify the consistency of feature matrix dimensions and data validity (no missing values and infinite values); when illegal data is detected, trigger the feature imputation function in unit 4.3.3 of this module to ensure the validity of the model input data.
[0138] 5.3 SOM Clustering Execution Unit; 5.3.1 Weighted distance calculation subunit; The logic for calculating weighted Euclidean distance is built based on the NumPy library: (1) Weighted Euclidean distance magnification of core features: In this embodiment, the magnification factor is set to 8, and the calculation formula is as follows: ; (2) Weighted Euclidean distance of other key features: The formula is as follows: ; The sample features are input vectors, and the node weights are SOM grid node weights; vectorization is used to improve computational efficiency and ensure clustering speed.
[0139] 5.3.2 Iterative Training Engine Subunit; Design the main loop logic for SOM iterative training: (1) Traverse the training samples according to a batch size of 32; (2) Calculate the weighted Euclidean distance between the sample and all grid nodes, and select the best matching unit (BMU) with the smallest distance. (3) Update the weights of BMU and neighboring nodes according to the formula: New weights = old weights + learning rate (t) × neighborhood function (t) × (sample features - old weights); Wherein, learning rate(t) = initial learning rate × exponent (-number of iterations / decay coefficient), neighborhood function(t) = exponent (-distance² / (2×neighborhood radius(t)²)), and neighborhood radius(t) = initial neighborhood radius × exponent (-number of iterations / radius decay coefficient); (4) Calculate the average weight change every 100 iterations. Terminate training when the weight change is less than the convergence threshold or when the maximum number of training rounds is reached. The quantization error (average distance from the sample to the BMU) is recorded in real time during the training process and stored in the training log.
[0140] 5.3.3 Convergence Judgment Subunit; Convergence determination logic is constructed. When the weight change is less than 1e-5 or the current iteration number is greater than or equal to the maximum number of training rounds, a convergence flag (True) is returned and training is terminated. An early termination mechanism is supported to reduce the time spent on invalid iterations.
[0141] 5.4 Post-processing unit for clustering results; 5.4.1 Merging sub-units of small clusters; Design a cluster merging logic, count the number of samples in each cluster and mark the clusters with less than 5 samples; calculate the weighted Euclidean distance between the clusters and other clusters, select the clusters with the smallest distance and a variance increase of ≤1.5 times after merging for merging; update the cluster center (weighted mean of samples from the two clusters) and the sample affiliation relationship after merging to optimize the rationality of the clustering results.
[0142] 5.4.2 Clustering Quality Assessment Subunit; Constructing a three-dimensional evaluation logic for cluster quality: (1) Cluster density: The average variance of key features within a cluster, ≤0.5 is considered excellent; (2) Feature concentration: The ratio of intra-cluster variance to global variance, <0.7 is considered good; (3) Quantization error: The average weighted distance from the sample to the center of its cluster, <0.3 is considered excellent; The evaluation results are stored in a report. Clustering is considered qualified if all superiority indicators are met, ensuring the reliability of the classification. 5.4.3 Model saving sub-units; The model storage logic is constructed by serializing the SOM model (weight matrix, cluster centers, feature weights, evaluation metrics) into files using the pickle library and storing them named "model version_training time"; the model version management table is updated synchronously to ensure model traceability and retrieval.
[0143] 6: Pollution pattern cause diagnosis module; The classification and diagnosis module, based on clustering results and a pollution cause rule base, realizes pollution pattern cause diagnosis. Its features include a pollution cause rule base loading unit, a diagnostic feature preprocessing unit, a multi-level diagnostic engine unit, and a diagnostic result generation unit. It primarily relies on toolkits such as pandas, numpy, scikit-learn, and Flask. The specific implementation is as follows: 6.1 Pollution Cause Rule Base Loading Unit; 6.1.1 Rule parsing subunit; The rule parsing logic is constructed based on the principle of Abstract Syntax Tree (AST), which converts natural language rules in the rule base into executable code objects. For example, "SO2 / NO2≤0.8 and the time period is the morning peak (7-9 am)" is converted into a logical judgment function. It supports AND (&), OR (|), and NOT (!) operators to verify the legality of rule syntax and prevent the risk of malicious code injection.
[0144] 6.1.2 Rule Priority Configuration Sub-unit; Design a four - level priority sorting logic for rules, read the priority fields in the rule library and sort them according to level 1 (data anomaly determination), level 2 (special pollution), level 3 (core pollution source), and level 4 (fallback processing); construct a priority determination queue to ensure that the rules are executed in the order of priority.
[0145] 6.1.3 Add a new pollution cause registration sub - unit; Build an HTTP POST interface ( / api / source_rule / register) based on the Flask framework, design the interface processing logic to support users to upload custom traceability rules in JSON format, including rule name, logical expression, threshold parameter, and priority; after verifying the syntax and parameter legality of the rules, write them into the pollution cause rule library table and synchronously update the determination queue; 6.2 Multi - level diagnostic engine unit; 6.2.1 Data anomaly determination sub - unit; Design a level - 1 priority data anomaly determination logic, including the following rules: (1) Equipment failure (meeting any condition): It is determined as equipment failure if any of the following conditions are met: ① CO is the main pollutant and CO > 4mg / m 3 and the average growth rate > 2; ② SO2 is the main pollutant and SO2 > 300μg / m 3 and the average growth rate > 2; ③ NO2 is the main pollutant and NO2 > 300μg / m 3 and the average growth rate > 2; ④ PM 10 is the main pollutant and PM 10 > 800μg / m 3 ; ⑤ PM 2.5 is the main pollutant and PM 2.5 > 800μg / m ' 3 ; [[ID=3 2]] (2) PM 2.5 inversion determination: If PM 2.5 is the main pollutant and the ratio of PM 2.5 / PM 10 > 1.25, it is determined as PM 2.5 inversion; (3) Ozone pollution (meeting any of the following conditions): O3 is the main pollutant, and 150 < O3 concentration ≤ 160μg / m 3 (ozone pollution); O3 concentration > 160μg / m 3 (ozone is high); In this embodiment, based on the requirements of air pollution prevention and control in Jiujiang City, data anomaly determination rules are set; the present invention can adaptively modify and supplement the data anomaly determination rules based on the prevention and control requirements of different regions.
[0146] 6.2.2 Special Pollution Matching Subunit; Construct a two-level priority special pollution matching logic and iterate through the rule matching: (1) Rules for determining biomass combustion sources (meeting all conditions): ①Time period conditions: Spring (March-May) & Autumn (September-November) & 18:00-24:00 daily | 5:00-9:00 daily; ② Main pollutant conditions: CO|PM 2.5 Must be the main pollutant & PM 10 Not a major pollutant ③ Component correlation conditions: CO and PM 2.5 The Pearson correlation coefficient is ≥0.6; ④ Specific pollutant characteristics: PM 2.5 / PM 10 The ratio is greater than 0.4.
[0147] (2) Rules for determining the source of catering fumes (meeting all conditions): ①Time period conditions: 11:00-13:00 | 18:00-20:00; ② Growth characteristic conditions: SO2 concentration growth rate ≤ 1.1 & NO2 concentration growth rate ≤ 1.1 (non-high growth); ③Main pollutant conditions: PM 2.5 The main pollutant; ④ Correlation condition: SO2-PM 2.5 NO2-PM 2.5 High correlation.
[0148] 6.2.3 Core Pollution Source Identification Sub-unit; The logic for determining core pollution sources is designed with a 3-level priority level, and the determination is performed by traversing the complete rules in the following order.
[0149] 6.2.3.1: Mobile source determination rules; NO2 is the main pollutant, with an SO2 / NO2 ratio ≤ 0.7, and meeting any of the following conditions: ① The process occurs during the morning peak (7-9 am) or evening peak (17-19 pm); ② The correlation coefficient between NO2 and CO is ≥ 0.6; ③ The CO growth rate is > 0.8.
[0150] 6.2.3.2: Rules for determining the source of fireworks displays; (1) Time period determination: The process occurs between January 1st and 10th (Spring Festival) or January 15th (Lantern Festival), and the proportion of cases in this period is >0.4% of the total cases; (2) SO2-dominant type: SO2 is the main pollutant, the SO2 / NO2 ratio is ≥0.8, and the time period is determined. At the same time, the meteorological conditions are humidity ≥60% and wind speed ≤2m / s. (3) PM 2.5 Dominant type: PM 2.5 The main pollutant is SO2, with an SO2 / NO2 ratio ≥ 0.8. SO2 is a high-growth pollutant, and the time period and the above meteorological conditions are met.
[0151] 6.2.3.3: Rules for determining dust sources; (1) PM 10 It is the main pollutant and meets all of the following conditions: ①PM 2.5 / PM 10 Ratio ≤ 0.7; ②PM 10 Concentration ≥115μg / m 3 ③ Wind speed ≥ 3m / s; (2) Meteorological conditions: ① Humidity is humid or high humidity; ② Wind speed is light or calm.
[0152] 6.2.3.4: Rules for determining coal combustion sources; (1) SO2 is the main pollutant, and the SO2 / NO2 ratio is ≥0.75 and the SO2 concentration is ≥15 μg / m³. 3 Or SO2, PM 2.5 Both are major pollutants and their correlation coefficient is ≥0.6; (2) Auxiliary conditions for determining coal combustion sources in winter: ① The process occurs in December or January; ② SO2 and PM 2.5 It has a high correlation; ③ SO2 concentration ≥ 20 μg / m 3 .
[0153] 6.2.3.5: Rules for determining industrial sources; NO2 is the main pollutant, the SO2 / NO2 ratio is <0.7, and any of the following conditions are met: ① not during morning or evening peak hours; ② no high correlation between NO2 and CO (<0.6); ③ CO is not a high-growth pollutant; ④ PM 2.5 As the main pollutant, PM 2.5 It is a high-growth pollutant.
[0154] 6.2.3.6: Rules for determining secondary sources; (1) PM 2.5 PM2.5 is the main pollutant. 2.5 / PM 10 The ratio is ≤1.2 and ≥0.7, and the humidity is ≥60%, while the wind speed is ≤1m / s; (2) Auxiliary conditions: ① O3 is a high-growth pollutant; ② The process occurs at night or in the early morning; ③ PM 2.5 Correlation with O3 > 0.5.
[0155] 6.2.4 Last-Choice Processing Subunit; A four-level priority fallback processing logic is constructed. When all rules are not triggered, the diagnostic result is marked as "unidentified". The process identifier, feature vector, and clustering results are stored in the pending review table to provide data support for expert feedback. Parallel matching of multiple rules is supported to return all special pollution labels that meet the conditions.
[0156] 6.3 Construction Unit for Historical Pollution Causes Case Database; Logic for designing and building a historical pollution cause case library: (1) Extraction of historical pollution cause case data: ① Feature data: Extract the feature matrix of high-value processes in the output historical data; ② Clustering data: Extract the cluster label column after SOM model clustering; ③ Pollution cause: Extract the pollution cause column corresponding to the high-value processes in the output historical data; (2) Construct a historical pollution cause case library based on feature data, cluster data and pollution causes.
[0157] 7: New high-value intelligent diagnostic module; 7.1: New Data Access and Preprocessing Unit; 7.1.1 New Data Access Subunit; The interface adaptation logic of the data acquisition layer module in step S1 is reused. In this embodiment, the station-level data and corresponding city-level data of the national control stations under the jurisdiction of Jiujiang City (including Xiyuan, Shili, Maoshantou, Petrochemical Plant, Foreign Language School, No. 8 Lianhua Avenue, and Comprehensive Industrial Park) are used as new data inputs; the time range is from January 1, 2025 to November 20, 2025, and the time resolution is 1 hour for national control stations; the monitoring data includes station code, monitoring time (accurate to the minute), concentration of six pollutants, and six meteorological parameters.
[0158] 7.1.2 New Data Preprocessing Subunit; Reuse step S2 data preprocessing pipeline: perform missing data repair (linear interpolation / forward filling) and invalid value filtering.
[0159] 7.2 High-value process identification unit; 7.2.1 Identification of discrete high values; Load the rule base configuration, execute the high value judgment logic on the preprocessed data time by time, mark the trigger rule identifier, and generate a new list of discrete high value time moments.
[0160] 7.2.2 Aggregation of continuous high-value processes; Following the logic in 3.4, aggregate the discrete high-value moments to form the complete pollution process, correct the process boundary, and mark the trend.
[0161] 7.2.3 Construction of the characteristic matrix for high-value processes; Using the feature engineering module, extract the key features of the new high values (high value features, core features, pollutant-related features, and growth features) to generate a feature matrix.
[0162] 7.3: Automated pollution pattern cause diagnosis unit for new high-value events; This unit uses a classification model to automatically diagnose the causes of pollution patterns in new high-value events. The XGBoost model is selected, which leverages its strong predictive classification capabilities to automatically classify the causes of pollution in new high-value events based on a historical pollution cause case library.
[0163] 7.3.1: XGBoost model initialization parameter settings; Configure the XGBoost framework to build a multi-class diagnostic model and initialize core parameters: (1) Objective function: multi:softprob (multi-class probability output); (2) Tree structure parameters: maximum depth = 6, learning rate = 0.1, subsample ratio = 0.8, feature sampling ratio = 0.8; (3) Training strategy: random seed = number of features, number of base trees = 100; After the initial training is completed, the model, feature name mapping table and label encoder are serialized and saved. Unless a full retraining is triggered, the classification of new high-value contamination causes will continue to use the trained XGBoost model.
[0164] 7.3.2: Intelligent Diagnosis Subunit for the Causes of New High-Value Pollution Events; This unit utilizes the XGBoost contamination pattern classification model to achieve rapid response in data-driven cause determination.
[0165] 7.3.2.1 XGBoost Model Execution Subunit; The model execution process includes: (1) Feature matching: Align the input data with the feature names used during model training to ensure dimensionality consistency; (2) Prediction execution: The XGBoost model is used to predict the classification of pollution causes; (3) Probability output: Provides the probability distribution of each pollution cause to support uncertainty assessment; (4) Label decoding: Convert the numerical prediction results back into the original text label form of pollution causes.
[0166] 7.3.2.3 Diagnostic quality assessment subunit; Constructing the logic for evaluating diagnostic results: (1) Accuracy verification: Compare the model prediction results with the manual verification results, and calculate the classification accuracy (≥85% is acceptable); (2) Confidence distribution: The percentage of results with high confidence (≥0.7) should be ≥90%; (3) Rule matching degree: Verify the consistency between the diagnostic results and the pollution cause rule base, with a deviation rate ≤5%; Generate diagnostic quality reports to provide a basis for model iterative optimization.
[0167] 8: Feedback and Learning Module; The feedback and learning module is used to process expert feedback data, incrementally update the model, and manage its version, ensuring continuous system optimization. It is characterized by including a feedback data processing unit, an incremental learning engine unit, a full retraining unit, a model version management unit, and a model deployment and monitoring unit. Its core dependencies include toolkits such as pandas, numpy, scikit-learn, XGBoost, and Celery. The specific implementation is as follows: 8.1 Feedback Data Processing Unit; 8.1.1 Feedback data access subunit; An HTTP / JSON interface ( / api / expert / feedback) is built based on the Flask framework. The interface processing logic is designed to support experts to upload review data in batches (≤100 records per batch). The data includes process identifiers, corrected cluster labels, corrected pollution source labels, and rule modification suggestions. The feedback expert identifier and feedback time are recorded and stored in the expert feedback table to ensure the traceability of feedback data.
[0168] 8.1.2 Data Deduplication Sub-unit; A deduplication logic for feedback data is constructed. Deduplication keys are built based on process identifiers, site codes, and core pollutants. Historical feedback data deduplication keys are cached in Redis. New feedback data is compared with the cache, and duplicate data is discarded directly to avoid model bias caused by repeated training.
[0169] 8.1.3 Validity verification subunit; Construct a logic to verify the validity of feedback data, verifying the legality of labels (within the preset label list or new labels) and the completeness of features (no missing key features); count the number of valid data in a single batch, and if there are ≥5, proceed to the learning process; otherwise, store them in the feedback sample library for accumulation to ensure the quality of training data.
[0170] 8.2 Incremental Learning Engine Unit; 8.2.1 Learning trigger condition determination sub-unit; Design incremental learning trigger logic: (1) If the number of new valid data entries in the monitoring feedback sample library is greater than or equal to the set threshold (the set threshold in this embodiment is 20 entries) and does not involve new labels, incremental learning will be automatically triggered; (2) If the data volume threshold is not met, but the amount of newly added valid data is greater than or equal to the time trigger threshold (the threshold is set to 10 in this embodiment), configure a timed trigger task to be executed once every 180 days to avoid performance degradation caused by the model not being updated for a long time.
[0171] 8.2.2 Loading sub-units into the model; The model loading logic is built to load the latest version of the trained XGBoost classification model from the storage directory, including the weight matrix, feature weights, and training parameters; it supports loading specific versions and provides a flexible interface for debugging and rollback.
[0172] 8.2.3 Incremental Training Execution Subunit; Set up incremental training logic: ① Use the XGBoost classification model loaded in step 7.2.2 on the feedback data; ② Adjust the number of decision trees. In this embodiment, it is set to increase by 10 (only update the model parameters, do not change the model structure).
[0173] 8.2.4 Sub-unit for evaluating incremental training results; Define the logic for evaluating incremental training results: (1) Classification accuracy: The number of correctly classified events / total number of events must be greater than the original accuracy + 2%; (2) Accuracy of pollution cause determination: The number of events with correct pollution cause determination / the total number of events must be ≥ the original accuracy + 2%; (3) New sample fit: The prediction accuracy of the newly added feedback samples is ≥90%.
[0174] 8.3 Full Retraining Unit; 8.3.1 Full Retraining Trigger Condition Determination Subunit; Design a full retraining trigger logic that is triggered when any of the following conditions are met: (1) Incremental learning is executed 30 times in total, and full retraining is automatically triggered to ensure the overall stability of the model; (2) When a new tag is detected and the expert completes the rule base update, full retraining is automatically triggered; (3) If the trigger threshold is not met, but the amount of newly added valid data is greater than or equal to the time trigger threshold (the threshold is set to 15 times in this embodiment), a full retraining will be automatically triggered every 360 days to avoid long-term lack of updates; Once triggered, a retraining task with a UUID identifier is generated and recorded in the retraining task table to ensure the traceability of the retraining process. 8.3.2 Constructing a sub-unit from the full retraining dataset; The logic for constructing the retraining dataset integration is as follows: (1) Integrate historical training data with all valid feedback samples, remove duplicate data by deduplication key, and ensure that the total number of samples is ≥100; (2) Perform stratified sampling at a ratio of 7:2:1 (to ensure that each label category is evenly distributed) and divide it into training set, validation set and test set to provide high-quality data for full training.
[0175] 8.3.3 Full training execution sub-unit; Design the logic for the full training process: (1) Re-execute the feature engineering process of the four modules of this system to generate a new feature matrix; (2) Re-execute module 6 of this system, initialize and retrain the XGBoost source model; (4) Detect the rule trigger rate, generate adjustment suggestions for rules with a trigger rate of <1%, store them in the rule optimization table for expert confirmation, and improve the effectiveness of the rules.
[0176] 8.3.4 Retraining Quality Assessment Subunit; Constructing a multi-dimensional retraining quality assessment logic: (1) Pollution cause assessment: The XGBoost model has a pollution cause classification accuracy greater than 85% and is greater than the original model's pollution cause classification accuracy; (2) Generalization ability assessment: The accuracy of the XGBoost model in classifying pollution causes on the test set decreased by ≤5% compared to the training set; If all indicators meet the requirements, an evaluation report will be generated for expert review to ensure the performance of the retrained model.
[0177] 8.4 Model Version Management Unit; Version 8.4.1 generates sub-units; The version number naming convention adopts the format "V[full retraining times].[incremental learning times]" (e.g., V2.3 represents 2 full retraining times and 3 incremental learning times). The version number is bound to the model file, evaluation report, and training data to ensure version traceability.
[0178] 8.4.2 Model storage sub-unit; A dual-mode model storage logic is constructed, supporting local server and cloud storage OSS storage. The stored content includes model weight files, feature weight vectors, training parameters, evaluation reports, and dataset snapshots. A 3-year storage period is configured, and the storage directory structure is built according to version number, supporting query and download by version number, ensuring model security and accessibility.
[0179] Version 8.4.3 backtracking subunit; The design incorporates version rollback logic, which queries the version history table to load the target version model file and configuration. It supports one-click rollback, ensuring that when the performance of the new version decreases by ≥5%, the rollback time is ≤30 seconds, thus guaranteeing system stability.
[0180] 8.5 Model Deployment and Monitoring Unit; 8.5.1 Gray-scale release scheduling subunit; Build a canary release strategy, deploy the new version model and then use a traffic distribution mechanism to route 10% of new events to the new version while keeping 90% of the old version; perform real-time statistics on the diagnostic accuracy and response latency of the two versions, set a 24-hour monitoring period, and provide a basis for decision-making for full switchover.
[0181] 8.5.2 Full-scale switching subunit; The design incorporates a full-scale switchover logic. During the canary release period, if the accuracy of the new version is ≥ the old version + 3% and the response latency is ≤ 3 seconds, a full-scale switchover will be automatically performed (100% of traffic will be routed to the new version). If the accuracy drops, a rollback mechanism will be triggered, and the reason for the failure will be recorded in the deployment log to ensure stable system operation.
[0182] 8.5.3 Real-time monitoring subunit; The system builds real-time monitoring logic based on the InfluxDB time-series database, collecting core indicators such as classification accuracy, source tracing accuracy, confidence distribution, and response latency every hour; and develops a web-based monitoring dashboard that supports indicator visualization and historical trend query, providing real-time data support for system operation and maintenance.
[0183] 8.5.4 Alarm and Report Generation Subunit; Construct a multi-level alarm logic. When an indicator fails to meet a threshold for three consecutive hours (e.g., accuracy < 85%), notify the administrator via SMS or email. Design a monthly report generation logic to compare the performance of different versions, the adoption rate of expert feedback, and the effectiveness of business applications (early warning accuracy), and generate a PDF report to provide a reference for subsequent system optimization.
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.
Claims
1. A method for diagnosing high-value pollution at national monitoring stations based on multimodal clustering, characterized in that: Includes the following steps: S1, Multi-source data standardization preprocessing, including: stable acquisition of multi-source data; missing data repair and removal; outlier identification and filtering; multi-scale data standardization; S2, Construction of the rule base for automatic high-value recognition, including: formal rule definition for high-value judgment; setting of combinational logic for high-value judgment; adaptive strategy for high-value judgment threshold; S3, High-value process identification and multimodal feature construction, including automatic identification and labeling of discrete high-value moments; aggregation of continuous high-value pollution processes; construction of multimodal diagnostic feature vectors; establishment of pollution pattern feature matrix; wherein, the steps of continuous high-value pollution process aggregation include: generation of initial seeds for high-value processes; determination of process merging based on pollutant continuity; intelligent correction of process boundaries and trend labeling; S4, intelligent classification of pollution patterns in high-value processes, includes: feature matrix reconstruction and enhancement; intelligent data imputation and feature standardization; SOM clustering model execution; post-processing and quality assessment of clustering results; and summarization of pollution pattern clustering results. S5, Pollution pattern cause classification, including: construction of a pollution cause classification rule base; setting the priority of cause determination; and summarizing pollution cause classification results; S6, intelligent diagnosis of new high-value events, including: new high-value event data input and preprocessing; high-value process identification and feature extraction; automated diagnosis of pollution pattern causes of new high-value events; and classification and determination of pollution pattern causes. S7, with its self-learning and continuous optimization model, enables intelligent diagnosis of high-value pollution at national monitoring stations.
2. The method for diagnosing high-value pollution at national monitoring stations according to claim 1, characterized in that: The steps for stable acquisition of multi-source data include: Repair threshold setting: Define the continuous missing duration threshold as 2 hours. When the continuous missing duration of a certain parameter is ≤2 hours, it is judged as a short missing. Repairing missing data: For short-term missing data, the repair method is adaptively selected according to the data type: linear interpolation is used for pollutant concentration data, and forward imputation is used for meteorological data; Removal rules: When the duration of consecutive missing data is greater than 2 hours, the data for that period is deemed invalid and the complete data sequence corresponding to that period is directly removed. If a high-value process has been preliminarily identified, the associated high-value process is removed simultaneously.
3. The method for diagnosing high-value pollution at national monitoring stations according to claim 1, characterized in that: The formalized rule definition for high-value determination adopts a multi-dimensional threshold collaborative determination mechanism. The pollutant concentration must meet any combination of conditions from the basic combination, the sudden combination, and the strong high-value rule to trigger the high-value marker. The basic combination is that the dynamic threshold, the static threshold, and the minimum threshold are all exceeded simultaneously. The sudden combination is when the static threshold, the minimum threshold, and the growth threshold are all exceeded simultaneously. The strong high value rule is when the value simultaneously exceeds the dynamic threshold, static threshold, minimum threshold, and growth threshold. The dynamic threshold is a relative threshold derived from historical data of the pollutants themselves, and the calculation formula is as follows: Dynamic threshold = Average concentration of all stations at the same time point × Threshold dynamic factor; The static threshold is an incremental threshold based on the average pollutant concentration at all monitoring stations at the current moment, and the calculation formula is as follows: Static threshold = average concentration of all stations at the same time point + threshold increment; The minimum threshold is the lowest trigger concentration set based on the background concentration characteristics of each pollutant; The growth threshold is a threshold based on the relative increase ratio of pollutant concentration, and the calculation formula is as follows: Growth rate = current concentration / previous concentration.
4. The method for diagnosing high-value pollution at national monitoring stations according to claim 1, characterized in that: The high-value determination combination logic is designed based on Boolean algebra to create a combination judgment process, including: S2.2.1: Composite rule construction, which combines a single condition to form a targeted composite rule by combining logical AND (&) and OR (|) operators. The composite rule includes regular high value rule, sudden high value rule, and strong high value rule, and satisfying any of the above rules is judged as a high value of the corresponding pollutant. S2.2.2: Dynamic extension rule setting, including dynamically adding custom extension rules according to actual monitoring needs while maintaining the original judgment logic, and adjusting the parameters of all rule expressions in combination with the regional pollution cause classification results and seasonal meteorological change patterns.
5. The method for diagnosing high-value pollution at national monitoring stations according to claim 1, characterized in that: The post-processing and quality assessment of the clustering results include: Cluster optimization and merging: Clusters with less than 5 samples are identified as small clusters. The weighted Euclidean distance between the small cluster and other clusters is calculated. If the distance is less than 0.1 and the variance of the key features increases by ≤1.5 times after merging, the small cluster is merged into the nearest large cluster. Clustering quality quantitative evaluation: The clustering effect is comprehensively evaluated using the following cluster density, cluster concentration, and quantification error indicators, among which: Cluster density: Measures the degree of concentration of samples within each cluster on key features. It is evaluated by calculating the average variance of key features within the cluster. The calculation formula is as follows: For each cluster ( ): Density ( )= ; It is a cluster The number of samples; It is a weighted Euclidean distance, calculated using the following formula: ; Cluster concentration: Calculate the ratio of the intra-cluster variance to the global variance for each key feature. The formula is as follows: For each key characteristic (f): Concentration ( )= ; It is the number of clusters; feature In cluster The variance in; yes Variance across the entire dataset; Quantization error: measures how well the SOM network fits the original data, i.e., the average distance of all samples to the weight vector of their respective best-fit unit (BMU). The calculation formula is as follows: Quantization error = ; It is the total number of samples; It uses the same weighted distance function as the S-cluster compactness function.
6. The method for diagnosing high-value pollution at national monitoring stations according to claim 1, characterized in that: The automated pollution pattern cause diagnosis of new high-value events uses a classification model to automatically diagnose the pollution pattern causes of new high-value events. The XGBoost model is used to classify the pollution causes of new high-value events based on a historical pollution cause case library. The steps of the automated pollution pattern cause diagnosis of new high-value events using the classification model include: inputting new high-value event data; executing the XGBoost model; and post-processing and quality assessment of the classification results.
7. The method for diagnosing high-value pollution at national monitoring stations according to claim 1, characterized in that: The model's self-learning and continuous optimization include: S7.1, Expert Feedback Data Collection and Standardization, including: feedback data sources and types; feedback data screening and validity determination; feedback data standardization processing; S7.2, Incremental Learning Triggering and Execution, including: Incremental learning triggering conditions; Incremental learning execution process; Evaluation of incremental training results; S7.3, Full Retraining Mechanism and Execution, including: Full Retraining Triggering Conditions; Full Retraining Data Preparation; Full Retraining Execution Process; Retraining Model Quality Assessment; S7.4, Model Version Management and Deployment, includes: version management rules; model deployment and switching; optimization effect monitoring and feedback.
8. The method for diagnosing high-value pollution at national monitoring stations according to claim 7, characterized in that: Incremental training results are evaluated using a test set to assess the performance of the updated model. Key metrics include: Classification accuracy: The number of correctly classified events / the total number of events must be greater than or equal to the original accuracy + 2%; Accuracy of pollution cause determination: The number of events with the correct pollution cause / the total number of events must be ≥ the original accuracy + 2%; New sample fit: Prediction accuracy of newly added feedback samples ≥ 90%.
9. A high-value pollution diagnostic system for national monitoring stations based on multimodal clustering, characterized in that: The system is used to execute the high-value pollution diagnosis method for national control stations as described in any one of claims 1 to 9, and includes a data acquisition layer module, a data preprocessing module, a high-value identification module, a feature engineering module, a pollution pattern clustering and typing module, a pollution pattern cause diagnosis module, a new high-value intelligent diagnosis module, and a feedback and learning module.
10. The high-value pollution diagnostic system for national monitoring stations according to claim 9, characterized in that: The data acquisition layer module includes a data source interface adaptation unit, a data acquisition scheduling unit, a raw data storage unit, and a data acquisition integrity verification unit. The data preprocessing module includes a preprocessing initialization unit, a missing data processing unit, an invalid value detection and filtering unit, and a data output unit. The high-value identification module includes a rule base management unit, a threshold rule setting unit, a high-value judgment engine unit, and a high-value process aggregation unit; The feature engineering module includes a multimodal feature extraction unit, a feature matrix reconstruction and core feature enhancement processing unit, and a feature matrix generation unit; The pollution pattern clustering and classification module includes a clustering model initialization unit, a feature preprocessing unit, an SOM clustering execution unit, and a clustering result postprocessing unit. The pollution pattern cause diagnosis module includes a pollution cause rule base loading unit, a multi-level diagnosis engine unit, and a historical pollution cause case library construction unit. The new high-value intelligent diagnostic module includes a new data access and preprocessing unit, a high-value process identification unit, and a new high-value event automated pollution pattern cause diagnosis unit. The feedback and learning module includes a feedback data processing unit, an incremental learning engine unit, a full retraining unit, a model version management unit, and a model deployment and monitoring unit.