Air quality high-value event identification method, device, equipment, medium and product
By aligning and fusing features of multi-source heterogeneous data in a spatiotemporal manner, and combining meteorological correction factors for dynamic threshold calculation and spatial clustering, the problem of accurately identifying and tracing the source of high air quality events has been solved, and efficient pollution source determination and optimized resource scheduling have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 重庆市生态环境监测中心
- Filing Date
- 2026-04-15
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies cannot accurately identify high-air-quality events, resulting in poor accuracy in tracing pollution sources and slow response times, making it difficult to meet the needs of precise pollution control.
By acquiring multi-source heterogeneous data, performing spatiotemporal alignment and feature fusion, combining meteorological correction factors for dynamic threshold calculation, utilizing spatial clustering for source tracing analysis, and comparing with historical event records to generate deduplicated event records, the data is then classified and scheduled according to preset rules.
It enables accurate identification and rapid response to high air quality events, improves the accuracy of pollution source determination and control efficiency, avoids resource waste, and ensures timely handling of pollution incidents.
Smart Images

Figure CN122065012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of air quality monitoring technology, and in particular to a method, device, equipment, medium and product for identifying high air quality events. Background Technology
[0002] With the continuous advancement of industrialization and urbanization, air pollution has become a key concern in ecological and environmental governance. Real-time air quality monitoring and rapid response to high-value air quality events are directly related to the improvement of ecological and environmental quality and the public's living and working experience. To address the complex pollution problems caused by multiple pollutants such as fine particulate matter and ozone, various regions have gradually established multi-dimensional air quality monitoring systems. The accurate identification, rapid source tracing, and efficient dispatch of high-value air quality events have become major requirements for improving governance efficiency and achieving precise pollution control in air pollution prevention and control. Constructing a scientific method for identifying high-value air quality events is of great significance for promoting the digitalization and intelligentization of air pollution prevention and control.
[0003] However, existing technologies cannot accurately identify high air quality events. Summary of the Invention
[0004] This application provides a method, apparatus, equipment, medium, and product for identifying high air quality events, which can achieve accurate identification of high air quality events.
[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, this application provides a method for identifying high air quality events, including: Acquire multi-source heterogeneous data for the target area; wherein, the multi-source heterogeneous data includes at least two of the following: air quality monitoring data, meteorological data, pollution source emission inventory data, video monitoring data, traffic data, electricity consumption data, radar data, and pollutant component data; Spatiotemporal alignment processing is performed on multi-source heterogeneous data to obtain multimodal feature vectors. The spatiotemporal alignment processing includes consistency verification of aligned data and spatiotemporal unified mapping of unaligned data based on latitude and longitude and hourly timestamps. Based on meteorological correction factors and multimodal feature vectors, dynamic threshold calculations are performed to obtain preliminary event information; among which, meteorological correction factors are determined based on the correlation between historical pollutant concentration data of the target area and meteorological baseline data. Spatial clustering is used to perform spatiotemporal correlation analysis on preliminary event information to obtain event tracing results; The event tracing results are compared with historical event records to generate duplicate event records; Based on the preset event classification rules, the deduplicated event records are classified into different levels, and scheduling results are generated.
[0006] In some possible implementations, spatial clustering is used to perform spatiotemporal correlation analysis on preliminary event information to obtain event tracing results, including: Based on preliminary event information, the corresponding locations of abnormal air quality concentrations were identified; Spatial clustering of air quality concentration anomalies yields at least one anomalous cluster; Based on the spatial location of the abnormal clusters and the real-time wind direction, the reverse trajectory of pollutants is extrapolated to determine the location of the pollution source; Spatial matching of pollution source locations with pollution source emission inventories yields a list of potential emission sources. The location of pollution sources and the list of potential emission sources are used as the results of incident tracing.
[0007] In some possible implementations, the event tracing results are compared with historical event records to generate deduplicated event records, including: Each event in the event tracing results is taken as the current event, and it is determined whether the pollutant type of the current event is consistent with that of the previously identified events in the historical event records; If the pollutant type of the current event is consistent with that of a previously identified historical event, then the spatiotemporal overlap and directional similarity between the current event and the previously identified historical event are calculated. If the spatiotemporal overlap is greater than the first preset threshold and the directional similarity is greater than the second preset threshold, then the current event and the historically identified events are determined to be the same continuous event, and the influence range, concentration change process, evolution state and duration of the historically identified events are updated; otherwise, the current event is stored as a new event in the historical event record to obtain a deduplicated event record.
[0008] Among the possible implementations are: When the preliminary suspected pollution source type in the preliminary event information is any one of open burning, traffic, or industrial emission sources, the identification results of video monitoring data, the congestion index of traffic data, the load anomaly characteristics of electricity data, and the particulate matter echo characteristics of radar data are jointly verified in a multimodal manner to correct the confidence level and update the grade of the preliminary suspected pollution source type, thus obtaining the corrected preliminary event information. Spatial clustering is used to perform spatiotemporal correlation analysis on the corrected preliminary event information to obtain the event tracing results.
[0009] In some possible implementations, dynamic threshold calculations are performed on multimodal feature vectors based on meteorological correction factors to obtain preliminary event information, including: Extract air quality concentration features of the target area from multimodal feature vectors; By using meteorological correction factors to correct the basic thresholds of pollutants, dynamic discrimination thresholds are obtained. The air quality concentration characteristics are compared with dynamic discrimination thresholds, and preliminary event information is generated based on the comparison results.
[0010] In some possible implementations, spatiotemporal alignment of multi-source heterogeneous data is performed to obtain multimodal feature vectors, including: A unified spatiotemporal grid for the target area is constructed based on latitude and longitude and hourly timestamps; Multi-source heterogeneous data is processed through at least one of the following methods: interpolation, spatial mapping, temporal aggregation, or inversion. The processing results are then mapped to a unified spatiotemporal grid to form modal features within the unified spatiotemporal grid. These modal features include air quality concentration features, meteorological features, video smoke features, traffic flow density features, electricity load features, radar echo features, and pollutant composition features. Modal features within a unified spatiotemporal grid are fused at the feature level to obtain multimodal feature vectors.
[0011] Secondly, this application provides a high-air-quality event identification device, comprising: The acquisition module is used to acquire multi-source heterogeneous data of the target area; wherein, the multi-source heterogeneous data includes at least two of the following: air quality monitoring data, meteorological data, pollution source emission inventory data, video monitoring data, traffic data, electricity consumption data, radar data, and pollutant component data; The data processing module performs spatiotemporal alignment on multi-source heterogeneous data to obtain multimodal feature vectors. This spatiotemporal alignment includes consistency verification of aligned data and unified spatiotemporal mapping of unaligned data based on latitude, longitude, and hourly timestamps. Based on meteorological correction factors, dynamic threshold calculations are performed on the multimodal feature vectors to obtain preliminary event information. These meteorological correction factors are determined based on the correlation between historical pollutant concentration data and meteorological baseline data for the target area. Spatial clustering is used to perform spatiotemporal correlation analysis on the preliminary event information to obtain event tracing results. Finally, the event tracing results are compared with historical event records to generate deduplicated event records. The scheduling module is used to classify deduplicated event records into different levels according to preset event classification rules and generate scheduling results.
[0012] Thirdly, this application provides a computing device, including a memory and a processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of the first aspects.
[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program for performing the method as described in any one of the first aspects.
[0014] Fifthly, this application provides a computer program product comprising one or more computer instructions, wherein when the computer instructions are executed by a computer, the computer performs the method as described in any one of the first aspects.
[0015] As can be seen from the above technical solution, this application has at least the following beneficial effects: This application overcomes the limitations of single data sources by acquiring at least two types of multi-source heterogeneous data from the target area, including air quality monitoring, meteorology, and pollution source emission inventories. It integrates real data from multiple dimensions, such as environmental monitoring, meteorological conditions, pollution source lists, and video surveillance. The fusion of multiple data types reflects air quality conditions and influencing factors from different perspectives, avoiding analytical biases caused by missing data dimensions. Furthermore, the complementarity of various data types allows subsequent event identification and source tracing to better align with actual pollution scenarios, improving the objectivity and accuracy of the overall analysis process.
[0016] This study performs spatiotemporal alignment processing on multi-source heterogeneous data, including consistency verification and unified spatiotemporal mapping, to obtain multimodal feature vectors. First, consistency verification eliminates contradictory and anomalous invalid information, ensuring data validity and rationality. Then, unified spatiotemporal mapping of latitude and longitude with hourly timestamps resolves the heterogeneity of multi-source data in both spatial location and time dimension, bringing data from different sources and formats into a unified spatiotemporal system and achieving data normalization. Furthermore, transforming the processed data into multimodal feature vectors extracts the main features of various data types, achieving feature-level fusion. This simplifies subsequent data analysis complexity while preserving key information from multiple sources, significantly improving analytical efficiency and accuracy.
[0017] Based on meteorological correction factors determined by combining historical data from the same period and meteorological benchmark data, dynamic threshold calculations are performed on multimodal feature vectors to obtain preliminary event information. This achieves dynamic adaptation of pollutant judgment thresholds, overcoming the drawback of fixed thresholds being unable to match different meteorological conditions. The meteorological correction factors allow thresholds to be adjusted according to real-time meteorological conditions, making the criteria for judging high air quality values more closely aligned with the actual environment of the target area, effectively avoiding misjudgments or omissions caused by meteorological factors. Simultaneously, by using dynamic threshold calculations to identify abnormal air quality conditions and compile preliminary event information such as the event's occurrence time, affected stations, and preliminary suspected pollution source types, basic information on pollution events can be quickly identified.
[0018] Spatial clustering was used to conduct spatiotemporal correlation analysis on preliminary event information and obtain event source tracing results. First, spatial clustering aggregated discrete concentration anomaly points into anomaly clusters, clearly defining the concentrated impact area of pollution and avoiding source tracing bias caused by analyzing scattered anomaly points individually. Then, by combining diffusion direction with real-time meteorological data for reverse trajectory extrapolation, the location of pollution sources could be scientifically and accurately determined. Finally, by spatial matching with pollution source emission inventories, a list of potential emission sources was screened, achieving accurate tracing from pollution phenomena to pollution sources. Combining temporal diffusion changes with spatial concentration distribution allows pollution source tracing to move beyond information from a single monitoring point and analyze the overall pollution pattern, significantly improving the accuracy of pollution source identification.
[0019] By comparing the event tracing results with historical event records and generating deduplicated event records, duplicate event records are effectively eliminated, avoiding repeated scheduling and control of the same ongoing pollution event. This reduces the waste of human and material resources and improves the efficiency of pollution control. Simultaneously, for cases determined to be the same ongoing event, information such as its impact range, concentration change process, and duration is updated in a timely manner, enabling the system to form a complete and continuous record of the evolution of the pollution event, facilitating staff to grasp the development trend of the pollution event. New events are included in the historical record, achieving dynamic updating and improvement of event records.
[0020] Based on preset event classification rules, duplicate event records are categorized and scheduling results are generated. By classifying events into Level 1, Level 2, and Level 3 based on pollution concentration, impact range, and pollution source confidence, high-value air quality events of varying severity can be differentiated. This allows for accurate allocation of control resources, with more control efforts deployed for high-level events with severe pollution and wide impact, and appropriate control measures implemented for low-level events with mild pollution. This avoids resource waste or inadequate control caused by a one-size-fits-all approach. Simultaneously, targeted scheduling content is developed based on event levels, clearly defining responsibility areas and control requirements. Information dissemination and task feedback are also planned, ensuring that scheduling instructions are accurately and quickly transmitted to relevant responsible parties. A complete task feedback channel is established, achieving closed-loop control from instruction issuance to task execution and result feedback. This significantly improves the efficiency and effectiveness of control response to high-value air quality events, ensuring timely and effective handling of various pollution events, helping to rapidly reduce pollutant concentrations and improve regional air quality. Ultimately, accurate identification of high-value air quality events enables efficient pollution control.
[0021] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0022] Figure 1 An application environment diagram for a method for identifying high air quality events provided in this application embodiment; Figure 2 A flowchart illustrating a method for identifying high air quality events provided in this application embodiment; Figure 3 A structural diagram of an air quality high-value event identification device provided in this application embodiment; Figure 4 This is a schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation
[0023] The terms "first," "second," and "third," etc., used in this application specification and accompanying drawings are used to distinguish different objects, not to limit a specific order.
[0024] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0025] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the related technologies is given first: Currently, in the field of high-air quality event identification, most technologies rely on single-dimensional data for event determination. Even when some solutions incorporate multi-source data, they only achieve simple information overlay without professional alignment and fusion processing tailored to the spatiotemporal characteristics of heterogeneous multi-source data. This results in the inability to effectively uncover correlations between data points, hindering accurate event identification and tracing. Furthermore, the industry commonly uses fixed thresholds for pollutant concentration anomalies without dynamically adjusting these thresholds based on meteorological conditions and regional geographical features. This makes event determination results susceptible to environmental interference, leading to frequent misjudgments and missed detections.
[0026] In the pollution source tracing and incident handling stages, the relevant technologies are mostly based on isolated analysis of individual pollutant concentration anomalies, lacking spatial clustering of anomalies and reverse extrapolation of pollutant diffusion trajectories, making it difficult to accurately locate pollution sources. Furthermore, the lack of a standardized event deduplication and classification mechanism makes it easy to repeatedly dispatch resources for the same persistent pollution event, resulting in a waste of control resources. At the same time, the event classification is highly subjective, making it impossible to accurately allocate control resources based on the degree of pollution and the scope of impact. Ultimately, this leads to low identification efficiency, poor source tracing accuracy, and slow response to high-value air quality events, making it difficult to meet the actual needs of precise and scientific pollution control in current air pollution prevention and control work.
[0027] In view of this, embodiments of this application provide a method for identifying high air quality events. To make the technical solution of this application clearer and easier to understand, the application scenarios of the technical solution are described below with reference to the accompanying drawings. Figure 1 As shown, this figure is an application environment diagram provided by an embodiment of this application.
[0028] In this application environment, the server continuously receives multi-source heterogeneous data from various monitoring devices in the target area. After completing the entire process of identifying, tracing, deduplicating, classifying, and generating scheduling results for high-value air quality events, the server pushes scheduling results, event details, pollution source tracing information, and other data to the terminal. The terminal receives and displays various information sent by the server for relevant operation and maintenance and control personnel to view. At the same time, personnel can upload pollution source inspection status, control execution results, task feedback information, etc. to the server through the terminal. The server receives and stores this feedback data for updating event records and evaluating control effectiveness, forming a two-way interactive closed loop of server data processing and instruction issuance, and terminal reception and display and information uploading.
[0029] To make the technical solution of this application clearer and easier to understand, the following describes a method for identifying high air quality events provided by an embodiment of this application, using server 104 as the execution subject, in conjunction with the above application scenario. Figure 2As shown, this figure is a flowchart illustrating a method for identifying high air quality events according to an embodiment of this application. The method for identifying high air quality events includes: S201. Obtain multi-source heterogeneous data of the target area.
[0030] Among them, multi-source heterogeneous data includes at least two of the following: air quality monitoring data, meteorological data, pollution source emission inventory data, video monitoring data, traffic data, electricity consumption data, radar data, and pollutant component data.
[0031] Air quality monitoring data refers to various types of data obtained by monitoring the hourly concentrations, concentration variations, and data continuity of pollutants such as PM2.5, PM10, ozone, SO2, NO2, and CO at various monitoring stations within the target area.
[0032] Meteorological data are relevant data reflecting the meteorological conditions of the target area, including meteorological indicators such as wind speed and wind direction that can help determine the direction of pollutant diffusion and the location of pollution sources.
[0033] Pollution source emission inventory data is a list of various pollution sources in the target area, covering dynamic updates of 1-5 sources (1 to 5 types of pollution sources, including domestic sources, industrial sources, agricultural sources, transportation sources and biomass combustion sources), elevated source lists, dust source lists, area source lists, and other lists of enterprises and pollution sources that emit relevant pollutants.
[0034] Video monitoring data is image data of the target area obtained by video monitoring equipment such as high-altitude observation, which can identify pollution-related information such as smoke from open burning.
[0035] Traffic data refers to relevant data reflecting traffic conditions within a target area, such as traffic flow, traffic density, and congestion index.
[0036] Electricity consumption data refers to electricity usage data such as electricity load and abnormal load characteristics of enterprises and other entities within the target area.
[0037] Radar data is data such as particulate matter echo characteristics acquired by radar equipment, which can be used in scenarios such as pollution process analysis.
[0038] Pollutant composition data is data on the composition of pollutants obtained through monitoring at component stations, which can assist in pollution source tracing and analysis.
[0039] For example, firstly, the target area for identifying high-air-quality events is clearly defined, and the corresponding geographical boundaries and monitoring coverage are delineated to ensure the accuracy of the data acquisition scope. Then, for different types of data sources, corresponding acquisition channels are established, connecting to the monitoring systems of air quality monitoring stations to collect real-time air quality monitoring data such as pollutant concentrations and concentration changes at each station. The meteorological monitoring platform is linked to obtain real-time meteorological data such as wind speed and direction within the target area. The pollution source emission inventory system of relevant ecological and environmental departments is retrieved, and the updated lists of various pollution sources and related emission information are synchronized to form pollution source emission inventory data. The backend systems of video monitoring equipment such as high-altitude observation devices are connected to capture real-time video images within the target area to obtain video monitoring data. In conjunction with traffic management platforms, traffic data such as traffic flow and congestion index are collected for each road segment within the target area. The electricity consumption monitoring system of the power sector is connected to obtain electricity consumption data such as electricity load and load changes for enterprises and related entities within the area. Radar monitoring equipment is connected to collect radar data such as particulate matter echo characteristics. Finally, the system of pollutant component monitoring stations is connected to collect pollutant component data.
[0040] During the data collection process, data will be acquired using real-time collection, scheduled synchronization, and on-demand retrieval methods, depending on the update frequency and monitoring requirements of various data types. Simultaneously, preliminary format identification and filtering will be performed on the collected data to remove obviously invalid data and ensure the validity of the acquired data. Finally, at least two types of data collected from different channels and in different ways will be aggregated to form a multi-source heterogeneous dataset for the target area.
[0041] S202. Perform spatiotemporal alignment processing on multi-source heterogeneous data to obtain multimodal feature vectors.
[0042] The spatiotemporal alignment process includes consistency verification of aligned data and spatiotemporal mapping of unaligned data based on latitude and longitude and hourly timestamps.
[0043] One feasible approach involves constructing a unified spatiotemporal grid for the target region based on latitude, longitude, and hourly timestamps; processing multi-source heterogeneous data through at least one of interpolation, spatial mapping, temporal aggregation, or inversion, and mapping the processing results to the unified spatiotemporal grid to form modal features within the unified spatiotemporal grid; wherein the modal features include air quality concentration features, meteorological features, video smoke features, traffic flow density features, electricity load features, radar echo features, and pollutant component features; and performing feature-level fusion of the modal features within the unified spatiotemporal grid to obtain a multimodal feature vector.
[0044] Spatiotemporal alignment is an operation that normalizes heterogeneous data from different sources and in different spatiotemporal dimensions, so that various types of data maintain a unified match in spatial location and temporal dimension, ensuring the spatiotemporal consistency of data analysis.
[0045] Consistency verification is the process of verifying the rationality of multi-source heterogeneous data that has been matched in the spatiotemporal dimension. It checks the logical relationship and numerical rationality between data and eliminates invalid data that is contradictory or abnormal.
[0046] Spatiotemporal unified mapping is a processing method that maps multi-source heterogeneous data that do not match in the spatiotemporal dimension to the same spatiotemporal system based on a unified spatial and temporal benchmark.
[0047] The unified spatiotemporal grid is a standardized spatiotemporal reference framework constructed for a target area using fixed latitude and longitude intervals as spatial units and hourly timestamps as time units. It serves as a unified carrier for various types of data.
[0048] Interpolation is a numerical processing method that estimates and completes missing spatial or temporal data, enabling the data to be fully mapped to a unified spatiotemporal grid.
[0049] Spatial mapping is the process of converting the original spatial coordinates of various types of data into unified spatiotemporal grid spatial coordinates, thereby achieving the unification of spatial dimensions.
[0050] Time-series aggregation is a time-dimension processing method that integrates and merges raw data of different time granularities according to hourly timestamps.
[0051] Inversion is a data analysis method that derives target parameters from observation data and completes the missing feature information in the grid.
[0052] Modal features are unique characteristics of various types of data presented within a unified spatiotemporal grid after processing of multi-source heterogeneous data. They are a concrete manifestation of the essential characteristics of data.
[0053] Air quality concentration characteristics are characteristic attributes that reflect the magnitude and range of change in the concentration of pollutants such as PM2.5, PM10, and ozone within a target area.
[0054] Meteorological characteristics are the characteristic attributes that reflect the meteorological conditions of the target area, such as wind speed and wind direction.
[0055] Video smoke features are smoke-related characteristic attributes extracted from video monitoring data that can characterize pollution phenomena such as open burning.
[0056] Traffic density features are characteristic attributes extracted from traffic data that reflect the density of vehicle distribution within a region.
[0057] Electricity load characteristics are features extracted from electricity data that reflect the magnitude and abnormal states of electricity load used by enterprises and other entities.
[0058] Radar echo characteristics are echo-related feature attributes extracted from radar data that reflect the distribution and concentration of particulate matter.
[0059] Pollutant component characteristics are characteristic attributes extracted from pollutant component data that reflect the composition of pollutant components.
[0060] Feature-level fusion is the process of integrating different types of modal features within a unified spatiotemporal grid through algorithms to form a feature set that incorporates multi-dimensional information.
[0061] Multimodal feature vectors are vectors formed by quantifying and representing fused multi-dimensional modal features. They integrate the core information of various heterogeneous data and serve as the basic data form for subsequent data analysis.
[0062] For example, the first step is to perform basic spatiotemporal alignment processing. This involves a preliminary screening of the collected multi-source heterogeneous data to distinguish between datasets that are naturally matched in spatiotemporal dimensions and those that are not. For aligned data, a consistency check is performed. Through logical comparison and numerical rationality analysis, the correlation and validity between data from different data sources are verified, and contradictory or abnormal data are eliminated to ensure the quality of the aligned data. For unaligned data, a unified spatiotemporal mapping is performed using latitude and longitude as the spatial reference and hourly timestamps as the time reference. This transforms the original spatiotemporal coordinates of various data types to the same spatiotemporal dimension, completing the initial spatiotemporal regularization.
[0063] Next, a unified spatiotemporal grid is constructed according to feasible methods. Based on the geographical latitude and longitude range of the target area, spatial grid units with fixed intervals are divided. At the same time, using hourly timestamps as time units, a standardized unified spatiotemporal grid covering the target area is built as a unified framework for carrying various types of data. Subsequently, the multi-source heterogeneous data that has completed preliminary spatiotemporal alignment is processed in a targeted manner. Depending on the data missingness and spatiotemporal granularity differences, one or more methods such as interpolation, spatial mapping, temporal aggregation, or inversion are used to complete missing data, transform spatial coordinates, integrate temporal granularity, and derive missing features. The processed data are then mapped to the pre-constructed unified spatiotemporal grid. Within the grid, various corresponding modal features such as air quality concentration features, meteorological features, and video smoke features are extracted, enabling different types of data to form their own feature expressions within the unified spatiotemporal grid.
[0064] Finally, feature-level fusion is carried out on all modal features within the unified spatiotemporal grid. Through a professional algorithm model, the scattered and different types of modal features are deeply integrated, the correlation information between each feature is mined, and the fused multi-dimensional feature information is then quantified and transformed into a structured multimodal feature vector.
[0065] For example, a target area can be defined. The spatiotemporal domain is constructed using latitude and longitude and hourly timestamps as a benchmark to create a unified spatiotemporal grid. :
[0066] in, It is to divide the target area into Each spatial grid cell Corresponding to a fixed latitude and longitude range; For time grid sets, ={1,2,…, } represents all time steps of the spatiotemporal grid; This represents the total number of steps on the timeline. Represents spatial grid cells and time , is a location combination that can be used to describe the value of a variable at which grid and at which time.
[0067] Assume there is a total Data source of various modalities , Each data source has an original spatiotemporal coordinate system. Through the spacetime mapping function Aligning it to a unified spatiotemporal grid can be represented as:
[0068] in, Depending on the data type, one or more of the following processing methods can be used, such as interpolation (to complete missing spatial or temporal values) and spatial mapping (to map the original spatial coordinates). Mapped to spatial grid cells ), time series aggregation (aggregating non-hourly data into hourly granularity), and inversion (using observation data to invert gridded features); Let k be the original spatiotemporal coordinate system of the k-th data source. The original spatial coordinates (such as latitude and longitude, station location). This is the original time coordinate (such as minutes, days, etc.).
[0069] After mapping, modal features within a unified spatiotemporal grid are obtained. such as air quality concentration characteristics Meteorological characteristics Video smoke characteristics Traffic flow density characteristics Electricity load characteristics Radar echo characteristics Pollutant composition characteristics .
[0070] In a unified spatiotemporal grid All modal features are fused at the feature level to form a multimodal feature vector. :
[0071] in, This is a feature-level fusion function, which can be achieved through concatenation, weighted combination, or attention-based fusion methods. The final result is:
[0072] in, Let be a 3D real vector space, and let represent the value space of the fused feature vector. The feature dimensions after fusion.
[0073] To ensure alignment quality, consistency check constraints can be introduced. :
[0074] in, Indicates a unified spatiotemporal grid cell Above, the first Features of the modality It is different from Another modality index, usually with They appear in pairs to represent two different data modes. That is, for the same spatiotemporal grid cell, the features of different modes satisfy consistency constraints. If the conflict level exceeds the conflict threshold If the data is abnormal, it will be marked as abnormal and removed or corrected.
[0075] S203. Based on meteorological correction factors and multimodal feature vectors, dynamic threshold calculation is performed to obtain preliminary event information.
[0076] Among them, the meteorological correction factor is determined based on the correlation between the historical pollutant concentration data of the target area and the meteorological baseline data. It is used to correct the basic pollutant threshold so that the threshold is adapted to the air quality monitoring needs under real-time meteorological conditions.
[0077] Preliminary event information is a set of basic information about high-value air quality events obtained through dynamic threshold calculations, including the event time, affected stations, pollutant types, concentration change rates, and preliminary suspected pollution source types.
[0078] One possible approach is to extract air quality concentration features of a target area from a multimodal feature vector; correct the basic pollutant threshold using a meteorological correction factor to obtain a dynamic discrimination threshold; compare the air quality concentration features with the dynamic discrimination threshold, and generate preliminary event information based on the comparison results.
[0079] Optionally, when the preliminary suspected pollution source type in the preliminary event information is any one of open burning, traffic, or industrial emission sources, the identification results of video monitoring data, the congestion index of traffic data, the load anomaly characteristics of electricity consumption data, and the particulate matter echo characteristics of radar data are jointly verified using multimodal methods. The confidence level of the preliminary suspected pollution source type is corrected and the level is updated to obtain the corrected preliminary event information. Based on the corrected preliminary event information, spatiotemporal correlation analysis is performed to obtain the event source tracing results.
[0080] Among them, dynamic threshold calculation is a calculation process that relies on meteorological correction factors to dynamically adjust the basic threshold of pollutants, and then combines multimodal feature vectors to carry out threshold comparison and judgment.
[0081] Historical pollutant concentration data refers to the monitoring data of various pollutant concentrations recorded in the target area during the same period in the past, and it is one of the basic data for deriving meteorological correction factors.
[0082] Meteorological baseline data are standardized basic meteorological monitoring data within the target area, covering key meteorological indicators such as wind speed and wind direction, and are used to analyze the correlation between meteorological conditions and pollutant concentrations.
[0083] The basic threshold for pollutants is a basic judgment limit set for various types of air quality pollutants, and serves as an initial reference standard for identifying high-value events.
[0084] The dynamic discrimination threshold is a pollutant judgment limit obtained after correction by meteorological correction factors. It can be dynamically adapted according to real-time meteorological conditions to improve the accuracy of event identification.
[0085] The hourly increase in concentration is the rate of increase in pollutant concentration over a one-hour period, and it is an important indicator for determining abnormal air quality concentrations.
[0086] Spatial gradient difference is the difference in pollutant concentration between different monitoring stations or areas, which reflects the spatial distribution differences in pollutant concentration.
[0087] The slope of concentration change is a trend indicator of how pollutant concentration changes over time, reflecting the rate of concentration change.
[0088] Multimodal joint verification is a process of cross-verifying the types of preliminary suspected pollution sources by combining feature information from multiple sources such as video, traffic, electricity consumption, and radar.
[0089] Confidence correction is an operation that adjusts the confidence level of the preliminary suspected pollution source type based on the results of multimodal joint verification.
[0090] The rating update is the process of reclassifying high-air quality events based on the results of multimodal joint verification.
[0091] The revised preliminary event information is high-level air quality event information after confidence level correction and level update, which is more in line with the actual pollution situation.
[0092] For example, firstly, meteorological correction factors are derived in advance, and historical pollutant concentration data and meteorological baseline data of the target area are sorted out. Through data analysis, the correlation between the two types of data is mined, and meteorological correction factors adapted to different meteorological conditions are calculated based on these patterns. Next, from the multimodal feature vectors obtained in the early stage, the air quality concentration characteristics of the target area are accurately screened out through feature extraction algorithms. These features include information such as the real-time concentration and concentration change trend of various pollutants, which are important bases for threshold determination. Subsequently, the basic thresholds corresponding to various pollutants are retrieved, and the meteorological correction factors are combined with the basic pollutant thresholds for calculation. The basic thresholds are dynamically corrected to obtain dynamic discrimination thresholds adapted to the current real-time meteorological conditions, making the threshold determination more in line with the actual environment.
[0093] After calculating the dynamic discrimination threshold, the extracted air quality concentration characteristics are comprehensively compared with the dynamic discrimination threshold. Simultaneously, a joint judgment is made by combining at least one of the following indicators: hourly concentration increase, spatial gradient difference, and concentration change slope. This multi-dimensional indicator approach comprehensively identifies whether air quality concentrations are abnormal. For identified anomalies, information such as the event occurrence time, affected monitoring stations, types of pollutants exceeding standards, and concentration change rate are compiled and recorded. Based on pollutant concentration change characteristics and preset pollution source identification rules, suspected pollution source types are initially identified, and this information is integrated to form preliminary event information.
[0094] If the preliminary suspected pollution source type in the initial event information is any one of open burning, traffic, or industrial emission sources, further multimodal joint verification will be conducted. This involves retrieving smoke identification results from video monitoring data, congestion index from traffic data, load anomaly characteristics from electricity consumption data, and particulate matter echo characteristics from radar data, and cross-referencing these data with the characteristics of the preliminary suspected pollution source type. Based on the verification results, the confidence level of the preliminary suspected pollution source type will be adjusted to improve the accuracy of pollution source identification. Simultaneously, the event level will be updated based on the verification results, forming a revised preliminary event information. Subsequent spatiotemporal correlation analysis will be conducted based on this revised information to obtain the event source tracing results. If the preliminary suspected pollution source type is not one of the above three categories, the preliminary event information will be directly used as the basis for subsequent spatiotemporal correlation analysis.
[0095] Optionally, the process of constructing meteorological correction factors can specifically include: defining pollutant types. Spatiotemporal grid unit The meteorological conditions are determined by at least one of the following: wind speed, wind direction, temperature, humidity, atmospheric pressure, boundary layer height, and precipitation intensity. After standardization and normalization, these factors form a multidimensional vector, which serves as the meteorological feature vector. Specifically, it can be represented as:
[0096] in, spatiotemporal grid unit The meteorological feature vector at time t; u(g,t) is the spatiotemporal grid cell. The real-time wind speed at time t; θ(g,t) is the spatiotemporal grid cell. The real-time wind direction at time t; T(g,t) is the spatiotemporal grid cell. The real-time temperature at time t; H(g,t) is the spatiotemporal grid cell. The relative humidity at time t; P(g,t) is the spatiotemporal grid cell. Atmospheric pressure at time t; Hbl(g,t) is the spatiotemporal grid cell. The boundary layer height at time t; R(g,t) is the spatiotemporal grid cell. Hourly precipitation / precipitation intensity at time t.
[0097] Based on historical data from the same period, a correlation model between pollutant concentration and meteorological conditions was established, and a meteorological correction factor was obtained:
[0098] in, For pollutant type p in spatial grid cells The meteorological correction factor at time t is used to correct the impact of current meteorological conditions on pollutant concentrations. Pollutant types from the same historical period In spatial grid units Concentration at time t; This represents the conditional expected concentration under the same meteorological conditions. This represents the overall average concentration for the same period in history.
[0099] This indicates that current meteorological conditions are unfavorable for dispersion, which could easily lead to concentration accumulation. This indicates that meteorological conditions are favorable for the dispersion of pollutants.
[0100] Furthermore, a dynamic discrimination threshold is generated, and the pollutant type is set. The basic threshold is Then the threshold is dynamically determined. for:
[0101] in, An adjustable sensitivity coefficient is used to control the tightness of the threshold, and can be configured according to regional pollution characteristics or season.
[0102] Furthermore, air quality concentration feature extraction and anomaly detection can be performed from multimodal feature vectors. Types of pollutants extracted In spatial grid units ,time Monitoring concentration Define a single point of failure indicator function. It can be represented as:
[0103] in, This is an indicator function; its value is 1 if the condition within the parentheses is true, and 0 otherwise.
[0104] like =0 (i.e.) ), directly determine the spatial grid cell At any moment If the condition is normal, there is no need to proceed to the subsequent comprehensive anomaly index calculation; if =1 (i.e.) If the location is marked as a suspected abnormal location, then that location will be identified.
[0105] Using the single-point anomaly indicator function After a rapid initial screening of monitoring data across the entire grid and all time periods, suspected anomaly points are identified. These suspected anomaly points are then input into the comprehensive anomaly index. Multi-dimensional feature fusion is performed to determine whether suspected anomalies are genuine anomalies by quantifying the degree of anomaly. A comprehensive anomaly index is defined. It can be represented as:
[0106] in, This represents the hourly increase in concentration. For spatial grid units The monitoring concentration of pollutant type p at time t. For spatial grid units The monitoring concentration of pollutant type p at time t-1; This represents the spatial gradient difference, reflecting the concentration difference between adjacent grid cells. The slope of the concentration change, i.e., the rate of change of concentration over time, reflects the speed and trend of concentration change. To monitor concentration in real time; These are the weighting coefficients.
[0107] The preliminary event determination rule can be expressed as:
[0108] in, This is a preliminary anomaly indication function used to determine spatial grid cells. At any moment Below, pollutant types Has a preliminary abnormal event been triggered? This is the threshold for anomaly detection. If... If this is triggered, an initial event is set up, and the event time is recorded. Spatial grid unit Pollutant types Concentration change rate and preliminary suspected pollution source types based on rule matching. Based on the hourly increase in concentration and the slope of concentration change, a comprehensive judgment is made according to the preset classification threshold, and an initial level is directly assigned to the current event. This initial level is the preliminary event level. For example, taking PM2.5 as an example, the hourly increase in concentration is set to ≥50 μg / m3 or the slope of concentration change is set to ≥15 μg / (m3). h), corresponding to level 3, 20 μg / m3 50 μg / m3 or 5 μg / (m3) h) slope of concentration change 15μg / (m3 h) corresponds to level 2, and the rest correspond to level 1. If the hourly increase in PM2.5 concentration in a certain grid unit is 60 μg / m3 and the slope of the concentration change is 8 μg / (m3) h), then a preliminary event level is assigned based on a comprehensive assessment. =3, meaning that when making a comprehensive judgment, the highest level corresponding to each indicator is taken as the final preliminary event level.
[0109] Furthermore, multimodal joint verification and confidence level correction can be performed. When the initial suspected pollution source type... In this case, multimodal joint verification is introduced. A multimodal verification vector is defined. :
[0110] in, It is the confidence score for video smoke recognition. It can be obtained by collecting real-time video images from high-altitude observation video monitoring equipment, reasoning through a smoke target detection model based on deep learning, outputting the probability of the category of smoke presence, and then normalizing it. The value range is [0,1]. It is a normalized value of the traffic congestion index. It can be obtained by using basic data such as real-time vehicle speed, traffic flow and actual travel time of road segments through traffic checkpoints, floating car GPS (Global Positioning System), microwave traffic radar and other equipment; combined with the historical traffic data of the road segment and reference speed, the original quantitative value of the road segment congestion degree is calculated and then normalized to the [0,1] interval through maximum and minimum normalization. It is an abnormal characteristic of electricity load. Hourly electricity load data can be obtained from smart meters and power acquisition terminals. By comparing it with the historical benchmark load curve for the same period, the load deviation can be calculated and the load deviation can be normalized to the [0,1] interval to directly obtain the abnormal characteristic of electricity load. The particulate matter echo intensity can be obtained by an atmospheric particulate matter lidar based on the Mie scattering principle. It emits a laser and receives the backscattered echo signal of aerosol particles. After distance correction, transmittance correction and geometric overlap factor correction, the normalized echo intensity is obtained. The value range is [0,1], which is the particulate matter echo intensity.
[0111] It should be noted that the initial event level Compared with the preliminary suspected pollution source type Based on multimodal feature vectors The comprehensive assessment based on pollutant concentration, spatiotemporal distribution, and trends aims to quickly identify anomalies and complete preliminary classification, but it does not perform specific verification for pollution source types. The subsequent multimodal joint verification, however, targets three typical pollution sources—open burning, transportation, and industrial emissions—using highly source-correlated features such as video smoke, traffic congestion, electricity load, and radar echoes to... The authenticity and confidence level are specifically verified and corrected, and the event level is updated based on the corrected confidence level. .therefore, It is a preliminary level based on the degree of pollution. It is a precise level based on the authenticity of the pollution source.
[0112] Constructing pollution source types The confidence correction function can be expressed as:
[0113] in, The confidence level is adjusted for pollution source type s. The initial confidence level for pollution source type s (which can be set according to historical rules); Types of pollution sources The corresponding weight vector reflects the contribution of each mode to the discrimination of the pollution source type; The transpose of is used to allow the two vectors to be used for the inner product.
[0114] After confidence level correction, the updated event level can be represented as:
[0115] in, This is the updated event level (the revised level). Initial event level; The increment is determined based on the confidence level, such as... , It rounds down; This is the highest event level.
[0116] The final preliminary event information was obtained, including the revised pollution source type. and the updated event level This is used for subsequent spatiotemporal correlation analysis and event tracing. Among them, Indicates the adjusted confidence level. The largest pollution source type s, which is the most likely pollution source type determined in the final analysis.
[0117] S204. Through spatial clustering, spatiotemporal correlation analysis is performed on the preliminary event information to obtain the event tracing results.
[0118] The results of the incident tracing include the location of the pollution source and a list of potential emission sources.
[0119] One feasible approach involves identifying corresponding air quality concentration anomaly locations based on preliminary event information; spatially clustering these anomaly locations to obtain at least one anomalous cluster; performing reverse trajectory extrapolation of pollutants based on the spatial location of the anomalous cluster and real-time wind direction to determine the location of the pollution source; spatially matching the pollution source location with the pollution source emission inventory to obtain a list of potential emission sources; and using the pollution source location and the list of potential emission sources as the event tracing results.
[0120] Among them, the diffusion direction is the direction in which pollutants spread in the atmosphere with the help of meteorological conditions and other factors, and it is an important basis for determining the location of pollution sources.
[0121] Spatial clustering is a data analysis method that categorizes and aggregates discrete points that are related in spatial location, highlighting concentrated areas of abnormal pollutant concentrations.
[0122] Spatiotemporal correlation analysis is an analytical process that combines changes in the time dimension and the distribution in the spatial dimension to mine correlations of information related to an event, and is used to trace the source of a pollution event.
[0123] The result of the incident tracing is a set of pollution source-related information obtained through spatiotemporal correlation analysis, such as the location of the pollution source and a list of potential emission sources.
[0124] Anomaly locations in air quality concentration are monitoring stations or spatial locations within a target area where the concentration of pollutants exceeds the dynamic discrimination threshold.
[0125] An anomalous cluster is a concentrated area of pollution formed by the aggregation of multiple adjacent points with abnormal air quality concentrations after spatial clustering.
[0126] Reverse trajectory extrapolation is an analytical method that uses the direction of pollutant diffusion and real-time meteorological conditions to reverse the trajectory of pollutant propagation, thereby determining the approximate location of the pollution source.
[0127] The location of the pollution source is the approximate geographical direction and spatial position of the pollution source obtained by reversing the trajectory.
[0128] The potential emission source list is a list of emission sources in a region that may produce corresponding pollution after spatially matching the location of pollution sources with the emission inventory of pollution sources.
[0129] Spatial matching is the process of comparing and filtering geographic location information with the location information in the pollution source emission inventory.
[0130] For example, based on the information recorded in the preliminary event information, such as the affected sites and abnormal pollutant concentration data, all abnormal air quality concentration points in the target area are accurately located and extracted, and the latitude and longitude, pollutant type, and concentration value of each abnormal point are determined. Then, a spatial clustering algorithm is used to analyze and process all discrete abnormal air quality concentration points. According to the spatial adjacency and the correlation of concentration anomalies, these points are aggregated into one or more abnormal clusters, clearly delineating the concentrated areas of pollutant concentration anomalies, and determining the location, coverage, and other important spatial characteristics of each abnormal cluster.
[0131] Subsequently, by combining the precise spatial location of the anomalous cluster, real-time meteorological data such as wind direction and speed in the target area were retrieved. Using the anomalous cluster as the endpoint, a reverse trajectory extrapolation method was employed to deduce the propagation trajectory of pollutants in the atmosphere. Combined with the analysis results of the diffusion direction, the approximate geographical location of the pollution source was determined, thus completing the determination of the pollution source location. Next, the pollution source emission inventory data for the target area was retrieved, and the determined pollution source location was spatially matched with the emission source spatial locations in the inventory. By comparing the pollutant emission type and emission scale of the emission sources in the inventory, emission sources that match the pollutant type of this pollution event and are located within the pollution source location range were selected, forming a list of potential emission sources. Finally, the pollution source location obtained through reverse trajectory extrapolation and the list of potential emission sources selected through spatial matching were integrated to form a complete event source tracing result, completing the entire spatiotemporal correlation analysis and pollution source tracing process.
[0132] Optionally, anomaly locations can be identified based on preliminary event information. Defined on the spacetime grid Exception indicator function , can be represented as:
[0133] in, This is an indicator function; its value is 1 when the condition within the parentheses is true, and 0 otherwise. Pollutant type In spatial grid units ,time The monitored concentration; For dynamic threshold determination.
[0134] All satisfied The grid points constitute the set of anomaly points. :
[0135] in, The spatial grid index and time index are for the i-th anomaly point.
[0136] Furthermore, the set of abnormal locations Spatial clustering is performed using density-based clustering methods (such as DBSCAN), and the clustering function is defined as follows:
[0137] in, Let M be the set of M anomalous clusters obtained after clustering, each It is a spatially adjacent subset of outlier points; This represents the total number of outlier clusters obtained after spatial clustering. ; For the first An anomalous cluster contains a set of spatially adjacent anomalous points; The radius of the spatial neighborhood; This represents the minimum number of points required to form a cluster.
[0138] Each anomalous cluster Its centroid coordinates can be used and coverage radius Characterization:
[0139]
[0140] in, The centroid coordinates of the m-th anomalous cluster; It is the first Anomalous clusters Number of outliers included; Let g be the center coordinates (latitude and longitude).
[0141] The location of the pollution source can be determined based on the inverse trajectory deduction. For each anomalous cluster, the real-time wind direction (represented by the wind direction at the cluster's centroid) is obtained, and a diffusion direction vector is defined:
[0142] in, For the first Real-time wind direction of an anomalous cluster; For the first The diffusion direction vector of an anomalous cluster.
[0143] The reverse trajectory extrapolation uses a Lagrange inversion model, considering wind speed and time decay factors, and defines the pollution source orientation vector:
[0144] in, Let be the source orientation vector of the m-th anomalous cluster; Let be the real-time wind speed scalar for the m-th anomalous cluster, represented by the wind speed at the cluster's centroid, characterizing the magnitude of pollutant diffusion velocity. This represents the total number of steps in the reverse deduction process; For the first A time decay factor with a time step size is used to correct for the concentration decay effect of pollutants over time, and can be expressed as: , The diffusion attenuation rate; For the first The time interval of each time step is the time discrete unit for reverse trajectory derivation.
[0145] The location of the pollution source can be further represented in polar coordinates:
[0146] in, Let be the azimuth angle (polar coordinates) of the pollution source of the m-th anomalous cluster. The four-quadrant arctangent function is used to calculate the azimuth vector of the pollution source. The angle between the x-axis and the positive x-axis returns the azimuth angle of the pollution source; The source orientation vector corresponding to the m-th anomalous cluster The y-axis (vertical axis) component, i.e. The vertical coordinate value in the figure; The pollution source orientation vector corresponding to the m-th anomalous cluster The x-axis (horizontal axis) component, i.e. The x-coordinate value in the graph.
[0147] Furthermore, the pollution source emission inventory is defined as follows: ,in, This is a pollution source emission inventory, containing records of L pollution sources. For the first The geographical coordinates of the pollution source For the first The types of pollution sources (such as industrial sources, transportation sources, open burning sources, etc.). For the first The emission intensity of each pollution source.
[0148] For anomalous clusters Calculate its spatial matching degree with each pollution source in the inventory:
[0149] in, For the m-th anomalous cluster and the list of... Spatial matching degree of individual pollution sources Location of pollution source With the inferred source location The Euclidean distance between them To smooth out parameters and avoid division by zero; For the first The emission intensity of each pollution source is positively correlated with the emission amount; Pollution source type and pollutant type The matching coefficient has a value range of [0,1].
[0150] The list of potential emission sources can be represented as:
[0151] in, This is the list of potential emission sources corresponding to the m-th anomalous cluster, i.e., the set of pollution sources with a matching degree exceeding a threshold. This is the matching threshold.
[0152] Integrate the results of event tracing; event tracing results It can be represented as:
[0153] in, The results of the incident tracing include the location of all pollution sources and a list of potential emission sources corresponding to all abnormal clusters; This represents the total number of outlier clusters obtained after spatial clustering. The pollution source orientation vector of the m-th anomalous cluster is obtained through reverse trajectory deduction; The list of potential emission sources corresponding to the m-th anomalous cluster is obtained by spatial matching.
[0154] That is, each anomalous cluster corresponds to a pollution source location and a list of potential emission sources.
[0155] S205. Compare the event tracing results with historical event records to generate deduplicated event records.
[0156] One possible approach is to treat each event in the event tracing results as the current event and determine whether the pollutant type of the current event is consistent with that of historically identified events. If the pollutant type of the current event is consistent with that of historically identified events, the spatiotemporal overlap and directional similarity between the current event and historically identified events are calculated. If the spatiotemporal overlap is greater than a first preset threshold and the directional similarity is greater than a second preset threshold, the current event and historically identified events are determined to be the same continuous event, and the influence range, concentration change process, evolution state, and duration of historically identified events are updated. Otherwise, the current event is stored as a new event in the historical event record to obtain a deduplicated event record.
[0157] Among them, the historical event record is a complete information archive of previously identified high air quality events that have been stored in the system, including the type of pollutants, the spatial and temporal range, pollution source information, and duration of the event.
[0158] Deduplicated event records are standardized event record sets formed by comparing the current event tracing results with historical event records, updating historical event records (such as updating the impact range and concentration change process of historically identified events, or adding new events), and integrating continuous event information.
[0159] The current event is the single high-value air quality event corresponding to the event tracing results obtained from this spatiotemporal correlation analysis.
[0160] Spatiotemporal overlap is a quantitative indicator used to measure the degree to which a current event overlaps with a previously identified historical event in terms of both the time frame and the spatial region in which the event occurred.
[0161] Locational similarity is a quantitative indicator used to determine the degree of similarity between the location of pollution sources and the direction of pollution-affected areas of current events and previously identified historical events.
[0162] The first preset threshold is a critical value set by the system to determine whether the spatiotemporal overlap has reached the standard of high overlap.
[0163] The second preset threshold is a critical value set by the system to determine whether the orientation similarity reaches the standard of high similarity.
[0164] The same ongoing event is determined to be a continuation or development of the same pollution event, and not a newly occurring high air quality event.
[0165] For example, comparing the event tracing results with historical event records to generate deduplicated event records first involves extracting all high-air quality events included in the current event tracing results, treating each event as a separate current event, and simultaneously retrieving all historical event records stored in the system to prepare data for subsequent comparisons. Next, for each current event, it is matched against previously identified historical events to determine the pollutant type, primarily checking whether the types of pollutants exceeding standards in the current event and historical events are consistent. If the pollutant types are inconsistent, the current event is directly classified as a new event and temporarily stored in the pending entry list.
[0166] If the current event shares the same pollutant type as a previously identified historical event, a specialized algorithm will calculate their spatiotemporal overlap and directional similarity. The spatiotemporal overlap comprehensively considers the proportion of overlapping time intervals and the intersection of the pollutant-affected spatial regions. The directional similarity compares the degree of agreement between the pollution source locations and the direction of pollution diffusion. After calculation, the spatiotemporal overlap is compared to a first preset threshold, and the directional similarity is compared to a second preset threshold. If both indicators exceed their respective preset thresholds, the current event and the previously identified historical event are determined to be the same ongoing event. The information for that historical event is then updated, supplementing the current event with information on the pollution impact range, pollutant concentration changes, and pollution evolution status. Simultaneously, the duration of the event is extended to ensure the completeness of the historical event record—that is, the historical event record is updated.
[0167] If the spatiotemporal overlap does not reach the first preset threshold, or the directional similarity does not reach the second preset threshold, and either condition is not met, the current event is judged as a new event and added to the entry list. After all current events have been compared and judged one by one with historical event records, all new events in the entry list are entered into the historical event records in a unified format, while the updated historical continuous event information is retained, and finally integrated to form a complete and standardized deduplicated event record.
[0168] Optionally, let the current event be... Historical events are Each event It can be represented as:
[0169] in, Type of pollutant; Let e be the time interval of the pollution event. Characterizes the time period in which the pollution event occurred. Let e be the start time of the pollution event. Let e be the end time of the pollution event. The spatial influence area (a set of grid points or a geometric polygon); The source location vector for the event pollution is obtained by aggregating the source location vectors of all abnormal clusters. Evolutionary states (such as generation, development, weakening, and dissipation); Duration.
[0170] Define pollutant type matching function , can be represented as:
[0171] in, The type of contaminant in the current event; The types of pollutants from historical events.
[0172] like If the event is not identified as a different type of event, it will directly enter the new event handling process.
[0173] If the pollutant types are the same, calculate the spatiotemporal overlap. .
[0174] Time overlap It can be represented as:
[0175] in, The time interval of the current event is represented as [ , ], The start time of the current event. The start time of the current event; The time interval of historical events is represented as [ , ], The start time of a historical event. The start time of a historical event; The length (duration) of the intersection of the two time intervals; The length of the union of the two time intervals; Indicates the length of the time interval.
[0176] Spatial overlap It can be represented as:
[0177] in, The spatial influence area of the current event can be a set of grid cells or a geometric polygon; The spatial impact area of historical events; The intersection area of the spatial regions (or the number of grid cells); The area of the union of spatial regions (or the number of grid cells); This indicates the area of a spatial region (or the number of grid cells).
[0178] Spatiotemporal overlap comprehensive It can be represented as:
[0179] in, This is the time weighting coefficient for spatiotemporal overlap, which can be adjusted according to the application scenario.
[0180] Furthermore, directional similarity calculation is performed. The directional vector of the event pollution source is defined. and the direction vector of the affected area (From the centroid of the spatial region to the location of the pollution source). Orientation similarity. Composed of two weighted parts, it can be expressed as:
[0181] in, , where is the cosine similarity of the source orientation vectors. This is the location vector of the pollution source for the current event (obtained through reverse trajectory deduction). This represents the location vector of the pollution source in historical events; To influence the direction cosine similarity of the region, This is the direction vector of the area affected by the current event, pointing from the centroid of the spatial region to the location of the pollution source. This represents the directional vector of the area of influence of a historical event. For source orientation weighting coefficients, Calculate the modulus (e.g., norm / absolute value).
[0182] That is, location similarity Based on pollution source orientation vector and the direction vector of the affected area Calculated, and It was derived from the reverse trajectory. Information such as the azimuth angle of the pollution source is already included in the vector.
[0183] Define the overall matching degree:
[0184] in, This is the overall matching degree function; This is the threshold for spatiotemporal overlap. This is the location similarity threshold.
[0185] like If the event is true, it is determined to be the same continuous event, and an update operation is performed; otherwise, it is determined to be a new event.
[0186] For cases determined to be the same continuous event, update the historical event information:
[0187] The current event is Historical events are , ( ) is the event update function; The assignment arrow indicates that the result of the expression on the right is assigned / written to the variable on the left, thus completing the data update and overwrite.
[0188] The update function is defined as follows:
[0189] in, The updated historical event time range is the union of the two time ranges (from the earliest start time to the latest end time). This represents the start time of the current event. This is the end time of the current event; The starting time of a historical event; The end time of a historical event; The updated spatial influence area of historical events is the union of the two spatial areas (the set of grid cells or the merging of geometric polygons). This is the updated pollution source location vector; The weighted average function assigns weights based on the duration of the event. The duration of a historical event; The duration of the current event; The updated duration is equal to the length of the new time interval; The updated evolutionary state (such as "generated", "developed", "continued", "weakening", "dissipated" etc.); This is a state update function that determines the evolutionary stage of an event based on temporal continuity and spatial extensibility.
[0190] All events identified as new are merged with the updated historical events to form a deduplicated event record. :
[0191] in, For all updated historical events (persistent events merged with the current event); The current event (a single event in this tracing result); For each historical event recorded in the historical event record; If the overall match between the current event and any historical event is 0, meaning it does not match any historical event, then the current event is considered a newly added event.
[0192] To improve the robustness and adaptability of the method, fuzzy matching of pollutant types can be further introduced, which can be expressed as:
[0193] in, This is a fuzzy type matching function used to determine whether two types of pollutants are similar (not necessarily identical). This is a pollutant type similarity matrix function that returns the similarity value between two pollutant types (e.g., ...). ), This is the type matching threshold; when the similarity exceeds this threshold, the types are considered to be consistent.
[0194] A dynamic threshold adjustment mechanism can also be introduced, which can be expressed as:
[0195] in, The spatiotemporal overlap threshold is dynamically adjusted (changing with time or event duration). This serves as the basic threshold for spatiotemporal overlap. The decay coefficient controls the rate at which the threshold decreases with the duration of the event; The duration of a historical event; The reference time scale (normalization factor) is used to make the duration dimensionless.
[0196] For events with a long duration, the spatiotemporal overlap requirement should be appropriately reduced to avoid the event being incorrectly segmented due to minor changes.
[0197] S206. According to the preset event classification rules, the deduplicated event records are classified into levels and the scheduling results are generated.
[0198] Among them, the event classification rules are standards pre-defined by the system to classify high-value air quality events, based on three dimensions: pollution concentration, scope of impact, and confidence level of pollution source.
[0199] Pollution concentration refers to the actual monitored concentration of pollutants exceeding the standard during an event and the extent to which it exceeds the dynamic discrimination threshold.
[0200] The scope of impact refers to the geographical area affected by a high air quality event, the number of monitoring stations involved, and other indicators that characterize the scale of the event's impact.
[0201] Pollution source confidence refers to the degree of credibility of the determination of pollution source type and potential emission source in the results of event tracing.
[0202] A Level 1 event is a high-value air quality event determined according to the classification rules, characterized by the most severe pollution, the widest impact range, and a high confidence level of the pollution source.
[0203] Level 2 events are high-quality air quality events with moderate levels of pollution, scope of impact, and confidence in the pollution source.
[0204] Level 3 events are high-quality air quality events with relatively low pollution levels and a small affected area.
[0205] The scheduling result is a set of information containing pollution control-related instructions based on the event level. It mainly includes scheduling content, information push and task feedback information. In other words, the scheduling result is the product of the preset event classification rules applied to deduplicated event records.
[0206] The dispatch content consists of specific instructions for pollution inspection, control, and treatment issued to the corresponding responsible areas in response to high-level air quality events of different levels.
[0207] Information push is the act of sending scheduling content to the terminals of relevant responsible units and maintenance personnel.
[0208] Task feedback information refers to the information reported by the relevant responsible units to their superiors after completing the dispatched tasks, including the inspection and control situation and the effectiveness of pollution control.
[0209] For example, the deduplicated event records are classified according to the preset event classification rules and the scheduling results are generated. First, the preset event classification rules in the system are retrieved to clarify the specific judgment criteria for Level 1, Level 2, and Level 3 events in dimensions such as the extent of pollution concentration exceeding the standard, the number of sites / area covered by the affected area, and the range of pollution source confidence values. At the same time, all event information in the deduplicated event records is extracted, and the classification basis such as the pollution concentration data, the details of the affected area, and the pollution source confidence results of each event are sorted out to prepare for the classification.
[0210] Next, for each event in the deduplicated event record, its pollution concentration, impact range, and pollution source confidence level are compared and matched with the pre-set event classification rules. Through comprehensive judgment of multi-dimensional indicators, a corresponding level is assigned to each event, namely, level one, level two, or level three event. Clear level labels are made for events of different levels, thus completing the level classification of all events.
[0211] Subsequently, based on the pre-defined event levels and specific information such as the type of pollution source, affected area, and types of pollutants exceeding standards, targeted dispatching measures were formulated. For high-level Level 1 events, stricter instructions for pollution source inspections, comprehensive pollution control, and rapid remediation were issued. For Level 2 events, routine and targeted inspection and control requirements were formulated. For Level 3 events, basic pollution source investigation and concentration reduction measures were formulated. At the same time, the responsible areas, responsible units, and completion deadlines for each event dispatching measure were clearly defined.
[0212] Following this, based on the responsibility attribution of the dispatched content, an information push operation is executed, sending the dispatch content for different events to the corresponding responsible units and maintenance personnel's work terminals to ensure that relevant entities receive control instructions in a timely manner. Simultaneously, the dispatch content clearly defines the requirements, channels, and time limits for task feedback, establishing an information channel for task feedback. This facilitates subsequent feedback from responsible units on inspection and control status, pollution control effectiveness, and other task feedback information. The dispatch content, the recipients and methods of information push, and the requirements for task feedback are integrated to form a complete dispatch result. Finally, the dispatch result is synchronized to the corresponding air quality control platform, achieving systematic storage and display of the dispatch results.
[0213] Optional, define each event in the deduplication event record. The rating is determined by a multi-dimensional index:
[0214] in, For the event The indicator vector contains evaluation indicators in five dimensions; The number of times the pollution concentration exceeds the standard. , Pollutant type The basic threshold, The measured maximum concentration is the maximum concentration of pollutants detected at the sampling point or within the monitoring area during this event monitoring. The extent of influence is expressed in terms of the number of grid cells or the coverage area. Normalized representation; For the confidence level of the pollution source, (From multimodal joint validation), take the confidence probability corresponding to all candidate pollution source types s. The maximum value in; Social impact factors, such as whether it involves sensitive areas (schools, hospitals) or transportation hubs; The degree of urgency is not judged by a single dimension, but by a combination of three important factors, such as the time dimension (duration) and the trend dimension (e.g., the trend of concentration change). ), and environmental dimensions (such as meteorological diffusion conditions). ).
[0215] The weighted composite scoring function can be expressed as:
[0216] in, This is the overall score for the event, used for subsequent rating. Let the weight of the j-th indicator satisfy... ; The normalized value of the j-th indicator is mapped to... The interval is normalized using the following function:
[0217] in, This represents the original value of the j-th indicator; This is the minimum value of the indicator within the entire sample or a preset range; This is the maximum value of the indicator across the entire sample or a preset range.
[0218] For indices with nonlinear sensitivity, S-shaped normalization can be used:
[0219] in, This is the sensitivity coefficient. This is the inflection point value.
[0220] Furthermore, preset event classification rules can be defined, and threshold vectors for classifying levels can be defined. :
[0221] in, For the event Overall score The minimum threshold (maximum standard) for Level 1 events, with a score ≥ The event is classified as a Level 1 event; This is the minimum threshold for a level 2 event. ≤Rating< The event is a level two event; This is the minimum threshold for Level 3 events. ≤Rating< The event is classified as a Level 3 event; rating < The event is classified as "Other" (can be ignored or only logged).
[0222] Threshold vector It is not a fixed basic threshold. Instead, it is dynamically replaced in actual operation. ( =1,2,3), this dynamic adjustment can be based on different regional and seasonal characteristics. Specifically, dynamic threshold adjustment can be introduced:
[0223] in, For the first Dynamic threshold adjusted for level events ( =1,2,3); For the first The base threshold (preset value) for level-3 events. For seasonal correction factors (such as the threshold being appropriately increased due to the frequent occurrence of smog in winter), Indicates the month or season; Regional adjustment factors (e.g., restrictions can be appropriately relaxed for industrial zones, but tightened for ecological protection zones). For example... yes The corresponding base threshold, yes The corresponding dynamically adjusted dynamic threshold.
[0224] Based on event level And event attributes, generate scheduling results :
[0225] Among them, Content refers to the scheduling content, that is, the specific control instructions; Push refers to information push, specifying the push target and push method; Feedback refers to task feedback information, specifying the feedback channel, feedback time limit, etc.
[0226] The scheduling content is populated by a basic template and dynamic parameters:
[0227] in, This is the standard instruction template for the corresponding level; For template filling operations; For the event The dynamic parameters include the area of responsibility, type of pollution source, scope of impact, and recommended measures.
[0228] The push target is determined by the responsible entity matching function, which can be expressed as:
[0229] in, For the event The optimal push target (responsible entity) is the entity with the lowest score after considering both spatial distance and functional matching. A set of responsible entities; an agency is a set of responsible entities. A single entity within; This represents the weighting coefficient for spatial distance; Spatial distance; The geographic location vector (such as latitude and longitude coordinates) of the responsible agency; For the event Geographical location vectors (such as the centroid coordinates of the polluted area); This is the weighting coefficient for the degree of job suitability; The degree of matching between the functions of the responsible entity and the type of pollution source; The type of function of the responsible agency (such as air pollution supervision, water pollution supervision, industrial enterprise supervision, etc.). For the event Types of pollution sources.
[0230] Push method Based on the level selection: Level 1 events use telephone + SMS + system push, Level 2 events use SMS + system push, and Level 3 events use system push.
[0231] You can also define a feedback closed-loop evaluation function:
[0232] in, The closed-loop evaluation score is used to quantify the execution effect of scheduled tasks. A higher value indicates better task execution quality. This represents the number of tasks that have been completed. To allocate the total number of tasks; This is the average response time; 1 represents the response time sensitivity coefficient.
[0233] The complete mapping from deduplicated event records to scheduling results can be represented as:
[0234] in, It is a complete mapping function from deduplicated event records to scheduling results, integrating the entire process such as level classification, dynamic threshold adjustment, push target matching, scheduling content generation, and feedback closed-loop evaluation; For deduplicating event records; This is the weight vector of the rating indicators; This is the threshold vector for classifying levels; This is a dynamic threshold correction factor; For the set of responsible parties; 2 represents the feedback closed-loop sensitivity coefficient.
[0235] To further improve the accuracy of classification, an adaptive weight adjustment mechanism can also be adopted:
[0236] in This is the updated weight vector for the rating indicators. Basic weights; The learning rate; s is the scheduling performance loss function (such as the difference between the actual response time and the expected response time); Softmax is the Softmax normalization function, used to map the calculation results into weight coefficients in the form of a probability distribution; For scheduling effect loss function The gradient with respect to the weights.
[0237] Simultaneously, machine learning models (such as gradient boosting trees) can be introduced to perform non-linear fitting of the comprehensive score:
[0238] in, For the event The overall score is used for grade classification; For the event The index vector; For pre-trained machine learning models, These are model parameters that can capture complex interaction effects between indicators.
[0239] Based on the above, the high-air quality event identification method overcomes the limitations of single data sources by acquiring at least two types of multi-source heterogeneous data from the target area, including air quality monitoring, meteorology, and pollution source emission inventories. It integrates real data from multiple dimensions, such as environmental monitoring, meteorological conditions, pollution source lists, and video surveillance. The fusion of multiple data types reflects air quality conditions and influencing factors from different perspectives, avoiding analytical biases caused by missing data dimensions. Furthermore, the complementarity of various data types allows subsequent event identification and source tracing to better reflect actual pollution scenarios, improving the objectivity and accuracy of the overall analysis process.
[0240] This study performs spatiotemporal alignment processing on multi-source heterogeneous data, including consistency verification and unified spatiotemporal mapping, to obtain multimodal feature vectors. First, consistency verification eliminates contradictory and anomalous invalid information, ensuring data validity and rationality. Then, unified spatiotemporal mapping of latitude and longitude with hourly timestamps resolves the heterogeneity of multi-source data in both spatial location and time dimension, bringing data from different sources and formats into a unified spatiotemporal system and achieving data normalization. Furthermore, transforming the processed data into multimodal feature vectors extracts the main features of various data types, achieving feature-level fusion. This simplifies subsequent data analysis complexity while preserving key information from multiple sources, significantly improving analytical efficiency and accuracy.
[0241] Based on meteorological correction factors determined by combining historical data from the same period and meteorological benchmark data, dynamic threshold calculations are performed on multimodal feature vectors to obtain preliminary event information. This achieves dynamic adaptation of pollutant judgment thresholds, overcoming the drawback of fixed thresholds being unable to match different meteorological conditions. The meteorological correction factors allow thresholds to be adjusted according to real-time meteorological conditions, making the criteria for judging high air quality values more closely aligned with the actual environment of the target area, effectively avoiding misjudgments or omissions caused by meteorological factors. Simultaneously, by using dynamic threshold calculations to identify abnormal air quality conditions and compile preliminary event information such as the event's occurrence time, affected stations, and preliminary suspected pollution source types, basic information on pollution events can be quickly identified.
[0242] Spatial clustering was used to conduct spatiotemporal correlation analysis on preliminary event information and obtain event source tracing results. First, spatial clustering aggregated discrete concentration anomaly points into anomaly clusters, clearly defining the concentrated impact area of pollution and avoiding source tracing bias caused by analyzing scattered anomaly points individually. Then, by combining diffusion direction with real-time meteorological data for reverse trajectory extrapolation, the location of pollution sources could be scientifically and accurately determined. Finally, by spatial matching with pollution source emission inventories, a list of potential emission sources was screened, achieving accurate tracing from pollution phenomena to pollution sources. Combining temporal diffusion changes with spatial concentration distribution allows pollution source tracing to move beyond information from a single monitoring point and analyze the overall pollution pattern, significantly improving the accuracy of pollution source identification.
[0243] By comparing the event tracing results with historical event records and generating deduplicated event records, duplicate event records are effectively eliminated, avoiding repeated scheduling and control of the same ongoing pollution event. This reduces the waste of human and material resources and improves the efficiency of pollution control. Simultaneously, for cases determined to be the same ongoing event, information such as its impact range, concentration change process, and duration is updated in a timely manner, enabling the system to form a complete and continuous record of the evolution of the pollution event, facilitating staff to grasp the development trend of the pollution event. New events are included in the historical record, achieving dynamic updating and improvement of event records.
[0244] Based on preset event classification rules, duplicate event records are categorized and scheduling results are generated. By classifying events into Level 1, Level 2, and Level 3 based on pollution concentration, impact range, and pollution source confidence, high-value air quality events of varying severity can be differentiated. This allows for accurate allocation of control resources, with more control efforts deployed for high-level events with severe pollution and wide impact, while appropriate control measures are implemented for low-level events with mild pollution. This avoids resource waste or inadequate control caused by a one-size-fits-all approach. Simultaneously, targeted scheduling content is developed based on event levels, clearly defining responsible areas and control requirements. Information dissemination and task feedback are also planned, ensuring that scheduling instructions are accurately and quickly transmitted to relevant responsible parties. A complete task feedback channel is established, achieving closed-loop control from instruction issuance to task execution and result feedback. This significantly improves the efficiency and effectiveness of control response to high-value air quality events, ensuring timely and effective handling of various pollution events, helping to rapidly reduce pollutant concentrations and improve regional air quality. Ultimately, accurate identification of high-value air quality events enables efficient pollution control.
[0245] The above text combined Figures 1 to 2 The method for identifying high air quality events provided in this application has been described in detail. The apparatus and equipment provided in this application will be described below with reference to the accompanying drawings.
[0246] This application also provides a high-air-quality event identification device, such as... Figure 3 As shown in the figure, this is a schematic diagram of an air quality high-value event identification device provided in an embodiment of this application. The device includes: The acquisition module 301 is used to acquire multi-source heterogeneous data of the target area; wherein, the multi-source heterogeneous data includes at least two of the following: air quality monitoring data, meteorological data, pollution source emission inventory data, video monitoring data, traffic data, electricity consumption data, radar data, and pollutant component data; The data processing module 302 is used to perform spatiotemporal alignment processing on multi-source heterogeneous data to obtain multimodal feature vectors. The spatiotemporal alignment processing includes consistency verification of aligned data and unified spatiotemporal mapping of unaligned data based on latitude, longitude, and hourly timestamps. Based on meteorological correction factors, dynamic threshold calculations are performed on the multimodal feature vectors to obtain preliminary event information. The meteorological correction factors are determined based on the correlation between historical pollutant concentration data and meteorological baseline data for the target area. Through spatial clustering, spatiotemporal correlation analysis is performed on the preliminary event information to obtain event tracing results. The event tracing results are compared with historical event records to generate deduplicated event records. The scheduling module 303 is used to classify deduplicated event records into levels according to preset event classification rules and generate scheduling results.
[0247] In some possible implementations, the data processing module 302 is specifically used for: Based on preliminary event information, the corresponding air quality concentration anomaly locations are identified; spatial clustering of the concentration anomaly locations is performed to obtain at least one anomalous cluster; based on the spatial location of the anomalous cluster and real-time wind direction, the reverse trajectory of pollutants is extrapolated to determine the location of the pollution source; the location of the pollution source is spatially matched with the pollution source emission inventory to obtain a list of potential emission sources; the location of the pollution source and the list of potential emission sources are used as the event source tracing results.
[0248] In some possible implementations, the data processing module 302 is specifically used for: Each event in the event tracing results is taken as the current event, and it is determined whether the pollutant type of the current event is consistent with that of the historically identified events. If the pollutant type of the current event is consistent with that of the historically identified events, the spatiotemporal overlap and directional similarity between the current event and the historically identified events are calculated. If the spatiotemporal overlap is greater than a first preset threshold and the directional similarity is greater than a second preset threshold, the current event and the historically identified events are determined to be the same continuous event, and the influence range, concentration change process, evolution state and duration of the historically identified events are updated. Otherwise, the current event is added as a new event and stored in the historical event record to obtain a deduplicated event record.
[0249] Among some possible implementations, the high-air-quality event identification device also includes: The correction module is used to perform multimodal joint verification of the identification results of video monitoring data, the congestion index of traffic data, the load anomaly characteristics of electricity data, and the particulate matter echo characteristics of radar data when the preliminary suspected pollution source type in the preliminary event information is any one of open burning, traffic, or industrial emission sources. This allows for confidence correction and level update of the preliminary suspected pollution source type, resulting in corrected preliminary event information. Based on the corrected preliminary event information, spatiotemporal correlation analysis is performed to obtain the event source tracing results.
[0250] In some possible implementations, the data processing module 302 is specifically used for: Air quality concentration features of the target area are extracted from the multimodal feature vector; the basic threshold of pollutants is corrected using meteorological correction factors to obtain a dynamic discrimination threshold; the air quality concentration features are compared with the dynamic discrimination threshold, and preliminary event information is generated based on the comparison results.
[0251] In some possible implementations, the data processing module 302 is specifically used for: Based on latitude, longitude, and hourly timestamps, a unified spatiotemporal grid for the target area is constructed. Multi-source heterogeneous data are processed through at least one of the following methods: interpolation, spatial mapping, temporal aggregation, or inversion, and mapped to the unified spatiotemporal grid to form modal features within the unified spatiotemporal grid. The modal features include air quality concentration features, meteorological features, video smoke features, traffic density features, electricity load features, radar echo features, and pollutant composition features. The modal features within the unified spatiotemporal grid are then fused at the feature level to obtain a multimodal feature vector.
[0252] The air quality high-value event identification device according to the embodiments of this application can correspond to the execution of the method described in the embodiments of this application, and the other operations and / or functions of each module / unit of the air quality high-value event identification device are respectively for implementing Figure 2 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.
[0253] This application also provides a computing device. For example... Figure 4 As shown in the figure, this is a schematic diagram of a computing device provided in an embodiment of this application. The computing device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other via the bus 401.
[0254] Bus 401 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0255] Processor 402 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0256] Communication interface 403 is used for communication with external devices.
[0257] Memory 404 may include volatile memory, such as random access memory (RAM). Memory 404 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0258] The memory 404 stores executable code, and the processor 402 executes the executable code to perform the aforementioned high air quality event identification method.
[0259] Specifically, in achieving Figure 3 In the case of the illustrated embodiment, and Figure 3 When the modules or units of the high-quality event identification device described in the embodiment are implemented by software, the following steps are performed: Figure 3 The software or program code required for the functions of each module / unit can be partially or entirely stored in memory 404. Processor 402 executes the program code corresponding to each unit stored in memory 404 to execute the aforementioned high air quality event identification method.
[0260] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to execute the aforementioned high-air-quality event identification method.
[0261] This application also provides a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.
[0262] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0263] When the computer program product is executed by a computer, the computer performs any of the aforementioned methods for identifying high air quality events. The computer program product can be a software installation package; when any of the aforementioned methods for identifying high air quality events needs to be used, the computer program product can be downloaded and executed on the computer.
[0264] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0265] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the scope of protection of this application.
Claims
1. A method for identifying high air quality events, characterized in that, The method includes: Acquire multi-source heterogeneous data of the target area; wherein, the multi-source heterogeneous data includes at least two of the following: air quality monitoring data, meteorological data, pollution source emission inventory data, video monitoring data, traffic data, electricity consumption data, radar data, and pollutant component data; The multi-source heterogeneous data is subjected to spatiotemporal alignment processing to obtain multimodal feature vectors; wherein, the spatiotemporal alignment processing includes consistency verification of aligned data and spatiotemporal unified mapping of unaligned data based on latitude and longitude and hourly timestamps; Based on the meteorological correction factor and the multimodal feature vector, dynamic threshold calculation is performed to obtain preliminary event information; Spatial clustering is used to perform spatiotemporal correlation analysis on the preliminary event information to obtain event tracing results; The event tracing results are compared with historical event records to generate duplicate event records; According to the preset event classification rules, the deduplicated event records are classified into levels, and scheduling results are generated; The meteorological correction factor is obtained in the following way: in, For pollutant type p in spatial grid cells Meteorological correction factor at time t; Pollutant types from the same historical period In spatial grid units Concentration at time t; spatiotemporal grid unit Meteorological feature vector at time t; This represents the conditional expected concentration under the same meteorological conditions. This represents the overall average concentration for the same period in history. A dynamic threshold is generated during the dynamic threshold calculation process. The dynamic threshold is obtained in the following way: in, For dynamic threshold determination; Pollutant type The basic threshold; This is an adjustable sensitivity coefficient.
2. The method according to claim 1, characterized in that, The process of performing spatiotemporal correlation analysis on the preliminary event information through spatial clustering to obtain event tracing results includes: Based on the preliminary event information, the corresponding locations of abnormal air quality concentrations were determined; Spatial clustering of air quality concentration anomalies yields at least one anomalous cluster; Based on the spatial location of the abnormal clusters and the real-time wind direction, the reverse trajectory of the pollutants is extrapolated to determine the location of the pollution source; Spatially match the locations of the pollution sources with the pollution source emission inventory to obtain a list of potential emission sources. The location of the pollution source and the list of potential emission sources are used as the results of the event tracing.
3. The method according to claim 1, characterized in that, The step of comparing the event tracing results with historical event records to generate deduplicated event records includes: Each event in the event tracing results is taken as the current event, and it is determined whether the pollutant type of the current event is consistent with that of the previously identified events in the historical event records; If the pollutant type of the current event is consistent with that of a previously identified historical event, then the spatiotemporal overlap and directional similarity between the current event and the previously identified historical event are calculated. If the spatiotemporal overlap is greater than the first preset threshold and the directional similarity is greater than the second preset threshold, then the current event and the historically identified events are determined to be the same continuous event, and the influence range, concentration change process, evolution state and duration of the historically identified events are updated; otherwise, the current event is stored as a new event in the historical event record to obtain a deduplicated event record.
4. The method according to claim 1, characterized in that, The method further includes: When the preliminary suspected pollution source type in the preliminary event information is any one of open burning, traffic, or industrial emission sources, the identification results of video monitoring data, the congestion index of traffic data, the load anomaly characteristics of electricity data, and the particulate matter echo characteristics of radar data are jointly verified in a multimodal manner to correct the confidence level and update the grade of the preliminary suspected pollution source type, thus obtaining the corrected preliminary event information. Spatial clustering is used to perform spatiotemporal correlation analysis on the corrected preliminary event information to obtain the event tracing results.
5. The method according to claim 1, characterized in that, The preliminary event information is obtained by dynamically thresholding the multimodal feature vector based on the meteorological correction factor, including: Extract the air quality concentration features of the target area from the multimodal feature vector; By using meteorological correction factors to correct the basic thresholds of pollutants, dynamic discrimination thresholds are obtained. The air quality concentration characteristics are compared with the dynamic discrimination threshold, and preliminary event information is generated based on the comparison results.
6. The method according to claim 1, characterized in that, The process of performing spatiotemporal alignment on the multi-source heterogeneous data to obtain multimodal feature vectors includes: A unified spatiotemporal grid for the target area is constructed based on latitude and longitude and hourly timestamps; Multi-source heterogeneous data is processed through at least one of interpolation, spatial mapping, temporal aggregation, or inversion, and the processing results are mapped to the unified spatiotemporal grid to form modal features within the unified spatiotemporal grid; wherein, the modal features include air quality concentration features, meteorological features, video smoke features, traffic flow density features, electricity load features, radar echo features, and pollutant composition features; The modal features within the unified spatiotemporal grid are fused at the feature level to obtain a multimodal feature vector.
7. A device for identifying high air quality events, characterized in that, The device includes: The acquisition module is used to acquire multi-source heterogeneous data of the target area; wherein, the multi-source heterogeneous data includes at least two of the following: air quality monitoring data, meteorological data, pollution source emission inventory data, video monitoring data, traffic data, electricity consumption data, radar data, and pollutant component data; The data processing module is used to perform spatiotemporal alignment processing on the multi-source heterogeneous data to obtain multimodal feature vectors. The spatiotemporal alignment processing includes consistency verification of aligned data and unified spatiotemporal mapping of unaligned data based on latitude, longitude, and hourly timestamps. Based on meteorological correction factors and the multimodal feature vectors, dynamic threshold calculation is performed to obtain preliminary event information. Through spatial clustering, spatiotemporal correlation analysis is performed on the preliminary event information to obtain event tracing results. The event tracing results are compared with historical event records to generate deduplicated event records. The scheduling module is used to classify the deduplicated event records into levels according to preset event classification rules and generate scheduling results. The data processing module is used to calculate meteorological correction factors and perform dynamic threshold calculations. The meteorological correction factors are obtained in the following ways: in, For pollutant type p in spatial grid cells Meteorological correction factor at time t; Pollutant types from the same historical period In spatial grid units Concentration at time t; spatiotemporal grid unit Meteorological feature vector at time t; This represents the conditional expected concentration under the same meteorological conditions. This represents the overall average concentration for the same period in history. A dynamic threshold is generated during the dynamic threshold calculation process. The dynamic threshold is obtained in the following way: in, For dynamic threshold determination; Pollutant type The basic threshold; This is an adjustable sensitivity coefficient.
8. A computing device, characterized in that, Including memory and processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes one or more computer instructions, which, when executed by a computer, perform the method as described in any one of claims 1 to 6.