A method and system for identifying and classifying risk points of air pollution caused by coal consumption

By using Transformer model and POI data, a spatial distribution prediction model for air pollutants is constructed, which solves the problem of identifying and classifying the risk points of air pollution caused by coal consumption in the prior art, and achieves high accuracy and efficient pollutant distribution prediction and risk point identification.

CN119760613BActive Publication Date: 2025-06-17SHANDONG UNIV

Patent Information

Application Number
CN202510252141.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify and classify the risk points of air pollution caused by coal consumption, and there are problems such as strong dependence on data quality, incomplete time and space coverage, and difficulty in reflecting geospatial information.

Method used

The Transformer model is used to construct a relationship model between remote sensing data and pollutants, generate a spatial distribution data set of atmospheric pollutants, and combine POI data to identify and classify the risk points of atmospheric pollution caused by coal consumption through deep learning models.

Benefits of technology

The spatial distribution prediction of atmospheric pollutants with continuous spatial and spatial coverage and high accuracy is achieved, the calculation efficiency, accuracy and accuracy of the model are improved, and the risk points of atmospheric pollution caused by coal consumption are accurately identified, and the prevention and control of heavy pollution weather and scientific emission reduction are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760613B_ABST
    Figure CN119760613B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of analysis of the causes of air pollution, and discloses a method and system for identifying and classifying risk points of air pollution caused by coal consumption, including obtaining multi-source data; extracting the features of the multi-source data to form data feature samples, using the ground monitoring data at the current moment of the grid point where the data feature sample is located as a label to train a deep learning model; using the trained model to predict the surface concentration of air pollutants; based on the prediction results, calculating the mean and standard deviation of the pollutant concentration of each grid point, screening out risk points; establishing a buffer zone centered on the risk points, analyzing the distribution characteristics of POIs in the buffer zone, and identifying and classifying pollution sources. The present invention uses POI points to achieve accurate identification of risk points of air pollution caused by coal consumption, provides data support for the prevention and control of heavy pollution weather, and promotes scientific emission reduction and refined governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of analysis of the causes of air pollution, and particularly to a method and system for identifying and classifying risk points of air pollution caused by coal consumption. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] The industrial structure dominated by heavy chemical industries, the energy structure dominated by coal, and the transportation structure dominated by road freight have led to a high-risk situation of frequent air pollution. In the current and long-term future, fossil energy dominated by coal will still occupy the dominant position in the energy structure. The air pollution problems caused by coal consumption activities cannot be ignored. For example, during the process of coal mining and utilization, a large amount of air pollutants such as SO2, particulate matter, soot, and dust are emitted. These pollutants will not only cause direct damage to the human respiratory system but also trigger disastrous climates such as the greenhouse effect, acid rain, photochemical smog, and haze, threatening the natural ecology and public health. Therefore, high attention should be paid to the air pollution caused by coal consumption activities. However, the concentration of air pollutants is affected by multiple factors, and it is difficult to screen out the air pollution caused by coal consumption from them. And air pollution source tracing is the basis for controlling air pollution caused by coal consumption and formulating coal emission reduction policies. It is urgent to identify the risk points of air pollution caused by coal consumption through technical means to provide means support for air pollution source tracing, control, etc.

[0004] Current methods for analyzing the causes of air pollution mainly include emission inventory method, statistical analysis method, chemical transport model method, and receptor model method, etc. The emission inventory method comprehensively reflects the emission characteristics and source composition of regional pollutants by systematically collecting and organizing pollution source emission data and using economic activity models and emission factors, etc. However, this method has limitations such as difficult data acquisition and long update cycle, and it is difficult to reflect the dynamic changes of short-term pollution. Moreover, its spatial resolution is often limited to the urban scale. The statistical analysis method identifies the causes of pollution by using statistical regression and other methods based on the spatio-temporal distribution characteristics of pollutant concentration data and the Air Quality Index (AQI). This method is relatively simple to operate and requires less data, making it suitable for preliminary pollution characteristic analysis. However, since the air pollution process often involves complex non-linear relationships and multiple interactions, simple statistical regression methods are difficult to accurately depict these complex relationships. At the same time, due to relying on ground monitoring station data, its spatio-temporal coverage is significantly limited. The receptor model method is based on environmental monitoring data and inversely calculates the contributions of pollution sources through mathematical methods such as chemical mass balance and factor analysis. This method performs well in identifying local pollution sources, but its accuracy depends to a large extent on the quality and representativeness of the monitoring data. The chemical transport model method can quantitatively evaluate the contributions of different sources by numerically simulating the physical and chemical processes of air pollutant emission, transport, and transformation. This method has a clear physical and chemical mechanism, but the model accuracy is not high, it cannot use existing ground observation information to further improve the model accuracy, requires a large amount of computing resources, and the result accuracy is affected by various input data and parameters. In recent years, researchers have proposed analysis methods using new monitoring means such as characteristic radar maps and lidar maps, which identify the causes of pollution according to the change characteristics of different pollutants in time series and space, and can achieve the analysis of the causes of air pollution at long time series or regional scales. However, it relies on station monitoring data and is difficult to accurately reflect the spatial distribution characteristics of pollution. Generally speaking, traditional methods for analyzing the causes of air pollution have disadvantages such as difficulty in obtaining complete data, strong dependence on data quality, incomplete spatio-temporal coverage, and difficulty in reflecting geographical spatial information. The air pollution caused by coal consumption has dual characteristics of point sources and area sources, and its spatial distribution shows complex diversity. Especially for the dispersed emissions such as residential coal burning, due to the randomness and uncertainty of its spatio-temporal distribution, it is difficult to effectively capture through fixed monitoring stations. This difficulty in identifying pollution sources greatly increases the costs of environmental supervision and law enforcement. At present, there is still a lack of a scientific and systematic method to accurately identify and delimit the high-value areas of air pollution caused by coal consumption, which poses challenges to the precise implementation and effective promotion of pollution prevention and control work.

[0005] With the development of satellite remote sensing technology, the cost of obtaining large-scale, long-time, and spatially continuous remote sensing observation data of pollutants is relatively low. With its powerful nonlinear fitting ability, deep learning algorithms can achieve the remote sensing inversion of pollutant concentrations, which has become an important way to obtain the complete distribution of pollutant concentrations in time and space and realize the analysis of the causes of atmospheric pollution with high temporal and spatial resolution. In addition, the widespread application of POI (Point of Interest) data provides a new idea for the identification of pollution sources. POI data contains rich attribute information such as location name, category, address, etc., which can reflect the characteristics of human activities and land use types in a specific area. Combining POI data with atmospheric pollutant distribution data with complete temporal and spatial coverage, according to the combined characteristics of different pollutant concentrations, a pollution source-receptor relationship based on geographical location can be established, and accurate identification and classification of atmospheric pollution risk points caused by coal consumption can be achieved. This method not only overcomes the limitations of traditional methods in data acquisition and computational efficiency, but also provides higher temporal and spatial resolution, providing more powerful technical support for the analysis of the causes of atmospheric pollution and the prevention and control of severe pollution. Summary of the invention

[0006] In order to solve the above problems, the present invention proposes a method and system for identifying and classifying air pollution risk points caused by coal consumption. The Transformer model is used to construct a relationship model between remote sensing data and pollutants, and a spatial distribution data set of air pollutants is generated. The buffer zone is calculated for the identified risk points according to the combination relationship of different pollutant concentrations. POI points are used to accurately identify air pollution risk points caused by coal consumption, provide data support for the prevention and control of heavy pollution weather, and promote scientific emission reduction and refined governance.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a method for identifying and classifying air pollution risk points caused by coal consumption, comprising the following steps:

[0009] Acquire multi-source data and pre-process them to obtain a multi-source dataset with consistent temporal and spatial resolution;

[0010] Feature extraction is performed on multi-source data sets to form data feature samples, and the ground monitoring data at the current moment of the grid point where the data feature samples are located is used as a label to train the deep learning model;

[0011] Use the trained deep learning model to predict the surface concentration of air pollutants and generate a data set of the spatial distribution of air pollutants;

[0012] Based on the spatially distributed data of air pollutants obtained from model predictions, calculate the mean and standard deviation of pollutant concentrations at each grid point on a time scale, and screen out the grid points that meet the set conditions as the air pollution risk points caused by coal consumption;

[0013] Taking the screened risk points as the center, establish a buffer zone, analyze the distribution characteristics of POIs within the buffer zone, calculate the density and proportion of different types of POIs, and identify and classify the pollution sources.

[0014] As an alternative implementation, the multi-source data includes meteorological reanalysis data, ground monitoring data of air pollutants, remote sensing column concentration data of air pollutants, digital elevation data, and land use data.

[0015] As an alternative implementation, the air pollutants for which the surface concentration is predicted using a deep learning model include SO2, PM 2.5 , O3, NO2, and CO.

[0016] As an alternative implementation, screening out the grid points that meet the set conditions specifically means:

[0017] Screen out the grid points where the SO2 concentration is higher than the mean and exceeds the mean plus one standard deviation, the CO concentration is higher than the mean, and the concentrations of PM 2.5 , O3, and NO2 are between the mean plus one standard deviation and the mean minus one standard deviation.

[0018] As an alternative implementation, the pollution source identification and classification criteria are specifically:

[0019] When the proportion of industrial POIs in the buffer zone is higher than other types, determine this risk point as an industrial source; when urban domestic POIs dominate, then determine this risk point as a domestic source.

[0020] As an alternative implementation, the deep learning model consists of an input layer, a Transformer module based on spatial attention, a feature mapping layer, a position encoding layer, six hidden layers, and an output layer. Among them, the hidden layers are composed of multi-head attention Transformer encoders, which are used to learn and extract the spatial correlation information between features. The output layer is composed of two fully connected layers, and the output units are five, corresponding to the concentrations of five air pollutants respectively.

[0021] In a second aspect, the present invention provides a system for identifying and classifying air pollution risk points caused by coal consumption, including:

[0022] A data acquisition module, configured to: acquire multi-source data and perform preprocessing to obtain a multi-source data set with consistent spatio-temporal resolution;

[0023] A model training module, configured to: extract features from a multi-source data set to form data feature samples, and use the ground monitoring data at the current moment of the grid point where the data feature sample is located as a label to train a deep learning model;

[0024] A model prediction module, configured to: use the trained deep learning model to predict the surface concentration of air pollutants and generate an air pollutant spatial distribution data set;

[0025] A risk point identification module, configured to: based on the air pollutant spatial distribution data obtained by model prediction, calculate the mean and standard deviation of the pollutant concentration at each grid point according to a time scale, and screen out the grid points that meet the set conditions as the air pollution risk points caused by coal consumption;

[0026] An identification and classification module, configured to: establish a buffer zone centered on the screened risk points, analyze the POI distribution characteristics in the buffer zone, calculate the density and proportion of different types of POIs, and identify and classify the pollution sources.

[0027] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0028] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0029] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0030] Compared with the prior art, the beneficial effects of the present invention are:

[0031] The present invention provides a method and system for identifying and classifying risk points of air pollution caused by coal consumption. By using surface monitoring data of air pollutants, remote sensing data, and meteorological reanalysis data, a deep learning model is used to learn the non-linear relationship between surface monitoring data and remote sensing data, realizing the prediction of the spatial distribution of air pollutants with continuous space, complete spatio-temporal coverage, and high accuracy. The calculation efficiency, precision, and accuracy of the model are improved. The monthly average value is calculated based on the predicted spatial distribution of pollutant concentrations, and risk points of air pollution caused by coal consumption are screened according to the combined characteristics of different pollutant concentrations. Further classification is carried out using POI information, making up for the shortcomings of traditional pollution identification methods caused by coal consumption, such as strong dependence on data quality, incomplete spatio-temporal coverage, and difficulty in reflecting geographical spatial information. The present invention uses POI points to accurately identify risk points of air pollution caused by coal consumption, providing data support for the prevention and control of heavy pollution weather, and promoting scientific emission reduction and refined governance. It meets the requirements of accuracy and practicality in specific applications, and provides an effective solution for the field of air pollution cause analysis.

[0032] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0034] Figure 1 It is a flowchart of a method for identifying and classifying risk points of air pollution caused by coal consumption provided in Embodiment 1 of the present invention;

[0035] Figure 2 It is a network structure diagram of the Transformer deep learning model provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0037] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0038] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0039] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0040] Embodiment 1

[0041] As Figure 1 shown, this embodiment provides a method for identifying and classifying risk points of air pollution caused by coal consumption, including the following steps:

[0042] S1. Obtain multi-source data and perform preprocessing to obtain a multi-source data set with consistent spatio-temporal resolution.

[0043] S2. Extract features from the multi-source data set to form data feature samples, and use the ground monitoring data at the current moment of the grid points where the data feature samples are located as labels to train the deep learning model.

[0044] S3. Use the trained deep learning model to predict the surface concentration of air pollutants and generate a spatial distribution data set of air pollutants.

[0045] S4. Based on the spatial distribution data of air pollutants predicted by the model, calculate the mean and standard deviation of the pollutant concentration at each grid point according to the time scale, and screen out the grid points that meet the set conditions as the risk points of air pollution caused by coal consumption.

[0046] S5. Establish a buffer zone with the screened risk points as the center, analyze the distribution characteristics of POIs in the buffer zone, calculate the density and proportion of different types of POIs, and identify and classify the pollution sources.

[0047] The multi-source data includes meteorological reanalysis data, ground monitoring data of air pollutants, remote sensing column concentration data of air pollutants, digital elevation data (DEM), and land use data.

[0048] Preprocess multi-source data: Unify the spatio-temporal resolutions of meteorological data, remote sensing data of atmospheric pollutants, DEM, and land use data through methods such as removing missing values and outliers, resampling, and interpolation. Specifically, create a standard grid with a spatial resolution of 1km * 1km, stack it into a two-dimensional array, and construct a list of geographical coordinates for all grid points. At the same time, extract feature information in the time and space dimensions through feature engineering to form a unified data format.

[0049] Create a KD tree (K-Dimensional Tree) based on the grid longitude and latitude information of this dataset to find the grid points closest to the atmospheric pollution monitoring stations, map the surface pollutant data of the stations to the grid, and generate a mask recording the station data to distinguish the grid points with measured data and those that need to be predicted. Finally, generate a training dataset containing meteorological data, surface monitoring station data of atmospheric pollutants, remote sensing data of atmospheric pollutants, DEM, and land use data, and a prediction dataset containing meteorological data, remote sensing data of atmospheric pollutants, DEM, and land use data.

[0050] During the feature generation process, first stack all environmental variables to form a basic feature layer (the basic feature layer is formed by merging and stacking all input variables while retaining the grid row and column relationships), use the observed values of the surface concentrations of atmospheric pollutants at the stations mapped to the grid and the station mask as additional feature layers, and finally use the calculated spatial encoding as an auxiliary feature layer. Concatenate these feature layers in the last dimension to form a complete input feature tensor, and use the surface monitoring data of atmospheric pollutant concentrations corresponding to the current moment as labels to train the deep learning model. Use the training dataset to train and validate the model, and the prediction dataset is used for prediction.

[0051] The atmospheric pollutants for which surface concentration predictions are made using the deep learning model include SO2, PM 2.5 , O3, NO2, and CO.

[0052] Use the trained deep learning model to predict the surface concentrations of SO2, PM 2.5 , O3, NO2, and CO to generate a spatial distribution dataset of atmospheric pollutants. Based on the spatial distribution data of atmospheric pollutants predicted by the model, calculate the mean and standard deviation of the pollutant concentrations at each grid point on a time scale, and select the grid points where the SO2 concentration is significantly higher than the mean and exceeds the mean plus one standard deviation, the CO concentration is higher than the mean, and the concentrations of the other three pollutants are within the interval between the mean plus one standard deviation and the mean minus one standard deviation, and identify them as the atmospheric pollution risk points caused by coal consumption.

[0053] To further clarify the pollution sources, a buffer zone with a radius of 1 km was established centered on the selected risk points, and the distribution characteristics of POIs within the buffer zone were analyzed. By calculating the density and proportion of different types of POIs, a pollution source discrimination criterion was established: when the proportion of industrial POIs in the buffer zone is significantly higher than other types, the risk point is determined as an industrial source; when urban living POIs such as residential areas and commercial areas dominate, it is determined as a living source.

[0054] Analyze the distribution characteristics of POIs in the buffer zone. By screening out POI points related to pollution within the defined buffer zone and classifying them according to their source information, the distribution characteristics of POIs are analyzed. The specific method is as follows:

[0055] 1. POI density calculation method:

[0056] Divide the number of relevant POIs in the buffer zone by the buffer zone area to obtain the POI density. The calculation formula is:

[0057] ;

[0058] This density value can be used as a reference for judgment to evaluate the possibility of high-value points in the buffer zone. The density index can be used as a reference for the degree of spatial distribution concentration. A high-density area usually indicates a strong spatial aggregation of this type of POI and can be considered as a potential area for high-value points. However, areas with low density still need to be comprehensively judged in combination with other factors and cannot be directly excluded.

[0059] 2. POI proportion calculation method:

[0060] Divide the number of POIs of different types by the total number of POIs to obtain the proportion of each type of source. The calculation formula is:

[0061]

[0062] According to the proportion results, select the source type with the highest proportion as the dominant emission source in this buffer zone.

[0063] Optionally, the meteorological reanalysis data includes: 10m wind speed, 2m dew point temperature, 2m temperature, surface pressure, total precipitation, boundary layer height, surface net solar radiation, total cloud cover, etc.; the remote sensing data includes: SO2 remote sensing column concentration data, atmospheric aerosol optical depth (AOD), O3 remote sensing column concentration data, NO2 remote sensing column concentration data, CO remote sensing column concentration data, and the surface monitoring data of atmospheric pollutants includes: SO2 surface concentration monitoring data, O3 surface concentration monitoring data, NO2 surface concentration monitoring data, CO surface concentration monitoring data, PM 2.5 Surface concentration monitoring data.

[0064] Optionally, the time and space information extracted by the feature engineering includes: spatial location information, Julian day, and the calculation formula for the spatial location information is:

[0065] ;

[0066] In the formula, is the longitude, is the latitude, is the spatial location information.

[0067] Optionally, the pre-trained deep learning model is a Transformer regression prediction neural network based on the self-attention mechanism, which is used for the spatial prediction of pollutant concentration.

[0068] Optionally, as Figure 2 shown, the pre-trained deep learning model consists of an input layer, a Transformer module based on spatial attention, a feature mapping layer, a position encoding layer, six hidden layers, and an output layer. The hidden layer is composed of an 8-head attention Transformer encoder, which is used to learn and extract the spatial correlation information between features; the output layer is composed of 2 fully connected layers, and the output units are 5, corresponding to the concentrations of 5 kinds of atmospheric pollutants respectively.

[0069] Optionally, the Transformer module based on spatial attention calculates the average value and maximum value of the channel dimension of the feature map, and after connection, uses a 7×7 convolution kernel and a sigmoid function to generate spatial attention weights to highlight the key spatial regions.

[0070] Optionally, the feature mapping layer of the model normalizes the input features through linear transformation and LayerNorm; the position encoding layer uses sine and cosine functions to generate position-aware information.

[0071] Optionally, during the training process of the model, dropout = 0.1 is used for regularization to prevent overfitting, and the learning rate is dynamically adjusted by ReduceLROnPlateau. When the validation loss does not decrease within 5 consecutive epochs, the learning rate is automatically reduced. In addition, the model supports DistributedDataParallel to achieve distributed training and uses mixed-precision training (AMP) to improve the training efficiency.

[0072] Optionally, the model uses the GELU function as the activation function:

[0073] ;

[0074] The loss is trained using the mean squared error loss function (MSE):

[0075] ;

[0076] Wherein, is the model predicted value, is the true value, is the sample size.

[0077] Optionally, the pollutant concentration obtained by model prediction is evaluated using the coefficient of determination , and the calculation formula of the coefficient of determination is:

[0078] ;

[0079] Wherein, is the number of two sets, is the true value, is the predicted value, is the mean of the true values, is the mean of the predicted values, is the standard deviation of the true values, is the standard deviation of the predicted values.

[0080] Optionally, the POI information calls the POI data interface of Amap to obtain the POI information in the buffer area, mainly including: industrial POIs such as industrial parks and factories, and urban life POIs such as residential areas and commercial areas.

[0081] Embodiment 2

[0082] This embodiment provides a system for identifying and classifying risk points of air pollution caused by coal consumption, including:

[0083] A data acquisition module, configured to: acquire multi-source data and perform preprocessing to obtain a multi-source data set with consistent spatio-temporal resolution;

[0084] A model training module, configured to: extract features from the multi-source data set to form data feature samples, and use the ground monitoring data at the current moment of the grid points where the data feature samples are located as labels to train the deep learning model;

[0085] A model prediction module, configured to: use the trained deep learning model to predict the surface concentration of air pollutants and generate a spatial distribution data set of air pollutants;

[0086] A risk point identification module, configured to: based on the spatial distribution data of air pollutants obtained by model prediction, calculate the mean and standard deviation of the pollutant concentration at each grid point according to the time scale, and screen out the grid points that meet the set conditions as the risk points of air pollution caused by coal consumption;

[0087] An identification and classification module, configured to: establish a buffer centered on the screened risk points, analyze the distribution characteristics of POIs in the buffer, calculate the density and proportion of different types of POIs, and identify and classify pollution sources.

[0088] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0089] In more embodiments, there is also provided:

[0090] An electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0091] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0093] A computer-readable storage medium, used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0094] The method in Embodiment 1 can be directly implemented by a hardware processor to complete, or implemented by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0095] A computer program product, including a computer program, which when executed by a processor, implements the method described in Embodiment 1.

[0096] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided among program modules as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.

[0097] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0098] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier such that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, sound, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0100] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for identifying and classifying air pollution risk points caused by coal consumption, characterized in that: The following steps are involved: Acquire multi-source data and pre-process them to obtain a multi-source dataset with consistent temporal and spatial resolution; Feature extraction is performed on multi-source data sets to form data feature samples, and the ground monitoring data at the current moment of the grid point where the data feature samples are located is used as a label to train the deep learning model; Use the trained deep learning model to predict the surface concentration of air pollutants and generate a data set of the spatial distribution of air pollutants; Based on the spatial distribution data of air pollutants predicted by the model, the mean and standard deviation of pollutant concentrations at each grid point are calculated according to the time scale, and grid points that meet the set conditions are selected as air pollution risk points caused by coal consumption; Establish a buffer zone with the selected risk points as the center, analyze the POI distribution characteristics within the buffer zone, calculate the density and proportion of different types of POIs, and identify and classify pollution sources; Filter out the grid points that meet the set conditions, specifically: Filter out SO2 concentrations above the mean and exceeding the mean plus one standard deviation, CO concentrations above the mean, PM 2.5 , O3 and NO2 concentrations are between the mean plus one standard deviation and the mean minus one standard deviation; The pollution source identification and classification criteria are as follows: When the proportion of industrial POIs in the buffer zone is higher than that of other types, the risk point is determined as an industrial source; when urban life POIs dominate, the risk point is determined as a life source; The deep learning model consists of an input layer, a Transformer module based on spatial attention, a feature mapping layer, a position encoding layer, six hidden layers and an output layer, wherein the hidden layer is composed of a multi-head attention Transformer encoder for learning to extract spatial correlation information between features, and the output layer is composed of two fully connected layers with five output units, corresponding to five concentrations of air pollutants; The Transformer module based on spatial attention calculates the average and maximum values ​​of the channel dimensions of the feature map, and then generates spatial attention weights using a 7×7 convolution kernel and a sigmoid function to highlight the key spatial areas. The feature mapping layer of the model normalizes the input features through linear transformation and LayerNorm; The position encoding layer generates position-aware information using sine and cosine functions; The model uses dropout=0.1 for regularization during training to prevent overfitting, and dynamically adjusts the learning rate through ReduceLROnPlateau. When the verification loss does not decrease within 5 consecutive epochs, the learning rate is automatically reduced. In addition, the model supports DistributedDataParallel to implement distributed training and uses mixed precision training to improve training efficiency.

2. A method for identifying and classifying air pollution risk points caused by coal consumption as claimed in claim 1, characterized in that: The multi-source data include meteorological reanalysis data, ground monitoring data of atmospheric pollutants, remote sensing column concentration data of atmospheric pollutants, digital elevation data and land use data.

3. A method for identifying and classifying air pollution risk points caused by coal consumption as claimed in claim 1, characterized in that: Atmospheric pollutants that use deep learning models to predict surface concentrations include SO2, PM 2.5 , O3, NO2 and CO.

4. A system for identifying and classifying air pollution risk points caused by coal consumption, characterized in that: include: The data acquisition module is configured to: acquire multi-source data and perform preprocessing to obtain a multi-source data set with consistent temporal and spatial resolution; The model training module is configured to: extract features from multi-source data sets to form data feature samples, and use the current ground monitoring data of the grid points where the data feature samples are located as labels to train the deep learning model; The model prediction module is configured to: use the trained deep learning model to predict the surface concentration of atmospheric pollutants and generate a spatial distribution dataset of atmospheric pollutants; The risk point identification module is configured to: calculate the mean and standard deviation of pollutant concentrations at each grid point according to the time scale based on the spatial distribution data of air pollutants predicted by the model, and select grid points that meet the set conditions as air pollution risk points caused by coal consumption; The identification and classification module is configured to: establish a buffer zone with the selected risk points as the center, analyze the POI distribution characteristics in the buffer zone, calculate the density and proportion of different types of POIs, and identify and classify the pollution sources; Filter out the grid points that meet the set conditions, specifically: Filter out SO2 concentrations above the mean and exceeding the mean plus one standard deviation, CO concentrations above the mean, PM 2.5 , O3 and NO2 concentrations are between the mean plus one standard deviation and the mean minus one standard deviation; The pollution source identification and classification criteria are as follows: When the proportion of industrial POIs in the buffer zone is higher than that of other types, the risk point is determined as an industrial source; when urban life POIs dominate, the risk point is determined as a life source; The deep learning model consists of an input layer, a Transformer module based on spatial attention, a feature mapping layer, a position encoding layer, six hidden layers and an output layer, wherein the hidden layer is composed of a multi-head attention Transformer encoder for learning to extract spatial correlation information between features, and the output layer is composed of two fully connected layers with five output units, corresponding to five concentrations of air pollutants; The Transformer module based on spatial attention calculates the average and maximum values ​​of the channel dimensions of the feature map, and then generates spatial attention weights using a 7×7 convolution kernel and a sigmoid function to highlight the key spatial areas. The feature mapping layer of the model normalizes the input features through linear transformation and LayerNorm; The position encoding layer generates position-aware information using sine and cosine functions; The model uses dropout=0.1 for regularization during training to prevent overfitting, and dynamically adjusts the learning rate through ReduceLROnPlateau. When the verification loss does not decrease within 5 consecutive epochs, the learning rate is automatically reduced. In addition, the model supports DistributedDataParallel to implement distributed training and uses mixed precision training to improve training efficiency.

5. An electronic device, characterized in that: The invention comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 3 is completed.

6. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 3.

7. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Near-surface atmospheric pollutant inversion method and system based on remote sensing image

    CN114926749A

  • Remote sensing traceability method for urban ozone over-standard pollution

    CN115453069A

Cited By

  • Distributed storage and data quality treatment method for coal market big data

    CN122045282A