Urban pollutant distribution prediction method based on dynamic and static feature fusion of multi-source data

By fusing dynamic and static features of multi-source data, combined with remote sensing images and population density data, the problem of insufficient accuracy in predicting air pollutant concentration distribution in existing technologies is solved, the needs of high-resolution air quality management are met, and the prediction accuracy and robustness of air pollutant concentration distribution are improved.

CN120496663BActive Publication Date: 2025-10-14XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510980988.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-14
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

In the existing technologies for predicting the distribution of urban air pollutant concentrations, the kriging interpolation method and the inverse distance weighted interpolation method fail to effectively integrate multivariate auxiliary information and ignore the heterogeneity of the geographical environment and pollution source distribution, resulting in poor interpolation effects in complex areas and an inability to meet the needs of high-resolution and high-precision air quality management.

Method used

A method based on dynamic and static feature fusion of multi-source data is adopted. By utilizing easily collected multi-source public data and combining remote sensing images, population density and altitude information, the concentration of air pollutants without monitoring stations is inferred through feature extraction and feature fusion, and the model parameters are optimized to improve prediction accuracy.

Benefits of technology

It achieves high temporal and spatial resolution monitoring of air pollutants, effectively solves the problem of predicting the concentration distribution of air pollutants in areas without monitoring stations, and enhances the robustness and prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496663B_ABST
    Figure CN120496663B_ABST
Patent Text Reader

Abstract

The application discloses a city pollutant distribution prediction method based on dynamic and static feature fusion of multi-source data, and comprises the following steps: S1, collecting multi-source heterogeneous data in a research area, and constructing a multi-source heterogeneous data set; S2, preprocessing the collected multi-source heterogeneous data; S3, processing remote sensing image blocks, population density information and altitude information by using a block feature extraction module to obtain regional long-term static features; S4, processing part of stations around the calculation target site by using an air data reference module to obtain air reference dynamic features of pollutant distribution from surrounding air monitoring stations; S5, performing feature fusion on the block long-term static features and the air reference dynamic features of the target to be predicted, and calculating the air pollutant concentration of the prediction target; and S6, optimizing the model to optimal parameters; the method uses easily collected multi-source public data, fuses dynamic and static features, and effectively solves the air pollutant concentration distribution prediction of regions without monitoring stations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of temporal and spatial distribution calculation of atmospheric pollutants, and specifically relates to a method for predicting urban pollutant distribution based on the fusion of dynamic and static features of multi-source data. Background Art

[0002] The monitoring and early warning of the spatiotemporal distribution of urban air pollutants is a core issue in the field of environmental governance. Air quality is closely related to people's lives. If the air contains fine particulate matter (PM 2.5 Excessive increases in the concentrations of various pollutants, such as nitrogen dioxide (NO2), ozone (O3), and others, can easily lead to a surge in respiratory diseases, increased pressure on the cardiovascular system, and impaired immune system function, severely eroding people's quality of life and health, and posing a heavy medical burden and profound health risks to society. Accurately measuring the spatiotemporal distribution of air pollutants, and thereby achieving precise and efficient pollution prevention and control and air quality improvement, has become a key technology urgently needed in the field of environmental management. However, achieving this goal in actual urban environmental monitoring and early warning work faces many complex and thorny challenges.

[0003] Spatial interpolation techniques such as Kriging and inverse distance weighted interpolation are widely used to infer the distribution of continuous variables based on data from a limited number of discrete sampling points. They were later introduced into the field of atmospheric science to estimate the distribution of air pollutant concentrations. Although Kriging takes into account the spatial correlation and variation structure of sample points, it is highly dependent on the layout of monitoring stations and can easily cause oversmoothing in areas with sparse stations, resulting in increased errors and distorted pollution distribution patterns. Inverse distance weighted interpolation assigns weights based on the distance between the monitoring station and the prediction point, with closer distances giving higher weights. However, it does not fully consider the impact of factors such as geographical features and land use types on pollutant propagation, resulting in poor interpolation results in areas with complex terrain or diverse land use types. These methods fail to effectively integrate multivariate auxiliary information and ignore factors such as the geographical environment, land use changes, and heterogeneous distribution of pollution sources. This makes it difficult to accurately depict the high-resolution spatiotemporal variations of pollutants and cannot meet the high-resolution and high-precision pollution distribution requirements required for refined air quality management. Summary of the Invention

[0004] To solve the above problems, the present invention proposes an urban pollutant distribution prediction method based on the fusion of dynamic and static features of multi-source data. This method uses easily collected multi-source public data, fuses dynamic and static features, and infers the air pollutant concentrations in the target area without monitoring stations, thereby realizing the calculation of the spatiotemporal distribution of air pollutants, effectively improving the temporal and spatial resolution of urban air pollutant monitoring, and effectively solving the problem of air pollutant concentration distribution prediction in areas without monitoring stations.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The urban pollutant distribution prediction method based on the fusion of dynamic and static features of multi-source data includes the following steps:

[0007] S1. Collect multi-source heterogeneous data in the study area and construct a multi-source heterogeneous dataset;

[0008] S2. Preprocess the collected multi-source heterogeneous data;

[0009] S3, using the block feature extraction module to process the remote sensing image blocks, population density information and altitude information to obtain the long-term static characteristics of the region;

[0010] S4. Using the air data reference module to process some stations around the target location, obtain the air reference dynamic characteristics of pollutant distribution from the surrounding air monitoring stations;

[0011] S5. Fusing the long-term static characteristics of the target block and the dynamic characteristics of the air reference to calculate the predicted target air pollutant concentration;

[0012] S6. Optimize the model to the optimal parameters.

[0013] Preferably, in step S1, the multi-source heterogeneous data includes regional satellite remote sensing image data, regional population density data, regional altitude data and hourly monitoring of air pollutant concentration data at air monitoring stations within the region; wherein the monitored air pollutants include sulfur dioxide, ozone, carbon monoxide, fine particulate matter PM 2.5 and inhalable particulate matter PM 10 .

[0014] Preferably, the pre-processing in step S2 comprises the following steps:

[0015] S21. Process and transform abnormal negative values ​​in regional population density data;

[0016] S22. Use interpolation to clean missing data from air monitoring stations and process them into hourly or daily data based on actual needs.

[0017] S23. Align the regional satellite remote sensing image data, regional population density data, and regional altitude data onto a 1 km × 1 km grid.

[0018] Preferably, the specific process of step S3 is:

[0019] S31, taking the location to be processed as the center, obtaining remote sensing image blocks within the regional range, and then sending the remote sensing image blocks to the semantic segmentation submodule and the feature representation submodule for dimensionality reduction, thereby obtaining remote sensing image features that describe the remote sensing image blocks from different angles;

[0020] S32, obtain the altitude data and population density data corresponding to the latitude and longitude area range in step S32, and obtain the altitude population feature by processing through the table feature extraction sub-network;

[0021] S33, splice and fuse the remote sensing image feature and the altitude population feature to obtain the regional long-term static feature.

[0022] Preferably, the specific process of step S4 is:

[0023] S41, select n air monitoring stations as reference air stations in the region with the coordinates of the target air station to be calculated as the center;

[0024] S42, calculate the distance information and angle information of the n air monitoring stations and the target air station respectively using the latitude and longitude information;

[0025] S43, splice and fuse the regional long-term static feature obtained in step S3, the distance information and angle information of the target air station, and the air pollutant concentration data of the air monitoring stations at time t to obtain n air reference dynamic features.

[0026] Preferably, the specific process of step S5 is:

[0027] S51, for the place to be calculated for the target air pollutant concentration, fuse the regional long-term static feature of the target place obtained from step S3 and the corresponding time t to be calculated with the n air reference dynamic features obtained from step S4 respectively to obtain n air concentration prediction values;

[0028] S52, for each prediction value, calculate a weight value according to the distance between the reference air station and the target air station, and then perform weighted summation to obtain the predicted target air pollutant concentration, and the calculation formula is: , wherein, is the predicted target air pollutant concentration; =1,2,……,n; is the i th weight value; is the total weight value; is the i th prediction value; is the i th reference air station and the target air station distance.

[0029] Preferably, it further includes step S6 for optimizing the model to the optimal parameters, and the specific process of step S6 is:

[0030] S61, use the coefficient of determination to measure the goodness of fit of the model: ​​​in, is the coefficient of determination; is the total number of air monitoring stations; For the air monitoring stations; For the The actual value of each air monitoring station; For the Inferred values ​​for air monitoring stations; is the average of the true values ​​of all air monitoring stations;

[0031] S62. Use mean absolute error and mean relative error to evaluate the performance of the model: Among them, MAE is the mean absolute error; MRE is the mean relative error;

[0032] S63. Use multi-loss parallel optimization: , , , ,in, is L1 loss; is the L2 loss; is the logarithmic loss; is the scale invariant loss;

[0033] S64. The total loss is: in, For the total loss.

[0034] After adopting the above technical solution, the present invention has the following beneficial effects: the present invention fully considers and combines the dynamic and static characteristics of multi-source heterogeneous data, so that the model can understand the propagation process of pollutants from more aspects; during the model training stage, when selecting neighboring reference points, a random selection strategy is used instead of fixed sites to enhance the robustness of the model and prevent overfitting; the data used are all public data and easy to obtain. Therefore, the present invention uses easily collected multi-source public data, integrates dynamic and static characteristics, and infers the concentration of air pollutants in the target area without monitoring stations, realizing the calculation of the spatiotemporal distribution of air pollutants, effectively improving the temporal and spatial resolution of urban air pollutant monitoring, and effectively solving the problem of predicting the distribution of air pollutant concentrations in areas without monitoring stations. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Flowchart of the present invention;

[0036] Figure 2 It is a flowchart of the present invention;

[0037] Figure 3 This is a schematic diagram of the structure of the regional long-term static feature extraction of the present invention;

[0038] Figure 4 It is a structural schematic diagram of the air reference dynamic feature extraction of the present invention;

[0039] Figure 5 It is a structural schematic diagram of the feature fusion of the present invention;

[0040] Figure 6 Schematic diagram of the relationship between the target air station and the reference air station of the present invention. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] like Figures 1 to 6 As shown in FIG, the urban pollutant distribution prediction method based on the fusion of dynamic and static features of multi-source data includes the following steps:

[0043] This example uses Huzhou as an example. The spatial scope covers the entire administrative area of ​​the demonstration zone, and the time period covers January 2022 to October 2024. The target spatial resolution is 1 km × 1 km, and the temporal resolution is hourly or daily.

[0044] S1. Collect multi-source heterogeneous data in the study area and construct a multi-source heterogeneous dataset;

[0045] In step S1, the multi-source heterogeneous data includes regional satellite remote sensing image data, regional population density data, regional altitude data and hourly monitoring of air pollutant concentration data of air monitoring stations within the region; wherein the monitored air pollutants include sulfur dioxide, ozone, carbon monoxide, fine particulate matter PM 2.5 and inhalable particulate matter PM 10 ;

[0046] S2. Preprocess the collected multi-source heterogeneous data;

[0047] The preprocessing in step S2 includes the following steps:

[0048] S21. Process and transform abnormal negative values ​​in regional population density data;

[0049] S22. Use interpolation to clean missing data from air monitoring stations and process them into hourly or daily data based on actual needs.

[0050] S23, aligning regional satellite remote sensing image data, regional population density data, and regional altitude data onto a 1 km × 1 km grid;

[0051] S3, using the block feature extraction module to process the remote sensing image blocks, population density information and altitude information to obtain the long-term static characteristics of the region;

[0052] The specific process of step S3 is:

[0053] S31, taking the location to be processed as the center, obtaining remote sensing image blocks within the regional range, and then sending the remote sensing image blocks to the semantic segmentation submodule and the feature representation submodule for dimensionality reduction, thereby obtaining remote sensing image features that describe the remote sensing image blocks from different angles;

[0054] S32, obtaining altitude data and population density data corresponding to the latitude and longitude area range of step S32, and processing them through the table feature extraction sub-network to obtain altitude population features;

[0055] S33, combining remote sensing image features with altitude population features to obtain long-term static features of the region;

[0056] S4. Using the air data reference module to process some stations around the target location, obtain the air reference dynamic characteristics of pollutant distribution from the surrounding air monitoring stations;

[0057] The specific process of step S4 is:

[0058] S41, taking the coordinates of the target air station to be calculated as the center, selecting n air monitoring stations in the area as reference air stations;

[0059] S42. Using the latitude and longitude information, calculate the distance information and angle information between the n air monitoring stations and the target air station;

[0060] S43: Use the regional long-term static features obtained in step S3, the distance information and angle information of the target air station, and the air pollutant concentration data of the air monitoring station at time t to perform splicing and fusion to obtain n air reference dynamic features;

[0061] S5. Fusing the long-term static characteristics of the target block and the dynamic characteristics of the air reference to calculate the predicted target air pollutant concentration;

[0062] The specific process of step S5 is:

[0063] S51. For the location where the target air pollutant concentration is to be calculated, the regional long-term static characteristics of the target location obtained in step S3 and the expected calculation corresponding time t are respectively integrated with the n air reference dynamic characteristics obtained in step S4 to obtain n predicted values ​​of air concentration;

[0064] S52. For each predicted value, a weight is calculated based on the distance between the reference air station and the target air station, and then weighted sum is performed to obtain the predicted target air pollutant concentration. The calculation formula is: , ,in, To predict the concentration of target air pollutants; =1,2,……,n; For the weights; is the total weight; For the predicted values; For the The distance between the reference air station and the target air station;

[0065] S6. Optimize the model to the optimal parameters;

[0066] The specific process of step S6 is:

[0067] S61. Use the coefficient of determination to measure the goodness of fit of the model: in, is the coefficient of determination; is the total number of air monitoring stations; For the air monitoring stations; For the The actual value of each air monitoring station; For the Inferred values ​​for air monitoring stations; is the average of the true values ​​of all air monitoring stations;

[0068] S62. Use mean absolute error and mean relative error to evaluate the performance of the model: Among them, MAE is the mean absolute error; MRE is the mean relative error;

[0069] S63. Use multi-loss parallel optimization: , , , ,in, is L1 loss; is the L2 loss; is the logarithmic loss; is the scale invariant loss;

[0070] S64. The total loss is: in, For the total loss.

[0071] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any changes or replacements within the technical scope disclosed by the present application, which can be easily thought by those skilled in the art, should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for predicting urban pollutant distribution based on the fusion of dynamic and static features of multi-source data, characterized by: The following steps are involved: S1. Collect multi-source heterogeneous data in the study area and construct a multi-source heterogeneous dataset; S2. Preprocess the collected multi-source heterogeneous data; S3, using the block feature extraction module to process the remote sensing image blocks, population density information and altitude information to obtain the long-term static characteristics of the region; The specific process of step S3 is: S31, taking the location to be processed as the center, obtaining remote sensing image blocks within the regional range, and then sending the remote sensing image blocks to the semantic segmentation submodule and the feature representation submodule for dimensionality reduction, thereby obtaining remote sensing image features that describe the remote sensing image blocks from different angles; S32, obtaining altitude data and population density data corresponding to the latitude and longitude area range of step S32, and processing them through the table feature extraction sub-network to obtain altitude population features; S33, combining remote sensing image features with altitude population features to obtain regional long-term static features; S4. Using the air data reference module to process some stations around the target location, obtain the air reference dynamic characteristics of pollutant distribution from the surrounding air monitoring stations; The specific process of step S4 is: S41, taking the coordinates of the target air station to be calculated as the center, selecting n air monitoring stations in the area as reference air stations; S42. Using the latitude and longitude information, calculate the distance information and angle information between the n air monitoring stations and the target air station; S43: Use the regional long-term static features obtained in step S3, the distance information and angle information of the target air station, and the air pollutant concentration data of the air monitoring station at time t to perform splicing and fusion to obtain n air reference dynamic features; S5. Fusing the long-term static characteristics of the target block and the dynamic characteristics of the air reference to calculate the predicted target air pollutant concentration; The specific process of step S5 is: S51. For the location where the target air pollutant concentration is to be calculated, the regional long-term static characteristics of the target location obtained in step S3 and the expected calculation corresponding time t are respectively integrated with the n air reference dynamic characteristics obtained in step S4 to obtain n predicted values ​​of air concentration; S52. For each predicted value, a weight is calculated based on the distance between the reference air station and the target air station, and then weighted sum is performed to obtain the predicted target air pollutant concentration. The calculation formula is: , ,in, To predict the concentration of target air pollutants; =1,2,……,n; For the weights; is the total weight; For the predicted values; For the The distance between the reference air station and the target air station; S6. Optimize the model to the optimal parameters; The specific process of step S6 is: S61. Use the coefficient of determination to measure the goodness of fit of the model: in, is the coefficient of determination; is the total number of air monitoring stations; For the air monitoring stations; For the The actual value of each air monitoring station; For the Inferred values ​​for air monitoring stations; is the average of the true values ​​of all air monitoring stations; S62. Use mean absolute error and mean relative error to evaluate the performance of the model: Among them, MAE is the mean absolute error; MRE is the mean relative error; S63. Use multi-loss parallel optimization: , , , ,in, is L1 loss; is the L2 loss; is the logarithmic loss; is the scale invariant loss; S64. The total loss is: in, For the total loss.

2. The urban pollutant distribution prediction method based on the fusion of dynamic and static features of multi-source data according to claim 1 is characterized by: In step S1, the multi-source heterogeneous data includes regional satellite remote sensing image data, regional population density data, regional altitude data and hourly monitoring of air pollutant concentration data of air monitoring stations within the region; wherein the monitored air pollutants include sulfur dioxide, ozone, carbon monoxide, fine particulate matter PM 2.5 and inhalable particulate matter PM 10 .

3. The urban pollutant distribution prediction method based on the fusion of dynamic and static features of multi-source data according to claim 1 is characterized in that: The preprocessing in step S2 includes the following steps: S21. Process and transform abnormal negative values ​​in regional population density data; S22. Use interpolation to clean missing data from air monitoring stations and process them into hourly or daily data based on actual needs. S23. Align the regional satellite remote sensing image data, regional population density data, and regional altitude data onto a 1 km × 1 km grid.

Citation Information

Patent Citations

  • Air pollution prediction method based on deep fusion of multi-source space-time big data

    CN112905560A

  • PM2.5 comprehensive domain space-time calculation inference method based on multi-source city big data

    CN113297527A