An air quality prediction method and system based on multi-source spatio-temporal data fusion

CN118861964BActive Publication Date: 2026-09-25ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410839715.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-09-25
Estimated Expiration
2044-06-26

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了一种基于多源时空数据融合的空气质量预测方法及系统,解决了现有技术缺少对多种数据同化以及衡量和选取同化数据的方法,导致预测数据缺乏精度和可靠性的问题

Benefits of technology

[0024]1、本发明通过ResNet提取地理环境特征计算出空间权重矩阵,并将数据输入SGWR模型(空间地理加权回归模型)中,进行数据加权回归分析,分析空气质量的空间组织,预测相关因素,成功构建了3D空间特征回归模型。并将卫星遥感的图像数值数据与同化后数据结合提供给空间地理加权回归模型作为污染物预测模型的参数变量,实现从公里级到百米级的精度跃升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861964B_ABST
    Figure CN118861964B_ABST
Patent Text Reader

Abstract

The application provides an air quality prediction method and system based on multi-source space-time data fusion, and relates to the technical field of environmental protection monitoring. The air quality prediction method based on multi-source space-time data fusion comprises the following steps: S1, real-time vertical profile data is obtained by using a laser radar, and then ground conventional monitoring network plane real-time data and satellite remote sensing plane real-time data of the national meteorological center are collected; S2, the three collected multi-source data are supplemented to a data set, and a three-dimensional distribution map of space-time pollutants is established; S3, the multi-source data are simultaneously input into a numerical model combined by WRF-Chem and CMAQ, and a source GSI assimilation system developed secondarily is adopted. The spatial accuracy is improved through ResNet and an SGWR model, the time accuracy is improved through a model combined by LSTM and CNN, the numerical model combined by WRF-Chem and CMAQ is used, and the source GSI assimilation system developed secondarily is adopted, so that the prediction data accuracy and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring technology, specifically to an air quality prediction method and system based on multi-source spatiotemporal data fusion. Background Technology

[0002] In the environmental protection industry, data needs to be obtained from air quality monitoring stations for analysis. Generally, there are six conventional atmospheric parameters (PM2.5, PM10, O3, SO2, NOx, CO) and five meteorological parameters (temperature, air pressure, humidity, wind direction, wind speed).

[0003] Currently, the country mainly uses two types of monitoring stations for ambient air quality: national control stations and small stations. National control stations offer reliable monitoring data, but are limited by high cost and limited coverage. Small stations are lower cost and easier to operate, but their monitoring data is relatively coarse.

[0004] Data assimilation, which combines meteorological data from different times, types, sources, and resolutions with background fields to create a dataset with temporal, spatial, and physical consistency, plays a crucial role in improving the accuracy of numerical atmospheric forecasts, particularly rainfall forecasts. However, current technologies rarely assimilate multiple data sets simultaneously, and there is no method for measuring and selecting assimilated data. This significantly limits data assimilation, failing to fully reflect the advantages of various data sources and resulting in a lack of accuracy and reliability in forecast data. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an air quality prediction method and system based on multi-source spatiotemporal data fusion. This solves the problem that existing technologies lack methods for assimilating multiple types of data and measuring and selecting assimilated data, resulting in a lack of accuracy and reliability in the predicted data.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an air quality prediction method based on multi-source spatiotemporal data fusion, comprising the following steps:

[0007] S1. Use lidar to acquire vertical profile data in real time, and then collect real-time plane data from the National Meteorological Center's ground-based conventional monitoring network and satellite remote sensing plane data.

[0008] S2. Supplement the dataset with the three types of multi-source data collected to establish a three-dimensional distribution map of spatiotemporal pollutants;

[0009] S3. Simultaneously, multi-source data are input into the numerical model combining WRF-Chem and CMAQ. A secondary-developed source GSI assimilation system is used to extend the ground assimilation of the chemical module to any location in space for 3D-Var assimilation processing. This completes the assimilation process of the lidar three-dimensional monitoring data, and provides a particulate matter reanalysis field constrained by the three-dimensional observation data, thus obtaining the dataset.

[0010] Preferably, the radar detection point range is 3km vertically and 100km wide, and the model is a Raman lidar, specifically a particulate lidar or an ozone lidar.

[0011] Preferably, the dataset extracts geographic environmental features (altitude, latitude, longitude, temperature, humidity, industrial areas, population density, etc.) using ResNet. A spatial weight matrix is ​​calculated based on geographic information to reflect factors such as spatial location, distance, and similarity. The data is then input into the SGWR model (Spatial Geographic Weighted Regression Model) for weighted regression analysis to analyze the spatial organization of air quality and predict related factors. A 3D spatial feature regression model has been successfully constructed. Satellite remote sensing image numerical data is combined with assimilated data and provided to the spatial geographic weighted regression model as parameter variables for the pollutant prediction model.

[0012] Preferably, a model combining LSTM and CNN is used. First, CNN convolutions are used to capture local and structural information of spatial features, and pooling layers are used to reduce dimensionality and retain the main salient features. Then, LSTM is used to model the long-term dependencies of time series and capture time series information. The predicted batch values ​​of the sequence are used to learn the time series information of the model. Finally, the two are combined to generate a prediction network model relative to spatial features and time. A high-precision dataset is input and the parameters are optimized by using the Adaptive Moment Estimation (Adam) optimizer.

[0013] Preferably, the specific process of the GSI assimilation system algorithm is to analyze the meteorological observation data onto the control variable field of the numerical prediction model by solving a least squares problem, thereby obtaining an analysis field with higher spatial resolution and more frequent temporal resolution than the observation data, as well as the corresponding analysis error covariance matrix. This analysis field and covariance matrix can be used to initialize the forecast field of the next numerical prediction model.

[0014] Preferably, the LSTM-CNN model processes data as follows: The data is processed by a convolutional neural network (CNN). First, the data passes through a convolutional layer to capture spatial features, then through a pooling layer to reduce dimensionality and reduce computational cost while retaining the main features. After all features are fused, the feature description of the convolutional neural network is obtained, and then the data is reshaped into the type processed by LSTM. At this point, the data is passed to the LSTM. After receiving the new input, the LSTM determines which data to keep and which to discard, using the sigmoid activation function. Then, the time series information of the data is obtained, and finally, the air quality data is predicted.

[0015] Preferably, the ratio of the model training set, validation set, and test set is 80%:10%:10%.

[0016] Preferably, the features extracted by CNN are air temperature, humidity, wind speed, and wind direction.

[0017] Preferably, the learning rate of the Adam algorithm is set to 0.01.

[0018] An air quality prediction system that integrates multi-source spatiotemporal data includes a multi-source data collection module, a multi-source data fusion module, and a prediction module;

[0019] The multi-source data collection module is used to collect real-time vertical profile data acquired by lidar, real-time planar data from the National Meteorological Center's conventional ground monitoring network, and real-time planar data from satellite remote sensing.

[0020] The multi-source data fusion module is used to extract geographic environmental features through ResNet, calculate the spatial weight matrix based on geographic information, input the data into the SGWR model for data weighted regression analysis, analyze the spatial organization of air quality, predict related factors, and successfully construct a 3D spatial feature regression model. At the same time, it uses a model combining LSTM and CNN to generate a prediction network model relative to spatial features and time.

[0021] The prediction module is used to input multi-source data into a numerical model combining WRF-Chem and CMAQ, and employs a secondary-developed source GSI assimilation system to provide a particulate reanalysis field constrained by three-dimensional observation data.

[0022] This invention provides an air quality prediction method and system based on multi-source spatiotemporal data fusion.

[0023] It has the following beneficial effects:

[0024] 1. This invention extracts geographic environmental features using ResNet to calculate a spatial weight matrix, and inputs the data into the SGWR model (Spatial Geographic Weighted Regression Model) for weighted regression analysis. This analysis examines the spatial organization of air quality, predicts relevant factors, and successfully constructs a 3D spatial feature regression model. Furthermore, satellite remote sensing image data is combined with assimilated data and provided to the spatial geographic weighted regression model as parameter variables for pollutant prediction, achieving a leap in accuracy from the kilometer level to the hundred-meter level.

[0025] 2. This invention employs a model combining LSTM and CNN. First, CNN convolutions are used to capture local and structural information of spatial features, while pooling layers reduce dimensionality and retain key salient features. Then, LSTM is used to model long-term dependencies in time series, capturing time series information. The predicted batch values ​​of the sequence are used to learn the time series information of the model. Finally, the two are combined to generate a predictive network model relative to spatial features and time. A high-precision dataset is input, and the parameters are optimized using the Adaptive Moment Estimation (Adam) optimizer, achieving a spatiotemporal resolution improvement from hourly to minute-level.

[0026] 3. Using a numerical model combining WRF-Chem and CMAQ, and employing a secondary-developed source GSI assimilation system, the ground assimilation of the chemical module is extended to any location in space for 3D-Var assimilation processing. This completes the assimilation process of lidar three-dimensional monitoring data, providing a particulate matter reanalysis field constrained by the three-dimensional observation data, resulting in the largest high-precision dataset in China. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the GS I assimilation system technical solution of the present invention;

[0028] Figure 2 This is a schematic diagram of the main components of the lidar of the present invention;

[0029] Figure 3 This is a schematic diagram illustrating the basic principle of the WRF-Chem invention.

[0030] Figure 4 This is a schematic diagram of the multi-source data quality control system for lidar according to the present invention;

[0031] Figure 5 This is a schematic diagram illustrating the working principle of the lidar of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Example:

[0034] Please see the appendix Figure 1 -Appendix Figure 5 This invention provides an air quality prediction method based on multi-source spatiotemporal data fusion, comprising the following steps:

[0035] S1. Real-time vertical profile data is acquired using lidar, followed by real-time planar data from the National Meteorological Center's conventional ground monitoring network and satellite remote sensing. This integrates the information advantages of lidar stereo detection, ground monitoring networks, and satellite remote sensing. For the networked lidar system in the field, real-time atmospheric environmental monitoring is conducted 24 hours a day under unattended conditions. Raw data generated by the acquisition software on remote machines is automatically uploaded to the laboratory's local FTP server. Simultaneously, a web server is built specifically for lidar data processing and result display. The web server first automatically downloads raw data from each site of the networked lidar system from the FTP server, then uses the Ferna ld method to invert the data, obtaining the extinction coefficient, and further obtaining useful information such as depolarization, signal-to-noise ratio, boundary layer height, and optical thickness. The establishment of the web server enables consistent management of networked lidar data, including establishing a unified data format, unified data storage, unified data inversion, and unified data quality control requirements. Finally, the inversion results are automatically and in real-time uploaded to the MySQL database. Simultaneously, the PM2.5 vertical profiles for each station are obtained using the extinction coefficient-PM2.5 conversion model and converted to the BUFR format required by the assimilation system. The files are then uploaded and input into the assimilation system. Finally, a Linux server is set up to run WRF-Chem, the assimilation system, and post-processing. A Crontab task in Linux automatically downloads the latest FNL and other data at 22:00 daily. After successful download, the automatic execution of WRF-Chem mode is triggered, including WPS preprocessing, real.exe, and wrf.exe. After execution, the PM2.5 BUFR data file and wrfout nc file generated by the lidar are read and input into the lidar assimilation system to begin the assimilation process. After assimilation is complete, ncl is used to process and plot the results, saving them to the corresponding location or writing them to the MySQL database.

[0036] S2. Supplement the dataset with the three types of multi-source data collected to establish a three-dimensional distribution map of spatiotemporal pollutants;

[0037] S3. Simultaneously, multi-source data is input into a numerical model combining WRF-Chem and CMAQ. A secondary-developed source GSI assimilation system is used to extend the ground-based assimilation of the chemical module to any spatial location for 3D-Var assimilation processing. This completes the assimilation process of the lidar three-dimensional monitoring data, providing a particulate matter reanalysis field constrained by the three-dimensional observation data, thus obtaining the dataset. Multi-source data assimilation includes data at different temporal and spatial resolutions, combining observational and model forecast data at different scales. It fuses high-resolution observational and model data with low-resolution data, thereby improving the system's accuracy and reliability. The basic idea of ​​the 3D-Var assimilation method is to estimate the three-dimensional variables (temperature, humidity, wind) in the atmosphere. Based on the errors between meteorological observation data and the numerical forecast model, an optimal match is found between the observed and forecasted fields, making the model forecast results closer to the actual observation results. The WRF-Chem model adds chemical reaction and diffusion models to the WRF model, enabling it to forecast atmospheric pollutants from satellite data and meteorological observation data.

[0038] The 3D-Var assimilation method is characterized by its robustness, efficiency, and parallelizability, and has been successfully applied in data assimilation research and practical applications in meteorology, environment, and oceanography.

[0039] The radar detection point range is 3km vertically and 100km wide. It is a Raman lidar, specifically a particulate matter lidar and an ozone lidar. The lidar system structure diagram shows it mainly consists of four subsystems: a laser emission subsystem, a laser receiving subsystem, a signal processing and control subsystem, and a data inversion subsystem. The lidar uses a laser as its light source to detect the backscattered echo signal from the interaction between the laser and atmospheric aerosols. It extracts aerosol backscattering and extinction characteristics using the Ferna ld method, and utilizes the polarization detection principle to obtain the aerosol depolarization characteristics, distinguishing dust storms from other aerosol particles. When solving the Ferna ld method, the aerosol extinction backscattering ratio is a significant source of error, depending on the scale spectrum distribution, refractive index, morphology, and composition of the aerosol particles, and its variation range is very wide.

[0040] The dataset extracts geographic environmental features (altitude, latitude, longitude, temperature, humidity, industrial areas, population density, etc.) using ResNet. A spatial weight matrix is ​​calculated based on geographic information to reflect factors such as spatial location, distance, and similarity. The data is then input into the SGWR model (Spatial Geographic Weighted Regression Model) for weighted regression analysis to analyze the spatial organization of air quality and predict related factors. A 3D spatial feature regression model was successfully constructed. Satellite remote sensing image numerical data and assimilated data are combined and provided to the spatial geographic weighted regression model as parameter variables for the pollutant prediction model.

[0041] This model combines LSTM and CNN. First, CNN convolutions are used to capture local and structural information of spatial features, and pooling layers reduce dimensionality to retain the main salient features. Then, LSTM is used to model the long-term dependencies of time series and capture time series information. The predicted batch values ​​of the sequence are used to learn the time series information of the model. Finally, the two are combined to generate a prediction network model relative to spatial features and time. A high-precision dataset is input and the parameters are optimized by using the Adaptive Moment Estimation (Adam) optimizer.

[0042] The specific process of the GSI assimilation system algorithm is to solve a least squares problem to analyze meteorological observation data onto the control variable field of the numerical prediction model, thereby obtaining an analysis field with higher spatial resolution and more frequent temporal resolution than the observation data, as well as the corresponding analysis error covariance matrix. This analysis field and covariance matrix can be used to initialize the forecast field of the next numerical prediction model.

[0043] The LSTM-CNN model processes data as follows: Data is processed by a Convolutional Neural Network (CNN). First, convolutional layers capture spatial features, followed by pooling layers to reduce dimensionality and retain key features. All features are then fused to obtain the CNN's feature description, which is then reshaped into the LSTM processing type. The data is then fed back to the LSTM. The LSTM receives the new input and determines which data to keep and discard using the Sigma-Aldrich activation function. Finally, it extracts the time-series information of the data and predicts air quality data. Using LSTM-CNN for air quality prediction has many unique advantages. It better handles time-series data and spatial correlations, improves model generalization and prediction accuracy, and enhances the ability to predict long-term dependencies in time-series data. Therefore, it can be considered the preferred model for air quality prediction tasks. When using a hybrid LSTM-CNN model, the goal is not to continuously generate single-value outputs, but rather to predict batch values ​​of a sequence to forecast air quality over a future period. The advantage of the LSTM-CNN hybrid model over other neural network models lies in its better ability to model long-term dependencies in time-series data and its superior extraction of spatial features. By combining the characteristics of CNN and LSTM, patterns in time series data can be better identified, and prediction accuracy and generalization can be improved. At the same time, it avoids the training instability problems associated with traditional RNN models. Furthermore, in the choice of RNN, this approach uses LSTM instead of SimpleRNN to address the gradient vanishing problem in long-term series.

[0044] The ratio of the model training set, validation set, and test set is 80%:10%:10%.

[0045] The features extracted by CNNs are air temperature, humidity, wind speed, and wind direction. The number of CNN parameters = kernel size × kernel depth × number of kernel groups = kernel size × input feature matrix depth × output feature matrix depth. The Convolutional Neural Network (CNN) part uses a Conv1 D structure for feature extraction. Because the air quality-related features between adjacent stations exhibit cross-influence, the features of multiple stations are unfolded into a two-dimensional structure integrating the data from each station during modeling. Then, CNNs are used for feature extraction in a multi-site, multi-feature space. In the field of air quality prediction, the advantages of CNNs mainly include the following: Advantages in extracting spatial sequence features: The convolutional layers of CNNs can automatically extract features. When the convolutional kernels are convolved with the input data, they can automatically extract features at different time or spatial scales at the same spatial location without manual feature engineering. This characteristic can be used to discover and analyze the complex and uneven distribution characteristics of pollutants. Adaptive characteristics: The network structure of CNNs is well-suited for processing data features with varying complexity, where the features of each layer's activation values ​​can be adaptively trained and learned. This feature improves the model's robustness in handling large datasets and its performance during generalization. It is suitable for large-scale training sets: in scenarios with very large training datasets, CNNs can significantly save training time and reduce computational load, making the model training process more efficient.

[0046] The learning rate for the Adam algorithm was set to 0.01.

[0047] An air quality prediction system that integrates multi-source spatiotemporal data includes a multi-source data collection module, a multi-source data fusion module, and a prediction module;

[0048] The multi-source data collection module is used to collect real-time vertical profile data acquired by lidar, real-time planar data from the National Meteorological Center's conventional ground monitoring network, and real-time planar data from satellite remote sensing.

[0049] The multi-source data fusion module is used to extract geographic environmental features through ResNet, calculate the spatial weight matrix based on geographic information, input the data into the SGWR model for data weighted regression analysis, analyze the spatial organization of air quality, predict related factors, and successfully construct a 3D spatial feature regression model. At the same time, it uses a model combining LSTM and CNN to generate a prediction network model relative to spatial features and time.

[0050] The prediction module is used to input multi-source data into a numerical model combining WRF-Chem and CMAQ, and uses a secondary developed source GSI assimilation system to provide a particulate reanalysis field constrained by three-dimensional observation data.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An air quality prediction method based on multi-source spatiotemporal data fusion, characterized in that, Includes the following steps: S1. Use lidar to acquire vertical profile data in real time, and then collect real-time plane data from the National Meteorological Center's conventional ground monitoring network and satellite remote sensing plane data. S2. Supplement the dataset with the three types of multi-source data collected to establish a three-dimensional distribution map of spatiotemporal pollutants; S3. Simultaneously, multi-source data are input into the numerical model combining WRF-Chem and CMAQ. A secondary-developed source GSI assimilation system is used to extend the ground assimilation of the chemical module to any location in space for 3D-Var assimilation processing. This completes the assimilation process of the lidar three-dimensional monitoring data, and provides a particulate matter reanalysis field constrained by the three-dimensional observation data, thus obtaining the dataset. The dataset extracts geographic environmental features using ResNet and calculates a spatial weight matrix based on geographic information, reflecting spatial location, distance, and similarity factors. The data is then input into the SGWR model for weighted regression analysis to analyze the spatial organization of air quality and predict related factors. A 3D spatial feature regression model was successfully constructed, and satellite remote sensing image numerical data was combined with assimilated data and provided to the spatial geographic weighted regression model as parameter variables for the pollutant prediction model. Using the LSTM-CNN model, we first use CNN convolution to capture local and structural information of spatial features, pooling layers reduce dimensionality to retain the main salient features, then use LSTM to model the long-term dependencies of time series, capture time series information, predict the batch values ​​of the sequence to learn the time series information of the model, and finally combine the two to generate a prediction network model relative to spatial features and time. We input a high-precision dataset and optimize the parameters by using the adaptive moment estimation (Adam) optimizer. The LSTM-CNN model processes data as follows: Data is processed by a convolutional neural network (CNN). First, the data passes through a convolutional layer to capture spatial features, then through a pooling layer to reduce dimensionality, reduce computational cost, and retain the main features. After all features are fused, the feature description of the CNN is obtained, and then it is reshaped into the type processed by LSTM. At this point, the data is passed to LSTM. After receiving the new input, LSTM determines which data to keep and which to discard, using the sigmoid activation function. Then, the time series information of the data is obtained, and finally, the air quality data is predicted.

2. The air quality prediction method based on multi-source spatiotemporal data fusion according to claim 1, characterized in that, The radar detection point range is 3km vertically and 100km in width. It is a Raman lidar, and the radar is a particulate lidar and an ozone lidar.

3. The air quality prediction method based on multi-source spatiotemporal data fusion according to claim 1, characterized in that, The specific process of the GSI assimilation system algorithm is to solve a least squares problem to analyze meteorological observation data onto the control variable field of the numerical prediction model, thereby obtaining an analysis field with higher spatial resolution and more frequent temporal resolution than the observation data, as well as the corresponding analysis error covariance matrix. This analysis field and covariance matrix can be used to initialize the forecast field of the next numerical prediction model.

4. The air quality prediction method based on multi-source spatiotemporal data fusion according to claim 1, characterized in that, The ratio of the model training set, validation set, and test set is 80%:10%:10%.

5. The air quality prediction method based on multi-source spatiotemporal data fusion according to claim 1, characterized in that, The features extracted by CNN are air temperature, humidity, wind speed, and wind direction.

6. The air quality prediction method based on multi-source spatiotemporal data fusion according to claim 1, characterized in that, The learning rate for the Adam algorithm was set to 0.

01.

7. An air quality prediction system based on multi-source spatiotemporal data fusion according to the method of claim 1, characterized in that, It includes a multi-source data collection module, a multi-source data fusion module, and a prediction module; The multi-source data collection module is used to collect real-time vertical profile data acquired by lidar, real-time planar data from the National Meteorological Center's conventional ground monitoring network, and real-time planar data from satellite remote sensing. The multi-source data fusion module is used to extract geographic environmental features through ResNet, calculate the spatial weight matrix based on geographic information, input the data into the SGWR model for data weighted regression analysis, analyze the spatial organization of air quality, predict related factors, and successfully construct a 3D spatial feature regression model. At the same time, it uses a model combining LSTM and CNN to generate a prediction network model relative to spatial features and time. The prediction module is used to input multi-source data into a numerical model combining WRF-Chem and CMAQ, and uses a secondary developed source GSI assimilation system to provide a particulate reanalysis field constrained by three-dimensional observation data.

Citation Information

Patent Citations

  • Method and visualization system for improving atmospheric transmission quantization capability by fusing multi-source data

    CN115203189A

  • Data collection system and method for feeding aquatic animals

    US20200113158A1