Atmospheric pollutant grid data downscaling method and device and electronic equipment

By combining air quality forecasting models and super-resolution pollutant deep fusion models, the problem that traditional air quality grid data cannot accurately capture pollutant concentrations has been solved, achieving high-resolution pollutant concentration prediction and improving prediction accuracy.

CN121880757APending Publication Date: 2026-04-17INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI
Filing Date
2025-11-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional air quality grid data cannot accurately capture the true patterns of pollutant concentration changes, resulting in poor accuracy in pollutant concentration prediction.

Method used

The first simulated concentration field is generated by calling the air quality forecasting model system, and then fused with the actual monitoring values. The grid data is processed by the block learning sub-model in the super-resolution pollutant deep fusion model to generate high-resolution pollutant concentration prediction results.

Benefits of technology

It achieves precise downscaling of atmospheric pollutant grid data, improves the accuracy of pollutant concentration prediction, and can better reflect the local pollution characteristics of small areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880757A_ABST
    Figure CN121880757A_ABST
Patent Text Reader

Abstract

The invention provides an atmospheric pollutant grid data downscaling method and device and electronic equipment, and the method comprises the steps: generating a first model concentration field of various pollutants through meteorological condition information, emission source information and geographical condition information which are preset in an air quality prediction model system and influence the pollutants; acquiring various actual monitoring values, fusing the actual monitoring values into the first simulation concentration field to generate a second simulation concentration field, and inputting the second simulation concentration field into a super-resolution pollutant deep fusion model, each block learning sub-model in the super-resolution pollutant deep fusion model respectively processes a single grid under the second simulation concentration field, and pollutant concentration prediction results under each grid in the second simulation concentration field are output. According to the method, large-scale grid data can be disassembled by means of each block learning sub-model of the super-resolution pollutant deep fusion model, and a high-resolution pollutant concentration prediction result under a single grid is generated, so that the effect of downscaling the atmospheric pollution grid data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of atmospheric pollution data processing technology, and in particular to a method, apparatus and electronic equipment for downscaling atmospheric pollutant grid data. Background Technology

[0002] Air quality grid data is obtained by dividing a certain area (such as a city, region, or the whole country) into regular grids (such as 1km×1km, 3km×3km) according to fixed spatial intervals. Then, through air quality numerical simulation, multi-source observation data assimilation and other methods, the core air quality indicators and spatiotemporal distribution data within each regular grid are obtained.

[0003] Because low-resolution grid data (e.g., 3km×3km) of air quality data averages out small-scale pollution differences—for example, the pollution concentration near an overpass is averaged out by the pollution concentration in a park—traditional air quality grid data cannot accurately reflect the local pollution characteristics of small areas. Therefore, downscaling air quality grid data is essential. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, and electronic device for downscaling atmospheric pollutant grid data to solve the problem that traditional air quality grid data downscaling schemes cannot accurately capture the true variation patterns of pollutant concentrations, resulting in poor pollutant concentration prediction accuracy.

[0005] In a first aspect, embodiments of this application provide a method for downscaling atmospheric pollutant grid data, wherein the method includes: A pre-set air quality forecasting model system is invoked to generate a first simulated concentration field for various pollutants. The air quality forecasting model system is pre-set with meteorological conditions, emission source information, and geographical conditions that affect the various pollutants. Acquire various actual monitoring values ​​collected from each monitoring station, and integrate the actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field; The second simulated concentration field is input into the super-resolution pollutant deep fusion model to obtain the pollutant concentration prediction results under each grid in the second simulated concentration field output by the super-resolution pollutant deep fusion model. The super-resolution pollutant deep fusion model includes several block learning sub-models, and each block learning sub-model processes the pollutant data of various types under a single grid in the second simulated concentration field.

[0006] Secondly, embodiments of this application provide an apparatus for downscaling atmospheric pollutant grid data, wherein the apparatus includes: The simulation module is used to call a pre-set air quality forecasting model system to generate a first simulated concentration field of various pollutants. The air quality forecasting model system is pre-set with meteorological conditions, emission source information and geographical conditions that affect the various pollutants. The fusion module is used to acquire various actual monitoring values ​​collected by each monitoring station and fuse the actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field. The downscaling module is used to input the second simulated concentration field into the super-resolution pollutant deep fusion model, and obtain the pollutant concentration prediction results of each grid in the second simulated concentration field output by the super-resolution pollutant deep fusion model. The super-resolution pollutant deep fusion model includes several block learning sub-models, and each block learning sub-model processes the pollutant data of each type in a single grid of the second simulated concentration field.

[0007] Thirdly, embodiments of this application provide an electronic device, wherein the electronic device includes: a processor; and a memory storing a program; wherein the program includes instructions, which, when executed by the processor, cause the processor to perform the atmospheric pollutant grid data downscaling method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method for downscaling atmospheric pollutant grid data as described in the first aspect.

[0009] The beneficial effects of this application are: This application provides a method, apparatus, and electronic device for downscaling atmospheric pollutant grid data. The method involves calling a pre-set air quality forecasting model system and generating a first model concentration field for various pollutants based on pre-set meteorological conditions, emission source information, and geographical conditions that may affect pollutants. Then, it acquires various actual monitoring values ​​collected from various monitoring stations and fuses these values ​​with the first simulated concentration field to generate a second simulated concentration field. This second simulated concentration field is then input into a super-resolution pollutant deep fusion model. Each block learning sub-model in the super-resolution pollutant deep fusion model processes individual grids under the second simulated concentration field, thereby outputting the pollutant concentration prediction results for each grid in the second simulated concentration field.

[0010] In this way, the large-scale grid data can be decomposed by learning sub-models in each block of the super-resolution pollutant deep fusion model, generating high-resolution pollutant concentration prediction results under a single grid, so as to achieve the effect of downscaling the atmospheric pollution grid data. Attached Figure Description

[0011] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 This paper illustrates a flowchart of a method for downscaling atmospheric pollutant grid data provided in this application. Figure 2 This application provides a schematic diagram of a process for generating a first simulated concentration field. Figure 3 This application provides a schematic diagram of a process for generating a second simulated concentration field. Figure 4 This paper presents a schematic diagram illustrating the effect of the second simulated concentration field generated by the present application. Figure 5 This paper illustrates a schematic diagram of a model processing flow for the super-resolution pollutant deep fusion model provided in this application. Figure 6 This illustration shows a schematic diagram of the effect of downscaling the atmospheric pollutant grid data provided in this application; Figure 7 A schematic diagram of one device structure for the atmospheric pollutant grid data downscaling device provided in this application is shown. Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of this application is shown. Detailed Implementation

[0012] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0013] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0015] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0016] The existing AI-based downscaling technology for air quality grid data is DeepCAMS. DeepCAMS is a deep learning-based air quality refinement technology specifically designed to increase the resolution of low-resolution PM2.5 grid data. Its approach is based on deep learning technology, training the model on 3-hourly, 0.75° baseline data. The trained model then processes the low-resolution data (which can be visualized as coarse data) into high-resolution data updated hourly, with each grid cell measuring only 0.25° (approximately 28 kilometers).

[0017] DeepCAMS is similar to an air quality data retouching tool. It uses deep learning technology to correct slow-updating, low-resolution data to make it faster-updating and higher-resolution. The basic data used in this scheme is accurate at 3-hour intervals, which needs to be downscaled to hourly. However, this scheme focuses on data patterns and does not consider that air pollution is a dynamic process with a constant physical logic of pollution diffusion. This downscaling scheme cannot accurately reflect the dynamic changes in air pollution. Therefore, the air quality data obtained by downscaling based on this scheme cannot guarantee the accuracy of subsequent air pollution predictions.

[0018] In view of this, this application provides a method, apparatus and electronic device for downscaling atmospheric pollutant grid data, and provides a novel approach to downscaling atmospheric grid data. It introduces actual monitoring data into the air quality data downscaling process to assist the model in learning the dynamic diffusion process of air pollution, thereby helping the air quality data obtained by the model downscaling to help with subsequent air pollution prediction.

[0019] In one aspect, this application provides a method for downscaling atmospheric pollutant grid data. This method can be applied to any electronic device with the function of downscaling atmospheric pollutant grid data, including but not limited to personal mobile terminals, computers, supercomputers, or servers. Figure 1 As shown, the method includes the following steps: S11. Call the preset air quality forecasting model system to generate the first simulated concentration field of various pollutants; The air quality forecasting model system is pre-loaded with meteorological conditions, emission source information, and geographical conditions that affect various pollutants. S12. Obtain various actual monitoring values ​​collected from each monitoring station, and integrate the actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field; S13. Input the second simulated concentration field into the super-resolution pollutant deep fusion model, and obtain the pollutant concentration prediction results of each grid in the second simulated concentration field output by the super-resolution pollutant deep fusion model; The super-resolution pollutant deep fusion model includes several block learning sub-models, each of which processes the pollutant data of various types under a single grid of the second simulated concentration field.

[0020] In this embodiment, the method calls a pre-set air quality forecasting model system. Based on the meteorological conditions, emission source information, and geographical conditions that may affect pollutants, which are pre-set in the air quality forecasting model system, a first model concentration field for various pollutants is generated. Then, various actual monitoring values ​​collected by each monitoring station are obtained and fused with the first simulated concentration field to generate a second simulated concentration field. The second simulated concentration field is then input into a super-resolution pollutant deep fusion model. Each block learning sub-model in the super-resolution pollutant deep fusion model processes a single grid under the second simulated concentration field, and then outputs the pollutant concentration prediction results under each grid in the second simulated concentration field.

[0021] In this way, the large-scale grid data can be decomposed by learning sub-models in each block of the super-resolution pollutant deep fusion model, generating high-resolution pollutant concentration prediction results under a single grid, so as to achieve the effect of downscaling the atmospheric pollution grid data.

[0022] The following will provide a detailed explanation of steps S11 to S14 with specific examples: It is understood that this application is still a specific solution combining artificial intelligence and air pollution data processing. The purpose of downscaling the air pollution grid data in this application is to obtain more refined (which can be understood as small grid) air pollution data so that in the subsequent data application process, the refined air pollution data can be used to predict air pollution and provide strong data assistance for subsequent air pollution prevention and control.

[0023] Similar to how image editing software refines individual pixels in a large image, downscaling air pollution data can be visualized as refining the data to improve the accuracy of air pollution data at a smaller scale (i.e., a smaller grid area). To achieve high accuracy, the underlying data must first be sufficiently accurate, much like a clear image negative, so that the refined pixels are also clear. Compared to the traditional DeepCAMS, which uses 3-hourly, 0.75° underlying data for downscaling, this application uses data from a pre-built air quality forecasting model system. This system includes pre-set information on meteorological conditions affecting various pollutants, emission sources affecting various pollutants, and geographical conditions affecting various pollutants.

[0024] As a preferred implementation, the underlying data can be data from the NAQPMS numerical air quality model. The Nested Air Quality Prediction Modeling System (NAQPMS) is a third-generation numerical air quality model independently developed by the applicant. NAQPMS integrates complex processes such as multi-scale meteorological fields, atmospheric chemical dynamics, aerosol physicochemistry, and dry and wet deposition. For example, the NAQPMS numerical model can be as follows: Figure 2 As shown, by acquiring topographic data, meteorological data, and emission data, and based on atmospheric pollution advection processes, vertical diffusion, dry and wet deposition, meteorological chemistry, liquid phase chemistry, aerosol chemistry, etc., the following numerical air quality model process was constructed:

[0025] in, This represents the change in concentration of the i-th pollutant. Characterizing the advection term, The model characterizes the diffusion term, R represents chemical formation, E represents the emission process, and S represents the dry and wet deposition process. This numerical model uses meteorological principles and mathematical equations to reflect the influence of advection, convection, chemical processes, dry and wet deposition, and emission processes on pollutant concentration changes. For example, assuming i represents PM2.5, this numerical model can reflect the influence of various processes on PM2.5 concentration changes. Specific details about the aforementioned air quality numerical model can be found in the applicant's technical documentation on NAQPMS; it is not the core focus of this paper and will not be described in detail here. Then, gridded air quality data is generated based on this air quality numerical model and serves as the first simulated concentration field provided in this application.

[0026] In this application, during step S11, simulated concentration fields of various pollutants can be generated by acquiring the NAQPMS numerical model as the first simulated concentration field. The types of pollutants include, but are not limited to: sulfur dioxide, nitrogen oxides, carbon monoxide, ozone, PM2.5, PM10, volatile organic compounds (VOCs), etc.

[0027] Specifically, based on meteorological conditions, emission source information, and geographical conditions included in the NAQPMS numerical model, a first simulated concentration field for various pollutants is generated. Meteorological conditions include wind speed, wind direction, temperature, humidity, etc., which determine the transport and diffusion paths of pollutants. Emission source information includes the emission intensity, location, and types of pollutants from various pollution sources such as industrial, transportation, and agricultural emissions. Geographical conditions include topography, land use type, and geological type, which affect the pollutant diffusion efficiency in local areas; for example, the diffusion efficiency of PM10 in desertified areas is much higher than that in coastal areas.

[0028] The simulated concentration field refers to the set of spatiotemporal distribution data of the simulated pollutant concentration values. When executing step S11, NAQPMS is called to obtain meteorological driving data, emission source data, geographical data and other information to calculate the initial emission intensity of various pollutants. Then, meteorological data is collected to simulate and calculate the transport, diffusion, physicochemical reaction and dry and wet deposition processes of pollutants in the atmosphere. Through gridded calculation, a region is divided into several grid regions, and the concentration values ​​of different pollutants in each grid are calculated to obtain the first simulated concentration field.

[0029] As a preferred implementation, the process of generating the first simulated concentration field mainly involves data simulation in two dimensions: meteorological driving process and pollutant process simulation.

[0030] The simulation of meteorological driving processes involves using the Weather Research and Forecasting (WRF) model to identify the underlying factors determining pollutant transport pathways, diffusion ranges, and chemical reactions. The following parameterization scheme from the WRF model is employed: Microphysics scheme selection: WSM3; Boundary layer scheme selection: YSU; Longwave radiation scheme selection: RRTM longwave radiation; Shortwave radiation scheme selection: Dudhi; Road surface process selection: Noah.

[0031] For pollutant process simulation: NAQPMS is used to perform detailed simulations of transport, diffusion, emission, wet and dry deposition, and chemical processes, fully recreating the evolution of pollutants in the atmosphere. The following models and mechanisms are employed to achieve this detailed simulation: Simulation choice for dry settlement: Wesely resistance model; Selection of wet sedimentation and liquid phase chemistry: A scheme based on an improved RADM model; Gas-phase chemical selection: CBM-Z carbon bond reaction mechanism; The aerosol process uses the aerosol thermodynamics module ISORROPIA 1.7 to handle the gas-particle distribution and thermodynamic equilibrium of sulfates, nitrates, and ammonium salts. In generating the first simulated concentration field, this application considers 28 heterogeneous chemical reactions for the interaction between gases and aerosols.

[0032] In one implementation, the horizontal resolution of the simulation grid of the first simulated concentration field is 15km, and the number of grids is 432*399. Each grid corresponds to a spatial unit. By calling NAQPMS, the concentrations of various pollutants in each grid are calculated, and finally a gridded first simulated concentration field is formed.

[0033] By selecting the various parameterization schemes provided above, the simulation bias of the NAQPMS numerical model on the pollutant concentration field in meteorological situations can be effectively reduced, thereby providing high-precision input data for subsequent model processing.

[0034] In this embodiment, in step S12, each monitoring station specifically refers to meteorological monitoring stations distributed throughout the country. The actual monitoring values ​​collected by these stations include: air quality monitoring data, high-resolution numerical simulation data, population data, road network density data, nighttime light intensity data, terrain monitoring data, and land use type data from multiple sources. Further, step S12 standardizes the data, fusing the obtained actual monitoring values ​​of various types into the first simulated concentration field to generate a second simulated concentration field.

[0035] One implementation involves resampling the actual monitored values ​​using nearest-neighbor interpolation to map the actual monitored values ​​belonging to the same grid resolution to the grid corresponding to the first simulated concentration field. Specifically, for a given target grid, for each point within that grid, the actual monitored value from the monitoring station spatially closest to it is directly selected as the interpolation result for that grid and mapped and saved to the target grid corresponding to the first simulated concentration field. For example, the spatial matching relationship between the grid and the monitoring stations can be determined first. For each target grid, the monitoring station closest to the grid center is calculated, and the actual monitored value from the selected nearest monitoring station is directly assigned to the corresponding target grid in the first simulated concentration field to complete the mapping of the actual monitored values.

[0036] During the process, the improved Z-Score algorithm can be used to clean the data before performing nearest neighbor interpolation mapping, which can further reduce the interference of abnormal data on the data in the grid and ensure the accuracy and consistency of the data subsequently input into the super-resolution pollutant deep fusion model.

[0037] As one implementation, the second simulated concentration field can be a comprehensive data set of the first simulated concentration field and the mapped actual monitoring values. In some possible embodiments, the second simulated concentration field can be a comprehensive data set obtained after correction based on the actual monitoring values. Specifically, data correction can be completed through the following steps: Based on the actual monitoring values, the simulated concentration data of each grid in the first simulated concentration field are optimally interpolated and corrected, and the second simulated concentration field is generated based on the corrected simulated concentration data of the grid.

[0038] Since existing AI downscaling techniques fail to incorporate constraints from ground-based observation data, the results of pollutant concentration downscaling differ significantly from reality. Furthermore, they are poor at characterizing areas with sparse monitoring stations or heavy pollution events. In this embodiment, the optimal interpolation algorithm constrains the simulated data with actual monitoring values, continuously supplements the simulated space, and optimizes statistical errors, effectively reducing the difference between the simulation results and the actual monitoring values.

[0039] The error between simulation results and actual monitored values ​​typically stems from two sources: background error and observation error. Background error refers to the simulation error in the NAQPMS numerical simulation, while observation error refers to the error in the actual observed values ​​from the monitoring stations. The NAQPMS numerical simulation error mainly arises from factors such as simplification of physical processes and grid resolution limitations, leading to a systematic deviation between the simulated concentration field and the actual monitored values. Observation error primarily originates from the accuracy of the measuring instruments themselves and insufficient spatiotemporal representativeness (e.g., in areas with sparse monitoring stations), resulting in localized deviations in the actual monitored values.

[0040] To overcome the aforementioned two types of errors, this application innovatively proposes an optimal interpolation correction logic. By quantitatively analyzing the statistical characteristics of systematic and local biases, and then calculating the optimal weighting coefficients for the actual monitored values ​​and simulated data based on the bias statistics, this optimal weighting coefficient ensures that the target data obtained after fusing the actual monitored values ​​and the data from the simulated concentration field minimizes the total error. Finally, the actual monitored values ​​and the data simulated by the NAQPMS numerical model are fused according to the optimal weights to generate high-precision second simulated concentration field data.

[0041] Specifically, as one implementation method, the simulated concentration data can be optimally interpolated and corrected through the following steps: Based on the actual monitored values, the simulated concentration data of each grid in the first simulated concentration field are optimally interpolated and corrected using the following formula:

[0042]

[0043] in, The simulated concentration data for the corrected grid. Let K be the simulated concentration data of each grid in the first simulated concentration field, K be the gain matrix, y be the actual monitored value, H be the observation operator, B be the error covariance of the first simulated concentration field, and R be the error covariance of the actual monitored value.

[0044] Using the above optimal interpolation formula, the optimal linear estimate of the first simulated concentration field and the actual monitored value with minimum variance is calculated. Assuming that the first simulated concentration field and the actual monitored value are independent of each other, the observation error covariance matrix is ​​a diagonal matrix. Assuming that the control error variance is 10% of the observed value, and assuming that the background error covariance matrix is ​​in Balgovind form, the error covariance B of the first simulated concentration field can be determined by the following formula:

[0045] in, This represents the error covariance between the i-th and j-th grids. This represents the distance between the i-th grid and the j-th grid, where L is the constraint assimilation radius. In a preferred embodiment, L is set to 80 km. It is the product of the prior covariance and the prior errors of the i-th and j-th grids. The estimated prior error for each grid point is assumed to be 30% of the grid point's forecast value plus 10% of the average forecast value for the same region.

[0046] Through this error covariance, it can be understood that the closer the distance between grids, the stronger the correlation of errors. Once the distance exceeds the constraint assimilation radius L, the error correlation between grids will disappear accordingly. This can fully reflect the physical law that atmospheric pollutants have strong local diffusion capabilities and weak correlation over long distances.

[0047] For specific details, please refer to, for example Figure 3 The processing flow shown involves obtaining the first simulated concentration field (e.g., ... Figure 3 Simulated concentration data for each grid in the Basesimulation (in the database). The actual monitored value y (which can be represented in vector form), i.e., the ground monitored value (e.g. Figure 3 The observation operator H is set, and the background error covariance matrix B is set according to the formula for the background error covariance matrix B. The observation error covariance matrix R is also set, and the gain matrix is ​​calculated using the formula... The gain matrix K of the second simulated concentration field was calculated, and optimal interpolation was used according to the formula. Generate the analysis field to obtain simulated concentration data for the corrected grid. Specifically, the difference between the actual monitored value and the first simulated concentration field is used to reanalyze the air quality in the first simulated concentration field according to the optimal weight correction value quantized by the gain matrix K, resulting in a fused, high-precision second simulated concentration field (e.g., ...). Figure 3 Reanalysis in (the context of the problem).

[0048] In this way, the spatial continuity of the first simulated concentration field can be used to fill the observation gaps in spatially sparse areas of actual monitoring data. At the same time, the local authenticity of actual monitoring values ​​can be used to correct the background bias generated during the simulation of the first simulated concentration field by the air quality numerical model. Finally, through continuous statistical optimization using this minimum variance, the fused second simulated concentration field is made to conform to atmospheric physical laws and approximate the actual pollution distribution as closely as possible, thereby reducing the difference between simulation and reality.

[0049] For specific results, please refer to the following: Figure 4As shown, the middle Base simulation is a schematic diagram of the first simulated concentration field generated by the numerical model. It can be seen that the numerical model significantly overestimates the PM2.5 concentration in the Sichuan-Chongqing region. The right Observation is a schematic diagram of the actual monitoring value distribution. The actual monitoring results show that the PM2.5 concentration in the Sichuan-Chongqing region is not as high as shown in the first simulated concentration field. By using the actual monitoring values ​​to perform optimal interpolation correction on the first simulated concentration field, the PM2.5 concentration distribution diagram shown on the left Reanalysis is obtained. The concentration in the first simulated concentration field is constrained by the actual monitoring values, so that the data in the second simulated concentration field can more accurately reflect the spatial distribution of PM2.5 in my country, providing accurate model input data for subsequent model downscaling.

[0050] Further, in step S13, the second simulated concentration field is input into the super-resolution pollutant deep fusion model to obtain the pollutant concentration prediction results under each grid obtained by the learning sub-model of each block in the super-resolution pollutant deep fusion model. As one implementation, model training data can be constructed using the same data processing path as in steps S11 and S12. That is, historical actual monitoring values ​​are fused into the first simulated concentration field to obtain the second simulated concentration field. The data of this second simulated concentration field is then stored separately in a database. In subsequent model training phases, the data in this database can be used as training sample data to train the model.

[0051] In some possible embodiments, the learning sub-models of each block are independent of each other. The step of inputting the second simulated concentration field into the super-resolution pollutant deep fusion model to obtain the pollutant concentration prediction results for each grid within the second simulated concentration field output by the super-resolution pollutant deep fusion model includes: The data from each grid of the second simulated concentration field is split and input into each of the block learning sub-models. Each block learning sub-model learns the pollution patterns based on the received data, and generates and outputs the pollutant concentration prediction results for each grid based on the pollution patterns.

[0052] As one implementation method, the super-resolution pollutant deep fusion model is obtained by training an initial super-resolution pollutant deep fusion model using training sample data from the information database. During training, the model parameters are continuously adjusted based on the output of the loss function until the loss function converges. The model framework of the super-resolution pollutant deep fusion model can be as follows: Figure 5 As shown, it includes: an input layer, multiple blocks ( Figure 5The blocks 1, 2, and 3 shown are only examples and are not the only examples. The output layer is also shown. Each block corresponds to a learning sub-model. The input layer includes a regularization layer and a fully connected layer. Each learning sub-model includes at least a regularization layer, a fully connected layer, and a temporary fallback layer. The output layer includes a fully connected layer.

[0053] Each fully connected layer includes n=512 hidden layers to ensure the model's ability to fit complex data. A dropout probability P=0.5 is set for the temporary defitting layer to randomly discard 50% of the neurons. The regularization layer performs batch normalization on the input data, adjusting the mean and variance to make the data distribution more stable. The fully connected layer performs non-linear deep learning on the features obtained from the normalization process, deeply extracting and fusing features from the input data. Finally, the temporary defitting layer forces the model to learn more generalized feature patterns, rather than over-relying on the output of a few neurons, to prevent overfitting.

[0054] Each block serves as a training unit and is trained independently. The number of blocks can be dynamically increased or decreased. For applications with high complexity and high accuracy requirements, the number of blocks can be increased adaptively, but this will increase the consumption of computing resources accordingly. Conversely, a model with fewer blocks can be selected for downscaling.

[0055] The data input to this super-resolution pollutant deep fusion model is the input dataset, which mainly includes low-resolution data, specifically air quality reanalysis data, i.e., various pollutants (PM2.5). 2.5 PM 10 The input dataset includes concentration data for (SO2, NO2, CO, O3). It also includes high-resolution data, specifically subdivided into: nighttime light pollution, population density, ground elevation, land type, road network density, etc. As mentioned earlier, the low-resolution and high-resolution data are optimized using an optimal interpolation algorithm to obtain the second simulated concentration field. Therefore, as an implementation method, the input dataset can directly be the data from the second simulated concentration field.

[0056] The training process of this super-resolution pollutant deep fusion model was designed with an 80% training set and a 20% test set. A mean squared error (MSE) loss function was used to evaluate the model accuracy across the entire super-resolution pollutant deep fusion data. Ultimately, the model learns the correlation between input data and pollutant concentrations. Actual training revealed that after partitioning the data into different grids for each block, the model can achieve a training cycle of 7 days. The output layer combines simulated concentrations to generate and output high-resolution fusion data for the entire country, hourly, with a 1km*1km grid, resulting in a high-precision pollutant concentration prediction dataset. The parameters set during model training included: 799,755 model parameters (approximately 800,000 parameters); a training time of 500 epochs (20 minutes) for a single block learning sub-model; processing approximately 30 million 1km grids; and an average data generation time of 30 seconds per hour.

[0057] In this application, a multi-block collaborative training mode enables the model to adapt to the resources of the computing platform. Each block includes a fully connected network, a residual neural network, regularization, and temporary regression. The interconnection of multiple blocks enhances the learning ability of the nonlinear relationship between multi-source data and training labels. Furthermore, this super-resolution pollutant deep fusion model incorporates high-resolution multi-source big data such as population, road network density, nighttime light, topography, and land use types, improving the representation ability of downscaling results in local details. Specific effects can be seen in [reference needed]. Figure 6 As shown, low-resolution atmospheric pollution grid data from across the country can be downscaled and then displayed in high resolution as atmospheric pollution grid data from North China.

[0058] Based on the method provided in the first aspect, and in the second aspect, this application provides an apparatus for downscaling atmospheric pollutant grid data, which can, as... Figure 7 As shown, the device 70 includes: The simulation module 701 is used to call a preset air quality forecasting model system to generate a first simulated concentration field of various pollutants. The air quality forecasting model system is preset with meteorological condition information, emission source information and geographical condition information that affect the various pollutants. The fusion module 702 is used to acquire various actual monitoring values ​​collected by each monitoring station and fuse the actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field; The downscaling module 703 is used to input the second simulated concentration field into the super-resolution pollutant deep fusion model, and obtain the pollutant concentration prediction results under each grid in the second simulated concentration field output by the super-resolution pollutant deep fusion model. The super-resolution pollutant deep fusion model includes several block learning sub-models, and each block learning sub-model processes the pollutant data of various types under a single grid in the second simulated concentration field.

[0059] In some possible embodiments, acquiring the various actual monitoring values ​​collected by each monitoring station includes: Acquire the actual monitoring values ​​of air quality monitoring data, high-resolution numerical simulation data, population data, road network density data, nighttime light, topography, and land use type collected by each of the aforementioned monitoring stations; The actual monitored values ​​are resampled using nearest neighbor interpolation, and the actual monitored values ​​belonging to the same grid resolution are mapped to the grid corresponding to the first simulated concentration field.

[0060] In some possible embodiments, the step of acquiring various actual monitoring values ​​collected from each monitoring station and fusing the actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field includes: Based on the actual monitoring values, the simulated concentration data of each grid in the first simulated concentration field are optimally interpolated and corrected, and the second simulated concentration field is generated based on the corrected simulated concentration data of the grid.

[0061] In some possible embodiments, the step of performing optimal interpolation correction on the simulated concentration data of each grid in the first simulated concentration field based on the actual monitored values ​​includes: Based on the actual monitored values, the simulated concentration data of each grid in the first simulated concentration field are optimally interpolated and corrected using the following formula:

[0062]

[0063] in, The simulated concentration data for the corrected grid. Let K be the simulated concentration data of each grid in the first simulated concentration field, K be the gain matrix, y be the actual monitored value, H be the observation operator, B be the error covariance of the first simulated concentration field, and R be the error covariance of the actual monitored value.

[0064] In some possible embodiments, the error covariance B of the first simulated concentration field is determined by the following formula:

[0065] in, This represents the error covariance between the i-th and j-th grids. L represents the distance between the i-th grid and the j-th grid, where L is the constraint assimilation radius. It is the product of the prior covariance and the prior errors of the i-th grid and the j-th grid.

[0066] In some possible embodiments, the learning sub-models of each block are independent of each other. The step of inputting the second simulated concentration field into the super-resolution pollutant deep fusion model to obtain the pollutant concentration prediction results for each grid within the second simulated concentration field output by the super-resolution pollutant deep fusion model includes: The data from each grid of the second simulated concentration field is split and input into each of the block learning sub-models. Each block learning sub-model learns the pollution patterns based on the received data, and generates and outputs the pollutant concentration prediction results for each grid based on the pollution patterns.

[0067] In some possible embodiments, the preset air quality forecasting model system is an embedded air quality forecasting model system, and the horizontal resolution of the simulation grid of the first simulated concentration field is 15km, with a grid number of 432*339.

[0068] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0069] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0070] Thirdly, exemplary embodiments of this application also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this application.

[0071] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.

[0072] An exemplary embodiment of this application also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of this application.

[0073] refer to Figure 8 The present invention describes a structural block diagram of an electronic device 800 that can serve as a server or client of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein.

[0074] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM 802) or a computer program loaded from a storage unit 808 into a random access memory (RAM 803). The RAM 803 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output interface (I / O interface 805) is also connected to the bus 804.

[0075] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information to electronic device 800. Input unit 806 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 808 may include, but is not limited to, disks and optical discs. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0076] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the aforementioned method for downscaling atmospheric pollutant grid data can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the aforementioned method for downscaling atmospheric pollutant grid data by any other suitable means (e.g., by means of firmware).

[0077] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0078] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0079] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0080] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0081] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0082] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

Claims

1. A method for downscaling atmospheric pollutant grid data, characterized in that, The method includes: A pre-set air quality forecasting model system is invoked to generate a first simulated concentration field for various pollutants. The air quality forecasting model system is pre-set with meteorological conditions, emission source information, and geographical conditions that affect the various pollutants. The actual monitoring values ​​collected from each monitoring station are obtained, and the actual monitoring values ​​are integrated into the first simulated concentration field to generate a second simulated concentration field; The second simulated concentration field is input into the super-resolution pollutant deep fusion model to obtain the pollutant concentration prediction results under each grid in the second simulated concentration field output by the super-resolution pollutant deep fusion model. The super-resolution pollutant deep fusion model includes several block learning sub-models, and each block learning sub-model processes the pollutant data of various types under a single grid in the second simulated concentration field.

2. The method according to claim 1, characterized in that, The acquisition of various actual monitoring values ​​collected from each monitoring station includes: Acquire the actual monitoring values ​​of air quality monitoring data, high-resolution numerical simulation data, population data, road network density data, nighttime light, topography, and land use type collected by each of the aforementioned monitoring stations; The actual monitored values ​​are resampled using nearest neighbor interpolation, and the actual monitored values ​​belonging to the same grid resolution are mapped to the grid corresponding to the first simulated concentration field.

3. The method according to claim 1, characterized in that, The step of acquiring various actual monitoring values ​​collected from each monitoring station and integrating these actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field includes: Based on the actual monitoring values, the simulated concentration data of each grid in the first simulated concentration field are optimally interpolated and corrected, and the second simulated concentration field is generated based on the corrected simulated concentration data of the grid.

4. The method according to claim 3, characterized in that, The step of performing optimal interpolation correction on the simulated concentration data of each grid in the first simulated concentration field based on the actual monitored values ​​includes: Based on the actual monitored values, the simulated concentration data of each grid in the first simulated concentration field are optimally interpolated and corrected using the following formula: in, The simulated concentration data for the corrected grid. Let K be the simulated concentration data of each grid in the first simulated concentration field, K be the gain matrix, y be the actual monitored value, H be the observation operator, B be the error covariance of the first simulated concentration field, and R be the error covariance of the actual monitored value.

5. The method according to claim 4, characterized in that, The error covariance B of the first simulated concentration field is determined by the following formula: in, This represents the error covariance between the i-th and j-th grids. L represents the distance between the i-th grid and the j-th grid, where L is the constraint assimilation radius. It is the product of the prior covariance and the prior errors of the i-th grid and the j-th grid.

6. The method according to claim 1, characterized in that, Each of the block learning sub-models is independent of the others. The step of inputting the second simulated concentration field into the super-resolution pollutant deep fusion model to obtain the pollutant concentration prediction results for each grid within the second simulated concentration field output by the super-resolution pollutant deep fusion model includes: The data from each grid of the second simulated concentration field is split and input into each of the block learning sub-models. Each block learning sub-model learns the pollution patterns based on the received data, and generates and outputs the pollutant concentration prediction results for each grid based on the pollution patterns.

7. The method according to claim 1, characterized in that, The preset air quality forecasting model system is an embedded air quality forecasting model system. The horizontal resolution of the simulation grid of the first simulated concentration field is 15km, and the number of grids is 432*339.

8. A device for downscaling atmospheric pollutant grid data, characterized in that, The device includes: The simulation module is used to call a pre-set air quality forecasting model system to generate a first simulated concentration field of various pollutants. The air quality forecasting model system is pre-set with meteorological conditions, emission source information and geographical conditions that affect the various pollutants. The fusion module is used to acquire various actual monitoring values ​​collected by each monitoring station and fuse the actual monitoring values ​​into the first simulated concentration field to generate a second simulated concentration field. The downscaling module is used to input the second simulated concentration field into the super-resolution pollutant deep fusion model, and obtain the pollutant concentration prediction results of each grid in the second simulated concentration field output by the super-resolution pollutant deep fusion model. The super-resolution pollutant deep fusion model includes several block learning sub-models, and each block learning sub-model processes the pollutant data of each type in a single grid of the second simulated concentration field.

9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing a program; wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.