High-temporal-spatial-resolution atmospheric particulate source analysis method based on observation constraint

By simulating and assimilating observational data, a machine learning model was constructed. Combined with meteorological and socioeconomic data, the problem of insufficient spatiotemporal resolution and representativeness in the source apportionment of atmospheric particulate matter was solved, achieving accurate analysis with high spatiotemporal resolution and supporting the formulation of scientific governance strategies.

CN121960144APending Publication Date: 2026-05-01NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANKAI UNIV
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing atmospheric particulate matter source apportionment techniques have low spatiotemporal resolution and low representativeness of source apportionment results, relying on a single model and data statistics.

Method used

A method based on observation constraints is adopted. By simulating particulate matter concentrations at a set time and region, observational data and model data are assimilated to construct a machine learning model. Meteorological, geographical and socioeconomic data are used for high spatiotemporal resolution analysis, and the model is optimized to improve the accuracy and representativeness of the analysis.

Benefits of technology

It achieves source apportionment results with high spatiotemporal resolution, improves the representativeness and accuracy of the apportionment results, provides a scientific basis for air pollution prevention and control policies, optimizes governance strategies and reduces governance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960144A_ABST
    Figure CN121960144A_ABST
Patent Text Reader

Abstract

The invention discloses a high-temporal-spatial-resolution atmospheric particulate matter source analysis method based on observation constraint. The method comprises the following steps: simulating set time, and setting regional particulate matter PM2.5 and concentration of each component thereof; assimilating the observed PM2.5 and component concentration data of the PM2.5 in the set time and the set region and the PM2.5 and component concentration data of the PM2.5 obtained through CMAQ simulation; performing source analysis based on the assimilated data; and constructing a machine learning model by taking the source contribution of the source analysis as a dependent variable and taking the meteorological data, the geographic and social economic data and the road traffic data in each grid of the set time and the set region as independent variables, and simulating the source contribution concentration of different sources through the constructed model in a scale reduction manner. By adopting the method provided by the invention, an accurate source analysis result with high temporal-spatial resolution can be obtained, so that a scientific basis is provided for accurately formulating an atmospheric pollution prevention and control policy, a treatment strategy is optimized, the treatment cost is reduced, and the treatment efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Observation-Constrained High Spatiotemporal Resolution Apportionment Method for Atmospheric Particulate Matter Sources Technical Field

[0001] This invention relates to the field of air pollution research, and in particular to a high spatiotemporal resolution method for source apportionment of atmospheric particulate matter based on observational constraints. Background Technology

[0002] Existing atmospheric particulate matter source apportionment techniques generally employ receptor models (such as the Chemical Mass Balance (CMB) model, the Positive Matrix Factorization (PMF) model, or the Multiple Linear Model (ME2)) to apportion the sources of atmospheric particulate matter. These methods analyze the chemical composition characteristics of particulate matter collected from single or multiple locations, extract key factors, and quantitatively assess the contribution of pollution sources. However, the results obtained by this method have low spatiotemporal resolution, and because they rely on a single model and statistical data, the representativeness of the source apportionment results is low. Summary of the Invention

[0003] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a high spatiotemporal resolution source apportionment method for atmospheric particulate matter based on observation constraints. This method can not only obtain source apportionment results with high spatiotemporal resolution, but also the source apportionment results have high representativeness.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0005] A high spatiotemporal resolution method for apportioning atmospheric particulate matter sources based on observation constraints includes:

[0006] Simulation of particulate matter PM2.5 in a set time and area 2.5 and the concentrations of its components;

[0007] Assimilate the PM observed at the specified time and in the specified region 2.5 Its component concentration data and the simulated PM 2.5 and its component concentration data;

[0008] Source parsing is performed based on the assimilated data;

[0009] Using the source contribution from the source analysis as the dependent variable, and meteorological data, geographical and socioeconomic data, and road traffic data within each grid of the set time and set region as independent variables, a machine learning model is constructed. The constructed model is then scaled down to simulate the concentration of source contributions from different sources.

[0010] Preferably, the simulation sets the time and the area for particulate matter (PM). 2.5 Methods for determining the concentrations of PM2.5 and its components include: simulating particulate matter PM2.5 using a combination of the WRF meteorological model and the CMAQ air quality model. 2.5and the concentrations of its components.

[0011] Preferably, the particulate matter PM 2.5 The components include water-soluble ions, potentially toxic elements, and elemental carbon (EC) and organic carbon (OC); wherein the water-soluble ions include Cl. − SO4 2- NO3 - and NH4 + The potentially toxic elements include Mg, Al, Si, Ca, K, Ti, Fe, As, Co, Cr, Cu, Pb, Zn, Mn, and Ni.

[0012] Preferably, the assimilation of PM observations from the set time and set region... 2.5 Its component concentration data and PM obtained from CMAQ simulation 2.5 The method for obtaining PM and its component concentration data includes: using the simulated PM 2.5 The concentrations of PM2.5 and its components were used as independent variables, with the observed PM2.5 concentrations as the basis. 2.5 The concentrations of its components were used as dependent variables and standardized.

[0013] Ideally, multiple different machine learning models are used for standardization, and then the results are compared. If the assimilation result R of a certain machine learning method is higher... 2 If all values ​​are greater than 0.9, then the data assimilated by this machine learning model will be used as the basis for subsequent source analysis.

[0014] Preferably, the method for constructing the machine learning model includes:

[0015] Using machine learning methods, with the source contribution of the source resolution as the dependent variable, and meteorological data, geographical and socio-economic data, and road traffic data in each grid of the set time and set region as independent variables, a machine learning model is trained.

[0016] During the training process, statistically significant dependent and independent variables are selected, and then combined with their actual physical meaning, the machine learning model is optimized based on the statistically significant and physically meaningful dependent and independent variables.

[0017] Preferably, the method for selecting statistically significant dependent and independent variables is as follows: performing importance and correlation analysis on the dependent and independent variables, comparing the strength of their interaction, and selecting statistically significant dependent and independent variables based on the analysis and comparison results.

[0018] Ideally, multiple machine learning models should be constructed, and each model should be used to perform downscaling simulations of source contribution concentrations from different sources.

[0019] Preferably, the acquisition of meteorological data, geographical and socioeconomic data, and road traffic data in each grid within the specified time and region involves first acquiring the raw data of meteorological data, geographical and socioeconomic data, and road traffic data of individual grids at different scales and resolutions, and then preprocessing the raw data using GIS tools to obtain data patterns that can be used to build machine learning models.

[0020] Preferably, the geographic and socioeconomic data includes land use data, road traffic data, topographic elevation data, population density data, vegetation index data, fire point data, and point of interest data.

[0021] Because the present invention adopts the above-described technical solution, it has the following beneficial effects:

[0022] 1. The method of this invention can obtain source apportionment results with high spatiotemporal resolution and high representativeness, thereby providing a more scientific basis for the accurate formulation of air pollution prevention and control policies, optimizing governance strategies, reducing governance costs, and improving governance efficiency;

[0023] 2. The dependent variable source contribution used in constructing the machine learning model in this invention is PM based on observations from a set time and a set region that have been assimilated. 2.5 Its component concentration data and the simulated PM 2.5 The data obtained from analyzing the concentration data of PM2.5 components means that the original data in this application is a fusion of actual and simulated data. This effectively combines the advantages of observational data and model simulation, providing more representative PM2.5 data. 2.5 Source analysis results provide more accurate dependent variables for building machine learning models. As a result, the simulation performance of the machine learning model is better, and the source analysis results obtained from the simulation are more representative and accurate.

[0024] 3. When assimilating data, multiple different machine learning models are used for standardized assimilation processing. The processing results are then compared, and the best data is selected as the basis for subsequent source analysis. This improves the representativeness and accuracy of the assimilated data, making the source analysis results based on the assimilated data more accurate and representative. Consequently, the high spatiotemporal resolution source analysis results obtained by the invention are more accurate and representative.

[0025] 4. During the construction of the machine learning model, the machine learning model is optimized simultaneously, so that the source analysis results obtained through the machine learning model are more accurate; and the data used to optimize the model has both statistical significance and physical practical meaning, so the model constructed has more practical application value, the source analysis structure simulated by the constructed machine learning model is also more scientific and practical, and the pollution prevention and control measures based on this data are also more practical.

[0026] 5. This invention constructs multiple different machine learning models, each used to perform downscaling simulations of source contribution concentrations from different sources. This allows for cross-validation of model performance based on the source resolution results of each model, enabling the adjustment of model parameters to improve resolution accuracy and generalization ability, thereby further enhancing the accuracy and representativeness of the source resolution results. Attached Figure Description

[0027] Figure 1 is a flowchart illustrating the high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints of the present invention.

[0028] Figure 2 shows the PM of the present invention. 2.5 Source resolution results;

[0029] Figure 3 shows the PM of the present invention. 2.5 Source analysis result factor spectrum;

[0030] Figure 4 shows the spatial distribution characteristics of the source analysis model using GAM with high spatiotemporal resolution in this invention.

[0031] Figure 5 shows the spatial distribution characteristics of the model analyzed by the high spatiotemporal resolution source analysis of RF in this invention;

[0032] Figure 6 shows the spatial distribution characteristics of the source analysis model using XGBoost's high spatiotemporal resolution;

[0033] Figure 7 shows the spatial distribution characteristics of the source analysis model using Bayesian high spatiotemporal resolution in this invention. Detailed Implementation

[0034] As shown in Figures 1 to 6, this invention discloses a high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints, which includes:

[0035] Simulation of particulate matter PM2.5 in a set time and area 2.5 and the concentrations of its components;

[0036] Assimilate the PM observed at the specified time and in the specified region 2.5 Its component concentration data and the simulated PM 2.5 and its component concentration data;

[0037] Source parsing is performed based on the assimilated data;

[0038] Using the source contribution from the source analysis as the dependent variable, and meteorological data, geographical and socioeconomic data, and road traffic data within each grid of the set time and set region as independent variables, a machine learning model is constructed. The constructed model is then scaled down to simulate the concentration of source contributions from different sources.

[0039] In this embodiment, the particulate matter PM 2.5 The components are its conventional components, including water-soluble ions, potentially toxic elements, and elemental carbon (EC) and organic carbon (OC). Among them, the water-soluble ions include Cl... − SO4 2- NO3 - and NH4 + The potentially toxic elements include Mg, Al, Si, Ca, K, Ti, Fe, As, Co, Cr, Cu, Pb, Zn, Mn, and Ni.

[0040] Preferably, in this embodiment, the simulation setting time and setting area particulate matter PM 2.5 Methods for determining the concentrations of PM2.5 and its components include: simulating particulate matter PM2.5 using a combination of the WRF meteorological model and the CMAQ air quality model. 2.5 and the concentrations of its components.

[0041] The specific method for combining the WRF meteorological model with the CMAQ air quality model is as follows:

[0042] The WRF meteorological model first acquires the necessary initial and boundary condition data, such as GFS (Global Forecast System) or NCEP (National Center for Environmental Prediction), and then analyzes the data; next, it sets the domain, that is, selects an appropriate horizontal resolution and vertical hierarchy according to the study area, and defines the computational domain of the model; then it runs WRF, that is, after configuring the physical options of the WRF meteorological model, it runs the WRF model to generate high-resolution meteorological data (including temperature, humidity, wind speed, wind direction, etc.), which will be used as input for CMAQ;

[0043] The CMAQ air quality model first obtains a pollution source emission inventory, then selects an appropriate chemical reaction mechanism according to the requirements to accurately describe the transformation process of chemical substances in the atmosphere; and then initializes the CMAQ model using the meteorological field output by WRF and appropriate boundary conditions.

[0044] Coupling WRF and CMAQ: Convert the meteorological data output from WRF into a format readable by CMAQ using tools such as MCIP (Meteorology-Chemistry Interface Processor); then perform a data consistency check to ensure the consistency between the meteorological field provided by WRF and the grid, time step, etc. used by CMAQ.

[0045] Simulation execution: After preparing all input data, run the CMAQ model for a predetermined time period to simulate PM. 2.5 and the concentrations of its components.

[0046] PM obtained from simulation 2.5 The concentration of each component is 9km x 9km with low spatiotemporal resolution.

[0047] In this embodiment, the PM observed at the set time and in the set area 2.5 The PM2.5 concentration data and its component concentration data are the actual data for the specified time and region. This can be achieved by first acquiring PM2.5 concentration data synchronously offline from multiple locations within the specified region over a specified time period. 2.5 The sample was then analyzed for PM. 2.5v The concentration of the component and the concentration of its individual components are the actual data for a given time and region. This analysis can be performed using any known and feasible method, and the analytical method is not an innovation of this invention, so it will not be described in detail here.

[0048] The assimilation observation of PM 2.5 Its component concentration data and PM obtained from CMAQ simulation 2.5 The method for obtaining PM and its component concentration data includes: using the simulated PM 2.5 The concentrations of PM2.5 and its components were used as independent variables, with the observed PM2.5 concentrations as the basis. 2.5 The concentrations of its components were used as dependent variables and standardized.

[0049] In the specific standardization process, machine learning methods can be used. This involves importing the aforementioned independent and dependent variable data into a machine learning model, which then performs the standardization. The machine learning model can be a generalized additive model (GAM), a random forest (RF), or XGBoost.

[0050] In this embodiment, preferably, multiple different machine learning models are used for standardized assimilation processing, and then the processing results are compared. If the assimilation result R of a certain machine learning model is... 2If all values ​​are greater than 0.9, the data assimilated by this machine learning method is used as the basis for subsequent source analysis. This improves the representativeness and accuracy of the assimilated data, resulting in more accurate and representative source analysis results based on this assimilated data. Consequently, the high spatiotemporal resolution source analysis results obtained by the invention are also more accurate and representative.

[0051] Regarding the aforementioned machine learning models, GAM can flexibly capture nonlinear relationships, making it particularly suitable for complex simulation error correction; RF, by integrating multiple decision trees, provides high prediction accuracy and robustness, especially performing well in situations with high data noise; XGBoost, an algorithm for gradient boosting, has significant advantages in handling large-scale data and complex patterns, particularly in feature selection and model tuning, effectively improving the representativeness and accuracy of simulation results. The assimilation methods used in these machine learning approaches are well-known technologies and not innovative points of this invention, therefore they will not be elaborated upon here.

[0052] In this embodiment, the method for source resolution based on assimilated data includes: processing the assimilated PM... 2.5 The concentration data of the particulate matter and its components were incorporated into the receptor model. The source of atmospheric particulate matter was analyzed and the source contribution was quantified through the receptor model.

[0053] The receptor model can be a chemical mass balance model (CMB), a positive matrix factorization model (PMF), or a multivariate linear model (ME2).

[0054] In this embodiment, the acquisition of meteorological data, geographical and socioeconomic data, and road traffic data in each grid within a set time and set region involves first acquiring the raw data of meteorological data, geographical and socioeconomic data, and road traffic data of individual grids at different scale resolutions, and then preprocessing the raw data using GIS tools (such as ArcGIS, QGIS) to obtain data patterns that can be used to build machine learning models.

[0055] The meteorological data includes the daily maximum temperature (tmax), the number of extreme high-temperature days (tmax_days), rainfall (pre), relative humidity (rh), atmospheric pressure (p), wind speed (ws), and wind direction (wd) for a specified time and region. The geographic and socioeconomic data includes land use data, road traffic data, topographic elevation data, population density data, vegetation index data, fire point data, and point of interest data. In the above meteorological data, temperatures greater than or equal to 35 degrees Celsius are considered extreme high temperatures; the others are daily averages. The meteorological data, geographic and socioeconomic data, and road traffic data are multi-source, high-resolution spatial data.

[0056] Preferably, the method for constructing the machine learning model includes:

[0057] Using machine learning methods, with the source contribution of the source resolution as the dependent variable, and meteorological data, geographical and socioeconomic data, and AOD data in each grid of the set time and set region as independent variables, a machine learning model is trained.

[0058] During model training, statistically significant independent variables are selected and then combined with their actual physical meaning. The machine learning model is trained based on these statistically significant and physically meaningful independent and dependent variables. The specific method is conventional and not an inventive point, so it will not be elaborated here.

[0059] After constructing the machine learning model using the above method, the model is used to simulate the source contribution concentrations from different sources at a reduced scale, i.e., to simulate the source contribution concentrations from different sources at a spatial resolution of 1 km * 1 km, thereby obtaining source resolution results with high spatiotemporal resolution. The statistical significance refers to the contribution and importance of the independent variable to the dependent variable exceeding a set threshold (this threshold can be set according to research needs). Preferably, the method for selecting statistically significant independent variables is as follows: performing importance and correlation analysis on the dependent and independent variables, comparing the strength of their interactions, and selecting statistically significant independent variables based on the analysis and comparison results.

[0060] Importance analysis involves analyzing the degree of influence of each independent variable on the dependent variable (i.e., the source contribution), thereby identifying factors affecting PM. 2.5 The key influencing factors and their contribution to the source; the correlation analysis is to analyze the degree of linear or nonlinear association between each different independent variable, so as to identify the importance of the relationship between different independent variables in affecting the dependent variable; the comparison of interaction strength is to compare the strength of the interaction effects between different independent variables and between them and the dependent variable, so as to discover the combination of variables (i.e., the combination of independent and dependent variables) that need to be considered simultaneously to accurately reflect the phenomenon.

[0061] The actual physical meaning refers to the fact that the variable can be reasonably explained in the real world. For example, in the study of PM... 2.5 When considering source contributions, if the emissions activities of an industrial zone are closely related to high concentrations of specific pollutant components, then this relationship has clear practical physical significance.

[0062] Training machine learning models using statistically significant and physically meaningful independent and dependent variables ensures that the selected independent variables are not merely chosen because of data significance, but are based on actual scientific principles and domain knowledge. This results in models with greater practical application value, source analysis structures simulated by the constructed machine learning models that are more scientific and meaningful, and pollution prevention measures based on this data that are more practically applicable.

[0063] You can build only one machine learning model, or you can build multiple different machine learning models. In this embodiment, it is preferable to build multiple different machine learning models, such as GAM (Generalized Additive Model), Random Forest, Extreme Gradient Boosting (XGBOOST), or Bayesian Spatiotemporal Model.

[0064] Multiple machine learning models are constructed, and each model is used to perform downscaling simulations of source contribution concentrations from different sources. This allows for cross-validation of model performance based on source apportionment results, enabling parameter adjustments to improve accuracy and generalization ability. This further enhances the representativeness and accuracy of the source apportionment results. Furthermore, the parameters of the machine learning models themselves can be adjusted based on geographical information and pollution characteristics of different regions, demonstrating strong applicability and flexibility. Therefore, the method of this invention can adapt to source apportionment in more diverse regions with varying pollution characteristics, resulting in more accurate and representative source apportionment results.

[0065] Furthermore, this application simulates the setting of time and area for particulate matter PM2.5. 2.5 When determining the concentrations of its components, the method employs a combination of the WRF meteorological model and the CMAQ air quality model, rather than using only the CMAQ air quality model. In other words, the method of this invention incorporates meteorological data from the outset; and the independent variables include meteorological data, geographical and socioeconomic data. As a result, this invention not only improves the accuracy of analysis as mentioned above, but also enhances the ability to identify complex pollution sources, especially performing exceptionally well in the analysis of secondary and mixed sources.

[0066] Because of the above-mentioned features, the source apportionment results obtained by this invention can provide a more scientific basis for the accurate formulation of air pollution prevention and control policies, thereby facilitating the optimization of governance strategies, reducing governance costs, and improving governance efficiency.

[0067] The following examples illustrate the observation-constrained high spatiotemporal resolution atmospheric particulate matter source apportionment method of this invention:

[0068] 1. Simulate particulate matter PM2.5 at a specified time and location using a Common Air Quality Model (CMAQ). 2.5 and the concentration of each component

[0069] The WRF meteorological model was used in conjunction with the CMAQ air quality model to simulate PM2.5 levels in Chengdu during the winter (January) and summer (July) seasons of 2021–2023. 2.5 and PM 2.5Concentrations of each component: The WRF 4.0 model was used, with initial meteorological and boundary conditions derived from the FNL reanalysis dataset released by the National Center for Environmental Prediction (NCEP). The temporal resolution was 6 hours, and the spatial resolution was 1°×1°. The horizontal coordinates of the WRF model used the Lambert projection coordinate system, with standard latitudes set at 25N° and 47N°, and the center coordinates at (30.17073°N, 102.9891°E). The CMAQ 5.3.1 model was used, with the simulation area, nesting layers, and parameters basically consistent with the WRF model. During MCIP runtime, three grid points at each of the southeast, northwest, and northeast grid boundaries were removed to reduce model errors caused by lateral boundaries. The chemical mechanism used in the CMAQ 5.3.1 model simulation was the cb6r3_ae7_aq mechanism. The pollution source emission inventory was compiled by the ISAT model based on the MEIC inventory, including five industry sectors: industrial, transportation, power, civil, and agricultural.

[0070] 2. Assimilate the PM observed at the specified time and location. 2.5 Its component concentration data and PM obtained from CMAQ simulation 2.5 and its component concentration data

[0071] The observed PM 2.5 The PM2.5 concentration data were collected offline simultaneously from multiple locations in Chengdu during January and July of 2021-2023. 2.5 Sample analysis revealed that approximately 15 samples were collected from each location per month, with each sample taken over 22 hours at a flow rate of 100 L / min. Experimental analysis of each sample determined the daily PM2.5 concentration at each location. 2.5 and the concentration of each component.

[0072] The PM obtained from the simulation 2.5 The concentrations of PM2.5 and its components were used as independent variables, with the observed PM2.5 concentrations as the basis. 2.5 The assimilation results, with the component concentrations as dependent variables, were standardized using machine learning methods such as Generalized Additive Model (GAM), Random Forest (RF), and XGBoost, and the results of different methods were compared. The assimilation results show that the result obtained based on Random Forest is R0. 2 All values ​​were greater than 0.9 (as shown in Figure 2), and the data after RF assimilation was then used for source parsing.

[0073] 3. Perform source analysis based on the assimilated data.

[0074] PM after assimilation 2.5Concentration data of PM2.5 and its components were incorporated into the receptor model (ME2). The receptor model analyzed the sources of atmospheric particulate matter and quantified source contributions. Specifically, five factors were identified: dust sources, other sources, vehicle sources, combustion and industrial sources, and secondary sources. These five factors accounted for a significant portion of PM2.5. 2.5 The contribution rates (i.e., source contributions) were 17.71%, 8.62%, 15.71%, 13.73%, and 44.22%, respectively (as shown in Figures 3 and 4).

[0075] 4. Construct a machine learning model and use the constructed model to simulate the source contribution concentrations from different sources using downscaling.

[0076] Using the source contribution data obtained from the source apportionment as the dependent variable, and the meteorological, geographical and socio-economic, and road traffic data of Chengdu during the winter (January) and summer (July) of 2021-2023 (preprocessed by ArcGIS) as independent variables, GAM (Generalized Additive Model), Random Forest, Extreme Gradient Boosting (XGBOOST), and Bayesian spatiotemporal models were constructed using the aforementioned methods. These models were used to simulate the source apportionment results with high spatiotemporal resolution, and high-resolution source contribution distribution maps were drawn using GIS software to facilitate an intuitive understanding of the spatial distribution characteristics of pollution sources, as shown in Figures 4-7.

[0077] The embodiments of the present invention have been described in detail above, but the content described is only a preferred embodiment of the present invention and should not be considered as limiting the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the patent coverage of the present invention.

Claims

1. A method for analyzing the sources of atmospheric particulate matter with high spatiotemporal resolution based on observation constraints, characterized in that: include: Simulation of particulate matter PM2.5 in a set time and area 2.5 and the concentrations of its components; assimilated the PM2.5 observed at the specified time and location. 2.5 Its component concentration data and the simulated PM 2.5 and its component concentration data; Source analysis is performed based on the assimilated data; with the source contribution from the source analysis as the dependent variable and the meteorological data, geographical and socio-economic data, and road traffic data in each grid of the set time and set region as independent variables, a machine learning model is constructed, and the source contribution concentration from different sources is simulated by scaling down the constructed model.

2. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 1, characterized in that: The simulation set time and set area particulate matter PM 2.5 Methods for determining the concentrations of PM2.5 and its components include: simulating particulate matter PM2.5 using a combination of the WRF meteorological model and the CMAQ air quality model. 2.5 and the concentrations of its components.

3. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 1 or 2, characterized in that: The particulate matter PM 2.5 The components include water-soluble ions, potentially toxic elements, and elemental carbon (EC) and organic carbon (OC); wherein the water-soluble ions include Cl. − SO4 2- NO3 - and NH4 + The potentially toxic elements include Mg, Al, Si, Ca, K, Ti, Fe, As, Co, Cr, Cu, Pb, Zn, Mn, and Ni.

4. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 1, 2, or 3, characterized in that: The assimilation of PM observations from the set time and set region 2.5 Its component concentration data and PM obtained from CMAQ simulation 2.5 The method for obtaining PM and its component concentration data includes: using the simulated PM 2.5 The concentrations of PM2.5 and its components were used as independent variables, with the observed PM2.5 concentrations as the basis. 2.5 The concentrations of its components were used as dependent variables and standardized.

5. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 4, characterized in that: The assimilation result R of a certain machine learning method is standardized using multiple different machine learning models, and then the results are compared. 2 If all values ​​are greater than 0.9, then the data assimilated by this machine learning model will be used as the basis for subsequent source analysis.

6. The method for high spatiotemporal resolution atmospheric particulate matter source apportionment based on observation constraints according to any one of claims 1 to 5, characterized in that: The method for constructing the machine learning model includes: using machine learning methods, with the source contribution of the source analysis as the dependent variable, and meteorological data, geographical and socio-economic data, and road traffic data in each grid of the set time and set region as independent variables, to train the machine learning model; during the training process, selecting dependent and independent variables with statistical significance, and then combining them with actual physical meaning, optimizing the machine learning model based on dependent and independent variables that are both statistically significant and have actual physical meaning.

7. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 6, characterized in that: The method for selecting statistically significant dependent and independent variables is as follows: perform importance and correlation analysis on the dependent and independent variables, compare the strength of their interaction, and select statistically significant dependent and independent variables based on the analysis and comparison results.

8. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 6 or 7, characterized in that: Multiple machine learning models were constructed, and each model was used to perform downscaling simulations of source contribution concentrations from different sources.

9. The method for high spatiotemporal resolution atmospheric particulate matter source apportionment based on observation constraints according to any one of claims 1 to 8, characterized in that: The acquisition of meteorological, geographical and socioeconomic, and road traffic data within each grid of a specified time and region involves first acquiring the raw data of meteorological, geographical and socioeconomic, and road traffic data of individual grids at different scales and resolutions, and then preprocessing the raw data using GIS tools to obtain data patterns that can be used to build machine learning models.

10. The high spatiotemporal resolution atmospheric particulate matter source apportionment method based on observation constraints according to claim 9, characterized in that: The geographic and socioeconomic data include land use data, road traffic data, topographic elevation data, population density data, vegetation index data, fire point data, and points of interest data.