Satellite observation-based inversion method and device for anthropogenic nitrogen oxide emission, equipment and storage medium
Patent Information
- Application Number
- CN202610757299.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]然而,上述“自上而下”反演方法虽然具有更好的排放动态表征能力,但需要反复进行大规模数值模拟,计算资源消耗大、计算周期长,难以满足快速响应和近实时排放反演的实际需求
本发明实施例中,将卫星观测数据与掩码数据结合,生成关联训练样本,并利用质量平衡法的物理约束,确定每日的氮氧化物的排放量比值作为标签,基于关联训练样本以及标签构建样本集对初始排放量反演模型进行训练,得到训练好的排放量反演模型,这样,在模型推理时,可以直接利用排放量反演模型进行反演预测,与相关技术中依赖于大规模的数值模拟的反演方法相比,可以减少计算量,从而可以提升反演速度。
Smart Images

Figure CN122817686A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pollutant emission monitoring technology, and more specifically, to a method, apparatus, equipment, and storage medium for retrieving anthropogenic nitrogen oxide emissions based on satellite observations. Background Technology
[0002] Nitrogen oxides are important precursors to air pollution, originating widely from human activities such as fossil fuel combustion, industrial production, and transportation. Nitrogen oxides directly participate in the formation of secondary pollutants such as ozone and fine particulate matter, significantly impacting regional air quality, the ecological environment, and human health.
[0003] In related technologies, anthropogenic nitrogen oxide emissions can be obtained through emission inventories constructed using a bottom-up approach. However, this relies on statistical data, updates slowly, lacks observational constraints, and has a certain degree of uncertainty. Another top-down emission retrieval method integrates observational information. It typically relies on atmospheric chemical transport models such as the Global Atmospheric Chemical Transport Model (GEOS-Chem) and the Community Multiscale Air Quality (CMAQ) model for simulation, and combines Bayesian inversion methods such as Kalman filtering to iteratively correct emissions by minimizing the difference between the simulated model concentration and the satellite-observed concentration.
[0004] However, while the aforementioned "top-down" inversion method has a better ability to characterize emission dynamics, it requires repeated large-scale numerical simulations, consumes a lot of computational resources, and has a long calculation cycle, making it difficult to meet the actual needs of rapid response and near real-time emission inversion. Summary of the Invention
[0005] In view of this, the present invention provides a method, apparatus, device and storage medium for inverting anthropogenic nitrogen oxide emissions based on satellite observations, so as to at least solve the problems existing in related technologies.
[0006] Specifically, the present invention is achieved through the following technical solution: This invention provides a method for inverting anthropogenic nitrogen oxide emissions based on satellite observations, comprising: Acquire target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the target area during the historical target year, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical reference year prior to the historical target year, reference meteorological data for each day in the historical reference year, and reference anthropogenic nitrogen oxide emissions for each day in the historical reference year. Using the mass balance emission inversion method, based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions, the daily emission ratio of nitrogen oxides is determined; the emission ratio is used to characterize the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year to the historical reference year. The target satellite observation data is masked to obtain target mask data, and the reference satellite observation data is masked to obtain reference mask data; Based on the target satellite observation data, target meteorological data, reference satellite observation data, and reference meteorological data, an observation sample is constructed, and based on the target mask data and the reference mask data, a mask sample is constructed. Based on the observation sample and the mask sample, an associated training sample is generated, and the corresponding emission ratio is used as the sample label. A sample set is generated based on the daily associated training samples and corresponding sample labels. The initial emission inversion model is trained based on the sample set to obtain the trained emission inversion model. Acquire daily correlation data for the target inversion area and input the correlation data into the trained emission inversion model to obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data. Based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions, determine the target anthropogenic nitrogen oxide emissions for each day within the target inversion period.
[0007] The present invention also provides an inversion device for anthropogenic nitrogen oxide emissions based on satellite observations, comprising: The data acquisition module is used to acquire target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year for the target area, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical reference year prior to the historical target year, reference meteorological data of daily in the historical reference year, and reference anthropogenic nitrogen oxide emissions of daily in the historical reference year. The ratio determination module is used to determine the daily emission ratio of nitrogen oxides based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions using the mass balance emission inversion method; the emission ratio is used to characterize the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year to the historical reference year. The masking module is used to perform masking processing on the target satellite observation data to obtain target masking data, and to perform masking processing on the reference satellite observation data to obtain reference masking data; The sample generation module is used to construct observation samples based on the target satellite observation data, target meteorological data, reference satellite observation data, and reference meteorological data, and to construct mask samples based on the target mask data and the reference mask data, and to generate associated training samples based on the observation samples and the mask samples, and to use the corresponding emission ratio as the sample label. The model training module is used to generate a sample set based on the daily associated training samples and corresponding sample labels, and to train the initial emission inversion model based on the sample set to obtain the trained emission inversion model. The emission inversion module is used to acquire daily correlation data for the target inversion area, input the correlation data into the trained emission inversion model, obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data, and determine the target anthropogenic nitrogen oxide emissions for each day within the target inversion period based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions.
[0008] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations as described in any of the foregoing embodiments.
[0009] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations as described in any of the foregoing embodiments.
[0010] The present invention also provides a computer program product, including a computer program that, when run by a processor, performs the steps of any of the possible methods described above for inverting anthropogenic nitrogen oxide emissions based on satellite observations.
[0011] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: In this embodiment of the invention, satellite observation data and masked data are combined to generate associated training samples. The physical constraints of the mass balance method are used to determine the daily nitrogen oxide emission ratio as a label. Based on the associated training samples and the label, a sample set is constructed to train the initial emission inversion model, resulting in a trained emission inversion model. In this way, during model inference, the emission inversion model can be directly used for inversion prediction. Compared with inversion methods in related technologies that rely on large-scale numerical simulations, the amount of computation can be reduced, thereby improving the inversion speed.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present invention of a method for retrieving anthropogenic nitrogen oxide emissions based on satellite observations; Figure 2 This is a schematic diagram illustrating an exemplary embodiment of the present invention of an inversion process for anthropogenic nitrogen oxide emissions based on satellite observations; Figure 3 This is a flowchart illustrating the determination of the nitrogen oxide emission ratio according to an exemplary embodiment of the present invention; Figure 4 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present invention; Figure 5 This is a schematic diagram of the network structure of a downsampling network shown in an exemplary embodiment of the present invention; Figure 6 This is a schematic diagram of the model architecture of an emission inversion model according to an exemplary embodiment of the present invention; Figure 7 This is a schematic diagram illustrating a convolutional attention module according to an exemplary embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an inversion device for anthropogenic nitrogen oxide emissions based on satellite observation, as shown in an exemplary embodiment of the present invention; Figure 9 This is a hardware structure diagram of a computer device illustrated in an exemplary embodiment of the present invention. Detailed Implementation
[0014] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0015] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0016] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Nitrogen oxides are important precursors to air pollution, originating widely from human activities such as fossil fuel combustion, industrial production, and transportation. Nitrogen oxides directly participate in the formation of secondary pollutants such as ozone and fine particulate matter, significantly impacting regional air quality, the ecological environment, and human health.
[0018] In related technologies, anthropogenic nitrogen oxide emissions are mainly obtained through two techniques: "bottom-up" and "top-down." The "bottom-up" method uses statistical data to estimate nitrogen oxide emissions through activity level data and emission factors, and is currently the mainstream method for emissions accounting, with the advantage of broad coverage. However, it does not integrate atmospheric observation information, lacks dynamic characterization of diurnal emissions, and has significant uncertainties at fine spatiotemporal resolution. The "top-down" emissions retrieval method starts from atmospheric concentration observation data, typically relying on atmospheric chemical transport models such as the Global Atmospheric Chemical Transport Model (GEOS-Chem) and the Community Multiscale Air Quality (CMAQ) model for simulation, and combines Bayesian inversion methods such as Kalman filtering to iteratively correct emissions by minimizing the difference between the simulated model concentration and the satellite-observed concentration. While this type of method has high inversion accuracy, it requires repeated large-scale numerical simulations, consuming significant computational resources and having long computation cycles, making it difficult to meet the application requirements of rapid response and near-real-time emissions characterization.
[0019] To address the aforementioned problems, this invention provides a method for inverting anthropogenic nitrogen oxide emissions based on satellite observations. This method first acquires target satellite observation data of daily tropospheric nitrogen dioxide column concentrations for a target region during a historical target year, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentrations for a historical reference year prior to the target year, reference meteorological data for each day in the historical reference year, and reference anthropogenic nitrogen oxide emissions for each day in the historical reference year. Then, using a mass balance emission inversion method, based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions, a daily emission ratio of nitrogen oxides is determined. This emission ratio characterizes the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year to those in the historical reference year. Next, the target satellite observation data is masked to obtain target masked data, and the reference satellite observation data is then used to determine the daily emission ratio of nitrogen oxides. Satellite observation data is masked to obtain baseline mask data. Then, based on the target satellite observation data, target meteorological data, baseline satellite observation data, and baseline meteorological data, observation samples are constructed, and mask samples are constructed based on the target mask data and the baseline mask data. Correlation training samples are generated based on the observation samples and the mask samples, and the corresponding emission ratios are used as sample labels. A sample set is generated based on the daily correlation training samples and corresponding sample labels, and the initial emission inversion model is trained based on the sample set to obtain a trained emission inversion model. Finally, daily correlation data for the target inversion area is obtained and input into the trained emission inversion model to obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data. Based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions, the target anthropogenic nitrogen oxide emissions for each day within the target inversion period are determined.
[0020] In this embodiment of the invention, satellite observation data and masked data are combined to generate associated training samples. The physical constraints of the mass balance emission inversion method are used to determine the daily nitrogen oxide emission ratio as a label. Based on the associated training samples and the label, a sample set is constructed to train the initial emission inversion model, resulting in a trained emission inversion model. In this way, during model inference, the emission inversion model can be directly used for inversion prediction. Compared with inversion methods in related technologies that rely on large-scale numerical simulations, this can save a lot of computing resources, reduce the calculation cycle, and thus improve the inversion speed.
[0021] To facilitate understanding of this embodiment, a detailed description of the method for retrieving anthropogenic nitrogen oxide emissions based on satellite observations disclosed in this invention will be provided first. The execution entity of this method is generally a computer device, which can be a server. This server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. In other embodiments, the computer device can also be a terminal device, which can be a mobile device, terminal, handheld device, computing device, vehicle-mounted device, etc.
[0022] In other embodiments, the method can also be applied to an implementation environment consisting of computer equipment and servers, or an implementation environment consisting of terminal equipment and servers. Furthermore, this method for retrieving anthropogenic nitrogen oxide emissions based on satellite observations can also be implemented by a processor calling computer-readable instructions stored in memory.
[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0024] Please see the appendix Figure 1 The flowchart illustrates an exemplary embodiment of the present invention, showing a method for retrieving anthropogenic nitrogen oxide emissions based on satellite observations. Figure 1 As shown, the method for retrieving anthropogenic nitrogen oxide emissions based on satellite observations in this embodiment of the invention includes the following steps S101~S106: S101: Acquire target satellite observation data of daily tropospheric nitrogen dioxide column concentration for the target region during the historical target year, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentration during the historical reference year prior to the historical target year, reference meteorological data for each day during the historical reference year, and reference anthropogenic nitrogen oxide emissions for each day during the historical reference year.
[0025] In this embodiment of the invention, the target area is a national-level area. In other embodiments, the target area may also be a state-level area, a national-level area, a city cluster area, etc., which are not limited here.
[0026] Here, the historical base year refers to a historical year that serves as a reference for emission changes, such as 2019. The historical target year refers to any year or many years after the historical base year, such as 2020 or 2020-2023.
[0027] The target satellite observation data refers to the tropospheric nitrogen dioxide column concentration observed by a satellite (such as TROPOMI). Similarly, the reference satellite observation data refers to the tropospheric nitrogen dioxide column concentration observed by a satellite (such as TROPOMI) on the same day.
[0028] The target meteorological data and the baseline meteorological data may include wind speed and direction (e.g., wind speed can preferably be represented by a 10m wind field), temperature (e.g., surface temperature at 2m), surface pressure, planetary boundary layer height, relative humidity, surface shortwave radiation, and precipitation, among other variables that characterize nitrogen oxide transport, diffusion, and photochemical conditions. In some embodiments, the above meteorological data can be obtained from MERRA-2, GEOS-FP, or other equivalent meteorological products.
[0029] Anthropogenic baseline nitrogen oxide emissions refer to anthropogenic nitrogen oxide emissions in a historical baseline year, which can be obtained from MEIC, CEDS, satellite inversion products, data assimilation products, or other equivalent historical baseline year emission inventories.
[0030] It should be noted that the target satellite observation data, baseline satellite observation data, and baseline anthropogenic nitrogen oxide emissions obtained in step S101 are all grid-level data. Specifically, the target area can be divided into multiple grids (the spatial resolution of the grids should be consistent with the resolution of the atmospheric chemical transport model, for example, it could be...). The target satellite observation data and the baseline satellite observation data each include the tropospheric nitrogen dioxide column concentration observation values corresponding to multiple grids, and the baseline anthropogenic nitrogen oxide emissions include nitrogen oxide emissions corresponding to multiple grids.
[0031] S102: Using the mass balance emission inversion method, based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions, determine the daily emission ratio of nitrogen oxides; the emission ratio is used to characterize the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year and the historical reference year.
[0032] Among them, the mass balance method is a calculation method based on the principle of conservation of matter. Its core idea is that the total input of matter in the system minus the total output and loss is the actual content or emission of the target substance in the system.
[0033] Here, the ratio of nitrogen oxide emissions refers to the ratio of anthropogenic nitrogen oxide emissions from the historical target year to the same day of the historical baseline year. If the ratio of nitrogen oxide emissions is 1.2, it means that emissions have increased by 20%.
[0034] The detailed determination process for step S102 will be described in detail later.
[0035] S103: Perform masking processing on the target satellite observation data to obtain target masking data, and perform masking processing on the reference satellite observation data to obtain reference masking data.
[0036] It is understandable that some grids may have missing observations due to cloud cover or failure to meet quality requirements. Therefore, masking can be used to identify whether each grid has observations.
[0037] As mentioned earlier, since satellite observation data is grid-dimensional, the corresponding target mask data and reference mask data are both data with the same grid resolution as the satellite observation data.
[0038] Specifically, the target mask data includes target mask values corresponding to multiple grids, and the reference mask data includes reference mask values corresponding to multiple grids. For each grid, if the grid has an observation value, its mask value is the first mask value (e.g., 1); if the grid lacks an observation value, its mask value is the second mask value (e.g., 0).
[0039] Here, for grids lacking observations, after constructing the mask data, the corresponding missing positions are filled with the second mask value (i.e., 0).
[0040] S104: Based on the target satellite observation data, target meteorological data, reference satellite observation data, and reference meteorological data, construct observation samples, and based on the target mask data and the reference mask data, construct mask samples, and generate associated training samples based on the observation samples and the mask samples, and use the corresponding emission ratio as the sample label.
[0041] In this step, the observation sample refers to the multi-channel two-dimensional grid data generated from the target satellite observation data, the reference satellite observation data, the target meteorological data, and the reference meteorological data, which serves as the feature part of the model input. The mask sample refers to the binary mask with the same spatial dimension as the observation sample. By pairing the observation sample with the mask sample, an associated training sample can be obtained, and the emission ratio of the corresponding day can be used as the sample label.
[0042] In some implementations, logarithmic or piecewise logarithmic transformation methods can be used to standardize the target satellite observation data, the reference satellite observation data, and the emission ratio. The z-score method can be used to standardize the target meteorological data and the reference meteorological data, thereby enhancing the convergence of model training.
[0043] S105: Generate a sample set based on the daily associated training samples and corresponding sample labels, and train the initial emission inversion model based on the sample set to obtain the trained emission inversion model.
[0044] The initial emissions retrieval model is built on a U-Net neural network, which includes an encoder, a bottleneck layer, and a decoder.
[0045] In this way, a sample set can be generated based on the associated training samples and corresponding sample labels of each day, and supervised training can be performed on the initial emission inversion model to obtain a trained emission inversion model.
[0046] Typically, a sample set can include a training set, a validation set, and a test set. If the historical target years are 2020-2023, then the associated training samples from any number of years (e.g., 2021-2023) can be selected to generate the training set and validation set, and the associated training samples from other years (e.g., 2020) can be used to generate the test set.
[0047] The model training process for step S105 will be described in detail later (see [link]). Figure 4 ).
[0048] S106: Obtain daily correlation data for the target inversion area, input the correlation data into the trained emission inversion model, obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data, and determine the target anthropogenic nitrogen oxide emissions for each day within the target inversion period based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions.
[0049] Here, the target inversion area refers to the geographical area where emissions inversion needs to be performed. The target inversion area can be the same as or different from the target area in step S101. The target inversion area can be any of the following: intercontinental area, national area, provincial area or urban agglomeration area. No limitation is made here.
[0050] In this step, the associated data refers to the data input during the inference stage that has the same structure as the associated training samples. Specifically, the associated data includes multiple observation data and mask data associated with each observation data. The observation data includes the first satellite observation concentration data and the first meteorological data of nitrogen oxides for one day in the target inversion time period, the second satellite observation concentration data and the second meteorological data of nitrogen oxides for the corresponding day in the inversion reference time period, and the mask data includes the first mask data corresponding to the first satellite observation concentration data and the second mask data corresponding to the second satellite observation concentration data.
[0051] Here, the target inversion time period refers to the time range during which the model inference stage needs to perform inversion and emission. The target inversion time period can be one year (e.g., 2025). In other implementations, the target inversion time period can also be a time period in units of days (e.g., March 1 to March 7, 2026) or months (e.g., January to May, 2026).
[0052] Preferably, the target inversion time period can be a near real-time time period, thus enabling near real-time emission inversion.
[0053] Thus, the associated data is input into the trained emission inversion model to obtain the target emission ratio of nitrogen oxides. Based on the target emission ratio and the anthropogenic nitrogen oxide emissions during the inversion reference time period, the daily anthropogenic nitrogen oxide emissions of the target inversion region during the target inversion time period indicated in the associated data are determined, as shown in formula (1): (1) in, For grid coordinates, It refers to one day within the corresponding time period. anth The representative is the source. To determine the daily anthropogenic nitrogen oxide emissions within the target inversion period, The ratio of the target emissions, This represents the baseline anthropogenic nitrogen oxide emissions for each day within the baseline time period.
[0054] In this embodiment of the invention, satellite observation data and mask data are combined to generate associated training samples. Using mass balance constraints, the daily nitrogen oxide emission ratio is determined as a label. Based on the associated training samples and the label, a sample set is constructed to train the initial emission inversion model, resulting in a trained emission inversion model. In this way, during model inference, the emission inversion model can be directly used for inversion prediction. Compared with inversion methods in related technologies that rely on large-scale numerical simulations, the amount of computation can be reduced, thereby improving the inversion speed.
[0055] In some implementations, the process of training the initial emission inversion model based on the sample set to obtain a trained emission inversion model also includes a model fine-tuning stage. Specifically, Based on the sample set, the initial emission inversion model is trained to obtain an intermediate model. Fine-tuning correlation training samples are obtained, and the model parameters of the intermediate model are adjusted based on the fine-tuning correlation training samples to obtain a trained emission inversion model.
[0056] Here, in order to eliminate the data distribution offset caused by the inconsistency between the meteorological driving sources in the training phase and the operational phase, this invention further performs transfer adaptation to achieve model fine-tuning after offline training based on the sample set.
[0057] Specifically, the third meteorological data for each day of the historical target year and the fourth meteorological data for each day of the historical baseline year can be obtained from near real-time meteorological products (such as GEOS-FP, ERA5T or other near real-time meteorological products). Then, following the steps of constructing associated training samples mentioned above, fine-tuned associated training samples are constructed based on the third meteorological data and the fourth meteorological data.
[0058] During fine-tuning, while keeping the main network structure of the intermediate model unchanged, the parameters of the first two encoders or some convolutional layers are frozen, and only the Convolutional Block Attention Module (CBAM), decoder and normalization parameters are updated, so as to quickly correct near-real-time meteorological biases while retaining historical emission knowledge.
[0059] In some implementations, the learning rate parameter used for the fine-tuning process is typically smaller than the learning rate parameter used during the training phase, for example... to .
[0060] Through the above migration process, the trained model can directly receive near-real-time satellite observations and near-real-time meteorological data, and quickly output daily anthropogenic nitrogen oxide emissions, thereby avoiding the need to re-execute large-scale simulations of atmospheric chemical transport models (such as GEOS-Chem) for each inversion, and realizing near-real-time operational applications.
[0061] In some implementations, the target satellite observation data described in step S101 can be determined through the following steps (I) to (III): (I) Obtain the first raw satellite observation data of the daily tropospheric nitrogen dioxide column concentration in the historical target year, and resample the first raw satellite observation data according to the spatial resolution of the grid to obtain the resampled first satellite observation data.
[0062] Here, the first raw satellite observation data refers to the raw data of tropospheric nitrogen dioxide column concentration directly detected and retrieved by satellites (such as the TROPOMI satellite), usually with raw spatial resolution (e.g., ...). ) and quality labels (such as cloud coverage, observation quality, etc.).
[0063] After acquiring the first raw satellite observation data, the first raw satellite data is converted from the raw resolution to the target grid resolution (e.g., This is the resampling process, which yields the first satellite observation data after resampling.
[0064] Common resampling methods include bilinear interpolation and nearest neighbor area weighted averaging, with the aim of making satellite data have the same spatial grid as baseline emissions, meteorological data, etc.
[0065] In some implementations, before resampling the first raw satellite observation data, the first raw satellite observation data may be subject to quality control screening. Specifically, data with low reliability and high error may be removed based on at least one condition, such as observation quality (qa_value) and / or cloud coverage (cloud_fraction).
[0066] Specifically, a higher observation quality value indicates greater reliability. For example, the Tropomi product recommends using pixels with a qa_value > 0.75. Clouds can block the sensor from detecting near-surface air columns, leading to biases in the inversion of tropospheric nitrogen dioxide column concentration. Data with excessive cloud cover usually needs to be removed (e.g., data with cloud_fraction > 0.3).
[0067] After filtering and resampling under at least one of the above conditions, invalid grid data is marked as missing (mask value set to 0). This step ensures that the satellite observation data entering the subsequent process has high reliability and avoids the deviation caused by low-quality data to the inversion results.
[0068] (II) The first satellite observation data after resampling is taken based on a preset sliding time window, and the average value of the first satellite observation data within the sliding time window is determined as the satellite observation data of the tropospheric nitrogen dioxide column concentration on the central day of the sliding time window.
[0069] Here, the preset time window refers to a fixed-length time interval (e.g., 7 days, 11 days, etc.). The preset time window is slid along the time axis day by day, with each window covering a continuous number of days (including the center day and several days before and after). This is used to statistically analyze the satellite observations within the preset time window (e.g., take the average value) to smooth out noise and fill in gaps on single days. In other words, the average value of the satellite observation data within the sliding time window is determined as the satellite observation data of the tropospheric nitrogen dioxide column concentration on the center day of the sliding time window.
[0070] (III) Based on the satellite observation data of tropospheric nitrogen dioxide column concentration for each central day, generate target satellite observation data of tropospheric nitrogen dioxide column concentration for each day of the historical target year.
[0071] After resampling and sliding window averaging, the final generated daily satellite tropospheric nitrogen dioxide column concentration data for the historical target year is the target satellite observation data of the daily tropospheric nitrogen dioxide column concentration in the historical target year.
[0072] In this embodiment, the above pretreatment process has the following beneficial effects: (1) Improve data spatial consistency: By resampling, satellite data is unified to the target grid, ensuring strict spatial alignment with input data such as emission inventories and meteorological fields, avoiding feature misalignment or calculation errors caused by resolution differences, and providing a reliable observation benchmark for grid-by-grid emission inversion.
[0073] (2) Effectively reduce random sampling noise: Random errors such as instrument noise and inversion uncertainty are unavoidable in satellite observation. The sliding window averaging uses information from multiple days to smooth out random fluctuations, making the concentration value of the central day of the window more stable and reliable, which is conducive to the model learning the real emission concentration relationship rather than noise interference.
[0074] (3) Significantly fills the observation gaps caused by cloud cover: Cloud cover is a major interference factor in satellite remote sensing, often leading to large-area missing daily data. By using sliding window averaging, as long as there are some valid observations within the window, a central day estimate can be generated by averaging, which significantly reduces the number of missing pixels, improves the utilization rate of training samples and the adaptability of the model to real scenes.
[0075] Similarly, the reference satellite observation data mentioned in step S101 can also be determined in a manner similar to steps (I) to (III) above, specifically including (i) to (iii): (i) Obtain the second raw satellite observation data of the daily tropospheric nitrogen dioxide column concentration in the reference target year, and resample the second raw satellite observation data according to the spatial resolution of the grid to obtain the resampled second satellite observation data.
[0076] (ii) The second satellite observation data after resampling is taken based on a preset sliding time window, and the average value of the satellite observation data within the sliding time window is determined as the satellite observation data of the tropospheric nitrogen dioxide column concentration on the central day of the sliding time window.
[0077] (iii) Based on satellite observation data of tropospheric nitrogen dioxide column concentration for each central day, generate baseline satellite observation data of daily tropospheric nitrogen dioxide column concentration for the historical baseline year.
[0078] The details of steps (i) to (iii) are similar to those of steps (I) to (III), and will not be repeated here.
[0079] Please see Figure 2 This is a schematic diagram illustrating an inversion process for anthropogenic nitrogen oxide emissions based on satellite observations, provided as an exemplary embodiment of the present invention. Figure 2 As shown, the overall inversion process includes sample set construction, model training, model inference, and model transfer.
[0080] For sample set construction: Obtain raw satellite observation data (meteorological data, emissions, and satellite observation data) for historical baseline years and historical target years. Perform quality control and preprocessing on the raw observation data to obtain data with controllable quality standards (i.e., target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year, daily target meteorological data, baseline satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical baseline years prior to the historical target year, baseline meteorological data in the historical baseline year, and baseline anthropogenic nitrogen oxide emissions in the historical baseline year). Then, perform masking processing to obtain target mask data and baseline mask data. Based on the data with controllable quality standards and the mask data, construct associated training samples to generate a sample set, and divide the sample set into a training set, a test set, and a validation set.
[0081] For model training: read the associated training samples from the training set, input the associated training samples into the initial emission inversion model for model prediction, calculate the model loss, perform backpropagation to adjust the model parameters, and use the validation set to evaluate the model to obtain the intermediate model.
[0082] During training, the optimizer is used to update the model parameters, and the iteration stops when the training requirements are met.
[0083] For model fine-tuning, near-real-time meteorological data is acquired, and fine-tuning correlation training samples are constructed based on the near-real-time meteorological data. Then, the intermediate model is fine-tuned based on the fine-tuning correlation training samples to obtain a trained emission inversion model.
[0084] For model inference, near real-time correlation data is obtained and input into the trained emission inversion model for prediction, to obtain the target emission ratio of nitrogen oxides per day within the target inversion period, and then the target anthropogenic nitrogen oxide emission amount per day within the target inversion period is determined.
[0085] Please see Figure 3 This is a flowchart illustrating the determination of a nitrogen oxide emission ratio, provided as an exemplary embodiment of the present invention. Figure 3As shown, for step S102, when determining the daily emission ratio of nitrogen oxides based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions using the mass balance emission inversion method, steps S1021~S1023 are included: S1021: Using an atmospheric chemical transport model, concentration simulation is performed based on the baseline anthropogenic nitrogen oxide emissions and the target meteorological data to obtain the target tropospheric nitrogen dioxide column concentration; the target tropospheric nitrogen dioxide column concentration is used to characterize the influence of the meteorological conditions corresponding to the target meteorological data on the tropospheric nitrogen dioxide column concentration when the baseline anthropogenic nitrogen oxide emissions remain unchanged.
[0086] S1022: Using the atmospheric chemical transport model, concentration simulation is performed based on the baseline anthropogenic nitrogen oxide emissions and the baseline meteorological data to obtain the baseline tropospheric nitrogen dioxide column concentration.
[0087] Among them, atmospheric chemical transport models (such as GEOS-Chem) are numerical tools used to simulate the emission, transport, chemical transformation and deposition processes of chemical substances in the atmosphere (such as ozone, aerosols, nitrogen oxides, etc.). Based on given emission sources, meteorological fields and initial / boundary conditions, they extrapolate changes in pollutant concentrations according to physicochemical laws.
[0088] In this step, a fixed baseline of human-origin nitrogen oxide emissions is used as input. Historical meteorological data for both the target year and the baseline year are then input, and two independent forward simulations are performed. This yields two simulated concentrations: the target tropospheric nitrogen dioxide column concentration (reflecting the tropospheric nitrogen dioxide column concentration under the target meteorological data conditions and with constant emissions) and the baseline tropospheric nitrogen dioxide column concentration (reflecting the tropospheric nitrogen dioxide column concentration under the baseline meteorological data conditions). The ratio of these two simulated concentrations can quantify the impact of meteorological changes on the tropospheric nitrogen dioxide column concentration, thus separating it from the total changes observed by satellite in subsequent steps.
[0089] The above simulation process is shown in formulas (2) to (3): (2) in, For the first t Latitude and longitude coordinates of the day Reference tropospheric nitrogen dioxide column concentration, This is a forward simulation model of the Geos-Chem atmospheric chemistry model. As the baseline meteorological data, The baseline anthropogenic nitrogen oxide emissions for the historical baseline year.
[0090] (3) in, The target tropospheric nitrogen dioxide column concentration, For target meteorological data.
[0091] S1023: Based on the target satellite observation data, the target tropospheric nitrogen dioxide column concentration, the reference satellite observation data, and the reference tropospheric nitrogen dioxide column concentration, the emission ratio of the nitrogen oxides is determined using the mass balance emission inversion method.
[0092] Specifically, it may include the following steps (1) to (3): (1) The ratio of the target satellite observation data to the reference satellite observation data is determined as the observation factor.
[0093] Here, the observation factor is determined based on the ratio of the target satellite observation data to the reference satellite observation data, which can reflect the changes in the target satellite observation data compared to the reference satellite observation data.
[0094] (2) The ratio of the target tropospheric nitrogen dioxide column concentration to the reference tropospheric nitrogen dioxide column concentration is determined as the meteorological effect influencing factor; the meteorological effect influencing factor is used to characterize the influence of changes in meteorological conditions on the tropospheric nitrogen dioxide column concentration.
[0095] Here, the meteorological effect influencing factor is determined based on the ratio of the target tropospheric nitrogen dioxide column concentration to the baseline tropospheric nitrogen dioxide column concentration, and is used to characterize the influence of meteorological conditions on the tropospheric nitrogen dioxide column concentration.
[0096] (3) The difference between the observed factor and the meteorological effect influencing factor is determined as the net change in column concentration on the same day between the historical target year and the historical base year. Based on the net change in column concentration and the preset response coefficient of tropospheric nitrogen dioxide column concentration change to anthropogenic emission change, the emission ratio of nitrogen oxides is determined.
[0097] In this way, based on the difference between the observed factors and the meteorological effect factors, the relative change in tropospheric nitrogen dioxide column concentration caused solely by emission changes can be obtained, that is, the net change in column concentration on the same day between the historical target year and the historical baseline year.
[0098] Furthermore, the emission ratio of nitrogen oxides can be determined based on the net change in column concentration and the preset response coefficient of tropospheric nitrogen dioxide column concentration change to emission change.
[0099] For the process of determining the above nitrogen oxide emission ratio, please refer to formulas (4) to (5): (4) in, This represents the net change in column concentration.
[0100] (5) in, This represents the ratio of nitrogen oxide emissions. This represents the preset response coefficient of tropospheric nitrogen dioxide column concentration to changes in emissions. The baseline anthropogenic nitrogen oxide emissions can be determined through positive and negative perturbation sensitivity experiments, historical sample fitting, or table lookup using GEOS-Chem, and no specific limitations are imposed here.
[0101] In this embodiment of the invention, by using a simulation strategy that fixes emissions and only replaces the meteorological field, the contribution of meteorological conditions (wind, temperature, humidity, boundary layer height, etc.) to the tropospheric nitrogen dioxide column concentration can be isolated, avoiding misjudging meteorological fluctuations as emission changes and improving the physical accuracy of emission inversion.
[0102] Please see Figure 4 This is a flowchart of model training provided as an exemplary embodiment of the present invention. Figure 4 As shown, for step S105, when training the initial emission inversion model based on the sample set to obtain the trained emission inversion model, steps S1051~S1054 are included: S1051: For each associated training sample, the encoder is used to encode the associated training sample to obtain encoded features.
[0103] In this embodiment, the encoder adopts a four-level downsampling network with 64, 128, 256 and 512 channels respectively. Each level of the downsampling network includes two partial convolutional layers, a normalization layer and a nonlinear activation layer.
[0104] In this step, the encoder first downsamples each associated training sample step by step. Through a four-level downsampling network, the encoder gradually compresses the spatial size and increases the channel dimension. At the same time, it retains effective observation location information through a mask propagation mechanism, so that the final output encoded features contain both the details of local emission hotspots and the abstract semantics of regional transport background.
[0105] Specifically, for step S1051, the following steps are included: (1) to (4): (1) Input the associated training samples into the primary downsampling network, perform partial convolution processing on the observed samples based on the mask samples to obtain intermediate features, and update the mask samples based on the preset mask update mechanism to obtain updated mask samples.
[0106] As can be seen from step S1051, each downsampling network includes two partial convolutional layers. Therefore, here, the partial convolution processing of the observed samples based on the mask samples is the processing process of the first partial convolutional layer, which yields intermediate features.
[0107] Furthermore, after obtaining the intermediate features, the mask sample is also updated synchronously, resulting in an updated mask sample.
[0108] Specifically, it may include the following steps (1.1) to (1.4): (1.1) Slide the convolution kernel of the preset size under the grid dimension of the observed sample, and determine whether the sum of the mask values of each grid in the convolution region currently covered by the convolution kernel is greater than 0.
[0109] Here, according to the principle of convolution, the convolution result is for the center position of the convolution kernel. In order to enable the convolution kernel to cover every grid, the observation sample can be pre-filled with 0 before sliding the convolution kernel of the preset size in the grid dimension of the observation sample. Then, the observation sample after filling is slid.
[0110] In this step, the primary downsampling network refers to a downsampling network with 64 channels.
[0111] In this embodiment, the convolution kernel of the preset size is... In other implementations, the size of the convolution kernel can be set according to actual needs, for example... The specific value depends on the grid resolution and the desired receptive field, which is not limited here.
[0112] The phrase "sliding the convolution kernel of a preset size along the grid dimension" means that the convolution kernel moves gradually along the two-dimensional grid (longitude and latitude), and each movement covers a new convolution region.
[0113] As mentioned above, the mask value includes a first mask value or a second mask value, wherein the first mask value indicates that the observation value of the grid is not empty, and the second mask value indicates that the observation value of the grid is empty.
[0114] The sum of the mask values within the convolution region refers to the cumulative sum of the mask values (binary: 1 represents a valid observation, 0 represents a missing observation) of all grid cells within the convolution region. Its value ranges from 0 to the total number of pixels in the convolution kernel (e.g., ...). (0~9), and greater than 0 indicates that at least one grid in the region has a valid observation.
[0115] (1.2) If so, perform convolution processing on the observations of the target grid with the mask value of the first mask value in the convolution region to obtain the convolution result.
[0116] In this step, if the sum is greater than zero (i.e., at least one valid observation exists), convolution is performed: only observations on the grid marked with a mask of 1 are used in the convolution operation, ignoring missing locations marked with a mask of 0. In actual calculation, the weighted sum of valid observations is usually scaled (multiplied by the ratio of the total number of pixels in the convolution kernel to the number of valid pixels) to compensate for the output amplitude attenuation caused by the reduction in valid pixels.
[0117] In other implementations, if the sum of the mask values within the convolution region is zero (indicating that there are no valid observations in each grid within the convolution region), the convolution result is directly output as 0, and the mask values of each grid within the convolution region are also set to 0.
[0118] (1.3) Based on the convolution result, the sum of the mask values of each grid in the convolution region, the convolution kernel, the sum of all-1 matrices with the same size as the convolution kernel, and the bias coefficient, determine the encoding value of the center grid of the convolution region.
[0119] As shown in formula (6): (6) in, The coordinates (latitude and longitude coordinates) of the center grid of the convolution region. That is, the encoded value of the center grid of the convolution region. For the input observations, The input mask. This is the transpose of the convolution kernel. The bias coefficient, It is the summation of a matrix of all ones with the same size as the convolution kernel. For example, if the convolution kernel size is... ,but Take 9, It is the summation of the masks corresponding to the convolution kernels.
[0120] According to the above formula (6), the core feature of partial convolution is that convolution is only performed when there are valid observed pixels (mask sum greater than 0) in the window.
[0121] (1.4) Determine the intermediate features based on the encoding values corresponding to the center grids of each convolutional region.
[0122] In this way, after the sliding ends, the center grid of each convolution region is determined to have a corresponding encoding value. Thus, based on the encoding values corresponding to the center grid of each convolution region, the result of the first partial convolution processing can be obtained, which is the intermediate feature.
[0123] Simultaneously, as the encoding values of each grid are updated, the mask values corresponding to each grid also need to be updated. Specifically, based on a preset mask update mechanism, when updating the mask sample to obtain the updated mask sample, if the sum of the mask values of each grid in the convolution region currently covered by the convolution kernel is greater than 0, the mask value of the center grid of the convolution region is set to 1; if the sum of the mask values of each grid in the convolution region is not greater than 0, the mask value of the center grid of the convolution region is set to 0.
[0124] Finally, the updated mask sample is determined based on the mask value of the center grid of each convolutional region.
[0125] As shown in formula (7), if the sum of the mask values in the convolution region is greater than 0, then the mask value of each grid in the convolution region is set to 1, otherwise it is set to 0.
[0126] (7) in, This is the mask value for the center grid within the updated convolution region.
[0127] (2) Based on the updated mask sample, perform partial convolution processing on the intermediate features to obtain the first-level encoded features, and update the updated mask sample based on the mask update mechanism to obtain the first-level mask sample.
[0128] In this step, since the convolution process of the two partial convolutions is the same, only the input and output data are different, that is, for the second partial convolution, the output of the first partial convolution and its corresponding mask are used as input, and the second partial convolution is performed again. The convolution kernel also slides in space to reprocess the intermediate features obtained from the first partial convolution to obtain the first-level encoded features. The specific processing process will not be described in detail.
[0129] Similarly, the process of updating the first-level mask sample based on the mask update mechanism is similar to the process of obtaining the updated mask sample, and will not be described in detail here.
[0130] Thus, after two partial convolutions, the first-level encoded features and the corresponding first-level mask samples can be obtained.
[0131] (3) The first-level coding features and the first-level mask samples are downsampled respectively to obtain the first-level network output results.
[0132] Here, downsampling is used to increase the number of channels, so that the output of the first-level network can be obtained. In this embodiment, the output of the first-level network is the encoded features and mask samples with 128 channels.
[0133] (4) Input the output of the first-level network into the next-level downsampling network and pass it level by level until all levels of the downsampling network have completed the convolution process to obtain the encoded features.
[0134] Then, the output of the first-level network is input into the next-level downsampling network, which performs the same processing flow as the first-level downsampling network (including bipartial convolution and downsampling).
[0135] Following this pattern, the output features pass through the third level (256 channels) and the fourth level (512 channels) sequentially. Each level performs a sliding window, two partial convolutions, mask updates, and optional downsampling operations. Finally, the deepest layer (fourth level) outputs the encoded features, which have the smallest spatial resolution (e.g., 1 / 16 of the original input), the largest number of channels (512), and carry mask information that has been propagated and gradually filled in through multiple layers. This encoded feature will serve as the input to the bottleneck layer, used to further extract large-scale contamination transmission background and mitigate gradient decay.
[0136] Please see Figure 5 This is a schematic diagram of a downsampling network structure provided as an exemplary embodiment of the present invention. Figure 5 As shown, the downsampling network includes a first part of convolutional layers and a second part of convolutional layers. Both the first part of the convolutional layers and the second part of the convolutional layers include components based on convolutional kernels (…). The network performs partial convolutions, batch normalization (BN), and activation functions (such as LeakyReLU, ReLU, etc.). Finally, the downsampling network outputs the encoded features of the corresponding layer.
[0137] In this implementation, through step-by-step propagation, the encoder gradually abstracts from local details to regional semantics, enabling the model to simultaneously perceive small-scale emission hotspots and large-scale pollution plumes. Meanwhile, the mask follows the features... Figure 1 The data is passed down level by level and updated at each layer to ensure that the deep network still knows which grid observations are reliable.
[0138] Furthermore, since some convolutions gradually fill in the missing measurement regions as the layers progress, the network is able to learn potential emission patterns in a continuous spatial field while maintaining the observation confidence level.
[0139] S1052: Using the bottleneck layer, perform double convolution processing on the encoded features to obtain bottleneck layer encoded features, and sum the residuals of the encoded features and the bottleneck layer encoded features to obtain bottleneck layer output features.
[0140] In this step, the bottleneck layer is the intermediate layer connecting the encoder and decoder, located at the deepest part of the network. Its core function is to further process the encoded features, expand the receptive field, and alleviate gradient decay.
[0141] In this embodiment, the bottleneck layer adopts a double convolution and residual summation structure. This design retains some of the ability of convolution to process missing data, and also alleviates the gradient vanishing problem through the residual path.
[0142] That is, in this step, the bottleneck layer is used to perform double convolution on the encoded features output by the encoder to obtain the bottleneck layer encoded features. Then, the encoded features and the bottleneck layer encoded features are summed by residuals to obtain the bottleneck layer output features. The residual connection allows the gradient to directly bypass the convolution operation and propagate backward, effectively alleviating the gradient decay problem of deep networks.
[0143] S1053: Using the decoder, the output features of the bottleneck layer are decoded to obtain the predicted emission ratio.
[0144] It can be understood that the decoder is the sub-network responsible for gradually recovering the spatial resolution from the output features of the bottleneck layer and generating emission ratio predictions.
[0145] In this embodiment, the decoder consists of a multi-level upsampling network. Each level of the upsampling network includes operations such as upsampling, skip connections (fusion with the corresponding layer features of the encoder), attention mechanism, and partial convolution. The multi-level upsampling network corresponds one-to-one with the multi-level downsampling network in the encoder (with the same number of channels).
[0146] Please see Figure 6 This is a schematic diagram of the model architecture of an emission inversion model provided as an exemplary embodiment of the present invention. Figure 6 As shown, the system includes an encoder, a decoder, and a bottleneck layer. First, the associated training samples are input into the encoder. The multi-layer downsampling network in the encoder encodes the samples sequentially until the 512-channel downsampling network outputs the encoded features. Then, the encoded features output by the encoder are input into the bottleneck layer to obtain the bottleneck layer output features. The bottleneck layer output features are then input into the decoder. The multi-layer upsampling network in the decoder decodes the samples sequentially (a channel spatial attention mechanism is used to achieve skip connections during the decoding process) until the predicted emission ratio is output.
[0147] Specifically, regarding step S1053, when using the decoder to decode the output features of the bottleneck layer to obtain the predicted emission ratio, it can be achieved through the following steps (1) to (4): (1) The bottleneck layer output features are upsampled using the lowest level upsampling network to obtain upsampled features.
[0148] Here, the lowest level upsampling network of the decoder is a 256-channel upsampling network.
[0149] Please continue reading Figure 5 Since the bottleneck layer output features are still the corresponding 512 channels, after the bottleneck layer output features are input into the decoder, they will be directly upsampled using a 256-channel upsampling network to obtain upsampled features.
[0150] (2) Obtain the coding features of the corresponding level in the encoder, and use the convolutional attention module to enhance the coding features of the corresponding level to obtain enhanced coding features.
[0151] Here, for the current decoder level (the lowest level), the same level (i.e., 256 levels) of coding features are extracted from the encoder. Channel attention (learning the importance of each channel) and spatial attention (learning the importance of each position) are applied to the extracted coding features (256 channels) in sequence, and the output is an enhanced coding feature of the same size as the input (still 256 channels, with the same spatial resolution).
[0152] (3) The upsampling features and the enhanced coding features are fused to obtain fused features, and the fused features are passed sequentially level by level until all levels of the upsampling network have completed the decoding process to obtain the predicted emission ratio.
[0153] This step, fusing the upsampled features and the enhanced coding features, can refer to concatenating the upsampled features and the enhanced coding features to obtain a fused feature. This fused feature is then used as the input to the next-level upsampled network, and the above process is repeated. Once the resolution is restored to the original input size, the decoded features are then processed in the final stage. The convolutional mapping is converted into a single-channel feature map, and a non-negative activation function (such as Softplus or ReLU) is applied to finally output the predicted emission ratio.
[0154] It should be noted that, in repeating the above process, the previous level upsampling network will use 128 channels, 64 channels, etc., and will obtain features from the corresponding encoder levels (128, 64) and perform CBAM enhancement respectively.
[0155] Here, the deep semantics from the bottleneck layer are fused with the shallow details from the encoder to perform multi-scale information fusion, so that the predicted emission ratio of the final output conforms to the large-scale transmission law and retains the local hotspots.
[0156] Please see Figure 7 This is a schematic diagram of a convolutional attention module provided in an embodiment of the present invention. Figure 7As shown, the Convolutional Attention Module (CBAM) employs channel attention and spatial attention mechanisms. For encoded features, the channel attention mechanism is used to process them to obtain channel attention weights. The channel attention weights are then multiplied with the encoded features channel by channel to obtain channel-enhanced features. The channel-enhanced features are then processed using the spatial attention mechanism to obtain spatial attention weights. The spatial attention weights are then multiplied with the channel-enhanced features spatially to obtain the decoded features for that level.
[0157] S1054: Based on the predicted emission ratio and the corresponding sample labels, determine the model loss, and adjust the model parameters of the initial emission inversion model based on the model loss to obtain the emission inversion model.
[0158] In this step, after obtaining the predicted emission ratio, the model loss can be determined based on the predicted emission ratio and the corresponding sample labels. Then, the model parameters of the initial emission inversion model are adjusted based on the model loss to obtain the emission inversion model.
[0159] Here, the predicted emission ratio output by the model and the sample labels (the true values calculated by the mass balance emission inversion method) can be input into the composite loss function to determine the model loss.
[0160] Specifically, it includes the following steps (A) to (D): (A) Determine the relative error loss based on the relative error between the predicted emission ratio and the corresponding sample labels.
[0161] As shown in formula (8), the expression for the relative error loss is: (8) in, For relative error loss, To predict the emissions ratio, For sample labels, To prevent constants with a denominator of zero.
[0162] This relative error loss is used to constrain the relative error across the entire region, preventing high-emission grids from completely dominating the training process.
[0163] (B) Determine the emission contribution loss based on the relative contribution weights of the observation data of each grid in the target region and the relative error.
[0164] Considering that anthropogenic nitrogen oxide emissions have obvious hotspot characteristics, this embodiment constructs an emission contribution weighted loss. To impose stronger constraints on high-emission grids and areas near hotspots, the expression for the emission contribution loss is shown in Equation (9): (9) in, Contributing to losses due to emissions, Indicates the first The relative contribution weight of the grid in the total regional emissions on that day.
[0165] (10) in, For grid Nitrogen oxide emissions.
[0166] This weight design avoids the model from excessively pursuing low error in the background area during training, thus ignoring key emission areas such as urban clusters, industrial zones, and major transportation routes.
[0167] (C) Based on the difference between the predicted emission ratio and the first-order difference of the sample labels in the latitude and longitude directions, determine the spatial gradient consistency loss.
[0168] In this step, spatial gradient consistency loss is introduced. The first-order difference between the predicted emission ratio and the sample label in the latitude and longitude directions is constrained as shown in formula (11): (11) For spatial gradient consistency loss, The first gradient in the longitude direction. It represents the first-order gradient in the latitudinal direction.
[0169] (D) Determine the model loss based on the relative error loss, the emission contribution loss, and the spatial gradient consistency loss.
[0170] Here, an adaptive uncertainty weighting method can be used for fusion to obtain the model loss, as shown in formula (12): (12) in, For model loss, Learnable scalar parameters trained in sync with network parameters, used to automatically adjust the weights of each loss term. To avoid the subjectivity brought about by human experience-based empowerment For the first n Each loss item (including) , , ), This is a regularization term.
[0171] In this implementation, the dimensions and numerical ranges of the various losses may differ greatly. Through adaptive weights, the weights can be dynamically adjusted according to the convergence of each loss during the training process, thus avoiding any one loss from dominating the total loss.
[0172] In actual training, the model is preferably trained using the Adam or AdamW optimizer, and the total loss function is minimized using the backpropagation algorithm. The initial learning rate is preferably 1e. -4 up to 5e -4 The batch size is 4 to 16, the training rounds are 50 to 200, and the convergence stability can be improved by combining learning rate decay strategy, gradient pruning and weight decay.
[0173] To improve adaptability to real-world scenarios with missing data, random block occlusion enhancement can be applied to the input mask during the training phase, or a small noise perturbation can be applied to the nitrogen dioxide observation.
[0174] Meanwhile, R-squared (R²) is monitored in real time on the independent validation set. The system measures indicators such as root mean square error (RMSE), mean absolute error (MAE), total regional error, and high-emission zone error. When the verification loss no longer decreases within a preset number of rounds, it triggers early termination and saves the optimal model parameters.
[0175] Corresponding to the aforementioned embodiments of the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations, the present invention also provides embodiments of an apparatus for inverting anthropogenic nitrogen oxide emissions based on satellite observations.
[0176] Please see Figure 8 This is a schematic diagram illustrating the structure of an inversion device for anthropogenic nitrogen oxide emissions based on satellite observations, as shown in an exemplary embodiment of the present invention. Figure 8 As shown, the inversion device 800 for anthropogenic nitrogen oxide emissions based on satellite observations includes: The data acquisition module 810 is used to acquire target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year for the target area, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical reference year prior to the historical target year, reference meteorological data of daily in the historical reference year, and reference anthropogenic nitrogen oxide emissions of daily in the historical reference year. The ratio determination module 820 is used to determine the daily emission ratio of nitrogen oxides based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions using the mass balance emission inversion method; the emission ratio is used to characterize the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year and the historical reference year. The masking processing module 830 is used to perform masking processing on the target satellite observation data to obtain target masking data, and to perform masking processing on the reference satellite observation data to obtain reference masking data; The sample generation module 840 is used to construct observation samples based on the target satellite observation data, target meteorological data, reference satellite observation data, and reference meteorological data, and to construct mask samples based on the target mask data and the reference mask data, and to generate associated training samples based on the observation samples and the mask samples, and to use the corresponding emission ratio as the sample label. The model training module 850 is used to generate a sample set based on the daily associated training samples and corresponding sample labels, and to train the initial emission inversion model based on the sample set to obtain the trained emission inversion model. The emission inversion module 860 is used to acquire daily correlation data for the target inversion area, input the correlation data into the trained emission inversion model, obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data, and determine the target anthropogenic nitrogen oxide emissions for each day within the target inversion period based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions.
[0177] In some embodiments, the ratio determination module 820 is specifically used for: Using an atmospheric chemical transport model, concentration simulations were performed based on the baseline anthropogenic nitrogen oxide emissions and the target meteorological data to obtain the target tropospheric nitrogen dioxide column concentration. This target tropospheric nitrogen dioxide column concentration characterizes the impact of the meteorological conditions corresponding to the target meteorological data on the tropospheric nitrogen dioxide column concentration, assuming the baseline anthropogenic nitrogen oxide emissions remain constant. Using the atmospheric chemical transport model, a concentration simulation was performed based on the baseline anthropogenic nitrogen oxide emissions and the baseline meteorological data to obtain the baseline tropospheric nitrogen dioxide column concentration; Based on the target satellite observation data, the target tropospheric nitrogen dioxide column concentration, the reference satellite observation data, and the reference tropospheric nitrogen dioxide column concentration, the daily emission ratio of nitrogen oxides is determined using the mass balance method.
[0178] In some embodiments, the ratio determination module 820 is specifically used for: The ratio of the target satellite observation data to the reference satellite observation data is determined as the observation factor; The ratio of the target tropospheric nitrogen dioxide column concentration to the baseline tropospheric nitrogen dioxide column concentration is determined as the meteorological effect influencing factor; the meteorological effect influencing factor is used to characterize the impact of changes in meteorological conditions on the tropospheric nitrogen dioxide column concentration. The difference between the observed factor and the meteorological effect influencing factor is determined as the net change in column concentration on the same day between the historical target year and the historical baseline year. Based on the net change in column concentration and the preset response coefficient of tropospheric nitrogen dioxide column concentration change to anthropogenic emission change, the daily emission ratio of nitrogen oxides is determined.
[0179] In some implementations, the target area is divided into multiple grids; when acquiring target satellite observation data on daily tropospheric nitrogen dioxide column concentrations in historical target years, the data acquisition module 810 is specifically used for: The raw satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year are obtained, and the raw satellite observation data is resampled according to the spatial resolution of the grid to obtain the resampled satellite observation data. The resampled satellite observation data is evaluated based on a preset sliding time window, and the average value of the satellite observation data within the sliding time window is determined as the satellite observation data of the tropospheric nitrogen dioxide column concentration on the central day of the sliding time window. Based on satellite observation data of tropospheric nitrogen dioxide column concentration for each central day, target satellite observation data of daily tropospheric nitrogen dioxide column concentration for the historical target year are generated.
[0180] In some embodiments, the initial emission inversion model is a U-Net neural network, which includes an encoder, a bottleneck layer, and a decoder; the model training module 850 is specifically used for: For each associated training sample, the encoder is used to encode the associated training sample to obtain encoded features; Using the bottleneck layer, the encoded features are subjected to double convolution to obtain the bottleneck layer encoded features, and the residuals of the encoded features and the bottleneck layer encoded features are summed to obtain the bottleneck layer output features. The decoder is used to decode the output features of the bottleneck layer to obtain the predicted emission ratio. Based on the predicted emission ratio and the corresponding sample labels, the model loss is determined, and the model parameters of the initial emission inversion model are adjusted based on the model loss to obtain the emission inversion model.
[0181] In some embodiments, the encoder includes a multi-level downsampling network, each level of which employs a two-layer partial convolution mechanism; the model training module 850 is specifically used for: The associated training samples are input into the primary downsampling network. Based on the mask samples, the observed samples are partially convolutional to obtain intermediate features. The mask samples are then updated based on a preset mask update mechanism to obtain updated mask samples. Based on the updated mask sample, the intermediate features are partially convolutional to obtain the first-level encoded features, and the updated mask sample is updated based on the mask update mechanism to obtain the first-level mask sample. The first-level encoded features and the first-level mask samples are downsampled to obtain the output of the first-level network. The output of the first-level network is input into the next-level downsampling network, and this process is repeated level by level until all levels of the downsampling network have completed convolution processing to obtain the encoded features.
[0182] In some implementations, the target region is divided into multiple grids; the observation samples include observation values corresponding to each of the multiple grids, and the mask samples include mask values corresponding to each of the multiple grids. Each mask value includes a first mask value or a second mask value, where the first mask value indicates that the observation value of the grid is not empty, and the second mask value indicates that the observation value of the grid is empty. The model training module 850 is specifically used for: Slide a convolution kernel of a preset size along the grid dimension of the observed sample. For the convolution region currently covered by the convolution kernel, determine whether the sum of the mask values of each grid in the convolution region is greater than 0. If so, perform convolution processing on the observations of the target grid with the mask value of the first mask value within the convolution region to obtain the convolution result; Based on the convolution result, the sum of the mask values of each grid in the convolution region, the convolution kernel, the sum of all-one matrices of the same size as the convolution kernel, and the bias coefficient, the encoding value of the center grid of the convolution region is determined. The intermediate features are determined based on the encoding values corresponding to the center grids of each convolutional region.
[0183] In some implementations, the model training module 850 is specifically used for: For the convolution region currently covered by the convolution kernel, if the sum of the mask values of each grid in the convolution region is greater than 0, the mask value of the center grid of the convolution region is set to 1. The updated mask sample is determined based on the mask value of the center grid of each convolution region.
[0184] In some implementations, the model training module 850 is specifically used for: If the sum of the mask values within the convolution region is not greater than zero, the encoding value of the center grid of the convolution region is set to 0, and the mask value of the center grid of the convolution region is set to 0.
[0185] In some implementations, the decoder includes a multi-level upsampling network, with each upsampling network corresponding one-to-one with a downsampling network; when the model training module 850 uses the decoder to decode the output features of the bottleneck layer to obtain the predicted emission ratio, it is specifically used for: The bottleneck layer output features are upsampled using the lowest level upsampling network to obtain upsampled features; The coding features of the corresponding layer in the encoder are obtained, and the coding features of the corresponding layer are enhanced using the convolutional attention module to obtain enhanced coding features; The upsampling features and the enhanced coding features are fused to obtain fused features, and the fused features are passed sequentially level by level until all levels of the upsampling network have completed the decoding process to obtain the predicted emission ratio.
[0186] In some implementations, the model training module 850 is specifically used to: determine the relative error loss based on the relative error between the predicted emission ratio and the corresponding sample labels; Based on the relative contribution weights of the observation data of each grid in the target region and the relative error, the emission contribution loss is determined; Based on the difference between the predicted emission ratio and the first-order difference of the sample labels in the latitude and longitude directions, the spatial gradient consistency loss is determined. The model loss is determined based on the relative error loss, the emission contribution loss, and the spatial gradient consistency loss.
[0187] In some implementations, the associated data includes observation data and associated mask data; the observation data includes first satellite observation concentration data and first meteorological data of nitrogen dioxide for one day in the target inversion time period, and second satellite observation concentration data and second meteorological data of nitrogen dioxide for the corresponding day in the inversion reference time period; the mask data includes first mask data corresponding to the first satellite observation concentration data and second mask data corresponding to the second satellite observation concentration data.
[0188] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0189] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0190] Corresponding to the above-described method for inverting anthropogenic nitrogen oxide emissions based on satellite observations, this embodiment of the invention also provides a computer device, such as... Figure 9 The diagram shown is a structural schematic of a computer device provided in an embodiment of the present invention. Figure 9 As shown, the computer device 900 includes a processor 910, an internal bus 920, memory 930, a network interface 940, and non-volatile memory 950, and may also include other hardware required for its functions. One or more embodiments of this specification can be implemented in software, for example, the processor 910 reads the corresponding computer program from the non-volatile memory 950 into the memory 930 and then runs it. Of course, besides software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0191] The memory 930, also known as internal memory, is used to temporarily store the computational data in the processor 910, as well as the data exchanged with non-volatile memory 950 such as hard disk. The processor 910 exchanges data with the non-volatile memory 950 through the memory 930.
[0192] In this embodiment of the invention, memory 930 is specifically used to store application code that executes the solution of the present invention, and its execution is controlled by processor 910. That is, when the computer device is running, processor 910 communicates with network interface 940, memory 930 and non-volatile memory 950 through internal bus 920, so that processor 910 executes the application code stored in memory 930 and non-volatile memory 950, and then executes the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations as described in the above method embodiment.
[0193] Processor 910 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware microservices. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0194] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the computer device 900. In other embodiments of the present invention, the computer device 900 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0195] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program performs the steps of the method for retrieving anthropogenic nitrogen oxide emissions based on satellite observations described in the above-described method embodiments. The storage medium can be either volatile or non-volatile computer-readable storage.
[0196] This invention also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0197] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0198] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0199] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0200] Computers suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. Basic computer microservices include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0201] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0202] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0203] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and microservices in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program microservices and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0204] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0205] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for inverting anthropogenic nitrogen oxide emissions based on satellite observations, characterized in that, include: Acquire target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the target area during the historical target year, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical reference year prior to the historical target year, reference meteorological data for each day in the historical reference year, and reference anthropogenic nitrogen oxide emissions for each day in the historical reference year. Using the mass balance emission inversion method, based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions, the daily emission ratio of nitrogen oxides is determined; the emission ratio is used to characterize the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year to the historical reference year. The target satellite observation data is masked to obtain target mask data, and the reference satellite observation data is masked to obtain reference mask data; Based on the target satellite observation data, target meteorological data, reference satellite observation data, and reference meteorological data, an observation sample is constructed, and based on the target mask data and the reference mask data, a mask sample is constructed. Based on the observation sample and the mask sample, an associated training sample is generated, and the corresponding emission ratio is used as the sample label. A sample set is generated based on the daily associated training samples and corresponding sample labels. The initial emission inversion model is trained based on the sample set to obtain the trained emission inversion model. Acquire daily correlation data for the target inversion area and input the correlation data into the trained emission inversion model to obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data. Based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions, determine the target anthropogenic nitrogen oxide emissions for each day within the target inversion period.
2. The method according to claim 1, characterized in that, The method of using mass balance emission inversion to determine the daily emission ratio of nitrogen oxides based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions includes: Using an atmospheric chemical transport model, concentration simulations were performed based on the baseline anthropogenic nitrogen oxide emissions and the target meteorological data to obtain the target tropospheric nitrogen dioxide column concentration. This target tropospheric nitrogen dioxide column concentration characterizes the impact of the meteorological conditions corresponding to the target meteorological data on the tropospheric nitrogen dioxide column concentration, assuming the baseline anthropogenic nitrogen oxide emissions remain constant. Using the atmospheric chemical transport model, concentration simulations were performed based on the baseline anthropogenic nitrogen oxide emissions and the baseline meteorological data to obtain the baseline tropospheric nitrogen dioxide column concentration. Based on the target satellite observation data, the target tropospheric nitrogen dioxide column concentration, the reference satellite observation data, and the reference tropospheric nitrogen dioxide column concentration, the daily emission ratio of nitrogen oxides is determined using the mass balance method.
3. The method according to claim 2, characterized in that, The determination of the daily emission ratio of nitrogen oxides using the mass balance method, based on the target satellite observation data, the target tropospheric nitrogen dioxide column concentration, the reference satellite observation data, and the reference tropospheric nitrogen dioxide column concentration, includes: The ratio of the target satellite observation data to the reference satellite observation data is determined as the observation factor; The ratio of the target tropospheric nitrogen dioxide column concentration to the baseline tropospheric nitrogen dioxide column concentration is determined as the meteorological effect influencing factor; the meteorological effect influencing factor is used to characterize the impact of changes in meteorological conditions on the tropospheric nitrogen dioxide column concentration. The difference between the observed factor and the meteorological effect influencing factor is determined as the net change in column concentration on the same day between the historical target year and the historical baseline year. Based on the net change in column concentration and the preset response coefficient of tropospheric nitrogen dioxide column concentration change to anthropogenic emission change, the daily emission ratio of nitrogen oxides is determined.
4. The method according to claim 1, characterized in that, The target area was divided into multiple grids; the target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year were determined through the following steps: The raw satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year are obtained, and the raw satellite observation data is resampled according to the spatial resolution of the grid to obtain the resampled satellite observation data. The resampled satellite observation data is evaluated based on a preset sliding time window, and the average value of the satellite observation data within the sliding time window is determined as the satellite observation data of the tropospheric nitrogen dioxide column concentration on the central day of the sliding time window. Based on satellite observation data of tropospheric nitrogen dioxide column concentration for each central day, target satellite observation data of daily tropospheric nitrogen dioxide column concentration for the historical target year are generated.
5. The method according to claim 1, characterized in that, The initial emission inversion model is a U-Net neural network, which includes an encoder, a bottleneck layer, and a decoder; the process of training the initial emission inversion model based on the sample set to obtain a trained emission inversion model includes: For each associated training sample, the encoder is used to encode the associated training sample to obtain encoded features; Using the bottleneck layer, the encoded features are subjected to double convolution to obtain the bottleneck layer encoded features, and the residuals of the encoded features and the bottleneck layer encoded features are summed to obtain the bottleneck layer output features. The decoder is used to decode the output features of the bottleneck layer to obtain the predicted emission ratio. Based on the predicted emission ratio and the corresponding sample labels, the model loss is determined, and the model parameters of the initial emission inversion model are adjusted based on the model loss to obtain the emission inversion model.
6. The method according to claim 5, characterized in that, The encoder includes a multi-level downsampling network, each level of which employs a two-layer partial convolution mechanism; the process of using the encoder to encode the associated training samples to obtain encoded features includes: The associated training samples are input into the primary downsampling network. Based on the mask samples, the observed samples are partially convolutional to obtain intermediate features. The mask samples are then updated based on a preset mask update mechanism to obtain updated mask samples. Based on the updated mask sample, the intermediate features are partially convolutional to obtain the first-level encoded features, and the updated mask sample is updated based on the mask update mechanism to obtain the first-level mask sample. The first-level encoded features and the first-level mask samples are downsampled to obtain the output of the first-level network. The output of the first-level network is input into the next-level downsampling network, and this process is repeated level by level until all levels of the downsampling network have completed convolution processing to obtain the encoded features.
7. The method according to claim 6, characterized in that, The target area is divided into multiple grids; the observation sample includes the observation values corresponding to the multiple grids respectively, and the mask sample includes the mask values corresponding to the multiple grids respectively. The mask value includes a first mask value or a second mask value, where the first mask value indicates that the observation value of the grid is not empty, and the second mask value indicates that the observation value of the grid is empty. The step of performing partial convolution processing on the observed samples based on the masked samples to obtain intermediate features includes: Slide a convolution kernel of a preset size along the grid dimension of the observed sample. For the convolution region currently covered by the convolution kernel, determine whether the sum of the mask values of each grid in the convolution region is greater than 0. If so, perform convolution processing on the observations of the target grid with the mask value of the first mask value within the convolution region to obtain the convolution result; Based on the convolution result, the sum of the mask values of each grid in the convolution region, the convolution kernel, the sum of all-one matrices of the same size as the convolution kernel, and the bias coefficient, the encoding value of the center grid of the convolution region is determined. The intermediate features are determined based on the encoding values corresponding to the center grids of each convolutional region.
8. The method according to claim 7, characterized in that, The method of updating the mask sample based on a preset mask update mechanism to obtain an updated mask sample includes: For the convolution region currently covered by the convolution kernel, if the sum of the mask values of each grid in the convolution region is greater than 0, the mask value of the center grid of the convolution region is set to 1. The updated mask sample is determined based on the mask value of the center grid of each convolution region.
9. The method according to claim 8, characterized in that, The method further includes: If the sum of the mask values within the convolution region is not greater than zero, the encoding value of the center grid of the convolution region is set to 0, and the mask value of the center grid of the convolution region is set to 0.
10. The method according to claim 5, characterized in that, The decoder includes a multi-level upsampling network, with each upsampling network corresponding one-to-one with a downsampling network; the step of using the decoder to decode the output features of the bottleneck layer to obtain the predicted emission ratio includes: The bottleneck layer output features are upsampled using the lowest level upsampling network to obtain upsampled features; The coding features of the corresponding layer in the encoder are obtained, and the coding features of the corresponding layer are enhanced using the convolutional attention module to obtain enhanced coding features; The upsampling features and the enhanced coding features are fused to obtain fused features, and the fused features are passed sequentially level by level until all levels of the upsampling network have completed the decoding process to obtain the predicted emission ratio.
11. The method according to claim 5, characterized in that, The determination of model loss based on the predicted emission ratio and the corresponding sample labels includes: The relative error loss is determined based on the relative error between the predicted emission ratio and the corresponding sample labels; Based on the relative contribution weights of the observation data of each grid in the target region and the relative error, the emission contribution loss is determined; Based on the difference between the predicted emission ratio and the first-order difference of the sample labels in the latitude and longitude directions, the spatial gradient consistency loss is determined. The model loss is determined based on the relative error loss, the emission contribution loss, and the spatial gradient consistency loss.
12. The method according to claim 1, characterized in that, The associated data includes observation data and associated mask data; the observation data includes first satellite observation concentration data and first meteorological data of nitrogen dioxide for one day in the target inversion time period, and second satellite observation concentration data and second meteorological data of nitrogen dioxide for the corresponding day in the inversion reference time period; the mask data includes first mask data corresponding to the first satellite observation concentration data and second mask data corresponding to the second satellite observation concentration data.
13. A device for retrieving anthropogenic nitrogen oxide emissions based on satellite observations, characterized in that, include: The data acquisition module is used to acquire target satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical target year for the target area, daily target meteorological data, reference satellite observation data of daily tropospheric nitrogen dioxide column concentration in the historical reference year prior to the historical target year, reference meteorological data of daily in the historical reference year, and reference anthropogenic nitrogen oxide emissions of daily in the historical reference year. The ratio determination module is used to determine the daily emission ratio of nitrogen oxides based on the target satellite observation data, the target meteorological data, the reference satellite observation data, the reference meteorological data, and the reference anthropogenic nitrogen oxide emissions using the mass balance emission inversion method; the emission ratio is used to characterize the ratio of anthropogenic nitrogen oxide emissions on the same day in the historical target year to the historical reference year. The masking module is used to perform masking processing on the target satellite observation data to obtain target masking data, and to perform masking processing on the reference satellite observation data to obtain reference masking data; The sample generation module is used to construct observation samples based on the target satellite observation data, target meteorological data, reference satellite observation data, and reference meteorological data, and to construct mask samples based on the target mask data and the reference mask data, and to generate associated training samples based on the observation samples and the mask samples, and to use the corresponding emission ratio as the sample label. The model training module is used to generate a sample set based on the daily associated training samples and corresponding sample labels, and to train the initial emission inversion model based on the sample set to obtain the trained emission inversion model. The emission inversion module is used to acquire daily correlation data for the target inversion area, input the correlation data into the trained emission inversion model, obtain the target emission ratio of nitrogen oxides for each day within the target inversion period indicated by the correlation data, and determine the target anthropogenic nitrogen oxide emissions for each day within the target inversion period based on the target emission ratio and the baseline anthropogenic nitrogen oxide emissions.
14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations as described in any one of claims 1-12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method for inverting anthropogenic nitrogen oxide emissions based on satellite observations as described in any one of claims 1-12.