Atmospheric multi-pollutant forecast correction method and system integrating multi-source observation and mode output

By constructing a background reference field-reliability adaptive fusion link and a multi-pollutant coupling constraint model, the problem of fusing multi-source pollutant monitoring and model forecast information was solved, achieving stability, robustness and consistency of multi-pollutant forecast results, and improving forecast accuracy and interpretability.

CN122087280APending Publication Date: 2026-05-26黑龙江省生态环境监测中心 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
黑龙江省生态环境监测中心
Filing Date
2026-02-05
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to reliably integrate multi-source pollutant monitoring and model forecasting information under uncertainty differences and consistency constraints, resulting in insufficient accuracy and stability of multi-pollutant forecasts.

Method used

By constructing a background reference field-reliability adaptive fusion link, combining the historical error statistics and real-time deviation information of each data source, dynamic reliability weights are generated, and an interactive correction model with multi-pollutant coupling constraints is constructed to correct the physicochemical consistency and statistical correlation of multi-pollutant forecast results.

Benefits of technology

It has achieved stability, robustness and consistency of multi-pollutant forecast results, improved forecast accuracy and interpretability, and ensured the spatiotemporal coherence and reliability of correction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087280A_ABST
    Figure CN122087280A_ABST
Patent Text Reader

Abstract

The invention discloses an atmospheric multi-pollutant forecast correction method and system fusing multi-source observation and mode output, and relates to the field of atmospheric pollution intelligent monitoring, and the method comprises the steps: mapping ground / satellite and other multi-source observation and mode output to a unified space-time reference, and forming comparable lattice point data; performing data assimilation on the mode initial field under observation constraint to obtain an assimilation analysis field with both spatial continuity and observation consistency as a background reference; combining historical error statistics of each data source and real-time deviation information of a relative background field, adaptively generating a dynamic reliability weight according to grid points and pollutants, realizing weighted fusion of multi-source observation and the background field, and obtaining a multi-pollutant preliminary fusion field; and introducing a multi-pollutant coupling constraint interaction correction model, cooperatively correcting multi-pollutant forecast and keeping physical and chemical consistency and statistical correlation. Therefore, service-oriented multi-source consistent fusion and multi-pollutant collaborative correction are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent air pollution monitoring technology, and in particular to a method and system for forecasting and correcting multiple air pollutants by integrating multi-source observations and model outputs. Background Technology

[0002] With the deepening of urbanization and industrialization, air pollution monitoring and accurate forecasting have become key aspects of environmental protection. In air pollution monitoring and forecasting operations, to achieve early warning, information sources such as numerical model forecasts, ground station monitoring, and satellite remote sensing are usually relied upon simultaneously, with low-cost sensors introduced as a supplement in some areas.

[0003] However, various data sources have inherent differences in spatial coverage, temporal resolution, observation error patterns, and operational availability: model outputs cover the entire region but are affected by prior uncertainties; ground stations are more accurate but have sparse coverage; satellite remote sensing has advantages in terms of wide range but factors such as resolution and cloud cover can limit stable availability, which in turn leads to inconsistencies in multi-source information and affects forecast accuracy and interpretability.

[0004] Currently, chemical transport models (such as those driven by meteorological inputs and emission inventories) are susceptible to hourly pollutant concentration predictions due to factors such as emission inventory lags, insufficient characterization of meteorological processes, and parameterization of chemical reactions, resulting in a coexistence of systematic biases and random errors. In engineering, statistical or machine learning post-processing is often used to reduce bias; however, these methods are typically geared towards single pollutants or single data sources and are largely based on empirical relationships, making it difficult to maintain necessary physical consistency constraints when merging multiple sources. Summary of the Invention

[0005] This application provides a method, system, storage medium, computer program product, and electronic device for atmospheric multi-pollutant forecast correction that integrates multi-source observations and model outputs, in order to at least solve the problem that in current related technologies, it is difficult to reliably integrate multi-source pollution monitoring and model forecast information under uncertainty differences and consistency constraints, resulting in insufficient accuracy and stability of multi-pollutant forecasts.

[0006] In a first aspect, embodiments of this application provide a method for atmospheric multi-pollutant forecast correction that integrates multi-source observations and model outputs. The method includes: acquiring multi-source atmospheric pollutant data and numerical model output data for a target area at a time to be processed and its corresponding preset forecast period; the multi-source atmospheric pollutant data being used to characterize observational information of at least two types of atmospheric pollutants, and including at least ground monitoring data and / or satellite remote sensing data; the numerical model output data being used to characterize model forecast information of the at least two types of atmospheric pollutants; converting the multi-source atmospheric pollutant data and the numerical model output data to a unified preset spatiotemporal reference to obtain a multi-source gridded data set corresponding to the preset spatiotemporal reference; and performing data assimilation processing on the numerical model output data based on the multi-source gridded data set to update the initial state of the model under observational constraints and generate a pollutant assimilation analysis field reflecting the spatial distribution characteristics of pollutants, and determining the pollutant assimilation analysis field as a method for characterizing the at least two types of atmospheric pollutants. A background reference field is established to represent the spatiotemporal distribution of at least two types of air pollutants. Based on the multi-source gridded data set, for each grid point and each pollutant in the preset spatiotemporal reference, the reliability index of each data source at the current moment is evaluated by combining the historical error statistical characteristics of each data source and the real-time deviation information of each data source relative to the background reference field. Dynamic reliability weights for each data source are generated based on these reliability indexes. Based on these dynamic reliability weights, a weighted fusion process is performed between the background reference field and the observation data in the multi-source gridded data set to generate a preliminary multi-pollutant fusion field containing at least two types of air pollutants. An interactive correction model containing multi-pollutant coupling constraints is constructed, and pollutant interactive correction processing is performed on the preliminary multi-pollutant fusion field using the interactive correction model to generate multi-pollutant corrected forecast results. The multi-pollutant coupling constraints are used to characterize the interaction relationships between pollutants and constrain the physicochemical consistency and statistical correlation between forecast results of different pollutants.

[0007] Secondly, embodiments of this application provide an atmospheric multi-pollutant forecast correction system integrating multi-source observations and model outputs. The system includes: a data acquisition unit, used to acquire multi-source atmospheric pollutant data and numerical model output data for a target area at the time to be processed and its corresponding preset forecast period. The multi-source atmospheric pollutant data is used to characterize observation information of at least two types of atmospheric pollutants, and includes at least ground monitoring data and / or satellite remote sensing data; the numerical model output data is used to characterize model forecast information of the at least two types of atmospheric pollutants; a spatiotemporal reference conversion unit, used to convert the multi-source atmospheric pollutant data and the numerical model output data to a unified preset spatiotemporal reference to obtain a multi-source gridded data set corresponding to the preset spatiotemporal reference; and an assimilation reference field generation unit, used to perform data assimilation processing on the numerical model output data based on the multi-source gridded data set, to update the initial state of the model under observation constraints and generate a pollutant assimilation analysis field reflecting the spatial distribution characteristics of pollutants, and to determine the pollutant assimilation analysis field as a model for characterizing the at least two types of atmospheric pollutants. The system comprises: a background reference field for the spatiotemporal distribution of two types of air pollutants; a dynamic reliability weighting unit, used to evaluate the reliability index of each data source at the current moment based on the multi-source gridded data set, for each grid point and each pollutant in the preset spatiotemporal reference, combining the historical error statistical characteristics of each data source and the real-time deviation information of each data source relative to the background reference field, and generating dynamic reliability weights for each data source based on the reliability indexes; a multi-source weighted fusion unit, used to perform weighted fusion processing on the background reference field and the observation data in the multi-source gridded data set based on the dynamic reliability weights, to generate a preliminary multi-pollutant fusion field containing at least two types of air pollutants; and an interactive correction solution unit, used to construct an interactive correction model containing multi-pollutant coupling constraints, and use the interactive correction model to perform pollutant interactive correction processing on the preliminary multi-pollutant fusion field to generate multi-pollutant corrected forecast results; the multi-pollutant coupling constraints are used to characterize the interaction relationship between pollutants and constrain the physicochemical consistency and statistical correlation between different pollutant forecast results.

[0008] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the atmospheric multi-pollutant forecast correction method fusion of multi-source observations and model outputs according to any embodiment of the present application.

[0009] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of the atmospheric multi-pollutant forecast correction method that integrates multi-source observations and model outputs according to any embodiment of this application.

[0010] Fifthly, embodiments of this application provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the atmospheric multi-pollutant forecast correction method that fuses multi-source observations and model outputs according to any embodiment of this application.

[0011] The atmospheric multi-pollutant forecasting correction method and system that integrates multi-source observations and model outputs provided in this application can achieve at least the following technical effects: (1) By constructing an integrated correction link of “background reference field-reliability adaptive fusion”, multi-source observations and model outputs form a sustainable iterative constraint system under a unified spatiotemporal reference. Specifically, on the one hand, the analysis field updated by the observation constraints is used as the background reference, so that the correction results maintain a stable field structure in terms of spatial continuity and temporal evolution; on the other hand, historical error statistics of each data source and real-time deviation information relative to the background field are introduced to form a reliability evaluation oriented towards grid points, time and pollutants, and dynamic reliability weights are generated accordingly. Thus, the fusion process can adaptively respond to fluctuations in data source availability, local anomalies and changes in error morphology, avoid the disproportionate pull of a single data source on the results under specific conditions, and achieve a reasonable balance between continuity, stability and reliability in the initial fusion field of multiple pollutants.

[0012] (2) Furthermore, by introducing an interactive correction mechanism with multi-pollutant coupling constraints, the corrections of different pollutants are no longer independent parallel corrections, but rather coordinated adjustments are made under the joint constraints of the interaction relationships and statistical correlation structures among pollutants. This mechanism can maintain the consistency and self-consistency of the correction results at the multi-pollutant level, reduce the risk of cross-pollutant corrections canceling each other out or the structure being destroyed, and make the final multi-pollutant correction forecast results not only more robust in reducing single-pollutant errors, but also more in line with the expected co-evolution characteristics at the joint distribution and correlation structure levels, thereby improving the interpretability and operational usability of the results.

[0013] This technical solution uses an "interpretable background reference field" as an anchor point and achieves adaptive fusion of multi-source information through dynamic reliability weights. It then applies interactive correction of "multi-pollutant coupling constraints" to the fusion result, so that the correction output simultaneously meets the comprehensive goals of stability (spatiotemporal coherence), robustness (insensitivity to data source fluctuations and anomalies), and consistency (cross-pollutant collaborative self-consistency). This results in a highly reliable fusion forecast result with a stable spatiotemporal structure and cross-pollutant consistency constraints for multi-pollutant forecast correction tasks. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating an example of an atmospheric multi-pollutant forecast correction method that integrates multi-source observations and model outputs according to an embodiment of this application is shown. Figure 2 A flowchart illustrating an example of generating a pollutant assimilation analysis field in a method according to an embodiment of this application is shown. Figure 3 This paper illustrates an example operational mechanism diagram of an atmospheric multi-pollutant forecast correction method that integrates multi-source observations and model outputs according to an embodiment of this application. Figure 4 A comparison chart of the root mean square error distribution of four comparative schemes for different pollutant forecast accuracies is shown. Figure 5 A two-dimensional sensitivity heatmap of the influence of key algorithm parameters on prediction error in the method according to an embodiment of this application is shown; Figure 6 A structural block diagram of an example atmospheric multi-pollutant forecast correction system that integrates multi-source observations and model outputs according to an embodiment of this application is shown. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] It should be noted that, at the level of specific data processing and correction algorithms, current technologies mainly improve forecast quality through data assimilation, linear deviation correction, and data fusion, but they still have limitations in handling the dynamic characteristics and multivariate coupling of complex atmospheric environments.

[0018] Firstly, regarding the optimization of numerical models, current techniques often utilize assimilation frameworks such as three-dimensional variational or ensemble Kalman filtering to incorporate aerosol optical depth (AOD) data from satellite remote sensing or ground monitoring data into the model. This improves forecast performance by updating the initial field or retrieving emission sources. Although related studies indicate that joint assimilation of multi-source observation data can effectively reduce root mean square error, current assimilation systems often focus on updating the state of a single pollutant (e.g., only aerosols or nitrogen dioxide), and the output results still highly depend on the structure of the chemical transport model, lacking an independent correction mechanism for nonlinear chemical reaction processes between multiple pollutants.

[0019] Secondly, regarding the post-processing of model outputs, some forecasting systems employ linear bias correction methods based on Kalman filtering, updating the bias term of the time series by establishing a state-space model. While such methods can reduce forecast errors to some extent, they typically assume that the statistical characteristics of the error are stationary over time, making it difficult to adapt to the non-stationary error characteristics caused by rapid changes in the atmospheric environment and sensor performance drift. Furthermore, these methods often model single pollutants independently, failing to address the complex coupling issues between different pollutants.

[0020] Furthermore, in the field of data fusion, some scholars have proposed using geostatistical methods to combine model, satellite, and ground data to generate high-resolution concentration distribution products and attempt to quantify uncertainty. However, current data fusion methods are mostly applied to static analysis of historical data or estimation for specific scenarios, lacking optimization for real-time operational forecasting scenarios, and often failing to fully consider the dynamic changes in the real-time reliability of data sources when allocating weights.

[0021] Furthermore, significant physicochemical coupling relationships exist among air pollutants. For example, the correlation between nitrogen oxides, ozone, and fine particulate matter fluctuates with seasonal, meteorological, and time-scale variations. Related studies indicate that ignoring these dynamic environmental interactions during the correction process and simply pursuing the statistical optimum of a single indicator can easily disrupt the consistency of the physicochemical logic of the forecast results, and may even introduce new systematic errors while correcting for biases in a particular pollutant.

[0022] It should be understood that the above description of the relevant technologies is intended only to help the public better understand the inventive spirit and motivation of this application, and is not intended to limit this application. Furthermore, the technical solutions described in the above-mentioned relevant technologies are not prior art, and may also be undisclosed technical solutions, such as those under research or in the laboratory stage.

[0023] The technical solutions in this application, including the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, comply with relevant laws and regulations and do not violate public order and good morals.

[0024] Figure 1 A flowchart illustrating an example of an atmospheric multi-pollutant forecast correction method that integrates multi-source observations and model outputs according to an embodiment of this application is shown.

[0025] Regarding the execution entity of the method in this application embodiment, it can be any controller or processor with computing or processing capabilities, such as a forecast correction platform controller deployed in an air quality forecasting operational system. The forecast correction platform controller can be implemented by a server, cloud computing node, or edge computing node, and establishes a connection with a multi-source data access interface through a communication network to obtain multi-source atmospheric pollutant data and numerical model output data of the target area at the time to be processed and within the preset forecast period from the ground monitoring station data interface, satellite remote sensing data service interface, and numerical model output interface. In addition, the forecast correction platform controller can also output the generated multi-pollutant correction forecast results to the forecast release terminal, the early warning operational terminal, and / or the visualization display terminal to support air quality early warning and coordinated response.

[0026] In some examples, it may be integrated into an electronic device or terminal through software, hardware, or a combination of both, and the type of terminal or electronic device may be diverse.

[0027] like Figure 1 As shown, in step S110, multi-source atmospheric pollutant data and numerical model output data of the target area during the time to be processed and the corresponding preset forecast period are obtained.

[0028] Here, multi-source atmospheric pollutant data are used to characterize observational information of at least two types of atmospheric pollutants, and include at least ground monitoring data and / or satellite remote sensing data; numerical model output data are used to characterize model forecast information of at least two types of atmospheric pollutants.

[0029] More specifically, the target area can be determined by administrative region, urban agglomeration, watershed, or operational grid domain; the time to be processed can be the current analysis time (e.g., the hour) or the operational trigger time; the preset forecast period is used to cover the forecast needs for several hours to several days (e.g., short-term 0~24h or 0~72h). Based on this, multi-source observation data for at least two types of pollutants (e.g., any two or more of PM2.5, PM10, O3, NO2, SO2, and CO) are acquired. The data source includes at least ground station monitoring data and / or satellite remote sensing data; optionally, low-cost sensor, mobile monitoring, or regional station data can also be introduced as supplements. To ensure subsequent quantifiable evaluation and fusion, each observation record can also simultaneously carry: observation timestamp, spatial location (latitude / longitude / pixel range / station altitude), pollutant type and unit, quality identifier (e.g., station validity marker, satellite cloud detection / quality level), and observation uncertainty or confidence index (e.g., inversion error, calibration coefficient, equipment status, etc.).

[0030] Simultaneously, numerical model output data is acquired to characterize model forecast information for the same pollutant set within a preset forecast period. The model can be a chemical transport model or the post-processing output of an operational numerical model. Its data may include: hourly (or higher frequency) concentration forecasts for each pollutant on the model grid, necessary meteorological field driving variables (wind field, boundary layer height, temperature, humidity, etc.), and forecast start time information. Furthermore, preliminary consistency data preprocessing can be performed, such as unifying pollutant units (recording the conditions required for volume fraction / mass concentration conversion), initial screening of outliers (e.g., exceeding physically reasonable ranges, significant jumps, consecutive missing measurements), and data timeliness screening (removing data with excessive lag or unusable delays), etc. This elevates the subsequent correction target from a single pollutant / single source to a unified problem definition of "multiple pollutants - multiple sources - covering the forecast period."

[0031] In step S120, the multi-source atmospheric pollutant data and numerical model output data are converted to a unified preset spatiotemporal reference to obtain a multi-source grid data set corresponding to the preset spatiotemporal reference.

[0032] Here, by constructing a unified, pre-defined spatiotemporal reference, ground points, satellite pixels, and model grids can be aligned under the same spatiotemporal coordinates. The spatial reference can adopt a target operational grid (e.g., a fixed latitude and longitude resolution or a projected coordinate grid), clearly defining the grid range, resolution, grid indexing rules, and representative height (e.g., near-surface layer); the temporal reference can adopt an integer sequence or a fixed time step sequence, clearly defining alignment rules (nearest neighbor alignment, time window averaging, or conversion based on observation integration periods). Subsequently, separate "mapping / observation operators" are constructed for different data sources: ground station data can be mapped from point to grid point with representative error control (e.g., grid placement, neighborhood weighting, or weighting based on the station's influence radius); satellite remote sensing data can be aggregated by area weighting based on the area overlap relationship between pixels and grid points, and usable pixels are filtered in conjunction with quality labels; model data undergoes grid resampling (bilinear interpolation, area conservation interpolation, or nearest neighbor) to match the unified spatial grid.

[0033] The resulting multi-source gridded data set can be organized into a unified data cube structure, for example, using (pollutant p, grid point g, time t, source s) as an index to store the gridded observation / model values ​​and their quality attributes. When a source is missing data at a certain grid point, it is explicitly marked as missing and the reason for the missing data (cloud cover, site outage, delay, etc.) is retained so that a consistent missing data handling strategy can be adopted in subsequent weight calculations. Thus, the "structural incomparability" defect of multi-source data is eliminated by unifying the spatiotemporal benchmark.

[0034] In step S130, based on the multi-source grid data set, data assimilation processing is performed on the numerical model output data to update the initial state of the model under observation constraints and generate a pollutant assimilation analysis field that reflects the spatial distribution characteristics of pollutants. The pollutant assimilation analysis field is then determined as a background reference field for characterizing the spatiotemporal distribution of at least two types of atmospheric pollutants.

[0035] Here, model forecasts are used as the prior background. The initial state of the model is updated through observation constraints to generate an assimilation analysis field that is closer to the real spatial distribution. This assimilation analysis field is then used as the background reference field. For example, variational assimilation (such as 3D-Var / 4D-Var) or ensemble filtering assimilation (such as EnKF / LETKF) can be used to achieve this.

[0036] In some implementations, the assimilation state vector comprises the concentration fields of at least two types of pollutants on a unified grid (which can be expanded to relevant meteorological variables or chemical state parameters if necessary), and the observation vector is derived from gridded ground / satellite observations. Simultaneously, an observation error description (synthesized from observation uncertainty, representativeness error, and mapping error) and a background error description (which can be obtained from historical model error statistics, ensemble perturbations, or empirical parameterization) are constructed. The assimilation process may also include consistent quality control: for example, threshold testing based on the innovation quantity (observation-background difference), neighborhood testing based on spatial consistency, and strategies to reduce or eliminate the impact of high-uncertainty observations, thereby preventing anomalous observations from skewing the analytical field.

[0037] It should be noted that the pollutant assimilation analysis field output by assimilation can have a clear spatiotemporal meaning: for example, as the analysis field at the time to be processed (t0 analysis), and can be extended to generate analysis sequences covering short time windows; since this analysis field inherits the spatial continuity and dynamic consistency of the model, and is also subject to observational constraint correction, it can be defined as a background reference field for subsequent real-time bias assessment and fusion anchoring.

[0038] In this embodiment, a reference benchmark that combines continuity and observation constraints is formed at the front end of the forecast chain, so that the bias assessment of all subsequent data sources has a unified reference. This avoids both the solidification of bias caused by relying entirely on the model as the benchmark and the spatial fragmentation caused by relying on discrete observations as the benchmark. At the same time, by assimilating and updating the initial state, the risk of cumulative propagation of the initial forecast bias to subsequent forecast periods can be reduced.

[0039] In step S140, based on the multi-source grid data set, for each grid point and each pollutant in the preset spatiotemporal reference, combined with the historical error statistics of each data source and the real-time deviation information of each data source relative to the background reference field, the reliability index of each data source at the current moment is evaluated, and the dynamic reliability weight of each data source is generated according to the reliability index.

[0040] Here, instead of setting fixed weights based solely on experience, a dynamic reliability assessment is performed on each data source at each grid point, for each pollutant, and at each time point. Specifically, historical error statistics are established for each data source. Statistical measures can be maintained using a rolling window or exponential forgetting mechanism, such as seasonal / weather-specific / regional deviations (systematic errors) and dispersion (random errors). These can be combined with availability, latency, and quality label stability to form a long-term stability profile. At the real-time level, the deviation information of the data source grid point observations relative to the background reference field is calculated. The deviation can simultaneously consider the absolute difference, relative difference, and the deviation strength after uncertainty normalization (to avoid incomparability caused by differences in the magnitude of different pollutants).

[0041] In some implementations, historical error terms and real-time deviation terms can be combined into a reliability index according to a preset fusion rule, and quality identification gating (such as satellite cloud occlusion / low-quality pixels being directly downweighted or set to zero) can be introduced. At the same time, smoothing or limiting constraints are applied to the weights in time and space to avoid drastic weight jumps that cause jitter in the fused field.

[0042] It should be noted that the generation of dynamic reliability weights can satisfy interpretability and computability: for example, first obtain the reliability scores of each source, and then normalize them into weight allocations on the same grid point / contaminant / time dimension; at the same time, the weight term corresponding to the "background reference field" should be explicitly included, so that the fusion will not lose continuity when observations are sparse, nor will it ignore the observation advantages when observations are reliable. In the case of missing measurements, the weights should be automatically redistributed (the weight of the missing source is zero and absorbed by the other sources and background terms), thereby ensuring the closure of the entire process.

[0043] In this embodiment, the fusion process is made capable of adaptively responding to fluctuations in data source quality, local anomalies, cloud obstruction, and missing site measurements, avoiding excessive influence of a single source on the results under unreliable conditions; at the same time, the combination of "historical + real-time" evaluation improves the stability and portability of weight allocation, enabling the system to maintain controllable robustness in long-term operation.

[0044] In step S150, based on dynamic reliability weights, the background reference field and the observation data in the multi-source grid data set are subjected to weighted fusion processing to generate a multi-pollutant preliminary fusion field containing at least two types of atmospheric pollutants.

[0045] In some implementations, the fusion process is performed grid-by-grid and pollutant-by-pollutant on a unified grid: when observations are available, the observations are weighted according to the weights of each source, while retaining the background reference field as a continuity anchor; when observations are unavailable or have low reliability, the fusion result automatically reverts to the background reference field, ensuring that the output field does not exhibit large-area voids or fragmentation in space. Simultaneously, the preliminary fused field can output accompanying quality characterization information, such as fusion uncertainty estimates for each grid point, identifiers of major contribution sources (for traceability), and weight distribution summaries (for auditing and maintenance), enabling interpretable, traceable, and diagnosable results in operational scenarios. Thus, without sacrificing spatial continuity, reliable observation constraints are introduced, enabling the output field to simultaneously possess coverage (guaranteed by the background field), enhanced realism (pulled back by reliable observations), and stability (controlled by dynamic weights).

[0046] In step S160, an interactive correction model containing multi-pollutant coupling constraints is constructed, and the interactive correction model is used to perform pollutant interactive correction processing on the preliminary fusion field of multi-pollutants to generate multi-pollutant correction prediction results.

[0047] Here, multi-pollutant coupling constraints are used to characterize the interaction relationships between pollutants and to constrain the physicochemical consistency and statistical correlation between forecast results of different pollutants.

[0048] In some implementations, the inputs to the interactive calibration model may include: a preliminary fusion field of multiple pollutants, the hourly forecast trajectory of the numerical model within the forecast period, and auxiliary factors related to pollution evolution (such as boundary layer height, wind field, temperature and humidity, radiation, key emission period indications, etc., as optional enhancement features).

[0049] It should be noted that interactive correction does not involve independently fitting a correction value for each pollutant. Instead, it establishes a multi-task joint correction structure: simultaneously outputting correction increments or correction coefficients for multiple pollutants at the same time point, and linking these outputs through "coupling constraints." The implementation of coupling constraints can employ engineering-feasible constraint design, such as: constraining the statistical correlation structure between correction results for different pollutants (preserving the shape of covariance / correlation coefficients), constraining the directional consistency of typical interactions (e.g., the synergistic change trends of precursors and secondary products under specific conditions), and constraining the temporal continuity within the forecast period (avoiding irregular jumps in hourly correction values).

[0050] In some implementations, this can be accomplished through frameworks such as constrained optimization, multi-output regression / machine learning models with regularization terms, or multivariate correction incorporating Kalman update principles. The key is that constraints are integrated into the model training and solution process. After completing the interactive correction, multi-pollutant correction prediction results are output, along with a coupled consistency evaluation index or constraint satisfaction summary, to demonstrate the consistency and interpretability of the correction results across multiple pollutants.

[0051] Through the embodiments of this application, the pollutant forecast correction is improved from "local optimum of a single pollutant" to "overall correction of multiple pollutants in a unified manner". This can improve the accuracy of forecasts for each pollutant while avoiding cross-pollutant correction conflicts, destruction of related structures, or illogical combination results, thereby enhancing the credibility and interpretability of multi-pollutant joint early warning operations.

[0052] Regarding the implementation details of step S120, in some examples of embodiments of this application, quality control is performed on ground monitoring data and satellite remote sensing data in multi-source atmospheric pollutant data, outliers are removed based on sliding time window statistical thresholds, and abrupt change records are removed in combination with spatial consistency tests of adjacent stations, so as to obtain a cleaned observation dataset.

[0053] First, rigorous quality control is implemented on multi-source atmospheric pollutant data to eliminate noise caused by sensor malfunctions or environmental interference. In this embodiment, a cleaning mechanism based on spatiotemporal constraints is constructed for both ground monitoring data and satellite remote sensing data: in the temporal dimension, a sliding time window (e.g., the past 24 hours) is set, and the statistical characteristics of the data within the window (e.g., mean and standard deviation) are calculated to remove outliers exceeding a preset statistical threshold (e.g., 3 times the standard deviation); in the spatial dimension, a spatial consistency test is performed between adjacent stations, and isolated abrupt changes are identified and removed by comparing the deviation of the target station's observation mean with that of other stations in its neighborhood. This effectively filters out outliers caused by equipment malfunctions, data transmission errors, or local transient interference, providing a high-confidence observation dataset for subsequent processing.

[0054] Then, a model-assisted vertical transformation is performed on the satellite remote sensing data. The vertical profile factor is calculated using the vertical distribution information of aerosols in the numerical model output data. Combined with the planetary boundary layer height parameter, the vertical column concentration or optical thickness of the satellite remote sensing data is inverted and mapped to near-surface concentration data.

[0055] In some implementations, to address the issue that satellite remote sensing data typically only provides vertical column concentration or optical thickness (AOD) and cannot directly characterize near-surface concentration, a model-assisted vertical transformation is performed. Specifically, the vertical profile information of each grid point is extracted using the three-dimensional aerosol distribution field output by the numerical model, and the vertical profile factor is calculated by combining it with the planetary boundary layer height (PBLH). This factor characterizes near-surface concentration The ratio to the total column concentration can be expressed, for example, as: The denominator represents the total column concentration simulated in the model. The vertically distributed concentration is simulated by the numerical model. This is the upper limit of the integration height; this factor is used to determine the vertical column concentration observed by satellite. Inversion mapping to near-surface concentration data By introducing the physical hierarchical structure information of the model, the problem of dimensional differences in satellite data, which "sees the surface but not the points," is solved, significantly improving the usability of satellite data in near-surface air quality assessment.

[0056] Subsequently, adaptive spatial reconstruction was performed on the cleaned ground monitoring data and near-ground concentration data. Based on the grid distribution of the preset spatiotemporal reference, a spatial interpolation algorithm that considers distance weight and terrain elevation differences was used to map the non-uniformly distributed observation points to the grid center to form gridded observation data.

[0057] Here, adaptive spatial reconstruction is performed on the cleaned ground monitoring data and the inverted satellite near-ground data to address the problem of uneven spatial distribution of observations. In this embodiment, instead of simple geometric distance interpolation, a spatial interpolation algorithm that simultaneously considers distance weights and terrain elevation differences (such as 3D inverse distance weighted IDW or co-kriging) is used. When calculating the contribution weight of non-uniformly distributed observation points to the target grid center, the algorithm not only inversely proportional to the horizontal distance but also introduces an elevation difference penalty term, significantly reducing the weight of observation points with large elevation differences (such as mountaintop stations to valley floor grids). This maps discrete observations to a unified grid center, forming gridded observation data, effectively solving the spatial interpolation distortion caused by neglecting vertical diffusion differences in complex terrain.

[0058] Furthermore, spatiotemporal alignment and unification are performed on the gridded observation data and the numerical model output data. Each data is resampled according to the temporal resolution of the time to be processed. In the case that the original grid of the numerical model output data is inconsistent with the preset spatiotemporal reference, spatial resampling is performed on the numerical model output data, thereby generating a multi-source gridded data set with consistent spatiotemporal dimensions.

[0059] Here, the generated gridded observation data and numerical model output data undergo final spatiotemporal alignment and unification. In the time dimension, based on the time to be processed (e.g., the forecast start time), each data stream is resampled or aggregated according to a preset time resolution (e.g., 1 hour). In the spatial dimension, if the numerical model output (usually a Lambert projection or other irregular grid) is inconsistent with the preset spatiotemporal reference (e.g., a regular latitude and longitude grid), spatial resampling of the model data is performed using bilinear interpolation or mass conservation remapping algorithms. This constructs a multi-source gridded data set that is fully aligned in both time series and spatial grid, and has unified physical dimensions, eliminating structural differences between the multi-source data.

[0060] Figure 2 A flowchart illustrating an example of generating a pollutant assimilation analysis field in a method according to an embodiment of this application is shown.

[0061] like Figure 2 As shown, in step S210, for the ground monitoring data in the multi-source grid data set, super-observation construction processing is performed. Within the preset time assimilation window, based on the spatial resolution of the preset spatiotemporal reference, multiple high-frequency observation samples falling into the same grid are spatiotemporally aggregated to generate representative observation samples that match the grid scale, thereby reducing the representativeness error and spatial correlation of the observation data.

[0062] Here, over-observation construction processing is performed on ground monitoring data to address the spatial scale mismatch between ground station "point observations" and numerical model "area (grid)" forecasts. In this embodiment, a preset spatiotemporal grid with the same resolution as the numerical model is defined, along with a preset temporal assimilation window centered on the assimilation time (e.g., ...). (minutes). For those falling into the same grid within A number of high-frequency ground observation samples are used, and their average value is calculated as the representative observation sample for that grid. The representativeness error is estimated based on the dispersion between samples. This process not only smooths out subgrid noise that only represents the local microenvironment, but also significantly reduces the spatial correlation of observation errors between adjacent stations, avoiding spurious local overfitting in the assimilation results caused by high-density stations.

[0063] In step S220, an assimilation observation system is constructed, which assembles representative observation samples and satellite remote sensing data from multi-source grid data sets into observation vectors, and constructs an observation operator for mapping the numerical model state space to the observation space, and configures corresponding observation error covariance parameters for different observation types.

[0064] Here, an assimilation observation system is constructed to prepare the mapping from the physical observation space to the model state space. Specifically, the representative observation samples (usually near-ground concentrations) and satellite remote sensing data (usually column concentrations or AOD) are vector-assembled to form a unified observation vector. Simultaneously, an observation operator is constructed. This operator is used to map state variables (such as three-dimensional concentration fields) from numerical models to the observation space; for ground data, It is a bilinear interpolation operator; for satellite data, This is for complex operators that include vertical integrals or radiative transfer models. Furthermore, the observation error covariance matrix is ​​configured for different observation types. The diagonal elements represent the sum of instrument error and representativeness error. This unifies multi-source data with vastly different physical properties onto a reference surface that can be mathematically compared with model states.

[0065] In step S230, a three-dimensional variational assimilation model is constructed, the numerical model output data is used as the model background field, the weight relationship between the model background field and the observation vector is defined using the background error covariance matrix and the observation error covariance matrix, an objective functional containing background terms and observation terms is constructed, and the objective functional is iteratively minimized to obtain the analysis increment.

[0066] Here, when the observation operator is nonlinear, linearization is performed on the observation operator during the iteration process.

[0067] More specifically, the core of solving for the optimal analysis increment using a three-dimensional variational (3D-Var) assimilation model is to find a state that is close to both the model background field and the observation. For example, a model containing a background term is constructed... and observation items Target functional : Equation (1) In the formula, The analytical increment to be solved (i.e., the difference between the analytical field and the background field). This is the background error covariance matrix, used to propagate observed information to unobserved regions through statistical laws. To observe innovation; This is a tangent linear operator for the observation operator. To address the nonlinear characteristics of the satellite observation operator (such as the nonlinear relationship between AOD and concentration), the operator is linearized during the minimization iteration process. The solution is obtained using optimization algorithms such as L-BFGS. Minimized This achieves optimal statistical fusion of multi-source information while taking into account the error weights of each party.

[0068] In step S240, the analytical increment is superimposed onto the model background field to generate a pollutant assimilation analytical field.

[0069] Specifically, the analytical increment obtained from the iterative solution Background field superimposed on numerical mode This generates the final pollutant assimilation analysis field. This analytical field not only inherits the physical continuity of the numerical model in terms of dynamic transmission and chemical evolution, but also incorporates real environmental information contained in multi-source observation data, correcting the initial bias of the model.

[0070] Furthermore, the assimilation analysis field was directly determined as the background reference field for subsequent operations. This not only provided a high-quality "anchor point" that had been observed and corrected for subsequent real-time bias assessment, but also effectively suppressed the cumulative effect of model systematic errors as the forecast lead time increased, significantly improving the baseline accuracy of the entire forecast correction system.

[0071] Regarding the implementation details of generating dynamic reliability weights in step S140, in some examples of embodiments of this application, for any target grid point and any target contaminant in a preset spatiotemporal reference, the first... Data sources at time Observations And obtain the reference values ​​of the background reference field at the corresponding grid points. Calculate the instantaneous residual .

[0072] Here, to quantitatively evaluate the real-time performance of multi-source data, a unified comparison benchmark is established. Specifically, the background reference field has undergone data assimilation processing and possesses high spatial continuity and physical consistency. Using it as the "relative truth" can effectively capture the degree to which each independent data source deviates from the overall system state at the current moment. The instantaneous residual is calculated to provide the basic input for subsequent error decomposition.

[0073] Then, based on the instantaneous residual Using an exponentially weighted moving average algorithm that includes a forgetting factor, the th Observation error variance of each data source and systematic bias Perform online recursive updates to track the time-varying error characteristics of the corresponding data source: Equation (2) Equation (3) In the formula, A forgetting factor between 0 and 1, used to adjust the memory depth of historical error statistics for the current assessment; and Each is the previous moment The observation error variance and systematic bias, and and It is initialized based on the residual statistical results within a preset historical statistical period.

[0074] Here, considering the non-stationary nature of atmospheric sensor errors over time, an exponentially weighted moving average (EWMA) algorithm incorporating a forgetting factor is used to update the error statistics online, instead of using fixed historical statistical values.

[0075] Specifically, systematic bias is updated using equation (2), which incorporates the forgetting factor. By balancing the effects of historical bias memory and current residuals, zero-point drift caused by equipment aging or environmental changes can be sensitively tracked. Simultaneously, the observation error variance is updated using equation (3), which utilizes the bias-free residuals in equation (3). The calculation uses the square of the error, thus decoupling precision from accuracy—that is, measuring only the degree of discrete fluctuation in the data, without being affected by the magnitude of systematic deviations. Through an online recursive mechanism, the system can grasp the latest error status of each data source in real time without storing massive amounts of historical data.

[0076] Furthermore, a weighted model based on the inverse of the total mean square error is constructed, utilizing the updated... and Calculate the first Dynamic reliability weights for each data source This enables adaptive configuration for data sources with varying levels of reliability. Equation (4) In the formula, This represents the total number of valid data sources for the target pollutant at the target grid point, and ; Iterate through variables for the data source index. This is a pre-defined numerical stability term to prevent the denominator from being zero.

[0077] Here, a weighted model based on Bayesian principles is constructed, and the aforementioned separated bias and variance terms are recombined to generate the final dynamic reliability weights. Specifically, the first weight is calculated according to equation (4). Weights of each data source In the numerator of equation (4), The total mean square error of the data source represents the total uncertainty of the data; taking its reciprocal means that the smaller the error, the higher the score of the data source. As a numerically stable term, it prevents the anomaly of a zero denominator in the ideal case (zero error). The denominator is a normalization term, ensuring that all valid data sources ( The sum of the weights of each of the data sources is 1. In this way, the system can automatically assign high weights to high-quality data sources and extremely low weights to satellite data affected by cloud obstruction or faulty ground station data, thus achieving the adaptability and robustness of multi-source data fusion configuration.

[0078] Regarding the implementation details of generating the preliminary fused field for multiple pollutants in step S150, in some examples of the embodiments of this application, for observation grid points with direct observation data in the preset spatiotemporal reference and arbitrary target pollutants, the debiased observations from each data source are weighted and fused based on dynamic reliability weights to obtain the observation-side fused estimate. and observation uncertainty variance : Equation (5) Equation (6) Here, for grid points with direct observation data, the first level of observation-side fusion is performed, aiming to aggregate multi-source observation data of different sources and varying quality into a unified observation estimate for that grid point. Specifically, the observation-side fusion estimate is calculated using Equation (5). This formula uses the dynamic reliability weights generated in the previous steps to linearly combine the biased observation values, ensuring that high-quality data sources dominate the fusion result. At the same time, Equation (6) is used to calculate the corresponding observation uncertainty variance based on the error propagation law, thereby effectively solving the problem of conflict between satellite and ground monitoring data at a single grid point and outputting a set of comprehensive observation data with both high reliability and clear error limits.

[0079] Then, based on the reference values ​​of the background reference field at the corresponding grid points. and its background uncertainty variance Calculate the background fusion coefficient And generate point fusion estimates. and the corresponding fusion uncertainty variance : Equation (7) Equation (8) Equation (9) In the formula, It is obtained from the assimilation output error of the background reference field.

[0080] Here, the second-level "background-observation fusion" is performed, using the background reference field to perform a Bayesian update on the observation estimates to obtain the statistically optimal estimate. Specifically, the background fusion coefficient is first calculated using equation (7). This coefficient is essentially based on the observed variance. vs. background variance The relative magnitudes of the weights are adaptively determined: when the observation uncertainty is high, The value approaches 1, indicating that the system relies more on the background field; conversely, it approaches 0. Subsequently, equation (8) is used to generate the point fusion estimate. The fusion uncertainty variance is updated using equation (9). This ensures that even when observational data is present but noisy, the fusion results will not blindly deviate from the physically consistent background reference field, thus guaranteeing data stability.

[0081] Furthermore, for non-observation grid points lacking direct observation data in the preset spatiotemporal reference and for target pollutants, residual space mapping based on the background reference field is performed. For example, point fusion estimates are calculated at the observation grid points. Compared with background reference field values Fusion correction increment between Then, the uncertainty variance will be integrated. As an observation noise parameter, the fusion correction increment is calculated using Kriging interpolation, which considers the spatial covariance structure. The data is mapped to each grid point covering the entire domain, and the mapped increments are superimposed onto the background reference field to generate a preliminary fusion field of multiple pollutants that includes the global concentration distribution and the global uncertainty distribution.

[0082] In this embodiment, for non-observation grid points lacking direct observation data, residual space mapping based on the background reference field is performed to achieve full coverage. Unlike direct interpolation concentration, this embodiment first calculates the fusion correction increment at the observation grid points. This refers to the "correction amount for the background by the observation". Subsequently, the Kriging interpolation method is used to map this increment to the entire domain. In this process, an innovative approach is proposed to calculate the fusion uncertainty variance. This serves as the observation noise parameter in the Kriging model. This means that during the interpolation process, the influence of fusion points with high uncertainty on surrounding grid points will be automatically weakened, thus achieving noise-resistant and smooth interpolation. Finally, the interpolated incremental field is superimposed back onto the background reference field, and the generated multi-pollutant preliminary fusion field not only approximates the true value near the observation point, but also perfectly preserves the fine physical structures such as terrain obstruction and diffusion plume simulated by the numerical model (background field) in the unobserved area.

[0083] Regarding the implementation details of constructing the interactive calibration model in step S160, in some examples of embodiments of this application, the multi-pollutant coupling constraint includes at least a chemical equilibrium constraint based on the photochemical reaction mechanism.

[0084] Specifically, photochemically associated components in at least two classes of air pollutants are identified, including nitric oxide. Nitrogen dioxide ,ozone and volatile organic compounds Based on the photochemical quasi-steady-state assumption, for target grid points in a preset spatiotemporal reference where the solar radiation intensity is higher than a preset threshold, the current solar radiation intensity and temperature parameters are obtained, and the photolysis rate constant of nitrogen dioxide is calculated accordingly. and the reaction rate constant of ozone and nitric oxide. .

[0085] In some implementations, the interactive correction model does not blindly correlate all pollutants, but rather identifies photochemically correlated components with strong coupling characteristics based on atmospheric chemical mechanisms. In this embodiment, the system focuses on nitric oxide. Nitrogen dioxide ,ozone and volatile organic compounds As the core calibration target, based on the photochemical quasi-steady-state (PSS) assumption, the system first performs validity checks on each grid point in the preset spatiotemporal reference: only for solar radiation intensity higher than a preset threshold (e.g., The target lattice (characterized during daytime) is activated with chemical constraints. Subsequently, the current solar radiation intensity and temperature parameters of this lattice are obtained, and two key reaction kinetic parameters are calculated based on these parameters: the radiation-driven photolysis rate constant of nitrogen dioxide. and the temperature-driven rate constant of the ozone-nitric oxide reaction. This ensures that the introduction of chemical constraints is environmentally adaptive and avoids the fallacy of forcibly imposing photochemical constraints at night or when weather conditions are not met.

[0086] Then, a function describing the chemical equilibrium deviation of photochemically related components under steady-state light conditions is constructed. To characterize the multi-pollutant concentration vector to be corrected Degree of deviation from Leiden relation and its modified form: Equation (10) In the formula, For inclusion , and The concentration variable to be corrected; To characterize the equivalent reaction rate term for the conversion of nitric oxide by peroxy radicals generated from the oxidation of volatile organic compounds, this term is based on The grid point concentration and parameterization scheme for peroxy radical generation were determined.

[0087] Here, by constructing a chemical equilibrium deviation function Based on the modified Leiden relation, as shown in equation (10), the numerator term Characterized Photolysis generation and Rate (source term); principal term in denominator Characterized Titration generate The rate (sum term). In addition, the function also introduces... (Equivalent reaction rate term), which characterizes the reaction rate caused by... Oxidation produces peroxide free radicals ( , )right Transform into Additional contributions.

[0088] In one optional implementation of this application, a parameterized formula based on photochemical activity is used for calculation: Equation (11) In the formula, The predefined equivalent conversion coefficient of organic peroxide radicals is used to characterize the unit concentration. Under specific light conditions, it transforms into peroxide free radicals and oxidizes. The comprehensive capability, this coefficient can be determined based on the target area. Localized calibration of component characteristics (such as the ratio of anthropogenic to natural sources); Used as a reference photolysis rate constant (e.g., the maximum value at noon). This constitutes the photochemical activity factor, characterizing the degree to which the current radiation intensity drives the free radical generation reaction; This refers to the total concentration of volatile organic compounds (or the concentration of active components). The variable is the concentration of nitric oxide to be corrected.

[0089] In equation (11), right The conversion rate essentially depends on three factors: "fuel" ( "Power" (solar radiation) ) and "reaction object" ( Using this parameterization formula, the model can approximate the nonlinear enhancement effect of ozone formation under "strong radiation + high VOCs" conditions with extremely low computational cost, without needing to run time-consuming fully coupled chemical mechanisms (CBM or SAPRC), thus correcting the dependence solely on inorganic reactions. Forecast deviations caused by ( ).

[0090] It should be noted that in traditional Leighton relationships, this is often overlooked. The effect of this leads to calculation deviation, while this embodiment uses... Concentration and introduction determined by parameterization This significantly improves the model's applicability in areas with high volatile organic compound emissions, such as urban centers. When the atmosphere is in photochemical equilibrium, the ratio of this fraction should approach 1, meaning the function value should approach 0.

[0091] Furthermore, a chemical constraint penalty term is constructed and introduced into the objective function of the interactive calibration model, enabling the interactive calibration model to correct the concentration vector while maintaining data fidelity. Apply constraints to cause chemical equilibrium to deviate to a certain degree. It tends to minimize, thereby ensuring the physicochemical consistency among the concentrations of multiple pollutants.

[0092] Here, the aforementioned deviation function is transformed into a chemical constraint penalty term and introduced into the objective function of optimization, requiring the corrected concentration vector to be... Make the norm Minimize. Therefore, it does not mandate that data strictly obey chemical equations (allowing for certain measurement errors and non-equilibrium perturbations), but it will affect correction results that severely violate physical laws (e.g., extremely high levels of radiation under high-radiation conditions). Accompanied by extremely high (In cases where) huge penalties are imposed.

[0093] In this way, the interactive calibration model can "force" the values ​​of multiple pollutants to compromise with each other during the adjustment process, thereby ensuring that the final output forecast results are consistent in physicochemical logic and effectively eliminating the phenomenon of "accurate data but absurd mechanism" that may occur with traditional statistical methods.

[0094] Regarding the implementation details of constructing the interactive correction model in step S160, in some examples of the embodiments of this application, the multi-pollutant coupling constraint includes at least the statistical correlation constraint based on historical observation data.

[0095] Specifically, the historical long-term observation sequences and corresponding meteorological field data of the target area are obtained. Clustering algorithms are used to divide the historical time periods into multiple typical meteorological circulation patterns, and multi-pollutant empirical correlation coefficient matrices are calculated for each meteorological circulation pattern to form a historical correlation map library. Then, the meteorological parameters at the current time to be processed are obtained, and their similarity to each meteorological circulation pattern in the historical correlation map library is calculated. Based on the principle of maximum similarity, the target correlation coefficient matrix applicable to the current time is obtained. .

[0096] Here, a historical correlation map database categorized by meteorological type is constructed to address the challenge of dynamically changing correlations between air pollutants under varying meteorological conditions. In this embodiment, the system first acquires long-term historical observation sequences and corresponding meteorological field data for the target area. Using clustering algorithms such as K-Means, the historical time periods are divided into multiple typical meteorological circulation patterns (e.g., "summer ozone type," "winter stable type"). For each type, a multi-pollutant empirical correlation coefficient matrix is ​​calculated (e.g., summer...). and (A negative correlation is observed, while a weak positive correlation may be observed in winter). During real-time processing, the system acquires the meteorological parameters for the current time, calculates their similarity (e.g., cosine similarity) with typical circulation patterns in the map library, and matches them according to the maximum similarity principle to obtain the target correlation coefficient matrix most suitable for the current time. This ensures that the statistical laws referenced by the correction model are adaptive to different scenarios, avoiding constraint failure caused by using a single fixed historical average correlation.

[0097] For example, a clustering algorithm is used to divide historical time periods into A typical meteorological circulation pattern is identified, and the empirical correlation coefficient matrix for each type is calculated. During real-time processing, the meteorological feature vector at the current moment is obtained. (Including wind speed, humidity, boundary layer height, etc.), calculate its relationship with the first image in the map library. The central eigenvector of each circulation pattern similarity between : Equation (12) System selection makes The matrix corresponding to the largest circulation pattern is used as the target correlation coefficient matrix at the current time. This ensures that the model always calls upon historical statistical patterns most similar to the current atmospheric diffusion conditions, achieving adaptive matching.

[0098] Next, construct the real-time correlation coefficient matrix function. Define a context containing the current pending time. and the past A sliding time window is used to construct a multi-pollutant concentration vector to be corrected at the current moment. Compared with historical concentration sequences that have been corrected within the sliding time window The real-time sequence is composed of these sequences, and the real-time correlation coefficient matrix is ​​calculated accordingly. Furthermore, a statistical constraint penalty term is established to constrain... The value of makes the real-time correlation coefficient matrix... Correlation coefficient matrix with target The distance between the Frobenius norms tends to be minimized, thus ensuring the consistency of the correction results with historical statistical patterns under specific meteorological conditions.

[0099] Here, the function for constructing the real-time correlation coefficient matrix is ​​described. This allows for dynamic statistical constraints on the current correction results. Since correlation cannot be calculated from single-time data, this embodiment employs a sliding time window strategy, defining a window that includes the current time frame. and the past A time window (e.g., the past 24 hours) is used to construct a vector to be corrected at the current time. Compared with the historical sequence that has been calibrated within the window The system generates a mixed sequence. Based on this mixed sequence, it calculates the correlation coefficient matrix in real time. .

[0100] It should be noted that, since statistical correlation cannot be calculated at a single moment, this embodiment constructs a system that includes the current moment to be processed. and the past At any given moment (e.g.) The sliding time window sequence matrix : Equation (13) In the formula, This is the corrected historical concentration vector. This is the concentration vector to be solved. Based on this matrix... Calculate the Pearson correlation coefficients between each column (i.e., different pollutants) to generate a real-time correlation coefficient matrix. By introducing historical sequences, a random sequence was constructed. The changing dynamic matrix allows correlation constraints to be mathematically transformed into constraints on the current variables. Constraints.

[0101] In addition, to handle cold start or data loss issues, when If insufficient, adaptively introduce preliminary fused field or background reference field data from the same period to fill the gap. Finally, establish a statistical constraint penalty term, requiring the optimized... Make and The Frobenius norm distance between them is minimized. This forces the temporal evolution trend of the correction results to conform to the historical statistical regularity under specific meteorological conditions, effectively suppressing the physical-logical inconsistency caused by abrupt changes in a single pollutant.

[0102] For example, a statistical constraint penalty term is established and introduced into the objective function. To quantify the difference between the real-time matrix and the target matrix, this embodiment uses the Frobenius norm as a distance metric to construct the penalty term: Equation (14) In the formula, For the types and quantities of pollutants, Characterizing the first Class and the The correlation coefficients between pollutant classes. The optimization solver forces the corrected [contaminant] to [perform a certain function] by minimizing this penalty term. Numerically, it is possible to maintain the covariant relationships of multiple pollutants that should exist under these meteorological conditions (e.g., forced covariance during high summer temperatures). and (Exhibiting a negative correlation) effectively eliminates outlier corrected solutions that do not conform to statistical laws.

[0103] Regarding the implementation details of the interactive correction process in step S160, in some examples of the embodiments of this application, the fusion uncertainty variance of the generated multi-pollutant preliminary fusion field at each grid point is extracted, and a diagonal matrix is ​​constructed with the reciprocal of the sum of the fusion uncertainty variance and the preset numerical stability term as its diagonal elements, forming the observation confidence weight matrix. .

[0104] In some implementations, the fusion uncertainty variance at each grid point generated in step S150 can be extracted. And construct a diagonal matrix whose diagonal elements are Thus, based on the principle of inverse variance weighting: grid points with larger variances indicate less reliable data, and therefore should have smaller weights in the optimization objective; conversely, grid points with high determinism will have large weights, exerting strong constraints on the correction results. This is achieved by introducing a pre-defined numerical stability term. This avoids the computational risk of weights tending to infinity when the variance is extremely small. This matrix... It will serve as the core metric for data fidelity in subsequent optimization functions, thus realizing a logical closed loop from "uncertainty assessment" to "correction decision".

[0105] Then, construct a multi-objective joint optimization loss function. The loss function consists of a data fidelity term representing the degree of observation deviation, a chemical constraint term representing the degree of violation of physical laws, and a statistical constraint term representing the inconsistency of statistical characteristics. , In the formula, The solution to be found at time 1 Multi-pollutant corrected concentration vector, For the initial fusion field of multiple pollutants at a certain moment The vector representation of , for The corresponding deviation from chemical equilibrium, and These are the regularization parameters used to balance data fidelity terms and model constraint strength. This is a vector transpose operation.

[0106] Here, a multi-objective joint optimization loss function is constructed. To find the optimal balance between data observation and physical laws, a multi-objective joint optimization loss function is employed. It mainly consists of three parts: The first term is the data fidelity term: essentially the square of the weighted Mahalanobis distance, used to constrain the vector to be solved. Do not deviate from the initial fusion field Too far, and the cost of deviation is determined by the weight matrix. Decide; The second term is the chemical constraint term: it uses the chemical equilibrium deviation function to penalize solutions that violate the Leiden relation and its modified form, forcing... and Maintain photochemical balance; The third term is the statistical constraint term: using the previously defined norm distance penalty for solutions that violate historical statistical patterns under the current weather conditions.

[0107] also, and As regularization parameters (similar to Lagrange multipliers), they are used to adjust the strength of physical mechanism constraints and statistical regularity constraints relative to data observations, respectively, and can be regarded as a balancing lever between "physical knowledge" and "data-driven".

[0108] Furthermore, the loss function of the multi-objective joint optimization is iteratively solved to search for the optimal solution. The optimal solution vector that reaches the minimum value and the optimal solution vector The results were determined to be multi-pollutant corrected forecasts; among them, the criteria for determining the termination of iterative convergence include the change in the objective function value being lower than a preset threshold and / or the change in the solution vector between two adjacent iterations being lower than a preset threshold.

[0109] Specifically, the multi-objective loss function defined above is numerically solved to obtain the final prediction result. Since the chemical and statistical constraints introduce nonlinear characteristics, this embodiment employs iterative algorithms suitable for high-dimensional nonlinear optimization, such as the conjugate gradient method or quasi-Newton methods (e.g., L-BFGS). Thus, from the initial fusion field... Starting from this point, iteratively search in the opposite direction of the gradient of the loss function, continuously updating... Continue until the convergence condition is met. The convergence criterion is set as the objective function value. The decrease in the solution vector is lower than a preset threshold, and / or the change in the solution vector between two adjacent iterations is lower than a preset threshold. The final output is the optimal solution vector. This means that the multi-pollutant correction forecast results retain the information from multiple sources to the greatest extent and maintain a high degree of consistency in physicochemical properties and statistical characteristics, effectively solving the technical problem of "single index is accurate but multiple indexes are contradictory" that often occurs in traditional methods.

[0110] Figure 3 The diagram illustrates an example of the operational mechanism of an atmospheric multi-pollutant forecast correction method that integrates multi-source observations and model outputs according to an embodiment of this application. It visually demonstrates the entire data flow and logical loop of multi-source data from acquisition to final correction output in the embodiment of this application.

[0111] like Figure 3As shown, at the input end (left side), the system aggregates multi-source heterogeneous data, including aerosol optical thickness and vertical profile information provided by satellite observations, background meteorological and concentration fields output by numerical models (such as WRF-CMAQ), and ground observation data covering standard air quality monitoring stations and low-cost sensors. WRF-CMAQ is a coupled numerical simulation system consisting of the Weather Research and Forecasting (WRF) model and the Community Multiscale Air Quality (CMAQ) model. After these data enter the core processing stage, a physically continuous background reference field is first generated through the data assimilation module (using Kalman filtering or three-dimensional variational techniques). Subsequently, a preliminary multi-pollutant fused field is generated in the data fusion and weighting module by combining dynamic reliability assessment.

[0112] By employing a deep coupling between the cross-pollutant interaction (APIC) module and the "optimization and reliability weighting" module (middle section), the system introduces photochemical reaction mechanisms and historical statistical patterns as constraints at this stage to construct a multi-objective joint optimization model and perform physicochemical consistency correction on the initial fusion field. Finally, the corrected multi-pollutant forecast results (covering...) are generated at the output (right side). , , , (and other components), and simultaneously outputs the uncertainty assessment index of the forecast result, thus realizing a highly reliable air quality forecast correction scheme that takes into account data accuracy, physical mechanisms and statistical laws.

[0113] To verify the effectiveness of the atmospheric multi-pollutant forecast correction method that integrates multi-source observations and model outputs proposed in the embodiments of this application, a comparative experiment with four schemes was designed: Scheme 1 is the "raw numerical model output (Baseline)", which uses only WRF. CMAQ simulation results; Scheme 2 is "traditional data assimilation correction", which only performs three-dimensional variational assimilation to update the initial field; Scheme 3 is "conventional multi-source data fusion", which performs fusion interpolation based on fixed weights on the basis of assimilation but does not introduce multi-pollutant coupling constraints; Scheme 4 is "the method of this application embodiment (full-process interactive correction)".

[0114] The experiment selected a multi-source dataset spanning one year from a typical region, covering ground-based national monitoring stations, low-cost sensor networks, and FY data. 4 / Himawari 8 satellites invert AOD and WRF CMAQ model forecast data. Following a time-series priority principle, the dataset is divided into a historical database construction phase (first 9 months) and an online validation phase (last 3 months): the former is mainly used to build historical correlation maps for different weather types and initialize error statistics parameters for each data source; the latter is used to simulate real-time rolling forecast correction under actual operational scenarios. PM 2.5 PM 10 NO2 and O3 were used as key evaluation indicators. The root mean square error (RMSE) and mean absolute error (MAE) were used to quantitatively evaluate the correction accuracy and physical consistency of each scheme under different pollutants and different forecast lead times.

[0115] Figure 4 A comparative graph showing the root mean square error (RMSE) distribution of four different pollutant forecasting schemes is presented. This graph, presented as a violin plot, visually illustrates the statistical distribution characteristics of the errors of the four schemes during the validation period. The horizontal axis corresponds to the original numerical model (model only), data assimilation correction, conventional data fusion, and the method described in this application (the method described herein). The vertical axis represents the magnitude of the root mean square error (RMSE). Different pollutants are represented using their commonly used concentration units, with particulate matter (PM2.5) listed in the figure. 2.5 PM 10 )by The graph indicates that gaseous pollutants (NO2, O3, etc.) are represented by ppb, the width of the graph represents the sample distribution density, and the height represents the discrete fluctuation range of the error.

[0116] like Figure 4 It can be seen that the "model-only" output without post-processing has the highest mean RMSE and the widest distribution range, indicating that the original forecast has a large systematic bias and random error. Although the "data assimilation" and "data fusion" schemes reduce the overall error level, they still show large dispersion in some periods (the graphic shape is relatively thin and long). In contrast, the method of this application's embodiment shows the best performance on all four types of pollutants: its RMSE distribution position shifts to the lowest level overall, indicating that the accuracy of the corrected forecast is significantly improved; at the same time, the graphic shape corresponding to this method is the most convergent and concentrated in the vertical direction, which means that the variance of the error distribution is the smallest. Figure 4 The experimental results corroborate that by introducing a dynamic reliability weighting – adaptive pollutant interaction correction (DRW-APIC) mechanism, the method provided in this application not only effectively reduces the systematic bias of multi-source data, but also significantly enhances the forecast stability and robustness of the system in the face of complex atmospheric environmental changes.

[0117] Figure 5 A two-dimensional sensitivity heatmap of the influence of key algorithm parameters on prediction error in the method according to an embodiment of this application is shown. To verify the stability of the algorithm under different configurations, this experiment plotted parameter sensitivity analysis heatmaps. Figure 5 Middle ordinate The horizontal axis represents the reliability assessment parameter (corresponding to the forgetting factor in equations (2) and (3) in some embodiments), used to adjust the model's memory depth of historical error statistics; This represents the cross-pollutant coupling strength parameter, which is the regularization strength in the multi-objective joint optimization loss function, specifically a chemical constraint parameter. With statistical constraint parameters Both are used to control the weight of physical / statistical laws relative to the data fidelity term; in this sensitivity experiment, the following settings are used: and Maintain synchronous changes (i.e.) To examine the impact of comprehensive constraint strength on the correction effect.

[0118] like Figure 5 As shown, the low-value areas of the root mean square error (RMSE) (i.e., the darkest purple area in the figure) are concentrated in the central region of the parameter space, roughly corresponding to the vertical axis. The x-coordinate is in the interval [0.5, 0.8]. The range is within the [0.4, 0.8] interval. This distribution characteristic reveals the optimal parameter configuration logic: on the one hand, the reliability assessment parameter (corresponding to the forgetting factor) needs to be kept moderate, so as to smooth the random noise of the observation data by retaining historical memory, and to maintain sufficient sensitivity to track the real-time drift of equipment error; on the other hand, the cross-pollutant coupling strength (corresponding to the regularization parameter) also needs to be moderate, so as to effectively introduce physical and statistical constraints while fully respecting the fidelity of the observation data, and avoid the mechanism failure caused by excessively strong constraints masking the true observation or by insufficiently weak constraints.

[0119] In particular, Figure 5 The low-to-medium error region exhibits a wide, gently sloping "basin-like" distribution rather than steep local extreme points. This strongly demonstrates that the method in this application has significant robustness to parameter selection. Even if the parameters fluctuate within a certain range in actual business applications, the model can still maintain stable high-precision forecasting performance, reflecting strong generalization ability.

[0120] In this embodiment, by constructing a multi-level data assimilation and fusion system, the limitations of insufficient spatiotemporal coverage of single data sources and lack of physical constraints for single pollutant correction in current related technologies are significantly overcome. At the data level, this method not only achieves multi-source complementarity from satellite remote sensing, ground-based national control stations, and low-cost sensors, but also innovatively introduces a dynamic reliability weighting algorithm (DRW) based on the forgetting factor. This algorithm can track and adaptively compensate for error drift of each data source (such as sensor aging or environmental interference) in real time, ensuring that the system maintains extremely high robustness and correction accuracy when facing a variable atmospheric observation environment.

[0121] At the mechanistic level, the core cross-pollutant interaction correction (APIC) module of this invention completely overcomes the drawbacks of independent correction for each pollutant in traditional methods. By embedding chemical equilibrium constraints based on photochemical reaction mechanisms (such as modified Leighton relations) and historical statistical correlation constraints based on meteorological classifications into the multi-objective optimization model, this method forces multi-pollutant forecast results to numerically approximate the observed true values ​​while strictly adhering to the evolutionary laws of atmospheric physicochemical processes. This effectively eliminates the phenomenon of "accurate data but contradictory mechanisms" caused by observational noise or algorithm overfitting, ensuring... , , The key components exhibit a high degree of consistency in their material transformation logic.

[0122] Furthermore, thanks to the end-to-end variance recursive calculation, this invention possesses the ability to quantify inherent uncertainty, which can assist in outputting the reliability distribution of forecast results. In summary, this scheme achieves deep coupling between statistical optimization techniques and atmospheric physical mechanisms, demonstrating significant advantages in improving forecast accuracy, stability, and physical interpretability, and providing strong technical support for refined air pollution prevention and control.

[0123] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0124] Figure 6 A structural block diagram of an example atmospheric multi-pollutant forecast correction system that integrates multi-source observations and model outputs according to an embodiment of this application is shown.

[0125] like Figure 6 As shown, the atmospheric multi-pollutant forecast correction system 600, which integrates multi-source observations and model outputs, includes a data acquisition unit 610, a spatiotemporal reference conversion unit 620, an assimilation reference field generation unit 630, a dynamic reliability weighting unit 640, a multi-source weighted fusion unit 650, and an interactive correction solution unit 660.

[0126] The data acquisition unit 610 is used to acquire multi-source air pollutant data and numerical model output data of the target area at the time to be processed and its corresponding preset forecast period. The multi-source air pollutant data is used to characterize the observation information of at least two types of air pollutants, and includes at least ground monitoring data and / or satellite remote sensing data. The numerical model output data is used to characterize the model forecast information of the at least two types of air pollutants.

[0127] The spatiotemporal reference conversion unit 620 is used to convert the multi-source atmospheric pollutant data and the numerical model output data to a unified preset spatiotemporal reference, so as to obtain a multi-source grid data set corresponding to the preset spatiotemporal reference.

[0128] The assimilation reference field generation unit 630 is used to perform data assimilation processing on the numerical model output data based on the multi-source grid data set, so as to update the initial state of the model under observation constraints and generate a pollutant assimilation analysis field that reflects the spatial distribution characteristics of pollutants, and to determine the pollutant assimilation analysis field as a background reference field for characterizing the spatiotemporal distribution of the at least two types of atmospheric pollutants.

[0129] The dynamic reliability weighting unit 640 is used to evaluate the reliability index of each data source at the current moment based on the multi-source grid data set, for each grid point and each pollutant in the preset spatiotemporal reference, combined with the historical error statistics of each data source and the real-time deviation information of each data source relative to the background reference field, and generate the dynamic reliability weight of each data source according to the reliability index.

[0130] The multi-source weighted fusion unit 650 is used to perform weighted fusion processing on the background reference field and the observation data in the multi-source grid data set based on the dynamic reliability weight, so as to generate a multi-pollutant preliminary fusion field containing the at least two types of atmospheric pollutants.

[0131] The interactive calibration solver unit 660 is used to construct an interactive calibration model containing multi-pollutant coupling constraints, and to perform pollutant interactive calibration processing on the preliminary fusion field of the multi-pollutant using the interactive calibration model to generate multi-pollutant calibration forecast results; the multi-pollutant coupling constraints are used to characterize the interaction relationship between pollutants and constrain the physicochemical consistency and statistical correlation between different pollutant forecast results.

[0132] In some embodiments, this application provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions. The execution instructions can be read and executed by electronic devices (including but not limited to computers, servers, or network devices) to perform the steps of any of the above-described atmospheric multi-pollutant forecast correction methods that integrate multi-source observations and model outputs.

[0133] In some embodiments, this application also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of any of the above-described methods for correcting atmospheric multi-pollutant forecasts by fusing multi-source observations and model outputs.

[0134] In some embodiments, this application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of an atmospheric multi-pollutant forecast correction method that fuses multi-source observations and model outputs.

[0135] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0136] The electronic devices in this application can exist in various forms, including but not limited to: mobile communication devices, ultra-mobile personal computer devices, portable entertainment devices, or other airborne electronic devices with data interaction functions.

[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for correcting atmospheric multi-pollutant forecasts by integrating multi-source observations and model outputs, characterized in that, The method includes: The system acquires multi-source atmospheric pollutant data and numerical model output data for the target area at the time to be processed and its corresponding preset forecast period. The multi-source atmospheric pollutant data is used to characterize the observation information of at least two types of atmospheric pollutants and includes at least ground monitoring data and / or satellite remote sensing data. The numerical model output data is used to characterize the model forecast information of the at least two types of atmospheric pollutants. The multi-source air pollutant data and the numerical model output data are converted to a unified preset spatiotemporal reference to obtain a multi-source grid data set corresponding to the preset spatiotemporal reference. Based on the multi-source grid data set, data assimilation processing is performed on the numerical model output data to update the initial state of the model under observation constraints and generate a pollutant assimilation analysis field that reflects the spatial distribution characteristics of pollutants. The pollutant assimilation analysis field is then determined as a background reference field for characterizing the spatiotemporal distribution of the at least two types of atmospheric pollutants. Based on the multi-source grid data set, for each grid point and each pollutant in the preset spatiotemporal reference, combined with the historical error statistics of each data source and the real-time deviation information of each data source relative to the background reference field, the reliability index of each data source at the current moment is evaluated, and the dynamic reliability weight of each data source is generated according to the reliability index. Based on the dynamic reliability weight, the background reference field and the observation data in the multi-source grid data set are subjected to weighted fusion processing to generate a preliminary multi-pollutant fusion field containing at least two types of atmospheric pollutants. An interactive calibration model containing multi-pollutant coupling constraints is constructed, and the interactive calibration model is used to perform pollutant interactive calibration processing on the preliminary fusion field of the multi-pollutant to generate multi-pollutant calibration forecast results. The multi-pollutant coupling constraints are used to characterize the interaction relationship between pollutants and constrain the physicochemical consistency and statistical correlation between forecast results of different pollutants.

2. The method according to claim 1, characterized in that, The step of converting the multi-source atmospheric pollutant data and the numerical model output data to a unified preset spatiotemporal reference to obtain a multi-source gridded data set corresponding to the preset spatiotemporal reference includes: Quality control was performed on the ground monitoring data and satellite remote sensing data in the multi-source atmospheric pollutant data. Outliers were removed based on the statistical threshold of the sliding time window, and abrupt change records were removed by combining the spatial consistency test of adjacent stations to obtain the cleaned observation dataset. A model-assisted vertical transformation is performed on the satellite remote sensing data. The vertical profile factor is calculated using the vertical distribution information of aerosols in the numerical model output data. Combined with the planetary boundary layer height parameter, the vertical column concentration or optical thickness of the satellite remote sensing data is inverted and mapped to near-surface concentration data. For the cleaned ground monitoring data and the near-ground concentration data, an adaptive spatial reconstruction is performed. Based on the grid distribution of the preset spatiotemporal reference, a spatial interpolation algorithm that considers distance weight and terrain elevation differences is used to map the non-uniformly distributed observation points to the grid center to form gridded observation data. Spatiotemporal alignment and unification are performed on the gridded observation data and the numerical model output data. Each data is resampled temporally according to the temporal resolution of the time to be processed. In the case that the original grid of the numerical model output data is inconsistent with the preset spatiotemporal reference, spatial resampling is performed on the numerical model output data, thereby generating a multi-source gridded data set with consistent spatiotemporal dimensions.

3. The method according to claim 1, characterized in that, The process involves performing data assimilation processing on the numerical model output data based on the multi-source gridded data set. This updates the initial state of the model under observation constraints and generates a pollutant assimilation analysis field reflecting the spatial distribution characteristics of pollutants. This pollutant assimilation analysis field is then used as a background reference field to characterize the spatiotemporal distribution of at least two types of atmospheric pollutants, including: For the ground monitoring data in the multi-source grid data set, super-observation construction processing is performed. Within a preset time assimilation window, based on the spatial resolution of the preset spatiotemporal reference, multiple high-frequency observation samples falling into the same grid are spatiotemporally aggregated to generate representative observation samples that match the grid scale, thereby reducing the representativeness error and spatial correlation of the observation data. An assimilation observation system is constructed, which assembles the representative observation samples and satellite remote sensing data in the multi-source grid data set into an observation vector, and constructs an observation operator for mapping the numerical model state space to the observation space, and configures corresponding observation error covariance parameters for different observation types. A three-dimensional variational assimilation model is constructed, using the numerical model output data as the model background field. The weight relationship between the model background field and the observation vector is defined using the background error covariance matrix and the observation error covariance matrix. An objective functional containing background and observation terms is constructed, and the objective functional is iteratively minimized to obtain the analysis increment. Where the observation operator is nonlinear, linearization is performed on the observation operator during the iteration process. The analytical increment is superimposed onto the model background field to generate the pollutant assimilation analytical field.

4. The method according to claim 1, characterized in that, Based on the multi-source grid data set, for each grid point and each contaminant in the preset spatiotemporal reference, and combining the historical error statistics of each data source and the real-time deviation information of each data source relative to the background reference field, the reliability index of each data source at the current moment is evaluated, and a dynamic reliability weight of each data source is generated according to the reliability index, including: For any target grid point and any target pollutant in the preset spatiotemporal reference, obtain the first... Data sources at time Observations And obtain the reference value of the background reference field at the corresponding grid point. Calculate the instantaneous residual ; Based on instantaneous residual Using an exponentially weighted moving average algorithm that includes a forgetting factor, the first... Observation error variance of each data source and systematic bias Perform online recursive updates to track the time-varying error characteristics of the corresponding data source: , , In the formula, A forgetting factor between 0 and 1, used to adjust the memory depth of historical error statistics for the current assessment; and Each is the previous moment The observation error variance and systematic bias, and and It is initialized based on the residual statistical results within a preset historical statistical period; Construct a weighted model based on the inverse of the total mean square error, and utilize the updated... and Calculate the first Dynamic reliability weights of each data source This enables adaptive configuration for data sources with varying levels of reliability. , In the formula, The total number of valid data sources for the target pollutant at the target grid point, and ; Iterate through variables for the data source index. This is a pre-defined numerical stability term to prevent the denominator from being zero.

5. The method according to claim 4, characterized in that, The step of performing weighted fusion processing on the background reference field and the observation data in the multi-source gridded data set based on the dynamic reliability weight to generate a preliminary multi-pollutant fusion field containing at least two types of atmospheric pollutants includes: For observation grid points with direct observation data in the preset spatiotemporal reference and any target pollutant, the bias-free observations from each data source are weighted and fused based on the dynamic reliability weight to obtain the observation-side fused estimate. and observation uncertainty variance : , ; Based on the reference values ​​of the background reference field at the corresponding grid points and its background uncertainty variance Calculate the background fusion coefficient And generate point fusion estimates. and the corresponding fusion uncertainty variance : , , , In the formula, It is estimated from the assimilation output error of the background reference field; For the non-observation grid points in the preset spatiotemporal reference lacking directly observed data and the target pollutant, perform residual space mapping based on the background reference field: The point fusion estimate is calculated at the observation grid point. Compared with background reference field values Fusion correction increment between ; The fusion uncertainty variance As an observation noise parameter, the fusion correction increment is calculated using Kriging interpolation, which considers the spatial covariance structure. The data is mapped to each grid point covering the entire domain, and the mapped increments are superimposed onto the background reference field to generate a preliminary fusion field of multiple pollutants that includes the global concentration distribution and the global uncertainty distribution.

6. The method according to claim 5, characterized in that, The multi-pollutant coupling constraint includes at least a chemical equilibrium constraint based on the photochemical reaction mechanism; The construction of the interactive correction model containing multi-pollutant coupling constraints includes: Identify photochemically correlated components in at least two classes of air pollutants, including nitric oxide. Nitrogen dioxide ,ozone and volatile organic compounds ; Based on the photochemical quasi-steady-state assumption, for target grid points in the preset spatiotemporal reference where the solar radiation intensity is higher than a preset threshold, the current solar radiation intensity and temperature parameters are obtained, and the photolysis rate constant of nitrogen dioxide is calculated accordingly. and the reaction rate constant of ozone and nitric oxide. ; Construct a function describing the chemical equilibrium deviation of the photochemically correlated components under steady-state light conditions. To characterize the multi-pollutant concentration vector to be corrected Degree of deviation from Leiden relation and its modified form: , In the formula, For inclusion , and The concentration variable to be corrected; To characterize the equivalent reaction rate term for the conversion of nitric oxide by peroxy radicals generated from the oxidation of volatile organic compounds, this term is based on The grid point concentration and parameterization scheme for peroxy radical generation were determined. A chemical constraint penalty term is constructed and incorporated into the objective function of the interactive calibration model, enabling the interactive calibration model to correct the concentration vector while maintaining data fidelity. Apply constraints to cause the chemical equilibrium to deviate by a certain degree. It tends to minimize, thereby ensuring the physicochemical consistency among the concentrations of multiple pollutants.

7. The method according to claim 6, characterized in that, The multi-pollutant coupling constraint includes at least a statistical correlation constraint based on historical observation data; The construction of the interactive correction model containing multi-pollutant coupling constraints includes: The historical long-term observation sequence and corresponding meteorological field data of the target area are obtained. The historical period is divided into multiple typical meteorological circulation patterns using a clustering algorithm. The empirical correlation coefficient matrix of multiple pollutants under each meteorological circulation pattern is calculated to form a historical correlation map library. Obtain the meteorological parameters at the current time to be processed, calculate their similarity with each meteorological circulation pattern in the historical correlation map library, and obtain the target correlation coefficient matrix applicable to the current time based on the principle of maximum similarity. ; Constructing a real-time correlation coefficient matrix function Define a context containing the current pending time. and the past A sliding time window is used to construct a multi-pollutant concentration vector to be corrected at the current moment. Compared with historical concentration sequences that have been corrected within the sliding time window The real-time sequence is composed of these sequences, and the real-time correlation coefficient matrix is ​​calculated accordingly. ; Establish statistical constraint penalty terms to constrain... The value of makes the real-time correlation coefficient matrix... Correlation coefficient matrix with the target The distance between the Frobenius norms tends to be minimized, thus ensuring the consistency of the correction results with historical statistical patterns under specific meteorological conditions.

8. The method according to claim 7, characterized in that, The interaction correction model is used to perform pollutant interaction correction processing on the preliminary fusion field of multiple pollutants to generate multi-pollutant correction forecast results, including: Extract the fusion uncertainty variance of the generated preliminary multi-pollutant fusion field at each grid point, and construct a diagonal matrix with the reciprocal of the sum of the fusion uncertainty variance and the preset numerical stability term as its diagonal elements to form the observation confidence weight matrix. ; Constructing a multi-objective joint optimization loss function The loss function is composed of a data fidelity term characterizing the degree of observation deviation, a chemical constraint term characterizing the degree of violation of physical laws, and a statistical constraint term characterizing the inconsistency of statistical characteristics. , In the formula, The solution to be found at time 1 Multi-pollutant corrected concentration vector, For the initial fusion field of multiple pollutants at a certain moment The vector representation of , for The corresponding deviation from chemical equilibrium, and These are the regularization parameters used to balance data fidelity terms and model constraint strength. This is a vector transpose operation; The multi-objective joint optimization loss function is iteratively solved to search for the optimal solution. The optimal solution vector that reaches the minimum value and the optimal solution vector The result is determined to be the multi-pollutant correction forecast result; wherein, the determination condition for the termination of iterative convergence includes the change in the objective function value being lower than a preset threshold and / or the change in the solution vector between two adjacent iterations being lower than a preset threshold.

9. An atmospheric multi-pollutant forecasting and correction system integrating multi-source observations and model outputs, characterized in that, The system includes: The data acquisition unit is used to acquire multi-source atmospheric pollutant data and numerical model output data of the target area at the time to be processed and its corresponding preset forecast period. The multi-source atmospheric pollutant data is used to characterize the observation information of at least two types of atmospheric pollutants, and includes at least ground monitoring data and / or satellite remote sensing data. The numerical model output data is used to characterize the model forecast information of the at least two types of atmospheric pollutants. The spatiotemporal reference conversion unit is used to convert the multi-source atmospheric pollutant data and the numerical model output data to a unified preset spatiotemporal reference, so as to obtain a multi-source grid data set corresponding to the preset spatiotemporal reference. The assimilation reference field generation unit is used to perform data assimilation processing on the numerical model output data based on the multi-source grid data set, so as to update the initial state of the model under observation constraints and generate a pollutant assimilation analysis field that reflects the spatial distribution characteristics of pollutants, and to determine the pollutant assimilation analysis field as a background reference field for characterizing the spatiotemporal distribution of the at least two types of atmospheric pollutants. The dynamic reliability weighting unit is used to evaluate the reliability index of each data source at the current moment based on the multi-source grid data set, for each grid point and each pollutant in the preset spatiotemporal reference, combined with the historical error statistics of each data source and the real-time deviation information of each data source relative to the background reference field, and generate the dynamic reliability weight of each data source according to the reliability index. A multi-source weighted fusion unit is used to perform weighted fusion processing on the background reference field and the observation data in the multi-source grid data set based on the dynamic reliability weight, so as to generate a multi-pollutant preliminary fusion field containing at least two types of atmospheric pollutants. The interactive calibration solution unit is used to construct an interactive calibration model containing multi-pollutant coupling constraints, and to perform pollutant interactive calibration processing on the preliminary fusion field of the multi-pollutant using the interactive calibration model to generate multi-pollutant calibration forecast results; the multi-pollutant coupling constraints are used to characterize the interaction relationship between pollutants and constrain the physicochemical consistency and statistical correlation between forecast results of different pollutants.