Heterogeneous data fusion and trend evaluation method for long-term atmospheric background observation

By combining dynamic Bayesian calibration and cross-platform consistency constraints with physical constraint embedding, this method solves the problems of multi-source data bias and deep learning fusion in long-term atmospheric background observations, achieving high-precision multi-factor attribution and trend assessment, and providing accurate support for climate change attribution.

CN122196927APending Publication Date: 2026-06-12METEOROLOGICAL BUREAU OF DIQING TIBETAN AUTONOMOUS PREFECTURE YUNNAN PROVINCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-09
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing technologies suffer from systematic biases in multi-source data, insufficient reliability of deep learning fusion, and low accuracy of multi-factor attribution in long-term atmospheric background observations. These issues result in insufficient accuracy of fused data and high uncertainty in trend assessment results, making it difficult to meet the needs of accurate climate change attribution and policy formulation.

Method used

Data calibration is performed using a dynamic Bayesian calibration model and a cross-platform consistency constraint method. Combined with a deep learning model that uses physical constraint embedding and attention weights for interpretability, multi-source data fusion and trend assessment are conducted. Graph neural networks are used to capture nonlinear coupling relationships, enabling multi-factor driven analysis and quantification of attribution results.

Benefits of technology

It significantly improves the stability and accuracy of attribution results, can accurately characterize the nonlinear relationships between factors, reduce attribution errors, and provide accurate support for climate change attribution and policy making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122196927A_ABST
    Figure CN122196927A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of atmospheric background analysis, in particular to a heterogeneous data fusion and trend evaluation method for long-term atmospheric background observation, comprising: acquiring multi-source atmospheric background observation data; calibrating the atmospheric background observation data to obtain multi-source calibration data; performing deep spatio-temporal fusion on the multi-source calibration data to obtain fused point data; performing multi-factor driving analysis on the fused point data to output attribution results; performing trend evaluation analysis on the attribution results and the fused point data to obtain trend evaluation results of the atmospheric background. The present application can solve the problem that the existing attribution method lacks a dynamic quantitative mechanism for factor contribution, cannot accurately capture the dominant role of factors in different periods and different regions, and leads to high uncertainty of the attribution results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of atmospheric background analysis technology, specifically to a method for heterogeneous data fusion and trend assessment for long-term atmospheric background observation. Background Technology

[0002] Long-term atmospheric background observation is a core foundation for understanding global climate change and assessing the impact of human activities on the atmospheric environment. Its core objective is to accurately characterize the spatiotemporal distribution characteristics and long-term trends of atmospheric background components (such as CO2, CH4, and aerosols) by integrating multi-source heterogeneous data from ground stations, satellites, reanalysis data, and model outputs. Current mainstream technologies mainly cover three core stages: data preprocessing, heterogeneous data fusion, and trend assessment. In the data preprocessing stage, a quality control process is established, employing methods such as interpolation and resampling to achieve spatiotemporal alignment. In the fusion stage, a technical hierarchy combining statistical methods, machine learning, and deep learning has been formed. In the trend assessment stage, the Mann-Kendall test and Sen's slope estimation are used as core methods, combined with wavelet transform and multiple regression to achieve trend detection and attribution.

[0003] However, with the surge in observational data, the diversification of observation platforms, and the increasing demands for assessment accuracy, existing technologies are gradually revealing their shortcomings in addressing key issues such as systematic biases in multi-source data, the reliability of deep learning fusion, and the accuracy of multi-factor attribution. These problems directly lead to insufficient accuracy of fused data and questionable scientific validity of trend assessment results, making it difficult to meet the core needs of accurate climate change attribution and policy formulation.

[0004] Specifically, this manifests as significant systematic biases across different observation platforms (ground stations, satellites, etc.) due to differences in observation principles, instrument drift, and inconsistent calibration standards. For example, the XCO2 concentration retrieved by satellite can deviate from the direct measurement by ground stations by 1-2 ppm, and differences in calibration standards among different institutions further exacerbate the bias. Simultaneously, the accumulated biases over time create nonlinear drift, making it difficult for traditional single-point and segmented calibration methods to cover the bias characteristics across the entire time series and region. This results in systematic shifts in the fusion results, directly affecting the accuracy of long-term trend assessments; for instance, the CO2 trend slope deviation can exceed 0.5 ppm / year.

[0005] While deep learning models such as CNN-LSTM and Transformer possess powerful capabilities for extracting deep spatiotemporal features, they suffer from a "black box" defect, failing to quantify the contribution weights of each data source to the fusion result and making it difficult to explain the decision-making logic. More importantly, the model training process does not fully incorporate atmospheric physics principles, such as the law of conservation of mass, atmospheric diffusion equations, and radiative transfer equations, easily leading to results that violate physical common sense, such as significantly higher local CO2 concentrations than the surrounding area without supporting emission sources. Furthermore, the dependence of deep learning models on massive amounts of data results in a significant decrease in fusion accuracy in small sample areas such as polar regions and oceans, limiting the technology's global applicability.

[0006] Furthermore, atmospheric background trends are driven by both anthropogenic factors (fossil fuel emissions, vegetation destruction, etc.) and natural factors (solar radiation, volcanic activity, ENSO cycle, etc.), and there are complex nonlinear coupling effects among these factors (e.g., rising temperatures exacerbate methane emissions from wetlands). Existing attribution methods struggle to characterize these nonlinear coupling relationships, achieving only qualitative or semi-quantitative attributions. Simultaneously, the lack of dynamic quantification mechanisms for factor contributions prevents the accurate capture of the dominant roles of each factor in different time periods and regions, leading to high uncertainty in attribution results and hindering support for climate change accountability and policy formulation.

[0007] Therefore, this invention provides a heterogeneous data fusion and trend assessment method for long-term atmospheric background observations to solve the above problems. Summary of the Invention

[0008] In order to overcome the shortcomings of existing technologies, this invention provides a heterogeneous data fusion and trend assessment method for long-term atmospheric background observations. This method addresses the problem that existing attribution methods lack a dynamic quantification mechanism for factor contributions, making it impossible to accurately capture the dominant role of factors in different time periods and regions, resulting in high uncertainty in attribution results.

[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A heterogeneous data fusion and trend assessment method for long-term atmospheric background observation includes: acquiring multi-source atmospheric background observation data; calibrating the atmospheric background observation data to obtain multi-source calibration data; performing deep spatiotemporal fusion of the multi-source calibration data to obtain fused gridded data; performing multi-factor driven analysis on the fused gridded data to output attribution results; and performing trend assessment analysis on the attribution results and the fused gridded data to obtain the trend assessment results of the atmospheric background.

[0010] Preferably, the calibration of atmospheric background observation data to obtain multi-source calibration data includes: processing the atmospheric background observation data using a data standardization mapping method and an association rule mining method to obtain a deviation source database; performing deviation dominant factor analysis and deviation feature analysis on the deviation source database to obtain deviation dominant factors and deviation features; using a dynamic Bayesian calibration model to analyze the deviation dominant factors and deviation features to obtain initial multi-source calibration data; and performing constraint verification on the initial multi-source calibration data to output multi-source calibration data.

[0011] Preferably, the step of performing deviation dominance factor analysis and deviation feature analysis on the deviation source database to obtain deviation dominance factors and deviation features includes: identifying deviation dominance factors in the deviation source database using a random forest feature importance ranking method; extracting temporal features of deviations from the deviation source database using a sliding window method based on the deviation dominance factors; and extracting regional heterogeneity features of deviations in the deviation source database using a spatial interpolation method.

[0012] Preferably, the deep spatiotemporal fusion of multi-source calibration data to obtain fused gridded data includes: spatiotemporal alignment and feature standardization of the multi-source calibration data to obtain standardized feature data; feature enhancement of the standardized feature data based on a preset physical constraint mechanism to obtain an enhanced fused feature set; deep spatiotemporal feature extraction and fusion of the enhanced fused feature set to obtain spatiotemporal fused feature data; quantification of the attention weights of the spatiotemporal fused feature data using the SHAP value decomposition method to generate a heatmap; and verification of the spatiotemporal fused feature data to obtain fused gridded data.

[0013] Preferably, the step of enhancing the standardized feature data based on a preset physical constraint mechanism to obtain an enhanced fusion feature set includes: applying a gradient penalty mechanism to perform hard constraint processing on the standardized feature data based on the mass conservation regularization term to obtain the hard constraint layer output; applying an attention mechanism to perform soft constraint based on the CTM data in the standardized feature data to obtain the soft constraint layer output; and performing dimensionality enhancement on the soft constraint layer output to obtain the enhanced fusion feature set.

[0014] Preferably, the multi-factor driven analysis of the fused gridded data and the output of attribution results includes: obtaining a preliminary core attribution factor set; analyzing the preliminary core attribution factor set and the fused gridded data using a causal discovery algorithm to determine the core attribution factor set; capturing the nonlinear coupling relationship between the core attribution factor set and the fused gridded data using a dynamic coupling attribution model based on a graph neural network to obtain initial multi-factor attribution results; and performing contribution quantification processing on the initial multi-factor attribution results using a Bayesian inference method to generate attribution results.

[0015] Preferably, the step of using a causal discovery algorithm to analyze the initial core attribution factor set and the fused gridded data to determine the core attribution factor set includes: constructing a factor causal relationship graph between the initial core attribution factor set and the fused gridded data using a causal discovery algorithm to identify confounding factors and mediating factors; using a conditional independence test to eliminate spurious correlations between factors and determine direct causal paths; and screening the initial core attribution factor set based on the direct causal paths to obtain the core attribution factor set.

[0016] Preferably, the dynamic coupling attribution model based on graph neural networks captures the nonlinear coupling relationship between the core attribution factor set and the fused grid data to obtain initial multi-factor attribution results, including: constructing a spatiotemporal heterogeneous graph by taking observation stations as nodes and spatiotemporal correlations between stations as edges; using graph convolutional layers to capture the nonlinear coupling relationship between factors to obtain nonlinear coupling relationship data; dynamically adjusting the edge weights of the spatiotemporal heterogeneous graph based on a temporal attention mechanism; and analyzing the nonlinear coupling relationship data, the edge weights of the spatiotemporal heterogeneous graph, and the fused grid data based on a preset factor coupling function to obtain initial multi-factor attribution results.

[0017] Preferably, the trend assessment analysis of the attribution results and fused grid data to obtain the trend assessment result of the atmospheric background includes: performing improved nonparametric trend detection and abrupt change point identification on the fused grid data to determine the abrupt change point; performing multi-scale and multi-dimensional trend characterization on the attribution results and fused grid data to obtain the characterization result; and performing triple verification on the characterization result to obtain the trend assessment result of the atmospheric background.

[0018] Preferably, the triple verification of the characterization results to obtain the trend assessment results of the atmospheric background includes: internal verification using a time-stratified cross-validation method to verify the consistency of the trend assessment results in different time periods; external verification to verify the absolute accuracy of the assessment results; and physical consistency verification to verify the matching of the trend results with the attribution logic.

[0019] The beneficial effects of this invention are as follows: 1. The causal inference of this invention can effectively eliminate spurious correlations and significantly improve the stability of attribution results; dynamic coupling modeling can accurately characterize the nonlinear relationship between factors and reduce attribution error; and it can separate the direct and indirect contributions of multiple factors. The three-dimensional attribution report can clearly identify the dominant factors in different time periods and regions, providing accurate support for climate change attribution and policy making. It helps to solve the problem that existing attribution methods lack a dynamic quantification mechanism for factor contributions, cannot accurately capture the dominant role of each factor in different time periods and regions, and thus lead to high uncertainty in attribution results.

[0020] 2. This invention employs a fusion framework combining physical constraint embedding, interpretable attention weights, and few-sample adaptive methods to address the bottleneck issues of traditional deep learning's "black box" nature and lack of physicality. Specifically, the physical embedding mechanism completely avoids fusion results that violate atmospheric physics laws, verifies the physical consistency of the fused data, and the attention weights and SHAP values ​​clearly quantify the contribution proportion of each data source, thus resolving the "black box" problem in deep learning.

[0021] 3. This invention uses a method based on dynamic Bayesian calibration and cross-platform consistency constraints to achieve systematic bias calibration of multi-source data, which significantly reduces the systematic bias of multi-source data. Attached Figure Description

[0022] Figure 1 This is a schematic flowchart of a heterogeneous data fusion and trend assessment method for long-term atmospheric background observation according to the present invention. Detailed Implementation

[0023] The following will refer to the attached reference. Figure 1 The various embodiments of the present invention will be described in detail below. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0024] A heterogeneous data fusion and trend assessment method for long-term atmospheric background observations, as shown in the appendix. Figure 1 As shown, it includes the following steps: Step S11: Acquire multi-source atmospheric background observation data; multi-source atmospheric background observation data includes, but is not limited to, ground station observation data (CO2 / CH4 / aerosol concentration time series, instrument parameters, calibration history, etc.), satellite inversion data (XCO2 isoconcentration grid data, satellite observation geometric parameters, inversion algorithm parameters), reanalysis data (global atmospheric composition grid data, meteorological driving field data), metadata (observation station altitude / temperature and humidity, satellite transit time, reanalysis data assimilation scheme), and global benchmark station ground truth data (long-term observation data from Wariguan and Mauna Loa stations).

[0025] Step S12: Calibrate the atmospheric background observation data to obtain multi-source calibration data; complete the deviation tracing and accurate correction of the original multi-source data to provide basic data with spatiotemporal consistency and accuracy for subsequent fusion.

[0026] Step S13: Perform deep spatiotemporal fusion on the multi-source calibration data to obtain fused gridded data; after completing the deep spatiotemporal fusion of the multi-source data after calibration, it breaks through the traditional fusion "black box" and physical missing problems, and outputs high-precision, interpretable fused gridded data that conforms to the laws of atmospheric physics.

[0027] Step S14: Perform multi-factor driven analysis on the fused gridded data and output the attribution results; that is, complete the multi-factor driven analysis of the fused data, accurately identify the anthropogenic / natural driving factors of the atmospheric background trend, quantify the nonlinear coupling contribution of each factor, and output the three-dimensional attribution results of "factor-time period-region".

[0028] Step S15: Perform trend assessment analysis on the attribution results and fused grid data to obtain the trend assessment results of atmospheric background; that is, based on the attributed purification fused data, complete the accurate detection, characterization and verification of the long-term trend of atmospheric background components, and output multi-scale, multi-dimensional trend assessment results and credibility reports.

[0029] The causal inference of this invention can effectively eliminate spurious correlations and significantly improve the stability of attribution results; dynamic coupling modeling can accurately characterize the nonlinear relationship between factors and reduce attribution error; and it can separate the direct and indirect contributions of multiple factors. The three-dimensional attribution report can clearly identify the dominant factors in different time periods and regions, providing precise support for climate change attribution and policy making. It helps to solve the problem that existing attribution methods lack a dynamic quantification mechanism for factor contributions, cannot accurately capture the dominant role of each factor in different time periods and regions, and thus lead to high uncertainty in attribution results.

[0030] This invention employs a fusion framework combining physically constrained embedding, interpretable attention weights, and few-sample adaptive methods to address the bottlenecks of traditional deep learning's "black box" nature and lack of physicality. Specifically, the physically embedded mechanism completely avoids fusion results that violate atmospheric physics laws, verifying the physical consistency of the fused data; attention weights and SHAP values ​​clearly quantify the contribution proportion of each data source, resolving the "black box" problem of deep learning.

[0031] This invention employs a method based on dynamic Bayesian calibration and cross-platform consistency constraints to calibrate the systematic bias of multi-source data, significantly reducing the systematic bias of multi-source data.

[0032] In one embodiment of the present invention, the calibration of atmospheric background observation data to obtain multi-source calibration data includes: processing the atmospheric background observation data using a data standardization mapping method and an association rule mining method to obtain a deviation source database; performing deviation dominant factor analysis and deviation feature analysis on the deviation source database to obtain deviation dominant factors and deviation features; using a dynamic Bayesian calibration model to analyze the deviation dominant factors and deviation features to obtain initial multi-source calibration data; and performing constraint verification on the initial multi-source calibration data to output multi-source calibration data.

[0033] Specifically, a data standardization mapping method is adopted to unify the format, units, and spatiotemporal resolution of different data sources; observation data and metadata are integrated through association rule mining to establish a deviation traceability database of "observation value-instrument parameter-environmental factor-calibration history"; database cleaning is completed based on missing value imputation and outlier removal to provide complete and related basic data for the identification of deviation-dominant factors and eliminate the disconnect between metadata and observation data.

[0034] Furthermore, deviation source databases are analyzed for dominant deviation factors and characteristics to obtain these factors and characteristics. A dynamic Bayesian calibration model is then used to analyze these factors and characteristics, yielding initial multi-source calibration data. Specifically, the true value of the reference station is used as the dependent variable, and the dominant deviation factors are used as covariates to construct a dynamic Bayesian calibration model. Gaussian process priors are used to characterize the spatiotemporal variation of the deviation. Polynomial basis functions are introduced to extend the model's nonlinear fitting capability and adapt to the nonlinear drift of the deviation. The model parameters are updated in real time using a Markov chain Monte Carlo (MCMC) algorithm (e.g., ≥1000 iterations, convergence threshold ≤1e-6) to perform grid-by-grid and time-series dynamic deviation correction on satellite and reanalysis data. Ground station data is corrected using single-point calibration to eliminate instrument drift errors, achieving accurate deviation correction across the entire time series and region, thus solving the problem that traditional calibration methods cannot cover nonlinear drift.

[0035] Constraint verification is performed on the initial multi-source calibration data to output multi-source calibration data. Specifically, this includes: constructing a cross-platform consistency loss function, for example, using "daily average deviation of different data sources ≤ 0.3ppm and annual trend slope deviation ≤ 0.1ppm / year" as hard constraints, embedding the constraints into the model optimization objective using the Lagrange multiplier method, and performing secondary optimization on the calibrated data; using three indicators, root mean square error (RMSE), trend slope deviation, and spatiotemporal correlation coefficient, to conduct cross-platform consistency verification, eliminating abnormal data segments that do not meet the constraints, ensuring that the spatiotemporal distribution characteristics of each data source are consistent after calibration, and avoiding systematic shifts during the fusion process.

[0036] In one embodiment of the present invention, the step of performing deviation dominance factor analysis and deviation feature analysis on the deviation source database to obtain deviation dominance factors and deviation features includes: identifying deviation dominance factors in the deviation source database using a random forest feature importance ranking method; extracting temporal features of deviations from the deviation source database using a sliding window method based on the deviation dominance factors; and extracting regional heterogeneity features of deviations in the deviation source database using a spatial interpolation method.

[0037] Specifically, using the calibrated base station data as a reference, random forest feature importance ranking is used to identify the dominant bias factors (instrument age, observation altitude, temperature and humidity, cloud cover retrieved from satellite, etc.); the temporal features (linear / nonlinear drift) of the bias are extracted using the sliding window method (e.g., window size of 1 month, step size of 1 week); spatial interpolation (ordinary kriging) is used to extract the regional heterogeneity features of the bias; and the features are normalized (Min-Max) to generate a standardized feature set.

[0038] Through the configuration method of this embodiment, the present invention can clearly identify the core driving factors of deviation formation, extract the spatiotemporal distribution law of deviation, and provide prior feature information for dynamic Bayesian calibration model.

[0039] In one embodiment of the present invention, the deep spatiotemporal fusion of multi-source calibration data to obtain fused gridded data includes: spatiotemporal alignment and feature standardization of the multi-source calibration data to obtain standardized feature data; feature enhancement of the standardized feature data based on a preset physical constraint mechanism to obtain an enhanced fused feature set; deep spatiotemporal feature extraction and fusion of the enhanced fused feature set to obtain spatiotemporal fused feature data; quantification of the attention weights of the spatiotemporal fused feature data using the SHAP value decomposition method to generate a heatmap; and verification of the spatiotemporal fused feature data to obtain fused gridded data.

[0040] Specifically, resampling (e.g., linear interpolation) is used to unify the temporal resolution (daily / monthly scale) of all input data; spatial coordinate system unification is achieved through projection transformation (e.g., WGS84 to UTM); spatial interpolation (e.g., co-kriging method) is used to complete the spatial grid of discrete ground station data; Z-score standardization is performed on all feature data to eliminate dimensional differences; a set of standardized feature data consisting of "spatiotemporal points - concentration values ​​- physical features - CTM output values" is constructed to achieve spatiotemporal dimension unification of multi-source data, providing standardized features for the input of deep learning models and eliminating the impact of spatiotemporal misalignment on the fusion results.

[0041] Based on a preset physical constraint mechanism, the standardized feature data is enhanced to obtain an enhanced fusion feature set, which avoids outputting results that violate the common sense of atmospheric physics and improves the physical rationality of the fusion data.

[0042] Furthermore, deep spatiotemporal feature extraction and fusion are performed on the enhanced fusion feature set to obtain fused gridded data. This involves constructing a deep learning model with a data source attention layer, a spatial attention layer, and a temporal attention layer, quantifying the contribution weights of different data sources, spatial regions, and temporal stages to the fusion result. Using calibrated multi-source data as the training set, the model is trained using mini-batch stochastic gradient descent (SGD) with an early stopping mechanism to avoid overfitting. For small sample regions such as polar / oceanic areas, transfer learning (pre-training-fine-tuning) and meta-learning (MAML algorithm) are introduced to transfer the general spatiotemporal features pre-trained in the land region to the target domain. A small number of samples are used to quickly fine-tune the model parameters, achieving deep spatiotemporal feature extraction and fusion of multi-source data. This also solves the problem of low fusion accuracy in small sample regions, making it applicable to the entire domain.

[0043] Three layers of attention weights are extracted to generate a spatiotemporal distribution map of the contribution weights of each data source, clarifying the contribution ratio of ground station / satellite / reanalysis data in different regions / time periods. The SHAP value decomposition method is used to quantify the feature contribution of the model output, generating a SHAP value heatmap to clarify the degree of influence of each input feature on the fusion result. A data source contribution ratio report is output, highlighting the core contribution of ground station data in near-ground concentration fusion, in order to solve the "black box" problem of deep learning, realize the interpretability quantification of the fusion result, and improve the scientific credibility of the result.

[0044] Physical consistency is verified using atmospheric physical laws (such as the matching of concentration gradients with meteorological fields and the rationality of concentrations around emission sources), with a pass rate of ≥98%. Ten-fold cross-validation is used to calculate the accuracy indicators of the fusion results, such as RMSE and R², ensuring R² ≥ 0.95 and RMSE reduced by 40%-50% compared to traditional methods. For small sample areas, the accuracy of the fusion results is verified by benchmark station interpolation comparison, ensuring R² ≥ 0.85, guaranteeing high accuracy and physical rationality of the fused data, and providing high-quality basic data for subsequent trend attribution. The output includes high-precision fused gridded data of atmospheric background components, spatiotemporal distribution maps of contribution weights from each data source, heatmaps of SHAP value feature contribution, fused complete data for small sample areas (polar / oceanic), a physical consistency verification report of the fusion results, and a deep learning model parameter library.

[0045] In one embodiment of the present invention, the step of enhancing the standardized feature data based on a preset physical constraint mechanism to obtain an enhanced fusion feature set includes: applying a gradient penalty mechanism to perform hard constraint processing on the standardized feature data based on a mass conservation regularization term to obtain the hard constraint layer output; applying an attention mechanism to perform soft constraint based on the CTM data in the standardized feature data to obtain the soft constraint layer output; and performing dimensionality enhancement on the soft constraint layer output to obtain the enhanced fusion feature set.

[0046] Specifically, atmospheric physical laws are embedded from both hard and soft constraints. In the hard constraint layer, a mass conservation regularization term is introduced, and a gradient penalty mechanism ensures that the output of the hard constraint layer satisfies the "balance between the regional concentration change rate and the surrounding diffusion flux". In the soft constraint layer, CTM data from the standardized feature data is used as auxiliary features, and an attention mechanism guides the learning of spatiotemporal distribution features dominated by physical laws. The feature set is dimensionality-enhanced by fusing spatiotemporal features, physical features, and model simulation features to generate a physically enhanced fused feature set. This avoids the output of this method from violating the common sense of atmospheric physics and improves the physical rationality of the fused data.

[0047] In one embodiment of the present invention, the multi-factor driven analysis of fused grid data and the output of attribution results include: obtaining a preliminary core attribution factor set; analyzing the preliminary core attribution factor set and the fused grid data using a causal discovery algorithm to determine the core attribution factor set; capturing the nonlinear coupling relationship between the core attribution factor set and the fused grid data using a dynamic coupling attribution model based on a graph neural network to obtain initial multi-factor attribution results; and performing contribution quantification processing on the initial multi-factor attribution results using a Bayesian inference method to generate attribution results.

[0048] Specifically, based on knowledge in the field of atmospheric science, core attribution factors are initially screened, and redundant factors that are not significantly related to the trend of atmospheric background are removed. Spatiotemporal matching of factor data is performed to unify the spatiotemporal resolution and coordinate system of factor grid data and fused data. Factor data cleaning is completed by standardization (Min-Max) and outlier removal (box plot method) to generate a standardized set of core attribution factors for initial screening. This reduces the interference of invalid factors on the attribution results, ensures the spatiotemporal consistency of factor data and fused data, and provides a high-quality set of core attribution factors for initial screening for causal relationship identification.

[0049] A causal discovery algorithm was used to analyze the initial core attribution factor set and the fused grid data to determine the core attribution factor set. A dynamic coupling attribution model based on graph neural networks was used to capture the nonlinear coupling relationship between the core attribution factor set and the fused grid data to obtain the initial multi-factor attribution results.

[0050] Furthermore, Bayesian inference methods were employed to quantify the contributions of multi-factor attribution results, generating attribution outcomes. Specifically, Bayesian inference was used to quantify the contribution distribution of each factor, outputting the contribution percentage for different confidence intervals (95% CI). A Bootstrap resampling method (≥500 resampling times) was introduced to assess the stability of the attribution results, calculating the coefficient of variation of contributions. Marginal effect analysis was used to separate the direct and indirect contributions of each factor (e.g., the indirect contribution of temperature to CO2 through its influence on vegetation respiration), quantifying the contribution percentage of coupling effects. The contributions were stratified and statistically analyzed by climate zone / land-sea region and interannual / decadal time period, generating three-dimensional attribution data. This achieved accurate quantification and uncertainty assessment of multi-factor contributions, clarified the impact of coupling effects, and provided factor-driven basis for subsequent trend assessment.

[0051] The causal consistency test was used to verify the matching between the attribution results and atmospheric science theories; the accuracy of the attribution model was evaluated by leave-one-out cross-validation to ensure that the attribution error was ≤10%; the three-dimensional attribution data was visualized to generate a three-dimensional attribution report of "factor-time period-region" to identify the dominant driving factors in different scenarios, ensuring the reliability and accuracy of the attribution results, and providing a clear factor-driven basis for trend assessment and policy making.

[0052] Furthermore, the method of using a causal discovery algorithm to analyze the initial core attribution factor set and the fused gridded data to determine the core attribution factor set includes: constructing a causal relationship graph between the initial core attribution factor set and the fused gridded data using a causal discovery algorithm to identify confounding factors and mediating factors; using a conditional independence test to eliminate spurious correlations between factors and determine direct causal paths; and screening the initial core attribution factor set based on the direct causal paths to obtain the core attribution factor set.

[0053] Specifically, a causal discovery algorithm, namely the PC algorithm, is employed. Using the atmospheric background concentration from the fused data as the target variable and the standardized attribution factor set as the explanatory variable, a causal relationship graph between factors is constructed. Confounding factors (such as temperature simultaneously affecting vegetation respiration and methane emissions) and mediating factors in the graph are identified. Conditional independence tests are used to eliminate spurious correlations between factors, clarifying the direct causal path between each factor and the atmospheric background trend. Factors with insignificant causal paths (p>0.05) are removed, generating the final core attribution factor set to avoid attribution bias caused by spurious correlations, ensuring the causal relationship between attribution factors and the atmospheric background trend, and improving the scientific rigor of the attribution results.

[0054] Furthermore, the dynamic coupling attribution model based on graph neural networks captures the nonlinear coupling relationship between the core attribution factor set and the fused grid data to obtain initial multi-factor attribution results. This includes: constructing a spatiotemporal heterogeneous graph by using observation stations as nodes and spatiotemporal correlations between stations as edges; using graph convolutional layers to capture the nonlinear coupling relationship between factors to obtain nonlinear coupling relationship data; dynamically adjusting the edge weights of the spatiotemporal heterogeneous graph based on a temporal attention mechanism; and analyzing the nonlinear coupling relationship data, the edge weights of the spatiotemporal heterogeneous graph, and the fused grid data based on a preset factor coupling function to obtain initial multi-factor attribution results.

[0055] Specifically, a dynamic coupling attribution model based on graph neural networks (GNNs) is constructed, using observation stations as nodes and spatiotemporal relationships between stations as edges to build a spatiotemporal heterogeneous graph. The graph convolutional layers of the GNN capture the nonlinear coupling relationships between factors. A temporal attention mechanism is introduced into the dynamic coupling attribution model to dynamically adjust the weights according to the driving strength of factors in different time periods, achieving dynamic attribution in three dimensions: time series, space, and factors. Customized factor coupling functions are designed for the characteristics of different atmospheric background components such as CO2, CH4, and aerosols (e.g., CO2 focuses on the coupling between emissions and diffusion, while aerosols focus on the coupling between source emissions and meteorological conditions). Using fused gridded data as the dependent variable and the core attribution factor set as the independent variable, the Adam optimizer is used to train the model, and a loss function (mean squared error MSE) is set to optimize the model accuracy, accurately characterizing the nonlinear coupling effect between factors. This breaks through the limitations of traditional linear attribution methods and achieves dynamic and accurate multi-factor attribution.

[0056] In one embodiment of the present invention, the trend assessment analysis of the attribution results and fused grid data to obtain the trend assessment result of the atmospheric background includes: performing improved nonparametric trend detection and abrupt change point identification on the fused grid data to determine the abrupt change point; performing multi-scale, multi-dimensional trend characterization on the attribution results and fused grid data to obtain the characterization result; and performing triple verification on the characterization result to obtain the trend assessment result of the atmospheric background.

[0057] Specifically, based on the multi-factor attribution results, abnormal data segments corresponding to extreme driving factors (such as sudden volcanic eruptions, super ENSO events, etc.) are removed to avoid interference from extreme events with long-term trend signals. Wavelet low-pass filtering (filter scale ≥ 5 years) or 5-year moving average method is used to perform time-series decomposition on the purified fused data, separating short-term fluctuations (seasonal fluctuations, random noise) from long-term trend components (interannual, interdecadal), retaining core trend signals, eliminating interference from extreme events and short-term fluctuations on trend assessment, extracting pure long-term trend signals, and improving the stability of trend assessment.

[0058] An improved nonparametric trend detection and abrupt change point identification method is used for fused gridded data to determine abrupt change points. This mainly includes: employing a modified Mann-Kendall (MK) test with pre-whitening to eliminate the interference of autocorrelation in time series data on trend significance judgment; calculating the Z-statistic and p-value (p < 0.05 indicates a significant trend); combining this with Sen's slope estimation to accurately quantify the trend magnitude, such as the annual growth slope of CO2 concentration; for long-term time series data, a segmented MK test (e.g., a sliding window of 10 years) is used to locate trend abrupt change points, and the timing of these abrupt change points is cross-referenced with the factor contribution changes in the attribution results (e.g., a surge in industrial emissions, the implementation of major environmental protection policies) to eliminate misjudged abrupt change points, thereby achieving significant detection and quantitative characterization of long-term trends, accurately identifying trend abrupt change points, and improving the accuracy of trend detection.

[0059] Multi-scale and multi-dimensional trend characterization was performed on the attribution results and fused gridded data to obtain the characterization results. Specifically, ensemble empirical mode decomposition (EEMD) was used to decompose the time series data into seasonal, interannual, and interdecadal scale components. MK test and Sen's slope estimation were performed on each scale component to assess the trend characteristics at each scale. Combining the multi-factor three-dimensional attribution results, the dominant driving factors of trends at different scales were identified, and a three-dimensional correlation characterization system between "scale-trend-factor" was constructed. Regional trend assessments were conducted according to climate zones / land-sea divisions, and quantile trend assessments were conducted according to the 5% extreme low value, 50% mean, and 95% extreme high value. The trend slope and significance differences under different regions and extreme conditions were statistically analyzed to comprehensively capture the multi-scale and multi-dimensional characteristics of atmospheric background trends, clarify the correlation between trends and driving factors, and achieve refined trend characterization.

[0060] The characterization results undergo triple verification to obtain the trend assessment results of the atmospheric background, including: internal verification using a time-stratified cross-validation method to verify the consistency of trend assessment results across different time periods; external verification to verify the absolute accuracy of the assessment results; and physical consistency verification to verify the matching between the trend results and the attribution logic. In other words, triple verification is conducted: ① Internal verification uses time-stratified cross-validation (dividing the training set / validation set by time period) to ensure the consistency of trend assessment results across different time periods; ② External verification compares the assessment results with long-term observed trends from global atmospheric background benchmark stations (Wariguan, Mauna Loa, etc.), calculating trend slope deviation and correlation coefficients to verify the absolute accuracy of the assessment results; ③ Physical consistency verification checks the matching between trend results and attribution logic (e.g., a positive correlation between the CO2 trend slope and the contribution of anthropogenic emissions), avoiding physical contradictions and comprehensively ensuring the scientific validity, accuracy, and credibility of the trend assessment results, providing a reliable basis for accurate attribution of climate change.

[0061] The multi-scale, regional, and quantile trend results, abrupt change point identification results, and "scale-trend-factor" correlations are visualized to generate a long-term trend map of atmospheric background components. All trend detection, characterization, and verification results are integrated to generate a final trend assessment report, clarifying core trend parameters, abrupt change characteristics, dominant driving factors, and the credibility of the results. This forms a standardized and visualized trend assessment result, providing intuitive and accurate decision support for climate change attribution and policy formulation.

[0062] Based on the attributed purification fusion data, the system accurately detects, characterizes, and validates the long-term trends of atmospheric background components, outputting multi-scale, multi-dimensional trend assessment results and credibility reports. In one embodiment of the present invention, the main implementation idea is as follows: First, a two-dimensional calibration system of "dynamic Bayesian calibration model + cross-platform consistency constraint" is constructed to achieve accurate deviation correction across all time series and all regions. Specifically, as follows: First, a multi-source data deviation traceability database is constructed, integrating metadata such as instrument parameters, observation environment, and calibration history. A random forest model is used to identify the dominant factors of deviation (such as instrument age, observation altitude, temperature, and humidity). The temporal characteristics (linear / nonlinear drift) and spatial characteristics (regional heterogeneity) of the deviation are extracted using the sliding window method to provide prior information for the calibration model.

[0063] Secondly, a dynamic Bayesian calibration model is constructed using baseline station observation data as the true value. The bias-dominant factor is used as a covariate, and a Gaussian process prior is employed to characterize the spatiotemporal variation of the bias. The model parameters are updated in real time using a Markov chain Monte Carlo algorithm to achieve dynamic bias correction for different data sources (satellite and reanalysis data). To address nonlinear drift, a polynomial basis function is introduced to extend the model's nonlinear fitting capability, ensuring calibration accuracy covers the entire time series.

[0064] Furthermore, a cross-platform data consistency loss function is established, incorporating the absolute value of the deviation and the difference in trend slope of data from different observation platforms as constraints into the calibration model. For example, for satellite and ground station data, the daily average deviation is constrained to be ≤0.3ppm and the annual trend slope deviation is constrained to be ≤0.1ppm / year after calibration. The constraints are embedded into the model optimization objective using the Lagrange multiplier method to ensure that the multi-source data has spatiotemporal consistency after calibration.

[0065] The second step involves constructing a fusion framework of "physical constraint embedding + interpretable attention weights + few-sample adaptation" to enhance the scientific validity and credibility of the fusion results, overcoming the bottlenecks of traditional deep learning's "black box" nature and lack of physicality. Specifically: Atmospheric dynamics and chemical laws are transformed into hard and soft constraints for the model. At the hard constraint level, a mass conservation regularization term is introduced into the model output layer (e.g., the rate of change of CO2 concentration in a certain region must be balanced with the surrounding diffusion flux), and a gradient penalty mechanism is used to ensure that the output results satisfy physical laws. At the soft constraint level, the output of the atmospheric chemical transport model is used as an auxiliary feature input into the deep learning model, and an attention mechanism is used to guide the model to learn the spatiotemporal distribution characteristics dominated by physical laws, thereby improving the physical rationality of the fusion results.

[0066] A "multi-layer attention weight visualization + contribution quantification" approach is adopted. The model structure incorporates a data source attention layer, a spatial attention layer, and a temporal attention layer to quantify the contribution weights of different data sources, spatial regions, and temporal stages to the fusion result. The model output is decomposed using SHAP values ​​to generate a heatmap of the contribution of each input feature, clarifying the decision-making logic of the fusion result. For key atmospheric background components (such as CO2), a report on the contribution percentage of the data source is output to improve the scientific interpretability of the results.

[0067] For adaptive optimization in small sample regions, transfer learning and meta-learning mechanisms are introduced to address the problem of insufficient data in small sample regions. Using land regions with abundant data as the source domain and small sample polar / ocean regions as the target domain, general spatiotemporal features learned from the source domain are transferred to the target domain through pre-training and fine-tuning. The meta-learning algorithm (MAML) is used to train the model's rapid adaptability, and the model parameters are rapidly fine-tuned using a small amount of target domain data, ensuring that the fusion accuracy of small sample regions is R²≥0.85.

[0068] The third step is to construct a full-chain attribution system that integrates "causal relationship identification, dynamic coupling modeling, and contribution quantification" to achieve accurate quantification of multi-factor nonlinear contributions and break through the linear limitations of traditional attribution methods.

[0069] First, core attribution factors are selected based on domain knowledge (anthropogenic factors: fossil fuel emissions, industrial activities, vegetation NPP; natural factors: solar radiation, volcanic aerosols, ENSO index, sea level pressure). A causal discovery algorithm (PC algorithm) is used to construct a causal graph between factors, identify potential confounding factors (such as temperature affecting both vegetation respiration and methane emissions) and mediating factors, eliminate spurious correlations, and clarify the direct causal path between each factor and the atmospheric background trend.

[0070] Secondly, a dynamic coupling attribution model based on graph neural networks (GNNs) is constructed. A spatiotemporal heterogeneous graph (nodes representing observation stations and edges representing spatiotemporal correlations) is built using attribution factors and atmospheric background concentration data. The nonlinear coupling relationships between factors are captured through GNNs. A temporal attention mechanism is introduced to dynamically adjust the weights of each factor at different time periods, achieving dynamic attribution across the three dimensions of "time-space-factor". Factor coupling functions are designed to address the characteristics of different atmospheric background components (such as CO2 and aerosols). For example, CO2 attribution focuses on the coupling between emissions and diffusion, while aerosol attribution focuses on the coupling between source emissions and meteorological conditions.

[0071] Furthermore, Bayesian inference is used to quantify the contribution distribution of each factor, outputting the contribution percentage for different confidence intervals (95% CI); the Bootstrap resampling method is introduced to assess the stability of the attribution results; for coupling effects, the direct and indirect contributions of each factor are separated through marginal effect analysis (such as the indirect contribution of temperature to CO2 by affecting vegetation respiration), and finally a three-dimensional attribution report of "factor-time period-region" is generated to clarify the dominant factors in different scenarios.

[0072] The fourth step involves constructing a "multi-scale, multi-dimensional, and strongly validated" trend assessment system based on the purification fusion data after multi-factor dynamic attribution. This system aims to accurately characterize and reliably guarantee long-term atmospheric background trends, and connects the process with the output of attribution results. The specific steps are as follows: First, abnormal data segments corresponding to extreme driving factors (such as sudden volcanic eruptions and super ENSO events) identified during the attribution process are removed to avoid interference from extreme events with long-term trend signals. Wavelet low-pass filtering or 5-year moving average is used to separate short-term fluctuations (seasonal fluctuations and random noise) from long-term trend components, retaining core trend signals at the decadal and interannual scales, and improving the stability of trend assessment.

[0073] A modified Mann-Kendall (MK) test with pre-whitened data was used to eliminate the interference of time series data autocorrelation on the judgment of trend significance. The significance of the trend was determined by the Z statistic and p-value (p<0.05 is significant). Sen's slope estimation was used to accurately quantify the magnitude of the trend (e.g., the annual growth slope of CO2 concentration). For long time series data, a segmented MK test was used to locate trend abrupt change points, and these were corroborated with the factor contribution abrupt changes in the attribution results (e.g., a surge in industrial emissions or the implementation of major environmental protection policies in a certain period) to reduce the misjudgment rate of abrupt change points.

[0074] The time-series data is decomposed into seasonal, interannual, and interdecadal scale components using ensemble empirical mode decomposition (EEMD) to assess the trend characteristics at each scale. The dominant driving factors of trends at different scales are identified by combining the attribution results (e.g., fossil fuel emissions are the dominant factor of interdecadal CO2 trend, and ENSO cycle is the dominant factor of interannual trend), forming a three-dimensional correlation characterization in the form of "scale-trend-factor". At the same time, regional trend assessments (divided by climate zone / land-sea) and quantile trend assessments (e.g., 5% extreme low value, 50% mean, and 95% extreme high value) are carried out to comprehensively capture the trend differences in different regions and under different extreme conditions.

[0075] During the validation process, internal validation employs time-stratified cross-validation to ensure consistency of trend assessment results across different time periods; external validation compares the results with long-term observation trends from global atmospheric baseline stations (such as Wariguan and Mauna Loa) to verify the absolute accuracy of the assessment results; and physical consistency validation checks the matching between trend results and attribution logic (e.g., the slope of the CO2 trend is positively correlated with the contribution of anthropogenic emission factors) to avoid physical contradictions.

[0076] The causal inference of this invention can effectively eliminate spurious correlations and significantly improve the stability of attribution results; dynamic coupling modeling can accurately characterize the nonlinear relationship between factors and reduce attribution error; and it can separate the direct and indirect contributions of multiple factors. The three-dimensional attribution report can clearly identify the dominant factors in different time periods and regions, providing precise support for climate change attribution and policy making. It helps to solve the problem that existing attribution methods lack a dynamic quantification mechanism for factor contributions, cannot accurately capture the dominant role of each factor in different time periods and regions, and thus lead to high uncertainty in attribution results.

[0077] This invention employs a fusion framework combining physically constrained embedding, interpretable attention weights, and few-sample adaptive methods to address the bottlenecks of traditional deep learning's "black box" nature and lack of physical consistency. The physically embedded mechanism completely avoids fusion results that violate atmospheric physics laws, verifying the physical consistency of the fused data. Attention weights and SHAP values ​​clearly quantify the contribution proportion of each data source, resolving the "black box" problem of deep learning. Furthermore, a method based on dynamic Bayesian calibration and cross-platform consistency constraints is used to calibrate systematic biases in multi-source data, significantly reducing systematic biases.

[0078] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0079] It should be noted that in the description of this invention, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0080] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0081] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0082] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0083] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0084] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0085] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A method for heterogeneous data fusion and trend assessment based on long-term atmospheric background observations, characterized in that, include: Acquire multi-source atmospheric background observation data; The atmospheric background observation data were calibrated to obtain multi-source calibration data; Deep spatiotemporal fusion of multi-source calibration data is performed to obtain fused gridded data; Perform multi-factor driven analysis on the fused grid data and output the attribution results; Trend assessment analysis was performed on the attribution results and fused gridded data to obtain the trend assessment results of the atmospheric background.

2. The heterogeneous data fusion and trend assessment method according to claim 1, characterized in that, The calibration of atmospheric background observation data to obtain multi-source calibration data includes: processing atmospheric background observation data using data standardization mapping and association rule mining methods to obtain a deviation source database; performing deviation dominant factor analysis and deviation feature analysis on the deviation source database to obtain deviation dominant factors and deviation features; using a dynamic Bayesian calibration model to analyze the deviation dominant factors and deviation features to obtain initial multi-source calibration data; and performing constraint verification on the initial multi-source calibration data to output multi-source calibration data.

3. The heterogeneous data fusion and trend assessment method according to claim 2, characterized in that, The aforementioned deviation source tracing database undergoes deviation dominant factor analysis and deviation feature analysis to obtain deviation dominant factors and deviation features, including: identifying deviation dominant factors in the deviation source tracing database using a random forest feature importance ranking method; extracting temporal features of deviations from the deviation source tracing database based on the deviation dominant factors using a sliding window method; and extracting regional heterogeneity features of deviations in the deviation source tracing database using a spatial interpolation method.

4. The heterogeneous data fusion and trend assessment method according to claim 1, characterized in that, The method of performing deep spatiotemporal fusion of multi-source calibration data to obtain fused gridded data includes: performing spatiotemporal alignment and feature standardization on the multi-source calibration data to obtain standardized feature data; performing feature enhancement on the standardized feature data based on a preset physical constraint mechanism to obtain an enhanced fused feature set; performing deep spatiotemporal feature extraction and fusion on the enhanced fused feature set to obtain spatiotemporal fused feature data; quantifying the attention weights of the spatiotemporal fused feature data using the SHAP value decomposition method to generate a heatmap; and verifying the spatiotemporal fused feature data to obtain fused gridded data.

5. The heterogeneous data fusion and trend assessment method according to claim 4, characterized in that, The method of enhancing the standardized feature data based on a preset physical constraint mechanism to obtain an enhanced fusion feature set includes: applying a gradient penalty mechanism to the standardized feature data based on the mass conservation regularization term to obtain the hard constraint layer output; applying an attention mechanism to the CTM data in the standardized feature data to obtain the soft constraint layer output; and increasing the dimensionality of the soft constraint layer output to obtain the enhanced fusion feature set.

6. The heterogeneous data fusion and trend assessment method according to claim 1, characterized in that, The multi-factor driven analysis of fused gridded data and the output of attribution results include: obtaining a preliminary core attribution factor set; analyzing the preliminary core attribution factor set and fused gridded data using a causal discovery algorithm to determine the core attribution factor set; capturing the nonlinear coupling relationship between the core attribution factor set and fused gridded data using a dynamic coupling attribution model based on a graph neural network to obtain initial multi-factor attribution results; and performing contribution quantification processing on the initial multi-factor attribution results using a Bayesian inference method to generate attribution results.

7. The heterogeneous data fusion and trend assessment method according to claim 6, characterized in that, The method of using a causal discovery algorithm to analyze the initial core attribution factor set and the fused gridded data to determine the core attribution factor set includes: constructing a causal relationship graph between the initial core attribution factor set and the fused gridded data using a causal discovery algorithm to identify confounding factors and mediating factors; using a conditional independence test to eliminate spurious correlations between factors and determine direct causal paths; and filtering the initial core attribution factor set based on the direct causal paths to obtain the core attribution factor set.

8. The heterogeneous data fusion and trend assessment method according to claim 6, characterized in that, The aforementioned dynamic coupling attribution model based on graph neural networks captures the nonlinear coupling relationship between the core attribution factor set and the fused grid data to obtain initial multi-factor attribution results. This includes: constructing a spatiotemporal heterogeneous graph by using observation stations as nodes and spatiotemporal correlations between stations as edges; using graph convolutional layers to capture the nonlinear coupling relationship between factors to obtain nonlinear coupling relationship data; dynamically adjusting the edge weights of the spatiotemporal heterogeneous graph based on a temporal attention mechanism; and analyzing the nonlinear coupling relationship data, the edge weights of the spatiotemporal heterogeneous graph, and the fused grid data based on a preset factor coupling function to obtain initial multi-factor attribution results.

9. The heterogeneous data fusion and trend assessment method according to claim 1, characterized in that, The aforementioned trend assessment analysis of the attribution results and fused grid data to obtain the trend assessment results of the atmospheric background includes: performing improved nonparametric trend detection and abrupt change point identification on the fused grid data to determine the abrupt change points; performing multi-scale and multi-dimensional trend characterization on the attribution results and fused grid data to obtain the characterization results; and performing triple verification on the characterization results to obtain the trend assessment results of the atmospheric background.

10. The heterogeneous data fusion and trend assessment method according to claim 9, characterized in that, The aforementioned triple verification of the characterization results to obtain the trend assessment results of the atmospheric background includes: internal verification using a time-stratified cross-validation method to verify the consistency of the trend assessment results in different time periods; external verification to verify the absolute accuracy of the assessment results; and physical consistency verification to verify the matching of the trend results with the attribution logic.