Satellite XCO2 Prediction Method and System Based on Multi-Source Data Fusion

Through a deep learning prediction model based on multi-source data fusion, the integration of meteorological, vegetation and man-made activity data is solved, and real-time dynamic monitoring and high-precision prediction of carbon dioxide emissions are achieved.

CN120012028BActive Publication Date: 2025-06-20NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510491472.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-20
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time dynamic monitoring of large-scale carbon dioxide emissions, and the satellite XCO2 data prediction method does not fully integrate socio-economic activities and emission drivers, resulting in low accuracy of prediction results.

Method used

Using a satellite XCO2 prediction method based on multi-source data fusion, a prediction model is constructed using a deep learning algorithm to build a prediction model, fuse a variety of data such as meteorology, vegetation, and man-made activities, perform bilinear interpolation and minimum-maximum normalization processing, and train the prediction model to improve the prediction accuracy of XCO2 data.

Benefits of technology

It significantly improves the prediction accuracy and reliability of XCO2 data, realizes real-time dynamic monitoring of carbon emissions, enhances the ability to evaluate carbon source sinks, and supports climate change research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012028B_ABST
    Figure CN120012028B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting satellite XCO2 based on multi-source data fusion, belonging to the technical field of atmospheric environment monitoring and remote sensing data processing, and is used to solve the problem of spatio-temporal discontinuity of existing satellite XCO2 data caused by reasons such as cloud cover and sensor errors. By integrating multi-source satellite remote sensing data and human activity data, a spatio-temporal coupled deep learning model is constructed to achieve high-precision filling of missing XCO2 values and prediction of future time series. Finally, the present invention breaks through the lag limitation of the traditional carbon emission inventory method, can support the real-time monitoring requirements of dynamic inversion of carbon emissions, and provides support for environmental monitoring, carbon emission assessment and policy formulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of atmospheric environment monitoring and remote sensing data processing, and specifically relates to a satellite XCO2 prediction method and system based on multi-source data fusion, which is applicable to global / regional carbon emission real-time monitoring, carbon source and sink assessment, and climate change research. Background Art

[0002] The emissions of greenhouse gases are rapidly accumulating in the atmosphere, which further intensifies the greenhouse effect and promotes the trend of global warming. Atmospheric carbon dioxide is one of the most important greenhouse gases and is produced by both human activities and natural phenomena. The core concern in this context is the emission and absorption of carbon dioxide.

[0003] Previously, research on carbon emission changes mainly used traditional anthropogenic carbon emission accounting methods based on emission inventories. This method relies on statistical yearbooks and human activity data, has a long update cycle (usually 1 - 2 years), and a low spatial resolution (provincial / city level), and cannot achieve real-time dynamic monitoring. Therefore, using atmospheric carbon dioxide concentration data to invert the spatio-temporal changes of land surface carbon emissions has become a hot method. In the traditional way, obtaining atmospheric carbon dioxide concentration data usually uses on-site measurement methods, but the obtained data is often sparse and most of the sites are concentrated in the United States and Europe, making it difficult to achieve wide-area monitoring over a large range. Therefore, the emergence and application of "carbon satellites" have gradually made up for the deficiencies of ground monitoring. Satellite remote sensing can provide real-time and high-precision atmospheric carbon dioxide column-averaged dry air mole fraction XCO2 data, and has quickly become a globally recognized new generation of global carbon inventory method. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a dynamic XCO2 prediction method that fuses multi-source data (meteorology, vegetation, human activities, etc.); based on existing data processing technologies and deep learning algorithms, it improves the accuracy and reliability of using XCO2 for carbon emission monitoring.

[0005] The satellite XCO2 prediction method based on multi-source data fusion of the present invention includes the following steps:

[0006] Step 1, constructing a data set, including historical XCO2 data and auxiliary variables ;

[0007] Step 2, constructing a prediction model for predicting XCO2 data, and the input data of the prediction model is historical XCO2 data and auxiliary variables ;

[0008] Before the data is input into the prediction model, use bilinear interpolation method to interpolate the input XCO2 data And auxiliary variables The spatio-temporal resolution of is unified to 16 days and 0.05°×0.05°. Secondly, min-max normalization is performed, and each variable is scaled to the interval [0, 1];

[0009] Step 3: Use the data set to train the prediction model to obtain a trained prediction model. Use the trained prediction model to predict the XCO2 data. Use the trained prediction model to predict the XCO2 data.

[0010] Furthermore, the auxiliary variables include meteorological data, environmental factors, human factors, social economy, and spatial characteristics;

[0011] The meteorological data includes relative humidity RH, temperature TEM, U-wind, and V-wind;

[0012] The environmental factors include vegetation cover, ground elevation, and land cover;

[0013] The vegetation cover is specifically the normalized difference vegetation index NDVI;

[0014] The spatial characteristics include latitude, longitude, and spatial average, and the spatial average refers to the average XCO2 concentration within the region.

[0015] Furthermore, the prediction model is expressed as:

[0016] ;

[0017] Where, represents the data of historical XCO2 at the past time steps, with features, , represents the data of auxiliary variables at the past time steps, with features, is the predicted value of the predicted XCO2 data at the future time steps.

[0018] Furthermore, the process of the prediction model processing the input data is as follows:

[0019] Step 2.1: Perform mean normalization on the historical XCO2 data in the time dimension to obtain ;

[0020] Step 2.2: Use one-dimensional convolution to perform sliding aggregation, extract local features by weighted averaging the data of surrounding time points, and combine the convolution result with the input sequence Perform a residual connection to obtain the output result , as the information of the target feature;

[0021] The one-dimensional convolution uses zero padding, and the convolution kernel size of the one-dimensional convolution is ;

[0022] Step 2.3: Apply the channel attention ECA mechanism to the auxiliary feature, generate channel weights through global average pooling and adaptive one-dimensional convolution, and the feature tensor after channel weighting is rearranged back to its original shape to obtain the output result of the ECA module, the weighted auxiliary feature ;

[0023] Step 2.4: Cross-attention fuse the output result of Step 2.2 with the weighted auxiliary feature to obtain the target variable containing auxiliary variable information ;

[0024] Step 2.5: Predict the target variable containing auxiliary variable information to obtain the final prediction result .

[0025] Furthermore, the specific steps of Step 2.3 are as follows:

[0026] Apply the channel attention mechanism to perform a transpose operation on the input auxiliary feature to obtain , and then apply global average pooling GAP to each channel to calculate the average value of each channel in the time series dimension;

[0027] The global average pooling is expressed as:

[0028]

[0029] where represents the value of the th channel at the th time step, c ∈ [ 1 , C 2 ] , represents the global average value of the channel in the time dimension. Through the global average pooling GAP operation, the input feature is compressed into a vector with global time features, that is, global information is extracted in the time dimension.

[0030] Apply a one-dimensional convolution operation on the channel dimension to calculate the weighted coefficients for each channel, and dynamically adjust the weights of the auxiliary features according to the importance of the features.

[0031] The convolutional kernel dynamically generates the attention weights for each channel by adaptively modeling the feature importance of different channels;

[0032] The size of the convolutional kernel is , and at the same time, padding is used to ensure that the input and output have the same channel dimension;

[0033]

[0034] Among them, and are learnable parameters used to control the size of the convolutional kernel, represents the distance to the nearest odd number; at the same time, is constrained to be odd to ensure symmetry;

[0035] After the convolution operation, the generated channel weights are processed by the Sigmoid activation function to map the weight values to the range of [0, 1]. The feature tensor weighted by the channels will be rearranged back to its original shape to obtain the output result of the ECA module .

[0036] Furthermore, the specific steps of step 2.4 are as follows:

[0037] Perform dimensional mapping on and to expand their feature space and obtain the mapped target features and auxiliary features , is the embedding dimension;

[0038] The mapped target feature sequence is used as the query vector Query, and the mapped auxiliary feature sequence is used as the key vector Key and value vector Value, which are expressed as follows:

[0039]

[0040]

[0041]

[0042] Among them, , and are learnable linear transformation matrices used to map the target variable and the auxiliary variable to the query, key, and value spaces, respectively;

[0043] The query , the key and the value are divided into multiple attention heads, and the dimension of each attention head is :

[0044] where is the number of attention heads, and new representations are obtained;

[0045] Calculate the attention scores of each attention head, convert them into a probability distribution, obtain the attention weights, and perform weighted summation on the value to obtain the output of each head, which is expressed as:

[0046] O head = [ Softmax( Q ′ ⋅ ( K ′ ) Τ d h ) ] ⋅ V ′

[0047] where is the transpose of the key, is the scaling factor;

[0048] Concatenate the outputs of each attention head, and then map them through a linear layer to obtain the target variable containing the information of the auxiliary variable.

[0049] Furthermore, the specific steps of step 2.5 are as follows:

[0050] Regard the target variable containing the information of the auxiliary variable as a univariate time series with a known period , and downsample the target variable containing the information of the auxiliary variable into multiple subsequences according to a preset period, and each subsequence contains complete period information;

[0051] Use the time-domain convolutional network TCN model to independently predict each subsequence, obtain the prediction results of each subsequence, merge the prediction results of the subsequences, and through a periodic upsampling operation, map the prediction results to the target prediction step to generate the final prediction result .

[0052] Furthermore, the time-domain convolutional network TCN model includes four cascaded TCN_Blocks;

[0053] Each TCN_Block has the same structure, including: Dilated Causal Convolution, Layer Normalization, ReLU activation function, Dropout, and an optional residual connection.

[0054] Furthermore, in the TCN_Block, intervals are inserted between the convolutional kernel elements. For a given dilation factor and convolutional kernel size , the calculation formula of the dilated convolution can be expressed as the formula:

[0055]

[0056] where is the output at time , is the -th weight of the convolutional kernel, represents causal access to historical time , is the dilation factor, which controls the sampling interval between convolutional kernel elements;

[0057] For each time step , the calculation of layer normalization is represented by the following formula:

[0058]

[0059] where is the normalized output, is the original input at time step , is the mean of all time steps, is the standard deviation of all time steps, and are learnable scaling factors and bias terms for adjusting the normalized output;

[0060] After layer normalization, the ReLU activation function is used to increase the non-linear transformation, and a Dropout layer is introduced after the activation function;

[0061] In the residual connection, a 1×1 convolution is added to align the number of channels to ensure that the residual connection can be smoothly applied between different layers.

[0062] Furthermore, the present invention also provides a satellite XCO2 prediction system based on multi-source data fusion, including:

[0063] A data processing module for cleaning multi-source data, filling missing values, unifying the spatio-temporal resolution using bilinear interpolation, and normalizing; the multi-source data includes historical XCO2 data and auxiliary variables ;

[0064] A prediction model that performs XCO2 prediction on the historical XCO2 data after data processing and auxiliary variables for XCO2 prediction;

[0065] The prediction model includes a hybrid attention module and a prediction module;

[0066] The hybrid attention module integrates the channel attention ECA mechanism and the cross-attention mechanism. The channel attention ECA mechanism is used to enhance the features of the auxiliary variables for feature enhancement; the cross-attention mechanism is used to perform feature fusion on the historical XCO2 data and the enhanced auxiliary variables to obtain a target variable containing auxiliary variable information ;

[0067] A prediction module that downsamples the target variable containing auxiliary variable information at a preset period to generate multiple subsequences, independently predicts each subsequence using a temporal convolutional network TCN, combines the prediction results of the subsequences, and performs periodic upsampling, and maps them to the final XCO2 prediction value through a fully connected layer.

[0068] Advantageous effects: The method of the present invention can alleviate two key engineering problems existing in the fields of existing XCO2 prediction and carbon emission monitoring:

[0069] [1] Data acquisition and fusion: The current carbon remote sensing field mainly focuses on filling in the missing values in satellite monitoring data. Only a small number of studies have carried out prediction operations by means of deep learning models on the basis of the complete XCO2 data constructed in previous studies, so as to obtain spatiotemporally continuous XCO2 data. However, the existing XCO2 prediction methods do not fully integrate emission driving factors such as social and economic activities and fossil fuel combustion, resulting in the inability to obtain high-precision prediction results. Therefore, the present invention collects impact factors related to human activities, and for different data sources, the present invention adopts different processing methods. At the same time, on the basis of previous studies, the present invention collects more external auxiliary data. By collecting, cleaning, interpolating, unifying the spatio-temporal resolution and the format of multi-source data, a high-quality comprehensive dataset is constructed to ensure the consistency and integrity of the data;

[0070] [2] Efficient model: The present invention proposes an efficient hybrid prediction model, which fully considers the characteristics of XCO2 itself, adopts the smallest amount of calculation and an efficient prediction algorithm, and realizes fast and accurate XCO2 prediction. Brief Description of the Drawings

[0071] Figure 1 is the overall process framework diagram of the present invention;

[0072] Figure 2 is the data acquisition and processing diagram of the present invention;

[0073] Figure 3 is the structure diagram of the entire model of the present invention;

[0074] Figure 4 is the structure diagram of the attention mechanism in the model structure diagram of the present invention;

[0075] Figure 5 is the structure diagram of the TCN prediction module in the model structure diagram of the present invention. Detailed implementation manners

[0076] The following further explains the content of the present invention in conjunction with the accompanying drawings and specific implementation manners.

[0077] The satellite XCO2 prediction method and system based on multi-source data fusion of the present invention adopt an efficient hybrid prediction model, which not only fuses multi-source data, but also fully considers the characteristics of XCO2 itself. Using the smallest amount of calculation and an efficient prediction algorithm, it realizes fast and accurate prediction, effectively reflects the spatio-temporal variation characteristics of XCO2, and significantly improves the accuracy and reliability of carbon emission monitoring. Among them, the satellite XCO2 prediction system based on multi-source data fusion of the present invention includes a data processing module, a hybrid attention module, and a prediction module, and its overall process is as Figure 1 shown.

[0078] The data processing module is used to preprocess and fuse multi-source data to solve the problems of data quality and spatio-temporal unity. The data processing module constructs a high-quality comprehensive multi-source data set by cleaning, interpolating, and fusing the collected data, etc.;

[0079] The hybrid attention module is based on the processed data set, uses the hybrid attention mechanism to extract and weight multi-feature data, enhances the model's attention to important features, and improves the prediction accuracy. The hybrid attention module combines channel attention and cross-attention mechanisms to ensure that the prediction module can make full use of the feature information in the data;

[0080] The prediction module is used to predict XCO2 for the data processed by the attention mechanism. The prediction module constructs a complex prediction model based on the features extracted by the hybrid attention module to perform accurate prediction of XCO2. The prediction results will be used to guide the formulation of carbon emission reduction policies and measures.

[0081] Figure 1 Shows the satellite XCO2 prediction method based on multi-source data fusion of the present invention, specifically as follows:

[0082] Step 1: First, start from the variable collection part. The target variable is the XCO2 data obtained by the satellite, which is the core target. To enhance the prediction ability of the model, a series of external factors are introduced in the invention, including meteorological data, environmental factors, human factors, and socioeconomic factors. The external factors provide rich background information for the model and help capture multiple factors affecting the change of XCO2. In addition, the present invention also constructs spatial features as model inputs, which emphasizes the importance of geographical location in XCO2 prediction. Latitude and longitude provide specific geographical coordinates, while spatial averaging is used to describe the average XCO2 concentration within a specific area, so as to analyze local and regional changes. The external factors and spatial features together constitute auxiliary variables as part of the model inputs.

[0083] Step 2: Process the data collected in Step 1;

[0084] The data processing part shows the unified information of all processed data, ensuring the meticulousness and accuracy of the data. Since data from different sources have different spatio-temporal resolutions, bilinear interpolation is used to unify all features to the same spatio-temporal resolution (time resolution is 16 days, and spatial resolution is 0.05°×0.05°). Secondly, feature normalization operations are performed on all input data, and min-max normalization is used to scale each variable to the [0, 1] interval, which can accelerate the model convergence speed and prevent uneven feature weights.

[0085] Step 3: Input the processed data into the prediction model and output the prediction results of the XCO2 concentration;

[0086] The prediction model part is the core of the invention. The model is trained using the XCO2 time series and auxiliary variables. An attention module is introduced to enable the model to focus on important features and improve the prediction effect. The model also includes downsampling and upsampling steps to process sequence periodic information. The final prediction module is responsible for outputting the prediction results of the XCO2 concentration.

[0087] Step 4: Verify the prediction model;

[0088] Finally, there is the Validation section. To ensure the effectiveness and reliability of the model, the TCCON site was used for validation, specifically located in Hefei, China (117.17° E, 31.9° N). The validation methods include ablation experiments, where certain variables are removed to test the robustness of the model; comparative experiments, where the superiority of the invented model is demonstrated by comparing it with existing time series prediction models; and validation based on samples, time, and space, where data from different sample data, time periods, and spatial regions are compared respectively to ensure the generalization of the model under various conditions.

[0089] In summary, Figure 1 A systematic XCO2 prediction process framework is presented, covering the complete process from data collection, processing to model construction and validation, fully considering various influencing factors, aiming to achieve accurate prediction of XCO2. In this way, complete XCO2 data can be obtained, which can provide data support for effectively reflecting the spatio-temporal variation characteristics of carbon emissions.

[0090] Figure 2 The data processing framework is presented, specifically including key elements such as data sources, formats, processing methods, and their spatial and temporal characteristics. The following is a detailed introduction to each part of the figure:

[0091] Step 1, collect multi-source data in the region (27° - 35° N, 115° - 123° E), which provides the basis for the prediction of XCO2. First is the satellite XCO2 data, which is sourced from the Orbiting Carbon Observatory-2 (OCO-2) launched by the United States in 2014. The high-precision complete dataset that has been constructed is used in the invention. This dataset has been strictly screened to ensure that data with a quality flag (xco2_quality_flag = 0) is retained to provide the most accurate XCO2 observation information. The XCO2 data for validating the model is sourced from the TCCON Hefei site.

[0092] Secondly, there are external factors. Existing research shows that meteorological data, vegetation conditions, anthropogenic emissions, etc. have a significant impact on atmospheric carbon dioxide concentration.

[0093] Meteorological data includes relative humidity (RH), temperature (TEM), etc. These data are sourced from the fifth-generation global climate and weather reanalysis dataset (ERA5) of the European Centre for Medium-Range Weather Forecasts (ECMWF). The data format is nc. In particular, data at 13:00 and 14:00 Beijing time every day are selected to match the transit time window of the carbon satellite, and further processed into daily averages.

[0094] The Normalized Difference Vegetation Index (NDVI), which is used to measure vegetation cover and health, can provide information on vegetation growth and distribution. In this invention, the MOD13Q1 vegetation product from the Moderate Resolution Imaging Spectroradiometer (MODIS) is adopted, and the storage format is nc. In addition, the SRTM3 version in the Shuttle Radar Topography Mission (SRTM) dataset is used, with the format of Tiff, which effectively depicts the surface morphology.

[0095] The land cover data used in this invention is sourced from the China Land Cover Dataset (CLCD), with the storage format of Tiff, which details the distribution of natural ecology and human activities.

[0096] The increase in carbon dioxide concentration in the atmosphere mainly stems from human activities, especially the burning of fossil fuels. Therefore, two key carbon emission datasets are also selected in the input data: the Emissions Database for Global Atmospheric Research (EDGAR) and the Open-Data Inventory for Anthropogenic Carbon dioxide (ODIAC). The data format of this dataset is nc, which is crucial for understanding the changes in XCO2.

[0097] Socio-economic data involves socio-economic factors such as GDP, population density, and energy consumption, which may affect anthropogenic emissions and ecosystems. This data is sourced from the statistical yearbook data of each city in the study area and is saved in csv format during the collection process.

[0098] Detailed information such as all datasets and their corresponding spatial resolution, temporal resolution, and data sources is shown in the following table.

[0099] As can be seen from the table content, each data source is saved in a different format, so corresponding processing and conversion operations are required. For satellite XCO2, NDVI, meteorological data, and anthropogenic factors stored in nc format, the export method is the NetCDF library in Python. The Tiff format is commonly used in Geographic Information Systems (GIS) and is suitable for storing raster data. Therefore, for land use data and elevation data, the GDAL geospatial data processing library in Python is used for processing. Finally, all data in different formats will be converted and saved in csv file format for use as input to the model.

[0100] Table 1 Dataset Details

[0101]

[0102] To ensure the accuracy and quality of external auxiliary data, the missing values in the data are first processed. For the missing values, the Kriging interpolation method is adopted in the present invention. Kriging interpolation can provide accurate interpolation results while maintaining the spatial structure characteristics. In addition, to improve the reliability and stability of the data, the Local Outlier Factor (LOF) method is used to detect and remove outliers from the data. The LOF method effectively identifies and removes noise and outliers by calculating the difference between the local density of the data points and the density of their neighbors, ensuring the data quality for subsequent model training.

[0103] Finally, bilinear interpolation is used for spatio-temporal matching to unify the time dimension of all data into 16 days and the spatial dimension into 0.05°×0.05°. In this way, a complete multi-source time-series dataset is constructed.

[0104] Figure 3 The overall structure of the prediction model in the present invention is shown. In many studies, the XCO2 time series has been found to have significant periodic and seasonal variations. Based on this characteristic, the present invention proposes a hybrid prediction model - HATSNet (Hybrid Attention Temporal Sequence Network). In the design process of the HATSNet model, the periodic characteristics in the XCO2 time series are fully considered and utilized, that is, the XCO2 time series has a known period .

[0105] The purpose of the present invention is to use the entire multi-source historical data to predict future results, and find a suitable function through learning for accurate prediction, that is , where represents the data of the target variable at the past time steps, with features (here , representing XCO2), represents the data of the auxiliary variable at the past time steps, with features, is the predicted value of the target variable to be predicted at the future time steps.

[0106] The HATSNet model combines the characteristics of time-series data and spatial information, and has high computational performance and high prediction accuracy. Specifically, the process of the HATSNet model for processing input data is as follows:

[0107] Step 1, the HATSNet model first receives the target variable (XCO2 concentration) and auxiliary variables (such as meteorological data, vegetation data, longitude and latitude, etc.). The input data is first processed by MinMaxScaler normalization to ensure that all features have the same scale. For a given dataset , where each data point belongs to , through this formula, all data points will be linearly mapped into the interval [0, 1]. The specific operation formula is as follows:

[0108]

[0109] where is the normalized data point, and are the minimum value and the maximum value in the dataset, respectively.

[0110] Step 2, in order to eliminate the distribution bias of the sequence and make the data in different time periods have a more consistent distribution, the target feature is normalized by the mean along the time dimension . For each sample , it is normalized along the dimension of time steps. Then, the calculated mean is used to normalize the data at each time step. The normalized result, denoted as , has the same dimension, only removing the mean shift in the time dimension and adding it back after the model output.

[0111] In addition, in order to capture the local dependencies in the time series and reduce the impact of outliers, one-dimensional convolution is used to perform sliding aggregation. This convolution operation uses zero-padding and sets the convolution kernel size to , where is the sequence period of XCO2. Local features are extracted by weighted averaging the data of surrounding time points. At the same time, the convolution result is connected with the input sequence by residual connection to retain the original information and avoid over-transformation, thereby improving the prediction accuracy and robustness, and obtaining the output result .

[0112] Step 3, the Attention Mechanism is a core technology in deep learning, aiming to improve the computational efficiency and prediction accuracy of the model by adaptively focusing on the important information in the input data.

[0113] In the XCO2 prediction task, the model needs to integrate information from multiple data sources. Therefore, how to effectively process these different data becomes particularly crucial. To solve this problem, after the convolutional layer, the model uses an improved hybrid attention mechanism, which can adaptively adjust the feature weights globally and endow the model with flexible feature selection capabilities.

[0114] Step 4, as mentioned above, the XCO2 time series is found to have significant periodic and seasonal variations. Therefore, the XCO2 time series can be regarded as consisting of periodic components and trend components. Therefore, the core object of prediction changes to predicting the future trend component.

[0115] Based on the decomposition of the above-mentioned periodic features, the present invention downsamples the time series according to the period to extract multiple subsequences. On this basis, a Temporal Convolutional Network (TCN) is used to independently predict each subsequence. After passing through the TCN module, the predicted future values of each subsequence are obtained. Finally, the prediction results of all subsequences are merged, and the predicted value of the final prediction step is obtained through a periodic upsampling operation.

[0116] Figure 4 shows Figure 3 the structure diagram of the attention mechanism in the model structure diagram. This module enhances the reusability of features through channel-level feature connections, thereby improving the expression ability of the model. Its innovation lies in introducing a lightweight channel attention mechanism and combining a multi-head cross-attention mechanism, enabling the model to flexibly integrate auxiliary feature information during the prediction process. Although these auxiliary features do not directly participate in the final prediction result, they significantly enhance the representation of the target variable.

[0117] This module consists of two parts: the first part is the channel attention mechanism applied to the auxiliary features, and the second part is the cross-attention mechanism for fusing the information of the target features and the auxiliary features. To better understand the technical solution of the present invention, the following will detail these two parts and their working principles.

[0118] 1) Channel-Attention channel attention mechanism: Efficient Channel Attention (ECA) is a lightweight channel attention mechanism that enhances the model's attention to key features by adaptively adjusting the weights of each channel in the input features, thereby improving the expression ability of the model. Specifically, the Efficient Channel Attention (ECA) mechanism first processes the input auxiliary features The transpose operation is performed to obtain , so as to meet the input requirements of pooling and convolution operations. Then, global average pooling (GAP) is applied to each channel to calculate the average value of each channel in the time series dimension. Specifically, for each channel , the operation of global average pooling can be expressed by the following formula:

[0119]

[0120] where represents the value of the th channel at the th time step, c ∈ [ 1 , C 2 ] , represents the global average value of the th channel in the time dimension. Through the global average pooling GAP operation, the input feature is compressed into a vector with global time features, that is, global information is extracted in the time dimension.

[0121] In the Efficient Channel Attention (ECA) mechanism, a one-dimensional convolution operation is applied to the channel dimension to calculate the weighted coefficient of each channel, so as to dynamically adjust the weight of the auxiliary feature according to the importance of the feature. The convolutional kernel adaptively models the feature importance of different channels and dynamically generates the attention weight of each channel. The size of this convolutional kernel is , and at the same time padding is adopted to ensure that the input and output channel dimensions are consistent. To optimize channel feature selection, the Efficient Channel Attention (ECA) module adaptively adjusts the size of the convolutional kernel according to the input data. Among them and are learnable parameters used to control the size of the convolutional kernel, represents the odd number closest to . At the same time, the size of the convolutional kernel is constrained to be an odd number to ensure symmetry.

[0122] After the convolution operation, the generated channel weights are processed through the Sigmoid activation function to map the weight values to the range of [0, 1]. Then, through element-wise multiplication, the channel features with larger weights are enhanced, while the channel features with smaller weights are suppressed, thereby highlighting the important features that the model needs to focus on. Finally, the feature tensor weighted by channels is rearranged back to its original shape to obtain the output result of the Efficient Channel Attention (ECA) module 。

[0123] 2) Cross-Attention: To better capture the interaction relationship between the target variable and the auxiliary variable, first perform dimensional mapping on and to expand its feature space and obtain and , is the embedding dimension. Based on and , an effective information interaction is established between the target variable and the auxiliary variable by constructing a multi-head cross-attention mechanism

[0124] Specifically, the mapped target feature sequence is used as the query vector (Query), representing the information that is expected to be obtained from the auxiliary variable; while the mapped auxiliary feature sequence serves as the key vector (Key) and value vector (Value), providing useful context information for the target variable. These mapping operations can be represented by the following formulas

[0125]

[0126]

[0127]

[0128] Among them, , and are learnable linear transformation matrices, which are used to map the target variable and the auxiliary variable to the query, key, and value spaces respectively is the embedding dimension. Through these mappings, the target variable and the auxiliary variable are aligned in the same high-dimensional space to prepare for the subsequent cross-attention calculation

[0129] At the same time, the multi-head attention mechanism is also adopted in this application to calculate the outputs of multiple attention heads in parallel. In the multi-head attention mechanism, the query ,the key and the value will be divided into multiple attention heads, and the dimension of each attention head is , where is the number of attention heads, and thus a new representation is obtained.

[0130] For each attention head, the attention scores are calculated independently and a weighted output is generated. The attention scores are obtained by calculating the dot product between the query and the key , which represents the relevance between the query and each key. To convert the attention scores into a probability distribution, the calculated attention scores are subjected to a softmax operation to obtain the attention weights. These weights are used to weight the value matrix, and the normalized attention weights are used to perform a weighted sum on the values to obtain the output of each head. The specific formula is as follows:

[0131] O head = [ Softmax( Q ′ ⋅ ( K ′ ) Τ d h ) ] ⋅ V ′

[0132] where is the transpose of the key, is a scaling factor to prevent the dot product result from being too large or too small and to help the stability of the training process.

[0133] In the multi-head attention mechanism, each attention head will calculate a weighted sum result (i.e., ). Secondly, the outputs of all attention heads are concatenated together to form the final representation. To convert the concatenated high-dimensional representation back to the original dimension of the target variable, the concatenated output needs to be mapped back to the original dimension through a linear layer. Finally, is the final output after the linear transformation, representing the target variable (including auxiliary variable information) at each time step.

[0134] Figure 5 shows Figure 3 the structure diagram of the prediction module in the prediction model structure diagram, which is the core content of the entire invention. By introducing the above-mentioned hybrid attention mechanism, the target variable and the auxiliary variable can be effectively fused to obtain a representation containing rich information (where ). Therefore, its input can be regarded as a univariate time series with a known period .

[0135] This sequence can be regarded as consisting of a periodic component and a trend component , that is Therefore, the core object of prediction changes to the prediction of future trend components.

[0136] 1) Model Atructure: Based on the above decomposition of periodic characteristics, first downsample the time series according to the period to extract multiple subsequences. Each subsequence includes complete periodic information, thus ensuring that the periodic pattern can be fully learned and utilized. On this basis, use the improved Time Convolutional Network (TCN) model to independently predict each subsequence.

[0137] The periodic components in each subsequence are already implicit in the input data, and the task of the improved Time Convolutional Network (TCN) is to learn the trend changes within the period from these periodic subsequences. After the prediction by the improved Time Convolutional Network (TCN), the future value predictions of each subsequence are obtained. Finally, the prediction results of all subsequences are merged, and through the periodic upsampling operation, these prediction results are mapped to the target prediction step to generate the final prediction result .

[0138] 2) TCN: Since each subsequence contains periodic information, TCN can effectively model these subsequences as local time series. In each subsequence, the periodic components are already embedded in the data, so the task of TCN is to learn the long-term trend changes in each subsequence and focus on capturing the temporal dependencies therein.

[0139] The present invention proposes an improved TCN architecture, which consists of 4 TCN_Blocks. Each TCN_Block includes Dilated Causal Convolution, Layer Normalization, ReLU activation function, Dropout, and optional residual connections.

[0140] Classic TCN uses causal convolution to replace traditional one-dimensional convolution to ensure that the model does not violate the time order when performing time series prediction, that is, the output at the current moment only depends on the current and previous inputs and is not affected by future information. For each input subsequence , given the convolution kernel size of , the operation of causal convolution can be expressed as:

[0141]

[0142] where is the output at time , is the th weight of the convolution kernel, is the time input of.

[0143] To ensure the causality of the model and maintain the length of the time series, the input data needs to be properly padded. In the TCN model, the padding amount is related to the dilation factor and the convolutional kernel size . Therefore, the dilation factor and the convolutional kernel size jointly determine the size of the padding amount, ensuring that the convolutional operation can adapt to the dependencies of longer time series.

[0144] TCN adopts the dilated convolution mechanism. By inserting intervals (controlled by the dilation factor ) between the convolutional kernel elements, the receptive field is expanded to capture long-term time dependencies. Compared with traditional convolution, without increasing the number of parameters or the network depth, it flexibly expands the interval sampling range of the input data by adjusting the value, thereby enhancing the ability to model historical trends. Specifically, for a given dilation factor and convolutional kernel size , the calculation formula of the dilated convolution can be expressed as the formula:

[0145]

[0146] where is the output at time , is the th weight of the convolutional kernel, represents the causal access to the historical time , is the dilation factor, controlling the sampling interval between the convolutional kernel elements.

[0147] In addition, by adjusting the , the coverage range of the receptive field can be flexibly expanded. TCN stacks multiple layers of dilated convolution, making the receptive field grow exponentially, thereby efficiently capturing long-term dependencies. Let the dilation factor of the th layer be , then the receptive field of the th layer can be expressed by the following formula:

[0148]

[0149] where is the size of the receptive field of this layer, is the size of the receptive field of the upper layer, and the receptive field of the initial layer .

[0150] In traditional TCN implementations, batch normalization is usually adopted to accelerate training. However, the performance of batch normalization highly depends on the choice of batch size, which easily leads to unstable training processes. To address this issue, this model uses layer normalization instead of batch normalization. Layer normalization has significant advantages in time series data modeling: it normalizes independently along the feature dimension at each time step without relying on the statistical information of other samples within the batch. For each time step , the calculation of layer normalization can be expressed by the following formula:

[0151]

[0152] where is the normalized output, is the original input at time step , is the mean of all time steps, is the standard deviation of all time steps, and are learnable scaling factors and bias terms used to adjust the normalized output.

[0153] After layer normalization, the model usually uses an activation function to increase non-linear transformations, enabling the network to learn more complex patterns. In this model, the ReLU (Rectified Linear Unit) activation function is used, which is specifically expressed as the following formula:

[0154]

[0155] To prevent overfitting and improve the generalization ability of the model, this model introduces a Dropout layer after each convolutional block to randomly discard neuron outputs. Through this means, the robustness of the model is further enhanced, and overfitting to specific features is reduced.

[0156] In traditional TCN architectures, residual connections are usually relied on to help alleviate the vanishing gradient problem in deep networks. However, as the network depth increases, when the number of input and output channels is inconsistent, the effectiveness of residual connections may be affected. To solve this problem, this model innovatively uses 1×1 convolutions to achieve channel alignment, ensuring that residual connections can be smoothly applied between different layers. This design helps solve the residual connection problem when dimensions are inconsistent, thus ensuring unobstructed information flow between network layers.

[0157] After predicting each subsequence through the improved TCN module, the model is based on the period Upsample, and then map the prediction result to the final prediction step through a fully connected layer . Finally, add the mean of the time dimension of calculated in step 2 and perform min-max inverse normalization to obtain the final predicted output of XCO2 .

[0158] The data division and parameter settings in the present invention are as follows:

[0159] The data covers the time period from 2016 to 2022 and is scientifically divided into two subsets: the data from 2016 to 2020 is used to construct the training set and the validation set, and the validation set accounts for 20% and is used for the training and tuning of the model; the data from 2021 and 2022 is used as the test set to evaluate the prediction performance of the model. During the model training process, the learning rate is set to 0.001 to ensure that the model can stably and effectively learn the features in the data. In addition, the number of training epochs is set to 50, the batch size is 256, and the Adam optimization algorithm is adopted to optimize the model parameters with its efficient gradient descent ability and good hyperparameter adaptability. The period of XCO2 is set to 23, indicating that each period is one year. The embedding dimension of the hybrid attention module is set to 64, and 4 attention heads are selected. The parameters of the TCN module are set to [25, 128, 64, 32], and the convolution kernel size is 3.

[0160] To verify the superiority of the model of the present invention, multiple evaluation metrics (mean absolute error MAE, root mean square error RMSE, mean absolute percentage error MAPE, and coefficient of determination R 2 ) are used to comprehensively evaluate the performance of the model. The formulas of the evaluation metrics are as follows:

[0161]

[0162]

[0163]

[0164]

[0165]

[0166] Among them, is the true value, is the predicted value, is the mean of the true values, is the number of samples. The complexity during the training process makes the absolute value function have good robustness when dealing with outliers, but its derivative is discontinuous when the model output is close to the true value. On the contrary, the square function shows different characteristics in this regard. To address the limitations of these two functions, the present invention selects the Log-cosh loss function.

[0167] In addition, this study also conducted ablation experiments and comparison time to verify the superiority and generalization of the model from multiple aspects.

[0168] By implementing the present invention, users can alleviate the key engineering problems existing in the existing high-precision and spatio-temporally continuous XCO2 data acquisition and carbon emission monitoring technologies:

[0169] [1] Challenges in multi-source data acquisition and fusion: In the context of the lack of inventions for obtaining spatio-temporally continuous XCO2 data using deep learning algorithms, the present invention conducts deep learning prediction tasks based on the high-precision XCO2 data constructed by predecessors. At the same time, this study additionally collects external factors related to human activities. In addition, aiming at the spatio-temporal scale differences and missing value problems of satellite XCO2 data and multi-source heterogeneous data such as meteorology, vegetation, and socioeconomic, a unified data alignment and feature enhancement method is proposed. Through technologies such as dynamic spatio-temporal interpolation and multi-modal feature encoding, co-modeling of meteorological data, human factors, etc. with XCO2 data is realized, providing high-robustness input for the deep learning model;

[0170] [2] Challenges in lightweight high-precision modeling: Based on the seasonal and trend characteristics of XCO2 data, a lightweight spatio-temporal attention mechanism and multi-task learning framework are designed to fuse multi-source auxiliary variables, improve the prediction accuracy while reducing the computational overhead of the model, enable it to efficiently process large-scale spatio-temporal data and support real-time XCO2 assessment, and capture the spatio-temporal change trends of carbon emissions. Thereby significantly improving the spatio-temporal continuity of the near-surface carbon emission inversion results and enhancing the model's ability to analyze complex human and natural factors.

Claims

1. Satellite XCO2 prediction method based on multi-source data fusion, characterized in that: The steps include: Step 1: Build a dataset including historical XCO2 data and auxiliary variables ; Step 2: Build a prediction model for predicting XCO2 data, where the input data of the prediction model is historical XCO2 data. and auxiliary variables ; The prediction model is expressed as: ; in, Indicates historical XCO2 data in the past time steps of data, with Features, Indicates that the auxiliary variable is in the past time steps of data, with Features, It is the predicted XCO2 data in the future The predicted value of time steps; The process of the prediction model processing the input data is as follows: Step 2.1, historical XCO2 data By normalizing the mean value according to the time dimension, we get ; Step 2.2, use one-dimensional convolution to Perform sliding aggregation, extract local features by weighted averaging the data of surrounding time points, and combine the convolution result with the input sequence Perform residual connection to get the output result , as information of target features; The one-dimensional convolution uses zero padding, and the convolution kernel size of the one-dimensional convolution is ; Step 2.3, for auxiliary variables The channel attention ECA mechanism is used to generate channel weights through global average pooling and adaptive one-dimensional convolution. The channel-weighted feature tensor is rearranged back to its original shape to obtain the output result of the ECA module. The weighted auxiliary feature ; Step 2.4: Output the result of step 2.2 With the weighted auxiliary features Perform cross attention fusion to obtain the target variable containing auxiliary variable information ; Step 2.5, for the target variable containing auxiliary variable information Make predictions and get the final prediction results ; Step 3: Use the data set to train the prediction model to obtain a trained prediction model, and use the trained prediction model to predict the XCO2 data.

2. The satellite XCO2 prediction method based on multi-source data fusion according to claim 1 is characterized in that: The auxiliary variables include meteorological data, environmental factors, human factors, social economy, and spatial characteristics; The meteorological data include relative humidity RH, temperature TEM, U-shaped wind and V-shaped wind; The environmental factors include vegetation cover, ground elevation, and land cover; The vegetation coverage is specifically the Normalized Difference Vegetation Index NDVI; The spatial characteristics include latitude, longitude and spatial average, and the spatial average refers to the average XCO2 concentration in the region.

3. The satellite XCO2 prediction method based on multi-source data fusion according to claim 1 is characterized in that: The step 2.3 is as follows: Use channel attention mechanism to input auxiliary features Perform the transpose operation to obtain , then apply global average pooling GAP to each channel to calculate the average value of each channel in the time series dimension; A one-dimensional convolution operation is applied to the channel dimension to calculate the weight coefficient of each channel to dynamically adjust the weight of the auxiliary features according to the importance of the features. The convolution kernel dynamically generates the attention weight of each channel by adaptively modeling the feature importance of different channels; The size of the convolution kernel is , while using Padding to ensure that the channel dimensions of input and output are consistent; ; in and is a learnable parameter used to control the size of the convolution kernel. Indicates distance The nearest odd number; at the same time, is constrained to be an odd number to ensure symmetry; After the convolution operation, the generated channel weights are processed by the Sigmoid activation function to map the weight values ​​to the range of [0, 1]. Then, through element-by-element multiplication, the channel features with larger weights are enhanced, while the channel features with smaller weights are suppressed, thereby highlighting the important features that the model needs to pay attention to. Finally, the channel-weighted feature tensor is rearranged back to its original shape to obtain the output result of the ECA module. .

4. The satellite XCO2 prediction method based on multi-source data fusion according to claim 1 is characterized in that: The step 2.4 is as follows: right and Perform dimension mapping to expand its feature space. Get the mapped target features and auxiliary features , is the embedding dimension; The mapped target feature sequence Used as query vector Query, the auxiliary feature sequence after mapping Then the key vector Key and the value vector Value are expressed as follows: ; ; ; in, , and are learnable linear transformation matrices used to map target variables and auxiliary variables to query, key, and value spaces, respectively; The query ,key Sum Divided into multiple attention heads, the dimension of each attention head is : ; in, is the number of attention heads, and the new representation is obtained ; Calculate the attention score of each attention head and convert it into a probability distribution to get the attention weight. Perform weighted summation to get the output of each head, expressed as: ; in, is the transpose of the key, is the factor used for scaling; The output of each attention head is concatenated and then mapped through a linear layer to obtain the target variable containing auxiliary variable information. .

5. The satellite XCO2 prediction method based on multi-source data fusion according to claim 3 is characterized in that: Step 2.5 is as follows: The target variable containing the auxiliary variable information Considered as a known cycle The univariate time series , for the target variable containing auxiliary variable information , downsampled into multiple subsequences according to a preset period, each subsequence contains complete period information; The time domain convolutional network (TCN) model is used to independently predict each subsequence, obtain the prediction results for each subsequence, merge the prediction results of the subsequences, and map the prediction results to the target prediction step size through periodic upsampling operations. , thus generating the final prediction result .

6. The satellite XCO2 prediction method based on multi-source data fusion according to claim 5 is characterized in that: The time domain convolutional network TCN model includes four cascaded TCN_Blocks; Each TCN_Block has the same structure, including: dilated causal convolution, layer normalization, activation function, Dropout, and residual connection.

7. The satellite XCO2 prediction method based on multi-source data fusion according to claim 6 is characterized in that: In the TCN_Block, the convolution kernel elements are inserted with spacing, for a given dilation factor and the convolution kernel size , the calculation formula of the dilated convolution is expressed as: ; in, It is at the moment The output at is the convolution kernel weights, Expressing historical moments The causal access, is the dilation factor, which controls the sampling interval between convolution kernel elements; For each time step , the calculation of layer normalization is expressed by the following formula: ; in, is the normalized output, is the time step The original input at is the mean value over all time steps, is the standard deviation of all time steps, and is a learnable scaling factor and bias term used to adjust the normalized output; After layer normalization, the ReLU activation function is used to add nonlinear transformation, and the Dropout layer is introduced after the activation function; A 1×1 convolution is added to the residual connection.

8. Satellite XCO2 prediction system based on multi-source data fusion, characterized by: include: Data processing module, used to clean multi-source data, fill missing values, use bilinear interpolation to unify spatiotemporal resolution, and normalize; the multi-source data includes historical XCO2 data and auxiliary variables ; Prediction model, historical XCO2 data after data processing and auxiliary variables Conduct XCO2 forecasts; The prediction model includes a hybrid attention module and a prediction module; The hybrid attention module integrates the channel attention ECA mechanism and the cross attention mechanism, and uses the channel attention ECA mechanism to Perform feature enhancement; use the cross attention mechanism to integrate historical XCO2 data With the enhanced auxiliary variables Perform feature fusion to obtain the target variable containing auxiliary variable information ; The prediction module predicts the target variable containing auxiliary variable information Downsampling is performed according to a preset period to generate multiple subsequences, and a time domain convolutional network (TCN) is used to independently predict each subsequence. The prediction results of the subsequences are merged, and periodic upsampling is performed, and the results are mapped to the final XCO2 prediction value through a fully connected layer.