Satellite XCO2 prediction method and system based on multi-source data fusion
Through multi-source data fusion and deep learning algorithms, an efficient satellite XCO2 prediction system was built, which solved the problem of insufficient XCO2 monitoring accuracy and real-time performance in the existing technology, and achieved high-precision and real-time carbon emission monitoring.
Patent Information
- Application Number
- CN202510491472.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art is difficult to achieve high-precision and real-time monitoring of carbon dioxide (XCO2) and monitoring of dynamic changes in carbon emissions, especially in terms of data acquisition and fusion.
Using a satellite XCO2 prediction method based on multi-source data fusion, we use deep learning algorithms to build prediction models, perform data preprocessing and feature extraction, and combine channel attention and cross attention mechanisms to achieve efficient XCO2 prediction.
It significantly improves the accuracy and reliability of XCO2 monitoring, enables real-time and high-precision carbon emission monitoring, and enhances support for carbon source sink assessment and climate change research.
Smart Images

Figure CN120012028A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of atmospheric environment monitoring and remote sensing data processing, and specifically relates to a satellite XCO2 prediction method and system based on multi-source data fusion, which is suitable for real-time monitoring of global / regional carbon emissions, carbon source and sink assessment, and climate change research. Background Art
[0002] Greenhouse gas emissions are rapidly accumulating in the atmosphere, a trend that further exacerbates the greenhouse effect and drives global warming. Atmospheric carbon dioxide is one of the most important greenhouse gases, produced by both human activities and natural phenomena. The core issue of concern in this context is the emission and absorption of carbon dioxide.
[0003] Previously, the research on carbon emission changes mainly used the traditional anthropogenic carbon emission accounting method based on emission inventory. This method relies on statistical yearbooks and human activity data, has a long update cycle (usually 1-2 years), low spatial resolution (provincial / city level), and cannot achieve real-time dynamic monitoring. Therefore, using atmospheric carbon dioxide concentration data to invert the spatiotemporal changes of land surface carbon emissions has become a hot method. In traditional methods, atmospheric carbon dioxide concentration data is usually obtained by site measurement, but the data obtained are often sparse and the sites are mostly concentrated in the United States and Europe, making it difficult to achieve large-scale wide-area monitoring. Therefore, the emergence and application of "Carbon Satellite" has gradually made up for the shortcomings of ground monitoring. Satellite remote sensing can provide real-time and high-precision atmospheric carbon dioxide column average dry air mole fraction XCO2 data, which has quickly become an internationally recognized new generation of global carbon inventory methods. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a dynamic XCO2 prediction method that integrates multi-source data (meteorological, vegetation, human activities, etc.); based on existing data processing technology and deep learning algorithms, the accuracy and reliability of carbon emission monitoring using XCO2 are improved.
[0005] The satellite XCO2 prediction method based on multi-source data fusion of the present invention comprises the following steps:
[0006] Step 1: Build a dataset including historical XCO2 data and auxiliary variables ;
[0007] Step 2: Build a prediction model for predicting XCO2 data, where the input data of the prediction model is historical XCO2 data. and auxiliary variables ;
[0008] Before the data is fed into the prediction model, the input XCO2 data is interpolated using bilinear interpolation. and auxiliary variables The temporal and spatial resolutions were unified to 16 days and 0.05°×0.05°. Secondly, the minimum-maximum normalization was performed and each variable was scaled to the interval [0, 1].
[0009] Step 3: train the prediction model using the data set to obtain a trained prediction model, use the trained prediction model to predict the XCO2 data, and use the trained prediction model to predict the XCO2 data.
[0010] Furthermore, the auxiliary variables include meteorological data, environmental factors, human factors, social economy, and spatial characteristics;
[0011] The meteorological data include relative humidity RH, temperature TEM, U-shaped wind and V-shaped wind;
[0012] The environmental factors include vegetation cover, ground elevation, and land cover;
[0013] The vegetation coverage is specifically the Normalized Difference Vegetation Index NDVI;
[0014] The spatial characteristics include latitude, longitude and spatial average, and the spatial average refers to the average XCO2 concentration in the region.
[0015] Furthermore, the prediction model is expressed as:
[0016] ;
[0017] in, Indicates historical XCO2 data in the past time steps of data, with Features, , Indicates that the auxiliary variable is in the past time steps of data, with Features, It is the predicted XCO2 data in the future The predicted value for each time step.
[0018] Furthermore, the prediction model processes the input data as follows:
[0019] Step 2.1, historical XCO2 data By normalizing the mean value according to the time dimension, we get ;
[0020] Step 2.2, use one-dimensional convolution to Perform sliding aggregation, extract local features by weighted averaging the data of surrounding time points, and combine the convolution result with the input sequence Perform residual connection to get the output result , as information of target features;
[0021] The one-dimensional convolution uses zero padding, and the convolution kernel size of the one-dimensional convolution is ;
[0022] Step 2.3, the channel attention ECA mechanism is used for the auxiliary features. The channel weights are generated through global average pooling and adaptive one-dimensional convolution. The channel-weighted feature tensor is rearranged back to its original shape to obtain the output result of the ECA module. The weighted auxiliary features ;
[0023] Step 2.4: Output the result of step 2.2 With the weighted auxiliary features Perform cross attention fusion to obtain the target variable containing auxiliary variable information ;
[0024] Step 2.5, for the target variable containing auxiliary variable information Make predictions and get the final prediction results .
[0025] Furthermore, the step 2.3 is specifically as follows:
[0026] Use channel attention mechanism to input auxiliary features Perform the transpose operation to obtain , then apply global average pooling GAP to each channel to calculate the average value of each channel in the time series dimension;
[0027] The global average pooling is expressed as:
[0028]
[0029] in, express No. The channel in The value of the time step, c ∈ [ 1 , C 2 ] , Indicates The global average of the channel in the time dimension. Through the global average pooling GAP operation, the input feature is compressed into a vector with global time characteristics , that is, global information is extracted in the time dimension.
[0030] A one-dimensional convolution operation is performed on the channel dimension to calculate the weight coefficient of each channel to dynamically adjust the weight of the auxiliary features according to the importance of the features.
[0031] The convolution kernel dynamically generates the attention weight of each channel by adaptively modeling the feature importance of different channels;
[0032] The size of the convolution kernel is , while using Padding to ensure that the channel dimensions of input and output are consistent;
[0033]
[0034] in and is a learnable parameter used to control the size of the convolution kernel. Indicates distance The nearest odd number; at the same time, is constrained to be an odd number to ensure symmetry;
[0035] After the convolution operation, the generated channel weights are processed by the Sigmoid activation function to map the weight values to the range of [0, 1]. The channel-weighted feature tensor is rearranged back to its original shape to obtain the output of the ECA module. .
[0036] Furthermore, the step 2.4 is specifically as follows:
[0037] right and Perform dimension mapping, expand its feature space, and obtain the mapped target features and auxiliary features , is the embedding dimension;
[0038] The mapped target feature sequence Used as query vector Query, the auxiliary feature sequence after mapping Then the key vector Key and the value vector Value are expressed as follows:
[0039]
[0040]
[0041]
[0042] in, , and are learnable linear transformation matrices used to map target variables and auxiliary variables to query, key, and value spaces, respectively;
[0043] The query ,key Sum Divided into multiple attention heads, the dimension of each attention head is :
[0044] in, is the number of attention heads, and the new representation is obtained ;
[0045] Calculate the attention score of each attention head and convert it into a probability distribution to get the attention weight. Perform weighted summation to get the output of each head, expressed as:
[0046] O head = [ Softmax( Q ′ ⋅ ( K ′ ) Τ d h ) ] ⋅ V ′
[0047] in, is the transpose of the key, is the factor used for scaling;
[0048] The output of each attention head is concatenated and then mapped through a linear layer to obtain the target variable containing auxiliary variable information. .
[0049] Furthermore, the step 2.5 is specifically as follows:
[0050] The target variable containing the auxiliary variable information Considered as a known cycle The univariate time series , for the target variable containing auxiliary variable information , downsampled into multiple subsequences according to a preset period, each subsequence contains complete period information;
[0051] The time domain convolutional network (TCN) model is used to independently predict each subsequence, obtain the prediction results for each subsequence, merge the prediction results of the subsequences, and map the prediction results to the target prediction step size through periodic upsampling operations. , thus generating the final prediction result .
[0052] Furthermore, the time domain convolutional network TCN model includes four cascaded TCN_Blocks;
[0053] Each TCN_Block has the same structure, including: dilated causal convolution Dilated Gausal Conv, layer normalization LayerNorm, activation function ReLU, Dropout, and optional residual connection.
[0054] Furthermore, in the TCN_Block, the convolution kernel elements are spaced apart, and for a given dilation factor and the convolution kernel size , the calculation formula of the dilated convolution can be expressed as:
[0055]
[0056] in, It is at the moment The output at is the convolution kernel weights, Expressing historical moments The causal access, is the dilation factor, which controls the sampling interval between convolution kernel elements;
[0057] For each time step , the calculation of layer normalization is expressed by the following formula:
[0058]
[0059] in, is the normalized output, is the time step The original input at is the mean value over all time steps, is the standard deviation of all time steps, and is a learnable scaling factor and bias term used to adjust the normalized output;
[0060] After layer normalization, the ReLU activation function is used to add nonlinear transformation, and the Dropout layer is introduced after the activation function;
[0061] A 1×1 convolution is added to the residual connection to achieve alignment of the number of channels and ensure that the residual connection can be successfully applied between different layers.
[0062] Furthermore, the present invention also provides a satellite XCO2 prediction system based on multi-source data fusion, comprising:
[0063] Data processing module, used to clean multi-source data, fill missing values, use bilinear interpolation to unify spatiotemporal resolution, and normalize; the multi-source data includes historical XCO2 data and auxiliary variables ;
[0064] Prediction model, historical XCO2 data after data processing and auxiliary variables Conduct XCO2 forecasts;
[0065] The prediction model includes a hybrid attention module and a prediction module;
[0066] The hybrid attention module integrates the channel attention ECA mechanism and the cross attention mechanism, and uses the channel attention ECA mechanism to Perform feature enhancement; use the cross attention mechanism to integrate historical XCO2 data With the enhanced auxiliary variables Perform feature fusion to obtain the target variable containing auxiliary variable information ;
[0067] The prediction module predicts the target variable containing auxiliary variable information Downsampling is performed according to a preset period to generate multiple subsequences, and a time domain convolutional network (TCN) is used to independently predict each subsequence. The prediction results of the subsequences are merged, and periodic upsampling is performed, and the results are mapped to the final XCO2 prediction value through a fully connected layer.
[0068] Beneficial effects: The method of the present invention can alleviate two key engineering problems in the existing XCO2 prediction and carbon emission monitoring fields:
[0069] [1] Data acquisition and fusion: The current field of carbon remote sensing mainly focuses on filling in the missing values in satellite monitoring data. Only a small number of studies have implemented prediction operations based on the complete XCO2 data constructed by previous studies with the help of deep learning models to obtain spatiotemporally continuous XCO2 data. However, the existing XCO2 prediction methods do not fully integrate emission driving factors such as socio-economic activities and fossil fuel combustion, resulting in the inability to obtain high-precision prediction results. Therefore, the present invention collects the influencing factors of human activities, and for different data sources, the present invention adopts different processing methods. At the same time, based on previous studies, the present invention collects more external auxiliary data. By collecting, cleaning, interpolating, unifying the spatiotemporal resolution and format of multi-source data, a high-quality comprehensive data set is constructed to ensure data consistency and integrity;
[0070] [2] Efficient model: This paper proposes an efficient hybrid prediction model that fully considers the characteristics of XCO2 itself, uses minimal computational effort and an efficient prediction algorithm, and achieves fast and accurate XCO2 prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is the overall process framework diagram of the present invention;
[0072] Figure 2 is a data acquisition and processing diagram of the present invention;
[0073] Figure 3 It is a structural diagram of the entire model of the present invention;
[0074] Figure 4 It is a structural diagram of the attention mechanism in the structural diagram of the model of the present invention;
[0075] Figure 5 It is a structural diagram of the TCN prediction module in the model structural diagram of the present invention. DETAILED DESCRIPTION
[0076] The content of the present invention is further explained below in conjunction with the accompanying drawings and specific implementation methods.
[0077] The satellite XCO2 prediction method and system based on multi-source data fusion of the present invention adopts an efficient hybrid prediction model, which not only integrates multi-source data, but also fully considers the characteristics of XCO2 itself, adopts the minimum amount of calculation and efficient prediction algorithm, realizes fast and accurate prediction, effectively reflects the spatiotemporal variation characteristics of XCO2, and significantly improves the accuracy and reliability of carbon emission monitoring. Among them, the satellite XCO2 prediction system based on multi-source data fusion of the present invention includes a data processing module, a hybrid attention module and a prediction module, and its overall process is as follows: Figure 1 shown.
[0078] The data processing module is used to pre-process and fuse multi-source data to solve data quality and spatiotemporal unification issues. The data processing module cleans, interpolates and fuses the collected data to build a high-quality comprehensive multi-source data set;
[0079] Based on the processed data set, the hybrid attention module uses the hybrid attention mechanism to extract and weight the multi-feature data, enhance the model's attention to important features, and improve the accuracy of prediction. The hybrid attention module combines channel attention and cross attention mechanisms to ensure that the prediction module can fully utilize the feature information in the data;
[0080] The prediction module is used to predict XCO2 for data processed by the attention mechanism. Based on the features extracted by the hybrid attention module, the prediction module builds a complex prediction model to accurately predict XCO2. The prediction results will be used to guide the formulation of carbon emission reduction policies and measures.
[0081] Figure 1 The satellite XCO2 prediction method based on multi-source data fusion of the present invention is demonstrated as follows:
[0082] Step 1, first start with the variable collection part, the target variable is the XCO2 data obtained by the satellite, which is the core goal. In order to enhance the predictive ability of the model, a series of external factors are also introduced in the invention, including meteorological data, environmental factors, human factors and social economy. External factors provide rich background information for the model, which helps to capture the multiple factors affecting the change of XCO2. In addition, the present invention also constructs spatial features as model input, which emphasizes the importance of geographical location to XCO2 prediction. Latitude and longitude provide specific geographical coordinates, while spatial average is used to describe the average XCO2 concentration in a specific area, so as to analyze local and regional changes. External factors and spatial features together constitute auxiliary variables as part of the model input.
[0083] Step 2, processing the data collected in step 1;
[0084] The data processing stage shows the unified information of all processed data, ensuring the meticulousness and accuracy of the data. Since data from different sources have different spatiotemporal resolutions, bilinear interpolation is used to unify all features to the same spatiotemporal resolution (temporal resolution is 16 days, spatial resolution is 0.05°×0.05°). Secondly, all input data are normalized, and Min-MaxNormalization is used to scale each variable to the [0, 1] interval, which can speed up the model convergence and prevent feature weight imbalance.
[0085] Step 3, input the processed data into the prediction model and output the prediction result of XCO2 concentration;
[0086] The Model for Forecasting is the core of the invention. The model is trained using the XCO2 time series and auxiliary variables. The Attention Module is introduced to enable the model to focus on important features and improve the prediction effect. The model also includes downsampling and upsampling steps to process the periodicity information of the sequence. The final Prediction Module is responsible for outputting the prediction results of XCO2 concentration.
[0087] Step 4, verify the prediction model;
[0088] Finally, the validation part, in order to ensure the validity and reliability of the model, the TCCON site was used for validation, specifically located in Hefei, China (117.17° E, 31.9° N). The validation methods include ablation experiments, testing the robustness of the model by removing certain variables; comparative experiments, demonstrating the superiority of the invented model by comparing it with the existing time series prediction model; and sample, time and space-based validation, comparing data from different sample data, time periods and spatial regions, to ensure the generalization of the model under various conditions.
[0089] In summary, Figure 1 A systematic XCO2 prediction process framework is presented, covering the complete process from data collection, processing to model building and verification, taking into full account a variety of influencing factors, aiming to achieve accurate prediction of XCO2. In this way, complete XCO2 data can be obtained, which can provide data support for effectively reflecting the spatiotemporal variation characteristics of carbon emissions.
[0090] Figure 2 The data processing framework is shown, including key elements such as data source, format, processing method, and its spatial and temporal characteristics. The following is a detailed introduction to each part of the figure:
[0091] Step 1, collect multi-source data in the region (27°-35°N, 115°-123°E), which provide the basis for the prediction of XCO2. The first is satellite XCO2 data, which comes from the Orbiting Carbon Observatory-2 (OCO-2) satellite launched by the United States in 2014. The invention uses a constructed high-precision complete data set, which has been strictly screened to ensure that the data with the quality flag (xco2_quality_flag=0) is retained to provide the most accurate XCO2 observation information. The XCO2 data used to verify the model comes from the TCCON Hefei station.
[0092] The second is external factors. Existing research shows that meteorological data, vegetation conditions, human emissions, etc. have a significant impact on the atmospheric carbon dioxide concentration.
[0093] Meteorological data include relative humidity (RH), temperature (TEM), etc. These data are from the fifth generation of global climate and weather reanalysis dataset (ERA5) of the European Centre for Medium-Range Weather Forecasts (ECMWF). The data format is nc, and in particular, the data at 13:00 and 14:00 Beijing time every day are selected to match the transit time window of the carbon satellite, and further processed into daily averages.
[0094] Normalized Difference Vegetation Index (NDVI) is used to measure vegetation coverage and health, and can provide information about vegetation growth and distribution. The present invention uses the MOD13Q1 vegetation product from the Moderate Resolution Imaging Spectroradiometer (MODIS), which is saved in nc format. In addition, the present invention uses the SRTM3 version of the SRTM (Shuttle Radar Topography Mission) dataset, which is in Tiff format and effectively depicts the surface morphology.
[0095] The land cover data used in the present invention comes from the China Land Cover Dataset (CLCD), which is saved in Tiff format and reflects the distribution of natural ecology and human activities in detail.
[0096] The increase in atmospheric carbon dioxide concentration is mainly due to human activities, especially the burning of fossil fuels. Therefore, two key carbon emission datasets are also selected in the input data: Emissions Database for Global Atmospheric Research (EDGAR) and The Open-Data Inventory for Anthropogenic Carbon Dioxide (ODIAC). The data format of this dataset is nc, which is crucial for understanding the changes in XCO2.
[0097] Socioeconomic data involves socioeconomic factors such as GDP, population density, and energy consumption, which may affect anthropogenic emissions and ecosystems. The data is derived from the statistical yearbook data of cities in the study area and is saved in csv format during the collection process.
[0098] Detailed information on all datasets and their corresponding spatial resolution, temporal resolution, and data sources is shown in the table below.
[0099] As can be seen from the table content, each data source is saved in a different format, so corresponding processing and conversion operations are required. For satellite XCO2, NDVI, meteorological data and human factors stored in nc format, the export method is the NetCDF library in Python. Tiff format is commonly used in geographic information systems (GIS) and is suitable for storing raster data. Therefore, for land use data and elevation data, the GDAL geospatial data processing library in Python is used for processing. Finally, all data in different formats will be converted and saved in csv file format for use as input to the model.
[0100] Table 1 Dataset details
[0101]
[0102] In order to ensure the accuracy and quality of external auxiliary data, the missing values in the data are first processed. For missing values, the present invention adopts the Kriging interpolation method, which can provide accurate interpolation results while maintaining spatial structural characteristics. In addition, in order to improve the reliability and stability of the data, the local outlier factor (LOF) method is used to detect and remove outliers in the data. The LOF method effectively identifies and removes noise and outliers by calculating the difference between the local density of the data point and the density of its neighbors, thereby ensuring the data quality of subsequent model training.
[0103] Finally, bilinear interpolation was used for time-space matching, and the time dimension of all data was unified into 16 days and the spatial dimension was unified into 0.05°×0.05°. In this way, a complete multi-source time series dataset was constructed.
[0104] Figure 3 The overall structure of the prediction model in the present invention is shown. In many studies, the XCO2 time series is found to have significant periodicity and seasonal changes. Based on this feature, the present invention proposes a hybrid prediction model - HATSNet (Hybrid Attention Temporal Sequence Network). In the design process of the HATSNet model, the periodic characteristics in the XCO2 time series are fully considered and utilized, that is, the XCO2 time series has a known period. .
[0105] The purpose of this invention is to use the entire multivariate historical data to predict future results and to find a suitable function for accurate prediction through learning, i.e. ,in Indicates that the target variable is in the past time steps of data, with Features (here , represents XCO2), Indicates that the auxiliary variable is in the past time steps of data, with Features, is the target variable to be predicted in the future The predicted value for each time step.
[0106] The HATSNet model combines the characteristics of time series data and spatial information, and has efficient computing performance and high prediction accuracy. Specifically, the HATSNet model processes input data as follows:
[0107] Step 1: The HATSNet model first receives the target variable (XCO2 concentration) and auxiliary variables (such as meteorological data, vegetation data, longitude and latitude, etc.). The input data is first normalized by MinMaxScaler to ensure that all features have the same scale. For a given dataset , where each data point belong , through this formula, all data points will be linearly mapped to the interval [0, 1]. The specific operation formula is as follows:
[0108]
[0109] in, is the normalized data point, and are the smallest and largest values in the data set, respectively.
[0110] Step 2: In order to eliminate the distribution bias of the sequence and make the data in different time periods have a more consistent distribution, the target feature By time dimension For each sample, , normalized according to the dimension of the time step. Then, the calculated mean is used to normalize the data at each time step. The normalized result is recorded as , whose dimensions remain unchanged, only the mean shift in the time dimension is removed and added back after the model output.
[0111] In addition, in order to capture local dependencies in time series and mitigate the impact of outliers, a one-dimensional convolution is used. Sliding aggregation is performed. This convolution operation uses zero padding and sets the convolution kernel size to ,in is the sequence period of XCO2. Local features are extracted by weighted averaging the data of surrounding time points. At the same time, the convolution result is combined with the input sequence Perform residual connection to retain the original information and avoid excessive transformation, thereby improving the accuracy and robustness of the prediction and obtaining the output result .
[0112] Step 3: Attention Mechanism is a core technology in deep learning, which aims to improve the computational efficiency and prediction accuracy of the model by adaptively focusing on important information in the input data.
[0113] In the XCO2 prediction task, the model needs to integrate information from multiple data sources, so how to effectively process these different data becomes particularly critical. To solve this problem, after the convolution layer, the model uses an improved hybrid attention mechanism that can adaptively adjust feature weights globally, giving the model flexible feature selection capabilities.
[0114] Step 4, as mentioned above, the XCO2 time series is found to have significant periodic and seasonal variations. Therefore, the XCO2 time series can be viewed as a periodic component. and trend component Therefore, the core object of the forecast is transformed into the forecast of future trend components.
[0115] Based on the decomposition of the above periodic characteristics, the present invention is based on the period Downsample the time series to extract multiple subsequences. On this basis, use the Temporal Convolutional Network (TCN) to independently predict each subsequence. After the TCN module, the future value prediction of each subsequence is obtained. Finally, the prediction results of all subsequences are merged and the prediction value of the final prediction step is obtained through periodic upsampling operations.
[0116] Figure 4 Shown Figure 3 The structure diagram of the attention mechanism in the model structure diagram. This module enhances the reusability of features through channel-level feature connections, thereby improving the expressiveness of the model. Its innovation lies in the introduction of a lightweight channel attention mechanism and the combination of a multi-head cross attention mechanism, which enables the model to flexibly integrate auxiliary feature information during the prediction process. Although these auxiliary features do not directly participate in the final prediction results, they significantly enhance the representation of the target variable.
[0117] The module consists of two parts: the first part is the channel attention mechanism applied to auxiliary features, and the second part is the cross attention mechanism, which is used to fuse the information of target features and auxiliary features. In order to better understand the technical solution of the present invention, the following will describe these two parts and their working principles in detail.
[0118] 1) Channel-Attention: Efficient Channel Attention (ECA) is a lightweight channel attention mechanism that adaptively adjusts the weights of each channel in the input features to enhance the model's attention to key features, thereby improving the model's expressiveness. Specifically, the Efficient Channel Attention (ECA) mechanism first focuses on the auxiliary features of the input. Perform the transpose operation to obtain , to meet the input requirements of pooling and convolution operations. Then, global average pooling (GAP) is applied to each channel to calculate the average value of each channel in the time series dimension. Specifically, for each channel , the global average pooling operation can be expressed by the following formula:
[0119]
[0120] in, express No. The channel in The value of the time step, c ∈ [ 1 , C 2 ] , Indicates The global average of the channel in the time dimension. Through the global average pooling GAP operation, the input feature is compressed into a vector with global time characteristics , that is, global information is extracted in the time dimension.
[0121] In the Efficient Channel Attention (ECA) mechanism, a one-dimensional convolution operation is applied to the channel dimension to calculate the weight coefficient of each channel to dynamically adjust the weight of the auxiliary features according to the importance of the feature. The convolution kernel dynamically generates the attention weight of each channel by adaptively modeling the feature importance of different channels. The size of the convolution kernel is , while using Padding is performed to ensure that the channel dimensions of the input and output are consistent. To optimize channel feature selection, the Efficient Channel Attention (ECA) module adaptively adjusts the size of the convolution kernel based on the input data. .in and is a learnable parameter used to control the size of the convolution kernel. Indicates distance The nearest odd number. At the same time, the convolution kernel size is constrained to be an odd number to ensure symmetry.
[0122] After the convolution operation, the generated channel weights are processed by the Sigmoid activation function to map the weight values to the range of [0, 1]. Then, through element-by-element multiplication, the channel features with larger weights are enhanced, while the channel features with smaller weights are suppressed, thereby highlighting the important features that the model needs to pay attention to. Finally, the channel-weighted feature tensor is rearranged back to its original shape to obtain the output of the Efficient Channel Attention (ECA) module. .
[0123] 2) Cross-Attention: In order to better capture the interaction between the target variable and the auxiliary variable, we first and Perform dimension mapping and expand its feature space to obtain and , is the embedding dimension. and ,By constructing a multi-head cross-attention mechanism, an effective information interaction is established between the target variable and the auxiliary variables.
[0124] Specifically, the mapped target feature sequence It is used as a query vector (Query), which indicates the information you want to obtain from the auxiliary variables; and the mapped auxiliary feature sequence As the key vector (Key) and value vector (Value), it provides useful context information for the target variable. These mapping operations can be expressed by the following formula:
[0125]
[0126]
[0127]
[0128] in, , and are learnable linear transformation matrices that map target variables and auxiliary variables to query, key, and value spaces, respectively. is the embedding dimension. Through these mappings, the target variable and the auxiliary variable are aligned in the same high-dimensional space, preparing for the subsequent cross-attention calculation.
[0129] At the same time, this application also uses a multi-head attention mechanism to parallelly calculate the outputs of multiple attention heads. ,key Sum It will be divided into multiple attention heads, and the dimension of each attention head is , here is the number of attention heads, so we get a new representation .
[0130] For each attention head, an attention score is calculated independently and a weighted output is generated. The attention score is calculated by querying and key It is obtained by the dot product between , which represents the correlation between the query and each key. In order to convert the attention score into a probability distribution, the calculated attention score is subjected to a softmax operation to obtain the attention weights. These weights are used to weight the value matrix, and the normalized attention weights are used to weight the value Perform weighted summation to get the output of each head. The specific formula is as follows:
[0131] O head = [ Softmax( Q ′ ⋅ ( K ′ ) Τ d h ) ] ⋅ V ′
[0132] in, is the transpose of the key, It is a scaling factor that prevents the dot product result from being too large or too small, helping the stability of the training process.
[0133] In the multi-head attention mechanism, each attention head calculates a weighted sum (i.e. ). Secondly, the outputs of all attention heads are concatenated together to form the final representation. In order to convert the concatenated high-dimensional representation back to the original dimension of the target variable, the concatenated output needs to be mapped back to the original dimension through a linear layer. Finally, It is the final output after linear transformation and represents the target variable at each time step (including auxiliary variable information).
[0134] Figure 5 Shown Figure 3 The structure diagram of the prediction module in the prediction model structure diagram is the core content of the entire invention. By introducing the above-mentioned hybrid attention mechanism, the target variable can be effectively integrated. With auxiliary variables , thus obtaining a representation containing rich information (in ). Therefore, its input can be regarded as a known period The univariate time series .
[0135] The sequence can be viewed as consisting of periodic components and trend component Composition, that is Therefore, the core object of forecasting becomes the prediction of future trend components.
[0136] 1) Model Atructure: Based on the decomposition of the above periodic characteristics, firstly Downsample the time series and extract multiple subsequences. Each subsequence The complete periodic information is included to ensure that the periodic pattern can be fully learned and utilized. On this basis, the improved time domain convolutional network (TCN) model is used to predict each subsequence independently.
[0137] The periodic components in each subsequence are implicit in the input data, and the task of the improved time domain convolutional network TCN is to learn the trend changes within the period from these periodic subsequences. After the improved time domain convolutional network TCN prediction, the future value prediction of each subsequence is obtained. Finally, the prediction results of all subsequences will be merged and mapped to the target prediction step through periodic upsampling operations. , thus generating the final prediction result .
[0138] 2) TCN: Since each subsequence contains periodic information, TCN can effectively model these subsequences as local time series. In each subsequence, the periodic component is already embedded in the data, so the task of TCN is to learn the long-term trend changes in each subsequence and focus on capturing the temporal dependencies therein.
[0139] The present invention proposes an improved TCN architecture, which consists of 4 TCN_Blocks, each of which includes a dilated causal convolution (Dilated Gausal Conv), a layer normalization (LayerNorm), an activation function (ReLU), Dropout, and an optional residual connection.
[0140] The classic TCN uses causal convolution instead of traditional one-dimensional convolution to ensure that the model does not violate the time order when making time series predictions, that is, the output at the current moment only depends on the current and previous inputs and will not be affected by future information. , given the convolution kernel size is , the operation of causal convolution can be expressed as:
[0141]
[0142] in, It is at the moment The output at is the convolution kernel weights, It's time Input.
[0143] In order to ensure the causality of the model and maintain the length of the time series, the input data needs to be properly padded. In the TCN model, the padding and expansion factor and the convolution kernel size Therefore, the dilation factor and the convolution kernel size jointly determine the amount of padding, ensuring that the convolution operation can adapt to the dependencies of longer time series.
[0144] TCN uses a dilated convolution mechanism by inserting intervals (given by the dilation factor) between the convolution kernel elements. control), expanding the receptive field to capture long-term temporal dependencies. Compared with traditional convolution, it does not increase parameters or network depth by adjusting The value can flexibly expand the interval sampling range of the input data, thereby enhancing the modeling ability of historical trends. Specifically, for a given expansion factor and the convolution kernel size , the calculation formula of the dilated convolution can be expressed as:
[0145]
[0146] in, It is at the moment The output at is the convolution kernel weights, Expressing historical moments The causal access, is the dilation factor, which controls the sampling interval between convolution kernel elements.
[0147] In addition, by adjusting , which can flexibly expand the coverage of the receptive field. TCN stacks multiple layers of dilated convolution to increase the receptive field exponentially, thereby efficiently capturing long-term dependencies. Dilation factor of the layer , then Receptive field of the layer It can be expressed by the following formula:
[0148]
[0149] in, is the size of the receptive field of this layer, is the size of the upper layer receptive field, the initial layer receptive field .
[0150] Batch Normalization is usually used in traditional TCN implementations to speed up training. However, the performance of batch normalization is highly dependent on the choice of batch size, which can easily lead to unstable training process. To address this problem, this model uses layer normalization instead of batch normalization. Layer normalization has significant advantages in time series data modeling: it independently normalizes along the feature dimension at each time step without relying on the statistical information of other samples in the batch. For each time step , the calculation of layer normalization can be expressed by the following formula:
[0151]
[0152] in, is the normalized output, is the time step The original input at is the mean value over all time steps, is the standard deviation of all time steps, and are learnable scaling factors and bias terms used to adjust the normalized output.
[0153] After layer normalization, the model usually uses an activation function to add nonlinear transformations so that the network can learn more complex patterns. In this model, the ReLU (Rectified Linear Unit) activation function is used, which is expressed as the following formula:
[0154]
[0155] In order to prevent overfitting and improve the generalization ability of the model, this model introduces a Dropout layer after each convolution block to randomly discard neuron outputs. This method further enhances the robustness of the model and reduces overfitting of specific features.
[0156] In traditional TCN architectures, residual connections are often relied upon to help alleviate the vanishing gradient problem in deep networks. However, as the depth of the network increases, the effectiveness of residual connections may be affected when the number of input and output channels is inconsistent. To address this issue, this model innovatively uses 1×1 convolutions to align the number of channels, ensuring that residual connections can be applied smoothly between different layers. This design helps solve the problem of residual connections when the dimensions are inconsistent, thereby ensuring unimpeded flow of information between network layers.
[0157] After predicting each subsequence through the improved TCN module, the model Upsample, and then map the prediction results to the final prediction step size through the fully connected layer Finally, add the value calculated in step 2 The mean of the time dimension is taken and the maximum and minimum inverse normalization is performed to obtain the final prediction output of XCO2 .
[0158] The data division and parameter setting in the present invention are as follows:
[0159] The data covers the period from 2016 to 2022 and is scientifically divided into two subsets: the data from 2016 to 2020 is used to construct the training set and validation set, with the validation set accounting for 20% for model training and tuning; the data from 2021 and 2022 is used as the test set to evaluate the predictive performance of the model. During the model training process, the learning rate is set to 0.001 to ensure that the model can stably and effectively learn the features in the data. In addition, the training cycle (epochs) is set to 50, the batch size (batchsize) is 256, and the Adam optimization algorithm is used to optimize the model parameters with its efficient gradient descent ability and good hyperparameter adaptability. The cycle of XCO2 is set to 23, indicating that each cycle is one year. The embedding dimension of the hybrid attention module is set to 64, and 4 attention heads are selected. The parameters of the TCN module are set to [25, 128, 64, 32], and the convolution kernel size is 3.
[0160] In order to verify the superiority of the model of the present invention, multiple evaluation indicators (mean absolute error MAE, root mean square error RMSE, mean absolute percentage error MAPE and determination coefficient R 2 ) to comprehensively evaluate the performance of the model. The evaluation index formula is as follows:
[0161]
[0162]
[0163]
[0164]
[0165]
[0166] in, is the true value, is the predicted value, is the mean of the true values, is the number of samples. The complexity of the training process makes the absolute value function have better robustness when dealing with outliers, but its derivative is discontinuous when the model output is close to the true value. In contrast, the square function exhibits different characteristics in this regard. In order to address the limitations of these two functions, the present invention selects the Log-cosh loss function.
[0167] In addition, this study also conducted ablation experiments and compared time to verify the superiority and generalization of the model from multiple aspects.
[0168] By implementing the present invention, users can alleviate the key engineering problems existing in existing high-precision and spatiotemporally continuous XCO2 data acquisition and carbon emission monitoring technologies:
[0169] [1] Difficulties in acquiring and integrating multi-source data: In the context of a lack of inventions for acquiring spatiotemporal continuous XCO2 data using deep learning algorithms, this paper conducts deep learning prediction tasks based on the high-precision XCO2 data that has been constructed by predecessors. At the same time, this study additionally collects external factors related to human activities. In addition, in view of the spatiotemporal scale differences and missing value problems between satellite XCO2 data and multi-source heterogeneous data such as meteorology, vegetation, and socio-economics, a unified data alignment and feature enhancement method is proposed. Through dynamic spatiotemporal interpolation, multimodal feature encoding and other technologies, collaborative modeling of meteorological data, human factors, etc. with XCO2 data is achieved, providing highly robust input for deep learning models;
[0170] [2] Challenges of lightweight and high-precision modeling: Based on the seasonal and trend characteristics of XCO2 data, a lightweight spatiotemporal attention mechanism and multi-task learning framework are designed to integrate multi-source auxiliary variables, thereby reducing the model's computational overhead while improving prediction accuracy, enabling it to efficiently process large-scale spatiotemporal data and support real-time XCO2 assessment, capturing the spatiotemporal trend of carbon emissions. This significantly improves the spatiotemporal continuity of near-ground carbon emissions inversion results and enhances the model's ability to analyze complex human and natural factors.
Claims
1. Satellite XCO2 prediction method based on multi-source data fusion, characterized in that: The steps include: Step 1: Build a dataset including historical XCO2 data and auxiliary variables ; Step 2: Build a prediction model for predicting XCO2 data, where the input data of the prediction model is historical XCO2 data. and auxiliary variables ; The prediction model is expressed as: ; in, Indicates historical XCO2 data in the past time steps of data, with Features, Indicates that the auxiliary variable is in the past time steps of data, with Features, It is the predicted XCO2 data in the future The predicted value of time steps; Step 3: Use the data set to train the prediction model to obtain a trained prediction model, and use the trained prediction model to predict the XCO2 data.
2. The satellite XCO2 prediction method based on multi-source data fusion according to claim 1 is characterized in that: The auxiliary variables include meteorological data, environmental factors, human factors, social economy, and spatial characteristics; The meteorological data include relative humidity RH, temperature TEM, U-shaped wind and V-shaped wind; The environmental factors include vegetation cover, ground elevation, and land cover; The vegetation coverage is specifically the Normalized Difference Vegetation Index NDVI; The spatial characteristics include latitude, longitude and spatial average, and the spatial average refers to the average XCO2 concentration in the region.
3. The satellite XCO2 prediction method based on multi-source data fusion according to claim 1 is characterized in that: The process of the prediction model processing the input data is as follows: Step 2.1, historical XCO2 data By normalizing the mean value according to the time dimension, we get ; Step 2.2, use one-dimensional convolution to Perform sliding aggregation, extract local features by weighted averaging the data of surrounding time points, and combine the convolution result with the input sequence Perform residual connection to get the output result , as information of target features; The one-dimensional convolution uses zero padding, and the convolution kernel size of the one-dimensional convolution is ; Step 2.3, for auxiliary variables The channel attention ECA mechanism is used to generate channel weights through global average pooling and adaptive one-dimensional convolution. The channel-weighted feature tensor is rearranged back to its original shape to obtain the output result of the ECA module. The weighted auxiliary feature ; Step 2.4: Output the result of step 2.2 With the weighted auxiliary features Perform cross attention fusion to obtain the target variable containing auxiliary variable information ; Step 2.5, for the target variable containing auxiliary variable information Make predictions and get the final prediction results .
4. The satellite XCO2 prediction method based on multi-source data fusion according to claim 3 is characterized in that: The step 2.3 is as follows: Use channel attention mechanism to input auxiliary features Perform the transpose operation to obtain , then apply global average pooling GAP to each channel to calculate the average value of each channel in the time series dimension; A one-dimensional convolution operation is applied to the channel dimension to calculate the weight coefficient of each channel to dynamically adjust the weight of the auxiliary features according to the importance of the features. The convolution kernel dynamically generates the attention weight of each channel by adaptively modeling the feature importance of different channels; The size of the convolution kernel is , while using Padding to ensure that the channel dimensions of input and output are consistent; ; in and is a learnable parameter used to control the size of the convolution kernel. Indicates distance The nearest odd number; at the same time, is constrained to be an odd number to ensure symmetry; After the convolution operation, the generated channel weights are processed by the Sigmoid activation function to map the weight values to the range of [0, 1]. Then, through element-by-element multiplication, the channel features with larger weights are enhanced, while the channel features with smaller weights are suppressed, thereby highlighting the important features that the model needs to pay attention to. Finally, the channel-weighted feature tensor is rearranged back to its original shape to obtain the output result of the ECA module. .
5. The satellite XCO2 prediction method based on multi-source data fusion according to claim 3 is characterized in that: The step 2.4 is as follows: right and Perform dimension mapping to expand its feature space. Get the mapped target features and auxiliary features , is the embedding dimension; The mapped target feature sequence Used as query vector Query, the auxiliary feature sequence after mapping Then the key vector Key and the value vector Value are expressed as follows: ; ; ; in, , and are learnable linear transformation matrices used to map target variables and auxiliary variables to query, key, and value spaces, respectively; The query ,key Sum Divided into multiple attention heads, the dimension of each attention head is : ; in, is the number of attention heads, and the new representation is obtained ; Calculate the attention score of each attention head and convert it into a probability distribution to get the attention weight. Perform weighted summation to get the output of each head, expressed as: ; in, is the transpose of the key, is the factor used for scaling; The output of each attention head is concatenated and then mapped through a linear layer to obtain the target variable containing auxiliary variable information. .
6. The satellite XCO2 prediction method based on multi-source data fusion according to claim 4 is characterized in that: Step 2.5 is as follows: The target variable containing the auxiliary variable information Considered as a known cycle The univariate time series , for the target variable containing auxiliary variable information , downsampled into multiple subsequences according to a preset period, each subsequence contains complete period information; The time domain convolutional network (TCN) model is used to independently predict each subsequence, obtain the prediction results for each subsequence, merge the prediction results of the subsequences, and map the prediction results to the target prediction step size through periodic upsampling operations. , thus generating the final prediction result .
7. The satellite XCO2 prediction method based on multi-source data fusion according to claim 6 is characterized in that: The time domain convolutional network TCN model includes four cascaded TCN_Blocks; Each TCN_Block has the same structure, including: dilated causal convolution, layer normalization, activation function, Dropout, and optional residual connection.
8. The satellite XCO2 prediction method based on multi-source data fusion according to claim 7 is characterized in that: In the TCN_Block, the convolution kernel elements are inserted with spacing, for a given dilation factor and the convolution kernel size , the calculation formula of the dilated convolution can be expressed as: ; in, It is at the moment The output at is the convolution kernel weights, Expressing historical moments The causal access, is the dilation factor, which controls the sampling interval between convolution kernel elements; For each time step , the calculation of layer normalization is expressed by the following formula: ; in, is the normalized output, is the time step The original input at is the mean value over all time steps, is the standard deviation of all time steps, and is a learnable scaling factor and bias term used to adjust the normalized output; After layer normalization, the ReLU activation function is used to add nonlinear transformation, and the Dropout layer is introduced after the activation function; A 1×1 convolution is added to the residual connection.
9. Satellite XCO2 prediction system based on multi-source data fusion, characterized by: include: Data processing module, used to clean multi-source data, fill missing values, use bilinear interpolation to unify spatiotemporal resolution, and normalize; the multi-source data includes historical XCO2 data and auxiliary variables ; Prediction model, historical XCO2 data after data processing and auxiliary variables Conduct XCO2 forecasts; The prediction model includes a hybrid attention module and a prediction module; The hybrid attention module integrates the channel attention ECA mechanism and the cross attention mechanism, and uses the channel attention ECA mechanism to Perform feature enhancement; use the cross attention mechanism to integrate historical XCO2 data With the enhanced auxiliary variables Perform feature fusion to obtain the target variable containing auxiliary variable information ; The prediction module predicts the target variable containing auxiliary variable information Downsampling is performed according to a preset period to generate multiple subsequences, and a time domain convolutional network (TCN) is used to independently predict each subsequence. The prediction results of the subsequences are merged, and periodic upsampling is performed, and the results are mapped to the final XCO2 prediction value through a fully connected layer.
Citation Information
Patent Citations
Method for improving typhoon trajectory prediction
CN112527860A
Carbon emission analysis method of multi-scale double attention guidance fusion network model
CN117496179A
Debris flow disaster susceptibility assessment method based on multi-source data
CN117636071A
Photovoltaic power generation day-ahead prediction method based on causal inference and multi-scale feature fusion
CN119578669A
Upsampling of compressed financial time-series data using a jointly trained Vector Quantized Variational Autoencoder neural network
US12229679B1
Cited By
Multi-source data fusion method and device for road carbon emission monitoring
CN120372557A
Deep learning CO2 near-real-time inversion method and system fusing multi-scale features
CN121543465A
Deep learning co2 near real-time inversion method and system fusing multi-scale features
CN121543465B
Regional XCO2 prediction method based on multi-source feature uncertainty modeling
CN122451813A
A Regional XCO2 Prediction Method Based on Multi-Source Feature Uncertainty Modeling
CN122451813B