Atmospheric methane concentration prediction method based on multi-source data fusion
By fusing multi-source data and using a hybrid deep learning model, combined with meteorological and socio-economic data, the problem of single data source and high computational resource consumption in existing technologies for atmospheric methane concentration prediction has been solved, achieving efficient and accurate methane concentration prediction.
Patent Information
- Application Number
- CN202510495243.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing methods for predicting atmospheric methane concentration mainly rely on a single data source, which cannot effectively capture concentration fluctuations under the complex interaction of environment and society, resulting in reduced long-term prediction reliability and high computational resource requirements.
A multi-source data fusion method is adopted, combining meteorological elements and socio-economic data. A prediction model is generated by training a hybrid deep learning model, and methane concentration is predicted using climate and socio-economic driving factors. The prediction steps are simplified by downscaling and feature extraction.
It improves forecasting capabilities under climate and socioeconomic change conditions, reduces forecasting costs and complexity, enhances forecasting accuracy and model efficiency, and ensures data consistency and robustness.
Smart Images

Figure CN120629471B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental monitoring and prediction technology, specifically to a method for predicting atmospheric methane concentration based on multi-source data fusion. Background Technology
[0002] Methane is a greenhouse gas. An increase in atmospheric methane concentration contributes far more to increased atmospheric temperature than carbon dioxide at the same concentration. Atmospheric methane concentration is driven by a combination of biogeochemical processes and anthropogenic disturbances. Among natural factors, meteorological elements such as precipitation and temperature exhibit highly nonlinear characteristics in regulating methane emissions. Precipitation directly affects the source-sink balance of methane by altering soil moisture conditions. Furthermore, socioeconomic activities directly change the intensity of anthropogenic methane emissions by altering the area of vegetation cover and the intensity of resource utilization.
[0003] Currently, existing methods for predicting atmospheric methane concentration mainly include numerical climate models (CMIP6) and statistical models. The CMIP6 model requires simulating the interactions of multiple subsystems, including the atmosphere, ocean, land, and cryosphere, resulting in enormous computational demands. High-resolution and long-scale simulations further increase the demand for computational resources. Existing statistical methods, such as the Autoregressive Integral Moving Average (ARIMA) model and the Adaptive Threshold Autoregressive model, use historical time-series methane data as input to predict future methane concentration trends. However, these methods rely on a single data source, failing to capture concentration fluctuations under complex environmental and social interactions, thus reducing the reliability of long-term predictions. Summary of the Invention
[0004] The purpose of this invention is to provide an atmospheric methane concentration prediction method based on multi-source data fusion. By fusing multi-source data based on meteorological elements and socio-economic data, it overcomes the previous reliance on a single data source for prediction, thereby improving the prediction capability under conditions of simultaneous climate and socio-economic changes. In addition, machine learning algorithms simplify the complexity of methane prediction steps and processes, which helps to reduce prediction costs.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A method for predicting atmospheric methane concentration based on multi-source data fusion includes:
[0007] Acquire climate factors, socioeconomic indicators, and methane data for the target region, and preprocess the acquired climate factors and methane data to obtain preprocessed results;
[0008] Screen climate-driven factors and socio-economic-driven factors, and determine target climate-driven factors and target socio-economic-driven factors;
[0009] The hybrid deep learning model is trained based on the target climate driving factors and the target socio-economic driving factors, so that the trained hybrid deep learning model can generate the target prediction model.
[0010] Input climate and socioeconomic driving factor data under future scenarios into the prediction model to implement methane prediction and generate prediction results. Perform inverse PCA transformation and inverse standardization on the prediction data to obtain the predicted methane concentration data.
[0011] As a further aspect of the present invention: the method for acquiring climate factors, socioeconomic indicators, and methane data of the target area, and preprocessing the acquired climate factors and methane data to obtain preprocessing results includes:
[0012] The acquired climate factor and methane data are geospatially segmented to retain the climate factor and methane data for the target area;
[0013] The climate factors of the target area are downscaled to obtain downscaled climate factors.
[0014] The dimensionality of the downscaled climate factors and the acquired methane data was reduced.
[0015] As a further aspect of the present invention: the method for reducing the dimensionality of the downscaled climate factors and the acquired methane data includes:
[0016] The climate factors and methane data of the target area were unified to international standard units to obtain climate factors and methane data in standard units;
[0017] The missing values of climate factors and methane data in standard units were filled using the sampling strip method to obtain the filled climate factors and methane data;
[0018] The resolution of the filled climate factors was unified by using bicubic interpolation to obtain climate factors with the same resolution.
[0019] The K-nearest neighbor algorithm is used to correct the bias of climate factors with the same resolution to obtain the corrected climate factors;
[0020] The corrected climate factors, socioeconomic indicators, and methane data were standardized to obtain standardized climate factors, socioeconomic indicators, and methane data.
[0021] Principal component analysis (PCA) was used to reduce the dimensionality of the standard climate factors and methane data, resulting in dimensionality-reduced climate factors and methane data.
[0022] As a further aspect of the present invention: the method of using principal component analysis (PCA) to reduce the dimensionality of standard climate factors and methane data to obtain dimensionality-reduced climate factors and methane data further includes:
[0023] Principal components retaining 95% variance were extracted from the standard climate factors and methane data to obtain low-latitude climate factors and low-latitude methane data.
[0024] By extracting principal components that retain 95% of the variance from standard climate factors and methane data, it is possible to decompose and simplify the standard climate factors and methane data, reveal the temporal evolution and spatial distribution structure of standard climate factors, socioeconomic indicators and methane data, thereby obtaining low-latitude data.
[0025] As a further aspect of the present invention: the method for screening climate driving factors and socio-economic driving factors to determine the climate driving factors and socio-economic driving factors affecting the target area includes:
[0026] Based on low-latitude climate factors and low-latitude methane data, the monthly climate factors in the low-latitude climate factors are determined as the basic climate factor indicators.
[0027] Collect socioeconomic indicator data of the target area to construct basic socioeconomic parameters. The socioeconomic indicator data includes economic scale parameters, population parameters and environmental statistics parameters of the target area.
[0028] The Pearson correlation coefficient method was used to calculate the correlation between basic climate factor indicators and basic socioeconomic parameters, and to extract core climate parameters and core socioeconomic parameters.
[0029] Core climate parameters and core socioeconomic parameters are screened according to preset threshold rules, and the screened core climate parameters and core socioeconomic parameters are determined as the climate driving factors and socioeconomic driving factors of the target area.
[0030] As a further aspect of the present invention: the method for screening core climate parameters and core socioeconomic parameters according to preset threshold rules, and determining the screened core climate parameters and core socioeconomic parameters as climate driving factors and socioeconomic driving factors of the target area includes:
[0031] Set the preset threshold rule to |r|>0.5;
[0032] Core climate parameters and core socioeconomic parameters are selected to obtain the core climate parameters and core socioeconomic parameters, where |r| is an empirical threshold.
[0033] As a further aspect of the present invention: the method for training a hybrid deep learning model based on target climate driving factors and target socioeconomic driving factors, so that the trained hybrid deep learning model generates a prediction model, includes:
[0034] The training process for the hybrid deep learning model is executed based on low-latitude methane data, climate driving factors, and socio-economic driving factors to obtain the initial training model.
[0035] Based on preset initial values, perform parameter combination analysis on the parameters of the initial training model, output the optimal hyperparameter configuration, establish a high-precision methane concentration prediction model, and complete the generation of the prediction model.
[0036] As a further aspect of the present invention: the hybrid deep learning model includes a climate feature extraction and analysis pathway and a socio-economic feature analysis pathway. The climate feature extraction and analysis pathway includes a two-layer LSTM structure, and the socio-economic feature analysis pathway includes a fully connected layer with several neurons.
[0037] As a further aspect of the present invention: the method for training a hybrid deep learning model based on low-latitude methane data, climate driving factors, and socioeconomic driving factors to obtain an initial training model includes:
[0038] A dual-input architecture is adopted to achieve multimodal fusion, so as to take into account both temporal dependence and static feature association;
[0039] A two-layer LSTM structure is used to extract features hierarchically. The first LSTM layer is used to capture local short-term patterns in the sequence of climate driving factors, while the second LSTM layer further extracts global long-term dependence data of climate driving factors based on the output of the first LSTM layer.
[0040] A fully connected layer is used to map the features of socioeconomic driving factors to an abstract space;
[0041] Regularization methods are used to suppress outlier interference in the data of socioeconomic driving factors;
[0042] Set up splicing layers and integrate the characteristics of climate-driving factors and socio-economic-driving factors through channel splicing;
[0043] The climate-social synergy effect is synthesized by using a fully connected layer ReLU activation function, and a certain proportion of nodes are randomly shielded by a random deactivation layer to simulate the stability under data missing scenarios and improve the robustness of the data.
[0044] Set up an output layer to linearly output the continuous distribution characteristics of the target variable;
[0045] By performing the above steps multiple times, a target prediction model is obtained.
[0046] This hybrid neural network architecture significantly improves model efficiency and prediction accuracy while preserving the characteristics of multimodal data through heterogeneous data collaborative processing and division of labor.
[0047] By employing a two-layer LSTM structure to extract features in layers, the model can learn features at different time scales simultaneously.
[0048] By using a fully connected layer to map the features of socioeconomic drivers to an abstract space, we can avoid the dimensional redundancy caused by directly splicing the original features. By using regularization methods to suppress outlier interference in the data of socioeconomic drivers, we can independently process non-time series data and prevent conflicts with the time series assumptions of climate series.
[0049] By setting up splicing layers and integrating the features of climate-driven and socio-economic-driven factors through channel splicing, it is possible to preserve the decay memory of latent climate states and the independent distribution of social features, thus avoiding feature confusion.
[0050] As a further aspect of the present invention: the method for training a hybrid deep learning model based on target climate driving factors and target socioeconomic driving factors, so that the trained hybrid deep learning model generates a prediction model, further includes:
[0051] Equipped with a 20% random deactivation layer (Dropout layer), it captures the 12-step temporal dependence of the differential features of the principal components of precipitation.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] 1. This invention fuses multi-source data based on climate factor data and socioeconomic data to downscale the climate factors of the target area and unify the climate factors of the target area to the same resolution. At the same time, it reduces the dimensionality of the climate factors and methane data of the target area, simplifies the complexity of the methane prediction steps and process, and helps to reduce prediction costs.
[0054] 2. This invention uses bicubic interpolation to interpolate low-resolution data onto a high-resolution grid, which ensures data spatial consistency. By comparing the downscaled data with the measured data, systematic biases in the data can be identified, improving the accuracy of model predictions. By using the K-nearest neighbor algorithm to correct the biases in the downscaled data, the consistency between the downscaled data and the measured data can be ensured.
[0055] 3. This invention improves model efficiency and prediction accuracy while preserving the characteristics of multimodal data through heterogeneous data collaborative processing and division of labor mechanism. By using a two-layer LSTM structure to extract features in layers, the model can learn features at different time scales at the same time.
[0056] 4. This invention uses a fully connected layer to map the features of socioeconomic driving factors to an abstract space, which avoids the dimensional redundancy caused by directly splicing the original features; by using a regularization method to suppress outlier interference in the data of socioeconomic driving factors, it can independently process non-time series data and prevent conflicts with the time series assumptions of climate series.
[0057] 5. By setting up a splicing layer and integrating the features of climate driving factors and socio-economic driving factors through channel splicing, this invention can preserve the decay memory of latent climate states and the independent distribution of social features, thus avoiding feature confusion. Attached Figure Description
[0058] Figure 1 A flowchart of the atmospheric methane concentration prediction steps based on multi-source data fusion provided by the present invention;
[0059] Figure 2 This is a flowchart illustrating the steps for obtaining preprocessing results in an embodiment of the present invention.
[0060] Figure 3 This is a structural diagram of the hybrid deep learning model of the present invention. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Example:
[0063] Please see Figure 1 In this embodiment of the invention, the atmospheric methane concentration prediction method based on multi-source data fusion includes:
[0064] S1: Obtain climate factors, socioeconomic indicators, and methane data for the target area, and preprocess the obtained climate factors and methane data to obtain preprocessing results;
[0065] S2: Screen climate driving factors and socio-economic driving factors, and determine target climate driving factors and target socio-economic driving factors;
[0066] S3: Train the hybrid deep learning model based on the target climate driving factors and the target socio-economic driving factors, so that the trained hybrid deep learning model can generate the target prediction model.
[0067] S4: Input climate and socioeconomic driving factor data under future scenarios into the prediction model to implement methane prediction and generate prediction results. Perform inverse PCA transformation and inverse standardization on the prediction data to obtain the predicted methane concentration data.
[0068] Preferred, such as Figure 2 As shown, the preprocessing steps for the acquired climate factors and methane data are as follows:
[0069] S11: The missing values of the climate factor and methane data in the standard unit are filled using the sampling strip method to obtain the filled climate factor and methane data;
[0070] S12: The resolution of the filled climate factors is unified by using bicubic interpolation to obtain climate factors with the same resolution.
[0071] S13: The K-nearest neighbor algorithm is used to correct the bias of climate factors with the same resolution to obtain the corrected climate factors;
[0072] S14: Standardize the corrected climate factors, socioeconomic indicators and methane data respectively to obtain standard climate factors, socioeconomic indicators and methane data;
[0073] S15: Principal component analysis (PCA) is used to reduce the dimensionality of the standard climate factors and methane data, resulting in dimensionality-reduced climate factors and methane data.
[0074] Preferably, climate factors for the target region are obtained from the CMIP6 (Coupled Model Intercomparison VI) and ERA5 (Generation V Atmospheric Reanalysis Dataset) databases, methane remote sensing data are obtained from the Copernicus website, and socioeconomic indicators are obtained from statistical yearbooks. The aforementioned data are downloaded from the official data portals of the CMIP6 database, ERA5 database, and Copernicus website and stored in a local structured database.
[0075] The downloaded data is in NetCDF format to ensure compatibility with subsequent processing tools. After the data download is complete, data quality control is performed, including data integrity checks, data consistency checks, and data format conversion, to ensure data integrity and consistency.
[0076] The data and methane data obtained from the CMIP6 and ERA5 databases are segmented using vector files of the target region to obtain CMIP6 data, ERA5 data and methane data, ensuring that the CMIP6 data, ERA5 data and methane data only contain information within the target region.
[0077] The vector file defines the spatial boundary of the study area, and the segmented data will retain only the data points within that boundary.
[0078] Align the time periods and spatial resolutions of the segmented CMIP6 and ERA5 data, and unify the data units to international standard units to prepare for subsequent data preprocessing and analysis.
[0079] At this point, missing values are identified in the segmented CMIP6, ERA5, and methane data, marking the missing time and spatial points in the data.
[0080] The missing values are filled in both time and space using spline interpolation. The interpolation formula (1) is as follows:
[0081] f(t,x,y)=∑ i,j,k c i,j,k ×B i (t)×B j (x)×B k (y) (1)
[0082] Where f(t,x,y) is the interpolation function, B i (t), B j (x) and B k (y) is a basis function, c ij It is a coefficient.
[0083] Because CMIP6 and ERA5 data have different spatial resolutions, it is necessary to unify the data to the same resolution.
[0084] Based on the padded CMIP6 and ERA5 data, a bicubic interpolation method is used to interpolate the low-resolution data onto the high-resolution grid, ensuring spatial consistency of the data.
[0085] By using bicubic interpolation, all data are unified to a spatial resolution of 0.1° × 0.1°, and formula (2) is:
[0086]
[0087] Wherein, the target pixel is P(x,y), and the coordinates of its surrounding pixels are (x,y). i ,y j (where i and j range from 0 to 3), the corresponding pixel value is P(i,j), aij These are the interpolation coefficients, calculated using the values of the surrounding 16 pixels.
[0088] By comparing the downscaled data with the measured data, systematic biases in the data can be identified.
[0089] The K-nearest neighbor algorithm is used to correct the skewness of the downscaled data to ensure consistency between the data and the measured data.
[0090] Select bias-related features as input to the KNN algorithm, calculate the Euclidean distance between the downscaled data points and the measured data points, and select the K nearest neighbors.
[0091] Based on the measured data values of K neighbors, the downscaled data is corrected using formula (3):
[0092]
[0093] Among them, P corrected For the corrected value, P observed (k) represents the measured value of the kth neighbor.
[0094] After obtaining downscaled data using the above downscaling method, the dimensionality of the methane data and the downscaled climate factors is reduced respectively.
[0095] This embodiment uses principal component analysis to reduce dimensionality. First, the data is standardized. The specific formula (4) is as follows:
[0096]
[0097] X scaled Here, M represents the standardized value, and μ represents the standard deviation.
[0098] Then, the three-dimensional data (time × latitude × longitude) is reshaped into a two-dimensional matrix (time × spatial point), and each spatial point is standardized to eliminate dimensional differences.
[0099] Principal component directions were extracted using Singular Value Decomposition (SVD), and the top k principal components with a cumulative variance contribution rate of 95% were selected.
[0100] After projection, the principal component time series and spatial loading matrix are obtained: the time series characterizes the main temporal variation features of evaporation, while the spatial loading reveals the geographical distribution pattern of the corresponding principal component.
[0101] The specific formula (5) is as follows:
[0102] Z = X scaled ·V k (5)
[0103] Z represents the data after dimensionality reduction; V k The matrix composed of the first k principal component directions (extracted by SVD, retaining 95% variance) yields low-latitude climate factors and low-latitude methane data. By extracting principal components with 95% variance from the standard climate factors and methane data respectively, the standard climate factors and methane data can be decomposed and simplified, revealing the temporal evolution and spatial distribution structure of the standard climate factors, socioeconomic indicators and methane data, thereby obtaining low-latitude data.
[0104] Based on low-latitude climate factors and low-latitude methane data, the monthly climate factors in the low-latitude climate factors are determined as the basic climate factor indicators.
[0105] Collect socioeconomic indicator data of the target area to construct basic socioeconomic parameters. The socioeconomic indicator data includes economic scale parameters, population parameters and environmental statistics parameters of the target area.
[0106] The Pearson correlation coefficient method was used to calculate the correlation between basic climate factor indicators and basic socioeconomic parameters, and to extract core climate parameters and core socioeconomic parameters.
[0107] Set the preset threshold rule to |r|>0.5;
[0108] Core climate parameters and core socioeconomic parameters are selected to obtain the core climate parameters and core socioeconomic parameters, where |r| is an empirical threshold.
[0109] The Pearson correlation coefficient was used to quantify the linear correlation between each feature and the principal component of methane concentration. An empirical threshold of |r|>0.5 was set to screen core climate parameters and core socioeconomic parameters. The screened core climate parameters and core socioeconomic parameters were determined as the climate driving factors and socioeconomic driving factors of the target area.
[0110] The climate driving factor for the target area was determined to be rainfall.
[0111] The socioeconomic drivers of the target region are GDP per capita and resident population.
[0112] At this point, the hybrid deep learning model is further trained based on the target climate driving factors and the target socio-economic driving factors, so that the trained hybrid deep learning model can generate a prediction model.
[0113] Preferably, a training process is performed on the hybrid deep learning model based on low-latitude methane data, climate driving factors, and socio-economic driving factors to obtain an initial training model;
[0114] Based on preset initial values, perform parameter combination analysis on the parameters of the initial training model, output the optimal hyperparameter configuration, establish a high-precision methane concentration prediction model, and complete the generation of the prediction model.
[0115] Preferably, the hybrid deep learning model includes a dual-branch input architecture and a feature fusion mechanism;
[0116] The two-branch input architecture includes a first-branch architecture and a second-branch architecture;
[0117] The first branch architecture is used to process climate driving factor features using a two-layer long short-term memory network (LSTM layer), and the second branch architecture is used to process socio-economic driving factor features using a fully connected network (Dense layer).
[0118] The first and second branch architectures merge features through a concatenation layer.
[0119] The hybrid deep learning model was trained multiple times to obtain the initial training model.
[0120] Preferably, a dual-input architecture is used to achieve multimodal fusion, so as to take into account both temporal dependence and static feature association; a two-layer LSTM structure is used to extract features hierarchically. The two-layer LSTM structure is a two-layer long short-term memory network layer, which includes a first LSTM layer and a second LSTM layer. The first LSTM layer is used to capture local short-term patterns in the sequence of climate driving factors, and the second LSTM layer further extracts global long-term dependence data of climate driving factors based on the output of the first LSTM layer.
[0121] A fully connected layer is used to map the features of socioeconomic driving factors to an abstract space; a regularization method is used to suppress outlier interference in the data of socioeconomic driving factors; a splicing layer is set up, and the features of climate driving factors and socioeconomic driving factors are fused by channel splicing.
[0122] The climate-social synergy effect is synthesized by using a fully connected layer ReLU activation function, and a certain proportion of nodes are randomly shielded by a random deactivation layer to simulate the stability under data missing scenarios and improve the robustness of the data.
[0123] Set up an output layer to linearly output the continuous distribution characteristics of the target variable;
[0124] By performing the above steps multiple times, a target prediction model is obtained.
[0125] Preferred, such as Figure 3As shown, the hybrid deep learning model includes a climate feature extraction and analysis pathway and a socio-economic feature analysis pathway. The climate feature extraction and analysis pathway includes a two-layer LSTM structure, and the socio-economic feature analysis pathway includes a fully connected layer with 8 neurons. The climate feature extraction and analysis pathway uses a two-layer LSTM network (32→16 units) and is equipped with 20% of random deactivation layers (Dropout layers). The random deactivation layers (Dropout layers) are used to capture the 12-step temporal dependence of the differential features of the principal components of precipitation.
[0126] The socioeconomic characteristics analysis pathway processes population and GDP per capita indicators through a fully connected layer with 8 neurons and applies L2 regularization (λ = 0.01).
[0127] After the bimodal features are fused by the concatenate layer, they are connected to the 32-neuron joint inference layer, which includes a 30% random deactivation layer and L2 regularization to generate predicted values.
[0128] The model uses the Adam optimizer to minimize the mean squared error. The learning rate of the Adam optimizer is 0.001. An early stopping mechanism is implemented during training, with a patience value of 20 rounds.
[0129] After multiple training iterations, the best training model is selected. By predicting methane in the region, the predicted data is subjected to inverse PCA transformation and inverse standardization to obtain the predicted methane concentration data, thus generating the prediction results.
[0130] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for predicting atmospheric methane concentration based on multi-source data fusion, characterized in that, include: Acquire climate factors, socioeconomic indicators, and methane data for the target region, and preprocess the acquired climate factors and methane data to obtain preprocessed results; The method for acquiring climate factors, socioeconomic indicators, and methane data of the target area, and preprocessing the acquired climate factors and methane data to obtain preprocessing results includes: The acquired climate factors and methane data are geospatially segmented, retaining the climate factors and methane data for the target area; the climate factors in the target area are downscaled to obtain downscaled climate factors; the downscaled climate factors and the acquired methane data are then subjected to dimensionality reduction processing. The method for reducing the dimensionality of the downscaled climate factors and the acquired methane data includes: The climate factors and methane data of the target area were unified to international standard units to obtain climate factors and methane data in standard units; The missing values of climate factors and methane data in standard units were filled using the sampling strip method to obtain the filled climate factors and methane data; The resolution of the filled climate factors was unified by using bicubic interpolation to obtain climate factors with the same resolution. The K-nearest neighbor algorithm is used to correct the bias of climate factors with the same resolution to obtain the corrected climate factors; The corrected climate factors, socioeconomic indicators, and methane data were standardized to obtain standardized climate factors, socioeconomic indicators, and methane data. Principal component analysis was used to reduce the dimensionality of standard climate factors and methane data, resulting in dimensionality-reduced climate factors and methane data. Screen climate-driven factors and socio-economic-driven factors, and determine target climate-driven factors and target socio-economic-driven factors; The hybrid deep learning model is trained based on the target climate driving factors and the target socio-economic driving factors, so that the trained hybrid deep learning model can generate the target prediction model. Input climate and socioeconomic driving factor data under future scenarios into the prediction model to implement methane prediction and generate prediction results. Perform inverse PCA transformation and inverse standardization on the prediction data to obtain the predicted methane concentration data.
2. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 1, characterized in that: The method of using principal component analysis to reduce the dimensionality of standard climate factors and methane data to obtain dimensionality-reduced climate factors and methane data also includes: Principal components retaining 95% variance were extracted from the standard climate factors and methane data to obtain low-latitude climate factors and low-latitude methane data.
3. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 2, characterized in that: The methods for screening climate and socioeconomic drivers to determine the climate and socioeconomic drivers affecting the target region include: Based on low-latitude climate factors and low-latitude methane data, the monthly climate factors in the low-latitude climate factors are determined as the basic climate factor indicators. Collect socioeconomic indicator data of the target area to construct basic socioeconomic parameters. The socioeconomic indicator data includes economic scale parameters, population parameters and environmental statistics parameters of the target area. The Pearson correlation coefficient method was used to calculate the correlation between basic climate factor indicators and basic socioeconomic parameters, and to extract core climate parameters and core socioeconomic parameters. The core climate parameters and the core socio-economic parameters are screened according to a preset threshold rule, and the screened core climate parameters and the core socio-economic parameters are determined as the climate driving factors and the socio-economic driving factors of the target region.
4. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 3, characterized in that: The method for screening the core climate parameters and the core socio-economic parameters according to the preset threshold rule and determining the screened core climate parameters and the core socio-economic parameters as the climate driving factors and the socio-economic driving factors of the target region comprises: The preset threshold rule is set as |r|>0.5; The core climate parameters and the core socio-economic parameters are screened to obtain the core climate parameters and the core socio-economic parameters, wherein |r| is an empirical threshold value.
5. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 4, characterized in that: The method for training the mixed deep learning model based on the target climate driving factors and the target socio-economic driving factors, so that the trained mixed deep learning model generates a prediction model comprises: An initial training model is obtained by performing a training process on the mixed deep learning model based on low-latitude methane data, climate driving factors and socio-economic driving factors. Parameters of the initial training model are analyzed based on preset initial values to output an optimal hyperparameter configuration, and a high-precision methane concentration prediction model is established to complete generation of the prediction model.
6. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 5, characterized in that: The mixed deep learning model comprises a climate feature extraction and analysis path and a socio-economic feature analysis path, the climate feature extraction and analysis path comprises a double-layer LSTM structure, and the socio-economic feature analysis path comprises a fully connected layer having a plurality of neurons.
7. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 6, characterized in that: The method for performing a training process on the mixed deep learning model based on low-latitude methane data, climate driving factors and socio-economic driving factors to obtain an initial training model comprises: A double-input architecture is used to realize multi-modal fusion to take into account both time sequence dependence and static feature correlation; A double-layer LSTM structure is used to extract features in layers, wherein a first LSTM layer is used to capture local short-term patterns in a sequence of climate driving factors, and a second LSTM layer is used to further extract global long-term dependence data of the climate driving factors based on an output of the first LSTM layer; A fully connected layer is used to map features of socio-economic driving factors to an abstract space; A regularization method is used to suppress abnormal value interference in data of the socio-economic driving factors; A splicing layer is set, and features of the climate driving factors and the socio-economic driving factors are fused through channel splicing; A fully connected layer ReLU activation function is used to synthesize climate-society synergistic effects, a random inactivation layer is set to randomly shield a certain proportion of nodes to simulate stability in a data missing scenario, and robustness of the data is improved; An output layer is set to linearly output a continuous distribution characteristic of an adaptive target variable through the output layer; The target prediction model is obtained by performing multiple training through steps corresponding to the above method for obtaining the initial training model.
8. The atmospheric methane concentration prediction method based on multi-source data fusion according to claim 7, characterized in that: The method for training the mixed deep learning model based on the target climate driving factors and the target socio-economic driving factors, so that the trained mixed deep learning model generates a prediction model further comprises: A 20% random inactivation layer is provided to capture 12-step time sequence dependence of principal component difference features of precipitation.
Citation Information
Patent Citations
Multi-attribute fusion air quality forecasting method based on deep learning
CN114676822A
Methane concentration prediction method and system based on neural network hybrid model
CN119495376A