Multi-source heterogeneous data fusion method of "sky tower and ground well" carbon flux observation system

By employing adaptive correction and multi-level fusion algorithms, the problems of spatiotemporal mismatch and insufficient real-time processing in the observation of carbon emissions and carbon sinks in coal mining areas have been solved. This has enabled full-scale, dynamic, real-time, and high-precision observation, which is suitable for complex and ever-changing coal mining environments.

CN120408516BActive Publication Date: 2026-01-06SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510529657.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2026-01-06
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing technologies for monitoring carbon emissions and carbon sinks in coal mining areas suffer from problems such as spatiotemporal mismatch of data, insufficient correction methods, poor algorithm adaptability, and insufficient real-time processing capabilities. This makes it difficult to effectively integrate multi-source heterogeneous data, affecting the accuracy and real-time performance of the monitoring.

Method used

An adaptive correction and multi-level fusion algorithm is adopted, including data preprocessing, spatiotemporal registration, automatic correction, multi-level data fusion and decision-level fusion. Kalman filtering and particle filtering are used to remove noise, convolutional neural networks and recurrent neural networks are used to extract features, and ensemble learning is combined to perform dynamic weight optimization to achieve high-precision fusion of multi-source heterogeneous data.

Benefits of technology

It has achieved full-scale, dynamic, real-time, and high-precision carbon emission and carbon sink observation, breaking through the limitations of traditional observation technologies, adapting to the complex and ever-changing coal mining environment, and improving the accuracy and reliability of data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_4
    Figure QLYQS_4
Patent Text Reader

Abstract

The multi-source heterogeneous data fusion method of the sky tower and ground well carbon flux observation system disclosed by the application comprises the following steps: step 1, data preprocessing and automatic correction, solving the noise, space-time inconsistency, format difference and systematic error in the multi-source heterogeneous data acquisition process, wherein the multi-source heterogeneous data come from satellite remote sensing, unmanned aerial vehicle remote sensing, eddy covariance flux tower, ground observation and underground observation; step 2, multi-level data fusion, the multi-source heterogeneous data are gradually refined from the original signal to form a unified high-dimensional feature expression, and finally the real-time observation data of carbon emission and carbon sink in the coal mining area are output through a decision model; and step 3, comprehensive observation and application output, real-time analysis is carried out according to the real-time observation data, and detection results and early warning information are output. The advanced algorithm with adaptive correction and multi-level fusion capability is adopted, the data fusion of the five observation modules of the sky tower and ground well is realized, and thus the key bottleneck in the prior art is broken.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon neutrality and data fusion methods, and in particular to a multi-source heterogeneous data fusion method for a "sky tower ground well" carbon flux observation system. Background Technology

[0002] In recent years, the global demand for carbon emission and carbon sink monitoring has been increasing, especially in coal mining areas where carbon emissions from production, transportation, and energy conversion are particularly prominent. Meanwhile, the carbon sink capacity of underground ecosystems in coal mining areas plays a crucial role in regional carbon balance. Currently, carbon emission and carbon sink monitoring in coal mining areas mainly relies on single or localized technologies, each with its own limitations. While satellite remote sensing can achieve large-scale carbon flux acquisition, it is limited by resolution, timeliness, and terrain obstruction, making it difficult to meet the needs of localized high-precision observation. UAV remote sensing can conduct localized high-precision observations, but its flight time is limited, preventing long-term continuous observation. Eddy covariance flux towers can continuously collect atmospheric carbon flux data in real time, but their observation range is limited, only achieving mesoscale observations. While ground and underground sensors combined with manual sampling can achieve small-scale high-precision observations, their deployment range is limited, and manual sampling lacks continuity, making it difficult to cover the entire mining area for long-term, continuous, real-time, and high-precision observation.

[0003] The use of multiple observation methods for comprehensive observation is a novel technological approach. Therefore, a "sky-tower-ground-well" integrated carbon flux observation system is proposed. This system scientifically deploys five carbon emission and carbon sink observation methods—satellite remote sensing (sky), UAV remote sensing (air), eddy covariance flux tower observation (tower), ground sensors and manual sampling (ground), and underground sensors and manual sampling (well)—at different locations and regions within coal mining areas. This constructs a multi-level, multi-dimensional integrated observation network, enabling full-scale, dynamic, real-time, high-precision, and long-term observation and data acquisition of carbon emissions and carbon sinks in the mining area.

[0004] However, the multi-source heterogeneous data obtained through different observation methods exhibit significant differences in acquisition methods, data formats, spatiotemporal resolution, and accuracy, leading to data silos. This hinders the comprehensive utilization of information and further affects the data analysis and fusion of the "Sky Tower Ground Well" integrated observation system. Specifically, the following are the main problems that need to be addressed:

[0005] 1. Data spatiotemporal mismatch: There are significant differences in the time interval, spatial coverage and resolution of data collected by different observation modules. Traditional spatiotemporal registration algorithms are difficult to achieve accurate alignment, which affects the quality of data fusion.

[0006] 2. Insufficient calibration methods: Due to factors such as equipment errors, environmental interference, and transmission delays, systematic deviations exist between different data sources. Existing methods mostly rely on manual calibration or static calibration schemes, lacking the ability to adaptively calibrate data in real time under dynamic and multi-scale environments.

[0007] 3. Poor algorithm adaptability: Most data fusion algorithms are designed for specific scenarios and are difficult to cope with the complex and ever-changing environmental conditions in coal mining areas, let alone effectively integrate heterogeneous data from multiple sources in the air and underground.

[0008] 4. Lack of real-time processing capability: In practical applications, the data fusion process needs to process a large amount of data in a short period of time. Traditional algorithms are insufficient in terms of real-time performance and computational efficiency, making it difficult to meet the requirements of real-time monitoring and early warning. 5. Data incompatibility between observation modules: Data obtained by different observation methods have significant differences in terms of acquisition methods, data formats, spatiotemporal resolution and accuracy, forming data silos and hindering the comprehensive utilization of information. Summary of the Invention

[0009] To address the aforementioned issues, this invention provides a multi-source heterogeneous data fusion method for the "Sky Tower Ground Well" carbon flux observation system. It employs an advanced algorithm with adaptive correction and multi-level fusion capabilities, thereby overcoming key bottlenecks in existing technologies such as data silos, spatiotemporal mismatch, and insufficient real-time processing.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0011] The multi-source heterogeneous data fusion method of the "Sky Tower Ground Well" carbon flux observation system includes the following steps:

[0012] Step 1: Data preprocessing and automatic correction to resolve noise, spatiotemporal inconsistencies, format differences and systematic errors in the process of acquiring multi-source heterogeneous data. The multi-source heterogeneous data comes from five observation modules: satellite remote sensing, UAV remote sensing, eddy covariance flux tower, ground observation and downhole observation.

[0013] Step 2: Multi-level data fusion, which extracts multi-source heterogeneous data from the original signals step by step to form a unified high-dimensional feature expression, and finally outputs real-time observation data on carbon emissions and carbon sinks in coal mining areas through a decision model.

[0014] Step 3: Comprehensive observation and application output. Based on real-time observation data, perform real-time analysis and output carbon emission heat maps, carbon sink distribution maps, carbon emission and carbon sink reports, and early warning information.

[0015] Furthermore, step 1 includes:

[0016] Step 1.1: Data cleaning and noise suppression, including data redundancy removal, invalid value removal, missing value imputation, and noise suppression;

[0017] Step 1.2, Spatiotemporal Registration and Automatic Correction: The spatiotemporal registration transforms data from different data sources into a unified spatiotemporal coordinate system, eliminating data inconsistencies caused by sampling time deviations and spatial location errors. The automatic correction uses manually sampled data or high-precision reference station data to dynamically correct each data source, thereby improving the overall accuracy of the data.

[0018] Preferably, the method for removing data redundancy is as follows:

[0019] If the difference between the observations of the same parameter data from different sensors at timestamp t for i and j is within the threshold range (|d) i -d j If |<∈), then retain one data source and remove the remaining duplicate records, as shown in equation (1):

[0020]

[0021] Where: d i and d j These are observations from different sensors at the same timestamp, where ∈ represents a set error threshold, which can be set.

[0022] Preferably, the invalid value removal includes invalid data removal and outlier removal. The invalid data removal method is as follows: for each sensor, a reasonable working range is preset, and data outside the reasonable working range is removed.

[0023] The outlier removal method is as follows: invalid data is removed from the data after invalid data removal using statistical methods.

[0024] Preferably, the method for filling missing values ​​is as follows:

[0025] For numerical data, use mean filling, mode filling, or median filling;

[0026] For time-series data, missing time point data are filled by interpolation methods, including linear interpolation and / or spline interpolation.

[0027] Preferably, the noise suppression employs a combination of Kalman filtering and particle filtering to remove noise.

[0028] Preferably, the spatiotemporal registration and automatic correction include:

[0029] Spatiotemporal registration: For numerical data, a time synchronization method based on cross-correlation is first used. Given two time series {z}... i} and {w i}, calculate their cross-correlation function R zw(τ) is as follows:

[0030]

[0031] Where τ is the time offset, chosen to make R zw (τ) The largest τ0 is used as the time alignment offset between the two sequences;

[0032] Then, for time series data with discontinuities or uneven sampling, interpolation methods are used to resample the data onto a unified time axis. The interpolation methods include linear interpolation and / or spline interpolation.

[0033] Automatic correction, constructing the error model as follows:

[0034] Assuming that the sensor observation z has a systematic bias b, and the true value is x, then:

[0035] z = x + b (16)

[0036] By comparing with reference data x ref By comparison, an error model is established, and the statistical characteristics of the deviation are calculated using multiple sampling data to obtain an initial estimate b0, with reference data x. ref For industrial sampling data or high-precision reference station data;

[0037] Real-time sensor bias correction: An adaptive correction method based on Kalman filtering is adopted, assuming the current corrected estimated value is... The reference value is x ref,k Then, by correcting the update bias b k :

[0038] b k =(1-λ)b k-1 +λ(z k -x ref,k (17)

[0039] Where λ is the smoothing factor, ranging from 0 to 1, reflecting the weight of the reference data on the correction, and the corrected observations... The calculation is as follows:

[0040]

[0041] By continuously updating the deviation b k This allows the data from each sensor to gradually approach the true value and adapt to dynamic changes in the environment and equipment status.

[0042] Preferably, the multi-level data fusion includes:

[0043] The underlying data fusion uses interpolation techniques and normalization methods to construct a standard data matrix, eliminating differences in sampling frequency, spatial resolution, and units between different data sources;

[0044] Feature-level fusion is used to obtain a deep joint representation of multimodal data;

[0045] Decision-level fusion, based on the joint feature vector obtained by feature-level fusion, constructs a decision model to obtain the final prediction and early warning output.

[0046] More preferably, the underlying data fusion includes:

[0047] Spatiotemporal grid construction: Using a preset spatiotemporal grid, the target area is divided into fixed spatial units, and time is divided according to a fixed period. All data is mapped to the spatiotemporal grid, and for each grid unit, synchronous processing is performed according to the sampling time of the data source.

[0048] Data interpolation and normalization: For missing data points in the grid, linear interpolation or cubic spline interpolation is used to supplement them; at the same time, in order to eliminate the influence of different data sources, normalization is used to unify all data into the same numerical range.

[0049] More preferably, the feature-level fusion includes:

[0050] Multimodal feature extraction: For image data, convolutional neural networks are used to extract the spatial features of the image;

[0051] For time-series data, recurrent neural networks or long short-term memory networks are used to extract dynamic features;

[0052] The spatial features and dynamic features are fused using a fully connected neural network to form a joint feature vector.

[0053] Feature dimensionality reduction and optimization are performed by using principal component analysis or autoencoders to optimize the joint feature vector and extract the most representative feature components.

[0054] More preferably, the decision-level fusion includes:

[0055] Multi-model prediction is constructed, using several prediction models to predict carbon emissions and carbon sinks respectively;

[0056] Ensemble learning and dynamic weight optimization are employed to weight and fuse the prediction results of several models to obtain the final output value.

[0057] The beneficial effects of this invention include:

[0058] 1. Multi-source heterogeneous data fusion: This invention adopts multi-level data fusion technology, which can simultaneously process heterogeneous data from different observation modules (sky tower and ground well), breaking through the limitations of traditional observation technology.

[0059] 2. Dynamic adaptive correction and fusion mechanism: This invention introduces a dynamic weight calculation and automatic correction mechanism based on real-time feedback, which significantly improves the accuracy and reliability of data fusion, and is especially suitable for complex and variable coal mining environments.

[0060] 3. Fusion of carbon emissions and carbon sinks across all scales: This invention achieves the fusion of observation data across all scales, from large-scale to medium-scale to small-scale, providing a new and precise data processing method for comprehensive observation of carbon emissions and carbon sinks in coal mining areas. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below.

[0062] This embodiment proposes a multi-source heterogeneous data fusion method for a "sky-tower-ground-well" carbon flux observation system. This method is specifically designed for integrated carbon emission and carbon sink observation technologies in coal mining areas. By tightly fusing, supplementing, and correcting carbon emission and carbon sink observation data at different scales, it achieves full-scale, dynamic, real-time, and high-precision observation of carbon emissions and carbon sinks within coal mining areas. This embodiment includes the following main contents:

[0063] 1. System Architecture Design

[0064] This invention constructs an integrated carbon emission and carbon sink data acquisition and fusion system with a four-layer architecture, consisting of a "sky tower and ground well".

[0065] 1) The acquisition layers include the sky layer, tower layer, and well layer, each corresponding to different data observation and acquisition modules:

[0066] The "sky" level comprises two main modules.

[0067] Satellite remote sensing module: Utilizing multispectral, high-resolution remote sensing satellites, this module periodically scans the mining area and its surroundings to acquire environmental parameters such as vegetation index, soil moisture, and temperature, thereby predicting changes in carbon flux within the mining area. Particularly useful for large-scale, long-term dynamic observations, satellite remote sensing can provide carbon flux data for the mining area and its surrounding regions, filling gaps in large-scale mining area observation techniques and is suitable for continuously acquiring macroscopic data on carbon flux changes.

[0068] UAV remote sensing module: In areas with insufficient satellite data updates or complex local environments, UAVs equipped with high-precision sensors supplement the collection of local environmental images and radiation data. It is particularly suitable for short-term, dynamic observation of local areas in mining areas, and has unique advantages, especially in the detailed observation of local hotspots.

[0069] At the "tower" level: Eddy covariance flux towers are deployed in key areas of the mining area. Using equipment such as gas flow sensors, temperature and humidity sensors, and weather stations, real-time data on wind speed, temperature, carbon dioxide, and other greenhouse gas concentrations are collected, and the atmospheric carbon gas exchange flux is calculated. The eddy covariance flux towers can monitor changes in carbon emissions and carbon sinks within the mining area in real time. Their observation range is typically 1.5-2 times the tower height, depending on surface roughness and turbulence characteristics. By measuring high-frequency time-series data of wind speed and gas concentration and calculating their covariance, the vertical gas exchange rate in the turbulent flow can be obtained, thereby estimating the gas flux in the vertical direction.

[0070] The "well" level comprises two main modules.

[0071] Ground observation module: Deploying fixed or mobile sensor networks, combined with regular manual sampling, can accurately capture the carbon release and absorption processes within the mining area, including the impact of mining activities on the carbon sink capacity of soil and vegetation, and monitor key indicators of carbon emissions and carbon absorption in the surface environment in real time.

[0072] Downhole observation module: Similar to the surface module, the downhole observation module deploys explosion-proof carbon flux observation sensors that are resistant to high temperatures, water, moisture, and dust in key areas of the well. Combined with manual sampling, it obtains downhole carbon flux data and other relevant environmental parameters.

[0073] 2) The transmission layer uses wireless network and Internet of Things (IoT) technologies to transmit the data collected from each layer to the central data processing platform in real time for processing and analysis.

[0074] 3) The fusion layer performs preprocessing, spatiotemporal registration, automatic correction, and data fusion on the collected multi-source data. It achieves bottom-level synchronous fusion, feature-level fusion, and decision-level fusion, providing high-quality data for subsequent analysis and decision-making.

[0075] 4) The application layer uses the fused data for real-time monitoring, risk assessment, and visualization of carbon emissions and carbon sinks. Reports, early warnings, and decision support are generated based on the real-time monitoring results.

[0076] 2. Data preprocessing and automatic correction

[0077] To ensure the accuracy and consistency of multi-source heterogeneous data fusion, this embodiment proposes an innovative data preprocessing and automatic correction scheme. Through multiple steps and technical means, it solves problems such as noise, spatiotemporal inconsistency, format differences, and systematic errors in the data acquisition process, as detailed below:

[0078] 2.1 Data Cleaning and Noise Suppression

[0079] Data cleaning and noise suppression are crucial steps in data preprocessing, aiming to improve the accuracy and reliability of multi-source heterogeneous data fusion. Because the raw data provided by various sensors and observation modules in coal mining areas (such as satellite remote sensing, UAV imagery, weather towers, and ground sensors) contains redundancy, missing values, anomalies, and noise, precise cleaning and noise suppression techniques must be employed to ensure that all data meets fusion standards. This includes data redundancy removal, invalid value removal, missing value imputation, and noise suppression, as detailed below:

[0080] 2.1.1 Data Redundancy Removal

[0081] The first step is data redundancy removal, which aims to remove redundant, invalid, or erroneous data to ensure the validity and consistency of the remaining data. Due to multi-module data synchronization issues, duplicate records may exist with the same timestamp or in the same region. To avoid data redundancy affecting subsequent analysis, a timestamp deduplication method is used to remove redundant data:

[0082] For duplicate records from different sensors, their timestamps (t) are compared. i and t j ) and data values ​​(d i and d j Using baseline sensor data as the standard, duplicate data is removed. If the difference between the observations provided by sensors i and j at timestamp t is within a threshold range (|d)... i -d j If |<∈), then retain one data source and remove the remaining duplicate records, as follows:

[0083]

[0084] Where d i and d j The values ​​represent observations from different sensors at the same timestamp, and ∈ represents the set error threshold. The error threshold needs to be set based on the characteristics of the data source, data quality requirements, and the accuracy requirements of the application scenario. To ensure the scientific validity and rationality of the error threshold, it is set using statistical analysis. By statistically analyzing the error distribution of different data sources, and based on the statistical characteristics of the data such as standard deviation, mean, and skewness, the threshold can be set to 1 or 2 times the standard deviation.

[0085] In practical applications, data accuracy can fluctuate due to factors such as environment, equipment aging, and weather. Therefore, the error threshold needs to be dynamically adjusted through a real-time feedback mechanism. During operation, if the data error of a particular sensor or observation module remains consistently high, the system can automatically increase the error threshold for that sensor to minimize its impact on the data fusion results. Conversely, if the error of a data source decreases, the system can appropriately lower its error threshold, thereby improving data accuracy. The specific formula is as follows:

[0086] New Threshold=Old Threshold+δ×(Current Error-Expected Error) (2)

[0087] Where δ is the adjustment coefficient, Current Error is the error of the current data source, and Expected Error is the expected error. In the application of carbon emission and carbon sink monitoring in coal mining areas, the initial error threshold is set to 1% to 5% based on sensor type, data acquisition frequency, and actual application requirements, and is further adjusted using experimental data. For critical observation data (such as carbon emissions), a more stringent threshold (e.g., 0.1% to 1%) may be required to ensure high-precision observation results.

[0088] After data cleaning, spatial redundancy removal is necessary. When processing data from different modules (e.g., satellite remote sensing and UAV imagery), due to spatial overlap or redundant regions, duplicate spatial data needs to be removed to ensure that information from the same region is not calculated repeatedly. For the sky module (satellite remote sensing and UAV imagery), if the observation areas of the two overlap, the higher-resolution data is selected for retention based on the resolution of the overlapping region, while the lower-resolution data is discarded. Redundancy removal based on region matching involves calculating the area of ​​the overlapping region (e.g., by calculating the pixel area of ​​the overlapping region). If the overlapping area is large, the higher-resolution data is retained. The deduplication based on spatial overlap area uses the following formula:

[0089] Suppose we have two sets of data, D1 and D2, which correspond to satellite data and UAV data, respectively. If these two sets of data have an overlapping region R, and the area of ​​the overlapping region is A... R Greater than a certain threshold A threshold If the resolution is higher, then the higher-resolution data is retained. This can be represented as:

[0090] If A R >A threshold ,then keep D high res and discard D low res (3)

[0091] 2.1.2 Removal of Invalid Data and Outliers

[0092] During data acquisition, data values ​​may occur that exceed the normal operating range of the equipment or violate common physical principles. This type of data falls into two categories:

[0093] Invalid data: refers to data that is clearly inconsistent with the intended physical range or equipment operating standards (e.g., temperature readings of -100°C, or data that exceeds the instrument's maximum measurement range).

[0094] Outliers: These are values ​​that deviate significantly from the majority of data in a statistical distribution. Such data may be caused by accidental interference or systematic errors, but under certain conditions, they may also reflect real extreme phenomena.

[0095] This embodiment adopts a step-by-step elimination strategy, first eliminating obviously invalid data, and then using statistical methods to eliminate outliers from the remaining data.

[0096] When removing invalid data, set physical range limits: For each sensor, pre-define its reasonable operating range. For example, for a temperature sensor, the reasonable range is [T...]. min ,T max The formula is retained as follows:

[0097] If x i <T min or x i >T max (4)

[0098] For the retained data, outliers were detected using statistical methods, employing a dual selection approach based on both standard deviation and interquartile range (IQR) to ensure data stability.

[0099] First, the standard deviation method is used to statistically eliminate data with excessive deviation. Assuming the data follows a normal distribution, the mean μ and standard deviation σ of the dataset are calculated. For any data point x... i The judgment is as follows:

[0100] |x i -μ|>3σ (5)

[0101] If the above conditions are met, then x is considered to be... i This is invalid data and should be removed from the dataset.

[0102] After performing the standard deviation test, the IQR method is then used to calculate the first quartile Q1 and the third quartile Q3 of the data. The interquartile range is defined as:

[0103] IQR = Q3 - Q1 (6)

[0104] The criteria for outlier determination are:

[0105] If x i <Q1-1.5×IQR or x i >Q3+1.5×IQR (7)

[0106] After invalid values ​​are removed, the dataset is ensured to contain only data that conforms to physical meaning and statistical laws, providing a foundation for subsequent missing value supplementation and noise suppression.

[0107] 2.1.3 Missing Data Imputation

[0108] After removing invalid values, interpolation is used to impute missing values ​​in the dataset. Missing values ​​are data points missing due to various reasons during data acquisition (such as sensor malfunction, power outage, data transmission loss, etc.). The purpose of missing value imputation is to fill in the missing data using appropriate algorithms, ensuring the dataset remains complete and does not affect subsequent analysis. The method for imputing missing values ​​depends on the data type, the extent of data loss, and application requirements. For numerical data, if the amount of missing data is small, mean imputation, mode imputation, or median imputation can be used. For time series data, interpolation is used to fill in missing time point data, employing methods such as linear interpolation or spline interpolation. For a missing data point x... i Assume the known data before and after it is x. i-1 and x i+1 Then, linear interpolation is performed using the following formula:

[0109]

[0110] For nonlinear interpolation of data, cubic spline interpolation can be used:

[0111] S(x) = a + bx + cx 2 +dx 3 (9)

[0112] Here, a, b, c, and d are coefficients obtained by the least squares method or numerical optimization algorithm.

[0113] For time series data, interpolation is generally better at preserving the continuity and trend of the data than mean filling, especially when the changes between adjacent data points are relatively stable.

[0114] 2.1.4 Noise Suppression

[0115] Noise suppression is another crucial step in ensuring data quality, especially when processing high-frequency, dynamically changing data. This invention employs a combination of Kalman filtering and particle filtering to remove noise and ensure data smoothness and consistency.

[0116] The Kalman filter update formula is:

[0117]

[0118] in, z is the estimated value at the current time. k K represents the current observation value. k The Kalman gain is calculated using the following formula:

[0119]

[0120] Among them, P k-1 R is the covariance of the prior error, and R is the covariance of the observation noise.

[0121] Particle filtering is suitable for nonlinear and non-Gaussian noise data. By resampling multiple particles, particle filtering can more accurately estimate the system state. The particle filtering formula is as follows:

[0122]

[0123] Resampling steps: Based on weights i Perform particle resampling.

[0124] 2.2 Spatiotemporal Registration and Automatic Correction

[0125] In the process of multi-source data fusion, due to differences in sampling time, spatial resolution, and acquisition angle among different data acquisition modules (such as satellites, UAVs, ground and downhole sensors), precise spatiotemporal registration and automatic correction are necessary to ensure the uniformity and high accuracy of the fused data. To this end, this invention proposes the following two main modules: a spatiotemporal registration module and an automatic correction module.

[0126] 2.2.1 Spatiotemporal Registration

[0127] The purpose of spatiotemporal registration is to transform data from different data sources into a unified spatiotemporal coordinate system, thereby eliminating data inconsistencies caused by sampling time deviations and spatial location errors.

[0128] For time alignment, the sampling frequencies and timestamps of different data sources may differ. To ensure data comparability on the same time scale, a time synchronization method based on cross-correlation is adopted, supplemented by interpolation techniques. The specific steps are as follows:

[0129] Based on the cross-correlation time synchronization method, given two time series {z} i} and {w i}, calculate their cross-correlation function R zw (τ) is as follows:

[0130]

[0131] Where τ is the time offset. Select R such that... zw The largest τ0 is used as the time alignment offset between the two sequences.

[0132] Then, interpolation calculations are performed. For time series data with discontinuities or uneven sampling, linear interpolation or spline interpolation techniques are used to resample the data onto a unified time axis. The interpolation formulas are (8) or (9).

[0133] For aerial imagery data (such as satellite and UAV images) and other data containing geographic location information, a feature-matching-based spatial registration method is employed. The main steps include:

[0134] Feature point extraction and matching: Use mature algorithms (such as SIFT or SURF) to extract feature points from each image, and obtain a preliminary set of corresponding points through descriptor matching.

[0135] Robust solution to the transformation matrix: The Random Sample Consensus (RANSAC) algorithm is used to eliminate mismatches, and the homography matrix H between the two images is calculated. This matrix satisfies the following relationship:

[0136] x′=Hx (14)

[0137] Where x is a point in the original image (represented by homogeneous coordinates), and x′ is the corresponding point in the target image. To obtain the optimal H, the error function is usually minimized:

[0138]

[0139] Through the above time and space registration process, data from different data sources are all converted into a unified spatiotemporal coordinate system, ensuring that subsequent data fusion is based on consistent and accurate data.

[0140] 2.2.2 Automatic Correction

[0141] In multi-source data acquisition, different sensors may generate systematic errors due to system bias, environmental changes, and long-term drift. To eliminate these errors, this invention proposes an automatic correction method based on real-time reference feedback. The core idea is to dynamically correct each data source using manually sampled data or high-precision reference station data, thereby improving the overall accuracy of the data.

[0142] First, we construct an error model, assuming that the sensor observation z has a systematic bias b, and the true value is x, then we have:

[0143] z = x + b (16)

[0144] By comparing with reference data x refBy comparison, an error model is established. The statistical properties of the deviation are calculated using multiple sampling data to obtain an initial estimate b0.

[0145] To correct sensor bias in real time during data acquisition, this embodiment employs an adaptive correction method based on Kalman filtering. Besides noise reduction and smoothing, Kalman filtering can also be used to estimate system bias. Let the current corrected estimate be... The reference value is x ref,k Then, the update bias b is corrected:

[0146] b k =(1-λ)b k-1 +λ(z k -x ref,k (17)

[0147] Where λ is the smoothing factor, ranging from 0 to 1, reflecting the weight of the reference data on the correction. Corrected observations The calculation is as follows:

[0148]

[0149] By continuously updating the deviation b k This allows the data from each sensor to gradually approach the true value and adapt to dynamic changes in the environment and equipment status.

[0150] Considering the subtle errors that may exist during spatiotemporal registration, automatic correction is combined with spatiotemporal registration. For spatially registered image data or point cloud data, joint correction is achieved by comparing overlapping areas with reference map data. Let x... i For the registered data points, the corresponding reference data point is x. ref,i Let N be the total number of data points. Construct the joint correction error function E:

[0151]

[0152] Iterative optimization methods such as gradient descent or Levenberg-Marquardt algorithm are used to adjust the spatiotemporal registration parameters and sensor calibration parameters to minimize EEE, thereby achieving the best joint calibration effect.

[0153] 3. Multi-level data fusion algorithm

[0154] To fully utilize the high-quality data obtained after data cleaning, spatiotemporal registration, and automatic correction, this invention employs a multi-level data fusion algorithm. The basic idea is to progressively extract multi-source heterogeneous data from the original signals to form a unified high-dimensional feature representation. Finally, a decision model is used to output real-time observation and early warning information on carbon emissions and carbon sinks in coal mining areas. The algorithm mainly consists of three levels: bottom-level data fusion, feature-level fusion, and decision-level fusion.

[0155] 3.1 Underlying Data Fusion

[0156] In the underlying data fusion stage, data from various data sources, after preprocessing (including data cleaning, spatiotemporal registration, and automatic correction), are mapped onto a unified spatiotemporal grid. To eliminate differences in sampling frequency, spatial resolution, and units between different data sources, this invention employs interpolation techniques and normalization methods to construct a standard data matrix. Key technologies and steps are as follows:

[0157] Spatiotemporal grid construction: A preset spatiotemporal grid is used to divide the target area into fixed spatial units, with time divided according to a fixed period, and all data is mapped to this grid. For each grid unit, synchronization processing is performed according to the data source sampling time.

[0158] Data interpolation and normalization: For missing data points in the grid, linear interpolation or cubic spline interpolation is used to supplement them; simultaneously, to eliminate the influence of different data sources' dimensions, normalization is used to unify all data into the same numerical range. Standardization (z-score) is as follows:

[0159]

[0160] The mean is μ, and the standard deviation is σ. The goal of low-level fusion is to provide consistent and comparable data input for subsequent feature extraction.

[0161] 3.2 Feature-level fusion

[0162] Feature-level fusion is one of the core innovations of this invention, which mainly achieves deep joint representation of multimodal data through the following steps:

[0163] Multimodal feature extraction employs different deep neural network models to extract features from different data types. For the data obtained from the sky module (satellite remote sensing and UAV imagery), a convolutional neural network (CNN) is used to extract the spatial features of the images, expressed as follows:

[0164] f i =CNN(I i ) (twenty one)

[0165] Among them, I i f represents the image input from the i-th source.i It is its corresponding feature vector.

[0166] For time-series data (such as eddy covariance towers, surface and downhole sensor data), dynamic features are extracted using recurrent neural networks (RNNs) or long short-term memory networks (LSTMs), expressed as follows:

[0167] g i =LSTM(s i ) (twenty two)

[0168] Among them, s i G represents the time series data from the i-th source. i The extracted temporal features.

[0169] Joint feature representation and dynamic weighting: To fully utilize the complementary information from various data sources, this invention fuses the above features using a fully connected neural network (FCN) to form a joint feature vector. Let the joint features of the i-th data source be:

[0170] F i =[f i g i ] (twenty three)

[0171] To enable each data source to adaptively adjust its contribution during the fusion process, this embodiment introduces a dynamic weighting mechanism, with weight w. i The calculation formula is:

[0172]

[0173] Among them, F i (t) is the feature vector of the i-th data source at time t; F j (t) represents the feature vector of the j-th data source after feature extraction at time t; F ref (t) is the reference feature vector, which can be determined by high-precision data or preliminary fusion results; λ is the adjustment parameter; N is the total number of data sources.

[0174] Therefore, the fused joint feature vector F fused (t) is then given by the following formula:

[0175]

[0176] Feature dimensionality reduction and optimization: To reduce the computational complexity caused by high-dimensional features, dimensionality reduction techniques such as principal component analysis (PCA) or autoencoders are used to optimize the joint feature vector, extract the most representative feature components, and further improve the generalization ability and running efficiency of subsequent models.

[0177] 3.3 Decision-level integration

[0178] In the decision-level fusion stage, a decision model is constructed based on the joint feature vector obtained from feature-level fusion to achieve the final prediction and early warning output. Specific solutions include:

[0179] This invention employs multiple prediction models (random forest, support vector machine, and deep neural network, among others) to predict carbon emissions and carbon sinks. Let the output of the i-th model be y. i .

[0180] Ensemble learning and dynamic weight optimization. To improve prediction accuracy and robustness, an ensemble learning method is employed to weight and fuse the prediction results from each model. The final output value y fused Represented as:

[0181]

[0182] Where M is the total number of models, α i The weights are dynamically adjusted in real time using a Bayesian optimization method to adapt to data changes and ensure the accuracy of the final prediction.

[0183] 4. Comprehensive observation and application output

[0184] After data fusion, this invention provides the following outputs through real-time analysis: Carbon emission heat map: By fusing multi-source data, a carbon emission distribution map of the mining area is generated in real time, helping to observe carbon emission hotspots. Carbon sink distribution map: Displays the carbon sink situation in the coal mining area and predicts the carbon sink change trends in different areas. Carbon emission and carbon sink reports: Detailed carbon emission and carbon sink reports are automatically generated to support decision-making and policy formulation. Early warning system: Through real-time data observation, it provides early warnings of abnormal carbon emissions, enabling timely responses for environmental supervision in the mining area.

[0185] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A multi-source heterogeneous data fusion method for a "sky tower and ground well" carbon flux observation system, characterized in that, Comprise the following steps: Step 1, data preprocessing and automatic correction, solve the noise, space-time inconsistency, format difference and systematic error in the process of multi-source heterogeneous data acquisition, the multi-source heterogeneous data are respectively from satellite remote sensing, unmanned aerial vehicle remote sensing, eddy covariance flux tower, ground observation and underground observation five observation modules; Step 2, multi-level data fusion, multi-source heterogeneous data are extracted from the original signal step by step, form a unified high-dimensional feature expression, and finally output the real-time observation data of coal mining area carbon emission and carbon sink through decision model; Step 3, comprehensive observation and application output, according to the real-time observation data, real-time analysis is carried out, and carbon emission thermal map, carbon sink distribution map, carbon emission and carbon sink report and early warning information are output; Step 1 includes: Step 1.1, data cleaning and noise suppression, including data redundancy removal, invalid value elimination, missing value filling and noise suppression; Step 1.2, space-time registration and automatic correction, the space-time registration converts the data from different data sources into a unified space-time coordinate system, eliminates the data inconsistency caused by sampling time deviation and spatial position error, and the automatic correction uses artificial sampling data or high-precision reference station data to dynamically correct each data source, thereby improving the accuracy of the overall data; spatial registration: for numerical data, a cross-correlation based time synchronization method is first applied, given two time series { z i} and { w i}, the cross-correlation function R zw (τ) is calculated as follows: (13) wherein τ is a time offset, chosen so that R zw (τ) the largest τ0 as the time alignment offset of the two sequences; Then for the time series data with discontinuity or uneven sampling, interpolation method is used to resample the data to a unified time axis, the interpolation method includes linear interpolation and / or spline interpolation; Automatic correction, the error model is constructed as follows: Assume sensor observations z There is systematic bias b , the true value is x Then we have: , By comparing with reference data x ref Comparing, establishing error model, calculating statistical characteristics of deviation by using multiple sampling data, obtaining initial estimation b 0, Reference data x ref For working sampling data or high-precision reference station data; Sensor bias real-time correction: adopt adaptive correction method based on Kalman filter, set the current corrected estimated value as , the reference value is x ref,k , then update the bias b k : wherein λ is a smoothing factor, with a value ranging between 0 and 1, reflecting the influence weight of the reference data on the correction, the corrected observation is calculated as: by continuously updating the bias b k so that the sensor data gradually approaches the true value and can adapt to dynamic changes in the environment and device state.

2. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 1, characterized in that, The invalid value elimination includes invalid data elimination and outlier elimination, the method of invalid data elimination is that for each sensor, its reasonable working range is set in advance, and the data not in the reasonable working range is eliminated; The method of outlier elimination is that the data after invalid data elimination is eliminated by statistical method.

3. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 1, characterized in that, The method of missing value filling is: For numerical data, mean filling, mode filling or median filling is adopted; For time series data, the missing time point data is filled by interpolation method, the interpolation method includes linear interpolation and / or spline interpolation.

4. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 1, characterized in that, The noise suppression adopts the combination of Kalman filter and particle filter to remove noise.

5. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to any one of claims 1-4, characterized in that, The multi-level data fusion includes: Bottom layer data fusion, interpolation technology and normalization method are used to construct standard data matrix, and the differences in sampling frequency, spatial resolution and dimension between different data sources are eliminated; Feature level fusion is used to obtain the deep joint representation of multi-modal data; Decision level fusion is based on the joint feature vector obtained by feature level fusion, and a decision model is constructed to obtain the final prediction and early warning output.

6. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 5, characterized in that, The bottom layer data fusion includes: Space-time grid construction: a preset space-time grid is used to divide the target area into fixed spatial units, and the time is divided into fixed periods, all data are mapped to the space-time grid, and for each grid unit, the data are synchronized according to the sampling time of the data source; Data interpolation and normalization: for the missing data points in the grid, linear interpolation or cubic spline interpolation is used to supplement; at the same time, in order to eliminate the influence of the dimension of different data sources, normalization processing is adopted to unify the data to the same numerical interval.

7. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 5, characterized in that, The feature-level fusion includes: Multi-modal feature extraction, using convolutional neural network to extract spatial features of images for image data; For time series data, using recurrent neural network or long short-term memory network to extract dynamic features; Joint feature representation and dynamic weighting, using fully connected neural network to fuse the spatial features and dynamic features to form a joint feature vector; Feature dimension reduction and optimization, using principal component analysis or autoencoder to optimize the joint feature vector to extract the most representative feature components.

8. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 5, characterized in that, The decision-level fusion includes: Multi-model prediction construction, using several prediction models to predict carbon emissions and carbon sinks respectively; Integrated learning and dynamic weight optimization, using integrated learning method to weight and fuse the prediction results of the several models to obtain the final output value.

Citation Information

Patent Citations

  • Urban department carbon emission estimation method and system based on satellite observation and GIS

    CN116109191A

  • Carbon emission calculation method based on deep learning

    CN119312040A