Multi-source heterogeneous data fusion method of sky tower ground well carbon flux observation system
Through adaptive correction and multi-level fusion algorithm, the problem of spatial and temporal mismatch between carbon emissions and carbon sink observations in coal mining areas and insufficient real-time processing is solved, and full-scale, dynamic, real-time and high-precision data fusion is achieved, which is suitable for complex and changeable coal mining areas.
Patent Information
- Application Number
- CN202510529657.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing technology has problems such as space-time mismatch between data, insufficient correction methods, poor algorithm adaptability and insufficient real-time processing capabilities in the carbon emissions and carbon sink observations in coal mining areas, which makes it difficult to effectively integrate multi-source heterogeneous data, affecting observation accuracy and real-time performance.
Adaptive correction and multi-level fusion algorithms are adopted, including data preprocessing, spatiotemporal registration, automatic correction, multi-level data fusion and decision-making level fusion. Through Kalman filtering and particle filtering, denoising, features are extracted using convolutional neural networks and recurrent neural networks, unified high-dimensional feature expressions are constructed and dynamic weight optimization is performed, and real-time observation data of carbon emissions and carbon sinks are finally generated.
It realizes full-scale, dynamic, real-time and high-precision carbon emissions and carbon sink observations, breaks through the limitations of traditional observation technology, improves the accuracy and adaptability of data fusion, and is suitable for complex and changeable coal mining environments.
Smart Images

Figure BDA0005376279850000041 
Figure BDA0005376279850000051 
Figure BDA0005376279850000111
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon neutralization and data fusion methods, and in particular to a multi-source heterogeneous data fusion method for a "sky tower ground well" carbon flux observation system. Background Art
[0002] In recent years, the global demand for observing carbon emissions and carbon sinks has been continuously increasing. Especially in coal mining areas, the carbon emission problems caused by production, transportation, and energy conversion are particularly prominent. At the same time, the carbon sink capacity of the underground ecosystem in coal mining areas plays a key role in the regional carbon balance. Currently, the observation of carbon emissions and carbon sinks in coal mining areas mainly relies on single or local technical means, each with its own limitations. Although satellite remote sensing can achieve large-scale carbon flux collection, it is limited by resolution, timeliness, and terrain occlusion, and it is difficult to meet the needs of local high-precision observations; unmanned aerial vehicle (UAV) remote sensing can carry out local high-precision observations, but it is limited by flight time and cannot achieve long-term continuous observations; eddy covariance flux towers can continuously collect atmospheric carbon flux data in real time, but the observation range is limited and only mesoscale observations can be achieved; the combination of ground and underground sensors with manual sampling can achieve small-scale high-precision observations, but the layout range is limited, and manual sampling has no continuity, making it difficult to cover the entire mining area to achieve long-term, continuous, real-time, and high-precision observations.
[0003] Using multiple observation means for comprehensive observation is a new technical means. Therefore, a "sky tower ground well" carbon flux comprehensive observation system is proposed, which scientifically arranges five carbon emission and carbon sink observation means, namely satellite remote sensing (sky), UAV remote sensing (air), eddy covariance flux tower observation (tower), ground sensors and manual sampling (ground), and underground sensors and manual sampling (well), at different positions and regions in coal mining areas to construct a multi-level and multi-dimensional comprehensive observation network, so as to achieve full-scale, dynamic, real-time, high-precision, and long-term observation and data collection of carbon emissions and carbon sinks in the mining area.
[0004] However, the multi-source heterogeneous data obtained by different observation means have significant differences in collection methods, data formats, spatio-temporal resolutions, and accuracies, resulting in the formation of data islands, which hinders the comprehensive utilization of information and further affects the data analysis and fusion of the "sky tower ground well" comprehensive observation system. The following are the main problems that need to be solved specifically:
[0005] 1. Data spatio-temporal mismatch: There are significant differences in the time intervals, spatial coverage ranges, and resolutions of the data collected by different observation modules. Traditional spatio-temporal registration algorithms are difficult to achieve precise alignment, which affects the quality of data fusion.
[0006] 2. Insufficient calibration methods: Due to factors such as equipment errors, environmental interference, and transmission delays, there are systematic biases between different data sources. Existing methods mostly rely on manual calibration or static calibration schemes and lack the ability to perform real-time adaptive calibration of data in dynamic and multi-scale environments.
[0007] 3. Poor algorithm adaptability: Most current data fusion algorithms are designed for specific scenarios, making it difficult to cope with the complex and changeable environmental conditions in coal mining areas and even less able to effectively integrate multi-source heterogeneous data from the air and underground.
[0008] 4. Lack of real-time processing ability: In practical applications, the data fusion process needs to process a large amount of data in a short time. Traditional algorithms have deficiencies in real-time performance and computational efficiency and are difficult to meet the requirements of real-time monitoring and early warning. 5. Incompatibility of data between observation modules: Data obtained by different observation means have significant differences in acquisition methods, data formats, spatio-temporal resolutions, and accuracies, forming data islands and hindering the comprehensive utilization of information. Summary of the Invention
[0009] To solve the above problems, the present invention provides a multi-source heterogeneous data fusion method for a "sky-tower-ground-well" carbon flux observation system, which adopts an advanced algorithm with adaptive calibration and multi-level fusion capabilities, thereby breaking through key bottlenecks such as data islands, spatio-temporal mismatches, and insufficient real-time processing in the prior art.
[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0011] The multi-source heterogeneous data fusion method for a "sky-tower-ground-well" carbon flux observation system includes the following steps:
[0012] Step 1. Data preprocessing and automatic calibration, which solve the problems of noise, spatio-temporal inconsistency, format differences, and systematic errors in the process of collecting multi-source heterogeneous data. The multi-source heterogeneous data come from five observation modules: satellite remote sensing, unmanned aerial vehicle remote sensing, eddy covariance flux tower, ground observation, and underground observation.
[0013] Step 2. Multi-level data fusion, which gradually extracts multi-source heterogeneous data from the original signal to form a unified high-dimensional feature representation, and finally outputs real-time observation data on carbon emissions and carbon sinks in the coal mining area through a decision-making model.
[0014] Step 3. Comprehensive observation and application output, which performs real-time analysis based on the real-time observation data and outputs a carbon emission heat map, a carbon sink distribution map, a carbon emission and carbon sink report, and early warning information.
[0015] Further, Step 1 includes:
[0016] Step 1.1. Data cleaning and noise suppression, including removing data redundancy, eliminating invalid values, filling in missing values, and suppressing noise;
[0017] Step 1.2, spatio-temporal registration and automatic correction. The spatio-temporal registration converts data from different data sources into a unified spatio-temporal coordinate system to eliminate data inconsistencies caused by sampling time deviation and spatial position errors. The automatic correction dynamically corrects each data source using artificial sampling data or high-precision reference station data, thereby improving the accuracy of the overall data.
[0018] Preferably, the method for removing data redundancy is as follows:
[0019] If the difference between the observed values provided by the same parameter data from different sensors at timestamps i and j is within the threshold range (∣d i -d j ∣ < ∈), then retain one data source and remove the remaining duplicate records, as shown in Equation (1):
[0020]
[0021] where: d i and d j are the observed values of different sensors at the same timestamp, and ∈ is the set error threshold, and the error threshold can be set.
[0022] Preferably, the invalid value elimination includes invalid data elimination and outlier elimination. The method for invalid data elimination is as follows: For each type of sensor, preset its reasonable working range and eliminate data outside the reasonable working range;
[0023] The method for outlier elimination is as follows: Eliminate invalid data from the data after invalid data elimination through statistical methods.
[0024] Preferably, the method for filling missing values is as follows:
[0025] For numerical data, use mean filling, mode filling, or median filling;
[0026] For time series data, fill in the missing time point data through interpolation methods, and the interpolation methods include linear interpolation and / or spline interpolation.
[0027] Preferably, the noise suppression uses a combination of Kalman filtering and particle filtering to remove noise.
[0028] Preferably, the spatio-temporal registration and automatic correction include:
[0029] Spatio-temporal registration: For numerical data, first use a time synchronization method based on cross-correlation. Suppose there are two time series {z i} and {w i}, calculate their cross-correlation function R zw(τ) is as follows:
[0030]
[0031] where τ is the time offset, and the τ0 that makes R zw (τ) maximum is selected as the time alignment offset of the two sequences;
[0032] Then, for the time series data with discontinuities or uneven sampling, the interpolation method is used to resample the data onto a unified time axis, and the interpolation method includes linear interpolation and / or spline interpolation;
[0033] Automatic correction, constructing an error model as follows:
[0034] Assume that there is a systematic bias b in the sensor observation value z, and the true value is x, then there is:
[0035] z = x + b (16)
[0036] By comparing with the reference data x ref and establishing an error model, using the multi-sampling data to calculate the statistical characteristics of the bias, obtaining the initial estimate b0, and the reference data x ref is the industrial sampling data or the high-precision reference station data;
[0037] Real-time correction of sensor bias: Adopt an adaptive correction method based on Kalman filter. Let the current corrected estimated value be and the reference value be x ref,k , then the bias b k is corrected and updated as follows:
[0038] b k = (1 - λ)b k-1 + λ(z k - x ref,k ) (17)
[0039] where λ is the smoothing factor, and its value range is between 0 and 1, reflecting the influence weight of the reference data on the correction. The corrected observed value is calculated as:
[0040]
[0041] By continuously updating the bias b k , the data of each sensor gradually approaches the true value and can adapt to the dynamic changes of the environment and equipment status.
[0042] Preferably, the multi-level data fusion includes:
[0043] Underlying data fusion, using interpolation technology and normalization method to construct a standard data matrix to eliminate the differences in sampling frequency, spatial resolution and dimension between different data sources;
[0044] Feature-level fusion, used to obtain a deep joint representation of multi-modal data;
[0045] Decision-level fusion, based on the joint feature vector obtained from feature-level fusion, constructs a decision model to obtain the final prediction and warning output.
[0046] Further preferably, the underlying data fusion includes:
[0047] Space-time grid construction: Using a preset space-time grid, the target area is divided into fixed spatial units, and time is divided by a fixed period. All data is mapped to the space-time grid. For each grid unit, synchronization processing is performed according to the sampling time of the data source;
[0048] Data interpolation and normalization: For the missing data points in the grid, linear interpolation or cubic spline interpolation is used to supplement; at the same time, to eliminate the influence of the dimension of different data sources, normalization processing is used to unify all data into the same numerical interval.
[0049] Further preferably, the feature-level fusion includes:
[0050] Multi-modal feature extraction, for image data, a convolutional neural network is used to extract the spatial features of the image;
[0051] For time series data, a recurrent neural network or a long short-term memory network is used to extract dynamic features;
[0052] Joint feature representation and dynamic weighting, the spatial features and dynamic features are fused using a fully connected neural network to form a joint feature vector;
[0053] Feature dimensionality reduction and optimization, using principal component analysis or an autoencoder to optimize the joint feature vector and extract the most representative feature components.
[0054] Further preferably, the decision-level fusion includes:
[0055] Multi-model prediction construction, using several prediction models to predict carbon emissions and carbon sinks respectively;
[0056] Ensemble learning and dynamic weight optimization, using an ensemble learning method, the prediction results of the several models are weighted and fused to obtain the final output value.
[0057] The beneficial effects of the present invention include:
[0058] ; 1. Multi-source heterogeneous data fusion: The present invention adopts a multi-level data fusion technology, which can simultaneously process heterogeneous data from different observation modules (sky, tower, ground, well), breaking through the limitations of traditional observation technologies.
[0059] 2. Dynamic Adaptive Calibration and Fusion Mechanism: The present invention introduces a dynamic weight calculation and automatic calibration mechanism based on real-time feedback, significantly improving the accuracy and reliability of data fusion, especially suitable for complex and variable coal mining areas.
[0060] 3. Full-scale Carbon Emission and Carbon Sink Fusion: The present invention realizes the fusion of full-scale observation data from large scale, medium scale to small scale, providing a new precise data processing method for the comprehensive observation means of carbon emissions and carbon sinks in coal mining areas. Detailed Implementation Manner
[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below.
[0062] The embodiment proposes a multi-source heterogeneous data fusion method for the "sky tower ground well" carbon flux observation system, which is a multi-source heterogeneous data fusion method for the integrated carbon emission and carbon sink comprehensive observation technology in coal mining areas. By closely fusing, complementing and calibrating the carbon emission and carbon sink observation data at different scales, the full-scale, dynamic, real-time and high-precision observation of carbon emissions and carbon sinks in coal mining areas is realized. The main contents of this embodiment include:
[0063] 1. System Architecture Design
[0064] The present invention constructs an integrated carbon emission and carbon sink data acquisition and fusion system including a four-layer architecture of "sky tower ground well".
[0065] 1) The acquisition layer includes the sky layer, tower layer and ground well layer, corresponding to different data observation and acquisition modules respectively:
[0066] At the "sky" level, it includes two major modules.
[0067] Satellite remote sensing module: Using multi-spectral, high-resolution remote sensing satellites to periodically scan the mining area and its surrounding areas to obtain environmental parameters such as vegetation index, soil moisture, temperature, etc., and then predict the change of carbon flux in the mining area. Especially when conducting large-scale and long-term dynamic observations, satellite remote sensing can provide carbon flux data of the mining area and its surrounding areas, filling the large-scale observation technology of the mining area, and is suitable for continuously obtaining macroscopic data of carbon flux changes.
[0068] UAV remote sensing module: In areas where satellite data update is insufficient or the local environment is complex, use UAVs equipped with high-precision sensors to supplement the acquisition of local environmental images and radiation data. It is especially suitable for short-term and dynamic observations in local areas of the mining area, and has unique advantages in fine observations of local hot spots.
[0069] "Tower" level: Eddy covariance flux towers are deployed in key areas of the mining area. With the help of equipment such as gas flow sensors, temperature and humidity sensors, and weather stations, data on wind speed, temperature, carbon dioxide, and other greenhouse gas concentrations are collected in real time, and the carbon gas exchange flux in the atmosphere is calculated. The eddy covariance flux tower can observe the changes in carbon emissions and carbon sinks in the mining area in real time. Its observation range is usually 1.5 - 2 times the tower height, which depends on the surface roughness and turbulence characteristics. By measuring the time series data of wind speed and gas concentration at high frequencies and calculating their covariance, the vertical exchange rate of gas in the turbulence can be obtained, and then the gas flux in the vertical direction can be estimated.
[0070] "Ground well" level, including two major modules.
[0071] Ground observation module: By deploying a fixed or mobile sensor network and combining regular manual sampling, the release and absorption processes of carbon in the mining area can be accurately captured, including the impact of mining activities on the carbon sequestration capacity of soil and vegetation, and the key indicators of carbon emissions and carbon absorption in the surface environment can be observed in real time.
[0072] Underground observation module: The underground observation module is similar to the ground one. Explosion-proof carbon flux observation sensors that are resistant to high temperatures, waterproof and moisture-proof, and resistant to dust are deployed in key areas underground. Combining with manual sampling, underground carbon flux data and other relevant environmental parameters are obtained.
[0073] 2) The transmission layer uses wireless network and Internet of Things technologies to transmit the data collected by each layer to the central data processing platform in real time for processing and analysis.
[0074] 3) The fusion layer preprocesses, spatially and temporally registers, automatically corrects, and fuses the multi-source data collected. It realizes bottom-layer synchronous fusion, feature-level fusion, and decision-level fusion, providing high-quality data for subsequent analysis and decision-making.
[0075] 4) The application layer conducts real-time observation, risk assessment, and visual display of carbon emissions and carbon sinks based on the fused data. Reports, early warnings, and decision support are generated through the real-time observation results.
[0076] 2. Data preprocessing and automatic correction
[0077] To ensure the accuracy and consistency of multi-source heterogeneous data fusion, this embodiment proposes an innovative data preprocessing and automatic correction scheme. Through multiple steps and technical means, problems such as noise, spatio-temporal inconsistency, format differences, and systematic errors in the data collection process are solved, as follows:
[0078] 2.1 Data cleaning and noise suppression
[0079] Data cleaning and noise suppression are crucial steps in the data preprocessing process, aiming to improve the accuracy and reliability of multi-source heterogeneous data fusion. Since the raw data provided by various sensors and observation modules in coal mining areas (such as satellite remote sensing, UAV images, meteorological towers, ground sensors, etc.) have redundancy, missing values, anomalies, and noise, precise cleaning and noise suppression techniques must be used to ensure that all data meet the fusion standards, including removing data redundancy, eliminating invalid values, filling in missing values, and suppressing noise, as follows:
[0080] 2.1.1 Removal of data redundancy
[0081] First is the removal of data redundancy, whose goal is to remove redundant, invalid, or incorrect data and ensure the validity and consistency of the remaining data. Due to the problem of multi-module data synchronization, there may be duplicate records at the same timestamp or in the same area. To avoid the impact of data redundancy on subsequent analysis, the method of removing redundant data by timestamp de-duplication is adopted:
[0082] For duplicate records from different sensors, by comparing their timestamps (t i and t j ) and data values (d[[ID=']] i and d j ), using the data of the reference sensor as the standard, the duplicate data are deleted. If the difference in the observed values provided by sensors i and j at timestamp t is within the threshold range (∣d i -d j ∣<∈), then one data source is retained and the remaining duplicate records are removed. The method is as follows:
[0083]
[0084] where d i and d j are the observed values of different sensors at the same timestamp, and ∈ is the set error threshold. The setting of the error threshold needs to be determined according to the characteristics of the data source, the requirements for data quality, and the accuracy requirements of the application scenario. To ensure the scientificity and rationality of the error threshold, the threshold is set through statistical analysis. By statistically analyzing the error distribution of different data sources and based on the statistical characteristics such as the standard deviation, mean, and skewness of the data, the threshold can be set, and the error threshold can be set to 1 or 2 times the standard deviation.
[0085] In practical applications, due to factors such as the environment, equipment aging, and weather, the accuracy of data will change. Therefore, the error threshold needs to have the ability to be dynamically adjusted, and the adjustment of the error threshold is achieved through a real-time feedback adjustment mechanism. During the actual operation process, if the data error of a certain sensor or observation module continues to be too large, the system can automatically increase the data error threshold of this sensor to ensure that its impact on the data fusion result is small. At the same time, if the error of a certain data source becomes smaller, the system can appropriately reduce its error threshold to improve the accuracy of the data. The specific formula is as follows:
[0086] New Threshold=Old Threshold+δ×(Current Error-Expected Error) (2)
[0087] Among them, δ is the adjustment coefficient, Current Error is the error of the current data source, and Expected Error is the expected error. In the application of carbon emissions and carbon sink observations in coal mining areas, according to the sensor type, data acquisition frequency, and actual application requirements, the error threshold is initially set to 1% to 5%, and is further adjusted through experimental data. For key observation data (such as carbon emissions), more stringent thresholds (such as 0.1% to 1%) may need to be set to ensure high-precision observation results.
[0088] After completing data cleaning, spatial redundancy removal also needs to be carried out. When processing data from different modules (such as satellite remote sensing and UAV images), due to spatial overlap or redundant areas in the data, duplicate spatial data needs to be removed to ensure that the information of the same area is not calculated repeatedly. For the sky module (satellite remote sensing and UAV images), if the observation areas of the two overlap, the data with higher resolution is selected for retention according to the resolution of the overlapping area, and the low-resolution data is discarded. For redundancy removal based on regional matching, for data with overlapping areas, calculate the area of the overlapping area (for example, by calculating the pixel area of the overlapping area). If the overlapping area is large, the data with higher resolution is retained. The formula for deduplication based on the spatial overlapping area is as follows:
[0089] Suppose there are two types of data, D1 and D2, which correspond to satellite data and UAV data respectively. If there is an overlapping area R between the two, and the area A of the overlapping area R is greater than a certain threshold A threshold , then the data with higher resolution is retained. It can be expressed as:
[0090] If A R >A threshold ,then keep D high res and discard D low res (3)
[0091] 2.1.2 Rejection of Invalid Data and Outliers
[0092] During the data acquisition process, it is possible that the data values exceed the normal operating range of the device or violate physical common sense. Such data includes two categories:
[0093] Invalid data: refers to data that clearly does not conform to the predetermined physical range or device operating standards (e.g., a temperature reading of -100°C, or data exceeding the maximum measurement range of the instrument).
[0094] Outlier: refers to a value that significantly deviates from the majority of the data in the statistical distribution. Such data may be caused by accidental interference or systematic errors, but under certain conditions, it may also reflect real extreme phenomena.
[0095] In this embodiment, a step-by-step rejection strategy is adopted. First, the obvious invalid data is rejected, and then the statistical method is used to reject the outliers from the remaining data.
[0096] When rejecting invalid data, set the physical range limit: for each sensor, preset its reasonable operating range. For example, for a temperature sensor, the reasonable range is [T min , T max ; the retention formula is:
[0097] If x i < T min or x i > T max (4)
[0098] For the retained data, detect outliers through statistical methods, and adopt a dual selection based on the standard deviation method and the interquartile range (IQR) method to ensure the stability of the data.
[0099] First, conduct the standard deviation method and use statistical methods to reject the data with excessive deviation. Assume that the data follows a normal distribution, calculate the mean μ and standard deviation σ of the data set, and for any data point x i Make the following judgment:
[0100] |x i - μ| > 3σ (5)
[0101] If the above conditions are met, then x i is considered invalid data and should be removed from the data set.
[0102] After conducting the standard deviation method for detection, then conduct the IQR method, calculate the first quartile Q1 and the third quartile Q3 of the data, and define the interquartile range as:
[0103] IQR = Q3 - Q1 (6)
[0104] The outlier determination condition is as follows:
[0105] If x i <Q1 - 1.5×IQR or x i >Q3 + 1.5×IQR (7)
[0106] After eliminating invalid values, it is ensured that the dataset only contains data that conforms to physical meaning and statistical laws, providing a basis for subsequent missing value filling and noise suppression.
[0107] 2.1.3 Filling Missing Data
[0108] After eliminating invalid values, for the missing values in the dataset, interpolation methods are used for filling. Missing values are data point omissions caused by various reasons (such as sensor failures, power failure faults, data transmission losses, etc.) during the data acquisition process. The purpose of filling missing values is to fill the missing data through appropriate algorithms to keep the dataset complete and not affect the subsequent analysis process. The methods for filling missing values depend on the data type, data missing situation, and application requirements. For numerical data, if the amount of missing data is small, mean filling, mode filling, or median filling can be used. For time series data, interpolation methods are used to fill the missing time point data, such as linear interpolation or spline interpolation. For a missing data point x i , assuming its known data before and after are x i-1 and x i+1 , then linear interpolation is performed using the following formula:
[0109]
[0110] For non - linear interpolation of data, cubic spline interpolation methods can be used:
[0111] S(x) = a + bx + cx 2 + dx 3 (9)
[0112] Among them, a, b, c, d are coefficients solved by the least - squares method or numerical optimization algorithms.
[0113] For time - series data, interpolation methods can generally better maintain the continuity and trend of data than mean filling, especially when the changes between adjacent data points are relatively smooth.
[0114] 2.1.4 Noise Suppression
[0115] Noise suppression is another key step in ensuring data quality, especially important when dealing with high - frequency and dynamically changing data. The present invention adopts a combination of Kalman filtering and particle filtering to remove noise and ensure the smoothness and consistency of data.
[0116] The Kalman filter update formula is as follows:
[0117]
[0118] Wherein, is the estimated value at the current time, z k is the observed value at the current time, K k is the Kalman gain, and its calculation formula is:
[0119]
[0120] Wherein, P k-1 is the prior error covariance, and R is the covariance of the observation noise.
[0121] The particle filter is applicable to non-linear and non-Gaussian noise data. By resampling multiple particles, the particle filter can more accurately estimate the state of the system. The particle filter formula is as follows:
[0122]
[0123] Resampling step: According to the weight Weight i perform particle resampling.
[0124] 2.2 Spatiotemporal registration and automatic correction
[0125] In the process of multi-source data fusion, due to differences in sampling time, spatial resolution, and acquisition angle among different data acquisition modules (such as satellites, drones, ground and downhole sensors), precise spatiotemporal registration and automatic correction must be achieved to ensure the unity and high precision of the fused data. For this purpose, the present invention proposes the following two major modules: a spatiotemporal registration module and an automatic correction module.
[0126] 2.2.1 Spatiotemporal registration
[0127] The purpose of spatiotemporal registration is to convert data from different data sources into a unified spatiotemporal coordinate system and eliminate data inconsistencies caused by sampling time deviations and spatial position errors.
[0128] For time alignment, the sampling frequencies and timestamps of different data sources may have deviations. To ensure that the data is comparable on the same time scale, a time synchronization method based on cross-correlation is adopted, supplemented by interpolation techniques. The specific steps are as follows:
[0129] Based on the time synchronization method of cross-correlation, there are two time series {z i} and {w i}, and calculate their cross-correlation function R zw (τ) as follows:
[0130]
[0131] where τ is the time offset. Select τ0 that maximizes R zw (τ) as the time alignment offset of the two sequences.
[0132] Then interpolation calculation is carried out. For time series data with discontinuities or uneven sampling, linear interpolation or spline interpolation techniques are used to resample the data onto a unified time axis. The interpolation formula is (8) or (9).
[0133] For image data obtained in the air (such as satellite and drone images) and other data with geographical location information, a spatial registration method based on feature matching is adopted. The main steps include:
[0134] Feature point extraction and matching: Use mature algorithms (such as SIFT or SURF) to extract feature points in each image, and obtain a preliminary set of corresponding points through descriptor matching.
[0135] Robustly solve the transformation matrix: Use the Random Sample Consensus (RANSAC) algorithm to eliminate mis-matches and calculate the homography matrix H between the two images. This matrix satisfies the following relationship:
[0136] x′ = Hx (14)
[0137] where x is the point in the original image (represented in homogeneous coordinates) and x′ is the corresponding point in the target image. To obtain the optimal H, usually minimize the error function:
[0138]
[0139] Through the above time and space registration processes, data from different data sources are all transformed into a unified spatio-temporal coordinate system, ensuring that subsequent data fusion is based on consistent and accurate data.
[0140] 2.2.2 Automatic correction
[0141] In multi-source data acquisition, different sensors may produce systematic errors due to system deviations, environmental changes, and long-term drifts. To eliminate these errors, the present invention proposes an automatic correction method based on real-time reference feedback. The core idea is to use artificial sampling data or high-precision reference station data to dynamically correct each data source, thereby improving the accuracy of the overall data.
[0142] First, an error model is constructed. Assuming that there is a systematic deviation b in the sensor observation value z and the true value is x, then:
[0143] z = x + b (16)
[0144] By comparing with the reference data x refCompare and establish an error model. Calculate the statistical characteristics of the deviation using multiple sampled data to obtain the initial estimate b0.
[0145] In order to correct the sensor deviation in real time during the data acquisition process, this embodiment adopts an adaptive correction method based on Kalman filtering. In addition to noise reduction and smoothing, Kalman filtering can also be used to estimate the system deviation. Let the current corrected estimate be The reference value is x ref,k , then the deviation b is updated through correction:
[0146] b k =(1 - λ)b k-1 +λ(z k -x ref,k ) (17)
[0147] where λ is the smoothing factor, and its value range is between 0 and 1, reflecting the influence weight of the reference data on the correction. The corrected observed value is calculated as:
[0148]
[0149] By continuously updating the deviation b k , the data of each sensor gradually approaches the true value, and can adapt to the dynamic changes of the environment and equipment status.
[0150] Considering the possible slight errors in the spatio-temporal registration process, the automatic correction is combined with the spatio-temporal registration. For the image data or point cloud data after spatial registration, the joint correction is realized by comparing the overlapping area with the reference map data. Let x i be the data point after registration, and the corresponding reference data point is x ref,i , N is the total number of data points, and the joint correction error function E is constructed:
[0151]
[0152] Adopt iterative optimization methods such as gradient descent or Levenberg-Marquardt algorithm to adjust the spatio-temporal registration parameters and sensor correction parameters to minimize EEE, so as to achieve the best joint correction effect.
[0153] 3. Multi-level data fusion algorithm
[0154] To make full use of the high-quality data obtained after data cleaning, spatio-temporal registration, and automatic calibration, the present invention adopts a multi-level data fusion algorithm. The basic idea is to refine multi-source heterogeneous data step by step from the original signal to form a unified high-dimensional feature representation, and finally output real-time observation and early warning information on carbon emissions and carbon sinks in coal mining areas through a decision-making model. This algorithm is mainly divided into three levels: bottom-layer data fusion, feature-level fusion, and decision-level fusion.
[0155] 3.1 Bottom-layer data fusion
[0156] In the bottom-layer data fusion stage, the data of each data source after preprocessing (including data cleaning, spatio-temporal registration, and automatic calibration) is mapped into a unified spatio-temporal grid. To eliminate the differences in sampling frequency, spatial resolution, and dimension between different data sources, the present invention uses interpolation technology and normalization methods to construct a standard data matrix. The key technologies and steps are as follows:
[0157] Spatio-temporal grid construction: Using a preset spatio-temporal grid, the target area is divided into fixed spatial units, and time is divided according to a fixed period. All data is mapped to this grid. For each grid cell, synchronization processing is performed according to the sampling time of the data source.
[0158] Data interpolation and normalization: For the missing data points in the grid, linear interpolation or cubic spline interpolation is used to supplement them. At the same time, to eliminate the influence of the dimension of different data sources, normalization processing is used to unify all data into the same numerical interval. Standardization (z-score) is as follows:
[0159]
[0160] where the mean is μ and the standard deviation is σ. The goal of low-level fusion is to provide consistent and comparable data input for subsequent feature extraction.
[0161] 3.2 Feature-level fusion
[0162] Feature-level fusion is one of the core innovations of the present invention, and mainly realizes the deep joint representation of multi-modal data through the following steps:
[0163] Multi-modal feature extraction: Different deep neural network models are respectively used to extract features from different data types. For the data obtained from the sky module (satellite remote sensing and UAV images), a convolutional neural network (CNN) is used to extract the spatial features of the image, and the expression is:
[0164] f i =CNN(I i ) (21)
[0165] where, I i represents the image input from the i-th source, and fi is its corresponding eigenvector.
[0166] For time series data (such as eddy covariance tower, ground and downhole sensor data), a recurrent neural network (RNN) or long short-term memory network (LSTM) is used to extract dynamic features, and the expression is:
[0167] g i = LSTM(s i ) (22)
[0168] where s i represents the time series data from the i-th source, and g i is the extracted time series feature.
[0169] Joint feature representation and dynamic weighting. To make full use of the complementary information of each data source, the present invention fuses the above features through a fully connected neural network (FCN) to form a joint feature vector. Let the joint feature of the i-th data source be:
[0170] F i = [f i ; g i (23)
[0171] To enable each data source to adaptively adjust its contribution during the fusion process, this embodiment introduces a dynamic weighting mechanism, and its weight w i The calculation formula is:
[0172]
[0173] where F i (t) is the feature vector of the i-th data source at time t; F j (t) represents the feature vector of the j-th data source after feature extraction at time t; F ref (t) is the reference feature vector, which can be determined by high-precision data or preliminary fusion results; λ is a regulation parameter; N is the total number of data sources.
[0174] Therefore, the fused joint feature vector F fused (t) is given by the following formula:
[0175]
[0176] Feature dimensionality reduction and optimization. To reduce the computational complexity brought by high-dimensional features, dimensionality reduction techniques such as principal component analysis (PCA) or autoencoders are used to optimize the joint feature vector, extract the most representative feature components, and further improve the generalization ability and operating efficiency of the subsequent model.
[0177] 3.3 Decision-level fusion
[0178] In the decision-level fusion stage, based on the joint feature vector obtained from the feature-level fusion, a decision model is constructed to achieve the final prediction and early warning output. The specific solutions include:
[0179] Construction of multiple model predictions. In the present invention, multiple prediction models (selecting prediction models such as random forest, support vector machine, and deep neural network) are used to predict carbon emissions and carbon sinks respectively. Let the output of the i-th model be y i .
[0180] Ensemble learning and dynamic weight optimization. To improve the prediction accuracy and robustness, an ensemble learning method is adopted to perform weighted fusion on the prediction results of each model. The final output value y fused is expressed as:
[0181]
[0182] where M is the total number of models, and α i is the dynamically adjusted weight, which is adjusted in real time through the Bayesian optimization method to adapt to data changes and ensure the final prediction accuracy.
[0183] 4. Comprehensive observation and application output
[0184] After data fusion, the present invention provides the following outputs through real-time analysis: Carbon emission heat map: By fusing multi-source data, a carbon emission distribution map of the mining area is generated in real time to help observe the hot spots of carbon emissions. Carbon sink distribution map: Displays the carbon sink situation in the coal mining area and predicts the changing trend of carbon sinks in different regions. Carbon emission and carbon sink report: Automatically generates a detailed carbon emission and carbon sink report to support decision-making and policy-making. Early warning system: Through real-time data observation, provides early warnings of abnormal carbon emissions to provide timely responses for environmental protection supervision in the mining area.
[0185] Of course, the present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. Multi-source heterogeneous data fusion method for the "Sky Tower Ground Well" carbon flux observation system, characterized in that, It includes the following steps: Step 1, data preprocessing and automatic correction, to solve the noise, spatio-temporal inconsistency, format differences and systematic errors in the process of multi-source heterogeneous data acquisition. The multi-source heterogeneous data come from five observation modules: satellite remote sensing, UAV remote sensing, eddy covariance flux tower, ground observation and underground observation respectively; Step 2, multi-level data fusion, refining the multi-source heterogeneous data from the original signal step by step to form a unified high-dimensional feature expression, and finally outputting the real-time observation data of carbon emissions and carbon sinks in the coal mining area through a decision-making model; Step 3, comprehensive observation and application output, conducting real-time analysis based on the real-time observation data, and outputting a carbon emission heat map, a carbon sink distribution map, a carbon emission and carbon sink report, and early warning information.
2. The multi-source heterogeneous data fusion method of the "Sky Tower Ground Well" carbon flux observation system according to claim 1, characterized in that Step 1 includes: Step 1.1, data cleaning and noise suppression, including removing data redundancy, eliminating invalid values, filling in missing values and suppressing noise; Step 1.2, spatio-temporal registration and automatic correction. The spatio-temporal registration converts the data from different data sources into a unified spatio-temporal coordinate system to eliminate data inconsistency caused by sampling time deviation and spatial position error. The automatic correction uses artificial sampling data or high-precision reference station data to dynamically correct each data source, thereby improving the accuracy of the overall data.
3. The multi-source heterogeneous data fusion method of the "sky-tower-ground-well" carbon flux observation system according to claim 2, characterized in that, The method for removing data redundancy is: If the difference in the observed values of the same parameter data from different sensors at time stamp t for i and j is within the threshold range (∣d i -d j ∣ < ∈), then one data source is retained and the remaining duplicate records are removed, as shown in Equation (1): Cleaned Data={(t1,d1),(t2,d2),…,(t n ,d n )} where t i ≠t j for i≠j and |d i -d j |<ò(1) where: d i and d j are the observed values of different sensors at the same timestamp, ∈ is the set error threshold, and the error threshold can be set.
4. The multi-source heterogeneous data fusion method of the "Sky Tower Ground Well" carbon flux observation system according to claim 2, characterized in that, The elimination of invalid values includes the elimination of invalid data and the elimination of outliers. The method for eliminating invalid data is: for each sensor, its reasonable working range is preset in advance, and the data outside the reasonable working range is eliminated; The method for eliminating outliers is: the data after eliminating invalid data is eliminated through statistical methods.
5. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 2, characterized in that, The method for filling in missing values is: For numerical data, mean filling, mode filling or median filling is adopted; For time series data, the missing time point data is filled by interpolation method, and the interpolation method includes linear interpolation and / or spline interpolation.
6. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 2, characterized in that, The noise suppression adopts a combination of Kalman filter and particle filter to remove noise.
7. The multi-source heterogeneous data fusion method of the "sky-tower-ground-well" carbon flux observation system according to claim 2, characterized in that, Spatio-temporal registration: For numerical data, first use a time synchronization method based on cross-correlation. Suppose there are two time series {z i} and {w i}, and calculate their cross-correlation function R zw (τ) as follows: where τ is the time offset, and the τ0 that maximizes R zw (τ) is selected as the time alignment offset of the two sequences; Then, for the time series data with discontinuity or uneven sampling, the interpolation method is used to resample the data onto a unified time axis, and the interpolation method includes linear interpolation and / or spline interpolation; Automatic correction, constructing an error model as follows: Assume that there is a systematic deviation b in the sensor observation value z, and the true value is x, then: z = x + b (16) By comparing with the reference data x ref An error model is established, and the statistical characteristics of the deviation are calculated using multiple sampled data to obtain the initial estimate b0. The reference data x ref is the industrial sampled data or the high-precision reference station data; Real-time correction of sensor deviation: An adaptive correction method based on Kalman filtering is adopted. Let the current estimated value after correction be The reference value is x ref,k , then the deviation b is updated through correction k : b k = (1 - λ)b k-1 + λ(z k - x ref,k ) (17) where λ is the smoothing factor, with a value range between 0 and 1, reflecting the influence weight of the reference data on the correction, and the corrected observed value is calculated as: By continuously updating the deviation b k , the data of each sensor gradually approaches the true value, and can adapt to the dynamic changes of the environment and equipment status.
8. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to any one of claims 1-7, characterized in that, The multi-level data fusion includes: Underlying data fusion, using interpolation technology and normalization method to construct a standard data matrix to eliminate the differences in sampling frequency, spatial resolution and dimension between different data sources; Feature-level fusion, used to obtain the deep joint representation of multi-modal data; Decision-level fusion, based on the joint feature vector obtained from feature-level fusion, constructing a decision-making model to obtain the final prediction and early warning output.
9. The multi-source heterogeneous data fusion method of the "sky tower and ground well" carbon flux observation system according to claim 8, characterized in that, The underlying data fusion includes: Space-time grid construction: Using a preset space-time grid, the target area is divided into fixed spatial units, and time is divided into fixed periods. All data is mapped to the space-time grid, and for each grid cell, synchronous processing is performed according to the sampling time of the data source; Data interpolation and normalization: For the missing data points in the grid, linear interpolation or cubic spline interpolation is used for supplementation; at the same time, to eliminate the influence of the dimensions of different data sources, normalization processing is used to unify all data into the same numerical interval.
10. The multi-source heterogeneous data fusion method of the "Sky Tower and Ground Well" carbon flux observation system according to claim 8, characterized in that, The feature-level fusion includes: Multi-modal feature extraction. For image data, a convolutional neural network is used to extract the spatial features of the image; For time-series data, a recurrent neural network or a long short-term memory network is used to extract dynamic features; Joint feature representation and dynamic weighting. The spatial features and dynamic features are fused using a fully connected neural network to form a joint feature vector; Feature dimensionality reduction and optimization. Principal component analysis or an autoencoder is used to optimize the joint feature vector and extract the most representative feature components.
11. The multi-source heterogeneous data fusion method of the "Sky Tower - Ground Well" carbon flux observation system according to claim 8, characterized in that, The decision-level fusion includes: Multi-model prediction construction. Several prediction models are used to predict carbon emissions and carbon sinks respectively; Ensemble learning and dynamic weight optimization. An ensemble learning method is used to perform weighted fusion on the prediction results of the several models to obtain the final output value.
Citation Information
Patent Citations
Urban department carbon emission estimation method and system based on satellite observation and GIS
CN116109191A
Carbon emission calculation method based on deep learning
CN119312040A
Mobile internet-based monitoring and warning system for big data of global earthquake geomagnetic anomalies, and monitoring and warning method
WO2016201759A1
Cited By
Reservoir carbon flux monitoring system and method
CN121453123A
Urban ecological toughness monitoring and early warning system integrating multi-mode environment sensor and climate monitoring device
CN121475327A
Adaptive fusion and parameter correction method and system for multi-source monitoring data
CN121479152A
A multi-source monitoring data adaptive fusion and parameter correction method and system
CN121479152B
Multi-component greenhouse gas flux observation data self-adaptive correction and completion method
CN122045556A