Carbon emission monitoring methods and systems based on multi-source monitoring data
By aligning carbon emission monitoring data in time and space and enhancing cross-modal physical consistency, the problem of timestamp misalignment in multi-source monitoring data was solved, achieving high-precision prediction and physical rationality of carbon emission rates, and effectively distinguishing emission contributions at different spatial scales.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE SECOND CONSTR OF CHINA CONSTR EIGHTH ENG DIV
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing carbon emission monitoring methods cannot effectively address the timestamp misalignment problem of multi-source monitoring data. They neglect the physical transmission lag effect between carbon emissions at the source, atmospheric cumulative response, and ground observation, resulting in poor model robustness and a lack of physical interpretability in prediction results. They cannot explicitly model the physical correlation between 'source strength-concentration-flux', and the carbon emission drivers at different spatial scales are not effectively distinguished.
By enhancing the spatiotemporal alignment and cross-modal physical consistency of carbon emission physical transmission processes, and utilizing atmospheric transport lag effects for time compensation, transmission weights, diffusion weights, and flux modulation weights are constructed to constrain and enhance multi-source features. Furthermore, through multi-scale feature decomposition and sparse activation, a dynamic lag compensation vector is generated to achieve adaptive prediction of carbon emission rates.
It achieves timestamp alignment of multi-source monitoring data, enhances the interpretability and robustness of the model, can effectively distinguish carbon emission contributions at different spatial scales, and the prediction results conform to basic physical laws, with high accuracy and source apportionment significance.
Smart Images

Figure CN122087772A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of carbon emission monitoring technology, specifically relating to a carbon emission monitoring method and system based on multi-source monitoring data. Background Technology
[0002] Against the backdrop of increasingly severe global climate change, accurate and timely monitoring of carbon emissions has become a key technological support for addressing climate change, formulating emission reduction policies, and assessing progress toward "dual carbon" targets. Traditional carbon emission monitoring methods are mainly divided into two categories: First, bottom-up inventory statistics based on emission factor methods. This method estimates total carbon emissions by collecting statistical data such as industrial energy consumption, traffic flow, and population activity, multiplying them by corresponding emission factors. While this method has clear physical meaning and source apportionment capabilities, it heavily relies on the completeness and timeliness of statistical data, typically exhibiting a lag of one to several years, and has coarse spatial resolution, making it difficult to reflect dynamic changes and local details of carbon emissions. Second, top-down monitoring based on satellite remote sensing inversion. This method obtains atmospheric carbon dioxide column concentrations using satellite-borne instruments and then uses atmospheric transport models to infer surface fluxes. This method enables large-scale, periodic monitoring, but is limited by satellite transit time, cloud cover, and uncertainties in the inversion algorithm, resulting in low temporal resolution and insufficient sensitivity to near-surface anthropogenic emissions. In recent years, with the development of deep learning technology, some studies have begun to attempt to integrate activity data, remote sensing data and ground observation data for carbon emission estimation.
[0003] Existing technologies suffer from the following problems: Firstly, they typically splice together activity data, remote sensing inversion data, and ground observation data after simple interpolation and spatial resampling at a unified timestamp. This ignores the physical transport lag effect between carbon emissions at their source, atmospheric cumulative response, and ground observation, resulting in the three types of data at the same timestamp not corresponding to the same physical process stage. Consequently, the model cannot stably learn the complete chain of change from "anthropogenic emissions-atmospheric transport and diffusion-local flux response." Secondly, directly splicing multi-source features into the network fails to explicitly model the physical correlation between "source strength-concentration-flux." The differences in the expression domains of the three types of features make the model prone to learning spurious statistical correlations rather than true physical laws. When anomalies or noise exist in a certain modality of data, erroneous signals directly contaminate the fused features, leading to poor model robustness and unsatisfactory prediction results. The methods lack physical interpretability; they use a uniform convolution kernel or fully connected layer to process carbon emission drivers at different spatial scales, ignoring the fundamental differences between industrial point sources (strong locality), traffic emissions (regional diffusion), and background concentrations (large-scale, slow variation). This leads to strong local emission signals being easily averaged or submerged by the large-scale background, making it difficult for the model to simultaneously capture emission contribution patterns at different scales, and resulting in significant information overlap between features at different scales. Furthermore, using regression layers with fixed parameters results in significant differences in the proportion of various sources' influence on total carbon emissions across different geographical units, seasons, and activity backgrounds, making it impossible for fixed parameters to adapt to dynamic changes in emission patterns. Additionally, existing methods lack physical constraints on the prediction results, potentially leading to outputs that violate basic physical principles, such as negative industrial source contributions or serious contradictions between predicted values and source contribution decomposition results. Summary of the Invention
[0004] To achieve the above objectives, the present invention employs the following technical solution: This invention provides a carbon emission monitoring method based on multi-source monitoring data, comprising the following steps: S1. Collect multi-source carbon emission monitoring data; the multi-source carbon emission monitoring data includes activity data, remote sensing inversion data, and ground observation data; S2. Multi-source carbon emission monitoring data are processed by enhancing spatiotemporal alignment and cross-modal physical consistency for the physical transmission process of carbon emissions. Multi-source features are mapped to a unified carbon emission driver-response semantic space to obtain physically enhanced feature vectors. S3. Perform multi-scale feature decomposition, scale-adaptive sparse activation and source contribution compression on the physical enhancement feature vector to obtain the source contribution feature vector and the compressed low-dimensional feature vector. S4. Adaptive regression prediction and physical constraint optimization are performed on the source contribution feature vector and the compressed low-dimensional feature vector to obtain the predicted carbon emission rate.
[0005] Furthermore, activity data typically characterizes the intensity of emission sources, remote sensing inversion data typically characterizes the cumulative response of emissions in the atmosphere, and ground observation data typically characterizes local fluxes and micrometeorological conditions. The same timestamp of these three types of data is not equivalent to the same physical process stage.
[0006] This invention performs time compensation on activity data based on atmospheric transport lag effects, and then performs physical correlation encoding with remote sensing inversion data and ground observation data to obtain a physical lag aligned feature vector: The study area was divided into multiple geographical units, and the spatial attribution of various data types was standardized; the temporal resolution was standardized, and basic preprocessing was performed on the original sequences; at the current timestamp Next, construct the first The original stitched feature vector of each geographic unit, including the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector; for the first The first geographical unit is used to determine the set of atmospheric transport source regions; for the first... For each transmission path in the atmospheric transmission source region set corresponding to a geographic unit, the transmission distance and transmission delay are calculated. Based on the delay results of each transmission path, lagged activity data features are extracted from the historical activity data sequence, and distance attenuation and atmospheric diffusion modulation weights are constructed to obtain the lagged activity data weighted features after distance attenuation and atmospheric diffusion modulation. The lagged activity data weighted feature sequence obtained from the same geographic unit at multiple historical moments and multiple transmission paths is input into a long short-term memory network to generate a dynamic lag compensation vector. The dynamic lag compensation vector is used to modulate the activity data feature vector at the current moment dimension by dimension and then concatenate it with the remote sensing inversion feature vector and the ground observation feature vector to obtain the physical lag aligned feature vector. The basic preprocessing includes: outlier removal, interpolation repair, and normalization.
[0007] Furthermore, after time delay compensation, although the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector are in a more reasonable time correspondence, the three types of features still have differences in their expression domains.
[0008] To further explicitly model the physical correlation between "source strength-concentration-flux", a cross-modal physical consistency enhancement mechanism is constructed. This mechanism constrains and enhances the three types of data in the physical lag alignment feature vector by using transmission weights, diffusion weights, and flux modulation weights. The physical consistency error is then used to correct the features, resulting in a physically enhanced feature vector. From the physically lag aligned feature vector, active data feature vector, remote sensing inversion feature vector, and ground observation feature vector are separated. Three fully connected mapping layers are set up to project these three feature vectors onto the same comparison dimension, resulting in an active transport descriptor, a remote sensing diffusion descriptor, and a ground flux descriptor. A transport weight vector corresponding to the active data is generated based on boundary layer height and wind field information. A diffusion weight vector corresponding to the remote sensing inversion feature vector is generated based on the similarity between the active transport descriptor and the remote sensing diffusion descriptor. A flux modulation weight vector corresponding to the ground observation feature vector is generated based on the difference between the active transport descriptor and the ground flux descriptor. The active data feature vector, remote sensing inversion feature vector, and ground observation feature vector are multiplied element-wise by their respective transport weight vector, diffusion weight vector, and flux modulation weight vector, and then concatenated along the feature dimension to obtain a physically consistent enhanced feature vector. A physical consistency error is constructed, and this error is used to perform gradient correction on the physically lag aligned feature vector to obtain the physically enhanced feature vector. The specific process of constructing the physical consistency error is as follows: extract comparable source intensity, concentration, and flux representations from the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector, respectively, and calculate the consistency error among the three representations; use the gradient of the consistency error with respect to the physical lag alignment feature vector as the correction direction, and use the physical consistency enhancement feature vector as the correction amplitude modulation term to update the physical lag alignment feature vector in one or more steps.
[0009] Furthermore, to distinguish carbon emission contributions at different spatial scales, the physical enhancement feature vectors of multiple geographic units at the same time are rearranged into local feature tensors according to spatial adjacency. A multi-branch convolution decomposition module is then used to extract features at the local, regional, and background scales, resulting in three scale component feature vectors: Centered on the geographic unit to be predicted, physical enhancement feature vectors within the surrounding spatial neighborhood are extracted and rearranged into local feature tensors. These local feature tensors are simultaneously input into three parallel convolutional branches to extract local-scale, regional-scale, and background-scale features, respectively, resulting in local-scale carbon emission feature vectors, regional-scale carbon emission feature vectors, and background-scale carbon emission feature vectors. The kernel sizes of the three parallel convolutional branches are 3, 7, and 15, respectively, and all include convolutional layers, batch normalization layers, and ReLU activation functions. During training, inter-scale orthogonality constraints are introduced for feature vectors of different scale components.
[0010] Furthermore, carbon emission characteristics at different scales exhibit varying sparsity. Local-scale characteristics typically show significant activation only in a few locations, regional-scale characteristics are moderately sparse, and background-scale characteristics are usually relatively flat. Applying the same activation function to all three scales often fails to meet the representational needs of different scales.
[0011] A sparse threshold is adaptively set for the statistical features of local-scale carbon emission feature vectors, regional-scale carbon emission feature vectors, and background-scale carbon emission feature vectors, and they are fused according to their scale contribution to obtain a fused sparse feature vector: Statistical analysis was performed on the local-scale, regional-scale, and background-scale carbon emission feature vectors, calculating their mean and standard deviation in the current training batch, and then standardizing them. A sparse activation function based on a Sigmoid gate was constructed from the standardized local-scale, regional-scale, and background-scale carbon emission feature vectors to obtain the corresponding sparse activation feature vectors. The gate response was then compared with the... The sparsity thresholds at each scale are compared. Parts below the threshold are directly suppressed, while parts above the threshold are retained and treated as valid responses. The scale contribution scores are calculated for the sparse activation feature vectors at each of the three scales, and the scale contribution scores are normalized to the fusion weights. The sparse activation feature vectors at each of the three scales are multiplied by their corresponding fusion weights and then concatenated along the feature dimensions to obtain the fused sparse feature vector.
[0012] Furthermore, carbon emission source categories are predefined, and a source category index set is established; a source contribution projection matrix is constructed, and the fused sparse feature vector is mapped to the source category space to obtain the source contribution feature vector; the source contribution feature vector is input into the compression mapping layer to obtain the compressed low-dimensional feature vector.
[0013] Furthermore, by controlling the weights and biases of the regression network through the source contribution feature vectors, the regression network can adaptively adjust its prediction strategy as the source composition of the samples changes. Read the contribution values of each source category from the source contribution feature vector; construct weight modulation coefficients for each source contribution value: first calculate the deviation between the contribution value of that source category and its global average, then input the deviation into the Sigmoid gate function to obtain the weight modulation coefficients in the range of 0 to 1; set the basic weight matrix and the weight modulation matrix corresponding to each source category for the regression network, and perform weighted correction on the basic weight matrix according to the weight modulation coefficients of each source category to obtain the dynamic weight matrix; further construct bias modulation coefficients for each source contribution value: ... The source contribution value is multiplied by the corresponding bias sensitivity coefficient and then input into the hyperbolic tangent function to obtain the positive and negative bounded bias modulation coefficients. The basic bias vector and the bias modulation vector corresponding to each source category are set for the regression network. Based on the bias modulation coefficients of each source category, the basic bias vector is corrected to obtain the dynamic bias vector, which is used for the regression output bias term of the current sample.
[0014] Furthermore, after obtaining the dynamic weight matrix and dynamic bias vector, regression prediction is performed on the compressed low-dimensional feature vector: The compressed low-dimensional feature vector is input into the dynamic regression layer to calculate the predicted carbon emission rate. A physical prior prediction is constructed based on the source contribution feature vector: an emission coefficient is pre-set for each source type, based on historical statistical data, carbon emission accounting experience, or expert knowledge. The contribution values of each source type are multiplied by the corresponding emission coefficient and summed to obtain the physical prior prediction. A source contribution consistency constraint term is calculated to measure the difference between the predicted carbon emission rate and the physical prior prediction. A symbolic physical prior is set for different source categories, and a symbolic consistency constraint term is calculated. The source contribution consistency constraint term and the symbolic consistency constraint term are added to obtain a physical constraint regularization term, which is used to ensure that the prediction results are consistent with the source contribution decomposition results and with the physical direction prior of the source categories.
[0015] Furthermore, during training, the prediction error, physical constraint regularization term, and scale orthogonality constraint loss are combined to construct a composite loss function, which is then used to uniformly optimize all network parameters. A training sample set is constructed, with each training sample including at least the activity data feature vector, remote sensing inversion feature vector, ground observation feature vector, and real carbon emission label value. Forward propagation is performed on each training batch to obtain the carbon emission rate prediction value for each sample. The mean squared error loss is calculated based on the carbon emission rate prediction value and the real carbon emission label value. The physical constraint regularization term and the scale orthogonality constraint loss are calculated. The mean squared error loss, the physical constraint regularization term, and the scale orthogonality constraint loss are weighted and summed to obtain the composite loss function.
[0016] The present invention also provides a carbon emission monitoring system based on multi-source monitoring data, which executes the above-described carbon emission monitoring method based on multi-source monitoring data, including: Data acquisition module: used to collect multi-source carbon emission monitoring data; the multi-source carbon emission monitoring data includes activity data, remote sensing inversion data and ground observation data; Feature enhancement module: Processes multi-source carbon emission monitoring data by enhancing spatiotemporal alignment and cross-modal physical consistency for the physical transmission process of carbon emissions, and maps multi-source features to a unified carbon emission driver-response semantic space to obtain physically enhanced feature vectors; The eigenvalue decomposition and activation module is used to perform multi-scale eigenvalue decomposition, scale-adaptive sparse activation, and source contribution compression on the physically enhanced feature vector to obtain the source contribution feature vector and the compressed low-dimensional feature vector. Carbon emission rate prediction module: Used to perform adaptive regression prediction and physical constraint optimization on source contribution feature vectors and compressed low-dimensional feature vectors to obtain carbon emission rate prediction values.
[0017] The advantages of this invention are: This invention identifies the set of atmospheric transport source regions through backward trajectory analysis and wind field information, calculates the transmission distance and time lag of each transmission path, extracts lag features from historical activity data sequences, and generates dynamic compensation vectors using a long short-term memory network. The current activity data is then subjected to dimension-wise time compensation before being stitched together with remote sensing and ground data, achieving cross-modal temporal alignment based on atmospheric physical transport processes. This solves the temporal misalignment problem caused by transmission lag effects in the three types of data. Activity, remote sensing, and ground features are separated from the physical lag alignment features and projected onto a unified comparison dimension. Transmission weights are generated using boundary layer height and wind field information, diffusion weights are generated based on the similarity between activity and remote sensing descriptors, and flux modulation weights are generated based on the differences between activity and ground descriptors. Dimension-wise constraint enhancement is applied to the three types of data. Furthermore, the consistency error among "source strength-concentration-flux" is calculated, and this error is used to perform gradient correction on the features, explicitly adjusting the output features towards physically reasonable directions, enhancing the model's interpretability and robustness to anomaly observations. The invention also extracts features centered on the unit to be predicted. By taking the spatial neighborhood feature tensor, local scale, regional scale, and background scale features are extracted through three parallel convolutional branches. An orthogonality constraint between scales is introduced to force different branches to learn independent contribution patterns. Then, a sparse threshold is adaptively set based on the statistical characteristics of each scale for gated activation. Finally, the high-dimensional sparse features are mapped to predefined category spaces such as industrial sources, transportation sources, agricultural sources, and natural sources through the source contribution projection matrix, achieving a compact representation with source resolution significance and effectively distinguishing emission contributions at different spatial scales. Based on the contribution values of each source category in the source contribution feature vector, combined with the global average contribution level and learnable sensitivity coefficients, the weight matrix and bias vector of the regression layer are dynamically generated, enabling the regression network to adaptively adjust the prediction strategy according to changes in the source composition of the samples. Simultaneously, a source contribution consistency constraint is introduced into the loss function to ensure that the predicted value is consistent with the prior value of a linear combination of source contributions. A sign consistency constraint is introduced to penalize prediction results that violate prior physical directions such as positive for industrial sources and negative for natural sources, ensuring that the model output has both high accuracy and conforms to basic physical laws. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0019] Figure 1 This is a flowchart of the steps of the invention method; Figure 2 This is an example diagram of the atmospheric transport source region set and backward trajectory of the present invention; Figure 3 This is the heatmap corresponding to the physical hysteresis alignment feature vector of the present invention; Figure 4 An example diagram of the source contribution feature vector map for a sample of this invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1 In this embodiment, as Figure 1 As shown, this invention provides a carbon emission monitoring method based on multi-source monitoring data, the specific steps of which include: S1. Training Data Collection and Labeling for Multi-Source Carbon Emission Monitoring The carbon emission monitoring method based on multi-source monitoring data proposed in this invention requires the prior collection of a training dataset covering the study area and possessing spatiotemporal consistency. This dataset is used for subsequent spatiotemporal alignment, physical augmentation, multi-scale decomposition, and parameter learning of the adaptive regression network. The training data collection process revolves around three types of multi-source data: activity data, remote sensing inversion data, and ground observation data. Simultaneously, real carbon emission label values are collected for model supervised training.
[0022] The collection of activity data focuses on the intensity of anthropogenic emission sources, encompassing indicators that characterize the driving factors of carbon emissions from human activities, such as industrial heat source intensity, traffic network density, nighttime light index, and energy consumption intensity. Industrial heat source intensity can be calculated from continuous emission monitoring systems or hourly coal consumption, oil consumption, and natural gas consumption reported by enterprises. For point sources (such as factories and power plants), the hourly activity level is recorded according to their latitude and longitude coordinates. For linear sources (such as roads), traffic flow and average vehicle speed are calculated hourly for each road segment based on traffic flow monitoring equipment or vehicle GPS trajectory data, thereby estimating traffic emission intensity. For area sources (such as administrative regions), monthly or annual total energy consumption is obtained from statistical yearbooks or energy balance sheets, and then decomposed into hourly activity data using time downscaling models (such as using nighttime light index or temperature data).
[0023] b>The acquisition of remote sensing inversion data mainly relies on satellite remote sensing products, including environmental parameters such as carbon dioxide column concentration, chlorophyll fluorescence, land surface temperature, and boundary layer height. Common satellite data sources include OCO-2, GOSAT, MODIS, and Sentinel series satellites. During acquisition, raw raster data needs to be obtained. Each raster corresponds to a regular spatial resolution (e.g., 0.1 degrees × 0.1 degrees). The temporal resolution is usually once a day or once every few days. For missing data periods, the original timestamps are retained and uniformly supplemented in subsequent preprocessing.
[0024] c> The collection of ground observation data relies on flux towers, meteorological stations and air quality monitoring stations to obtain local observation information such as net ecosystem exchange, near-ground wind speed, wind direction, temperature and humidity. The data from ground stations are usually hourly or every three hours. When collecting data, it is necessary to record the latitude and longitude coordinates, observation time and measured values of each physical quantity for each station.
[0025] The collection and labeling of actual carbon emission values employs a combined approach of flux tower observation conversion and carbon emission accounting: For geographic units with flux towers, the net carbon exchange observed by the flux towers is converted into net ecosystem exchange using methods such as nighttime respiration separation. This is then combined with land use type to deduce the anthropogenic carbon emission rate, which serves as the actual carbon emission label for that unit. For areas without flux tower coverage, carbon emission accounting based on emission factors is used. This involves multiplying energy consumption in activity data by the corresponding emission factor (such as the default emission factor for coal, oil, and natural gas) to obtain the hourly carbon emission rate, which is then corrected for deviations from neighboring flux tower observations before being used as the label. The labeling category is a continuous numerical carbon emission rate, expressed in tons of carbon dioxide per hour. Each training sample corresponds to the actual carbon emission rate value of a geographic unit at a specific timestamp.
[0026] To ensure consistency between the labels and the input data in time and space, all collected activity data, remote sensing inversion data, ground observation data, and label values are pre-aligned according to the geographic unit division method (such as a 1 km × 1 km regular grid) and time resolution (such as 1 hour, 3 hours, or 1 day) set in the subsequent spatiotemporal alignment steps. Each sample contains three types of multi-source feature vectors under the same geographic unit and the same timestamp, as well as the corresponding real carbon emission label.
[0027] S2. Spatiotemporal alignment and heterogeneity perception enhancement of multi-source carbon emission monitoring data Multi-source carbon emission monitoring data includes activity data, remote sensing inversion data, and ground observation data. These three types of data differ significantly in terms of temporal sampling frequency, spatial resolution, physical meaning, and response time lag. Directly stitching these data into the model using only uniform temporal interpolation and uniform spatial resampling can easily confuse "anthropogenic emission-driven," "atmospheric cumulative response," and "local flux observation" information into the same category, causing the model to be unable to stably learn the chain of carbon emission changes from the source to transmission and then to the observation point.
[0028] Before the data enters the backbone network, this invention first performs spatiotemporal alignment and cross-modal physical consistency enhancement for the physical transmission process of carbon emissions, mapping multi-source features to a unified carbon emission driver-response semantic space to obtain physically enhanced feature vectors. The specific steps are as follows: S201, Spatiotemporal Alignment Feature Coding Based on Physical Transmission Pathways of Carbon Emissions Activity data typically characterizes the intensity of emission sources, remote sensing inversion data typically characterizes the cumulative response of emissions in the atmosphere, and ground observation data typically characterizes local fluxes and micrometeorological conditions. The same timestamp of these three types of data is not equivalent to the same physical process stage.
[0029] This invention utilizes the atmospheric transport lag effect to perform time compensation on activity data, and then performs physical correlation encoding with remote sensing inversion data and ground observation data to obtain a physically lag aligned feature vector. The specific steps are as follows: 1> Divide the study area into multiple geographical units and standardize the spatial attribution of various data types. Specifically, Geographic units can be regular grids, or they can be administrative regions, industrial parks, or watershed units; Spatial mapping is performed on activity data, remote sensing inversion data, and ground observation data respectively, so that the same geographic unit can generate corresponding feature records at the same timestamp.
[0030] In practical implementation, spatial mapping refers to assigning each data point to a corresponding geographic unit according to the geographic unit division method of multi-source data (activity data, remote sensing inversion data, and ground observation data). Specifically, Activity data exists in the form of point sources (factories, power plants), line sources (roads), or area sources (administrative region statistics). For point sources, the geographical unit they fall into is determined based on their latitude and longitude coordinates, and the emission intensity of that point source is accumulated to that unit. For line sources, the line segments are allocated to multiple covered units according to their length ratio. For area sources, they are directly mapped according to the administrative division correspondence. Remote sensing inversion data consists of regular rasters (e.g., 0.1°×0.1°). The intersection of the center point or coverage area of each raster with the geographic unit is calculated, and the raster value is assigned to the geographic unit using an area-weighted average method. Ground observation data consists of discrete stations. The observation values of each station are directly assigned to the geographic unit to which the station is located. If there are multiple stations within a unit, the average value can be taken.
[0031] In one embodiment, as an example, the study area is divided into a regular grid of 1km × 1km. A factory with coordinates (longitude 116.3°, latitude 39.9°) falls within the grid (105, 87). Therefore, the factory's hourly activity data (such as coal consumption) is added to the activity data feature vector of that grid. A remote sensing... The column density grid covers an area of (116.2°-116.3°, 39.8°-39.9°) and intersects with grid (105, 87) in an area of 0.05 km² (grid area 1 km²). Therefore, the grid value is multiplied by 0.05 and included in the remote sensing features of grid (105, 87).
[0032] 2. Unify the temporal resolution and perform basic preprocessing on the original sequence. Specifically, The time resolution can be standardized to 1 hour, 3 hours, or 1 day. Activity data is aggregated according to timestamps. For remote sensing inversion data, missing time periods are filled by the most recent time or by the average of the sliding window. For ground observation data, outliers are removed by thresholding and then interpolated for repair.
[0033] It should be noted that unified temporal resolution refers to standardizing the temporal sampling intervals of the three types of data to the same time step, so that feature concatenation can be performed at the same timestamp. Specifically, 1 hour: One data point every hour on the hour, suitable for high dynamic monitoring (such as industrial parks, urban traffic); 3-hour intervals: One data point every 3 hours (e.g., 0:00, 3:00, 6:00, etc.), suitable for monitoring in mesoscale regions; 1 day: One data point per day (e.g., daily average), suitable for long-term trend analysis or areas with sparse data.
[0034] In one embodiment, for example, regional activity data is hourly (with values every hour), remote sensing inversion data is retrieved daily (e.g., satellite transit times), and ground observation data is retrieved every 3 hours. If a uniform 3-hour resolution is chosen, then: Activity data: Average the hourly values over a 3-hour window to generate data points such as 0 o'clock (average of 0-2 o'clock) and 3 o'clock (average of 3-5 o'clock); Remote sensing data: Assign the daily value to all 3-hour time points of that day (e.g., 0:00, 3:00, ..., 21:00 all use the same daily value); Ground observation data: raw values are directly retained every 3 hours, and if a certain time is missing, it is filled in with the most recent time.
[0035] Furthermore, normalization is performed on the three types of features respectively. The normalization method can be min-max normalization or Z-score standardization to avoid training instability caused by directly inputting different units into the network.
[0036] 3> At the current timestamp Next, construct the first The original spliced feature vector of each geographic unit , represented as , dimension ,in, Indicates the first The activity data feature vector of each geographic unit (after normalization) has dimensions of [dimensions missing]. It characterizes anthropogenic driving factors such as industrial heat source intensity, traffic network density, nighttime light index, and energy consumption intensity; Indicates the first The remote sensing inversion feature vector of each geographic unit (after normalization), with dimensions of [dimensions missing]. , characterization Column concentration, chlorophyll fluorescence, surface temperature, boundary layer height and other remote sensing environmental parameters; Indicates the first The ground observation feature vector of each geographic unit (after normalization) has dimensions of [dimensions missing]. It represents ground-based observational information such as net ecosystem exchange, near-ground wind speed, wind direction, temperature, and humidity; This indicates a concatenation operation along the feature dimension; This represents a geographic unit index, with values... to , This indicates the total number of geographical units.
[0037] It should be noted that the original concatenated feature vector This will serve as the basis input for subsequent atmospheric transport lag compensation, whereby... Used to extract historical lagging activity data and generate dynamic compensation vectors. and Then after compensation and after modulation Reassemble.
[0038] 4> Regarding the first For each geographical unit, a set of atmospheric transport source regions that significantly influence remote sensing and ground-based observations is identified; this set is denoted as […]. , Indicates the first The set of atmospheric transport source regions corresponding to a geographical unit can be determined through backward trajectory analysis, prevailing wind direction statistics, or meteorological reanalysis data.
[0039] In one implementation, based on wind field data from 24 to 72 hours prior to the current timestamp, the main originating region of the air mass is tracked, and upwind geographical units that contribute significantly to the current geographical unit are added. .
[0040] In one embodiment, as an example, suppose the study area is a city, the geographic unit is a 1km grid, and the current focus grid is... (Located in the city center), wind field data (wind speed, wind direction) for the past 72 hours was obtained using meteorological reanalysis data (such as ERA5). The HYSPLIT model was used with a grid. The center point is the receptor point. The backward trajectory of the air mass is simulated, and one trajectory is output every 6 hours, resulting in a total of 12 trajectories. The upwind grids traversed by each trajectory are counted, and upwind grids that appear more than 3 times (or contribute more than 5%) are added to the set. For example, the trajectory shows that the air mass mainly originated from the industrial park in the southwest (grid). ) and the main traffic artery in the southeast direction (grid) ),but If the prevailing wind direction is stable, it can also be simplified to adding the grid within a certain distance upwind of the prevailing wind direction. .
[0041] 5> For each transmission path, calculate the transmission distance and transmission delay. Specifically, Transmission distance is denoted as , Indicates the first The first geographical unit and the first The spatial distance between transmission source areas can be expressed in kilometers. Transmission delay is denoted as , Indicates the first The effects of emissions from one transmission source region spread to the next... The time difference required for each geographical unit can be determined by comprehensively considering spatial distance, prevailing wind speed, wind direction consistency, and boundary layer stability.
[0042] In practice, the basic propagation time can be estimated based on wind speed first, and then the propagation efficiency can be corrected based on the boundary layer height. The higher the boundary layer height, the more complete the vertical mixing, and the propagation impact can be appropriately amplified. The lower the boundary layer height, the propagation impact can be appropriately weakened.
[0043] In one embodiment, for example, suppose a geographical unit (Receptor) and upwind unit Straight-line distance between (source regions) kilometer, prevailing wind speed m / s (18 km / h), wind direction consistency coefficient (Indicates that the angle between the wind direction and the direction of the line is less than 30°), boundary layer height Meters, then the basic propagation time hour, boundary layer correction factor (Vertical mixing is sufficient above 500 meters, resulting in high propagation efficiency), and the final transmission time delay. The time is rounded to the nearest 2 hours. If the boundary layer height is only 200 meters, then... , The hour indicates that propagation is slower when the atmosphere is stable.
[0044] 6> Based on the time delay results of each transmission path, extract the lagged activity data features from the historical activity data sequence, and construct distance attenuation and atmospheric diffusion modulation weights. Specifically, For the first Each geographical unit, reading time point The corresponding activity data feature vector, as along the first Historical emission characterization after propagation along each path; then, following the principle that "the greater the distance, the smaller the weight; the higher the boundary layer, the more significant the diffusion effect," a modulation weight is calculated for each path; the diffusion length scale can be denoted as... , This represents a hyperparameter that controls the spatial attenuation range; an example value is 50 kilometers.
[0045] 7> Input the weighted feature sequence of lagged activity data obtained from the same geographic unit at multiple historical moments and through multiple transmission paths into a long short-term memory network to generate a dynamic lag compensation vector. , dimension This is used to perform dimension-wise time compensation on the feature vector of the activity data at the current moment; Long Short-Term Memory (LSTM) networks are used to learn the accumulation, lag, and decay patterns of activity emission signals over time.
[0046] In practical implementation, sub-step 6 yields the weighted characteristics of the hysteresis activity data across multiple transmission paths, after distance attenuation and atmospheric diffusion modulation; that is, for each path... This yields a weighted vector of lagged activity data. ,in, To modulate the weights, then these Arranged chronologically (multiple paths may exist for the same timestamp) to form a sequence. It serves as the input to the Long Short-Term Memory (LSTM) network.
[0047] In one implementation, the input dimension of the Long Short-Term Memory network is... (Active data feature vector dimension), including a 2-layer stacked long short-term memory network (the output of the first layer is used as the input of the second layer), with an output dimension of (The dimension is consistent with the feature vector of the active data, used for dimensional compensation), the activation function adopts the hyperbolic tangent function (inside the long short-term memory network) and the sigmoid activation function (gating).
[0048] 8> Utilizing dynamic lag compensation vectors The feature vector of the activity data at the current moment (i.e. The feature vector is then modulated dimension-wise (by element-wise multiplication) and concatenated with the remote sensing inversion feature vector and the ground observation feature vector to obtain the physically lag aligned feature vector. , dimension This characterizes the multi-source fusion features after time-series compensation has been completed.
[0049] In practical implementation, the dimension-by-dimensional modulation method uses the feature vector of the active data. With dynamic hysteresis compensation vector Multiply element by element, then multiply and After splicing, nonlinear activation is performed by parameterizing the linear unit to maintain the main positive values of emission-related features while allowing a small number of negative values to pass through, characterizing carbon sink or net absorption.
[0050] It should be noted that remote sensing inversion data often does not reflect current instantaneous emissions, but rather the comprehensive response of several historical emissions after diffusion and transport. This invention does not simply perform the same timestamp splicing on the three types of data, but first performs physical time compensation on the activity data based on atmospheric transport lag relationships, and then performs cross-modal fusion. By compensating before fusion, the misalignment problem caused by "rapid changes in activity intensity, lag in remote sensing response, and localization of ground observation" can be significantly reduced, thereby improving the ability of subsequent regression models to model carbon emission transmission chains.
[0051] In one embodiment, such as Figure 2 As shown, this figure illustrates an example of an atmospheric transport source region set and backward trajectory. It presents the results of backward trajectory analysis based on the HYSPLIT model, used to identify the set of upwind source regions that significantly influence the recipient geographic unit (marked with a red asterisk). The horizontal and vertical axes represent spatial grid cells (unit: cells), with each blue dot representing a source region. Light blue arrows indicate the background wind field. In actual monitoring, the remote sensing response and ground observations of the current geographic unit are not solely determined by local emissions, but are influenced by the combined effects of multiple historical source regions at different distances.
[0052] S202, Enhanced Physical Constraints on Cross-Modal Carbon Emission Correlation Characteristics After time delay compensation, although the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector are in a more reasonable time correspondence, the three types of features still have differences in their expression domains.
[0053] To further explicitly model the physical correlation between "source strength-concentration-flux", this invention constructs a cross-modal physical consistency enhancement mechanism. This mechanism enhances the constraints on three types of data through transmission weights, diffusion weights, and flux modulation weights. Then, it uses physical consistency errors to correct the features, resulting in a physically enhanced feature vector. The specific steps are as follows: 1> Aligning feature vectors from physical lag The activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector are separated, and intermediate descriptors required for cross-modal comparison are constructed for each. Specifically, Three fully connected mapping layers are set up to project the feature vectors of active data, remote sensing inversion, and ground observation onto the same comparison dimension. We obtained the activity transport descriptor, the remote sensing diffusion descriptor, and the ground flux descriptor, among which, This indicates the dimension for cross-modal comparison; an example value is 32.
[0054] In practical implementation, due to It is obtained by splicing, and its dimensions are... During separation, it is only necessary to slice according to the known dimensions of each part, that is: Activity data feature vector: (forward (number of elements), among which, These are activity characteristics that have already undergone dynamic lag compensation, i.e. ; Remote sensing inversion feature vectors: ; Ground observation feature vector: .
[0055] It should be noted that the original dimensions of the feature vectors of activity data, remote sensing inversion feature vectors, and ground observation feature vectors are usually different. Directly subtracting or directly calculating the distance can easily lead to problems of inconsistent dimensions and unclear physical meaning. It is easier to calculate similarity or difference after projecting them to a unified comparison space.
[0056] 2> Generate a transmission weight vector corresponding to the activity data based on the boundary layer height and wind field information (including wind speed and direction). Specifically, the meteorological parameters are associated with the activity data feature vector dimension by dimension. A small fully connected network is used to generate a transmission weight vector with the same dimension as the activity features. The transmission weight vector is denoted as... , The dimension-wise modulation weights characterizing the feature vectors of activity data are used to emphasize the emission components of activities that are more likely to have a downwind impact.
[0057] It should be noted that the boundary layer height reflects the vertical mixing capacity of the atmosphere, and wind field information includes wind speed and wind direction (represented by sine and cosine components), which are key meteorological parameters affecting atmospheric transport efficiency.
[0058] 3> Based on the similarity between the activity transport descriptor and the remote sensing diffusion descriptor, generate a diffusion weight vector corresponding to the remote sensing inversion feature vector. The diffusion weight vector is denoted as... , The dimension-wise modulation weights characterize the remote sensing inversion feature vectors. Their dimensions are consistent with those of the remote sensing inversion feature vectors, reflecting the degree of matching between the remote sensing observation response and the activity emission drive.
[0059] In practical implementation, the activity transport descriptor is computed. With remote sensing diffusion descriptors difference vector (dimension is) Then, the difference vector is input into a small fully connected network (e.g., a single-layer linear mapping with Sigmoid activation), and the output dimension is... The vector, as the diffusion weight vector ,but Each element independently modulates the corresponding dimension of the remote sensing feature vector, reflecting the degree of consistency between emission drive and remote sensing response on different feature dimensions.
[0060] 4> Based on the difference between the active transport descriptor and the ground flux descriptor, a flux modulation weight vector corresponding to the ground observation feature vector is generated. This flux modulation weight vector is denoted as... , The dimension-wise modulated weights of the surface observation feature vectors are characterized by having the same dimension as the surface observation feature vectors, and are used to adjust the degree of contribution of surface observations to the final features.
[0061] In practical implementation, the difference between the active transport descriptor and the ground flux descriptor reflects the inconsistency between "emission-driven" and "local flux observation". The difference between the active transport descriptor and the ground flux descriptor is input into the fully connected layer and then constrained to the range of 0 to 1 by the Sigmoid function. When the active emission-driven and ground flux performance are consistent, the corresponding flux modulation weight is increased; when the difference between the two is large, the corresponding flux modulation weight is decreased to reduce the interference of abnormal ground observations on the fusion features.
[0062] 5> Multiply the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector element-wise with the corresponding transmission weight vector, diffusion weight vector, and flux modulation weight vector, respectively, and then concatenate them along the feature dimensions to obtain the physical consistency enhancement feature vector. , dimension It is used to strengthen the parts of the three types of data that have a physical correspondence and to suppress the parts that are inconsistent with each other.
[0063] 6> Construct a physical consistency error and use this error to perform gradient correction on the physical hysteresis alignment feature vector to obtain a physically enhanced feature vector. Dimensions This serves as the input for the subsequent multi-scale sparse decomposition module.
[0064] In practical implementation, firstly, comparable source strength, concentration, and flux representations are extracted from the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector, respectively. Then, the consistency error among the three is calculated. The larger the consistency error, the more significantly the current fusion result deviates from the physical chain of "source strength-concentration-flux". Next, the gradient of this consistency error with respect to the physically lag-aligned feature vector is used as the correction direction, and the physically consistent enhancement feature vector is used as the correction amplitude modulation term to update the physically lag-aligned feature vector in one or more steps, thereby making the output features closer to the physically consistent state. Preferably, the correction step size is denoted as... The example value is 0.01.
[0065] In the specific implementation, three independent fully connected mapping layers are used to map the feature vectors of the active data after dynamic hysteresis compensation. Remote sensing inversion feature vectors Ground observation feature vector Mapped to the same low-dimensional physical semantic space (dimension) (e.g., 8). The mapped vectors are denoted as follows: (Source strength characterization) (Concentration characterization) (Fluidity representation), the parameters of these mapping layers are all learnable parameters.
[0066] In one implementation, consistency error is defined as a measure of the difference among the three, expressed as: ,in, It is the L2 norm; the smaller this error, the more physically consistent the "source strength-concentration-flux" ratio is.
[0067] It should be noted that updating the physically lag aligned feature vector is similar to gradient descent, but it operates in the feature space rather than the parameter space. This is equivalent to performing a physically driven gradient descent step at the feature level, making the output features closer to a physically consistent state. Specifically, this can be understood as: calculating... about gradient The gradient indicates the direction of feature change that reduces the consistency error, and then... (The physically consistent part has been enhanced through weight modulation) as the adaptive learning rate or modulation factor, through the formula To implement the update, among which, This is element-wise multiplication.
[0068] It should also be noted that this invention directly applies physical consistency to the feature update process, causing the network to adjust towards a physically reasonable direction at the feature level. Introducing physical corrections in the intermediate feature layer allows subsequent multi-scale feature decomposition and regression prediction to be based on a more reasonable representation, making the model more robust to anomalous observations, local noise, and inter-modal mismatches, and improving the interpretability of carbon emission estimation results.
[0069] In one embodiment, such as Figure 3 As shown, the physical lag aligned feature vectors are analyzed using a heatmap. The horizontal axis represents the feature dimension index (stitched together in the order of activity, remote sensing, and ground observation), and the vertical axis represents the sample index (corresponding to different geographic units or timestamps). The colors represent feature values (normalized, dimensionless). The left area of the figure represents activity features, the middle area represents remote sensing features, and the right area represents ground observation features. There are obvious structural differences between different samples. After spatiotemporal alignment and atmospheric transport lag compensation, the three types of heterogeneous data are mapped to a unified physical lag aligned semantic space.
[0070] S3, Multiscale sparse decomposition and dynamic activation of carbon emission characteristics Physically enhanced feature vectors already contain alignment and physical consistency information from multi-source data, but they have high dimensionality, and different carbon emission drivers vary significantly across spatial scales. For example, industrial point sources are highly localized, traffic emissions are regionally diffuse, and background concentrations exhibit large-scale, gradual variability. Directly applying a uniform nonlinear transformation to all features can easily cause local emission signals to be overwhelmed by the large-scale background.
[0071] This invention transforms physically enhanced feature vectors into compact, low-dimensional representations with source resolution through three processes: multi-scale feature decomposition, scale-adaptive sparse activation, and source contribution compression. The specific steps are as follows: S301, Multi-scale Feature Decomposition Based on Carbon Emission Source Analysis To differentiate carbon emission contributions at different spatial scales, this invention rearranges the physical enhancement feature vectors of multiple geographic units at the same time into local feature tensors according to their spatial adjacency relationships. Then, a multi-branch convolution decomposition module is used to extract features at the local, regional, and background scales, respectively. The specific steps are as follows: 1> Taking the geographic unit to be predicted as the center, extract the physical augmentation feature vectors in its surrounding spatial neighborhood and rearrange them into local feature tensors. The size is It represents the distribution of multi-source fusion features in the neighborhood surrounding the current geographic unit; and These represent the height and width of the local neighborhood, with examples of values of 15 and 15; When the study area is not a regular grid, a local neighborhood can be constructed first based on the adjacency relationship of geographical units, and then an approximate regular tensor can be formed by interpolation or adjacency sorting.
[0072] In one embodiment, as an example, assume the study area is a regular grid with a size of [size missing]. (like For those located in grid coordinates The central geographic unit, whose surrounding neighborhood is a row ,List ,in ,Pick ,but Extract the Physical augmentation feature vector (dimension) of each grid in the sub-region ), and arrange them into a tensor in the original row and column order, with a size of .
[0073] Using the example from the previous embodiment, if the study area is an irregular grid (such as an administrative region), a graph adjacency relationship needs to be constructed first. Using the central cell as the root, a breadth-first search is performed, adding neighbors layer by layer until a neighborhood of 225 cells (including itself) is collected. These cells are then sorted according to their distance from the center, and then the results are applied through interpolation (such as inverse distance weighting) or by direct expansion. An approximate regular grid is used, for example, the nearest neighbor cell is placed in the center, the outer cells are filled to the surrounding positions according to the azimuth, and the gaps are filled with 0.
[0074] 2> Convert the local feature tensor Simultaneously, three parallel convolutional branches are input to extract local scale features, regional scale features, and background scale features, respectively. The scale index is denoted as... , The values are 1, 2, and 3, corresponding to the local scale, regional scale, and background scale, respectively. The kernel sizes of the three convolution branches are denoted as follows: , , ,in, This indicates the size of the local-scale convolution kernel, with an example value of 3, used to highlight point source emissions, road emissions, and local anomalous activity signals. This indicates the size of the region-scale convolution kernel, with an example value of 7, used to extract features of industrial areas, urban traffic, and mesoscale diffusion. This indicates the background-scale convolution kernel size, with an example value of 15, used to extract large-scale background density and regional transmission trends.
[0075] In the specific implementation, each convolutional branch adopts a "convolution + batch normalization + rectified linear unit (ReLU)" structure, with the following specific parameters: a> Local scale branch ( )include, Convolutional layer: Input channel Output channel convolution kernel Stride 1, padding 1 (preserving size); Batch normalization layer: feature dimension 32; ReLU activation; Optional second convolutional layer: kernel Output channels 64, step size 1, padding 1, followed by batch normalization and ReLU; b>Regional scale branch ( )include, Convolutional layer: Input channel Output channel convolution kernel Stride 1, padding 3 (preserves size); batch normalization layer; ReLU activation; optional second convolutional layer: kernel Output channels: 128; c> Background Scale Branch ( )include, Convolutional layer: Input channel Output channel convolution kernel Step size 1, padding 7 (maintain size); batch normalization layer; ReLU activation.
[0076] It should be noted that the output feature map size for each branch is [size missing]. ,in They are 32, 64, and 128 respectively (corresponding to) ).
[0077] Furthermore, spatial pooling is performed on the output feature maps of each convolutional branch (aggregating the two-dimensional feature maps into one-dimensional feature vectors through global average pooling) to obtain the component feature vectors at the corresponding scale. , Indicates the first The component feature vectors at each scale characterize the carbon emission contribution pattern at that scale.
[0078] 3> During training, inter-scale orthogonality constraints are introduced for feature vectors of different scale components. The orthogonality constraint loss is denoted as... , A correlation penalty term characterizing the features of components at different scales is used to reduce information overlap between local, regional, and background scales.
[0079] In practical implementation, to reduce the information overlap among the feature vectors of the local, regional, and background scales, and to ensure that each scale branch learns an independent carbon emission contribution pattern, it is necessary to introduce inter-scale orthogonality constraints during training. This is because the feature vectors of the three scales... , and Since the dimensions are different (32, 64, and 128 respectively), the relevance matrix or similarity cannot be directly calculated. Therefore, the following steps are used to implement the constraints: a) Project the model uniformly to a low-dimensional space. Specifically, set up a learnable fully connected mapping layer (i.e., a linear projection layer) for each scale. The input dimensions of these three mapping layers are the original dimensions of the feature vectors at each scale (32, 64, 128), and the output dimensions are uniformly set to a small common dimension (e.g., 16 or 32, usually much smaller than the original dimension). The parameters of each mapping layer are independent and trainable. b) Obtaining low-dimensional projection vectors: Specifically, in a training batch, for each sample, the feature vectors of the three scales are input into the corresponding mapping layer to obtain three low-dimensional projection vectors of the same dimension, denoted as follows: , , These projection vectors reside in the same semantic space and can be directly compared; c> Calculate the cosine similarity, specifically for each pair of projection vectors ( , , ), calculate the cosine similarity, the range of values for cosine similarity is: The closer the value is to 0, the more orthogonal the two vectors are (i.e., the less information overlap), and the closer the value is to ±1, the stronger the correlation. d> Constructing the orthogonality penalty term: Sum (or take the absolute value) the squares of the three pairs of cosine similarities of all samples in the batch, then divide by the number of samples to obtain the average penalty value of the batch. The squaring operation results in a positive penalty for both positive and negative similarities, with the penalty increasing the further away from 0. This penalty value is the inter-scale orthogonality constraint loss. During training, the optimizer will minimize This forces projection vectors of different scales to tend to be orthogonal, thereby indirectly enabling the original scale feature vectors to learn mutually independent information.
[0080] Based on this, the three scale branches are subject to regularization pressure during backpropagation to avoid extracting highly correlated features. For example, if the local scale branch and the regional scale branch learn similar edge responses, their cosine similarity after projection will be large, thus generating a penalty that forces them to focus on features of different spatial ranges. Ultimately, the local scale features highlight local signals such as point sources and roads, the regional scale features express the diffusion pattern of a region, and the background scale features capture large-scale gradual trends.
[0081] It should be noted that the output consists of component feature vectors at three scales. , and These correspond to carbon emission contribution models at the local, regional, and background scales, respectively. Local-scale carbon emission characteristic vector By local scale branch (convolution kernel) The output feature map is obtained by global average pooling, with dimension [missing information]. ; Regional-scale carbon emission characteristic vector By region-scale branch (convolution kernel) The output feature map is obtained by global average pooling, with dimension [missing information]. ; Background-scale carbon emission feature vector By background scale branch (convolution kernel) The output feature map is obtained by global average pooling, with dimension [missing information]. .
[0082] It should be noted that different emission sources naturally have different spatial ranges. If they are all mixed together for learning at a single scale, it is easy for local strong emissions to be averaged out by the background field. This invention is not intended to simply increase the network depth, but to enable the model to explicitly decompose the spatial contribution of carbon emissions into multiple scale levels. After multi-scale branch decomposition, the model can simultaneously retain three types of information: "local anomalies", "regional propagation" and "gradual background changes".
[0083] S302, Scale-Adaptive Sparse Activation and Feature Fusion Carbon emission characteristics exhibit varying sparsity at different scales. Local-scale characteristics typically show significant activation only in a few locations, regional-scale characteristics are moderately sparse, and background-scale characteristics are usually relatively flat. Applying the same activation function to all three scales often fails to meet the representational needs of each scale simultaneously.
[0084] This invention adaptively sets a sparse threshold based on the statistical characteristics of each scale component and fuses them according to their scale contribution. The specific steps are as follows: 1> For each scale ( ) component eigenvectors Statistical analysis was performed, calculating the mean and standard deviation of the feature vectors for each of the three scale components in the current training batch, and standardizing the feature vectors for each scale component. Specifically, No. The mean of each scale is denoted as , Characterizing the first The average activation level of each scale component feature vector in the training batch; No. The standard deviation of each scale is denoted as , Characterizing the first The degree of fluctuation of each scale component feature vector in the training batch.
[0085] It should be noted that during the inference phase, the mean and standard deviation of the moving average accumulated during the training phase can be used.
[0086] 2> For each scale component feature vector (after standardization), construct a sparse activation function based on a Sigmoid gate to obtain the corresponding sparse activation feature vector. , Characterizing the first The sparse activation feature vectors of each scale are the effective features after scale-adaptive sparse filtering.
[0087] In the specific implementation, first, the standardized first... Each scale component feature vector is multiplied by the learnable scale parameter. Then input the Sigmoid function to generate the gated response, where, Indicates the first The sensitivity control parameter for each scale to features deviating from the mean is such that the larger the value, the more sensitive the scale is to features with high response.
[0088] 3> Connect the gating response with the first Sparsity threshold at each scale The comparison is performed, and the portion below the threshold is directly suppressed, while the portion above the threshold is retained and treated as an effective response. Indicates the first The sparse threshold at each scale is determined by a learnable threshold parameter. Obtained through Sigmoid mapping, this ensures that the threshold falls within the range of 0 to 1. For use in controlling the first The sparse threshold at each scale is a learnable parameter.
[0089] It should be noted that features at different scales have different sparsity priors, therefore different initial sparsity thresholds should be set, for example, an initial sparsity threshold for local scales ( Since local emissions are typically sparse (with strong emissions in a few grids), it is desirable to preserve peak responses while suppressing most low values. Therefore, the initial threshold is set relatively high (e.g., 0.7), so that only features with gating responses greater than 0.7 are preserved. This initial sparsity threshold is set for the region scale. The initial value is set to 0.5, which represents medium sparsity and is the initial sparsity threshold for the background scale. The initial value is set to 0.3 to retain more smooth background information.
[0090] 4> Activate the sparse feature vectors at three different scales respectively ( , and Calculate the scale contribution score and normalize it into the fusion weight. , Characterizing the first The contribution of each scale to the final fusion is a scalar.
[0091] In practice, global average pooling can be performed on the sparse activation feature vectors of each scale first, and then the contribution score of that scale can be obtained through a linear mapping layer. Finally, the contribution scores of the three scales can be normalized so that the sum of the fusion weights of the three scales is 1.
[0092] 5> Activate the sparse feature vectors at three scales ( , and ) multiplied by the corresponding fusion weights ( , and After that, the features are concatenated along the feature dimension to obtain the fused sparse feature vector. , dimension , It is the sum of the dimensions of the sparse activation features at three scales.
[0093] It should be noted that by using the "weighted concatenation" method, instead of directly adding the three scale features with different dimensions, we can preserve the independent information of different scales at the same time and avoid structural errors caused by direct addition between different dimensions.
[0094] In one embodiment, such as Figure 4As shown, the source contribution feature vector of an example sample is analyzed. The horizontal axis represents the predefined carbon emission source categories (industrial, transportation, agricultural, and natural sources), and the vertical axis represents the contribution intensity (normalized, dimensionless). In the figure, industrial sources have the highest contribution (0.62), while natural sources have the lowest (0.01). This invention can transform high-dimensional fused sparse features into a compact representation with source resolution significance, demonstrating the interpretability of the model for carbon emission sources.
[0095] S303, sparsity compression of carbon emission drivers To further reduce the feature dimensionality and transform the high-dimensional fused sparse features into a more compact representation with greater source resolution significance, this invention first projects the fused sparse features onto a predefined carbon emission source category space, and then performs nonlinear compression to obtain a compressed low-dimensional feature vector. The specific steps are as follows: 1> Pre-define carbon emission source categories and create a source category index set. The total number of source categories is denoted as . , The number of predefined carbon emission source categories is used, for example, a value of 4 corresponds to industrial sources, transportation sources, agricultural sources, and natural sources.
[0096] In practical applications, the scope can be expanded to more granular categories such as residential sources, power sources, and port sources, depending on the characteristics of the monitoring area. This invention only describes four categories of carbon emission sources: industrial sources, transportation sources, agricultural sources, and natural sources.
[0097] 2> Construct the source contribution projection matrix And will fuse sparse feature vectors Mapping to the source class space yields the source contribution feature vector. , dimension This characterizes the distribution of the current sample's contribution across various carbon emission source categories, where... The source contribution projection matrix has a size of . This is used to transform high-dimensional fused sparse features into the contribution intensity of each source category; In the specific implementation, the source contribution projection matrix is constructed. In order to reduce the number of parameters while preserving the main correspondence between the source categories and the feature space, a size of [size missing] is directly used. The matrix has many parameters (if) If the parameter is 4000, then it is acceptable, but if If the value is very large, low-rank decomposition can be chosen. The source contribution projection matrix is constructed using low-rank decomposition; that is, low-rank decomposition will... Represented as the product of two smaller matrices Specifically, first train the low-rank factor matrix of the source categories. and characteristic low-rank factor matrix Then, the two are combined to form a projection matrix, and each row is normalized, where, Represents the source category semantic encoding matrix. Represents the semantic encoding matrix of the feature space. express The transpose of , with the lower-rank factor dimension denoted as , Denotes the dimension of the low-rank factor, satisfying that it is much smaller than and The example value is 16.
[0098] 3> Convert the source contribution feature vector Input to the compression mapping layer to obtain compressed low-dimensional feature vectors , dimension This is used as input for the subsequent regression prediction module; This represents the compressed feature dimension, with an example value of 64.
[0099] In the actual implementation, a compressed projection matrix is first used. and compressed bias vector The source contribution feature vector is linearly transformed and then nonlinearly compressed using a compression activation function.
[0100] In practical implementation, carbon emission drivers may manifest as either a significant increase in positive emissions or a subtle increase in a weakly changing background. Using only a single activation function may easily miss one type of information. The compressed activation function adopts a combination of hyperbolic tangent function and soft positive function to simultaneously preserve the positive and negative trends and the sensitivity to small positive values. Specifically, the result after linear transformation is subjected to hyperbolic tangent mapping on one hand to preserve the relatively symmetrical trend, and soft positive function mapping on the other hand to enhance the weak positive response. The two results are then added together as the final compressed output.
[0101] S4. Adaptive Regression Prediction and Physical Constraint Optimization for Carbon Emission Monitoring Tasks After multi-scale decomposition, sparse activation, and source contribution compression, compressed low-dimensional feature vectors and source contribution feature vectors have been obtained. The proportion of influence of various sources on total carbon emissions varies across different geographical units, seasons, and activity backgrounds. If fixed regression parameters are still used, it will be difficult to adapt to changes in emission patterns.
[0102] This invention dynamically generates regression parameters based on the source contribution feature vector and superimposes physical constraint regularization during training, so that the prediction results have both high accuracy and good physical rationality. The specific steps are as follows: S401, Dynamic parameter generation for source contribution awareness This invention controls the weights and biases of the regression network by using source contribution feature vectors, enabling the regression network to adaptively adjust its prediction strategy as the source composition of the samples changes. The specific steps are as follows: 1> Read the source contribution feature vector The contribution values of each source category in the data, the first The contribution value of the class source is denoted as , Characterizing the first The contribution intensity of carbon emission sources; Represents the source category index, with values ranging from 1 to... .
[0103] In practical implementation, feature vectors are contributed from the source. (dimension) Extract the contribution value of each source category from the contribution intensity of each element corresponding to different source categories. It's a vector, so you can read it directly by index. .
[0104] Furthermore, weighted modulation coefficients are constructed for each type of source contribution value. Specifically, the deviation between the source contribution value of that type and its global average value is first calculated, and then the deviation is input into the Sigmoid gate function to obtain the weighted modulation coefficients in the range of 0 to 1. No. The global average contribution value of the class source is denoted as , Characterizing the first The average contribution level of the class source in the training set, where a moving average is used to update the value during training, and a cumulative value is used during inference. ; No. The modulation sensitivity coefficient of the source is denoted as , Characterizing the first The sensitivity of a class source to the dynamic weight generation is a learnable sensitivity coefficient (scalar). The larger the value, the more easily changes in the contribution of that class source will cause adjustments to the regression weights.
[0105] In practical implementation, the degree of deviation is calculated. Its global average contribution The difference is then typically calculated by dividing by the standard deviation. The greater the deviation, the more significantly the current sample's contribution to that source category deviates from the average level.
[0106] 2> Set the basic weight matrix for the regression network and the weight modulation matrix corresponding to each source category Based on the weight modulation coefficients of each source type, the basic weight matrix is weighted and corrected to obtain the dynamic weight matrix. ,in, Represents a dynamic weight matrix with size . , used for carbon emission regression prediction of the current sample; This indicates the output dimension, which is set to 1 in this invention, corresponding to the predicted carbon emission rate. This represents the basic weight matrix, which consists of learnable parameters. Indicates the first The modulation matrix of the class source on the regression weights is a learnable parameter.
[0107] It should be noted that the regression network is a linear regression layer (fully connected layer), and the input is a compressed low-dimensional feature vector. (dimension) The output is a scalar. The weights and biases of this layer are dynamically generated, rather than fixed, allowing the network to adaptively adjust to different source configurations while maintaining training stability. This represents the predicted carbon emission rate, which is a scalar value, expressed in tons of carbon dioxide per hour.
[0108] 3> Further construct bias modulation coefficients for each type of source contribution value. Specifically, the first... Class source contribution value (i.e.) Multiply by the corresponding bias sensitivity coefficient Then, by inputting the hyperbolic tangent function, the positive and negative bounded bias modulation coefficients are obtained, where, Indicates the first The sensitivity of the class source to dynamic bias generation is a learnable bias sensitivity coefficient (scalar).
[0109] 4> Set the basic bias vector for the regression network and the bias modulation vectors corresponding to each source category Based on the bias modulation coefficients of various sources, the basic bias vector is corrected to obtain the dynamic bias vector. , dimension , used as the regression output bias term for the current sample, where, This represents the fundamental bias vector, which is a learnable parameter. Indicates the first The modulation vector of the source class to the regression bias is a learnable parameter.
[0110] It should be noted that the emission formation mechanisms of industrial-dominated regions and natural-dominated regions are different. The same regression mapping should not be used for the same input features. This invention does not retrain a new model for each sample. Instead, it generates a set of sample-adaptive lightweight regression parameters based on the source contribution feature vector on the same set of backbone network parameters. Through dynamic parameter generation, the model can better adapt to abrupt changes in emission patterns, seasonal changes and regional differences.
[0111] S402, Carbon Emission Rate Prediction and Physical Constraint Regularization After obtaining the dynamic weight matrix and dynamic bias vector, regression prediction can be performed on the compressed low-dimensional feature vector. Meanwhile, to ensure the physical plausibility of the output results, this invention introduces source contribution consistency constraints and sign consistency constraints, the specific steps of which are as follows: 1> Compress low-dimensional feature vectors Input the dynamic regression layer (i.e., the dynamic regression layer constructed by S401) to calculate the predicted carbon emission rate. , represented as ,in, This represents the predicted carbon emission rate, which is a scalar quantity, and the unit is tons of carbon dioxide per hour. This represents compressing a low-dimensional feature vector with dimension . ; Represents the dynamic weight matrix; This represents the dynamic bias vector.
[0112] 2. Construct physical prior prediction values based on the source contribution feature vectors. Specifically, pre-set an emission coefficient for each type of source. , Characterizing the first The prior value of the carbon emission rate corresponding to the unit contribution intensity of the source can be set based on historical statistical data, carbon emission accounting experience, or expert knowledge.
[0113] In one embodiment, as an example, suppose the source categories are industrial, transportation, agricultural, and natural sources, with contribution values of 0.6, 0.3, 0.1, and 0.05 (normalized), respectively. The prior emission coefficients... (Unit: tons of CO2 / hour / unit contribution) Determined based on historical statistics, i.e., industrial sources: (High emission intensity), traffic sources: Agricultural source: Natural source: (Natural sources may be carbon sinks, resulting in negative emissions), then, the physical prior prediction value... .
[0114] Then, the contribution values of each source are multiplied by the corresponding emission coefficient and summed to obtain the physical prior prediction value, which represents "the theoretically approximate emission level when the contribution of each source is linearly combined".
[0115] 3> Calculate the consistency constraint term of source contribution This constraint term is used to measure the predicted carbon emission rate. The difference between the prediction and the physical prior prediction is considered. A smaller difference indicates a better match between the prediction and the source contribution decomposition; a larger difference indicates a more significant deviation from the physically interpretable source contribution structure. .
[0116] 4. Set symbolic physical priors for different source categories and calculate symbolic consistency constraints. , represented as ,in, The function represents the maximum value, and its physical prior is denoted as . , Characterizing the first The sign and direction of the relationship between carbon emission sources and changes in carbon emissions; when When this occurs, it indicates that an increase in the contribution of this type of source typically corresponds to an increase in emissions; when When this occurs, it indicates that an increase in the contribution of this type of source usually corresponds to an increase in net absorption or a decrease in emissions.
[0117] In practice, industrial sources, transportation sources, and power sources correspond to positive signs; natural sink sources such as forest absorption and wetland absorption can correspond to negative signs. If the sign of the contribution value of a certain type of source is opposite to its physical prior direction, a penalty is added to that part.
[0118] 5> Incorporate source contribution consistency constraints and symbol consistency constraints Adding them together yields the physical constraint regularization term. This is used to constrain the prediction results to be consistent with the source contribution decomposition results and with the prior physical direction of the source category.
[0119] S403, Model Training and Parameter Optimization To achieve end-to-end training, this invention combines prediction error, physical constraint regularization term, and scale orthogonality constraint loss to construct a composite loss function, and uniformly optimizes all network parameters. The specific steps are as follows: 1> Construct a training sample set, each training sample including at least the activity data feature vector, remote sensing inversion feature vector, ground observation feature vector, and real carbon emission label value.
[0120] The actual carbon emission label value is recorded as , The actual carbon emission rates corresponding to the training samples can be derived from flux tower observation conversion results, inversion calibration results, or carbon emission accounting results.
[0121] It should be noted that, in order to ensure that the labels are consistent with the input, the temporal granularity and spatial unit division of the training samples need to be consistent with the alignment rules in S201.
[0122] 2> Perform forward propagation for each training batch, sequentially completing spatiotemporally aligned feature encoding, cross-modal physical constraint enhancement, multi-scale feature decomposition, scale-adaptive sparse activation, source contribution compression, dynamic parameter generation, and carbon emission rate prediction, to obtain the predicted carbon emission rate value for each sample. .
[0123] 3> Based on carbon emission rate predictions Compared with the actual carbon emission label value Calculate the mean squared error loss, denoted as . , The mean squared error, which characterizes the difference between the predicted carbon emission rate and the actual carbon emission label value, is used to measure the accuracy of the prediction.
[0124] 4> Calculate the physical constraint regularization term and scale orthogonality constraint loss ,in, Used to constrain the physical plausibility of prediction results, Used to constrain the independence of multi-scale decomposition results.
[0125] 5> Construct the total loss function , represented as ,in, This represents the total loss function for model training; This represents the physical constraint regularization coefficient, with an example value of 0.01. This represents the scale orthogonality constraint coefficient, with an example value of 0.01.
[0126] It should be noted that incorporating the scale orthogonality constraint loss into the total loss function can ensure that the multi-scale decomposition constraint introduced in S301 truly plays a role in training, avoiding the learning of highly overlapping features by different scale branches.
[0127] 6> An adaptive moment estimation optimizer is used to perform backpropagation updates on all learnable parameters, including long short-term memory network parameters, parameterized modified linear unit parameters, transmission weight generation parameters, diffusion weight generation parameters, flux modulation weight generation parameters, multi-scale convolution kernel and bias parameters, scale-adaptive sparse activation parameters, source contribution projection parameters, compression mapping parameters, dynamic regression weight generation parameters, dynamic regression bias generation parameters, and other learnable hyperparameters.
[0128] In one implementation, the learning rate can be set to... or The batch size can be 16, 32 or 64, and the number of training rounds can be 50 to 200.
[0129] 7> Monitor the change in the total loss function on the validation set during training. If the validation set loss no longer decreases within several consecutive rounds, perform early stopping or reduce the learning rate to prevent overfitting.
[0130] After training is complete, the model parameters corresponding to the minimum loss are saved to obtain the fully trained deep adaptive regression network.
[0131] S5, Carbon Emission Monitoring Methods After the model training is completed, it enters the practical carbon emission monitoring application stage. In the application process, the trained deep adaptive regression network is deployed to the monitoring system to perform forward inference on multi-source monitoring data collected in real time or near real time, and output the predicted carbon emission rate of each geographical unit at the current time.
[0132] In practice, the data acquisition process begins by following the same acquisition rules as during the training phase to obtain activity data, remote sensing inversion data, and ground observation data for the time to be monitored and historical times. Activity data must be backed up to at least 72 hours of historical data to meet the needs of subsequent atmospheric transport lag compensation. Remote sensing inversion data is taken from satellite transit products for the current time and the past few days. Ground observation data is taken from meteorological and flux observations for the current time and the past few hours.
[0133] All raw data undergoes spatial mapping according to the geographic unit division method set in the training phase (such as regular grids or administrative units). Point, line, and area source activity data are accumulated or assigned to corresponding geographic units. Remote sensing raster data are assigned to geographic units through area-weighted averaging, and ground station observations are assigned to the units where the stations are located. Then, the three types of data are resampled according to the unified time resolution (such as 3-hour intervals) in the training phase. The activity data is averaged using a sliding window, the remote sensing data is amended using the most recent time, and the ground observation data is interpolated or amended using the most recent time. All features are then subjected to min-max normalization or Z-score standardization with the same parameters as in the training phase to obtain the original stitched feature vector of each geographic unit at the current timestamp.
[0134] Then, the preprocessed multi-source feature vectors are input into the trained network model, and spatiotemporal aligned feature encoding, cross-modal physical constraint enhancement, multi-scale feature decomposition, scale-adaptive sparse activation, source contribution compression, and dynamic regression prediction are executed in sequence.
[0135] In the spatiotemporal alignment feature encoding stage, the system determines the set of atmospheric transport source areas for each geographic unit based on pre-calculated or real-time acquired wind field data (wind speed, wind direction) and boundary layer height, calculates the transmission distance and transmission lag of each transmission path, extracts lagged activity data features from historical activity data sequences, constructs distance attenuation and atmospheric diffusion modulation weights, inputs the weighted lagged activity data sequences into a long short-term memory network to generate a dynamic lag compensation vector, modulates the activity data feature vector at the current moment dimension by dimension, and then concatenates it with the remote sensing inversion feature vector and the ground observation feature vector to obtain the physical lag alignment feature vector.
[0136] In the cross-modal physical constraint enhancement stage, the system separates three types of features from the physical lag alignment feature vector, projects them onto a unified comparison dimension, generates a transmission weight vector using boundary layer height and wind field information, generates a diffusion weight vector based on the similarity between the active transmission descriptor and the remote sensing diffusion descriptor, and generates a flux modulation weight vector based on the difference between the active transmission descriptor and the ground flux descriptor. The three types of features are multiplied element-wise with their corresponding weights and then concatenated to obtain the physical consistency enhancement feature vector. The physical consistency error is then constructed and used to perform gradient correction on the physical lag alignment feature vector, outputting the physical enhancement feature vector.
[0137] In the subsequent multi-scale decomposition and sparse activation stages, the system extracts the physical enhancement feature vectors within the spatial neighborhood centered on the geographic unit to be predicted and rearranges them into local feature tensors. Through three parallel convolutional branches (with kernel sizes of 3, 7, and 15, respectively), carbon emission contribution patterns at the local, regional, and background scales are extracted. Global average pooling is performed on the output of each branch to obtain component feature vectors at the three scales. Then, the mean and standard deviation of each scale component are calculated. After standardization, scale-adaptive sparse filtering is performed using a sparse activation function based on Sigmoid gating. The contribution scores of each scale are calculated and normalized into fusion weights. The sparse activation feature vectors at the three scales are weighted and concatenated to obtain the fused sparse feature vector.
[0138] The system maps the fused sparse feature vectors to a predefined carbon emission source category space (such as industrial, transportation, agricultural, and natural sources) through a pre-trained source contribution projection matrix to obtain source contribution feature vectors. Then, it passes through a compression mapping layer (using a combination of hyperbolic tangent and soft positive functions for activation) to obtain compressed low-dimensional feature vectors.
[0139] The system dynamically generates the regression weight matrix and bias vector for the current sample based on the various source contribution values in the source contribution feature vector, combined with the global average contribution value, basic weight matrix and weight modulation matrix corresponding to each source, basic bias vector and bias modulation vector corresponding to each source saved during the training phase. The compressed low-dimensional feature vector is then input into the dynamic regression layer to calculate the predicted carbon emission rate of the geographic unit at the current moment, in tons of carbon dioxide per hour.
[0140] The entire inference process is executed in parallel across all geographical units within the study area. The system can output a spatial distribution map of carbon emission rates covering the entire region and supports continuous time series monitoring on an hourly or 3-hourly basis.
[0141] For newly arriving real-time data streams, the system repeats the aforementioned forward inference process to achieve rolling update predictions. No backpropagation or parameter updates are required during application; only forward computation is performed, thus meeting the timeliness requirements of near real-time monitoring. Monitoring results can be displayed through a visual interface and provide auxiliary functions such as exceedance warnings, source contribution analysis, and historical trend comparisons.
[0142] Example 2 This embodiment also provides a carbon emission monitoring system based on multi-source monitoring data, which executes the carbon emission monitoring method based on multi-source monitoring data described in Embodiment 1, including: Data acquisition module: used to collect multi-source carbon emission monitoring data; the multi-source carbon emission monitoring data includes activity data, remote sensing inversion data and ground observation data; Feature enhancement module: Processes multi-source carbon emission monitoring data by enhancing spatiotemporal alignment and cross-modal physical consistency for the physical transmission process of carbon emissions, and maps multi-source features to a unified carbon emission driver-response semantic space to obtain physically enhanced feature vectors; The eigenvalue decomposition and activation module is used to perform multi-scale eigenvalue decomposition, scale-adaptive sparse activation, and source contribution compression on the physically enhanced feature vector to obtain the source contribution feature vector and the compressed low-dimensional feature vector. Carbon emission rate prediction module: Used to perform adaptive regression prediction and physical constraint optimization on source contribution feature vectors and compressed low-dimensional feature vectors to obtain carbon emission rate prediction values.
[0143] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A carbon emission monitoring method based on multi-source monitoring data, characterized in that, Includes the following steps: S1. Collect multi-source carbon emission monitoring data; the multi-source carbon emission monitoring data includes activity data, remote sensing inversion data, and ground observation data; S2. Multi-source carbon emission monitoring data are processed by enhancing spatiotemporal alignment and cross-modal physical consistency for the physical transmission process of carbon emissions. Multi-source features are mapped to a unified carbon emission driver-response semantic space to obtain physically enhanced feature vectors. S3. Perform multi-scale feature decomposition, scale-adaptive sparse activation and source contribution compression on the physical enhancement feature vector to obtain the source contribution feature vector and the compressed low-dimensional feature vector. S4. Adaptive regression prediction and physical constraint optimization are performed on the source contribution feature vector and the compressed low-dimensional feature vector to obtain the predicted carbon emission rate.
2. The carbon emission monitoring method based on multi-source monitoring data according to claim 1, characterized in that, In step S2, time compensation is performed on the activity data, and then physical correlation encoding is performed with the remote sensing inversion data and ground observation data to obtain the physical lag alignment feature vector: The study area was divided into multiple geographical units, and the spatial attribution of various data types was standardized; the temporal resolution was standardized, and basic preprocessing was performed on the original sequences; at the current timestamp Next, construct the first The original stitched feature vector of each geographic unit, including the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector; for the first The first geographical unit is used to determine the set of atmospheric transport source regions; for the first... For each transmission path in the atmospheric transmission source region set corresponding to a geographic unit, the transmission distance and transmission delay are calculated. Based on the delay results of each transmission path, lagged activity data features are extracted from the historical activity data sequence, and distance attenuation and atmospheric diffusion modulation weights are constructed to obtain the lagged activity data weighted features after distance attenuation and atmospheric diffusion modulation. The lagged activity data weighted feature sequence obtained from the same geographic unit at multiple historical moments and multiple transmission paths is input into a long short-term memory network to generate a dynamic lag compensation vector. The dynamic lag compensation vector is used to modulate the activity data feature vector at the current moment dimension by dimension and then concatenate it with the remote sensing inversion feature vector and the ground observation feature vector to obtain the physical lag aligned feature vector. The basic preprocessing includes: outlier removal, interpolation repair, and normalization.
3. The carbon emission monitoring method based on multi-source monitoring data according to claim 2, characterized in that, In step S2, a cross-modal physical consistency enhancement mechanism is constructed. This mechanism uses transmission weights, diffusion weights, and flux modulation weights to constrain and enhance the three types of data in the physical hysteresis alignment feature vector. The physical consistency error is then used to correct the features, resulting in a physically enhanced feature vector. From the physically lag aligned feature vector, active data feature vector, remote sensing inversion feature vector, and ground observation feature vector are separated. Three fully connected mapping layers are set up to project these three feature vectors onto the same comparison dimension, resulting in an active transport descriptor, a remote sensing diffusion descriptor, and a ground flux descriptor. A transport weight vector corresponding to the active data is generated based on boundary layer height and wind field information. A diffusion weight vector corresponding to the remote sensing inversion feature vector is generated based on the similarity between the active transport descriptor and the remote sensing diffusion descriptor. A flux modulation weight vector corresponding to the ground observation feature vector is generated based on the difference between the active transport descriptor and the ground flux descriptor. The active data feature vector, remote sensing inversion feature vector, and ground observation feature vector are multiplied element-wise by their respective transport weight vector, diffusion weight vector, and flux modulation weight vector, and then concatenated along the feature dimension to obtain a physically consistent enhanced feature vector. A physical consistency error is constructed, and this error is used to perform gradient correction on the physically lag aligned feature vector to obtain the physically enhanced feature vector. The specific process of constructing the physical consistency error is as follows: extract comparable source intensity, concentration, and flux representations from the activity data feature vector, remote sensing inversion feature vector, and ground observation feature vector, respectively, and calculate the consistency error among the three representations; use the gradient of the consistency error with respect to the physical lag alignment feature vector as the correction direction, and use the physical consistency enhancement feature vector as the correction amplitude modulation term to update the physical lag alignment feature vector in one or more steps.
4. The carbon emission monitoring method based on multi-source monitoring data according to claim 1, characterized in that, In step S3, the physical enhancement feature vectors of multiple geographic units at the same time are rearranged into local feature tensors according to their spatial adjacency. The local scale, regional scale, and background scale features are then extracted using a multi-branch convolution decomposition module, resulting in three scale component feature vectors: Centered on the geographic unit to be predicted, physical enhancement feature vectors within the surrounding spatial neighborhood are extracted and rearranged into local feature tensors. These local feature tensors are simultaneously input into three parallel convolutional branches to extract local-scale, regional-scale, and background-scale features, respectively, resulting in local-scale carbon emission feature vectors, regional-scale carbon emission feature vectors, and background-scale carbon emission feature vectors. The kernel sizes of the three parallel convolutional branches are 3, 7, and 15, respectively, and all include convolutional layers, batch normalization layers, and ReLU activation functions. During training, inter-scale orthogonality constraints are introduced for feature vectors of different scale components.
5. The carbon emission monitoring method based on multi-source monitoring data according to claim 4, characterized in that, In step S3, the statistical features of the local-scale carbon emission feature vector, the regional-scale carbon emission feature vector, and the background-scale carbon emission feature vector are adaptively set with sparsity thresholds, and then fused according to scale contribution to obtain a fused sparse feature vector: Statistical analysis was performed on the local-scale, regional-scale, and background-scale carbon emission feature vectors, calculating their mean and standard deviation in the current training batch, and then standardizing them. A sparse activation function based on a Sigmoid gate was constructed from the standardized local-scale, regional-scale, and background-scale carbon emission feature vectors to obtain the corresponding sparse activation feature vectors. The gate response was then compared with the... The sparsity thresholds at each scale are compared. Parts below the threshold are directly suppressed, while parts above the threshold are retained and treated as valid responses. The scale contribution scores are calculated for the sparse activation feature vectors at each of the three scales, and the scale contribution scores are normalized to the fusion weights. The sparse activation feature vectors at each of the three scales are multiplied by their corresponding fusion weights and then concatenated along the feature dimensions to obtain the fused sparse feature vector.
6. The carbon emission monitoring method based on multi-source monitoring data according to claim 5, characterized in that, In step S3, carbon emission source categories are predefined and a source category index set is established; a source contribution projection matrix is constructed, and the fused sparse feature vector is mapped to the source category space to obtain the source contribution feature vector; the source contribution feature vector is input into the compression mapping layer to obtain the compressed low-dimensional feature vector.
7. The carbon emission monitoring method based on multi-source monitoring data according to claim 1, characterized in that, In step S4, the weights and biases of the regression network are controlled by the source contribution feature vectors, enabling the regression network to adaptively adjust its prediction strategy as the source composition of the samples changes. Read the contribution values of each source category in the source contribution feature vector; construct weight modulation coefficients for each source contribution value: first calculate the deviation between the source contribution value of that category and its global average value, then input the deviation into the Sigmoid gate function to obtain the weight modulation coefficients in the range of 0 to 1; To set up the basic weight matrix and the weight modulation matrix corresponding to each source category for the regression network, the basic weight matrix is weighted and modified according to the weight modulation coefficient of each source category to obtain the dynamic weight matrix. For each type of source contribution value, a bias modulation coefficient is further constructed: the first... The source contribution value is multiplied by the corresponding bias sensitivity coefficient and then input into the hyperbolic tangent function to obtain the bias modulation coefficient; Set a basic bias vector and a bias modulation vector corresponding to each source category for the regression network. Then, modify the basic bias vector according to the bias modulation coefficient of each source category to obtain a dynamic bias vector, which is used for the regression output bias term of the current sample.
8. The carbon emission monitoring method based on multi-source monitoring data according to claim 7, characterized in that, In step S4, after obtaining the dynamic weight matrix and dynamic bias vector, regression prediction is performed on the compressed low-dimensional feature vector: The compressed low-dimensional feature vector is input into the dynamic regression layer to calculate the predicted carbon emission rate. Physical prior predictions are constructed based on the source contribution feature vectors: an emission coefficient is pre-set for each source type, based on historical statistical data, carbon emission accounting experience, or expert knowledge. The contribution values of each source type are multiplied by the corresponding emission coefficient and summed to obtain the physical prior prediction. A source contribution consistency constraint term is calculated to measure the difference between the predicted carbon emission rate and the physical prior prediction. A signed physical prior is set for different source categories, and a signed consistency constraint term is calculated. Adding the source contribution consistency constraint term and the symbol consistency constraint term together yields the physical constraint regularization term.
9. The carbon emission monitoring method based on multi-source monitoring data according to claim 8, characterized in that, During training, the prediction error, physical constraint regularization term, and scale orthogonality constraint loss are combined to construct a composite loss function, which is then used to uniformly optimize all network parameters. Construct a training sample set, each training sample including at least the activity data feature vector, remote sensing inversion feature vector, ground observation feature vector, and real carbon emission label value; Forward propagation is performed on each training batch to obtain the predicted carbon emission rate for each sample; mean squared error loss is calculated based on the predicted carbon emission rate and the actual carbon emission label value; physical constraint regularization term and scale orthogonality constraint loss are calculated; the mean squared error loss, physical constraint regularization term and scale orthogonality constraint loss are weighted and summed to obtain the composite loss function.
10. A carbon emission monitoring system based on multi-source monitoring data, executing the carbon emission monitoring method based on multi-source monitoring data as described in claim 1, characterized in that, include: Data acquisition module: used to collect multi-source carbon emission monitoring data; The multi-source carbon emission monitoring data includes activity data, remote sensing inversion data, and ground observation data; Feature enhancement module: Processes multi-source carbon emission monitoring data by enhancing spatiotemporal alignment and cross-modal physical consistency for the physical transmission process of carbon emissions, and maps multi-source features to a unified carbon emission driver-response semantic space to obtain physically enhanced feature vectors; The eigenvalue decomposition and activation module is used to perform multi-scale eigenvalue decomposition, scale-adaptive sparse activation, and source contribution compression on the physically enhanced feature vector to obtain the source contribution feature vector and the compressed low-dimensional feature vector. Carbon emission rate prediction module: Used to perform adaptive regression prediction and physical constraint optimization on source contribution feature vectors and compressed low-dimensional feature vectors to obtain carbon emission rate prediction values.
Citation Information
Patent Citations
Carbon emission prediction method and system based on multi-source heterogeneous data contrast learning
CN119558678A
Carbon emission flow monitoring method, computer equipment and computer storage medium
CN119647760A
Method and system for dynamically monitoring carbon emission of Yellow River basin by using big data
CN119850229A
Monitoring system for forest ecological carbon sink calculation
CN120542737A
Road engineering carbon emission analysis method based on multi-source heterogeneous data fusion
CN121073006A