A multi-source heterogeneous carbon emission data fusion alignment method
Patent Information
- Application Number
- CN202610877501.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-22
AI Technical Summary
[0002]现有碳排放核算体系主要依赖单一来源的能耗统计或局部监测,在面对复杂工业场景时缺乏多源数据的有效协同
1、通过构建跨模态图谱与时序弹性匹配,将孤立的碳排放观测值精准映射至具体生产设备与生产批次,解决了数据溯源难题,为企业碳资产确权与精细化管理提供了强关联的业务依据。
Smart Images

Figure CN122797922A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of carbon emission management technology, specifically to a method for fusing and aligning multi-source heterogeneous carbon emission data. Background Technology
[0002] Existing carbon emission accounting systems primarily rely on single-source energy consumption statistics or localized monitoring, lacking effective coordination of multi-source data when facing complex industrial scenarios. On one hand, there is a severe disconnect between macroscopic space-based / airborne remote sensing observations and microscopic enterprise production and operational data in terms of spatiotemporal scale and business semantics. This makes it difficult to accurately trace environmental monitoring values to specific production assets or order batches, resulting in a "monitoring without accounting" dilemma and failing to meet the needs of refined management. On the other hand, existing methods often simply handle the inherent uncertainties of different observation methods, ignoring the impact of physical laws such as atmospheric transport on the values. They lack dynamic assessment and compliance screening mechanisms for data quality, resulting in accounting results that fail to meet the stringent standards of carbon trading audits and are unable to effectively support enterprises' compliance decisions and asset assessments.
[0003] To address the technical problem of difficulty in merging and aligning existing multi-source carbon monitoring data with enterprise production and business data in terms of spatiotemporal semantics, which leads to difficulties in tracing the source of carbon emission accounting results and low confidence, a multi-source heterogeneous carbon emission data fusion and alignment method is proposed. Summary of the Invention
[0004] This invention aims to provide a method for fusing and aligning multi-source heterogeneous carbon emission data. Through the dual constraints of physical mechanisms and business logic, it generates high-confidence, traceable carbon emission data products to support compliant carbon trading and asset management.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for fusing and aligning multi-source heterogeneous carbon emission data, the method comprising: Receive carbon emission observation data, including data from space-based satellites, airborne aviation, ground-based stations, IoT monitoring data, and resource planning business data; analyze production equipment entities and production batch events. Construct a cross-modal knowledge graph, use a temporal elastic matching algorithm to perform spatiotemporal semantic alignment, project carbon emission observation data onto a unified spatiotemporal grid system, and construct an emission contribution mapping relationship from the grid scale to the equipment scale based on the geographical coordinates of production equipment entities; The observation uncertainty of each data source at its respective spatiotemporal scale is quantified, and an observation error matrix is constructed. Based on the observation error matrix, a carbon asset confidence evaluation index is dynamically generated, and a high-confidence data interval that meets the preset carbon trading verification standard is identified and set as the spatiotemporal trajectory calibration domain. A hierarchical Bayesian model is constructed. Based on the evolution path of the spatiotemporal trajectory calibration domain, a nonlinear observation mapping is established using a two-stream neural network. The joint probability distribution is established by combining physical prior constraints and the observation error matrix. Based on the optimal estimation theory, the overall optimization objective function is constructed and solved iteratively using the expectation-maximization algorithm. The optimal reconstructed carbon emission field is obtained by inversion.
[0006] Preferably, the parsing of the production equipment entity and production batch event specifically includes: The received carbon emission monitoring data is cleaned, denoised, and preprocessed for time synchronization; production work order information, equipment ledgers, and process logs are extracted from the resource planning business data, and the relevant production equipment entities and their unique identifiers are determined through keyword matching and semantic recognition; the equipment power curves and operating status signals in the IoT monitoring data are matched with the planned time windows in the production work order information, and the peaks and troughs of the actual operation of the equipment are identified using edge detection algorithms to determine the actual start and end times of the production batch and generate production batch events.
[0007] Preferably, the specific steps of cross-modal knowledge graph entity alignment include: A two-layer knowledge graph containing geospatial entities and business asset entities is constructed, defining processing units, emission sources, and geographic polygon classes. Entity nodes and relational edges of the knowledge graph are automatically generated from the enterprise asset ledger. Geospatial entities are embedded in a first vector space, and business asset entities are embedded in a second vector space. Using known chimney emission outlet coordinates as a supervision signal, an alignment model is trained to maximize the corresponding likelihood probability of non-seed nodes. Combining spatial reasoning rules, erroneous matches from unrelated sources upstream of the wind field are eliminated. The emission intensity weights of each production equipment entity are calculated based on ground monitoring data. The weights are then used to downscale and decompose the spatiotemporal network data observed by satellite to quantify the emission contribution ratio of each enterprise asset entity.
[0008] Preferably, the construction of the cross-modal knowledge graph utilizes a temporal elastic matching algorithm to perform spatiotemporal semantic alignment, mapping carbon emission observation data to production equipment entities and production batch events. Specifically, this involves performing alignment with dual physical constraints in both spatial and temporal dimensions. In the spatial dimension, a two-layer knowledge graph of geographic and business entities is constructed and embedded into the vector space respectively; the known emission outlet coordinates are used as supervised seeds to train the alignment model, and the wind field transmission direction is introduced as a spatial reasoning rule to eliminate upstream mismatches. The mapping relationship between satellite observation pixels and enterprise assets is established based on the maximum likelihood probability. In the temporal dimension, a physical lag window based on process flow length and gas flow rate is set; within this window, a strip constraint is introduced, and a dynamic programming algorithm is used to search for the regular path with the minimum cumulative distance. Based on this, the emission data timestamp is reconstructed to achieve physical process-driven time-series alignment.
[0009] Preferably, the quantification of the observation uncertainty of each data source at its respective spatiotemporal scale and the construction of the observation error matrix specifically includes: using a Gaussian distribution model to model the uncertainty of the carbon emission observation data; for any observation value from any data source, its uncertainty is defined as the result of the actual carbon emission hidden state after being mapped by the observation operator and then superimposed with the observation error term; Specifically, the observation error term is assumed to follow a zero-mean Gaussian distribution, and a corresponding observation error matrix is constructed. The diagonal elements of the matrix are determined by three parts: instrument calibration error, which characterizes the measurement noise of the sensor hardware itself; inversion algorithm error, which characterizes the model bias in the process of converting the original spectral and / or voltage signals into carbon emission values; and representativeness error, which characterizes the scale difference uncertainty caused by the mismatch between the spatiotemporal resolution of the observation data and the spatiotemporal network system. In this way, the inverse matrix of the observation error matrix is used to adaptively adjust the weight contribution of each data source in the joint probability distribution.
[0010] Preferably, the step of dynamically generating carbon asset confidence evaluation indicators based on the observation error matrix, identifying high-confidence data intervals that meet preset carbon trading verification standards, and setting them as spatiotemporal trajectory calibration domains, specifically includes: The diagonal elements of the observation error matrix are extracted as variance features. Combined with instrument accuracy level and data integrity metadata, a mapping relationship from mathematical uncertainty to operational confidence score is constructed to classify carbon emission data within the spatiotemporal network. High-confidence data intervals that meet the maximum permissible error are selected using a preset carbon emission MRV standard and set as rigid constraint boundaries for Bayesian inversion. Within this boundary, the corresponding element values in the observation error matrix are reset to the preset minimum variance threshold, and the observation data in this boundary region are given the maximum update weight in the Bayesian update step. This forms a strong prior constraint, driving the evolution trajectory of the hidden state to converge toward high-confidence observations, thereby correcting the drift of the entire trajectory.
[0011] Preferably, the construction of the hierarchical Bayesian model, based on the evolution path of the spatiotemporal trajectory calibration domain, utilizes a two-stream neural network to establish a nonlinear observation mapping, and combines physical prior constraints and the observation error matrix to establish a joint probability distribution, specifically includes: A two-stream neural network comprising a mechanism calculation branch and a data-driven observation branch is constructed as a nonlinear observation operator. The mechanism calculation branch is a sequence of theoretical emission values calculated based on industry emission factors by temporally coupling the static carbon content of raw materials with the dynamic equipment feed rate or energy consumption curve. The data-driven observation branch takes the aligned carbon emission observation data as input and generates the observed emission values through a deep network. High-confidence data within the spatiotemporal trajectory calibration domain are set as strong constraint anchors for the evolution of the hidden state, limiting the output bias of the dual-stream network and preventing the inversion results from drifting in the observation blind zone. When establishing the joint probability distribution, the law of conservation of mass and the atmospheric transport convection-diffusion equation are introduced as physical prior constraints, requiring that the total output carbon element is balanced and that the emission field satisfies the laws of fluid dynamics in its spatiotemporal evolution. A likelihood function is constructed in conjunction with the observation error matrix to establish a joint probability distribution that includes the hidden state of the real carbon emission field, multi-source observation data, and physical constraints.
[0012] Preferably, the overall optimization objective function consists of the following weighted terms: Adaptive robust data fitting term: An adaptive robust kernel function is used to measure the deviation between the observed and predicted values. The penalty for outliers is dynamically adjusted according to the distribution characteristics of the data residuals to suppress the influence of non-Gaussian observation noise. Prior background field constraint: measures the degree of deviation between the inversion result and the historical prior background field; Graph Laplacian Regularization: Based on a two-layer knowledge graph, an adjacency matrix is constructed, the graph Laplacian matrix is calculated, numerical smoothing constraints are applied to production equipment nodes with semantic connections in the graph, and business logic is used to correct observation blind spots; Atmospheric transport PDE residuals: Based on the atmospheric convection-diffusion equation, the residuals of the partial differential equation are constructed, and the deviations between the time derivative of the concentration field, the convection term, the diffusion term and the source term are directly incorporated into the loss function to distinguish between local emissions and external transport.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By constructing cross-modal graphs and time-series flexible matching, isolated carbon emission observations are accurately mapped to specific production equipment and production batches, solving the data traceability problem and providing a strongly relevant business basis for enterprises' carbon asset confirmation and refined management.
[0014] 2. By using the observation error matrix and MRV verification standards to screen high-confidence calibration domains as rigid constraints for Bayesian inversion, interference from low-quality data is effectively eliminated, ensuring that the final accounting results meet the stringent audit requirements of the carbon trading market and reducing the risk of corporate compliance.
[0015] 3. By introducing a dual-stream neural network and atmospheric transport partial differential equations as strong prior constraints, a hierarchical probability model was constructed for iterative optimization, which effectively corrected the trajectory drift of the pure data-driven model in the observation blind zone and truly restored the spatiotemporal evolution and attribution of carbon emissions. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of a multi-source heterogeneous carbon emission data fusion and alignment method according to the present invention. Figure 2 This is a flowchart illustrating the process of parsing production batch events in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the spatiotemporal semantic alignment process in Embodiment 1 of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figures 1 to 3 This invention provides a method for fusing and aligning multi-source heterogeneous carbon emission data, referring to... Figure 1 The flowchart and technical solution are as follows: Receive carbon emission observation data, including data from space-based satellites, airborne aviation, ground-based stations, IoT monitoring data, and resource planning business data; analyze production equipment entities and production batch events. Construct a cross-modal knowledge graph, use a temporal elastic matching algorithm to perform spatiotemporal semantic alignment, project carbon emission observation data onto a unified spatiotemporal grid system, and construct an emission contribution mapping relationship from the grid scale to the equipment scale based on the geographical coordinates of production equipment entities; The observation uncertainty of each data source at its respective spatiotemporal scale is quantified, and an observation error matrix is constructed. Based on the observation error matrix, a carbon asset confidence evaluation index is dynamically generated, and a high-confidence data interval that meets the preset carbon trading verification standard is identified and set as the spatiotemporal trajectory calibration domain. A hierarchical Bayesian model is constructed. Based on the evolution path of the spatiotemporal trajectory calibration domain, a nonlinear observation mapping is established using a two-stream neural network. The joint probability distribution is established by combining physical prior constraints and the observation error matrix. Based on the optimal estimation theory, the overall optimization objective function is constructed and solved iteratively using the expectation-maximization algorithm. The optimal reconstructed carbon emission field is obtained by inversion.
[0019] Example 1 This embodiment is set in the application scenario of precise carbon emission monitoring and digital accounting in industrial parks and high-energy-consuming enterprises. In response to the technical bottlenecks such as inconsistent observation scales, spatiotemporal semantic breaks and strong heterogeneity of data sources among the sky, air, ground, Internet of Things and enterprise business systems, a deep fusion processing system is constructed, which integrates cross-modal entity alignment, physical process-driven time series regularization, multi-dimensional observation uncertainty quantification and Bayesian inversion under physical prior constraints.
[0020] The system receives multi-source carbon emission observation data through a preset data interface protocol. This data specifically includes gridded carbon dioxide column concentration data from the Sentinel-5P satellite, local emission plume image data collected by UAV-borne sensors, real-time concentration values transmitted from the CEMS flue gas online monitoring station deployed in the production plant, and power curves of PLC equipment and production resource planning records from the ERP system obtained through industrial Ethernet.
[0021] Furthermore, the received carbon emission observation data is input into a preprocessing module based on a 1D-CNN architecture to perform cleaning and denoising. This network includes an input layer with 32 convolutional kernels, two hidden layers with 64 neurons each using the ReLU activation function, and an output layer. By performing local mean filtering with a sliding window width of 50 sampling points, the high-frequency power jitter signal that originally contained electromagnetic interference is transformed into a smooth time-series curve.
[0022] Specifically, the interpolation alignment operator based on the UTC unified time reference is used to perform time synchronization on the processed data. The satellite observation grid data with a sampling frequency of once every 5000 milliseconds and the IoT monitoring power data with a frequency of once every 1000 milliseconds are subjected to third-order spline interpolation operation to output a standardized time series dataset with a unified time resolution of 60 seconds.
[0023] Furthermore, production work order information tables, equipment asset ledgers, and process operation logs from the DCS system are retrieved from the resource planning business data stored in the relational database. A semantic recognition model based on the Transformer architecture is used to extract entities from the text fields. This model adopts a 6-layer encoder structure and is configured with 8 parallel attention heads. The hidden layer dimension is set to 512, and the probability weight of each candidate string belonging to the production equipment entity is output through the Softmax activation function.
[0024] Specifically, by comparing the identification results with a probability weight exceeding a preset threshold of 0.95 with the existing equipment topology index of the factory, specific equipment entities involved in emission behavior are identified, such as the refining reactor with the unique identifier string RF-2024-001, and the rated emission coefficient 5.2 and installation geographical coordinate information associated with it are extracted simultaneously.
[0025] Furthermore, a Boolean correlation is performed between the real-time power signal sequence in the IoT monitoring data and the planned start and end time windows in the production work order information. An improved edge detection algorithm is used to perform local extremum retrieval on the first derivative of the power curve. The moment when the power change rate exceeds a preset threshold of 50kW per unit time is calculated as a potential state switching point.
[0026] Specifically, refer to Figure 2The flowchart of the production batch event is shown in the figure. The decision operator determines the actual start time of the production batch by matching the sampling segment where the power is continuously higher than the rated load by 30% in the identified set of state switching points, and defines the time when the power drops to 5% below the baseline standby level as the termination time. Finally, a production batch event object is generated, which includes a unique work order number, associated equipment identifier, actual start time of 08:30 on March 15, 2024 and termination time of 17:45.
[0027] By mining work order information and energy consumption waveform characteristics, actual production batch events can be accurately segmented. This effectively solves the problem of inconsistency between planned time and execution time, refines the accounting granularity to the single order level, and significantly improves the real-time performance and business relevance of carbon footprint tracking.
[0028] By using a pre-set data extraction and transformation tool, structured ledger forms containing a list of enterprise fixed assets and geographic information system vector data containing coordinates of factory boundaries are extracted to construct a two-layer knowledge graph. In the metadata layer of the graph, three types of core entities are defined, including processing units, fixed emission sources, and geospatial polygons. The equipment subordination and geographic association fields in the ledger are parsed using regular expressions to automatically generate topological directed edges connecting the nodes of each entity, including physical connectivity, process dependency, and spatial inclusion. Finally, the generated nodes and edges are loaded into a distributed graph database in the format of the Resource Description Framework (RDF).
[0029] Furthermore, the topological features of each node output from the graph database are input into an entity embedding model based on a dual-path graph convolutional neural network. Each path of this model consists of three stacked graph convolutional layers, each containing 256 hidden layer neurons and using LeakyReLU as a non-linear activation function. The first path receives the normalized geospatial entity features and maps them to a 256-dimensional first vector space. The second path synchronously receives the process parameter features of the business asset entities and maps them to a second vector space of the same dimension.
[0030] Specifically, the correspondence between the latitude and longitude coordinates of the three chimney emission outlets measured on-site and the ledger codes is extracted as the anchor point supervision signal. The Adam optimizer is used to train the dual-path graph convolutional neural network for 500 rounds of backpropagation iterations with an initial learning rate of 0.001. During the training process, the matching probability of entities is evaluated by calculating the cosine similarity of the embedding vectors of corresponding non-seed nodes in the first vector space and the second vector space. The model is made to converge in the direction of maximizing the alignment likelihood probability of non-seed nodes by minimizing the cross-entropy loss function. Finally, a probability matrix table containing the mapping relationship between geographical entities and business asset entities is output.
[0031] Furthermore, the probability matrix table is input into the spatial reasoning parser, and real-time meteorological tensor data containing easterly wind direction and wind speed of 5 m / s collected by the plant's meteorological micro-station is retrieved simultaneously. Spatial reasoning rules based on directed geometric vectors are constructed. By calculating whether the spatial coordinate vector of a specific production equipment entity is negatively multiplied by the current wind direction vector, it is determined whether it is upstream of the wind field. Incorrect matching pairs with no physical connection source upstream of the wind field are directly eliminated. Subsequently, the flue gas velocity of 20 m / s and carbon dioxide concentration of 50 mg / m³ measured by the ground CEMS online monitoring station are extracted. By dividing the real-time calculated emission flux of a single device by the total emission flux of the plant, a normalized value is obtained. Based on this, the emission intensity weight of each production equipment entity is calculated. For example, the weight score of a certain refining reactor is determined to be 0.35.
[0032] Specifically, the calculated emission intensity weight value is used as a decomposition factor and is used to perform element-wise dot product operation with the synchronously input satellite observation spatiotemporal grid carbon column concentration data matrix with a spatial resolution of 5 km × 5 km. Using a weight-based spatial downscaling decomposition algorithm, the coarse-grained satellite grid overall emission data is redistributed to the micro-level refining reactor asset entity nodes with a scaling factor of 0.35. The absolute emission value of the enterprise asset entity is accurately quantified and output as 120.5 kg per hour. Finally, the quantified emission flow data containing multi-dimensional traceability tags is pushed to the subsequent carbon asset accounting module for full life cycle management.
[0033] Furthermore, by utilizing the energy conservation constraints in the topology network, the energy consumption and emissions of devices without installed sensors are analyzed in reverse through the master table data, achieving a "soft measurement" style full-domain asset-level mapping; When constructing the business-related topology network, the system designates nodes with high-precision monitoring capabilities as "strong observation nodes" and branch equipment not covered by monitoring as "hidden nodes." Utilizing non-intrusive load decomposition technology, the system analyzes the high-frequency harmonic characteristics of the main incoming power meter and, combined with the equipment start-up and shutdown status identified by the operating condition fluctuation identification logic, decouples and distributes total energy consumption to each end-point emission source asset entity. Furthermore, the system combines the rated power and aging coefficient in the equipment ledger to construct a system of linear equations based on energy conservation. Using the decoupled virtual energy consumption data as input in the comprehensive metering mapping model, it calculates the independent carbon emission contribution value of each "hidden node" device. This is equivalent to installing a "virtual sub-meter" for each dummy device in the digital space, logically supplementing the blind spots of physical sensors. This method achieves granular coverage of all emission source assets in the plant at low cost without the need for large-scale addition of hardware sensors, solving the monitoring blind spot problem caused by wiring difficulties or cost constraints in industrial sites.
[0034] By utilizing a two-layer relational topology network and spatial transmission rules, macroscopic satellite grid data is accurately attributed to microscopic enterprise assets; environmental background interference is effectively eliminated, the uniqueness of the emission responsibility subject is established, and a strong logical mutual evidence is provided for the accurate and precise allocation of carbon assets.
[0035] By extracting geographic information system vector data containing the geographical coordinates of the enterprise's production workshop and text records of equipment ledgers stored in the resource planning system, a two-layer knowledge graph covering processing unit nodes, emission source nodes, and geospatial polygonal entities is constructed.
[0036] Furthermore, a graph convolutional neural network-based entity embedding model is used to perform feature vectorization processing on heterogeneous entities in the graph. The model architecture consists of three graph convolutional layers, each containing 512 neurons and using ReLU as the activation function. By receiving the adjacency matrix and feature matrix of the graph nodes as initial inputs, geospatial entities and business asset entities are mapped to the first vector space and the second vector space with 128 dimensions, respectively.
[0037] Specifically, the three-dimensional spatial coordinates of three pre-labeled fixed chimney emission outlets are extracted as supervision seed signals. The correlation probability distribution of corresponding nodes in the first vector space and the second vector space is calculated by introducing a contrastive loss function into the embedding model. The Adam optimizer is used to perform 500 rounds of gradient descent iterations with a learning rate of 0.001, aiming to minimize the distance between non-seed node pairs in the vector space, thereby outputting an initial mapping weight table with cross-modal alignment features.
[0038] Furthermore, referring to Figure 3 The flowchart of the spatiotemporal semantic alignment process introduces the wind field transmission direction as a hard physical constraint spatial reasoning rule in the spatial dimension alignment process. Specifically, it retrieves real-time meteorological tensor data containing a wind speed of 5 meters per second and a wind direction of 30 degrees north of east from the plant's micro-meteorological station, and constructs an association judgment operator based on directed geometric vectors.
[0039] The atmospheric physics basis is reasonable: when the wind speed is strong enough (usually >2m / s), the advection transport of pollutants is dominant.
[0040] Limitations of applicable conditions: (i) At low wind speeds (<2m / s), diffusion becomes dominant; (ii) The time scale of wind direction changes needs to be considered; (iii) Wind direction deflection and vortexes exist in complex terrain. Specifically, the decision operator receives the difference vector between the spatial coordinate vector of the production equipment entity and the center of the satellite observation pixel as input. By calculating the cosine of the angle between the difference vector and the wind direction vector, when the angle is in the obtuse angle range, it determines that the asset node is located in the upstream irrelevant region of the wind field of the satellite observation pixel, thereby directly resetting the matching probability weight of the pair of nodes to zero, and finally establishing the accurate spatial mapping matrix between the satellite observation pixel grid and the enterprise asset entity using the maximum likelihood probability.
[0041] Furthermore, during the time dimension alignment process, based on the extracted physical length of the production equipment outlet pipeline of 150 meters and the real-time monitored flow rate of the exhaust port of 15 meters per second, a physical lag window value containing a 10-second gas transport delay is obtained by performing a division operation.
[0042] Specifically, the system extracts the equipment power load curve recorded in the production batch event and the time series of satellite-observed emission intensity as dual input signals, and performs an initial time axis offset shift on the satellite observation signals according to the physical lag window value. In the time dimension, a dynamic alignment based on physical lag characteristics is adopted. First, the system obtains the theoretical residence time of the fluid inside the equipment from the production process specifications (calculated from the pipe length and flow velocity), and expands a specific physical lag search window before and after the equipment start-up and shutdown times accordingly.
[0043] Subsequently, within this window, a dynamic warping algorithm with band constraints is used to perform time-series alignment. The specific logic of this process is as follows: a difference matrix is established by calculating the distance between the device power load sequence and the satellite emission intensity sequence at each time point; to prevent the alignment process from deviating too much from the physical reality, a maximum number of offset points allowed to deviate from the diagonal is set, locking the search range within a preset physical lag time; a cumulative cost matrix is constructed using a recursive calculation method; and an optimal path with the minimum total difference is found from the endpoint using a backtracking technique. Finally, nonlinear time warp alignment is performed using this path to remap the timestamps of the original satellite observation data to the associated device power sampling times according to the correspondence in the normalized path.
[0044] Furthermore, a band constraint region with a width of ±300 seconds is introduced on the offset time axis, and a path regularization algorithm based on dynamic programming is used to search for a regularized path with the minimum cumulative distance.
[0045] Specifically, the path normalization algorithm constructs a local cost matrix with a dimension of 600×600, calculates the Euclidean distance between the equipment load value and the emission intensity value at each grid node, and uses recursive logic to find a path component that minimizes the sum of squared cumulative residuals of the entire sequence.
[0046] Furthermore, the timestamps of the original satellite observation emission data are reconstructed using the optimal regularized path obtained from the search, and the reconstructed observation time series data are mapped to the actual start and end time span of the production batch event.
[0047] Specifically, the reconstructed emission data stream with a unified time reference and the spatial mapping matrix are loaded into the carbon accounting processing engine, and the quantitative emission contribution results of each production equipment entity in a specific production batch are output, completing the physical process-driven alignment of carbon emission observation data from the macro grid to the micro production process.
[0048] By introducing dual physical constraints of wind field transmission and process lag, the spatiotemporal deviation between emission generation and monitoring response is corrected, spurious matching caused by simple statistical correlation is eliminated, and the data fusion process is ensured to fully conform to the objective evolution law of the physical world and business logic, thereby improving the accuracy of alignment.
[0049] Uncertainty modeling is performed on carbon emission observation data obtained from different monitoring nodes using a Gaussian distribution model. By inputting multi-source discrete sequences including satellite column concentration, UAV plume intensity and ground monitoring electrical signals, the error term of each observation data source is set to follow an independent Gaussian distribution with an expected value of 0, thereby establishing an initial fluctuation characteristic description for each data stream in the probability statistical space.
[0050] Furthermore, an observation error matrix is constructed in the logic processing unit to quantify the comprehensive deviation of different data sources. The diagonal elements of this matrix are jointly determined by three independent components—instrument calibration error, inversion algorithm error, and representativeness error—through a variance weighted superposition algorithm. This ensures that the matrix can fully map the comprehensive impact of hardware physical characteristics, software calculation deviations, and spatiotemporal scale differences on observation accuracy.
[0051] Specifically, the system obtains the instrument calibration error of 0.02 by parsing the XML format metadata provided by the sensor manufacturer. This value represents the inherent measurement noise and calibration residual of the sensor hardware circuit at 1200 degrees Celsius and a range of 50 milligrams per cubic meter, and uses this value as the first part of the basic increment of the diagonal element.
[0052] Furthermore, in order to determine the error components of the inversion algorithm, an error prediction model based on a multilayer perceptron architecture is constructed. This model includes an input layer consisting of 3 neurons to receive real-time signal strength, ambient temperature and humidity features, two hidden layers containing 64 neurons and 32 neurons respectively and using the ReLU activation function, and a single-neuron output layer.
[0053] Specifically, the error prediction model converts the input raw spectral absorbance or voltage signal into a corresponding calculated deviation component within the parameter space after 500 iterations of training by the Adam optimizer at a learning rate of 0.001. For example, for the ground-based online monitoring system, the output value is an algorithm deviation value of 0.15, and the square of it is used as the second part of the increment of the diagonal element.
[0054] Furthermore, when calculating the representative error components, the system constructs a scale difference evaluation function based on spatial variance weights by obtaining the 5000-meter spatial resolution value of the satellite observation grid and the 100-meter representative radius parameter of the ground monitoring station.
[0055] Specifically, the evaluation function calculates the spatial dispersion of each micro-sampling point within a satellite pixel relative to the pixel center value, quantifies the representative error of 0.08 caused by the mismatch between spatiotemporal resolution, and uses it as the third part of the increment of the diagonal elements, thereby completing the numerical filling of each diagonal element in the observation error matrix.
[0056] Uncertainty modeling of carbon emission observation data includes the following steps: Step 1: Uncertainty Modeling of the Observation Process. The system employs a Gaussian distribution model to model carbon emission observation data from various data sources, including space-based satellites, airborne UAVs, ground-based continuous emission monitoring systems, and IoT devices. For any data source's observation value at a specific time, the observation process is considered as the superposition of the result transformed from the hidden carbon emission field by the observation operator and the observation error term. The observation error term is set to a Gaussian distribution with a mean of zero, and its fluctuation is characterized by the observation error variance.
[0057] Step 2: Construction of the Observation Error Covariance Matrix. The system constructs a square matrix whose dimensions match the total number of data sources as the observation error covariance matrix. The off-diagonal elements of this matrix are set to zero, that is, based on the deployment location of each sensor and the principle of independent measurement, it is assumed that the observation errors of different data sources are independent and uncorrelated.
[0058] Specifically, the diagonal elements of the matrix represent the aggregate observation error variance of each data source, which is derived by linearly weighted summing of the following three independent error components: Instrument calibration error characterizes the inherent measurement bias of sensor hardware within a specified range. This value is extracted directly from the sensor's factory metadata or calibration documentation. For data given in the form of relative error (such as satellite observations), it needs to be multiplied by the order of magnitude of the current observation value to uniformly convert it to the dimension of absolute error.
[0059] The inversion algorithm error characterizes the model bias when converting raw sensor signals (such as spectral absorbance or voltage) into carbon emission values. This component is determined through offline calibration experiments: the model is trained by collecting multiple sets of high-precision reference standards and their corresponding raw sensor signals, and the standard deviation of the residuals between the model inversion values and the reference standard values is calculated, which is used as the algorithm error.
[0060] Representativeness error is used to characterize the scale difference caused by the mismatch between the spatiotemporal resolution of the observed data and the accounting grid system. During calculation, within a specific grid pixel range, all higher-resolution reference data within that range are extracted, and the variance of these reference data relative to the mean of that pixel is calculated. The square root of this variance is used as the representativeness error.
[0061] Step 3: Weighting and Integration: When calculating the variance of the integrated observation error, the system pre-sets weight coefficients for each error component. Specifically, the system pre-sets the weights for instrument calibration error, inversion algorithm error, and representativeness error to be 0.4, 0.3, and 0.3, respectively. These weight ratios are determined based on the degree of influence of each error term on data quality: since the inherent accuracy of the hardware sensors is the foundation of observation, instrument calibration error is given the highest weight; while the impact of algorithm conversion process and spatiotemporal scale differences on accuracy is relatively balanced, so the latter two errors are given equal weights to achieve an objective and comprehensive evaluation of the uncertainty of multi-source data. Furthermore, this weight can be dynamically adjusted according to actual conditions, proportionally weighting instrument calibration error, inversion algorithm error, and representativeness error. Through this modeling process, the quality level of carbon emission data from different sources before fusion can be accurately characterized, providing a data foundation for subsequent confidence assessment.
[0062] Furthermore, the constructed observation error matrix is loaded into the matrix operation engine, and the inverse matrix is calculated using the Cholesky decomposition algorithm.
[0063] Specifically, the inverse matrix is passed to the Bayesian joint probability distribution calculation module as a weight factor matrix. By performing a dot product operation between the inverse matrix and the innovation vector, the contribution ratio of data sources with large calibration errors or representativeness errors in the fusion process is adaptively reduced, thereby realizing the dynamic optimization allocation of the weights of each heterogeneous data source in the process of reconstructing the carbon emission trajectory.
[0064] By constructing a measurement error matrix to quantify data quality risks, adaptively reducing the weight of low-quality or scale-mismatched data, preventing individual faulty sensors from interfering with the overall accounting results, the robustness of the multi-source fusion system is significantly enhanced, and the credibility of comprehensive measurement data in complex environments is ensured.
[0065] The diagonal elements of the observation error matrix output by the preprocessing module are extracted as the core variance feature. This feature, in the form of a floating-point sequence, characterizes the degree of dispersion of each spatiotemporal sampling point at the physical observation level.
[0066] Furthermore, a confidence mapping model based on a multilayer perceptron architecture is constructed. The model architecture specifically includes an input layer with 3 neurons, two hidden layers with 64 and 32 neurons respectively and ReLU as the activation function, and an output layer with a single neuron.
[0067] Specifically, the input layer of the confidence mapping model receives the normalized variance features, the instrument accuracy level encoding value of 0.95, and the data integrity metadata reflecting the data packet successful reception rate of 0.98. The mathematical uncertainty components of the input are nonlinearly weighted by the hidden layer neurons, and the output layer is processed by the Sigmoid activation function to generate a business confidence score between 0 and 1. For example, the output score for the ground-based CEMS online monitoring point is 0.96.
[0068] Furthermore, the generated business confidence scores are used to perform asset classification operations on the carbon emission data within the spatiotemporal grid. Grid areas with scores above 0.9 are defined as primary core asset accounting areas, and grid areas with scores between 0.7 and 0.9 are defined as secondary auxiliary verification areas.
[0069] Specifically, the preset carbon emission monitoring, reporting and verification MRV standard is used as a filter operator. This operator sets a hard target of a maximum permissible error of 5%. It selects high-confidence data intervals that meet the error requirements (relative observation error is less than the permissible error of the industry MRV standard) and have a score higher than 0.9 from the asset classification results, and defines the coordinate sequence and time window of this interval as the spatiotemporal trajectory calibration domain.
[0070] The spatiotemporal grid point set that meets the following conditions is extracted as the spatiotemporal trajectory calibration domain: (1) confidence score greater than 0.90; (2) relative observation error meets industry MRV standards; (3) continuous distribution on the time axis (number of consecutive valid points greater than 3). The observations in this calibration domain will be used as strong constraint anchors in Bayesian inversion.
[0071] Furthermore, the observation data features within the spatiotemporal trajectory calibration domain are input into the error matrix reset module, which obtains the values of the diagonal elements in the corresponding observation error matrix.
[0072] Specifically, the error matrix reset module performs a logical branch decision to forcibly modify the values of elements falling within the calibration domain to the preset minimum variance threshold of 0.001, while keeping the original values of elements outside the domain unchanged, thereby outputting a corrected observation error matrix that has been non-uniformly enhanced.
[0073] Furthermore, the corrected observation error matrix and its inverse matrix are loaded as weighting factors into the update step operator of the Bayesian inversion engine.
[0074] Specifically, when the Bayesian inversion engine performs Kalman gain calculation, since the variance term in the calibration domain is reset to a minimum value of 0.001, the corresponding weighted gain coefficient tends to be maximized in matrix operations, so that the observations in this region account for more than 95% of the contribution ratio in the state update formula, thereby forming a strong prior constraint on the hidden state evolution trajectory.
[0075] Furthermore, the strong prior constraints are used to guide the hidden state evolution vector in the Bayesian inversion engine, and the carbon emission spatiotemporal trajectory in an uncertain drift state is forcibly pulled toward a high-confidence observation value through matrix dot product operations.
[0076] Specifically, the system compares the sum of squared residuals of the trajectory before and after calibration. When the residual value drops from the initial 15.2 to below 0.85, it completes the closed-loop correction of the spatiotemporal offset of the entire emission trajectory and outputs a quantitative carbon emission accounting report with high confidence labels.
[0077] Using a multilayer perceptron architecture, a network structure comprising an input layer, two hidden layers (with 64 and 32 neurons respectively, activated by a linear rectified function), and an output layer, nonlinearly maps observation errors, instrument accuracy, and data integrity. The normalized error compensation is then weighted and summed with various accuracy and integrity indicators according to preset weights, and a confidence score between 0 and 1 is output using a logistic regression function. Based on this score, carbon emission data is divided into three levels: areas with scores above 0.90 are designated as Level 1 core asset accounting areas, requiring high-precision monitoring and extremely low errors; areas with scores between 0.70 and 0.90 are designated as Level 2 auxiliary verification areas, used to meet routine verification needs; and areas with scores below 0.70 are marked as Level 3 early warning areas, considered substandard abnormal data, triggering a review mechanism.
[0078] By combining MRV standards to select high-confidence data as rigid calibration anchors, the accounting model is forced to lock the true value in key areas, effectively preventing data drift in sparse observation areas, ensuring that the final ledger meets the stringent audit requirements of carbon market trading level, and reducing corporate compliance risks.
[0079] A hierarchical Bayesian model is constructed to achieve deep fusion of multi-source heterogeneous carbon emission data. This model provides a probabilistic statistical framework for the inference of the entire carbon emission trajectory by establishing a hidden state evolution layer, a nonlinear observation mapping layer, and a data constraint layer.
[0080] Furthermore, in the nonlinear observation mapping layer of the hierarchical Bayesian model, a two-stream neural network containing a mechanistic calculation branch and a data-driven observation branch is constructed as a nonlinear observation operator. This operator is responsible for mapping the physical hidden state of carbon emissions to the multimodal observation space. The mechanistic calculation branch (MCB) adopts a multilayer perceptron architecture, comprising an input layer with four neurons (receiving raw material carbon content, equipment feed rate, equipment energy consumption, and equipment power load rate), two hidden layers each containing 64 neurons and employing the LeakyReLU activation function, and a single-neuron output layer. This branch couples the static raw material carbon content features with the dynamic equipment feed rate or energy consumption curve through a time-series point-by-point product (i.e., the raw material carbon content at the corresponding time is multiplied by the feed rate), sequentially performs nonlinear weight mapping through the hidden layers, and finally uses an industry emission factor table for linear scaling to calculate the theoretical emission value sequence of the equipment at each sampling time.
[0081] The Data-Driven Observation (DOB) branch employs a one-dimensional convolutional neural network (1D-CNN) architecture, comprising an input layer with 32 kernels, a kernel size of 3, and a stride of 1; an intermediate convolutional layer with 64 kernels, a kernel size of 3, and a stride of 1; a max-pooling layer with a sampling stride of 2; a fully connected layer with 128 neurons using the ReLU activation function; and finally, a single-neuron output layer (using linear activation). This branch receives spatiotemporally aligned carbon emission observation data sequences (time-by-time stitching of satellite, IoT, and CEMS observations), extracts multi-scale temporal features through convolution and pooling operations, and outputs real-time observed emission values.
[0082] The fusion of the two-stream network adopts a weighted fusion mechanism: the output of the MCB is the theoretical value vector, and the output of the DOB is the observation value vector. The two are fused together, with the fusion weights initially set to 0.6 and 0.4. They are dynamically adjusted based on the gradient descent method in the maximization step of the Bayesian inversion, so that the contribution ratio of the two branches to the final calculation result adapts to the quality of the observation data.
[0083] Furthermore, to ensure the training stability of the two-stream neural network, during the network initialization phase, the parameters of the mechanistic computation branch (MCB) are initialized using the Xavier method. The initial values of the weights are set within a specific range determined by the sum of the number of neurons in the input and output layers. Specifically, the range is based on the square root of the sum of the two neurons, with positive and negative values used as upper and lower bounds. The convolutional kernel weights of the data-driven observation branch (DOB) are initialized using the He method, with the bias term initialized to zero. Both branches are trained independently using the Adam optimizer, with an initial learning rate of 0.001 for both branches. An exponential decay strategy is adopted, multiplying the learning rate by 0.95 every 10 epochs. To prevent overfitting, a Dropout layer is introduced between the fully connected layer and the output layer, with a dropout ratio set to 0.3. The network training data comes from historically labeled carbon emission accounting datasets, using an 80% training and 20% validation data partitioning method.
[0084] Specifically, the mechanism calculation branch adopts a multilayer perceptron architecture, which includes an input layer consisting of 4 neurons, two hidden layers each containing 64 neurons and using the ReLU activation function, and a single-neuron output layer. This branch receives a static raw material carbon content feature of 0.85 extracted from the resource planning system, and performs time-series coupling calculation with the production equipment feed rate of 50.5 kg / s obtained in real time from IoT monitoring data. The hidden layer neurons are used to perform nonlinear weighted mapping on the input parameters, and output a theoretical emission value sequence based on the industry carbon balance logic.
[0085] Furthermore, the data-driven observation branch adopts a one-dimensional convolutional neural network architecture, which includes an input convolutional layer with 32 kernels and a kernel size of 3, an intermediate convolutional layer with 64 kernels and a kernel size of 3, a max pooling layer with a sampling stride of 2, and a fully connected layer containing 128 neurons. This branch receives carbon emission observation data sequences that have undergone spatiotemporal alignment processing as input.
[0086] Specifically, the data-driven observation branch uses the sliding window operation of the convolution kernel on the time axis to extract the emission plume fluctuation features in the observation data. After reducing the feature dimension through the pooling layer, it enters the fully connected layer and uses the Adam optimizer to perform backpropagation training with a learning rate of 0.001. Finally, the output layer generates an observed emission value that represents the real-time observation state.
[0087] Furthermore, high-confidence data within the spatiotemporal trajectory calibration domain identified by the preprocessing module are extracted as constraint components, and the observations within this region are defined as strong constraint anchor points for the evolution of hidden states by setting a hard physical threshold.
[0088] Specifically, when the system performs forward inference of the two-stream neural network, it calculates the Euclidean distance between the output value and the strong constraint anchor point as the residual bias term, and uses the penalty function to feed this bias term back to the network loss function. By limiting the output deviation of the two-stream network to within ±2%, it forces the evolution path of the inversion trajectory outside the observation blind zone to move closer to the high confidence anchor point, thus preventing the evolution results from undergoing physical spatiotemporal drift.
[0089] Furthermore, in establishing the joint probability distribution, the law of conservation of mass based on the principle of material balance and the atmospheric transport convection-diffusion equation based on atmospheric dynamics are introduced as physical prior constraints. The application conditions of the convection-diffusion equation are: the equation is applicable to environmental conditions where the diffusion coefficient changes relatively slowly (spatial scale > 100m) and the wind speed is relatively uniform (variance coefficient < 30%). However, for complex terrain or strong turbulence conditions, terrain correction factors or time-varying diffusion coefficient corrections need to be introduced.
[0090] Specifically, the physical constraint operator first calculates the difference between the total carbon input of 1000 kg of the production equipment entity and the total amount captured and emitted at the output end. It requires that this difference tends to zero after considering the measurement uncertainty. Simultaneously, meteorological parameters such as wind speed of 5 m / s and diffusion coefficient of 0.12 in the plant area are retrieved. The evolution process of the emission field is constrained by the fluid dynamics discretization operator, requiring that the mass flux between any adjacent space grid nodes satisfy the continuity condition of the convection-diffusion partial differential equation.
[0091] Furthermore, by combining the observation error matrix containing instrument calibration error and representative error constructed in the preceding module, the uncertainty weighting coefficient is obtained by performing an inversion operation on the matrix.
[0092] Specifically, the physical prior constraints, the observation mapping values output by the two-stream neural network, and the inverse of the observation error matrix are input into the Bayesian joint probability density function generator. By calculating the product of the likelihood function and the prior distribution, a joint probability distribution model containing hidden state variables of the real carbon emission field, multi-source heterogeneous observation datasets, and physical mechanism constraint terms is constructed.
[0093] Furthermore, the completed joint probability distribution model is loaded into the subsequent Bayesian inversion engine as the core logic.
[0094] Specifically, the inversion engine uses the Markov chain Monte Carlo algorithm to perform large-scale sampling in the joint probability distribution space. By calculating the posterior probability value of each set of sampled states and finding the global maximum likelihood solution, it finally outputs the quantified emission intensity of each production equipment entity in a specific production batch and the corresponding uncertainty confidence interval, thus completing the accurate alignment and fusion of multi-source carbon emission data under the dual drive of mechanism and data.
[0095] Furthermore, a four-dimensional balance verification mechanism is introduced, which includes raw material input, product output, energy consumption, and end-of-pipe emissions, to establish a shadow ledger independent of monitoring data and to identify and eliminate abnormal data caused by human tampering or equipment failure. When establishing the probability distribution of joint operations, the system does not rely solely on direct monitoring data but also operates a parallel "shadow accounting engine." This engine reads raw material purchase orders (to calculate input carbon) and finished product warehousing orders (to calculate solidified carbon) from resource planning business data, and calculates the theoretical "should-emit carbon" based on the law of conservation of mass. Simultaneously, it calculates "energy consumption-derived emissions" by combining electricity / gas meter readings from utilities. The "measured emissions," "material-derived emissions," and "energy consumption-derived emissions" are then triangulated. Only when the deviation among the three is within the allowable range of the preset measurement error matrix (e.g., ±5%) is the data marked as high confidence. If the monitored value is significantly lower than the material / energy consumption-derived value, the system automatically triggers an "anti-fraud warning" and forcibly lowers the carbon asset credibility score for that range. This first-principles-based cross-validation mechanism can effectively detect common carbon fraud behaviors such as "data laundering," "closing monitoring equipment," or "modifying sensor readings."
[0096] By integrating theoretical calculations with measured characteristics to construct a measurement model, and introducing material balance and fluid transport as strong priors, a reasonable distribution can be deduced based on business logic even in monitoring blind spots, ensuring the logical self-consistency of the ledger and realizing a leap from simple data fitting to physical business logic reasoning.
[0097] Furthermore, based on the optimal estimation theory, a joint optimization objective function is constructed, which is a weighted sum of an adaptive robust data fitting term, a prior background field constraint term, a graph Laplace regularization term, and an atmospheric transport PDE residual term. This objective function is then used as the core criterion for executing the expectation maximization iterative algorithm.
[0098] Specifically, the system extracts the initial theoretical emission value of 120.5 kg of carbon dioxide per hour generated by the preceding mechanism calculation branch for a specific refining reactor, as the zero-time initial distribution state of the hidden state vector in the Bayesian iteration sequence.
[0099] Further, the expectation step of the expectation-maximization algorithm is entered. The hidden state vector at the current time is mapped through the nonlinear observation operator in the two-stream neural network to obtain a set of predicted value sequences representing the grid predicted concentration. These predicted values are then subtracted from the measured carbon dioxide column concentration grid data to generate an information vector describing the degree of prediction bias.
[0100] Specifically, the adaptive robust data fitting term uses a robust kernel function with an adaptive kernel bandwidth parameter during the calculation process. By reading the residual value of each element in the innovation vector, when the residual value exceeds a preset threshold of 3 times the standard deviation, the penalty weight coefficient of the sampling point is automatically reduced linearly from 1.0 to 0.1, thereby suppressing the influence of non-Gaussian outlier observation noise introduced by extreme weather interference or instantaneous sensor interruption on the fitting results.
[0101] Furthermore, the prior background field constraint term is calculated by retrieving the average background field data of the previous month stored in the historical emission database, and calculating the Mahalanobis distance between the currently inverted carbon emission field hidden state vector and the historical background field vector, as a physical penalty term to measure the degree of deviation of the result.
[0102] Specifically, the graph Laplacian regularization term constructs an adjacency matrix based on the production equipment node associations defined in the aforementioned two-layer knowledge graph. By calculating the difference between the degree matrix of each node and the adjacency matrix, the graph Laplacian matrix is obtained. Second-order difference smoothing constraints are applied to production equipment nodes that have close semantic connections in geospatial space and process flow. Business logic is used to correct the observation blind spot data caused by sensor occlusion.
[0103] Furthermore, the atmospheric transport PDE residual term is constructed based on the principles of mass conservation and fluid dynamics. By extracting the horizontal wind speed vector of 5 m / s and the vertical diffusion parameter of 0.12 provided by the plant's micro-meteorological station, a partial differential equation residual containing the concentration field time derivative, three-dimensional convection term, grid diffusion term and equipment source term is constructed, and the residual value is directly accumulated into the total cost function.
[0104] Specifically, the partial differential equation residuals quantify and distinguish between source intensity components belonging to local equipment emissions and background concentration components belonging to external transport by comparing the inflow flux of the boundary grid with the emission source terms in the region, ensuring that the inversion results conform to the physical laws of material balance.
[0105] Further, the algorithm proceeds to the maximization step calculation process, where the Adam optimizer is used to perform gradient calculations on the joint optimization objective function, and the gradient components generated by each weighted term are superimposed as tensors in each iteration step.
[0106] Specifically, the algorithm logic performs hidden state vector correction with a step size of 0.005 along the direction of the synthesized negative gradient. By continuously adjusting the predicted emission flux of the production equipment entity, the total cost function converges to the global minimum in the parameter space. The objective function is solved by using the expectation-maximization algorithm. The optimal fusion estimate is obtained by alternately executing the E-step (expectation step, which calculates the expectation of the likelihood function under the current estimate) and the M-step (maximization step, which finds the parameter value that maximizes the expectation) until convergence.
[0107] A dynamic weight adjustment strategy is adopted in the fusion process, which adaptively adjusts the weight contribution of each data source in the objective function based on the uncertainty of each data source in the current spatiotemporal window and its consistency with other data sources. Iterative solution: Repeat the E-step and M-step until convergence to the optimal solution, which is the best estimated field after fusion.
[0108] Furthermore, to improve the robustness of the algorithm when there are large instantaneous errors in some data sources, a dynamic weight adjustment strategy is introduced.
[0109] In each iteration or each spatiotemporal window, calculate the dynamic weights of each data source; Calculate the median absolute deviation of each data source from all other data sources. If the median absolute deviation of a data source is significantly greater than that of the other sources, multiply its error variance by a penalty factor (e.g., 1.5), which is equivalent to reducing its weight in the objective function, thus obtaining the adjusted covariance matrix.
[0110] Furthermore, after each iteration, the relative rate of change of the cost function is calculated. When the decrease in the cost function for three consecutive iterations is less than the preset threshold of 0.0001 or the L2 increment of the hidden state vector is less than 0.01, the algorithm is deemed to meet the preset convergence condition.
[0111] Specifically, the system finally outputs the optimal carbon field reconstruction solution after multiple rounds of iterative optimization. This solution is represented as a numerical matrix with dimensions of 100×100×50×24, which is presented in the form of a structured grid as a dynamically evolving four-dimensional volumetric cloud map covering the longitude, latitude, altitude, and 24-hour time dimension of the entire plant area.
[0112] Furthermore, by utilizing high-precision micrometeorological data and fluid equations, a dynamic "responsibility exclusion domain" is delineated in the spatiotemporal network to distinguish between local emissions and external transmission interference; When constructing the spatiotemporal benchmark verification domain, the system integrates micro-meteorological station data (wind speed, wind direction, and air pressure) around the factory area, solves the inverse problem of the atmospheric transport convection-diffusion equation, calculates the theoretical contribution concentration of nearby upwind emission sources to the monitoring points in the factory area, and generates a "background subtraction mask". This mask is used to filter the grid data of space-based satellite or ground-based stations, removing pollutant concentrations that are external inputs. The clean data interval after background correction is marked as a high-confidence verification domain, serving as the "net value anchor" for Bayesian updates in the hierarchical probabilistic accounting model, forcing the model to learn parameters and update its state only for emissions actually generated locally. From a physical perspective, the system clarifies the emission responsibility boundaries between enterprises in dense industrial parks. By accurately separating background noise and external transmission, it ensures the fairness of the optimal carbon emission accounting, providing an operable technical criterion for solving complex regional cross-pollution attribution problems.
[0113] An efficient accounting closed loop is constructed using an iterative approximation algorithm to gradually correct hidden states to approximate the true emission distribution. The final output is a dynamic asset ownership ledger containing multi-dimensional spatiotemporal information, providing regulatory authorities with a visualized decision support tool for identifying abnormal emission behavior and analyzing diffusion paths. Furthermore, multi-dimensional compliance verification indicators are constructed to conduct automated audits from four dimensions: measurement consistency, historical benchmarks, asset association, and physical conservation. This ensures a logical closed loop for the accounting results across all dimensions, greatly reducing the cost of manual verification and significantly improving the efficiency and pass rate of carbon asset certification.
[0114] This embodiment realizes the entire accounting link from polymorphic monitoring signals to four-dimensional carbon field cloud maps by coupling a neural network with an uncertainty evolution engine. It uses robust kernel functions and graph regularization techniques to suppress noise interference and introduces partial differential equation residuals to ensure that the results follow physical laws such as mass conservation. This method can realize dynamic correction of emission trajectories while improving the audit strength and decision confidence of accounting data.
[0115] Example 2 This embodiment is set in a carbon emission monitoring scenario of an ultra-large integrated refining and chemical base, which includes high-altitude flare combustion, organized emissions from chimneys, and diffuse unorganized leaks in the production unit area.
[0116] Furthermore, in response to the transient nonlinear emission fluctuations generated during the controlled combustion process of the high-altitude torch, a physical information perception model based on the Transformer architecture is introduced to perform real-time inversion of combustion efficiency and emission factors. The model architecture consists of 4 layers of encoder units, each layer containing 8 independent self-attention heads. The hidden layer dimension is set to 128 and the Adam optimizer is used for weight update. Specifically, the physical information perception model uses the flame radiation energy tensor with a sampling frequency of 100Hz collected by the infrared thermal imaging sensor and the real-time gas mass flow rate data stream output by the ultrasonic flow meter as dual-modal inputs. The temporal features are mapped to a 128-dimensional vector space through the position encoding layer. The spatiotemporal correlation score between radiation intensity and combustion product concentration is calculated using a multi-head attention layer. A physical penalty term based on the energy balance principle is introduced into the loss function. By constraining the mass conservation relationship between combustion enthalpy and total product, the output layer generates the instantaneous emission flux value in each sampling period, realizing the second-level quantization of the dynamic emission process of the flare.
[0117] Furthermore, for low-concentration diffuse unorganized leakage caused by the failure of dynamic and static sealing points in the production unit area, a source resolution extraction model based on graph neural network is constructed. The model architecture includes 3 graph convolutional layers with 64 hidden neurons and uses ReLU as the activation function. It receives the carbon dioxide concentration time series matrix generated by 64 high-density IoT monitoring nodes in the plant area as input. Specifically, the source resolution extraction model synchronously loads the business asset feature vector containing reactor pressure and tower temperature parameters after alignment by the preceding module, uses the adjacency matrix to characterize the spatial topological relationship between monitoring points and physical equipment, aggregates neighborhood concentration gradient features through graph convolution operators, and uses a classification layer with a Softmax function to calculate the contribution weight of each potential leakage source to the grid concentration increase.
[0118] Furthermore, by performing a dot product operation between the calculated contribution weight score and the synchronously retrieved wind direction and speed vector data, specific equipment entities with a weight value exceeding 0.75 are identified. Specifically, the system uses this weight value to spatially redistribute the 500-meter grid concentration observed by satellite, accurately stripping away the diffuse grid volume and aggregating it to specific flange, valve, or pump body physical nodes, outputting the absolute emission contribution of this type of fugitive leakage source as 5.6 kg per hour. This embodiment achieves high-precision source tracing in complex multi-source emission scenarios at refining and chemical bases by coupling physical perception neural networks and graph neural networks; it uses the Transformer architecture to capture the transient evolution characteristics of flare combustion and combines spatial topological constraints to solve the problem of pinpoint quantification of unorganized leaks; the solution significantly improves the precision and real-time performance of carbon emission monitoring in complex industrial clusters through deep inference driven by multi-source data and strict alignment with physical boundary conditions.
[0119] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for fusing and aligning multi-source heterogeneous carbon emission data, characterized in that, The method includes: Receive carbon emission observation data, including data from space-based satellites, airborne aviation, ground-based stations, IoT monitoring data, and resource planning business data; analyze production equipment entities and production batch events. Construct a cross-modal knowledge graph, use a temporal elastic matching algorithm to perform spatiotemporal semantic alignment, project carbon emission observation data onto a unified spatiotemporal grid system, and construct an emission contribution mapping relationship from the grid scale to the equipment scale based on the geographical coordinates of production equipment entities; The observation uncertainty of each data source at its respective spatiotemporal scale is quantified, and an observation error matrix is constructed. Based on the observation error matrix, a carbon asset confidence evaluation index is dynamically generated, and a high-confidence data interval that meets the preset carbon trading verification standard is identified and set as the spatiotemporal trajectory calibration domain. A hierarchical Bayesian model is constructed. Based on the evolution path of the spatiotemporal trajectory calibration domain, a nonlinear observation mapping is established using a two-stream neural network. The joint probability distribution is established by combining physical prior constraints and the observation error matrix. Based on the optimal estimation theory, the overall optimization objective function is constructed and solved iteratively using the expectation-maximization algorithm. The optimal reconstructed carbon emission field is obtained by inversion.
2. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The parsing of production equipment entities and production batch events specifically includes: The received carbon emission monitoring data is cleaned, denoised, and preprocessed for time synchronization; production work order information, equipment ledgers, and process logs are extracted from the resource planning business data, and the relevant production equipment entities and their unique identifiers are determined through keyword matching and semantic recognition; the equipment power curves and operating status signals in the IoT monitoring data are matched with the planned time windows in the production work order information, and the peaks and troughs of the actual operation of the equipment are identified using edge detection algorithms to determine the actual start and end times of the production batch and generate production batch events.
3. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The specific steps for cross-modal knowledge graph entity alignment include: A two-layer knowledge graph containing geospatial entities and business asset entities is constructed, defining processing units, emission sources, and geographic polygon classes. Entity nodes and relational edges of the knowledge graph are automatically generated from the enterprise asset ledger. Geospatial entities are embedded in a first vector space, and business asset entities are embedded in a second vector space. Using known chimney emission outlet coordinates as a supervision signal, an alignment model is trained to maximize the corresponding likelihood probability of non-seed nodes. Combining spatial reasoning rules, erroneous matches from unrelated sources upstream of the wind field are eliminated. The emission intensity weights of each production equipment entity are calculated based on ground monitoring data. The weights are then used to downscale and decompose the spatiotemporal network data observed by satellite to quantify the emission contribution ratio of each enterprise asset entity.
4. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The construction of a cross-modal knowledge graph utilizes a temporal elastic matching algorithm to perform spatiotemporal semantic alignment, mapping carbon emission observation data to production equipment entities and production batch events. Specifically, this involves performing alignment with dual physical constraints in both spatial and temporal dimensions. In the spatial dimension, a two-layer knowledge graph of geographic and business entities is constructed and embedded into the vector space respectively; the known emission outlet coordinates are used as supervised seeds to train the alignment model, and the wind field transmission direction is introduced as a spatial reasoning rule to eliminate upstream mismatches. The mapping relationship between satellite observation pixels and enterprise assets is established based on the maximum likelihood probability. In the temporal dimension, a physical lag window based on process flow length and gas flow rate is set; within this window, a strip constraint is introduced, and a dynamic programming algorithm is used to search for the regular path with the minimum cumulative distance. Based on this, the emission data timestamp is reconstructed to achieve physical process-driven time-series alignment.
5. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The process of quantifying the observation uncertainty of each data source at its respective spatiotemporal scale and constructing an observation error matrix specifically includes: An uncertainty model is used to model the carbon emission observation data using a Gaussian distribution model, assuming that the observation error term follows a zero-mean Gaussian distribution. A corresponding observation error matrix is constructed, and the diagonal elements of the matrix are determined by three parts: instrument calibration error, which characterizes the measurement noise and calibration residual of the sensor hardware itself; inversion algorithm error, which characterizes the calculation deviation in the process of converting the original spectral and voltage signals into carbon emission values; and representativeness error, which characterizes the scale difference uncertainty caused by the mismatch between the spatiotemporal resolution of the observation data and the spatiotemporal network system. The inverse matrix of the observation error matrix is used to adaptively adjust the weight contribution of each data source in the joint probability distribution.
6. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The process of dynamically generating carbon asset confidence evaluation indicators based on the observation error matrix, identifying high-confidence data intervals that meet preset carbon trading verification standards, and setting them as the spatiotemporal trajectory calibration domain specifically includes: The diagonal elements of the observation error matrix are extracted as variance features. Combined with instrument accuracy level and data integrity metadata, a mapping relationship from mathematical uncertainty to operational confidence score is constructed to classify carbon emission data within the spatiotemporal network. High-confidence data intervals that meet the maximum permissible error are selected using a preset carbon emission MRV standard and set as rigid constraint boundaries for Bayesian inversion. Within this boundary, the corresponding element values in the observation error matrix are reset to the preset minimum variance threshold, and the boundary region observation data is given the maximum update weight in the Bayesian update step. This forms a strong prior constraint, driving the evolution trajectory of the hidden state to converge toward high-confidence observations, thereby correcting the overall trajectory drift.
7. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The construction of the hierarchical Bayesian model, which utilizes a two-stream neural network to establish a nonlinear observation mapping, establishes a joint probability distribution based on the evolution path of the spatiotemporal trajectory calibration domain and combines physical prior constraints with the observation error matrix, specifically includes: A two-stream neural network comprising a mechanism calculation branch and a data-driven observation branch is constructed as a nonlinear observation operator. The mechanism calculation branch is a sequence of theoretical emission values calculated based on industry emission factors by temporally coupling static raw material carbon content with dynamic equipment feed rate and / or energy consumption curves. The data-driven observation branch takes aligned carbon emission observation data as input and generates observed emission values through a deep network. High-confidence data within the spatiotemporal trajectory calibration domain are set as strong constraint anchors for the evolution of the hidden state, limiting the output bias of the dual-stream network and preventing the inversion results from drifting in the observation blind zone. When establishing the joint probability distribution, the law of conservation of mass and the atmospheric transport convection-diffusion equation are introduced as physical prior constraints, requiring that the total output carbon element is balanced and that the emission field satisfies the laws of fluid dynamics in its spatiotemporal evolution. A likelihood function is constructed in conjunction with the observation error matrix to establish a joint probability distribution that includes the hidden state of the real carbon emission field, multi-source observation data, and physical constraints.
8. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 1, characterized in that, The process of constructing a total optimization objective function based on optimal estimation theory, iteratively solving it using the expectation-maximization algorithm, and inverting to obtain the optimal reconstructed carbon emission field includes: Construct the overall optimization objective function and execute an iterative loop: set the theoretical values derived from the mechanism branch as the initial state; calculate the innovation vector between the predicted value and the actual observed value in the expectation step; use the gradient descent method in combination with the innovation vector to calculate the gradient, and correct the hidden state along the negative gradient direction to minimize the cost function in the maximization step; if the rate of change of the cost function and / or the state increment of the iterations meet the preset threshold, output the optimal carbon field reconstruction solution; otherwise, return to execute the expectation step; the optimal carbon field reconstruction solution is a dynamically evolving volumetric cloud map covering longitude, latitude, altitude and time.
9. The method for fusing and aligning multi-source heterogeneous carbon emission data according to claim 8, characterized in that, The overall objective function consists of the following weighted terms: Adaptive robust data fitting term: An adaptive robust kernel function is used to measure the deviation between the observed and predicted values. The penalty for outliers is dynamically adjusted according to the distribution characteristics of the data residuals to suppress the influence of non-Gaussian observation noise. Prior background field constraint: measures the degree of deviation between the inversion result and the historical prior background field; Graph Laplacian Regularization: Based on a two-layer knowledge graph, an adjacency matrix is constructed, the graph Laplacian matrix is calculated, numerical smoothing constraints are applied to production equipment nodes with semantic connections in the graph, and business logic is used to correct observation blind spots; Atmospheric transport PDE residuals: Based on the atmospheric convection-diffusion equation, the residuals of the partial differential equation are constructed, and the deviations between the time derivative of the concentration field, the convection term, the diffusion term and the source term are directly incorporated into the loss function to distinguish between local emissions and external transport.