Carbon emission data accounting method based on activity data tracing

By constructing a two-way mapping mechanism between activity semantic fingerprints and emission factor evolution maps, the problems of insufficient generalization ability and untimely dynamic adaptation of existing carbon emission factor matching technologies are solved, realizing lightweight and highly robust carbon emission accounting, which is suitable for efficient management of multi-industry and multi-granularity data.

CN122133881APending Publication Date: 2026-06-02CARBON LEAP FUTURE (QINGDAO) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CARBON LEAP FUTURE (QINGDAO) TECHNOLOGY CO LTD
Filing Date
2026-04-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing carbon emission factor matching technologies lack generalization capabilities, fail to adapt dynamically in a timely manner, rely on standardized formats for data preprocessing, and struggle to integrate multi-source, multi-granularity activity data in a hierarchical manner, resulting in low calculation accuracy and high costs.

Method used

By constructing a two-way mapping mechanism between activity semantic fingerprints and emission factor evolution maps, dynamic adaptation factors are generated, and core semantic elements are extracted using a lightweight coding method, enabling efficient matching and updating of cross-industry, multi-granular data.

Benefits of technology

It improves the accuracy and applicability of carbon emission accounting, reduces the cost of manual intervention, and supports efficient carbon emission management under varying operating conditions and dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133881A_ABST
    Figure CN122133881A_ABST
Patent Text Reader

Abstract

This invention relates to a carbon emission data accounting method based on activity data tracing, aiming to solve the problem of accurate adaptation and dynamic updating of emission factors under multi-source heterogeneous activity data. Its core technical solution includes: extracting and encoding core semantic elements from multi-source activity data to generate fixed-length activity semantic fingerprints; establishing a multi-dimensional emission factor evolution map based on historical emission factors and their change conditions; achieving dynamic screening and verification of candidate factors and data completion through bidirectional mapping between fingerprints and the map, and mapping the macroscopic and microscopic levels of the activity fingerprints to the map in parallel to achieve hierarchical weighted fusion output of standard carbon emission accounting factors; the system supports manual review signal feedback and real-time optimization and adjustment of map parameters. This solution improves the matching accuracy of emission accounting factors and the system's adaptability, supporting efficient dynamic carbon report generation and iterative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon emission data accounting and dynamic adaptation of emission factors, and in particular to a carbon emission data accounting method based on activity data tracing. Background Technology

[0002] With the advancement of the "dual carbon" policy and the development of the carbon trading market, carbon emission accounting systems are increasingly becoming a key tool for governments at all levels and enterprises to fulfill their energy management, emission disclosure, and compliance obligations. In the carbon emission accounting process, emission factors are the core parameters determining the accuracy of the accounting, regional adjustments, and traceability of process steps. Currently, mainstream carbon accounting systems generally adopt a method based on static matching of activity data with emission factor libraries. These factor libraries are typically developed according to IPCC guidelines, industry standards, or local supplementary documents, and are organized using a triplet (activity type, applicable conditions, factor value) or hierarchical dictionary structure. The system selects emission factors by matching activity message fields with applicable conditions in the library (such as region, equipment type, time period, etc.). This method can achieve automated preliminary accounting in scenarios with simple data types, clearly defined industry boundaries, or regulatory compliance.

[0003] However, as carbon accounting scenarios become increasingly detailed and enterprises demand more refined management, current mainstream carbon emission factor adaptation technologies are gradually revealing the following problems: First, existing emission factor models have limited generalization capabilities. Most systems in the industry only support standardized activity data for specific industries (such as thermal power, cement, and steel) or fixed granularity (enterprise annual, regional quarterly, etc.), and are poorly adapted to multi-source heterogeneous, cross-industry, and multi-granularity (such as enterprise, production line, equipment, and operating condition) activity data. Once activity processes or data collected from specific equipment across different industries occur, the system struggles to perform accurate and high-confidence automatic factor matching.

[0004] Secondly, the preprocessing and factor matching of activity data are highly dependent on data format specifications. Many existing solutions require raw data to be uploaded in a preset format or to undergo extensive formatting and structural reconstruction, increasing the data access threshold and the cost of manual intervention, making it difficult to adapt to the diverse data collection environments in actual business. In addition, for heterogeneous data from IoT terminals, manual reporting, production reports, and other channels, forced standardization often leads to information loss or semantic distortion, affecting the accuracy of accounting.

[0005] Third, existing systems lack effective tracking and intelligent adaptation mechanisms for the dynamic evolution of emission factors over time, region, policy, and technological upgrades. Current emission factor databases are mostly static tables or rule-based query models, lacking the ability to automatically manage changes in factor applicability conditions, version tracing, and dynamic updates. Changes in various factors, such as increased regional renewable energy share, equipment upgrades, and policy adjustments, are difficult to reflect in the accounting system in a timely manner relying solely on manual maintenance, leading to difficulties in tracing historical data and delays in adapting to dynamic scenarios.

[0006] Specifically, in terms of current technological implementation, existing related technical solutions often rely on heavy technical frameworks such as large-scale knowledge graph triple construction, semantic vector computation, or blockchain notarization in the process of mapping emission factors to activity data. These solutions require significant investment in knowledge annotation, sample training, and computing power, resulting in high barriers to implementation and promotion. They are also difficult to deploy efficiently and at low cost in application environments with multiple industries, processes, and granularities.

[0007] Furthermore, traditional matching strategies often lack hierarchical alignment and fusion mechanisms for multi-granularity data (such as enterprise-level monthly reports and device-level hourly data) for the same activity, resulting in limited responsiveness to the needs of differentiated business scenarios such as regulation and technological upgrades, and making it difficult to support innovative applications that require both macro-level compliance and micro-level evaluation.

[0008] Against this backdrop, the industry urgently needs a technical solution that can overcome the problems of insufficient generalization ability of existing emission factors, untimely dynamic adaptation, excessive reliance on format specifications for data preprocessing, and difficulty in hierarchically integrating multi-source and multi-granularity activity data. This invention addresses the shortcomings of existing technologies in intelligent matching of carbon emission factors by proposing an adaptive emission factor matching method based on activity semantic fingerprints and factor evolution maps. It aims to improve the accuracy and applicability of carbon emission accounting with lightweight, low-intervention, strong robustness, and high generalization ability, breaking through the technical bottlenecks of existing methods being "unusable, inflexible, and unintelligent." Summary of the Invention

[0009] This application provides a carbon emission data accounting method based on activity data tracing, which aims to solve one of the problems or issues of the prior art mentioned in the background section above.

[0010] The carbon emission data accounting method based on activity data tracing provided in this application specifically includes: S1: Acquire multi-source heterogeneous activity data composed of carbon emission data, extract the core semantic elements from the multi-source heterogeneous activity data, and discretize and encode the core semantic elements according to preset semantic slots to generate a fixed-length activity semantic fingerprint. S2: Based on historical carbon emission data, construct a directed time series network, and generate an emission factor evolution map based on the associated data in each node of the directed time series network; S3: Input the activity semantic fingerprint into the fingerprint graph bidirectional mapping engine, locate the set of initial nodes that can be activated in the emission factor evolution graph based on the spatiotemporal context and object attributes in the fingerprint, and generate a candidate factor node group; S4: Perform logical verification based on the candidate factor node group, remove unreachable paths that do not meet the implicit constraints, and generate a dynamic adaptation factor subset; S5: Using the dynamic adaptation factor subset to output factor adaptation feedback in reverse, identify the missing key semantic slots in the active semantic fingerprint and generate data completion guidance instructions. S6: Generate macro-level activity semantic fingerprints (including enterprise-level) and micro-level activity semantic fingerprints (including equipment-level) for the same activity data, and map them together to different level nodes of the emission factor evolution map to generate hierarchical projection matching results, thereby completing the carbon emission data accounting. S7: The hierarchical projection matching results are weighted and fused according to the preset business rules. If the business rules determine that it is a regulatory reporting scenario, macro-level node data is adopted first. If the business rules determine that it is a technical transformation assessment scenario, micro-level node data is adopted first. Finally, the carbon emission accounting factor is obtained. S8: Based on the final carbon emission accounting factor and data completion guidance instructions, update the emission factor evolution map to generate an iteratively optimized emission factor evolution map.

[0011] The carbon emission data accounting method based on activity data tracing provided in this application has the following beneficial effects: (1) To address the problems of existing carbon emission factor matching technologies that generally rely on large-scale semantic parsing models and static databases, resulting in slow system response, weak generalization ability, and high computational cost, this solution constructs a bidirectional mapping mechanism between activity semantic fingerprints and emission factor evolution maps, thereby achieving lightweight and robust semantic representation of complex production activities. This mechanism abandons the traditional natural language processing path and instead adopts a discretization encoding method based on preset semantic slots to extract core elements such as activity subject type, action verbs, object attributes, spatiotemporal context, and implicit constraints, generating fixed-length semantic fingerprints that do not depend on training corpora, effectively avoiding the bottlenecks of large language models in terms of deployment cost and data privacy; at the same time, the emission factors are organized into a directed temporal network with "changes in applicable conditions" as the edge, enabling each factor node to have dynamic context awareness, significantly improving the adaptability and accuracy of factor matching under changing working conditions, and is especially suitable for actual industrial scenarios where equipment operating status frequently changes and regional policies are dynamically adjusted.

[0012] (2) To address the problems of lagging updates, high manual maintenance costs, and difficulty in coping with cross-process chain linkages in traditional factor libraries, this solution innovatively introduces a dynamic pruning and closed-loop feedback collaborative mechanism to achieve intelligent path selection and data quality feedback during the fingerprint mapping process. When a new activity fingerprint is input, the system first activates multiple potential matching nodes in the graph and performs logical consistency verification by combining the implicit constraints that can be derived from the fingerprint (such as the load range deduced from the rated power). It dynamically eliminates unreachable paths that do not meet the operating boundary conditions, thereby avoiding mismatches to nominally similar but actually mismatched factor nodes. At the same time, the graph nodes can actively output "factor adaptation feedback", indicating that key fields are missing or data accuracy is insufficient, guiding users to supplement necessary information, forming a positive closed loop from matching failure to data completion, significantly improving the integrity and structural standardization of data collection, greatly reducing the intensity of manual intervention in subsequent accounting corrections, and enhancing the system's self-learning ability and engineering usability.

[0013] (3) To address the granularity conflict and inconsistency between macro-level reporting and micro-level monitoring in a multi-source heterogeneous data environment, this solution proposes a "fingerprint hierarchical projection" strategy and integrates a lightweight online learning mechanism, achieving a unified cross-scale data fusion and system continuous evolution capability. By generating semantic fingerprints at the macro (enterprise-year-province) and micro (equipment-hour-factory) levels for the same activity, and mapping them in parallel to different abstract level nodes of the graph, the solution automatically selects the best fusion output results according to preset business rules, effectively taking into account the differentiated needs of regulatory compliance and technological transformation assessment. At the same time, whenever the matching result is manually verified, the system only locally updates the weight parameters and trigger thresholds of the relevant graph edges, without needing to retrain the global model or reconstruct the knowledge structure, greatly reducing the computational resource consumption of model iteration and supporting long-term stable operation in edge or low-configuration server environments. The overall architecture completely avoids complex technical paths such as knowledge graph triple construction, tensor alignment, and digital twin modeling. It has strong practicality, high industry migration capability, and broad applicability, and is particularly suitable for application scenarios of accurate carbon emission accounting and dynamic management in high energy-consuming fields such as metallurgy, chemical industry, and cold chain transportation.

[0014] The aforementioned technical means jointly construct a semantically lightweight, structurally dynamic, and interactively closed-loop intelligent matching system for carbon emission factors. This not only significantly improves the relevance, timeliness, and interpretability of factor selection, but also fundamentally solves the systemic defects of existing technologies in terms of computing power dependence, update delay, multi-granularity conflicts, and cold start difficulties. It provides reliable technical support for achieving low-cost, high-efficiency, and wide-coverage digital governance of carbon footprint. Attached Figure Description

[0015] Figure 1 This is the main flowchart of the carbon emission data accounting method based on activity data tracing.

[0016] Figure 2 This is a sub-flowchart of a carbon emission data accounting method based on activity data tracing.

[0017] Figure 3 This is another sub-flowchart of the carbon emission data accounting method based on activity data tracing. Detailed Implementation

[0018] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0019] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0020] like Figure 1 As shown, this application provides a carbon emission data accounting method based on activity data tracing, specifically including: S1: Acquire multi-source heterogeneous activity data composed of carbon emission data, extract the core semantic elements from the multi-source heterogeneous activity data, and discretize and encode the core semantic elements according to preset semantic slots to generate a fixed-length activity semantic fingerprint. S2: Based on historical carbon emission data, construct a directed time series network, and generate an emission factor evolution map based on the associated data in each node of the directed time series network; S3: Input the activity semantic fingerprint into the fingerprint graph bidirectional mapping engine, locate the set of initial nodes that can be activated in the emission factor evolution graph based on the spatiotemporal context and object attributes in the fingerprint, and generate a candidate factor node group; S4: Perform logical verification based on the candidate factor node group, remove unreachable paths that do not meet the implicit constraints, and generate a dynamic adaptation factor subset; S5: Using the dynamic adaptation factor subset to output factor adaptation feedback in reverse, identify the missing key semantic slots in the active semantic fingerprint and generate data completion guidance instructions. S6: Generate macro-level activity semantic fingerprints (including enterprise-level) and micro-level activity semantic fingerprints (including equipment-level) for the same activity data, and map them together to different level nodes of the emission factor evolution map to generate hierarchical projection matching results, thereby completing the carbon emission data accounting. S7: The hierarchical projection matching results are weighted and fused according to the preset business rules. If the business rules determine that it is a regulatory reporting scenario, macro-level node data is adopted first. If the business rules determine that it is a technical transformation assessment scenario, micro-level node data is adopted first. Finally, the carbon emission accounting factor is obtained. S8: Based on the final carbon emission accounting factor and data completion guidance instructions, update the emission factor evolution map to generate an iteratively optimized emission factor evolution map.

[0021] Step S1: Acquire multi-source heterogeneous activity data composed of carbon emission data, extract core semantic elements from the multi-source heterogeneous activity data, and discretize and encode the core semantic elements according to preset semantic slots to generate a fixed-length activity semantic fingerprint. Specifically, this includes: S1.1: Obtain multi-source heterogeneous activity data, perform format parsing and field extraction on the original message based on the data source interface protocol, and generate a set of original semantic elements including activity subject type, action verbs, object attributes, spatiotemporal context and implicit constraints.

[0022] The system performs data source protocol identification on the input status of multi-source heterogeneous activity data sources to determine the interface protocol type and encoding method corresponding to each data source. For the identified interface protocol, the protocol parsing module is invoked to perform original message structure segmentation processing according to the protocol field definitions, splitting the message into several independently parsable logical units. Field extraction methods are applied to the segmented logical units, extracting the key business fields labeled in each data segment according to preset field mapping rules, and temporarily storing them in an intermediate buffer. Semantic tag binding is performed on the field data in the buffer, mapping the activity subject type field to the subject semantic tag set, the action verb field to the action tag set, the object attribute field to the attribute tag set, the spatiotemporal context field to the spatiotemporal tag set, and the implicit constraint field to the constraint tag set. Consistency verification is performed on the bound semantic tag set, removing tag entries with incorrect format or incomplete semantics to output the original semantic element set including activity subject type, action verb, object attribute, spatiotemporal context, and implicit constraints.

[0023] By using interface protocol-based parsing and field extraction, the multi-source heterogeneous raw messages from the previous step are transformed into a set of raw semantic elements with subsystem semantic integrity and subsequent normalization capabilities, thereby improving the semantic preprocessing capabilities of cross-source data.

[0024] For example, in a power company's carbon emission accounting system, three types of data sources are accessed: the plant boiler operation monitoring system transmits messages including fuel type, boiler model, and rated power via the Modbus protocol; the regional power grid dispatching system transmits power grid zone identifiers and real-time load interval messages via a custom TCP protocol; and the cold chain transportation management system transmits vehicle load capacity and temperature control interval parameter messages via HTTP / JSON interfaces. The parsing module calls the corresponding protocol parser to split the Modbus message into the coal type field "Coal Type A", the boiler model field "K-450", and the rated power field 450kW; it splits the power grid message into the power grid zone field "East China Zone 1" and the real-time load interval field 3.5 to 5MW; and it splits the JSON message into the vehicle load capacity field 15t and the temperature control interval field -18 to 0℃. Each field is bound to five semantic tag sets: subject, behavior, attribute, spatiotemporal, and constraint. Records with null values ​​or incorrect formats are removed after consistency checks. The final generated set of original semantic elements covers, at the macro level, enterprise-level activity subjects (boilers, vehicles), action verbs (combustion, transportation), object attributes (fuel type, load capacity), spatiotemporal context (power grid zone, temperature control zone), and implicit constraints (rated power corresponding to load zone). It can be directly normalized to standard domain terminology identifiers in the subsequent S1.2 step. Verification results show that the set has completeness and mappability under different data sources and different industry scenarios, which significantly improves the preparation efficiency of matching emission factors with multi-source heterogeneous data.

[0025] S1.2: Receive the original set of semantic elements, normalize various elements using a predefined semantic slot mapping table, convert unstructured text descriptions into standardized domain terminology identifiers, and generate a standard semantic element sequence.

[0026] The system receives the raw semantic element set generated by the preceding steps as input, parses the field types and source tags of each element in the set, and establishes an index association with a predefined semantic slot mapping table to ensure that elements from different data sources can correspond to a unified domain terminology library. For activity subject types, action verbs, object attributes, spatiotemporal contexts, and implicit constraints, the system calls the normalization rule set of their respective slots to perform condition-matched term replacement, transforming unstructured free text descriptions into standardized term identifiers, eliminating redundant expressions such as synonyms, aliases, or colloquialisms, and ensuring terminology uniqueness and comparability. For object attribute fields including numerical attributes, the system calls a numerical normalization function to adjust data with inconsistent units to a unified measurement system through conversion parameters, and records the conversion coefficients during the normalization process for subsequent coding. A multi-level priority matching method is used to determine fuzzy descriptions. When an element matches multiple standard terms, the optimal term identifier is selected based on a comprehensive ranking of industry priority, time label, and regional label, thereby eliminating coding ambiguity caused by polysemy. The normalized term identifiers are arranged in a preset order according to the semantic slots to form a standard semantic element sequence. Placeholders for missing slots are inserted into the sequence to maintain the consistency of slot positions in subsequent encoding steps. Through the above normalization process, the original semantic element set from the previous step is transformed into a standard semantic element sequence with standardization, uniqueness, and structural consistency, achieving low-error alignment of cross-source heterogeneous data within the same semantic space.

[0027] For example, in the scenario of processing activity data for enterprise boiler operation, the original semantic element set includes the subject type "coal-fired boiler", the action verb "combustion", the object attribute "coal GCV 4200 kcal / kg", the spatiotemporal context "East China Power Grid Zone, Altitude 70m", and the implicit constraint "no desulfurization device". The standard terminology library defining the subject type slots in the semantic slot mapping table includes "coal-fired boiler", "gas-fired boiler", etc., and the standard terminology library for the action verb slots includes "combustion", "heating", etc. The numerical normalization rules for the object attribute slots require that the uniform calorific value unit be MJ / kg.

[0028] S1.3: Based on the standard semantic element sequence, a discretization coding method is used to numerically map the term identifiers in each slot, converting continuous attribute values ​​or discrete category values ​​into fixed-width integer codes to generate a discretized semantic coding vector.

[0029] Based on the term identifiers of each slot in the standard semantic element sequence, a mapping control table from slot to numerical encoding is constructed to clarify the position range and bit width allocation rules of different categories of terms in the encoding space.

[0030] Continuous attribute values ​​are divided into intervals according to a preset quantization precision, and mapped to fixed-width integer codes through linear mapping or piecewise functions to ensure that the same attribute has comparability and consistency in different value ranges.

[0031] Discrete category terms are converted into integer encoded values ​​through hash mapping or dictionary indexing, and mutual exclusion and uniqueness between categories are maintained during the encoding process to avoid conflicts in encoding with the same bit width.

[0032] For composite slots that include both continuous and categorical attributes, hierarchical encoding is performed. First, the categorical part is integerized, and then the quantized encoding of continuous values ​​is embedded in the categorical encoding space to ensure the parsing of composite features.

[0033] A multi-slot encoding aggregation mechanism is adopted to store the integer codes generated by each slot in the discretized semantic encoding vector in slot order, and to record the slot index during the aggregation process for subsequent bit concatenation and alignment operations.

[0034] By using the above discretization coding method, the standard semantic element sequence of the previous step is transformed into a structured, computable discretized semantic coding vector, thereby achieving a unified numerical representation of multi-source heterogeneous semantic information and compatibility with subsequent fingerprint generation.

[0035] For example, in a cross-industry carbon management platform application, the activity subject slot "Coal-fired Boiler" is assigned an 8-bit width, corresponding to a category index code of 45; the "Fuel Calorific Value" in the object attribute slot is a continuous attribute, with a quantization step size set to 0.5 MJ / kg, an original value of 24.5 MJ / kg, and a calculated interval index of 49. The code value Code is calculated using the following formula: The index is 49, and the final encoded value is 49, stored in a continuous encoded segment with a bit width of 10 bits. The "altitude range" attribute in the spatiotemporal context slots uses segmented mapping; the segment index corresponding to an altitude of 520m is 3. The categorical slots are mapped to the encoded value 287 using a hash method. The above three types of codes are combined in slot order, i.e., coal-fired boiler category code 45, fuel calorific value code 49, and altitude range code 287 form a discrete semantic encoded vector [45, 49, 287]. In performance verification, this vector maintains uniqueness after subsequent bit concatenation and hash compression, and significantly improves mapping accuracy and recognition robustness across cross-industry data sources.

[0036] S1.4: Perform bit concatenation and padding operations on the discretized semantic encoding vector, align the encoding bits according to the preset fingerprint length specification, eliminate the dimension inconsistency problem caused by missing data, and generate a fixed-length semantic encoding string.

[0037] Bit width detection is performed on the discretized semantic coding vector at the input stage to identify the difference between the actual bit length of each semantic slot coding segment and the preset fingerprint bit width specification, so as to establish the initial reference parameters for coding alignment processing.

[0038] The detection results are processed by bit splicing order planning. Each slot encoding segment is arranged according to the semantic slot priority, and a bit splicing path table is constructed to ensure that the logical position of the semantic slot is not destroyed during the splicing process.

[0039] For coded segments that are too short, bit stuffing is performed. The stuffing value is selected from either zero stuffing or stuffing with a specific pattern code according to a preset stuffing mode. The overall bit length after stuffing is made consistent with the preset specification through a formula. For example: Where L is the current bit length, q is the number of padding bits, and 1 is the padding value.

[0040] Bit alignment is performed on the filled coding segments. The shift register model is used to adjust each coding segment to the starting position in the specification to avoid dimensional misalignment caused by missing data.

[0041] Perform a consistency check on the complete encoded string after the completion of bit concatenation and padding, calculate the check bits and compare them with the preset encoding check rules to confirm the validity and usability of the entire encoded structure.

[0042] By using the bit concatenation, bit padding and alignment methods described above, the discretized semantic encoding vector from the previous step is transformed into a fixed-length semantic encoding string, achieving the expected technical effect of consistent semantic slot structure and uniform bit width.

[0043] For example, in enterprise boiler activity data processing, the activity subject type slot encoding segment length is 4 bits, the action verb slot length is 3 bits, the object attribute slot length is 5 bits, the spatiotemporal context slot encoding segment length is 6 bits, the implicit constraint slot length is 2 bits, and the preset fingerprint length specification is 20 bits. If the action verb slot length is found to be insufficient, 1 bit needs to be padded, with a padded value of 0. After padding, the execution bit concatenation order is subject type → action verb → object attribute → spatiotemporal context → implicit constraint, and the initial length after concatenation is 20 bits. Using a shift register to align according to the slot start position, it is ensured that the data in each slot is within a fixed interval. When calculating the check bit, an even parity check is selected. The entire fixed-length string passes the check, and the final output fixed-length semantic encoding string can be directly used as a unique identifier input in the subsequent hash compression stage. The verification results show that this string can maintain parsing compatibility in cross-industry boiler and cold chain transportation cases and significantly improve the accuracy of subsequent matching factors.

[0044] S1.5: Based on the fixed-length semantic encoding string, a hash compression function is applied to perform feature dimensionality reduction and uniqueness verification, and finally outputs an activity semantic fingerprint with a fixed length and unique identification capability.

[0045] Step S2: Based on historical carbon emission data, a directed time-series network is constructed, and an emission factor evolution map is generated based on the correlation data in each node of the directed time-series network. Specifically, this includes: S2.1: Obtain the metadata set of historical emission factors, use natural language processing technology to extract entities and identify relationships from unstructured factor application condition text, extract multi-dimensional change triggering elements including time dimension, regional dimension and process dimension, and generate a standardized application condition change relationship sequence.

[0046] The system receives unstructured factor application condition text from the historical emission factor metadata set. A text preprocessing pipeline is established to remove redundant characters, standardize the encoding format, and segment conditional phrases to ensure the accuracy of subsequent entity extraction. A named entity recognition method based on a combination of lexical rules and statistical models is applied to the preprocessed text. In the time dimension, the factor effective date range and change trigger timestamp are labeled; in the geographical dimension, administrative division codes and power grid partition identifiers are labeled; and in the process dimension, equipment models, process names, and operating condition categories are labeled. A relationship recognition module is used to perform multidimensional association analysis on the labeled entities, identifying the conditional dependencies and triggering logic between time, geographical, and process entities, and mapping them to computable conditional expressions. Based on these conditional expressions, a multidimensional change trigger element set is constructed, and each element within the set undergoes standardized encoding processing. The time dimension uses dual encoding of absolute and relative time; the geographical dimension uses standard administrative division codes; and the process dimension uses equipment compatibility codes and operating condition category codes. The encoded multidimensional change triggering elements are sorted according to time sequence and causal dependency to generate a standardized applicable condition change relationship sequence with indexability. Serialization storage ensures that it can be directly called in subsequent graph modeling, achieving condition consistency matching across scenarios.

[0047] Through the above natural language processing and multidimensional factor association analysis, the unstructured text results of the previous step are transformed into a structured sequence of applicable condition change relationships, realizing the accurate extraction and standardized coding of factor applicable conditions, and providing unified and computable input data for constructing a directed temporal network skeleton.

[0048] For example, in the historical emission factor metadata set of a steel company, the applicable condition text includes "Since 2021, the proportion of newly installed wind power capacity in the East China Power Grid region has exceeded 30%, and the energy efficiency level of the electric arc furnace process has been upgraded to Level 2." Text preprocessing removes commas and redundant spaces, and unifies the encoding to UTF-8. The named entity recognition method labels the time dimension entity as "Since 2021," the geographical dimension entity as "East China Power Grid region," and the technological dimension entity as "Electric arc furnace process energy efficiency level upgraded to Level 2." The relationship recognition module generates a time-region-technical ternary association and extracts the trigger logic "the proportion of newly installed wind power capacity exceeds 30%." After encoding standardization, the time dimension is encoded as absolute time 20210101 and relative time T+0, the geographical dimension is encoded as administrative division 320000 and power grid partition code HDG, and the technological dimension is encoded as equipment model EAF and energy efficiency level code L2. The trigger logic is then converted into a conditional expression: Where P(wind) represents the proportion of wind power installed capacity, and 0.3 is the threshold. A standardized applicable condition change relationship sequence is generated according to time priority: (20210101, HDG, EAF, P(wind)>0.3), and stored in the condition sequence index library. This sequence can be directly matched with the corresponding factor nodes in the subsequent graph modeling process, ensuring the cross-industry availability and high-precision matching effect of change triggering conditions.

[0049] S2.2: Based on the applicable conditions change relationship sequence, the graph theory modeling method is used to instantiate each independent emission factor into a graph node object, and the extracted multidimensional change triggering elements are mapped into directed edge attributes connecting nodes to construct an initial directed temporal network skeleton including node topology and edge triggering logic.

[0050] The system receives a sequence of applicable condition change relationships as input. It initializes a container of factor node objects in memory, creating independent node instances grouped according to factor identifiers in the relationship sequence. Each instance corresponds to a unique factor identifier and reserves attribute fields for subsequent metadata enhancement. The system decomposes the multi-dimensional change triggering elements in the relationship sequence (time, region, and process dimensions) into features, mapping them to components of directed edge attribute values, and establishing a triggering condition parameter set to provide a parameter basis for subsequent edge logic judgment and temporal traversal. For each pair of factor nodes with a triggering relationship, it uses graph theory modeling methods to generate directed edge objects and writes the triggering condition parameter set into the edge's attribute structure, forming a connection path including triggering logic, ensuring that the edge can activate the node when the condition is met. It performs topology assembly on all nodes and edges, inserting the node object set and edge object set into the data structure of the directed temporal network skeleton, and generating an index table in the network for efficient retrieval of connection relationships between different factors. A topology integrity verification procedure checks the in-degree and out-degree distribution of nodes to confirm that the network skeleton does not have isolated nodes or circular dependencies, ensuring that the structure can be used for subsequent metadata enhancement and temporal consistency verification. Through the above chain-like derivation process, the standardized applicable condition change relationship sequence of the previous step is transformed into an initial directed temporal network skeleton with node topology and edge triggering logic, realizing the structured modeling of factor relationships and providing a basic framework for the construction of emission factor evolution map.

[0051] S2.3: Perform metadata enhancement processing on each node in the initial directed temporal network skeleton, and generate a set of enriched factor nodes carrying complete state description information by querying static attribute information such as fill factor value, confidence interval, effective period and device hardware compatibility list through the association database.

[0052] It receives the emission factor node objects output from the initial directed time-series network skeleton as the execution input for metadata enhancement processing, and loads the associated database index, including historical factor value records, confidence interval statistics table, periodic applicability condition list and equipment model compatibility mapping table.

[0053] The system invokes the foreign key matching mechanism between node identifiers and related databases, performs precise retrieval for each node, extracts the corresponding factor value fields, and synchronously reads the upper and lower limits of the confidence interval related to the factor. The correspondence between the values ​​and the intervals is confirmed through statistical consistency verification methods.

[0054] The time stamp field in the list of periodic applicable conditions is parsed, and the start and end times of the effective period of each node are mapped to the standard timestamp format. Combined with the international time zone conversion rules, the time is encoded in a consistent manner across regions, ensuring the uniformity of the effective period attribute in the global map retrieval.

[0055] Access the device model compatibility mapping table, read the list of compatible device models associated with the node factor, and serialize the model string into an integer code using the device classification coding rules to enable efficient execution of subsequent matching operations.

[0056] The aforementioned factor values, confidence intervals, effective periods, and device compatibility lists are encapsulated as static attributes into node objects. The integrity of the node status data structure is verified, and nodes with missing attributes or abnormal values ​​are removed and recorded in the difference log.

[0057] By associating metadata and populating attributes, the initial network skeleton built in the previous step is transformed into a set of enriched factor nodes with complete state description capabilities, thereby comprehensively enhancing node information and providing a sufficient data foundation for subsequent time-series consistency verification and conflict detection.

[0058] S2.4: Utilize the edge-triggered logic attribute in the enrichment factor node set to execute the temporal consistency check and conflict detection method, remove invalid connection edges that violate physical laws or temporal logic, and initialize the weights of the remaining valid edges to generate a standard directed temporal network that has undergone topology purification.

[0059] By utilizing the edge trigger logic attributes recorded in the enrichment factor node set, a precise matching operation is performed on the timestamp sequence and effective period of each directed edge to identify the temporal consistency between the triggering event and the applicable conditions of the factor. For cases where the time sequence is reversed, the triggering condition lags behind the start of the factor's effective period, or ends prematurely at the end of the factor's effective period, a physical rationality judgment is performed using the time logic rule base, and these are marked as invalid connecting edges. For marked invalid connecting edges, an edge deletion operation is performed in the network topology, and the in-degree and out-degree attributes of adjacent nodes are updated synchronously to ensure that the connectivity calculation in subsequent traversal processes is not disturbed. For the remaining valid connecting edges that have not been removed, a weight initialization parameter set is constructed based on the importance of the triggering condition, the triggering frequency, and the matching confidence, and an initial weight value is assigned to each edge using a normalized weighting strategy. To reduce the interference of noisy edge weights on network retrieval performance, variance smoothing is performed using the edge weight distribution characteristics to control the difference between discrete high weights and low weights within a preset range. By using the above-mentioned temporal consistency verification and conflict detection, topology purification and weight initialization processing methods, the enrichment factor node set of the previous step is transformed into a standard directed temporal network with a pure structure and ordered weights, thereby improving the accuracy and searchability of the connection relationships in the emission factor evolution map construction process.

[0060] S2.5: Based on the standard directed temporal network that has undergone topological purification, the node attributes and edge relationships are encapsulated into an indexable graph data structure using a graph serialization storage protocol, and finally an emission factor evolution graph with dynamic evolution capability and efficient retrieval is output.

[0061] like Figure 2 As shown, step S3 involves inputting the activity semantic fingerprint into the fingerprint graph bidirectional mapping engine, locating the initial set of activatable nodes in the emission factor evolution graph based on the spatiotemporal context and object attributes in the fingerprint, and generating a candidate factor node group. Specifically, this includes: The fingerprint bidirectional mapping engine includes: 1. Fingerprint parsing and query vector generation module The core is responsible for parsing the spatiotemporal context and object attribute encoding fields of the activity semantic fingerprint. Through bit segmentation and standard index transformation (power grid partition mapping, altitude range matching, and equipment model parsing), it extracts key-value index pairs and generates a graph query vector containing power grid partition identifiers, altitude range parameters, and equipment model codes, thus realizing the connection between the activity semantic fingerprint and the graph retrieval interface.

[0062] 2. Multidimensional Node Matching Module Based on the graph query vector, multi-dimensional matching is performed in the node metadata index of the emission factor evolution graph: the geographic granularity and equipment model dimension are split, and administrative division level matching and equipment compatibility list intersection operation are performed respectively. Then, high matching degree nodes are filtered through composite similarity calculation to generate a primary active node set and complete the initial screening of graph nodes.

[0063] 3. Timing Consistency Verification Module Parse the metadata of the activation period of the primary activation node, extract the start / end activation time, compare it with the current business timestamp in a time-series interval, remove expired and ineffective nodes, and generate a subset of time-series valid activation nodes to ensure the timeliness of the nodes and provide a valid foundation for subsequent expansion.

[0064] 4. Neighborhood Traversal and Association Extension Module Based on the directed edge connections in the emission factor evolution graph, a breadth-first neighborhood traversal is performed on the time-series valid nodes. The applicable conditions and triggering logic of the directed edges are parsed, and the triggering conditions are verified by combining the activity semantic fingerprint. The associated derived nodes are obtained, and after deduplication, an extended list of associated nodes is generated to achieve dynamic expansion of the node matching range.

[0065] 5. Node merging and confidence ranking module The primary valid nodes and extended associated nodes are hashed to remove duplicates and merged into a unified node set. The confidence score of each node (integrating geographical, device, and time series matching scores) is calculated using a segmented weighted cumulative algorithm. The nodes are sorted in descending order of confidence score and ascending order of factor value standard deviation, and then encapsulated into a candidate factor node group to complete the final output, forming a functional closed loop of "parsing-matching-filtering-expansion-merging".

[0066] 6. Bidirectional Mapping Core Support Module The built-in bidirectional index mapping mechanism, node metadata index library, and adjacent edge index structure enable forward retrieval from activity semantic fingerprints to graph nodes. On the other hand, it ensures the consistency between nodes and current activity semantics through node validity period and trigger condition verification, supporting bidirectional mapping throughout the entire process and avoiding matching deviations.

[0067] S3.1: Parse the spatiotemporal context encoding field and object attribute encoding field in the activity semantic fingerprint to extract key-value index pairs for graph retrieval and generate a graph query vector including power grid partition identifier, altitude range parameter and equipment model code.

[0068] The received activity semantic fingerprint encoding string is set to be parsed as a spatiotemporal context encoding field and an object attribute encoding field. The encoding string is separated into a semantic subset including power grid partition bits, altitude interval bits and equipment model bits by binary bit segmentation.

[0069] The extracted power grid partition field is standardized and indexed according to the power grid partition mapping table to generate a unique identifier for the corresponding geographical area in the table.

[0070] The altitude range field is mapped to a standard altitude range parameter set using a range segmentation matching method, and the range number is confirmed by a logical determination method. For example, if the altitude value is decoded to a value of 720, the range number is located in the "500-1000 meters" range according to the preset range table.

[0071] The equipment model field is parsed using the model identification rule table to generate the equipment model code, ensuring that the code is consistent with the equipment compatibility list index of the node metadata in the emission factor evolution map.

[0072] The above-mentioned power grid partition identifiers, altitude range parameters, and equipment model codes are combined into a query load according to a preset key-value pair format, and a graph query vector is constructed to perform multi-dimensional conditional retrieval in the node metadata index.

[0073] By parsing bit by bit and converting to a standard index, the fixed-length fingerprint data from the previous step is transformed into a structured query vector that can be used for emission factor evolution map retrieval, thus achieving precise integration between activity semantics and map retrieval interface.

[0074] For example, considering the semantic fingerprint encoding string of an activity reported by a company, the spatiotemporal context encoding field has a length of 12 bits, and the object attribute encoding field has a length of 10 bits. After bit segment separation, the power grid partition bit segment is parsed to obtain the original code "0110", which corresponds to the power grid partition identifier "CN-EC" in the mapping table. The altitude range bit segment is parsed to obtain the original code "1011", which is mapped to the altitude range parameter 750 meters and numbered "ALT-3" according to the range table. The equipment model bit segment is parsed to obtain the original code "100101", which corresponds to the model code "BLR-M420" in the model identification rule table. The three standardized results are assembled into a structured query vector {power grid partition identifier: CN-EC, altitude range parameter: ALT-3, equipment model code: BLR-M420}. During graph retrieval, this query vector is matched with the node metadata in the index, and nodes that meet the triple conditions of region, power grid, and equipment compatibility are output as primary matching candidates, which significantly improves retrieval efficiency and maintains matching accuracy in the high confidence range.

[0075] S3.2: Based on the graph query vector, perform multidimensional matching operations in the node metadata index of the emission factor evolution graph to filter out initial matching nodes that meet the geographical granularity constraints and equipment compatibility list requirements, and generate a primary set of activated nodes.

[0076] The graph query vector from S3.1 is decomposed into dimensions, and the geographic granularity field and the device model field are extracted as independent matching dimensions to form a matching input unit with multi-domain indexing capabilities.

[0077] A hierarchical matching method based on administrative division codes is applied to the geographic granularity field. The power grid partition identifier in the query vector is compared with the geographic granularity index field in the metadata of the emission factor evolution map nodes to obtain a pre-selected set of nodes that meet the regional coverage conditions.

[0078] Perform a set matching operation based on the device compatibility list on the device model field. Calculate the intersection of the device model code in the query vector and the model code in the device compatibility list in the node metadata record to filter out the node filter set that meets the hardware compatibility requirements.

[0079] Perform an intersection fusion operation on the regional matching result set and the device compatibility matching result set to form a candidate set of nodes that simultaneously meet the geographical granularity constraints and the device compatibility list requirements.

[0080] The index field and query vector in the node candidate set are combined to calculate the cosine similarity and Jaccard similarity coefficient, so as to prioritize the retention of node instances with a matching degree higher than a preset threshold, thus forming the initial active node set.

[0081] By using multidimensional matching operations and intersection fusion processing, the graph query vector generated in the previous step is transformed into a set of primary activation nodes that meet spatial constraints and hardware compatibility conditions, thereby achieving accurate initial screening of factor candidate nodes under cross-domain conditions.

[0082] S3.3: Perform time-series consistency comparison processing using the effective period metadata of each node in the primary activation node set and the current business timestamp to remove invalid nodes that are in an inactive state or have expired, and generate a subset of time-series valid activation nodes.

[0083] The initial set of activated nodes generated by S3.2 is parsed for its effective period metadata, extracting the start and end times of each node's effectiveness. The parsing results are then structured into a set of timestamp parameters for computation. Based on the current business timestamp input, a time consistency judgment matrix is ​​constructed, comparing the node's effective period with the business timestamp one by one to determine the node's effectiveness status. A time-series interval inclusion method is used to perform mathematical operations on the business timestamp and the node's start and end times to determine if the interval closure condition is met.

[0084] For example, for a set of primary active nodes from the East China Power Grid sub-region, the effective period of node A is from 00:00 on January 1, 2024 to 23:59 on December 31, 2024, the effective period of node B is from 00:00 on June 1, 2023 to 23:59 on December 31, 2023, and the effective period of node C is from 00:00 on March 1, 2024 to 23:59 on September 30, 2024. With the current business timestamp set to 10:30 AM on May 15, 2024, the effective period of each node is converted to a Unix timestamp. For example, node A starts at 1704067200 and ends at 1735660799; node B starts at 1685587200 and ends at 1704067199; node C starts at 1709251200 and ends at 1727711999. Substituting into the formula, node A satisfies the condition that T lies within a closed interval, node B is determined to be expired, and node C satisfies the condition that the interval is closed. Remove node B and calculate the time matching confidence for nodes A and C respectively. For example, node A has a high matching confidence with a coverage of 240 days, while node C has a low matching confidence with a coverage of 140 days, but is still valid. This forms a time-series valid active node subset {A,C}. This output is directly used for the neighborhood traversal expansion of S3.4, ensuring that the path expansion is based on the currently effective factor nodes, thereby significantly improving the timeliness and accuracy of subsequent matching results.

[0085] S3.4: Perform a neighborhood traversal operation based on the directed edge connection relationship of the time-series effective activated node subset in the emission factor evolution graph to obtain the associated derived nodes triggered by changes in applicable conditions, and generate an extended associated node list.

[0086] For each node in the time-series effective activation node subset, the adjacency edge index structure of the emission factor evolution graph is invoked to extract the directed edge element information that is directly connected to the node and whose edge attributes include the applicable condition change trigger logic.

[0087] Based on the extracted directed edge information, the time, region, and process dimension parameters in the triggering logic are parsed to generate an executable triggering condition judgment vector, and condition matching judgment is performed in combination with the encoded fields in the semantic fingerprint of the current activity.

[0088] The connecting edges that meet the triggering conditions are taken as valid extension paths. Their terminal node identifiers are added to the stack of nodes to be traversed. The weight values ​​and triggering time intervals on the extension paths are recorded for use in subsequent sorting of associated nodes.

[0089] A breadth-first traversal operation is performed on the node stack to visit nodes sequentially according to the topological hierarchy of the graph. This ensures that processed nodes are not visited repeatedly during the traversal process. The condition judgment process is triggered repeatedly when visiting each node to recursively obtain all reachable derived related nodes.

[0090] After the traversal is complete, all derived associated node sets obtained are deduplicated according to node identifiers, while retaining the triggering condition information of the corresponding edges, forming an extended associated node list for subsequent merging and sorting of candidate factor nodes.

[0091] By using a collaborative processing approach of neighborhood traversal and trigger condition determination, the direct matching results of the effective active node subset in the time series are expanded to include derived node data that includes potential adaptation paths, thereby achieving dynamic expansion and coverage improvement of the factor matching range.

[0092] For example, in a carbon emission accounting scenario for industrial boiler operation, the time-series effective active node subset includes node N1 (coal-fired boiler + East China power grid zone + medium-pressure steam), whose adjacent edge E1 trigger condition is "regional wind power installed capacity ratio ≥ 30% and boiler interlocking control modification completed", and the terminal node is N2 (coal-fired boiler + East China power grid zone + medium-pressure steam + high-load calibration). In the activity semantic fingerprint, the object attribute encoding corresponds to a boiler power value of 8MW, the power grid zone code in the spatiotemporal context is East China, the process status includes an interlocking control modification completion identifier, and the measured wind power installed capacity ratio is 35%, satisfying the E1 trigger condition. N2 is added to the node stack to be traversed, and when the adjacent edge E2 of N2 is judged, its trigger condition is "SCR is put into operation during the high-load calibration period", and the terminal node is N3 (coal-fired boiler + SCR operation status + high-load calibration). The operating load rate encoding in the activity semantic fingerprint is converted into a fixed interval, indicating that it is currently in the high-load segment and the SCR device is in operation, satisfying the E2 condition, and N3 is added to the extended associated node list. After breadth-first traversal, the final expanded list of associated nodes includes N2 and N3. After deduplication, the trigger condition description and edge weight information are retained, realizing the dynamic supplementation of the derived node set, which significantly improves the factor matching coverage and adaptability in this scenario.

[0093] S3.5: Perform deduplication and confidence sorting on the primary activated node subset and the extended associated node list to construct a final result set covering direct matching and indirect derivation paths, and generate a candidate factor node group.

[0094] Perform node uniqueness verification on the subset of time-series active nodes and the extended associated node list, apply hash fingerprint comparison method to detect node identifiers with consistent encoding, and move duplicate node entries into the redundancy buffer for removal.

[0095] A merged mapping table is constructed for the deduplicated node set. Key values ​​are summarized based on the factor values, effective periods, and device compatibility lists in the node metadata. Directly matched nodes and indirectly derived nodes are stored in the same index cluster to ensure complete coverage in subsequent sorting calculations.

[0096] The sorting weight calculation module generates a confidence value for each node in the merged mapping table. The calculation adopts a segmented weighted cumulative mode, in which the geographical granularity matching score, device compatibility matching score, and time series consistency score are accumulated according to a preset ratio after being normalized, as shown in the following formula: in, For geographic matching weight, Geographic matching score; Assign weights to devices. Assign a score to the device for matching. For time-series matching weights, The score is for time-series matching.

[0097] The confidence level is input as a sorting criterion into the multi-level sorting method module. This module first sorts the nodes in descending order of confidence level, and then sorts them in ascending order of standard deviation of factor values, so as to ensure that nodes with high confidence and stability are given priority in entering the result set.

[0098] The sorted node sequence is encapsulated into a candidate factor node group data structure, providing input objects for subsequent path pruning based on implicit constraints.

[0099] By using deduplication, merging, and confidence ranking methods, the direct matching and indirect derivation results from the previous step are transformed into a group of candidate factor nodes with complete coverage and clear priority, thereby achieving the expected technical effect of maximizing the matching range and prioritizing accuracy.

[0100] like Figure 3 As shown, step S4 involves performing logical verification based on the candidate factor node group, removing unreachable paths that do not satisfy the implicit constraints, and generating a dynamic adaptation factor subset. Specifically, this includes: S4.1: Obtain node metadata and object attributes from the activity semantic fingerprint in the candidate factor node group, use the inference big model to extract the nonlinear mapping relationship between the rated parameters of the equipment and the operating load, generate the implicit constraint derivation rule set, and establish the logical criterion basis from static attributes to dynamic operating conditions.

[0101] S4.2: Based on the implicit constraint derivation rule set, feature inversion calculation is performed on the spatiotemporal context of the activity semantic fingerprint, converting the discrete fingerprint encoding into a continuous device operating state interval, generating a derivable constraint quantification index to clarify the specific boundary conditions of the current activity in the actual physical scenario.

[0102] Each rule entry in the implicit constraint derivation rule set is parsed to identify rule conditions that are directly or indirectly related to the spatiotemporal context encoding field, and their mapping parameters are extracted to form a parameter index table that can be called for feature inversion.

[0103] The parameter index table is used to perform decoding mapping on the power grid zoning code, altitude range code and shift code in the activity semantic fingerprint, and the discretized symbol values ​​are converted into initial values ​​of physical quantities. For example, the power grid zoning code is mapped to the typical load curve, the altitude range code is mapped to the air pressure and fuel combustion efficiency curve, and the shift code is mapped to the operation duration curve.

[0104] The initial values ​​of the above physical quantities are input into the multidimensional state inversion model. The continuous operating state interval is calculated by combining the rated parameters of the equipment and the process mechanism equation. The state inversion model performs multi-parameter cross-fitting on the physical quantities according to the nonlinear mapping relationship defined in the rule set, and outputs a preliminary state matrix including load rate, fuel consumption rate, emission intensity, etc.

[0105] Through the above feature inversion and quantification methods, the rule parsing results of the previous step are transformed into deducible constraint quantification indicators with continuous value description capabilities, so as to achieve the expected technical effect that candidate factor nodes can be directly compared with physical boundary conditions in subsequent logical verification.

[0106] S4.3: Use derivable constraint quantification indicators to perform Boolean logic verification on the applicable condition edges of each node in the candidate factor node group, compare the effective period of the node annotation with the equipment compatibility list, and generate path reachability determination results to identify all potential conflicting paths that violate physical laws or operating condition limitations.

[0107] Based on the input data vector of derivable constraint quantification indicators, the applicable condition edges of each node in the candidate factor node group are verified one by one to form a set of consistency criteria between physical conditions and operating conditions. Using a Boolean logic comparison method, for the associated node attributes of each applicable condition edge, the effective period parameter is extracted and compared with the current business timestamp to determine the time interval inclusion, generating a time validity Boolean flag. For the device compatibility list, a list matching matrix is ​​established, and the device type code in the fingerprint is compared bit by bit with the node compatibility code field to generate a device compatibility Boolean flag.

[0108] S4.4: Based on the path reachability determination results, perform topology pruning on the connection edges in the directed time-series network, remove nodes marked as unreachable and their associated successor paths, and generate a subset of dynamic adaptation factors to ensure that the emission factors ultimately retained strictly meet the implicit constraints of the activity data.

[0109] The path reachability determination results are input into the directed temporal network topology processing module, selecting nodes marked as unreachable and their associated edges as pruning targets. Edge attribute clearing is performed on the pruning targets, and the node state table is called to obtain the successor node chain. Based on the chain dependency relationship, the node identifiers and edge indices of all successor paths are recursively identified. A topology chain-breaking method is applied to the recursively identified successor path set, removing the directed edges between the source node and successor nodes from the topology structure, while updating the network adjacency matrix to maintain the connectivity description of the residual network. An application condition integrity scan is performed on the node attribute set of the residual network, eliminating isolated nodes that have lost their association conditions due to successor path chain breaks, generating a basic set of dynamic adaptation factor nodes. The nodes in the basic set are sorted in descending order of confidence value, and nodes that are completely consistent with the implicit constraint quantification index are merged to obtain the final dynamic adaptation factor subset. Through this pruning process, the path reachability determination results of the previous step are transformed into a group of nodes that satisfy the implicit constraints, achieving precise constraint matching of the factor set.

[0110] Step S5: Using the dynamic adaptation factor subset to output factor adaptation feedback in reverse, identify the missing key semantic slots in the active semantic fingerprint and generate data completion guidance instructions. Specifically, this includes: S5.1: Parse and process the metadata constraint set of each dynamic adaptation factor node in the dynamic adaptation factor subset, and extract the mandatory constraint conditions, including the device compatibility list, effective period and geographical granularity, to construct the factor constraint feature vector, providing a benchmark reference object for subsequent consistency verification.

[0111] Receive a subset of dynamic adaptation factors from the output of step S4.4 as the input object of this sub-step, including a set of factor node metadata that has passed the implicit constraint logic verification.

[0112] For each factor node in the dynamic adaptation factor subset, the node metadata parsing module is invoked to perform field-level decomposition of its metadata structure to extract constraint-related information.

[0113] During the parsing process, the device compatibility list field is standardized and mapped using the device model index library to ensure that the model code is consistent with the industry-wide unified coding system and to avoid matching deviations caused by different record formats.

[0114] The effective period field is parsed for time intervals. A time series parsing engine is used to convert the original timestamp sequence into a continuous and comparable interval object, and its validity across years and quarters is verified to support subsequent time consistency comparison.

[0115] Geocoding conversion is performed on the regional granularity field. The geographic granularity mapping table is used to convert the free text regional description into administrative division code or power grid partition number, forming a standardized regional identifier that can be used for cross-comparison of multi-source data.

[0116] The three types of mandatory constraint fields mentioned above are combined into an input vector. According to the preset factor constraint feature extraction method, the normalized values ​​of each field are numerically encoded and arranged in slot order to generate factor constraint feature vectors.

[0117] In the numerical encoding process, a multi-hot encoding method is used for the device compatibility list field, an interval boundary value pair is used to represent the effective period field, and a unique integer mapping is used for the regional granularity field to ensure the structural consistency and comparability of the feature vectors.

[0118] By constructing factor constraint feature vectors, the constraint information in the dynamically adapting factor subset is transformed into numerical feature data that can be uniformly verified, thus achieving a benchmark reference for subsequent consistency verification.

[0119] S5.2: Based on the factor constraint feature vector, the preset semantic slot filling status of the original activity semantic fingerprint is mapped and compared, and the coverage index of each mandatory constraint in the preset semantic slot is calculated to generate a semantic missing difference matrix including uncovered constraint terms.

[0120] Based on the factor constraint feature vectors parsed from the dynamic adaptation factor subset, the preset semantic slot filling state matrix of the original activity semantic fingerprint is read as the comparison object.

[0121] Each mandatory constraint in the factor constraint feature vector is mapped to the corresponding semantic slot position index according to the constraint category, generating a slot mapping table.

[0122] According to the slot mapping table, perform condition coverage detection on the preset semantic slot filling state matrix, directly extract the encoded value of the filled field, and mark the missing field with logical null value and record the missing index.

[0123] Arrange the coverage indices of all mandatory constraints into an index vector, and perform differential marking on the constraints in the vector that are below the preset coverage threshold.

[0124] A semantic missing difference matrix is ​​constructed based on the difference tagging results. Each row corresponds to an uncovered constraint item, and each column corresponds to a slot index. The matrix cells are filled with missing tags or matching status values.

[0125] By calculating coverage and generating a difference matrix, the feature vectors of factor constraints are mapped and compared with the slot filling status of the original activity semantic fingerprint, clarifying the distribution of uncovered constraint terms and providing a quantitative basis for subsequent missing slot location.

[0126] S5.3: Utilize the semantic missing difference matrix to perform causal back-inference analysis on the uncovered constraint terms, and combine the triggering logic of the associated edges in the emission factor evolution graph to deduce the key semantic elements corresponding to the implicit deducible constraints, so as to locate the missing key semantic slot identifiers that lead to the decrease in matching confidence.

[0127] The set of uncovered constraint terms in the semantic missing difference matrix is ​​received as the input condition for causal back-inference analysis, and the applicable condition edge attributes of each node in the dynamic adaptation factor subset are bound to form a searchable triggerable logical index.

[0128] A causal path backtracking method is performed on the set of uncovered constraint terms. Based on the directed edge triggering conditions stored in the emission factor evolution graph, the set of potential triggering source nodes for each uncovered constraint term in the graph is established, and its associated spatiotemporal and process dimension labels are extracted to establish a triggering causal chain.

[0129] The implicit constraint matching operation is performed on the extracted set of trigger source nodes. The intersection calculation of the physical working condition boundary of the known dynamic adaptation factor subset and the triggering condition of the trigger source node is performed to screen out the candidate set of key semantic elements corresponding to the derivable constraints and remove elements that are inconsistent with the actual activity fingerprint boundary.

[0130] Extract the highest-scoring semantic elements from the confidence score ranking, generate a list of missing key semantic slot identifiers, and perform consistency verification with the semantic missing difference matrix to ensure the accuracy of the localization results.

[0131] By using causal back-inference analysis and trigger logic association matching, the uncovered constraint terms in the semantic missing difference matrix are transformed into verified missing key semantic slot identifiers, thereby achieving the expected technical effect of improving the confidence of factor matching.

[0132] For example, in the carbon emission factor matching process for a coal-fired boiler in a steel plant, the semantic missing difference matrix shows that the "denitrification process identifier" is not covered; the triggering logic of the dynamic adaptation factor subset node B is that a specific factor value is applied when "SCR commissioning status = yes"; the backtracking method locates the triggering source node of this constraint as equipment model M100 and East China power grid granularity conditions, and extracts the process dimension element "SCR commissioning status". The intersection calculation of the operating condition boundaries shows that the rated load rate of the boiler is in the high load range, which meets the triggering condition of node B.

[0133] S5.4: Based on the missing key semantic slot identifier, retrieve the preset data collection template library, match the data field definition and collection specification description corresponding to the missing key semantic slot identifier, and generate a standardized data completion requirement description text.

[0134] The system receives the set of missing key semantic slot identifiers generated by the preceding sub-step S5.3, confirming that each entry in the identifier set is the culprit for the reduced matching confidence identified by the causal back-analysis and trigger logic derivation in the previous step. The missing key semantic slot identifier is input as a search condition into the index retrieval module of the pre-set data acquisition template library. Template entries that are completely identical to the identifier or have a synonymous substitution relationship are retrieved, ensuring that the search covers synonyms, abbreviations, and cross-industry terminology mappings. Field definition mapping parsing is performed on the retrieved template entries to extract the standard data field composition and field type constraints under the template, forming a set of field definition prototypes. A one-to-one mapping relationship is established between this set and the missing slot identifiers. Based on the set of field definition prototypes, the system reads the acquisition specification descriptions in the template entries, including acquisition process steps, parameter acquisition accuracy requirements, applicable equipment category constraints, and spatiotemporal data synchronization strategies. The acquisition specifications are broken down into quantifiable indicator description units to achieve structured preparation for subsequent text splicing. The field definition prototype set and the quantitative indicator description unit are processed in the text generation engine in a preset order to perform operations such as field description splicing, data collection specification embedding, and logical paragraph alignment, generating standardized data completion requirement description text that strictly follows industry data collection standards and has embedded slot identifiers.

[0135] By using field definition mapping parsing and embedded processing of the collection specifications, the missing slot identifiers from the previous step are transformed into structured text that can directly guide the execution of the data collection end, achieving the expected technical effect of converting the abstract identifiers in the semantic missing difference matrix into operable completion instructions.

[0136] For example, in a carbon emission accounting scenario for a steel company, the missing key semantic slot identifier is "electric arc furnace commissioning status". When searching the data collection template library, the template entry "Smelting Equipment Operation Status Collection Table" is located. This template entry's field definition prototype set includes "Equipment ID", "Process Type", "Commissioning Status", "Commissioning Start Time", and "Commissioning End Time", with field type constraints of string, enumeration value, boolean value, timestamp, and timestamp, respectively. The collection specification requires that the commissioning status be represented by a boolean value, with a collection accuracy of seconds, applicable equipment category of electric arc furnace, and that timestamp data be generated synchronously from the plant's SCADA system's operation logs. The above information is quantified into indicator description units, including the commissioning status value range {0,1}, timestamp format YYYY-MM-DD hh:mm:ss, and SCADA system log synchronization period of 3600 seconds. By using a text generation engine to concatenate field definitions and indicator descriptions, the generated standardized data completion requirement description text explicitly lists: "Field [Operation Status]: Boolean, value 0 indicates shutdown, 1 indicates operation; Collection granularity: second-level; Synchronization source: plant SCADA operation log; Period: 3600 seconds; Field [Operation Start Time] / [Operation End Time] format YYYY-MM-DD hh:mm:ss". This description text is invoked by the data acquisition end during actual execution, forming instructions to complete the operation status, directly improving the accuracy of factor matching and the completeness of the calculation.

[0137] S5.5: Encapsulate the data completion requirement description text into a structured data completion guidance instruction and bind it to a unique session identifier of the current activity semantic fingerprint to output an interactive guidance signal that can trigger the front-end data collection interface or the manual input interface.

[0138] Step S6: For the same activity data, generate macro-level activity semantic fingerprints including enterprise-level and micro-level activity semantic fingerprints including equipment-level, and map them together to different level nodes of the emission factor evolution map to generate hierarchical projection matching results, thereby completing the carbon emission data accounting. Specifically, this includes: S6.1: Perform dimensional attribute parsing on the original multi-source heterogeneous activity data, extract the enterprise registration code, administrative division code and annual statistical cycle as macro dimension identifiers, and discretize the macro dimension identifiers based on the preset macro semantic slot template to generate macro activity semantic fingerprints including enterprise dimensions.

[0139] Using raw, multi-source, heterogeneous activity data as input, and relying on the semantic completion results from previous steps, data dimension attribute parsing is performed to lock in the required macro-level identifier variables. During parsing, the enterprise registration information parsing module is invoked to perform format verification and legality validation of the enterprise registration code in the original message fields, removing illegal characters and unifying the encoding format to ensure the stability of unique identifiers. Based on administrative division identification rules, administrative division coding mapping is performed on the regional description fields included in the data. A multi-level coding system is used to achieve hierarchical identification at the provincial, municipal, and county levels, and the consistency of regional codes is ensured through cross-comparison of the mapping results with the standard division database. Using statistical cycle identification methods, the annual statistical cycle is extracted from timestamps or annual description information. A cycle formatting function is used to convert yearly and quarterly information into a unified annual granularity value, providing a stable time index for the macro-dimensional time attributes. After obtaining the three macro-dimensional identifiers—enterprise registration code, administrative division code, and annual statistical cycle—a preset macro-semantic slot template is loaded, and the above identifiers are matched and deployed according to the slot order to fill the corresponding coding fields. A discretization encoding method is employed, mapping each macroscopic dimension identifier to a fixed-width integer encoding stack. A macroscopic semantic encoding vector is generated based on the correspondence between the encoding bit width and the slot template. After the encoding vector is constructed, bit alignment and padding are performed to ensure that the length of the encoded sequence strictly conforms to the specified length of the macroscopic activity semantic fingerprint, avoiding inconsistencies in bit length due to missing macroscopic dimension values. After generating the fixed-length encoding string, a hash compression function is called to perform feature dimensionality reduction and uniqueness verification, outputting a macroscopic activity semantic fingerprint that includes the enterprise dimension.

[0140] Through the chain-like processing described above, the multi-source activity data completed in the previous step is transformed into a stable, uniformly long, and uniquely identifiable enterprise-level macro-activity semantic fingerprint, enabling efficient retrieval and matching of macro-level activity data in the emission factor evolution map.

[0141] For example, in the carbon emission accounting scenario of a steel production enterprise, the raw activity data collected through an interface includes three fields: the enterprise registration code "91350200M00012345J", the geographical description "Quanzhou City, Fujian Province", and the time description "2023". In the enterprise registration information parsing module, the registration code is encoded as the integer value 102334 after illegal character removal and format verification. The administrative division identification module maps "Quanzhou City, Fujian Province" to the provincial code 350000 and the municipal code 350500, and combines them into the administrative division code 350500. The statistical period identification method formats the time description "2023" into the annual value 2023. The macro-semantic slot template defines the enterprise registration code slot width as 20 bits, the administrative division code slot width as 12 bits, and the year slot width as 8 bits. A discretization encoding method maps 102334, 350500, and 2023 to integer codes of corresponding widths, forming a macro-semantic encoding vector [00100110001010101110,0000110101010100,0001111111]. Bit alignment and padding ensure the vector length conforms to the preset 40-bit specification. A hash compression function compresses it to generate a fixed-length hash string "A39FBC12E8D7". The final output macro-activity semantic fingerprint can directly activate the corresponding enterprise-level coarse-grained nodes in the evolutionary graph retrieval, achieving rapid data matching and location at the macro level.

[0142] S6.2: Perform device-level feature mining on the original multi-source heterogeneous activity data, extract the device's unique identification code, real-time operating condition parameters, and hourly timestamp as micro-dimensional identifiers, and discretize and encode the micro-dimensional identifiers based on the preset micro-semantic slot template to generate a micro-activity semantic fingerprint including the device dimension.

[0143] The device-level data segment that receives raw multi-source heterogeneous activity data is used as the input object, including the device's unique identification code transmitted by the field acquisition system, the continuously sampled sequence of operating condition parameters, and the hourly timestamp signal corresponding to the operating status.

[0144] The device's unique identification code is verified for uniqueness and its format is parsed. Illegal codes are eliminated according to the preset hardware identification rules, and legal codes are mapped to standardized identification sequences to form the basic field of "device identifier" in the micro-semantic slot.

[0145] Multidimensional index decomposition is performed on the operating condition parameter sequence to decompose the original numerical signal into sub-indicators reflecting load level, energy consumption rate, temperature control deviation, etc. Based on the operating condition physical model, key parameter values ​​within the stable operating condition range are extracted and filled into the "operating status" field in the micro-semantic slot.

[0146] Perform timing accuracy verification and format normalization processing on hourly timestamp signals, unify the time formats from different sources into Coordinated Universal Time (UTC) millisecond precision, and map them to the "time context" field in the micro-semantic slot.

[0147] The three fields are encoded sequentially according to the preset micro-semantic slot template. The device identifier, operating status and time context are converted into fixed-width integer codes using a discretization mapping method. Continuous indicators in the operating status are divided into intervals and encoded using interval ID.

[0148] The encoded device identifier, operating status, and time context are bit-concatenated according to the template definition order to form a micro-activity encoding vector. Padding and alignment processes are performed to ensure that the vector length meets the specification length of the micro-activity semantic fingerprint.

[0149] By using a fixed discretization encoding method for micro-semantic slot templates, the original equipment-level data is transformed into a micro-activity semantic fingerprint that is structurally complete, of fixed length, and can uniquely identify the activity characteristics of the equipment, thereby realizing the basic data for factor matching at the equipment level.

[0150] For example, an industrial boiler equipment data acquisition system provides a unique equipment identification code "BLR-A1F38X", and a sequence of operating parameters including rated power of 500kW, real-time load rate of 0.68, fuel input mass of 32.5kg / h, and flue gas temperature of 180℃. The hourly timestamp format is local time 2024-05-15 14:00. The system maps the identification code to the code segment "000145" through rule matching, maps the load rate to the code "05" through interval division rules (interval ID of "05" for the 0.6-0.7 interval), and combines the fuel input mass and flue gas temperature to obtain the operating status code "1782". The timestamp is converted to UTC millisecond precision and mapped to the code segment "1623457800000". The encoded segments are concatenated according to the template order to form a micro-activity encoded vector [000145, 1782, 05, 1623457800000]. Alignment and padding processing is performed to output a fixed-length micro-activity semantic fingerprint "0001451782051623457800000". This fingerprint significantly improves the accuracy and accessibility of device-level factor localization during the map matching process, achieving a fine mapping effect of operating parameters.

[0151] S6.3: Based on the emission factor evolution map constructed in the previous steps, identify the enterprise-level coarse-grained node set and the equipment-level fine-grained node set defined in the map, and input the macro-activity semantic fingerprint into the top-level index structure of the map, and input the micro-activity semantic fingerprint into the bottom-level index structure of the map, to establish a parallel mapping channel between the two-layer fingerprint and the multi-level nodes.

[0152] With the macro-activity semantic fingerprints and micro-activity semantic fingerprints generated and possessing a standardized discretized encoding format, the emission factor evolution map constructed in the previous steps is used as the matching object. It receives macro-fingerprints including enterprise-level identifiers and micro-fingerprints including equipment-level identifiers as parallel input data. During execution, the map structure parsing module is first invoked to perform dimensional segmentation on the internal index structure of the evolution map, identifying the top-level node set defined for enterprise-level coarse-grained matching and simultaneously identifying the bottom-level node set defined for equipment-level fine-grained matching. Subsequently, the coding slots of the macro-activity semantic fingerprint are precisely aligned with the index keys of the top-level node set. During the alignment process, a hash index matching mechanism is used to compare the enterprise registration code, administrative division code, and annual statistical cycle fields to achieve a deterministic mapping from fingerprint slots to the top-level index keys. The mapping results are then written to the top-level index query queue. Next, the coding slots of the micro-activity semantic fingerprint are precisely aligned with the index keys of the bottom-level node set. During the alignment process, the device's unique identification code, real-time operating parameters, and hourly timestamp fields are combined, and a multi-dimensional attribute cross-matching mechanism is used to map the fingerprint slots to the bottom-level index keys. The mapping results are then written to the bottom-level index query queue. Next, a parallel channel is established between the top-level and bottom-level index query queues. A dual-queue synchronous scheduling strategy ensures that macro-fingerprint and micro-fingerprint mapping requests simultaneously enter the corresponding hierarchical index structure for node retrieval. During the establishment of the parallel channel, to avoid cross-layer node interference, a hierarchical label-based instruction isolation mechanism is set up. During execution, the retrieval results of each channel are limited to the corresponding hierarchical node set. By using parallel mapping of dual-layer fingerprints and multi-level nodes, macro-level category matching and micro-level working condition matching are ensured to be performed simultaneously, and the output includes a set of parallel mapping results including hierarchical distinguishing identifiers, thus achieving cross-dimensional and cross-granular synchronous matching capabilities.

[0153] For example, in the carbon management platform of an energy company, the system receives the semantic fingerprint of enterprise activities after macro-template encoding. The enterprise registration code field value is "9133XXXX", the administrative division code is "320100", the annual statistical cycle code is "2022", and the hash index uses a 40-bit fixed-length code, which is mapped to the node set number E-Top-102 in the top-level index structure of the graph. At the same time, the system receives the semantic fingerprint of equipment activities after micro-template encoding. The equipment unique identification code field value is "EQP-00982", the real-time operating condition parameter code is standardized to "Load-75", and the hourly timestamp code is "2022-08-13-15". Using a multi-dimensional attribute cross-matching mechanism, the system is mapped to the node set number E-Btm-457 in the bottom-level index structure of the graph. The top-level queue contains the macro-fingerprint index key {E-Top-102}, and the bottom-level queue contains the micro-fingerprint index key {E-Btm-457}. A dual-queue synchronous scheduling strategy is established, and hierarchical labels are set for isolation. This ensures that during the matching process, the macro-queue only searches for enterprise-level coarse-grained nodes, and the micro-queue only searches for device-level fine-grained nodes. After parallel mapping is completed, a result set is output. The macro-matching result set is identified as "MacroMatchID=MC102", and the micro-matching result set is identified as "MicroMatchID=MI457". Both have hierarchical differentiation labels, which are used to perform macro-similarity matching and micro-constraint verification respectively in step S6.4, significantly improving the synchronous matching accuracy of cross-industry multi-source data.

[0154] S6.4: Using the parallel mapping channel, the macro-similarity matching method is executed at the top layer of the graph to activate the candidate enterprise-level factor node group, and the micro-constraint verification method is executed at the bottom layer of the graph to activate the candidate device-level factor node group, generating a hierarchical projection matching result including the macro-matching result set and the micro-matching result set.

[0155] Under the established dual-layer fingerprint and multi-level node parallel mapping channel input conditions, enterprise-level coarse-grained matching operations are performed on the top-level mapping of macro-activity semantic fingerprints. Using the macro-dimensional encoding field as the search key, a macro-similarity matching method is invoked in the top-level node index library to calculate a macro-matching score vector. This score vector is weighted and accumulated based on the coverage and weight coefficients of the macro-semantic slots to form a preliminary enterprise-level factor node activation set. For the bottom-level mapping of micro-activity semantic fingerprints, device-level fine-grained constraint verification is performed. Using the micro-dimensional encoding field as the search key, a micro-constraint verification method is triggered in the bottom-level node index library. The implicit constraint quantification indicators derived from the device operating condition parameters are used to perform Boolean filtering on the applicable conditions of the bottom-level nodes, generating a preliminary device-level factor node activation set that has passed physical condition consistency verification.

[0156] A confidence ranking operation is performed by combining the macro-level matching score vector with the enterprise-level factor node activation set. The ranking result is stored as the index structure of the macro-level matching result set to support subsequent association comparison calls. A dynamic weight adjustment operation is performed by combining the micro-level constraint verification result with the device-level factor node activation set. The adjusted node list is stored as the index structure of the micro-level matching result set to ensure the accuracy of fine-grained matching. Through the structured encapsulation of the macro-level and micro-level matching result sets, a hierarchical projection matching result covering different levels of matching paths is formed, realizing parallel activation and result complementarity of enterprise-level and device-level factors under cross-granularity conditions.

[0157] S6.5: Perform an association consistency comparison between the macro-matching result set and the micro-matching result set in the hierarchical projection matching result, mark the semantic conflict intervals or data complementarity intervals between the two layers of nodes, and output the final hierarchical projection matching result with conflict and complementarity markers for subsequent weighted fusion processing by the business rule module.

[0158] Step S7: The hierarchical projection matching results are weighted and fused according to preset business rules. If the business rules determine it to be a regulatory reporting scenario, macro-level node data is prioritized; if the business rules determine it to be a technological upgrading assessment scenario, micro-level node data is prioritized, ultimately obtaining the carbon emission accounting factor. Specifically, this includes: The preset business rules include: 1. Scene type determination rules Rule 1.1: Parse the business context metadata in the hierarchical projection matching results and extract three core fields: "scenario purpose", "data destination" and "business subject"; Rule 1.2: If "Scenario Purpose = Regulatory Reporting" and "Data Destination = Environmental Protection Department / Carbon Exchange", then it is marked as a regulatory reporting scenario, and the identifier "SCENE_SUPERVISION" is generated; Rule 1.3: If "Scenario Purpose = Technological Upgrade Assessment" and "Business Entity = Enterprise Technological Upgrade Department", then it is marked as a technological upgrade assessment scenario, and the identifier "SCENE_TECH" is generated; Rule 1.4: If the above conditions are not met, the scenario will be judged by default according to the regulatory reporting scenario to ensure that no scenario is missed.

[0159] 2. Weighting rules Rule 2.1: The business rule knowledge base stores weighted strategies according to scenario identifiers, with independent rule entries corresponding to regulatory reporting scenarios and technical transformation assessment scenarios; Rule 2.2: Regulatory reporting scenario: Macro-level priority coefficient = 0.7, Micro-level priority coefficient = 0.3 (macro-level data is preferred); Rule 2.3: Technical upgrade assessment scenario: Macro-level priority coefficient = 0.3, Micro-level priority coefficient = 0.7 (micro-level node data is preferred); Rule 2.4: Weight normalization constraint: Macro coefficient + micro coefficient = 1.0, eliminating redundant weight configurations that do not match the current scenario; Rule 2.5: In special cases (such as a single level with no data), the weight is automatically adjusted to the level with data coefficient = 1.0 to ensure that fusion can be executed.

[0160] 3. Weighted fusion and interval processing rules Rule 3.1: Factor value alignment rule: Macro and micro factors must be aligned with the same business semantic slot (such as the same device or the same time period). Unaligned factors will not participate in the fusion for the time being. Rule 3.2: Arbitration rule for conflict intervals: For factor pairs marked as conflicting, the one with a confidence level ≥ 0.8 shall be selected first; if the confidence levels of both factors are below 0.8, they shall be reconciled according to the current scenario weight (e.g., for regulatory scenarios, macro × 0.7 + micro × 0.3). Rule 3.3: Complementary Interval Overlay Rule: Factor pairs marked as complementary are overlaid according to the formula "macroeconomic factor + microeconomic increment" (the increment is the absolute value of the difference between the microeconomic and macroeconomic factors), or by weighted overlay. Rule 3.4: The fused factor value must be within a preset reasonable range (e.g., carbon emission factor 0.1-10.0 tCO2 / e). Values ​​outside the range are temporarily stored as candidate outliers.

[0161] 4. Confidence check and anomaly removal rules Rule 4.1: Confidence Calculation Rule: Fusion Factor Confidence = (Macro Factor Confidence × Macro Weight) + (Micro Factor Confidence × Micro Weight); Rule 4.2: Threshold determination rule: The confidence threshold is set to 0.6. Candidate values ​​below this threshold are marked as abnormal fluctuation data and are removed. Rule 4.3: Anomaly Association Rule: Based on the semantic consistency comparison records, if an anomaly value corresponds to a conflict marker and there is no reasonable matching path, it is directly removed; Rule 4.4: Smoothing rule: After removing outliers, perform moving average smoothing on the remaining factor values ​​to reduce discrete gradients and ensure data stability.

[0162] 5. Encapsulation and Serialization Rules Rule 5.1: Source Traceability Rule: The encapsulated data packet must include the source node identifier, matching path, and confidence score of macro / micro factors to ensure traceability; Rule 5.2: Scenario Labeling Rule: Add labels according to scenario type (regulatory reporting / technical improvement assessment), and indicate the scope of application (e.g., "applicable to regulatory reporting in 2024"). Rule 5.3: Formatting rules: Serialize according to carbon report / trading settlement standards, using JSON format, with fields including "accounting factor value, scenario label, source information, and generation timestamp"; Rule 5.4: Validation rules: After serialization, the format integrity must be verified. If key fields are missing (such as source traceability), the data must be re-encapsulated.

[0163] S7.1: Obtain the final hierarchical projection matching result with conflict and complementary markers, and use the scene feature extraction method to parse the business context metadata bound therein to generate a scene type determination vector including regulatory reporting identifier or technical transformation assessment identifier.

[0164] S7.2: Based on the scenario type determination vector, retrieve the pre-set business rule knowledge base, call the corresponding weight allocation strategy model, and generate a dynamic weight configuration parameter set including macro-level priority coefficients and micro-level priority coefficients.

[0165] Based on the generated scenario type determination vector, the retrieval module performs a multi-index matching operation in the pre-built business rule knowledge base to locate the set of rule entries that completely correspond to the scenario identifier. For the retrieved set of rule entries, the macro-level and micro-level weight configuration information is parsed and transformed into a computable initial priority coefficient vector. Using the constraints in the rule entries and the business context parameters in the scenario type determination vector, a conditional filtering operation is performed to eliminate redundant configuration records that do not meet the characteristics of the current scenario, forming an effective weight subset for the current business context. The macro-level and micro-level priority coefficients in the effective weight subset are normalized to ensure that their sum is a constant and that the proportional relationship is maintained.

[0166] S7.3: Using the dynamic weight configuration parameter set, perform weighted fusion calculation on the numerical factor data in the macro matching result set and the micro matching result set, execute the conflict interval arbitration logic and the complementary interval superposition logic to generate preliminary fused carbon emission accounting factor candidate values.

[0167] The dynamic weight configuration parameter set is received, and the numerical weights corresponding to the macro-level priority coefficient and the micro-level priority coefficient are determined as the proportional parameters for subsequent numerical factor fusion.

[0168] The numerical factor fields in the macro and micro matching result sets are parsed to form a factor value matrix that can be directly used in the calculation. The matrix row and column structures are distinguished according to the hierarchical identifier to ensure that each factor is aligned with the same business semantic slot.

[0169] For factor value pairs marked as conflict intervals in the hierarchical projection matching results, arbitration logic is executed to select the party with higher confidence or to reconcile the conflict values ​​according to the business rule coefficient, thereby generating a set of fused factor values ​​after conflict resolution.

[0170] For factor value pairs marked as complementary intervals, superposition logic is executed to arithmetically superimpose or weighted increment the factor values ​​at the macro and micro levels in that slot, generating a fused factor value set after complementary processing.

[0171] The result sets after conflict resolution and complementation are merged to form a structurally complete fusion factor matrix that has eliminated contradictions, which serves as the preliminary candidate values ​​for carbon emission accounting factors.

[0172] By using the weighted fusion and interval processing methods described above, the hierarchical projection matching results from the previous step are transformed into candidate factor data covering both conflict arbitration and complementary overlay paths, achieving the expected technical effect of multi-source factor fusion that takes into account business rules.

[0173] S7.4: Perform confidence verification on the preliminary fused candidate values ​​of carbon emission accounting factors, and eliminate abnormal fluctuation data by combining the semantic consistency comparison records in the hierarchical projection matching results, so as to generate a standard carbon emission accounting factor that has been purified in quality.

[0174] Using the preliminary fusion of candidate carbon emission accounting factors as input, the semantic consistency comparison records in the hierarchical projection matching results are loaded to establish a confidence verification association index.

[0175] Locate the semantic association group of the macro-matching result data and micro-matching result data corresponding to the candidate values ​​in the consistency comparison record, and extract the matching path and conflict marker in the group as the verification benchmark.

[0176] The calculated confidence score is compared with a preset confidence threshold, and candidate values ​​that are below the threshold are marked as abnormal fluctuation data.

[0177] For data marked as abnormal fluctuations, anomaly removal processing is performed by combining the conflict marker positions in its semantic consistency comparison records, including removing the candidate value and its corresponding matching path reference.

[0178] For the remaining candidate values ​​that were not eliminated, the carbon emission accounting factor sequence was regenerated and the sequence was smoothed to reduce the discrete gradient between factor values.

[0179] By using the confidence verification and anomaly removal methods described above, the fusion candidate values ​​from the previous step are transformed into standard carbon emission accounting factors that have undergone quality purification, thereby enhancing the stability of the factor data and improving the accuracy of the accounting.

[0180] S7.5: Based on the standard carbon emission accounting factor, encapsulate a structured data packet including source traceability information and scenario applicability labels, perform a final format serialization operation to output a final carbon emission accounting factor that can be directly used for carbon report generation or transaction settlement.

[0181] Step S8: Based on the final carbon emission accounting factor and data completion guidance instructions, update the emission factor evolution map to generate an iteratively optimized emission factor evolution map. Specifically, this includes: S8.1: Obtain the manual review and correction signal of the final carbon emission accounting factor, extract the original activity semantic fingerprint, the initial matching factor node identifier before correction, and the target adaptation factor node identifier confirmed by manual verification, and construct a review and correction dataset including the difference feature vector to clarify the specific spectral path range that needs to be adjusted.

[0182] S8.2: Based on the difference feature vector in the verification and correction dataset, locate the directed temporal network edge connecting the initial matching factor node and the target adapting factor node in the emission factor evolution graph, extract the historical weight value and original trigger threshold parameter currently stored in the directed temporal network edge, generate the graph edge state information to be updated, and determine the specific parameter adjustment object.

[0183] After receiving the revised dataset including the difference feature vector, the difference feature vector is input into the emission factor evolution map node index retrieval module to extract the unique node index corresponding to the initial matching factor node identifier and the target fitting factor node identifier.

[0184] Based on the unique node index, a path lookup operation is performed in the topology table of a directed temporal network to locate all direct and indirect directed edges between two nodes.

[0185] Based on the query results, the edge attribute storage module is called to read the current historical weight value and original trigger threshold parameter of each edge, forming a temporary data pair set including two types of static attributes.

[0186] The temporary data set is encapsulated in a structured manner for each edge, and the historical weight values ​​and the original trigger threshold parameters are stored in the graph edge state data template to be updated in a field-based manner.

[0187] The encapsulated graph edge state data template to be updated is bound to the differential feature vector record to clarify that the edge is the parameter adjustment object and can be used for subsequent iterative optimization.

[0188] By using the above edge localization and attribute extraction processing methods, the difference feature vector results from the previous step are transformed into edge state information data, thereby locking the target parameters for subsequent weight and threshold optimization.

[0189] For example, in the verification and correction dataset including the difference feature vector, the initial matching factor node is identified as N105, the target adaptation factor node is identified as N210, and the difference feature vector includes the device type code D45 and the timestamp 2024-03-15 08:00. The system obtains the unique indices IDX105 and IDX210 of N105 and N210 respectively through the node index retrieval module. The path query results show that there is a direct directed edge E105-210 between the two nodes and an indirect edge E105-150-210 via N150. For the direct edge E105-210, the edge attribute storage module returns the historical weight value of 0.75 and the original trigger threshold parameter. For indirect edge E105-150, return the historical weight value of 0.62 and the original trigger threshold parameter. The system encapsulates these attributes into the edge state data template of the graph to be updated, and attaches the device type code and timestamp recorded in the difference feature vector, thus identifying these two edges as the objects of parameter adjustment. Through subsequent optimization, the weight values ​​and trigger thresholds of these edges will be iteratively adjusted based on manually corrected confidence levels and the magnitude of changes in implicit constraints, to achieve adaptive fine-tuning of the graph path.

[0190] S8.3: The historical weight values ​​in the edge state information of the graph to be updated are iteratively calculated using the preset incremental gradient descent method. New edge connection weight values ​​are generated by combining the manually corrected confidence coefficient. At the same time, the original trigger threshold parameter is dynamically adjusted according to the change range of implicit constraints in the activity semantic fingerprint to generate an optimized set of edge attribute parameters, so as to achieve adaptive fine-tuning of the graph topology.

[0191] The system receives historical weight values ​​and original trigger threshold parameters from the edge state information of the graph to be updated, and combines them with manually adjusted confidence coefficients as an influence factor to define the objective function for iterative weight updates.

[0192] The weight update results are combined with the threshold adjustment results to form an optimized set of edge attribute parameters, and it is ensured that the values ​​of each item in the parameter set meet the predetermined physical constraints and topological rules.

[0193] By using a combined approach of incremental gradient descent and dynamic threshold adjustment, the edge state information of the graph to be updated is transformed into an optimized set of edge attribute parameters, including new weight values ​​and new trigger thresholds, thereby enabling adaptive fine-tuning of the local topological relationships in the emission factor evolution graph.

[0194] S8.4: Write the optimized set of edge attribute parameters into the corresponding directed time-series network edge nodes in the emission factor evolution graph, replace the original historical weight values ​​and original trigger threshold parameters, and record the timestamp and version number of this update to generate a local update graph intermediate state with version traceability capability, so as to ensure the real-time performance and traceability of the graph data.

[0195] S8.5: Execute full update based on the intermediate state of the local update graph with version traceability capability. Figure 1 Consistency verification checks whether the updated edge connection weights and the new trigger threshold parameters satisfy the graph topology constraints. If the verification passes, the iteratively optimized emission factor evolution graph is output, completing a single lightweight online learning loop and ensuring the accuracy of subsequent matching processes.

[0196] Based on the intermediate state of the locally updated graph with version traceability capabilities, the full graph topology and node attribute indexes are loaded. Optimized edge connection weights and new trigger threshold parameters are used as verification inputs. The topology constraint rule set is invoked to perform parameter consistency verification on each edge of the directed temporal network, ensuring that the updated weights do not disrupt the priority relationships between nodes and that the trigger threshold parameters do not cause boundary condition conflicts. For node combinations with multiple paths, a global path reachability test is performed. The edge weights of each path are accumulated using depth-first traversal, and the accumulated results are compared with the maximum reachable weight limit in the topology constraints to ensure no exceedances occur. A temporal consistency detection method is used, comparing the updated trigger threshold with the effective period based on the node timestamp sequence in the network, eliminating illegal edges whose trigger time is earlier than the time the preceding condition is met. Physical constraint verification is introduced, matching the updated edge attributes with equipment specifications and process condition sets to ensure that paths violating physical laws are not generated. A set of qualification markers is generated based on the consistency verification results. If all verifications pass, the markers are written into the iteratively optimized emission factor evolution map, completing a single lightweight online learning loop to ensure the accuracy of subsequent matching processes and network stability. This is achieved through a full... Figure 1 The consistency verification process transforms the edge parameter update results from the previous step into verified topological structure data, thereby ensuring both the availability and matching accuracy of the updated emission factor evolution map.

[0197] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.

[0198] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.

[0199] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. Carbon emission data accounting methods based on activity data tracing, specifically including: S1: Acquire multi-source heterogeneous activity data composed of carbon emission data, extract the core semantic elements from the multi-source heterogeneous activity data, and discretize and encode the core semantic elements according to preset semantic slots to generate a fixed-length activity semantic fingerprint. S2: Based on historical carbon emission data, construct a directed time series network, and generate an emission factor evolution map based on the associated data in each node of the directed time series network; S3: Input the activity semantic fingerprint into the fingerprint graph bidirectional mapping engine, locate the set of initial nodes that can be activated in the emission factor evolution graph based on the spatiotemporal context and object attributes in the fingerprint, and generate a candidate factor node group; S4: Perform logical verification based on the candidate factor node group, remove unreachable paths that do not meet the implicit constraints, and generate a dynamic adaptation factor subset; S5: Using the dynamic adaptation factor subset to output factor adaptation feedback in reverse, identify the missing key semantic slots in the active semantic fingerprint and generate data completion guidance instructions. S6: Generate macro-level activity semantic fingerprints (including enterprise-level) and micro-level activity semantic fingerprints (including equipment-level) for the same activity data, and map them together to different level nodes of the emission factor evolution map to generate hierarchical projection matching results, thereby completing the carbon emission data accounting.

2. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, Step S6 is followed by: S7: The hierarchical projection matching results are weighted and fused according to the preset business rules. If the business rules determine that it is a regulatory reporting scenario, macro-level node data is adopted first. If the business rules determine that it is a technical transformation assessment scenario, micro-level node data is adopted first. Finally, the carbon emission accounting factor is obtained. S8: Based on the final carbon emission accounting factor and data completion guidance instructions, update the emission factor evolution map to generate an iteratively optimized emission factor evolution map.

3. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, The core semantic elements include the type of the activity subject, action verbs, object attributes, spatiotemporal context, and implicit constraints.

4. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, The associated data includes factor values, confidence intervals, effective periods, and a list of compatible devices.

5. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, Step S3 specifically includes: The spatiotemporal context encoding field and object attribute encoding field in the activity semantic fingerprint are parsed to extract key-value index pairs for graph retrieval and generate graph query vectors. Based on the graph query vector, a multidimensional matching operation is performed in the emission factor evolution graph to filter out the initial matching nodes that meet the geographical granularity constraints and equipment compatibility list requirements, and generate a primary set of activated nodes. The primary set of activated nodes is subjected to a time-series consistency comparison to remove invalid nodes that are inactive or have expired, thereby generating a subset of time-series valid activated nodes. Based on the directed edge connections of the time-series effective activated node subset in the emission factor evolution graph, a neighborhood traversal operation is performed to obtain the associated derived nodes triggered by changes in applicable conditions, and an extended associated node list is generated. The primary activation node subset and the extended associated node list are deduplicated, merged, and sorted by confidence to construct a final result set covering direct matching and indirect derivation paths, generating a candidate factor node group.

6. The carbon emission data accounting method based on activity data tracing according to claim 5, characterized in that, The vector information of the map query vector includes the power grid zone identifier, altitude range parameters, and equipment model code.

7. The carbon emission data accounting method based on activity data tracing according to claim 3, characterized in that, Step S4 specifically includes: Obtain node metadata and object attributes from the activity semantic fingerprint in the candidate factor node group, use the inference big model to extract the nonlinear mapping relationship between the equipment rated parameters and the operating load, and generate a set of implicit constraint derivation rules. Based on the implicit constraint derivation rule set, feature inversion calculation is performed on the spatiotemporal context in the activity semantic fingerprint, transforming the discrete fingerprint encoding into a continuous device operating state interval, and generating a derivable constraint quantification index. Boolean logic is used to verify the applicable condition edges of each node in the candidate factor node group by using derivable constraint quantification indicators, and the effective period of node annotation is compared with the device compatibility list to generate path reachability determination results. Based on the path reachability determination results, topology pruning is performed on the connection edges in the directed temporal network to remove nodes marked as unreachable and their associated successor paths, generating a dynamic adaptation factor subset.

8. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, The extraction of core semantic elements from the multi-source heterogeneous activity data includes automatically identifying the message structure and encoding method of the multi-source heterogeneous activity data using the data source interface protocol, extracting the activity subject type, action verbs, object attributes, spatiotemporal context and implicit constraints in segments, completing consistency verification, and generating core semantic elements.

9. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, The step of locating the set of initial nodes that can be activated in the emission factor evolution graph based on the spatiotemporal context and object attributes in the fingerprint includes determining the temporal consistency of the emission factor evolution graph based on the spatiotemporal context and object attributes in the fingerprint, retaining only the nodes within the effective period corresponding to the current business timestamp, and expanding to all derived associated nodes that meet the conditions through breadth-first traversal and trigger condition determination. The set of all derived associated nodes is the set of initial nodes that can be activated.

10. The carbon emission data accounting method based on activity data tracing according to claim 1, characterized in that, The term includes macro-level activity semantic fingerprints at the enterprise level and micro-level activity semantic fingerprints at the device level. The macro-level activity semantic fingerprints are generated by discrete encoding of macro-level dimensions such as enterprise registration code, administrative division code, and statistical period, while the micro-level activity semantic fingerprints are generated by discrete encoding of device identification code, real-time operating parameters, and hourly timestamp.