Carbon emission measurement and calculation data management system and method based on multi-source data
By dynamically adjusting carbon emission factors through knowledge graphs and reinforcement learning, combined with blockchain notarization and multi-stage feedback, the problems of insufficient integration of multi-source data and low credibility of cross-domain circulation are solved. This improves the accuracy of carbon emission calculation and the practicality of data asset application, supporting the efficient operation of the carbon trading market and corporate emission reduction decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG WELLSUN INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-04-21
AI Technical Summary
Existing carbon emission measurement technologies suffer from problems such as insufficient integration of multi-source data, inability of static application of carbon emission factors to adapt to changes in energy structure, lack of exploration of data asset value, and insufficient credibility of cross-domain circulation, resulting in reduced measurement accuracy and low data utilization.
By constructing a semantic association network using knowledge graphs to achieve multi-source data fusion, combining reinforcement learning to dynamically adjust carbon emission factors, establishing a two-dimensional evaluation model that links data quality with economic value, performing standardized processing and cross-validation, using blockchain for evidence storage to ensure the credibility of data circulation, and forming a closed-loop management through multi-stage feedback.
It significantly improves the accuracy of carbon emission measurement and the security of data circulation, supports the efficient operation of the carbon trading market and the scientific formulation of corporate emission reduction decisions, and realizes the practicality of data asset application and the precision of supervision.
Smart Images

Figure CN121901211A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of carbon emission measurement, and in particular to a carbon emission measurement data management system and method based on multi-source data. Background Technology
[0002] With the continued deepening of global carbon emission reduction targets and the dynamic adjustment of carbon trading markets and regional energy structures, accurate carbon emission measurement and efficient data management have become core links supporting low-carbon transformation and ensuring the orderly operation of the carbon market. Currently, existing carbon emission measurement and data management technologies generally adopt a core architecture of "single-source data acquisition - static factor calculation - centralized cloud processing": the data acquisition stage mainly relies on IoT sensors to collect single-dimensional data such as electrical parameters and basic environmental parameters, and transmits them to local or cloud nodes through communication protocols such as RS485 and LoRa; carbon emission factors are mostly assigned using fixed benchmark values published by the industry, and are only manually updated at the end of quarters or years based on industry statistical reports; the calculation process is completed through a linear formula of "energy consumption data × fixed emission factor", and some optimization schemes introduce simple environmental parameter correction coefficients, but a dynamic calculation model with multi-parameter coupling has not been established; the data management stage relies on cloud databases to store calculation results and output standardized statistical reports to support basic query and export needs.
[0003] The significance of this technical architecture lies in the fact that it has initially realized the automated collection and basic calculation of carbon emission data, replacing the inefficient mode of early manual statistics, and meeting the needs of basic data recording in early carbon emission monitoring scenarios. However, with the rapid increase in the proportion of new energy sources and the standardization of the carbon trading market, its limitations have gradually become apparent: First, there is insufficient integration of multi-source data. Business data such as carbon quotas and real-time carbon prices, as well as rule data such as accounting benchmarks and regional standards in policy texts, have not been incorporated into the system. Furthermore, there is a lack of semantic association construction for data from different sources, resulting in data fragmentation and low cross-scenario utilization. Second, the static application of carbon emission factors ignores the real-time fluctuations in energy structure and the dynamic changes in equipment operating status, making it unable to adapt to the differences in energy consumption in different regions and at different times. The accuracy of the calculation results continues to decrease over time. Third, the value of data as an asset has not been explored. Data is only used as a monitoring result rather than a tradable asset. A mechanism for linking data quality and economic value has not been established, making it difficult to support core scenarios such as carbon trading pricing. Fourth, the credibility of cross-regional circulation is insufficient. Data sharing relies on simple format conversion and lacks unified semantic standards and cross-validation mechanisms. Cross-enterprise and cross-regional data interaction is prone to misunderstandings and data distortion, failing to meet the current integrated needs of carbon management.
[0004] With the increasing number of participants in the carbon market and the growing complexity of the system, how to achieve deep integration of multi-source data, dynamic adaptation of carbon emission factors, and at the same time, explore the asset value of data and ensure the credibility of cross-regional circulation have become urgent technical challenges in the field of carbon emission data management. Summary of the Invention
[0005] To address the technical deficiencies in the background technology, this invention proposes a carbon emission measurement data management system and method based on multi-source data, which solves the aforementioned technical problems and meets practical needs. The specific technical solution is as follows: The carbon emission measurement data management method based on multi-source data includes the following steps: Collect basic data, business data, and rule data; construct a semantic association network through knowledge graph to achieve data fusion; and filter out effective data with quality labels through quality cleaning to form a structured data set. Based on a structured dataset, the initial carbon emission factor is determined by combining industry benchmarks and historical data. The factor is then dynamically adjusted using reinforcement learning with the measurement deviation as feedback. The optimized factor and adjustment log are generated and written to the blockchain for storage. Based on effective data and optimized factors, the value of data assets is calculated and classified by combining carbon price and emission reduction coefficient, forming a graded data asset with value label; The hierarchical data assets are standardized and metadata is attached. Cross-domain semantic consistency is verified based on semantic association network. After cross-validation and permission allocation, a trusted data flow is formed. High-value data from trusted data flows is pushed to the carbon trading platform, medium-value data is output to enterprise terminals, and all data is made available to regulatory interfaces, thus completing the closed-loop application of data.
[0006] Furthermore, the basic data is collected through an IoT interface, including electrical parameters and environmental parameters. The electrical parameters include voltage, current, power, and power flow direction, and the environmental parameters include temperature and humidity. The business data is collected through the carbon trading platform interface and includes carbon quotas, real-time carbon prices, and emission reduction coefficients. The rule data is collected from carbon policy texts through a policy parsing interface, and natural language processing technology is used to extract applicable industry and regional standards and accounting benchmark values. The entities in the semantic association network include data source identifiers, measurement object identifiers, policy document identifiers, industry types, and regional affiliations. The relationships in the semantic association network include the affiliation relationship between data sources and data types, the association relationship between data sources and data collection accuracy, the subordinate relationship between measurement objects and industry types, the geographical relationship between measurement objects and regional affiliations, the adaptation relationship between carbon policy constraints and applicable industries, and the correspondence relationship between carbon policy constraints and regional standards. A unified semantic label is added to the basic data through a semantic association network, where electrical parameters are associated with the energy consumption dimension label, environmental parameters are associated with the external impact dimension label, carbon quotas are associated with the policy constraint dimension label, and carbon prices are associated with the economic value dimension.
[0007] Furthermore, the quality cleaning is based on the comprehensive data quality index, which consists of three dimensions: completeness, accuracy, and timeliness. The completeness is calculated by the ratio of the number of valid data entries to the total number of data entries collected, and the ratio is required to be no less than 95%. For data that does not meet the completeness standard, linear interpolation of data from adjacent collection points in the same time period is used to complete it. The accuracy is calculated by the deviation rate between the data to be cleaned and the data collected by the reference device. The deviation rate is required to be no more than 3%. For data that does not meet the accuracy standard, a second collection process of the corresponding collection device is triggered. If the second collection still does not meet the standard, it is marked as invalid data and discarded. The timeliness is calculated by the difference between the time of data collection completion and the time of transmission to the processing node. The difference is required to be no more than 10 seconds. Data that does not meet the timeliness standard will be downgraded to historical data storage and used only for trend analysis and will not participate in real-time carbon emission calculation and asset value assessment. After cleaning, each valid data entry is labeled with specific values for completeness, accuracy, and timeliness, forming a quality information label.
[0008] Furthermore, the specific steps for generating the optimized factors are as follows: The industry type, historical factors, and error data of the measurement object are retrieved from the structured dataset, and the initial carbon emission factor is calculated by weighted average method after matching the industry benchmark value. Based on the initial carbon emission factor, real-time carbon emission calculation is performed using effective data to generate real-time calculated values. The real-time calculated values are then compared with the actual verified values for the same period to calculate the deviation rate. Using the bias rate as the state input for reinforcement learning, the factor adjustment amount is set as the action, a reward function is constructed, the action value function is iteratively optimized through the Q-Learning algorithm, the optimal adjustment amount is output, and the optimized factor is calculated in combination with the initial carbon emission factor. Key data from the optimization process are extracted, a factor adjustment log is generated, and the log is transmitted to the blockchain notarization module for hash calculation and on-chain storage, thus completing the data loop of the entire factor optimization process.
[0009] Furthermore, the value of the data assets is calculated using a two-dimensional quantification model. The first dimension is the basic value, the second dimension is the added value, and the total value is the sum of the basic value and the added value. The formula for calculating the basic value is: , in, Based on value, This refers to the comprehensive quality indicators corresponding to the quality label. For real-time carbon prices collected, Carbon emissions calculated based on optimized factors; Formula for calculating added value: , in, As added value, The emission reduction coefficients collected are determined based on the energy type used by the object being measured; The grading criteria for the value of the data assets are as follows: Data assets with a total value of ≥100,000 yuan are designated as Level 1 data assets and used for carbon trading settlement and cross-regional large-scale asset transfers. Data assets valued between RMB 10,000 and RMB 100,000 are designated as Level 2 data assets and used for corporate emission reduction decisions and medium-term carbon management planning. Data assets with a total value of less than 10,000 yuan are classified as Level 3 data assets and are used for internal statistical analysis and regulatory filing.
[0010] Furthermore, the specific steps for forming a trusted data flow are as follows: Based on tiered data assets with value identifiers, the data structure is converted according to JSON-LD format, and metadata is added synchronously to form a standardized data asset stream. The metadata includes quality tags, optimized factors, data collection timestamps, and measurement object identifiers, and the metadata is bound to the tiered data assets through a unique identifier. Based on standardized data asset streams, the semantic association network is invoked to extract core fields from the standardized data asset streams. These fields are then compared with the unified semantic labels defined in the semantic association network to verify the consistency of cross-domain data in core concepts and generate a semantic consistency verification result stream. Based on the semantic consistency verification result stream, if the verification passes, the standardized data asset stream is distributed to three or more cross-domain nodes. Each node independently calculates the carbon emission data of the same measurement object based on local data, generates node calculation results, compares the node calculation result stream with the carbon emission in the standardized data asset stream, calculates the deviation rate, and forms a cross-validation result stream. Based on the cross-validation result stream, if the validation passes, access permissions are allocated according to the hierarchical results. Specifically, Level 1 data assets are open to carbon exchange nodes, enterprise authorized nodes, and regulatory nodes; Level 2 data assets are open to enterprise management nodes; and Level 3 data assets are open to internal statistical nodes. The permission configuration and data assets are bound to an encryption key to form a trusted data flow with permission identifiers.
[0011] Furthermore, the specific process of the data closed-loop application is as follows: Based on the trusted data flow, it is divided into a first-level data sub-flow, a second-level data sub-flow, and a full data traceability flow; The primary data sub-stream is transmitted to the carbon trading platform. Based on the sub-stream data, the carbon trading platform calculates the trading price, generates a settlement data stream, and feeds it back to the system, thus completing the closed loop of trading data. The secondary data substream is transmitted to the enterprise terminal, which uses a built-in algorithm to generate a correlation curve between the investment amount for equipment upgrades, the expected emission reduction, and the increase in asset value, forming an emission reduction decision recommendation stream for enterprises to implement upgrade plans. By opening up the full data traceability stream to the regulatory interface, regulators can call the traceability stream through the interface to trace back the entire chain of any data from collection to application, generate audit result streams and feed them back to the system, thus achieving a closed loop of regulatory data.
[0012] Furthermore, feedback adjustments are achieved through closed-loop data applications, with the following specific steps: The system receives the audit result stream output from the regulatory interface. If the audit finds data quality issues, it extracts the problem data identifier, the ID of the data collection device involved, and the specific deviation value, generates a quality issue feedback stream, and transmits it to the data quality cleaning module. Based on the quality issue feedback stream, the data quality cleaning module adjusts the accuracy verification threshold of the corresponding data collection device, or increases the sampling frequency of the device, and optimizes the cleaning rules. The system receives the feedback stream of the transformation execution output from the enterprise terminal, transmits it to the reinforcement learning factor optimization module, calculates the deviation between the actual emission reduction and the predicted emission reduction in the decision suggestion stream, incorporates the deviation into the adjustment basis of the reward function, and iteratively optimizes the action value function of the Q-Learning algorithm. The system receives settlement feedback streams from the carbon trading platform and transmits them to the data asset valuation module. If the average deviation rate exceeds 10%, the weighting of the total value is adjusted or the standard for the emission reduction coefficient is corrected.
[0013] A carbon emission calculation data management system based on multi-source data includes a multi-source heterogeneous data acquisition and fusion module, a data quality cleaning module, a distributed asset pool module, a factor optimization module, a blockchain notarization module, a value assessment module, a cross-domain mutual recognition module, and a regulatory audit module. Each module forms a collaborative closed loop through hierarchical data flow and bidirectional feedback. The multi-source heterogeneous data acquisition and fusion module integrates IoT interfaces, carbon trading platform interfaces, and policy analysis interfaces. It is used to collect multi-source data, generate raw data streams with semantic tags through knowledge graphs, and connect to the data quality cleaning module through a wired communication link to transmit the raw data streams to the data quality cleaning module in real time. The data quality cleaning module receives the raw data stream, filters and completes the data according to completeness, accuracy and timeliness, generates a valid data stream with quality labels, and outputs it to the distributed asset pool module. The distributed asset pool module is used to store effective data streams and build multi-dimensional indexes, providing data query interfaces for the factor optimization module and the value assessment module; The factor optimization module is used to call data from the distributed asset pool module, calculate the initial factors, and output the optimized factors and adjustment logs through reinforcement learning, which are then output to the value assessment module and the blockchain evidence storage module, respectively. The blockchain evidence storage module is used to store the hash values of adjustment logs and cross-domain records, and provides an evidence storage query interface; The value assessment module is used to receive valid data streams and optimized factors, calculate the value of data assets and classify them, and output the classified asset streams with value labels to the cross-domain mutual recognition module. The cross-domain mutual recognition module is used to standardize the tiered asset flow and attach metadata, verify semantic consistency and data deviation, generate a trusted circulation data flow through access control, and output it to the carbon trading platform, enterprise terminal and regulatory interface respectively. The regulatory audit module is used to receive trusted data streams and blockchain-based evidence logs, provide a full-chain traceability interface, and report quality issues to the data quality cleaning module.
[0014] Furthermore, it also includes: The transaction support module is used to receive trusted circulation data streams and generate settlement data according to the pricing formula; The decision support module is used to receive secondary data from trusted data streams and generate emission reduction decision recommendations.
[0015] Compared with existing technologies, the carbon emission measurement data management system and method based on multi-source data provided by this invention have the following beneficial effects: This invention constructs a semantic association network of multi-source heterogeneous data through knowledge graphs, achieving deep integration of basic data, business data, and rule data. It combines reinforcement learning algorithms with real-time deviation measurement as feedback to dynamically optimize carbon emission factors, establishes a two-dimensional evaluation model linking data quality and economic value to complete data assetization grading, and constructs a cross-domain data trustworthy circulation mechanism through standardized processing, semantic verification, and cross-validation. It relies on blockchain for evidence storage to ensure full-chain traceability and immutability of data, and forms closed-loop management through multi-stage feedback adjustments. This effectively solves the problems of insufficient multi-source data integration, lagging factor adjustment, lack of data value, and low credibility of cross-domain circulation in existing technologies. It significantly improves the accuracy of carbon emission measurement, the security of data circulation, and the practicality of assetization applications, comprehensively supporting the efficient operation of the carbon trading market, the scientific formulation of corporate emission reduction decisions, and the precise supervision of regulatory authorities. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the carbon emission measurement data management method based on multi-source data in this invention.
[0017] Figure 2 This is a schematic diagram of the modules of the carbon emission measurement data management system based on multi-source data in this invention. Detailed Implementation
[0018] In the description of this invention, it should be understood that the terms "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "middle," and "inner," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, it should be noted that unless otherwise explicitly specified and limited, the terms "installed," "connected," and "joined" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention through specific circumstances.
[0019] The embodiments of the present invention will be described below with reference to the accompanying drawings and related examples. The embodiments of the present invention are not limited to the following examples, and the present invention relates to the relevant necessary components in this technical field, which should be regarded as well-known technology in this technical field and can be known and mastered by those skilled in this technical field.
[0020] See Figure 1 This invention provides a carbon emission measurement data management method based on multi-source data, characterized by the following steps: Step S100: Collect basic data, business data and rule data, construct a semantic association network through knowledge graph to achieve data fusion, and filter out effective data with quality labels through quality cleaning to form a structured data set; Basic data refers to raw physical quantity data reflecting energy consumption and the operating environment, originating from equipment or environmental monitoring systems. It can be used as the fundamental input for carbon emission calculations, supporting the quantitative calculation of the relationship between energy consumption and emissions. Business data refers to economic and managerial data reflecting the operational status of the carbon market and corporate behavior. It can provide market background and business context for carbon emission calculations, enhancing the data's applicability. Rule data can be policy and normative textual information guiding the implementation of carbon emission accounting methods and standards. It can be used to ensure that the carbon emission calculation process complies with current regulatory requirements and support cross-regional compliance assessments. A knowledge graph is a semantic network structure organized in the form of entity-relationship-attribute triples. It is used to represent the logical relationships between multi-source heterogeneous data. It can be used to achieve unified modeling and interconnection of different types of data at the semantic level, supporting cross-domain information fusion and contextual understanding. In this embodiment, the knowledge graph can be constructed based on a graph database, using key entities in basic data, business data, and rule data as nodes. Semantic parsing is used to establish relationships between them, forming a reasonable knowledge system. Semantic association networks are semantic connection systems between multi-source data driven by knowledge graphs. They reflect the logical dependencies and contextual relationships between data elements and can be used to break down data silos, achieve deep integration of basic, business, and rule data, and improve data interpretation capabilities and cross-scenario reusability. Quality tags are meta-identifiers attached to raw data to characterize the data's completeness, accuracy, and reliability level. They can be used to distinguish between high-reliability and low-quality data, providing a selection basis for subsequent analysis. Valid data is a qualified subset of data that has undergone quality cleaning and retains quality tags. It can be used for subsequent modeling and analysis, reducing noise interference and improving the stability and reliability of downstream models. Structured datasets are datasets organized in a unified format after fusion, cleaning, and tagging, supporting efficient querying and reuse.
[0021] Step S200: Based on the structured dataset, the initial carbon emission factor is determined by combining industry benchmarks and historical data. The factor is dynamically adjusted by using reinforcement learning to calculate the deviation as feedback. The optimized factor and adjustment log are generated and written to the blockchain for storage. Industry benchmarks are reference values for typical carbon emission intensity per unit of energy consumption published by authoritative institutions. They can be used as the starting point for initializing carbon emission factors, ensuring that the calculation results are policy compliant and comparable across different regions. Historical data are records of actual energy consumption and emissions accumulated over past periods, including timestamps and relevant contextual information. They can be used to identify emission trends, periodic characteristics, and abnormal patterns, assisting in initial factor setting and model training. Initial carbon emission factors are the initial emission conversion coefficients set before dynamic optimization begins. They are usually determined based on a combination of industry benchmarks and historical averages. They can be used as initial input parameters for reinforcement learning algorithms to initiate the factor adaptive adjustment process. Reinforcement learning algorithms are a type of learning algorithm framework that uses an agent to interact with the environment and continuously adjusts its strategy based on reward signals. They can be used to achieve online adaptive adjustment of carbon emission factors, allowing them to dynamically evolve with changes in energy structure and equipment operating status. Calculation deviation is the difference between the current carbon emission calculation result and the actual emission amount, usually expressed as an absolute difference or relative error. It can be used as the core feedback signal for reinforcement learning algorithms, driving factor parameters to adjust towards greater accuracy. The optimized factor is the latest carbon emission conversion coefficient obtained through reinforcement learning algorithms. It reflects the current actual emission characteristics and can be used to replace the static benchmark value for current carbon emission calculations, improving the accuracy and timeliness of the calculation. The adjustment log records the time, value, triggering reason, and algorithm decision basis for each change of the carbon emission factor. It can be used to provide audit clues for the factor evolution trajectory and support post-event traceability and accountability.
[0022] Step S300: Based on valid data and optimized factors, calculate the value of data assets by combining carbon price and emission reduction coefficient, and classify them to form graded data assets with value labels; Emission reduction coefficients are technical parameters that measure the contribution of a unit of data to emission reduction efforts. They are typically related to data accuracy and coverage, and can be used to quantify the functional value of data in promoting low-carbon transformation and participate in asset value calculations. Data asset value comprehensively reflects the overall value level of carbon emission data across both quality and economic dimensions, and can provide a quantitative basis for data grading and differentiated circulation. Graded data assets are ordered sets of data classified according to their data asset value scores. Different levels correspond to different usage permissions and circulation strategies, which can be used to achieve refined management of data resources and value-oriented distribution control.
[0023] Step S400: Standardize the hierarchical data assets and attach metadata. Verify cross-domain semantic consistency based on semantic association network. After cross-validation and permission allocation, a trusted data flow is formed. Metadata is supplementary information describing the content, structure, source, and semantic characteristics of data. It can be used to enhance the self-descriptive capabilities of data and support automated parsing and cross-system interoperability. Standardization is the process of unifying data formats, encoding rules, and units of measurement to predetermined specifications. It can be used to eliminate syntactic differences and improve cross-system compatibility and integration efficiency. Cross-domain semantic consistency is the state of consistent understanding of the same data concept among different organizations or regions. It can be used to prevent data misuse or decision-making bias due to misunderstandings of terminology. Cross-validation is a mechanism that compares the same data item from multiple independent sources or methods to confirm its authenticity and accuracy. It can be used to enhance data credibility and prevent false reporting or transmission errors. Access control is a management mechanism that grants different levels of data access and operation rights based on user identity, role, or organizational attributes. It can be used to ensure the security of sensitive data and prevent unauthorized access and abuse. Trusted data flow is a securely shareable data transmission flow formed after standardization, semantic verification, cross-validation, and access control. It can be used to achieve efficient and reliable data flow across enterprises and regions.
[0024] Step S500: Push high-value data from the trusted circulation data stream to the carbon trading platform, output medium-value data to enterprise terminals, and open all data to the regulatory interface to complete the data closed-loop application.
[0025] High-value data, representing the highest-level data assets in the tiered system, possesses both high quality and high economic impact. It can be prioritized for supply to carbon trading platforms for pricing reference and quota settlement. In this embodiment, high-value data can be obtained by selecting a subset of data assets with value scores in the top range from the tiered data assets. Medium-value data, representing the middle-level data assets in the tiered system, is suitable for internal management and general analysis. It can be pushed to enterprise terminals to support energy efficiency analysis and emission reduction path planning. For example, medium-value data can be obtained by extracting the portion with a middle value score from the tiered data assets. Full-volume data is the complete collection of all processed carbon emission-related data in the system, regardless of its quality or value level. It can be made available to regulatory authorities to meet the needs of comprehensive monitoring and compliance review. In a specific embodiment, full-volume data can be obtained by aggregating data assets from all levels, removing duplicates, and forming a unified view. Completing the data closed-loop application means achieving seamless integration and feedback loops from data collection to final distribution. Furthermore, the completion of data closed-loop applications can be achieved by setting up monitoring dashboards to track the status of each link and introducing user feedback to adjust the hierarchical strategy, thereby forming a sustainable and iterative carbon data governance system.
[0026] Taking the collaborative management of carbon emissions in industrial parks as an example, the carbon emission calculation data management method based on multi-source data in this embodiment can involve multiple manufacturing enterprises in an industrial park accessing sensors such as electricity meters and gas meters through a unified platform to collect basic energy consumption data. Simultaneously, it connects to the regional carbon exchange to obtain real-time carbon prices and imports the latest emission accounting rules issued by the provincial Department of Ecology and Environment as rule data. The system uses a knowledge graph to establish semantic links between the energy use behavior of each enterprise and the policy requirements of its industry, forming a high-quality structured dataset after cleaning. An initial carbon emission factor is set based on local historical emission characteristics and industry benchmark values. Subsequently, a reinforcement learning algorithm continuously optimizes the factor parameters based on the deviation between the daily CEMS measured values and the model output, and all adjustment records are stored on the blockchain. The system further combines carbon prices and emission reduction potential to assess the asset value of the data reported by each enterprise, classifying it into high, medium, and low levels. High-value data (such as precise emission flows from key emission units) are pushed to the provincial carbon trading platform for quota settlement; medium-value data (such as production line-level energy efficiency data) are pushed to the enterprise's own management system to support energy-saving technological transformation decisions; and full data is opened to the ecological and environmental departments through regulatory interfaces to support total emission control and law enforcement inspections, thereby achieving closed-loop governance of carbon data in the park.
[0027] In this invention, a knowledge graph serves as the underlying semantic modeling tool, unifying the representation of entities and relationships in three types of heterogeneous data to form a semantic association network, achieving deep fusion of cross-source data. Based on this, a quality cleaning mechanism assigns quality labels to the data, filtering out valid data and forming a structured data set to ensure the input quality for subsequent processing. The initial carbon emission factor is determined jointly by industry benchmarks and historical data, serving as the starting point for calculation. A reinforcement learning algorithm uses real-time calculation deviations as feedback signals to dynamically adjust factor parameters, generating optimized factors, and writing adjustment logs to the blockchain for full traceability. Based on the optimized factors and valid data, combined with carbon prices and emission reduction coefficients, a two-dimensional evaluation model calculates the value of data assets and completes grading, forming tiered data assets with value identifiers. To further support cross-domain circulation, the system standardizes the tiered assets and adds metadata, uses a semantic association network to verify terminology consistency, and ensures data authenticity and access security through cross-validation and permission allocation mechanisms, ultimately forming a trusted data flow. Based on different value levels, high-value data is pushed to the carbon trading platform to support market pricing, medium-value data is output to enterprise terminals to assist in emission reduction decisions, and all data is opened to regulatory interfaces to meet the needs of panoramic monitoring. The entire process achieves closed-loop management through multi-stage feedback (such as deviation feedback, permission feedback, and application feedback), systematically solving the problems of insufficient multi-source data integration, lagging factor adjustment, lack of data value, and low credibility of cross-domain circulation. This achieves the technical effect of transforming carbon emission measurement and management from static recording to dynamic intelligent governance.
[0028] In one embodiment of the present invention, the basic data is collected through an Internet of Things (IoT) interface and includes electrical parameters and environmental parameters. The electrical parameters include voltage, current, power, and power flow direction, and the environmental parameters include temperature and humidity. Employing a dual-mode IoT interface combining wired and wireless connections, the system utilizes an RS485 wired interface for industrial applications and a LoRa / NB-IoT wireless interface for civilian applications, ensuring stable data transmission across different scenarios. Both interfaces support Modbus-RTU / JSON protocols, ensuring compatibility with mainstream sensors and smart carbon meters. The electrical parameters collected include voltage, current, power, and power flow direction (a Hall effect sensor detects the current direction, outputting a "forward / reverse" binary signal to determine the power input / output status). The acquisition logic is as follows: the smart carbon meter collects electrical parameters in real time, generating one structured data entry (including acquisition timestamp, device ID, and parameter value) every 10 seconds. This data is transmitted to edge nodes via the IoT interface. Edge nodes perform preliminary filtering and then label the data with an "electrical parameter-device ID" association identifier. Environmental parameters collected include temperature and humidity, using an integrated temperature and humidity sensor deployed near the carbon meter to ensure spatial correlation between environmental and electrical parameters. The logic for collecting environmental parameters is as follows: the sensor collects data once every 30 seconds and binds it to the electrical parameters of the same device ID to avoid "spatiotemporal misalignment" between environmental parameters and electrical parameters, and to ensure that the impact of the external environment on carbon emissions can be accurately correlated in subsequent calculations.
[0029] The business data is collected through the carbon trading platform interface and includes carbon quotas, real-time carbon prices, and emission reduction coefficients. To establish communication with the carbon trading platform, a RESTful API interface is used. Before integration, enterprise identity authentication (submitting business license and carbon trading account information) must be completed, and an API key assigned by the platform must be obtained (valid for one year, updated every three months) to ensure secure data transmission. The interface supports HTTPS encryption protocol to prevent tampering or leakage during data transmission.
[0030] The rule data is collected from carbon policy texts through a policy parsing interface, and natural language processing technology is used to extract applicable industry and regional standards and accounting benchmark values. It should be noted that the process of extracting policy parameters using natural language processing technology may lead to ambiguity due to the diversity of policy text expressions and the contextual dependence of technical terms. To improve the accuracy and consistency of extraction, the system has a pre-set interface for manual review and correction. When the confidence level of the NLP model for key parameters (such as "accounting baseline values") falls below a preset threshold (e.g., 90%), a manual review process will be triggered. Domain experts will confirm or correct the extraction results, and the corrected samples will be fed back into the model training set to achieve continuous model optimization. Meanwhile, the domain ontology upon which the semantic association network established by the system relies is not static. The system has set up an ontology update mechanism based on time and events: when newly released policy texts are collected, or when changes in industry standard codes (such as national economic industry classification codes) are detected through the interface, an incremental ontology update process will be automatically initiated to ensure that the semantic association network is synchronized with the current institutional and technical context.
[0031] The policy parsing interface connects to government policy release platforms, supporting automatic retrieval of publicly available carbon policy texts at a frequency of once daily to ensure timely access to the latest policies. The retrieved policy text undergoes format conversion, noise reduction, and word segmentation. The BERT Named Entity Recognition (NER) model is used to extract three core parameters from the policy text. The extracted parameters are organized into JSON format according to the fields "Policy ID, Release Date, Applicable Industry, Regional Standard, Accounting Benchmark Value, and Expiration Date," stored in the "Rule Data Partition" of the distributed asset pool, and associated with the "Policy Document Identifier" entity in the semantic association network. Natural Language Processing (NLP) is a branch of artificial intelligence used to enable computers to understand and process human language (text). In this solution, it specifically refers to the technology of extracting structured parameters from unstructured policy text through word segmentation and Named Entity Recognition (NER). Named Entity Recognition is one of the core tasks of NLP, used to identify entities with specific meanings from text (such as "Applicable Industry" and "Accounting Benchmark Value" in this solution). This solution uses the BERT model to improve entity recognition accuracy and adapt to the professional terminology in carbon policies.
[0032] The entities in the semantic association network include data source identifiers, measurement object identifiers, policy document identifiers, industry types, and regional affiliations. The relationships in the semantic association network include the affiliation relationship between data sources and data types, the association relationship between data sources and data collection accuracy, the subordinate relationship between measurement objects and industry types, the geographical relationship between measurement objects and regional affiliations, the adaptation relationship between carbon policy constraints and applicable industries, and the correspondence relationship between carbon policy constraints and regional standards. The semantic association network comprises five core entity classes, each with a unique identifier and key attributes. Data source identifiers include a unique ID, data type, data acquisition device model, deployment location, and accuracy level. Measurement object identifiers include a unique ID, industry type, regional affiliation, and production scale. Policy document identifiers include a unique ID, issuing authority, effective date, expiration date, and scope of application. Industry types include a unique ID, industry code, and typical energy structure. Regional affiliation includes a unique ID, regional level, and regional energy policy. A graph database (such as Neo4j) is used to store the semantic association network. The construction process is as follows: extract entities from the collected multi-source data; establish relationships between entities according to predefined relationship rules; add attributes to the relationships; and support the Cypher query language, allowing subsequent steps (such as factor calculation and cross-domain validation) to retrieve entity and relationship data from the network via query statements. The semantic association network comprises six core relationships, each defined as "starting entity - ending entity - relationship attribute". Specifically: Attribution relationship: Starting entity (data source identifier) → Ending entity (data type), attribute is "data collection frequency"; Association relationship: Starting entity (data source identifier) → Ending entity (collection accuracy), attribute is "accuracy value"; Subordination relationship: Starting entity (measurement object identifier) → Ending entity (industry type), attribute is "subordination time"; Geographic relationship: Starting entity (measurement object identifier) → Ending entity (regional affiliation), attribute is "registered address"; Adaptation relationship: Starting entity (carbon policy constraint, i.e., policy document identifier) → Ending entity (applicable industry), attribute is "adaptation priority"; Correspondence relationship: Starting entity (carbon policy constraint, i.e., policy document identifier) → Ending entity (regional standard), attribute is "standard effective time".
[0033] A unified semantic label is added to the basic data through a semantic association network, where electrical parameters are associated with the energy consumption dimension label, environmental parameters are associated with the external impact dimension label, carbon quotas are associated with the policy constraint dimension label, and carbon prices are associated with the economic value dimension.
[0034] Adding unified semantic labels to data through the aforementioned semantic association network is fundamental to achieving cross-domain understanding. However, different regions or platforms may have terminological differences for the same concept. Therefore, the semantic vocabulary embedded in the system prioritizes or maps to internationally and domestically recognized standard terms (such as the terminology definitions in the ISO 14064 series of standards). Before performing cross-domain semantic consistency verification, the system first attempts to map the terms in the data to be verified to the local vocabulary. If a direct mapping is not possible, the system records the terminology differences and prompts both parties to conduct semantic negotiation and confirmation, serving as a supplementary mechanism for building a sustainable data ecosystem.
[0035] Based on the previously constructed semantic association network (including entities such as data source identifiers and measurement object identifiers, as well as relationships such as attribution and subordination), four core dimensions of semantic tags were specifically designed. Each tag is bound to a specific entity in the network: the "energy consumption dimension tag" is associated with entities corresponding to electrical parameter data; the "external impact dimension tag" is associated with entities corresponding to environmental parameter data; the "policy constraint dimension tag" is associated with entities corresponding to policy data such as carbon quotas; and the "economic value dimension tag" is associated with entities corresponding to business data such as carbon prices. This achieves dimensional unification of data semantics. When data is entered into the database, the system first queries the entity type corresponding to the data through the semantic association network, then matches the corresponding dimension tag and binds it to the data for storage. If the entity-relationship in the semantic association network is subsequently adjusted (such as adding new data types), the tag system will also expand the corresponding new tags synchronously to ensure that the tags and data types are always compatible. In the factor calculation stage, electrical parameter data can be quickly filtered through the "energy consumption dimension label" and the corresponding industry's policy benchmark value can be called through the "policy constraint dimension label". In the cross-domain verification stage, the consistency of labels in different domain data can be compared to ensure that the data semantics are unambiguous. In the value assessment stage, carbon price data can be located through the "economic value dimension label" and environmental parameter correction factors can be called through the "external impact dimension label".
[0036] It should be noted that the quality cleaning is based on the comprehensive data quality index, which consists of three dimensions: completeness, accuracy, and timeliness. The completeness is calculated by the ratio of the number of valid data entries to the total number of data entries collected, and the ratio is required to be no less than 95%. For data that does not meet the completeness standard, linear interpolation of data from adjacent collection points in the same time period is used to complete it. The comprehensive data quality index serves as the benchmark for data cleaning. It employs a design logic of "parallel constraints + priority ranking," prioritizing data in the order of "accuracy (core, directly affecting measurement precision) > completeness (foundation, ensuring data continuity) > timeliness (scenario-based, distinguishing between real-time and non-real-time applications)." Specific thresholds are set with reference to relevant national standards and the needs of the carbon emission measurement industry (completeness ≥ 95%, accuracy deviation rate ≤ 3%, timeliness delay ≤ 10 seconds). During calculation, data batches are divided into "device ID + 1 hour." Completeness is calculated by statistically analyzing the proportion of non-empty and legally formatted data; accuracy is determined by calculating the maximum deviation rate of the batch using data from certified benchmark equipment; and timeliness is verified by determining the maximum delay of the batch through the difference between the collection timestamp and the processing node's receiving timestamp. This ensures that the data meets the requirements for carbon emission measurement in terms of continuity, accuracy, and timeliness. Incomplete data integrity refers to "valid data percentage < 95%", often caused by temporary sensor malfunctions or momentary communication interruptions leading to data loss. The core of the solution is to "complete missing data and avoid data chain breaks". By analyzing the timestamp sequence of data within a batch, missing positions where "timestamps are continuous but data is empty" are identified. One valid data point before and after the missing position is selected as the "interpolation reference point". The following conditions must be met: the time interval between the reference point and the missing point is ≤ 60 seconds, and the accuracy deviation rate of the reference point is ≤ 2%. The "time-weighted linear interpolation method" is used to complete the missing values. The formula is: missing data value = previous reference point data value + (missing point timestamp - previous reference point timestamp) / (second reference point timestamp - previous reference point timestamp) × (second reference point data value - previous reference point data value). The completed data must be marked with an "interpolation completion" tag in the metadata to facilitate subsequent data source tracing.
[0037] The accuracy is calculated by the deviation rate between the data to be cleaned and the data collected by the reference device. The deviation rate is required to be no more than 3%. For data that does not meet the accuracy standard, a second collection process of the corresponding collection device is triggered. If the second collection still does not meet the standard, it is marked as invalid data and discarded. Accuracy failure refers to a batch maximum deviation rate > 3%, often caused by sensor drift or installation position deviation leading to data distortion. The core of the handling is "prioritize repair, and remove if repair is ineffective." Each data point within the batch is compared with the baseline equipment data. For "abnormal data" with a deviation rate > 3%, the sensor ID and acquisition timestamp corresponding to the abnormal data are recorded. A "secondary acquisition command" is sent to the sensor corresponding to the abnormal data. If the sensor is online and responds to the command, it returns secondary acquisition data within 30 seconds, and the deviation rate is recalculated. If the deviation rate is ≤ 3%, the original abnormal data is replaced with the secondary acquisition data. If the deviation rate is still > 3%, it is marked as "invalid data." If the sensor is offline (e.g., communication interruption) or does not respond to the command, the original abnormal data is directly marked as "invalid data." After removing invalid data, the batch integrity is recalculated (≥ 95% is required; if the integrity after removal is < 95%, "linear interpolation completion" is performed on the newly missing positions). Simultaneously, an "accuracy anomaly report" is generated, recording the sensor ID, anomaly time, deviation rate, and handling result (secondary acquisition meets standard / invalid data is removed), and pushed to the maintenance terminal to remind staff to check the sensors.
[0038] The timeliness is calculated by the difference between the time of data collection completion and the time of transmission to the processing node. The difference is required to be no more than 10 seconds. Data that does not meet the timeliness standard will be downgraded to historical data storage and used only for trend analysis and will not participate in real-time carbon emission calculation and asset value assessment. Timeliness failure refers to "maximum batch delay > 10 seconds", which is mostly caused by network congestion and excessive load on edge nodes, resulting in data transmission delay. The core of the solution is to "distinguish the purpose of the data and avoid the impact of delayed data on real-time calculations".
[0039] After cleaning, each valid data entry is labeled with specific values for completeness, accuracy, and timeliness, forming a quality information label.
[0040] Quality labels are the "structured output of the cleaning results". By adding a combination of dimension values and processing identifiers to each valid data point, subsequent modules (factor optimization, value assessment) can quickly identify data quality.
[0041] When calculating the data quality index, the values of completeness (weight 0.3), accuracy (weight 0.5), and timeliness (weight 0.2) in the labels are directly extracted without recalculation.
[0042] Taking a blast furnace energy efficiency monitoring system in a steel enterprise as an example, the carbon emission calculation data management method based on multi-source data in this embodiment can be as follows: a steel plant deploys multiple sets of current, voltage, and temperature sensors in the blast furnace area, reporting data once per second. The system continuously monitors data quality through a quality cleaning module: when a temperature sensor has no data for 3 consecutive seconds due to signal interference (completeness = 87%), the system calls data from adjacent sensors in the same time period for linear interpolation to complete the data; when the deviation between a meter reading and the on-site standard meter reaches 4.2% (>3%), the system immediately triggers a secondary data acquisition command for that meter; if the retest still exceeds the standard, it is marked as invalid and an alarm is triggered; for data with a transmission delay of 15 seconds due to network congestion, the system automatically classifies it into the historical data storage area, using it only for monthly energy consumption trend analysis, and not participating in the daily real-time carbon emission calculation and carbon asset valuation. After cleaning, all qualified data are attached with quality information tags containing 98% completeness, ±2.1% accuracy, and 8 seconds timeliness, serving as the input basis for subsequent factor optimization and asset classification, significantly improving the stability and reliability of carbon emission calculation.
[0043] This invention quantifies and differentiates multidimensional quality indicators through the above steps, enabling data loss to be repaired, measurement errors to be controlled, transmission delays to be classified and managed, and quality status to be explicitly labeled. It can systematically solve the problem of measurement distortion caused by data loss, errors, and delays, and provide a high-fidelity, structured input foundation for dynamic optimization of carbon emission factors, data asset value classification, and reliable cross-domain circulation, significantly enhancing the technical effect of improving the reliability of carbon emission management and the credibility of data assets.
[0044] In one embodiment of the present invention, the specific steps for generating the optimized factor are as follows: Step S201: Retrieve the industry type, historical factors and error data of the measurement object from the structured dataset, match the industry benchmark value and then use the weighted average method to calculate the initial carbon emission factor; The initial carbon emission factor serves as a dynamic optimization "starting point benchmark," balancing industry compliance with the individual characteristics of the measurement object. It retrieves three core data categories from a structured dataset through semantic association between "measurement object identifier and industry type." Specifically, the industry type of the measurement object is used to match the corresponding industry benchmark value; historical factors refer to carbon emission factor data from the past six months, selecting valid historical factors with a historical measurement error of less than 2%; and error data refers to historical measurement error data used to verify the reliability of the historical factors.
[0045] Based on the industry type of the measurement object, "industry benchmark emission factors" published by authoritative institutions are extracted from the rule data to ensure that the initial factors meet policy compliance requirements. A weighted strategy of "industry benchmark value as the main factor and historical effective factors as the supplementary factor" is adopted, and the calculation formula is as follows: Initial carbon emission factor = industry benchmark value × 0.6 + average historical effective factor × 0.4. The industry benchmark value represents the general standard required by the policy and accounts for 60% of the weight to ensure the compliance of the measurement; the historical effective factor reflects the actual emission characteristics of the measurement object and accounts for 40% of the weight to improve the individual adaptability of the factors.
[0046] Step S202: Based on the initial carbon emission factor, perform real-time carbon emission calculation using effective data, generate real-time calculation values, compare the real-time calculation values with the actual verification values of the same period, and calculate the deviation rate. The "actual verification value of the same period" is the core benchmark of the factor optimization feedback loop, and its source and accuracy directly affect the optimization effect. In practical applications, this value can come from one or more of the following channels, and the system supports configuring the priority level based on data availability: 1) High-frequency measured values: from continuous emission monitoring systems (CEMS) deployed at emission outlets and certified by metrology, which can provide near real-time high-precision data as the optimal feedback signal; 2) Periodic verification values: from monthly or quarterly verification reports issued by qualified third-party verification agencies. Although the timeliness is lower, the accuracy is authoritative and can be used as a low-frequency calibration benchmark; 3) Indirect accounting reference values: when the above direct data are lacking, high-frequency (such as daily) accounting values based on the material balance method can be used as an alternative reference, but their uncertainty level must be clearly marked in the metadata. The system allows setting different confidence weights for feedback signals from different sources, and when they are missing for a long time or fluctuate in quality, it can automatically pause the dynamic adjustment of factors and switch to using static industry benchmark factors to ensure the stability of the calculation system.
[0047] The deviation rate is the core "feedback signal" of reinforcement learning. It needs to be generated based on a precise comparison between real-time data and actual verification values. Based on the initial carbon emission factor F0, combined with the selected effective data (electrical parameters, environmental parameters), the real-time calculated value is generated according to the logic of "carbon emissions = energy consumption data × F0". The actual verification value is the "true carbon emission benchmark", which comes from two sources: carbon emission values in verification reports issued periodically (e.g., monthly) by qualified third-party institutions, serving as a low-frequency, high-precision benchmark; and real-time emission data collected by online monitoring equipment (e.g., CEMS system) deployed at the emission outlet of the measured object, serving as a high-frequency, real-time benchmark. Timestamp synchronization ensures that the real-time calculated value and the actual verification value are data from the same time period. The relative deviation rate is used to quantify the accuracy of the calculation, and the calculation formula is as follows: Deviation rate e = |real-time calculated value - actual verification value| / actual verification value × 100%.
[0048] Step S203: Using the bias rate as the state input for reinforcement learning, setting the factor adjustment amount as the action, constructing the reward function, iteratively optimizing the action value function through the Q-Learning algorithm, outputting the optimal adjustment amount, and calculating the optimized factor in combination with the initial carbon emission factor; The adaptive adjustment of factors is achieved through the "state-action-reward" cycle of reinforcement learning. The state space S in reinforcement learning uses the bias rate e as the core state, with a value range of [0%, +∞). In practical applications, the state intervals are divided according to "e≤2% (excellent), 2%<e≤5% (good), 5%<e≤10% (average), e>10% (poor)" to simplify the computational complexity of the algorithm. The action space A sets the factor adjustment amount ΔF as the action, with a value range of [-0.05F0, 0]. [0.05F0] means that each adjustment should not exceed 5% of the current factor value (to avoid excessive adjustment leading to measurement fluctuations), and the adjustment step size is 0.001F0 (e.g., F0 = 1.18 kg CO2 / kWh, step size = 0.00118 kg CO2 / kWh); the reward function R is a positive incentive function that constructs "the smaller the deviation rate, the higher the reward", and the formula is R = 10 × (1 - e / 100). When e ≤ 10%, R is positive; when e > 10%, R = -5.
[0049] The specific process of iterative optimization of the Q-Learning algorithm is as follows: Initialize the action value function Q(S,A), with the initial value of Q set to 0. Q(S,A) represents the expected cumulative reward after performing action A (adjustment amount ΔF) in state S. For each deviation rate e (state S) acquired, the algorithm randomly selects an initial action ΔF1 from the action space A, calculates the corresponding reward R1, and updates Q(S,ΔF1): Q(S,ΔF1) ← Q(S,ΔF1) + α × [R1 + γ × max] a Q(S',a)-Q(S,ΔF1)], where α is the learning rate, which controls the magnitude of each update, with a value of 0.1, γ is the discount factor, which controls the weight of future rewards, with a value of 0.9, and S' is the new state after executing action ΔF1; After 100 iterations, the algorithm selects the action ΔFopt (optimal adjustment) with the largest Q value under the current state S, ensuring that the factor adjustment evolves in the direction of "maximizing reward (minimizing deviation)".
[0050] The optimized carbon emission factor Ffinal = F0 + ΔFopt. For example, if F0 = 1.18 kg CO2 / kWh and the optimal adjustment ΔFopt = -0.02 kg CO2 / kWh, then Ffinal = 1.18 - 0.02 = 1.16 kg CO2 / kWh. If Ffinal exceeds ±10% of the industry benchmark value, a constraint mechanism is triggered, and Ffinal is forcibly corrected to ±10% of the industry benchmark value to ensure that the factor does not deviate from the policy compliance range.
[0051] Considering that the feedback signal (actual verification value) itself may contain noise or short-term fluctuations, to avoid the reinforcement learning algorithm overfitting to noise and causing unnecessary factor oscillations, an action exploration decay mechanism and a bias confidence interval judgment are introduced in the Q-Learning algorithm iteration process. Specifically, as the number of iterations increases, the probability of randomly exploring new actions gradually decreases. Meanwhile, large-scale factor adjustments are only performed when the bias rate e consistently exceeds a robustness threshold (e.g., 2%) and the confidence level of the feedback signal source is high; for occasional, small-amplitude bias fluctuations, the algorithm tends to maintain the current factor or make fine adjustments, thereby improving the robustness of the optimization process.
[0052] Step S204: Extract key data from the optimization process, generate factor adjustment logs, and transmit the adjustment logs to the blockchain notarization module for hash calculation and on-chain storage to complete the data closed loop of the entire factor optimization process.
[0053] Extract key data from the entire factor optimization process and generate a structured adjustment log containing the following fields: measurement object ID, adjustment timestamp (accurate to milliseconds), adjustment period, initial factor F0, optimal adjustment amount ΔF_opt, optimized factor Ffinal, deviation rate before adjustment e, reward value R, expected deviation rate after adjustment e', log generation node ID, and data signature.
[0054] The adjustment logs are converted into string format, and a unique hash value is calculated using the SHA-256 hash algorithm. Each hash value corresponds one-to-one with the log content; modifications to the log content result in a corresponding change in the hash value. The hash value and the original adjustment log text are simultaneously uploaded to the consortium blockchain. The consortium blockchain assigns a unique notarization ID to each log entry, which is associated with the measurement object ID and the adjustment timestamp, supporting queries based on multiple conditions. Nodes synchronously record transactions to ensure the logs are tamper-proof. The optimized factor Ffinal is fed back to the data asset value assessment step as a core parameter for calculating carbon emissions E. Simultaneously, the adjustment log notarization ID is attached to a quality tag, forming a closed-loop data process of "factor adjustment - measurement - value assessment - notarization," ensuring that subsequent traceability can link "factor data - measurement data - value data."
[0055] In one embodiment of the present invention, the value of the data asset is calculated using a two-dimensional quantification model, with the first dimension being the basic value and the second dimension being the added value, and the total value being the sum of the basic value and the added value. The dual-dimensional quantitative model is a mathematical modeling method that decomposes the value of data assets into two components: basic value and added value. These components are calculated independently and then summed for evaluation. This model can be used to characterize the multi-dimensional value of carbon emission data assets, improving the rationality and incentive orientation of valuation results. Basic value reflects the fundamental economic value of data under the coupled effects of quality, market conditions, and emission scale. It can be used to demonstrate the basic monetization capability of high-quality data in the carbon market. Added value is the incremental value that further amplifies or adjusts based on basic value. It reflects the additional contribution brought about by specific emission reduction behaviors or energy structure transformation. It can be used to provide value incentives to entities with high clean energy usage ratios or advanced emission reduction technologies, enhancing the policy guidance function of data asset assessment.
[0056] The formula for calculating the basic value is: , in, Based on value, The comprehensive quality index corresponding to the quality label needs to be calculated first through the "completeness (D1), accuracy (D2), and timeliness (D3)" in the quality label. The formula is: DQI=D1×0.3+D2×0.5+D3×0.2; The real-time carbon price is obtained from the collected business data and retrieved in real time through the carbon trading platform's API interface; The carbon emissions calculated based on the optimized factors (Ffinal) and effective data (energy consumption data) are given by the formula: E = energy consumption data × Ffinal. Formula for calculating added value: , in, As added value, The emission reduction coefficient is obtained from the collected business data and dynamically allocated by the carbon trading platform based on the proportion of clean energy use of the measured object. It has no unit and the value ranges from 0 to 0.3. It is determined according to the type of energy used by the measured object. The specific rules are as follows: Clean energy (solar, wind, hydro, etc.) > 50%: K = 0.2; Clean energy 30%-50%: K = 0.1; Clean energy < 30%: K = 0.05.
[0057] The grading criteria for the value of the data assets are as follows: Data assets with a total value of ≥100,000 yuan are designated as Level 1 data assets and used for carbon trading settlement and cross-regional large-scale asset transfers. Data assets valued between RMB 10,000 and RMB 100,000 are designated as Level 2 data assets and used for corporate emission reduction decisions and medium-term carbon management planning. Data assets with a total value of less than 10,000 yuan are classified as Level 3 data assets and are used for internal statistical analysis and regulatory filing.
[0058] In one embodiment of the present invention, the specific steps for forming a trusted data stream are as follows: Step S401: Based on the hierarchical data assets with value identifiers, convert the data structure according to JSON-LD format, and synchronously attach metadata to form a standardized data asset stream. The metadata includes quality tags, optimized factors, data collection timestamps, and measurement object identifiers, and the metadata is bound to the hierarchical data assets through a unique identifier. The JSON-LD format is adopted, which adds semantic linking capabilities to the traditional JSON. It supports automatic machine parsing and can be associated with a unified semantic standard. The hierarchical data assets are reconstructed according to the preset JSON-LD template. The template includes a core data area and a semantic linking area. The core data area stores core calculation and value information such as carbon emissions, data asset value V, and value level (such as "Level 1"). The semantic linking area is associated with a unified semantic vocabulary of the semantic association network through the "@context" field. Metadata is the "identity statement" of the data and needs to be bound to the standardized data assets through a unique identifier (UUID) to ensure traceability throughout the entire chain.
[0059] Step S402: Based on the standardized data asset stream, call the semantic association network to extract the core fields in the standardized data asset stream, compare them with the unified semantic labels defined in the semantic association network, verify the consistency of cross-domain data in core concepts, and generate a semantic consistency verification result stream. Three types of core semantic fields are extracted from the standardized data asset stream: carbon emission-related fields, factor-related fields, and value-related fields. A semantic association network is invoked to obtain the unified semantic labels defined in the network. The extracted core fields are compared one by one with the unified semantic labels to generate a semantic consistency verification result stream. The comparison rule is that the semantics of the core fields must match the unified semantic labels 100%. If there is any ambiguity, it is marked as semantically inconsistent and returned to the multi-source acquisition and fusion module for correction. Only when the semantic matching degree of all core fields is 100% and there are no ambiguities can the next step of cross-validation be carried out.
[0060] Step S403: Based on the semantic consistency verification result stream, if the verification passes, the standardized data asset stream is distributed to 3 or more cross-domain nodes. Each node independently calculates the carbon emission data of the same measurement object based on local data, generates node calculation results, compares the node calculation result stream with the carbon emission in the standardized data asset stream, calculates the deviation rate, and forms a cross-validation result stream. Select three or more cross-domain nodes with carbon accounting qualifications (such as regional carbon accounting centers, accounting departments of key emitting enterprises, and third-party verification agencies). Nodes must meet the following requirements: possess independent basic data collection capabilities; have a carbon accounting error rate of ≤3% in the past 12 months; and belong to different administrative regions or different entities. Synchronously distribute the semantically validated standardized data asset stream to the selected cross-domain nodes. Each node, based on its own collected "same-period basic data of the same accounting object" (such as electrical parameters and environmental parameters of a certain enterprise in the same hour), independently calculates carbon emissions using optimized factors consistent with the standardized data, generating node calculation results. Compare each node's calculation results with the "carbon emissions in the standardized data asset stream" to calculate the deviation rate. The single-node deviation rate ei = |node i's calculated value - standardized data emissions| / standardized data emissions × 100%, and the average deviation rate eavg = (e1 + e2 + ... + e n The cross-validation standard is as follows: the deviation rate of all single nodes ei ≤ 5%, and the average deviation rate eavg ≤ 3%. If the deviation rate of a node is > 5%, a deviation analysis report (such as its own basic data collection error) must be submitted. If the standard is still not met after recalculation, the node is removed and a new node is added for revalidation. Record the calculation results, deviation rates, and analysis reports of all nodes to form a cross-validation result stream, which is bound to the standardized data asset stream and serves as the basis for the credibility of subsequent permission allocation.
[0061] Step S404: Based on the cross-validation result stream, if the validation passes, access permissions are allocated according to the hierarchical results. Specifically, Level 1 data assets are open to carbon exchange nodes, enterprise authorization nodes, and regulatory nodes; Level 2 data assets are open to enterprise management nodes; and Level 3 data assets are open to internal statistical nodes. The permission configuration and data assets are bound to an encryption key to form a trusted circulation data stream with permission identifiers.
[0062] Establish a three-dimensional mapping relationship between data value levels, permission roles, and operation permissions, clarify the authorization scope of data at different levels, adopt attribute encryption technology, generate and distribute public and private keys according to role attributes, bind permission configurations and data assets with encryption keys, and combine standardized data asset streams and permission identifiers that have passed cross-validation to generate trusted circulation data streams with permission identifiers that include core data areas, verification areas, and permission areas. Finally, output to carbon trading platforms, enterprise terminals, and regulatory interfaces to achieve fine-grained security control and authorized access to cross-domain data.
[0063] Taking the collaborative reporting of carbon data among power companies across provinces as an example, the carbon emission calculation data management method based on multi-source data in this embodiment can be used when a large power generation group needs to report unit emission data to the carbon markets in East China and North China. The system converts its graded high-value data (Level 1) into JSON-LD format, adds metadata including quality tags, optimized factors, timestamps, and unit numbers, and binds it with a unique identifier. Then, it calls the semantic association network to extract core fields such as "fuel type," "installed capacity," and "calculation method," and compares them with the unified semantic tags in the regional standard to confirm that the terminology is consistent. After verification, the data is distributed to three cross-domain nodes: the group headquarters' local system, the provincial regulatory platform, and a third-party verification agency. Each node independently calculates the emissions for the same period based on its own monitoring data and returns the result stream. The system calculates the deviation rate; if all deviations are below 5%, the data is considered reliable. Finally, Level 1 data is pushed to the two regional carbon exchanges and the national regulatory node after being bound with an encryption key, Level 2 data is used by the group's internal management platform, and Level 3 data is retained for statistical analysis, realizing the circulation of high-confidence data across regions.
[0064] In one embodiment of the present invention, the specific process of the data closed-loop application is as follows: Based on the trusted data flow, it is divided into a first-level data sub-flow, a second-level data sub-flow, and a full data traceability flow; The Level 1 data sub-stream is a high-value, frequently updated subset of data separated from the trusted circulating data stream. It is specifically designed for carbon market trading scenarios and can serve as the core input for pricing and settlement on carbon trading platforms, supporting real-time and accurate market behavior decisions. The Level 2 data sub-stream is a medium-value data subset built to meet internal enterprise management needs. It includes equipment-level energy consumption and emission information and can be used to support enterprises in generating emission reduction pathway assessment models and assist in developing technological transformation plans. The full data traceability stream is a complete data link record stream containing all original collected data and its entire processing log. It can be used to provide regulatory agencies with end-to-end data traceability capabilities, supporting authenticity verification and compliance audits.
[0065] The primary data sub-stream is transmitted to the carbon trading platform. Based on the sub-stream data, the carbon trading platform calculates the trading price, generates a settlement data stream, and feeds it back to the system, thus completing the closed loop of trading data. The primary data sub-stream is pushed to the carbon trading platform via an encrypted communication link (TLS 1.3 protocol), with a blockchain-based notarized ID attached during transmission for the platform to verify the data's credibility. Based on the sub-stream data, the carbon trading platform uses a "value benchmark + quality fluctuation" pricing model to generate transaction settlement data. The trading platform then forms a settlement data stream (including actual transaction prices, transaction quantities, and settlement status) and feeds it back to the system's value assessment module for subsequent value model optimization.
[0066] The secondary data substream is transmitted to the enterprise terminal, which uses a built-in algorithm to generate a correlation curve between the investment amount for equipment upgrades, the expected emission reduction, and the increase in asset value, forming an emission reduction decision recommendation stream for enterprises to implement upgrade plans. Through a dedicated interface on the enterprise terminal (supporting both API and SDK access methods), the secondary data sub-stream is pushed to the enterprise's internal management system. The enterprise terminal has a built-in "input-emission reduction-value" correlation algorithm, which generates correlation curves for equipment modification investment amount, expected emission reduction, and asset value increase based on data such as emission change trends and equipment contribution ratio in the sub-stream. The enterprise will form a modification execution feedback stream based on the modification execution status (such as the modified equipment model and actual emission reduction), and feed it back to the system's reinforcement learning factor optimization module.
[0067] By opening up the full data traceability stream to the regulatory interface, regulators can call the traceability stream through the interface to trace back the entire chain of any data from collection to application, generate audit result streams and feed them back to the system, thus achieving a closed loop of regulatory data.
[0068] Through a dedicated regulatory interface (supporting multi-dimensional query and filtering), the full data traceability stream is made available to regulatory authorities. Regulators can query data based on criteria such as the measurement object, time range, and data level. Regulators can use the traceability stream to trace back the entire process of any data (collection → cleaning → factor adjustment → value assessment → cross-domain circulation) to verify the authenticity and compliance of the data. The regulatory authorities will generate an audit result stream (such as data quality issues and compliance conclusions) and feed it back to the system's data quality cleaning module.
[0069] It should be noted that feedback adjustments are achieved through data closed-loop applications, and the specific steps are as follows: The system receives the audit result stream output from the regulatory interface. If the audit finds data quality issues, it extracts the problem data identifier, the ID of the data collection device involved, and the specific deviation value, generates a quality issue feedback stream, and transmits it to the data quality cleaning module. Based on the quality issue feedback stream, the data quality cleaning module adjusts the accuracy verification threshold of the corresponding data collection device, or increases the sampling frequency of the device, and optimizes the cleaning rules. The system receives the audit result stream from the regulatory interface. This stream includes the problem data identifier, the ID of the involved acquisition device, the specific deviation value, and the problem type. The data quality cleaning module parses the quality problem feedback stream and adjusts the cleaning rules for the corresponding acquisition device accordingly: if the quality problem feedback stream shows that the device accuracy deviation exceeds 3% (e.g., long-term device drift), the accuracy verification threshold for that device is relaxed from 3% to 3.5% to avoid frequent triggering of invalid secondary acquisitions, and a device calibration reminder is generated and pushed to the maintenance terminal; if the problem is insufficient integrity, such as unstable communication, the sampling frequency of that device is increased from 10 seconds / time to 5 seconds / time. The adjusted cleaning rules take effect in real time, and a quality rule adjustment log is generated, associated with the problem data identifier and device ID, and stored in the distributed asset pool module for subsequent audit traceability.
[0070] The system receives the feedback stream of the transformation execution output from the enterprise terminal, transmits it to the reinforcement learning factor optimization module, calculates the deviation between the actual emission reduction and the predicted emission reduction in the decision suggestion stream, incorporates the deviation into the adjustment basis of the reward function, and iteratively optimizes the action value function of the Q-Learning algorithm. The system receives the retrofit execution feedback stream from the enterprise terminal, which includes the model of the retrofitted equipment and the actual emission reduction. The reinforcement learning factor optimization module compares the actual emission reduction with the predicted emission reduction from the emission reduction decision suggestion stream and calculates the deviation rate. This deviation rate is then incorporated into the reward function adjustment of the Q-Learning algorithm. If the actual emission reduction is lower than the predicted value, for example, if the deviation rate is greater than 10%, the reward value for the corresponding factor adjustment action is reduced. For example, the original reward R=9.8 is adjusted to R=7.5, guiding the algorithm to output more conservative factor adjustment amounts in the future and avoiding prediction deviations caused by excessive factor adjustment.
[0071] The system receives settlement feedback streams from the carbon trading platform and transmits them to the data asset valuation module. If the average deviation rate exceeds 10%, the weighting of the total value is adjusted or the standard for the emission reduction coefficient is corrected.
[0072] The system receives settlement feedback streams from the carbon trading platform, which include data asset identifiers, system-assessed value V, and actual transaction prices. The data asset valuation module calculates the deviation rate between the assessed value and actual transaction price of all primary data assets over the past 30 days. If the average deviation rate exceeds 10%, model adjustments are initiated: if the deviation stems from an overestimation of the basic value V1, the weighting of the basic value and added value is adjusted from "V1 accounts for 80%, V2 accounts for 20%" to "V1 accounts for 75%, V2 accounts for 25%"; if the deviation stems from an excessively high value of the emission reduction coefficient K, the value standard of K is corrected. The adjusted valuation model is backtested against historical data. If the average deviation rate drops below 10%, it officially takes effect, and a value model adjustment log is generated and stored on the blockchain.
[0073] The feedback adjustment achieved through data closed-loop application is the core of the system's self-optimization. To avoid parameter adjustment conflicts or system instability that may be caused by simultaneous feedback streams from different ports (regulatory, enterprise, and trading platforms), a central feedback coordinator is established. This coordinator is responsible for receiving all external feedback streams and scheduling them according to preset priority rules and effective windows. The rules are as follows: 1) Priority: Regulatory audit feedback (involving data quality and compliance) has the highest priority, followed by enterprise execution feedback (involving model prediction accuracy), and finally market transaction feedback (involving the economics of the value model). High-priority feedback can interrupt the adjustment process of low-priority feedback. 2) Effective window: For non-urgent parameter adjustments (such as value model weights), the feedback coordinator will place them in a pending queue and execute them in batches during a preset time window when the system load is low (such as early morning every day). The stability of the system output will be observed for a period of time after the adjustment before deciding whether to officially take effect.
[0074] Taking the closed-loop optimization of carbon management in steel enterprises as an example, the carbon emission calculation data management method based on multi-source data in this embodiment can be as follows: After a large steel group connects to this system, the regulatory department discovers periodic fluctuations in some blast furnace energy consumption data during quarterly audits and returns the audit result stream through the regulatory interface. The system automatically extracts the problem data identifier, the corresponding thermocouple sensor ID, and the maximum deviation value (±8%), generates a quality problem feedback stream, and transmits it to the data quality cleaning module. This module then increases the sampling frequency of such sensors to 5 times per minute and adjusts the fluctuation verification threshold from ±5% to ±3%. At the same time, it introduces a cross-validation rule for adjacent measuring points, effectively suppressing subsequent abnormal reporting. Meanwhile, the enterprise implements sintering waste heat recovery transformation and uploads the transformation execution feedback stream through the enterprise terminal. The system comparison finds that the actual emission reduction per ton of steel is 12% lower than the original prediction. Therefore, this deviation is included in the reward function, reducing the recommendation weight of similar energy-saving suggestions in future decisions, and iteratively optimizing the action value function of the Q-Learning algorithm, making it pay more attention to operational stability rather than theoretical efficiency in subsequent factor adjustments. In addition, feedback from the carbon trading platform shows that the average transaction price of carbon credit products based on the company's data assets is 9.7% lower than the estimated value, close to the preset threshold. The system automatically fine-tunes the carbon price weight ratio in the value assessment model and organizes the technical committee to revise the emission reduction coefficient value standards for different types of energy-saving projects, realizing a closed-loop feedback from regulation, implementation to the market.
[0075] See Figure 2This invention also provides a carbon emission measurement data management system based on multi-source data, characterized by comprising a multi-source heterogeneous data acquisition and fusion module, a data quality cleaning module, a distributed asset pool module, a factor optimization module, a blockchain notarization module, a value assessment module, a cross-domain mutual recognition module, and a regulatory audit module. Each module forms a collaborative closed loop through hierarchical data flow and bidirectional feedback, wherein: The multi-source heterogeneous data acquisition and fusion module integrates IoT interfaces, carbon trading platform interfaces, and policy analysis interfaces. It is used to collect multi-source data, generate raw data streams with semantic tags through knowledge graphs, and connect to the data quality cleaning module through a wired communication link to transmit the raw data streams to the data quality cleaning module in real time. The IoT interface collects electrical and environmental parameters; the carbon trading platform interface collects carbon quotas, real-time carbon prices, and emission reduction coefficients; the policy analysis interface collects carbon policy texts and extracts applicable industry and regional standards and accounting benchmarks through natural language processing. The multi-source heterogeneous data acquisition and fusion module incorporates a knowledge graph engine, defining entities such as "data source identifier - calculation object identifier - policy document identifier" and relationships such as "attribution relationship - subordinate relationship," generating a raw data stream with unified semantic tags. This raw data stream is transmitted in real-time to the data quality cleaning module via a wired communication link, with interface identity identifiers added during transmission to prevent data source forgery.
[0076] The data quality cleaning module receives the raw data stream, filters and completes the data according to completeness, accuracy and timeliness, generates a valid data stream with quality labels, and outputs it to the distributed asset pool module. After receiving the raw data stream, the data quality cleaning module verifies it according to three indicators: completeness, accuracy, and timeliness. The built-in quality verification engine automatically calculates the indicator values. For data with insufficient completeness, linear interpolation is used to complete it. Data with excessive accuracy is triggered for secondary collection. Data with excessive timeliness is downgraded to historical data. A valid data stream with quality labels is generated and output to the distributed asset pool module through a dedicated interface, while cleaning logs are recorded.
[0077] The distributed asset pool module is used to store effective data streams and build multi-dimensional indexes, providing data query interfaces for the factor optimization module and the value assessment module; The distributed asset pool module adopts a distributed storage architecture, storing effective data streams in categories of "basic data - business data - rule data", with each type of data associated with a unique identifier; it has a built-in index building unit that builds multi-dimensional indexes based on "measurement object ID, data type, collection time, and quality label", supporting queries based on multiple conditions; it provides a standardized data query interface to support data access for the factor optimization module and value assessment module, and the interface calls require verification of the module's identity.
[0078] The factor optimization module is used to call data from the distributed asset pool module, calculate the initial factors, and output the optimized factors and adjustment logs through reinforcement learning, which are then output to the value assessment module and the blockchain evidence storage module, respectively. The factor optimization module retrieves the industry type and historical factors (error < 2%) of the measurement object from the distributed asset pool. After matching the industry benchmark value, it calculates the initial factor using a weighted average method (industry benchmark value weight 0.6, historical factor weight 0.4). The built-in reinforcement learning engine (Q-Learning algorithm) constructs a reward function R = 10 × (1-e) (R = -5 when e > 10%) with the measurement deviation rate e as the state and the factor adjustment amount ΔF as the action. It iteratively optimizes and outputs the optimal ΔF, calculates the optimized factor, outputs the optimized factor to the value assessment module, and outputs the factor adjustment log to the blockchain evidence storage module.
[0079] The blockchain evidence storage module is used to store the hash values of adjustment logs and cross-domain records, and provides an evidence storage query interface; After receiving factor adjustment logs and cross-domain records, the blockchain evidence storage module uses the SHA-256 algorithm to generate a unique hash value through its built-in hash calculation unit. It adopts a consortium blockchain architecture to store the hash value and the original record in the consortium blockchain node. The evidence storage is completed after the node reaches consensus. It provides an evidence storage query interface to support querying the original record and hash value by evidence storage ID, which is used by regulatory agencies and trading platforms to verify the credibility of the data.
[0080] The value assessment module is used to receive valid data streams and optimized factors, calculate the value of data assets and classify them, and output the classified asset streams with value labels to the cross-domain mutual recognition module. The value assessment module obtains effective data streams and quality tags from the distributed asset pool, obtains optimized factors from the factor optimization module, calculates them according to the two-dimensional model, generates a hierarchical asset stream with value labels, and outputs it to the cross-domain mutual recognition module.
[0081] The cross-domain mutual recognition module is used to standardize the tiered asset flow and attach metadata, verify semantic consistency and data deviation, generate a trusted circulation data flow through access control, and output it to the carbon trading platform, enterprise terminal and regulatory interface respectively. After receiving the tiered asset stream, the cross-domain mutual recognition module converts the data structure into JSON-LD format, adds metadata, and binds the metadata and data with a unique identifier. It calls the semantic association network to verify the semantic consistency of the cross-domain data, distributes the data to more than 3 cross-domain nodes for independent measurement, calculates the deviation rate, uses attribute encryption technology to allocate permissions according to tiers, generates a trusted circulation data stream, and outputs it to the carbon trading platform, enterprise terminals, and regulatory interfaces.
[0082] The regulatory audit module is used to receive trusted data streams and blockchain-based evidence logs, provide a full-chain traceability interface, and report quality issues to the data quality cleaning module.
[0083] The regulatory audit module receives trusted circulation data streams and blockchain-based evidence logs, storing them in a dedicated regulatory database. It has a built-in traceability query unit that supports multi-dimensional queries of the entire data chain record by "object ID, time range, and data level." The built-in audit analysis unit generates regulatory audit reports, and if quality issues are found, it generates a quality issue feedback stream and pushes it to the data quality cleaning module.
[0084] It should be noted that this also includes: The transaction support module is used to receive trusted circulation data streams and generate settlement data according to the pricing formula; The transaction support module receives the trusted circulation data stream output by the cross-domain mutual recognition module, extracts the data asset value V and quality label DQI, calculates the transaction price according to the pricing formula "Pricing = V×(1+0.05×(DQI-0.8))," generates settlement data, pushes the settlement data to the carbon trading platform, and generates a settlement feedback stream, which is transmitted back to the value assessment module to optimize the value assessment model.
[0085] The decision support module is used to receive secondary data from trusted data streams and generate emission reduction decision recommendations.
[0086] The decision support module receives the trusted data stream output by the cross-domain mutual recognition module, extracts the carbon emission change trend, the carbon emission contribution ratio of equipment, and value assessment parameters, and has a built-in data analysis unit to generate a correlation curve of "equipment transformation investment amount - expected emission reduction - asset value increase" through a linear regression algorithm; it forms an emission reduction decision suggestion stream, pushes it to the enterprise terminal, and records the enterprise's transformation feedback at the same time.
[0087] The overall system operation process is as follows: The multi-source heterogeneous data acquisition and fusion module collects three types of data—basic, business, and rule-based—through the Internet of Things (IoT), carbon trading platforms, and policy analysis interfaces. It then uses knowledge graphs to construct a semantic association network, generating a raw data stream with semantic tags, which is transmitted in real-time to the data quality cleaning module. The data quality cleaning module filters and completes the data according to completeness, accuracy, and timeliness indicators, generating a valid data stream with quality tags. This stream is then output to the distributed asset pool module for storage and a multi-dimensional index is established. The distributed asset pool module provides data query support for the factor optimization and value assessment modules. The factor optimization module calculates the initial carbon emission factor using data and then dynamically optimizes it using a reinforcement learning algorithm, generating optimized factors and adjustment logs. The optimized factors are transmitted to the value assessment module, and the adjustment logs are transmitted to the blockchain notarization module for on-chain notarization. The value assessment module combines the valid data and optimized factors, calculates the data asset value using a two-dimensional model, and classifies it, outputting a tiered asset stream to the cross-domain mutual recognition module. The cross-domain mutual recognition module standardizes the tiered asset stream and adds metadata. After semantic consistency verification, cross-validation, and attribute encryption permission allocation, a trusted circulation data stream is generated and pushed to the carbon trading platform, enterprise terminals, and regulatory interfaces, while also distributing it to the supplementary module. The transaction support module receives data streams, extracts value and quality information, generates settlement data according to the pricing formula, pushes it to the carbon trading platform, and feeds back the settlement data stream to the value assessment module. The decision support module receives data streams, extracts trends and parameters, generates emission reduction decision recommendations, and pushes them to the enterprise terminal. The regulatory audit module receives trusted circulation data streams and blockchain-based evidence logs, provides full-chain traceability, and generates audit reports. If quality issues are found, the quality issue stream is fed back to the data quality cleaning module. The enterprise terminal then feeds back the implementation status of the improvements to the factor optimization module, ultimately forming a closed-loop collaborative process of "collection-cleaning-storage-optimization-assessment-mutual recognition-application-feedback".
[0088] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A carbon emission measurement data management method based on multi-source data, characterized in that, Includes the following steps: Collect basic data, business data, and rule data; construct a semantic association network through knowledge graph to achieve data fusion; and filter out effective data with quality labels through quality cleaning to form a structured data set. Based on a structured dataset, the initial carbon emission factor is determined by combining industry benchmarks and historical data. The factor is then dynamically adjusted using reinforcement learning with the measurement deviation as feedback. The optimized factor and adjustment log are generated and written to the blockchain for storage. Based on effective data and optimized factors, the value of data assets is calculated and classified by combining carbon price and emission reduction coefficient, forming a graded data asset with value label; The hierarchical data assets are standardized and metadata is attached. Cross-domain semantic consistency is verified based on semantic association network. After cross-validation and permission allocation, a trusted data flow is formed. High-value data from trusted data flows is pushed to the carbon trading platform, medium-value data is output to enterprise terminals, and all data is made available to regulatory interfaces, thus completing the closed-loop application of data.
2. The carbon emission measurement data management method based on multi-source data according to claim 1, characterized in that, The basic data is collected through an IoT interface and includes electrical parameters and environmental parameters. The electrical parameters include voltage, current, power and power flow direction, and the environmental parameters include temperature and humidity. The business data is collected through the carbon trading platform interface and includes carbon quotas, real-time carbon prices, and emission reduction coefficients. The rule data is collected from carbon policy texts through a policy parsing interface, and natural language processing technology is used to extract applicable industry and regional standards and accounting benchmark values. The entities in the semantic association network include data source identifiers, measurement object identifiers, policy document identifiers, industry types, and regional affiliations. The relationships in the semantic association network include the affiliation relationship between data sources and data types, the association relationship between data sources and data collection accuracy, the subordinate relationship between measurement objects and industry types, the geographical relationship between measurement objects and regional affiliations, the adaptation relationship between carbon policy constraints and applicable industries, and the correspondence relationship between carbon policy constraints and regional standards. A unified semantic label is added to the basic data through a semantic association network, where electrical parameters are associated with the energy consumption dimension label, environmental parameters are associated with the external impact dimension label, carbon quotas are associated with the policy constraint dimension label, and carbon prices are associated with the economic value dimension.
3. The carbon emission measurement data management method based on multi-source data according to claim 2, characterized in that, The quality cleaning is based on the comprehensive data quality index, which consists of three dimensions: completeness, accuracy, and timeliness. The completeness is calculated by the ratio of the number of valid data entries to the total number of data entries collected, and the ratio is required to be no less than 95%. For data that does not meet the completeness standard, linear interpolation of adjacent data points in the same time period is used to complete it. The accuracy is calculated by the deviation rate between the data to be cleaned and the data collected by the reference device. The deviation rate is required to be no more than 3%. For data that does not meet the accuracy standard, a second collection process of the corresponding collection device is triggered. If the second collection still does not meet the standard, it is marked as invalid data and discarded. The timeliness is calculated by the difference between the time of data collection completion and the time of transmission to the processing node. The difference is required to be no more than 10 seconds. Data that does not meet the timeliness standard will be downgraded to historical data storage and used only for trend analysis and will not participate in real-time carbon emission calculation and asset value assessment. After cleaning, each valid data entry is labeled with specific values for completeness, accuracy, and timeliness, forming a quality information label.
4. The carbon emission measurement data management method based on multi-source data according to claim 1, characterized in that, The specific steps for generating the optimized factors are as follows: The industry type, historical factors, and error data of the measurement object are retrieved from the structured dataset, and the initial carbon emission factor is calculated by weighted average method after matching the industry benchmark value. Based on the initial carbon emission factor, real-time carbon emission calculation is performed using effective data to generate real-time calculated values. The real-time calculated values are then compared with the actual verified values for the same period to calculate the deviation rate. Using the bias rate as the state input for reinforcement learning, the factor adjustment amount is set as the action, a reward function is constructed, the action value function is iteratively optimized through the Q-Learning algorithm, the optimal adjustment amount is output, and the optimized factor is calculated in combination with the initial carbon emission factor. Key data from the optimization process are extracted, a factor adjustment log is generated, and the log is transmitted to the blockchain notarization module for hash calculation and on-chain storage, thus completing the data loop of the entire factor optimization process.
5. The carbon emission measurement data management method based on multi-source data according to claim 1, characterized in that, The value of the data assets is calculated using a two-dimensional quantitative model. The first dimension is the basic value, and the second dimension is the added value. The total value is the sum of the basic value and the added value. The formula for calculating the basic value is: , in, Based on fundamental value, This refers to the comprehensive quality indicators corresponding to the quality label. For real-time carbon prices, Carbon emissions calculated based on optimized factors; Formula for calculating added value: , in, As added value, The emission reduction coefficients collected are determined based on the energy type used by the object being measured; The grading criteria for the value of the data assets are as follows: Data assets with a total value of ≥100,000 yuan are designated as Level 1 data assets and used for carbon trading settlement and cross-regional large-scale asset transfers. 10,000 yuan ≤ Total value Data assets valued at less than 100,000 yuan are designated as Level 2 data assets and used for enterprise emission reduction decisions and medium-term carbon management planning. Total value Data assets valued at less than 10,000 yuan are designated as Level 3 data assets for internal statistical analysis and regulatory filing.
6. The carbon emission measurement data management method based on multi-source data according to claim 1, characterized in that, The specific steps for forming a trusted data flow are as follows: Based on tiered data assets with value identifiers, the data structure is converted according to JSON-LD format, and metadata is added synchronously to form a standardized data asset stream. The metadata includes quality tags, optimized factors, data collection timestamps, and measurement object identifiers, and the metadata is bound to the tiered data assets through a unique identifier. Based on standardized data asset streams, the semantic association network is invoked to extract core fields from the standardized data asset streams. These fields are then compared with the unified semantic labels defined in the semantic association network to verify the consistency of cross-domain data in core concepts and generate a semantic consistency verification result stream. Based on the semantic consistency verification result stream, if the verification passes, the standardized data asset stream is distributed to three or more cross-domain nodes. Each node independently calculates the carbon emission data of the same measurement object based on local data, generates node calculation results, compares the node calculation result stream with the carbon emission in the standardized data asset stream, calculates the deviation rate, and forms a cross-validation result stream. Based on the cross-validation result stream, if the validation passes, access permissions are allocated according to the hierarchical results. Specifically, Level 1 data assets are open to carbon exchange nodes, enterprise authorized nodes, and regulatory nodes; Level 2 data assets are open to enterprise management nodes; and Level 3 data assets are open to internal statistical nodes. The permission configuration and data assets are bound to an encryption key to form a trusted data flow with permission identifiers.
7. The carbon emission measurement data management method based on multi-source data according to claim 1, characterized in that, The specific process of the data closed-loop application is as follows: Based on the trusted data flow, it is divided into a first-level data sub-flow, a second-level data sub-flow, and a full data traceability flow; The primary data sub-stream is transmitted to the carbon trading platform. Based on the sub-stream data, the carbon trading platform calculates the trading price, generates a settlement data stream, and feeds it back to the system, thus completing the closed loop of trading data. The secondary data substream is transmitted to the enterprise terminal, which uses a built-in algorithm to generate a correlation curve between the investment amount for equipment upgrades, the expected emission reduction, and the increase in asset value, forming an emission reduction decision recommendation stream for enterprises to implement upgrade plans. By opening up the full data traceability stream to the regulatory interface, regulators can call the traceability stream through the interface to trace back the entire chain of any data from collection to application, generate audit result streams and feed them back to the system, thus achieving a closed loop of regulatory data.
8. The carbon emission measurement data management method based on multi-source data according to claim 4, characterized in that, The feedback adjustment is achieved through closed-loop data application, and the specific steps are as follows: The system receives the audit result stream output from the regulatory interface. If the audit finds data quality issues, it extracts the problem data identifier, the ID of the data collection device involved, and the specific value of the deviation, generates a quality issue feedback stream, and transmits it to the data quality cleaning module. Based on the quality issue feedback stream, the data quality cleaning module adjusts the accuracy verification threshold of the corresponding data collection device or increases the sampling frequency of the device and optimizes the cleaning rules. The system receives the feedback stream of the transformation execution output from the enterprise terminal, transmits it to the reinforcement learning factor optimization module, calculates the deviation between the actual emission reduction and the predicted emission reduction in the decision suggestion stream, incorporates the deviation into the adjustment basis of the reward function, and iteratively optimizes the action value function of the Q-Learning algorithm. The system receives settlement feedback streams from the carbon trading platform and transmits them to the data asset valuation module. If the average deviation rate exceeds 10%, the weighting of the total value is adjusted or the standard for the emission reduction coefficient is corrected.
9. A carbon emission measurement data management system based on multi-source data, characterized in that, It includes modules for multi-source heterogeneous data acquisition and fusion, data quality cleaning, distributed asset pooling, factor optimization, blockchain notarization, value assessment, cross-domain mutual recognition, and regulatory auditing. These modules form a collaborative closed loop through hierarchical data flow and bidirectional feedback. The multi-source heterogeneous data acquisition and fusion module integrates IoT interfaces, carbon trading platform interfaces, and policy analysis interfaces. It is used to collect multi-source data, generate raw data streams with semantic tags through knowledge graphs, and connect to the data quality cleaning module through a wired communication link to transmit the raw data streams to the data quality cleaning module in real time. The data quality cleaning module receives the raw data stream, filters and completes the data according to completeness, accuracy and timeliness, generates a valid data stream with quality labels, and outputs it to the distributed asset pool module. The distributed asset pool module is used to store effective data streams and build multi-dimensional indexes, providing data query interfaces for the factor optimization module and the value assessment module; The factor optimization module is used to call data from the distributed asset pool module, calculate the initial factors, and output the optimized factors and adjustment logs through reinforcement learning, which are then output to the value assessment module and the blockchain evidence storage module, respectively. The blockchain evidence storage module is used to store the hash values of adjustment logs and cross-domain records, and provides an evidence storage query interface; The value assessment module is used to receive valid data streams and optimized factors, calculate the value of data assets and classify them, and output the classified asset streams with value labels to the cross-domain mutual recognition module. The cross-domain mutual recognition module is used to standardize the tiered asset flow and attach metadata, verify semantic consistency and data deviation, generate a trusted circulation data flow through access control, and output it to the carbon trading platform, enterprise terminal and regulatory interface respectively. The regulatory audit module is used to receive trusted data streams and blockchain-based evidence logs, provide a full-chain traceability interface, and report quality issues to the data quality cleaning module.
10. A carbon emission measurement data management system based on multi-source data according to claim 9, characterized in that, Also includes: The transaction support module is used to receive trusted circulation data streams and generate settlement data according to the pricing formula; The decision support module is used to receive secondary data from trusted data streams and generate emission reduction decision recommendations.
Citation Information
Patent Citations
Method for producing lipopeptide antibiotics by optimizing bacillus subtilis with response surface method
CN111304127A
Metering-based carbon emission accounting system
CN119443533A
Carbon emission data management method and system and readable storage medium
CN119886525A
Carbon emission digital management system based on cloud platform
CN120671999A
Industrial enterprise carbon emission metering method and system based on multi-source data fusion
CN121480922A