Carbon emission factor database data quality control method

By constructing a domain ontology of carbon emission factors and a dynamically coupled weight matrix, the problem of inconsistent formats and semantics of multi-source heterogeneous data in the data quality control of the carbon emission factor database was solved. Dynamic adaptation and closed-loop optimization of data quality were achieved, improving the accuracy and adaptability of the data and meeting the needs of different application scenarios.

CN121935239APending Publication Date: 2026-04-28CHINA IND INTERNET RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA IND INTERNET RES INST
Filing Date
2026-03-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing carbon emission factor database data quality control methods fail to fully consider the correlation effects between various quality dimensions, ignore the spatiotemporal validity of data, and lack an effective closed-loop optimization mechanism, resulting in a disconnect between quality assessment results and actual data characteristics, making it difficult to meet the differentiated needs of different industries and life cycle stages.

Method used

We construct an ontology for carbon emission factors, perform automatic format mapping and semantic alignment of multi-source heterogeneous data, calculate basic scores for four dimensions: consistency, commutativity, integrity, and transparency, construct a dynamically coupled weight matrix, combine comprehensive attenuation factors and scenario adaptation factors to calculate the comprehensive quality score of data entering the database, and optimize quality control parameters through closed-loop iteration.

Benefits of technology

It achieves the unification of data format and semantics, dynamically adjusts weights to reflect the correlation strength between dimensions, adapts to the spatiotemporal validity of data, improves the accuracy and adaptability of data quality, meets the needs of different application scenarios, and promotes the reliability and long-term availability of carbon accounting and related decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935239A_ABST
    Figure CN121935239A_ABST
Patent Text Reader

Abstract

The invention discloses a data quality control method for a carbon emission factor database, and relates to the technical field of carbon emission data quality control. When warehousing data information is called by a specific application scene of an application end, the warehousing comprehensive quality score of the data information and characteristic parameters of the specific application scene are calculated; and expanding the representativeness of the evaluation data, obtaining a representative score, calculating an accuracy score, fusing the warehousing comprehensive quality score, the normalized representative score and the accuracy score, and calculating an application end comprehensive quality score in combination with a scene sensitivity coefficient. According to the method, the problem that data formats and semantics are inconsistent is effectively solved, and the association strength and scene requirements of all quality dimensions are fully considered by constructing the dynamic coupling weight matrix; the evaluation deviation caused by the dimension difference is effectively eliminated by means of a collaborative correction coefficient; precise adaptation of data space-time validity is realized based on the comprehensive attenuation factor, so that the warehousing comprehensive quality score can truly reflect the inherent quality of the data and the actual application adaptation potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon emission data quality control technology, and in particular to a method for controlling the data quality of a carbon emission factor database. Background Technology

[0002] Carbon emission factor databases are the core foundation for carbon accounting, carbon management, and related decision-making, and their data quality directly affects the accuracy and reliability of carbon accounting results. With the diversification of carbon accounting needs, the sources of carbon emission factor data are becoming increasingly complex, covering various types such as enterprise measurement, industry statistics, and literature reports, resulting in prominent issues such as heterogeneous data formats and semantic ambiguity.

[0003] Existing data quality control methods often use fixed weights to independently evaluate a single dimension, failing to fully consider the correlation effects between various quality dimensions, resulting in a disconnect between quality assessment results and actual data characteristics.

[0004] Meanwhile, existing methods lack mechanisms to adapt to the spatiotemporal validity of data, ignoring quality changes caused by data decay over time and geographical differences. Furthermore, the quality assessment is poorly matched to application scenarios, making it difficult to meet the differentiated needs of different industries and lifecycle stages. In addition, existing technologies generally lack effective closed-loop optimization mechanisms, failing to continuously improve the accuracy of data quality control based on application feedback, thus limiting the long-term availability and application value of carbon emission factor databases. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a method for controlling the data quality of a carbon emission factor database. The technical solution adopted is as follows: A method for controlling the data quality of a carbon emission factor database includes the following steps: Step 1: Construct an ontology for carbon emission factors, perform automatic format mapping and semantic alignment on multi-source heterogeneous data, and clean up missing and outlier values ​​to generate standardized data; Step 2: Calculate the base scores for the four dimensions of consistency, exchangeability, integrity, and transparency; construct a dynamic coupling weight matrix; correct the base scores for each dimension based on the collaborative correction coefficient to obtain the corrected scores; and calculate the overall quality score for data entry by combining the comprehensive attenuation factor and the scenario adaptation factor. Step 3: Set differentiated qualification thresholds according to the data source type, conduct quality rating based on the comprehensive quality score of the incoming goods, execute the hierarchical incoming goods review process, and generate complete data information, which includes the comprehensive quality score of the incoming goods and rating information. Step 4: When the data information is called by the application terminal in a specific application scenario, based on the comprehensive quality score of the data information and the characteristic parameters of the specific application scenario, the representativeness of the evaluation data is expanded and a representativeness score is obtained. The accuracy score is calculated, and the comprehensive quality score of the data information, the normalized representativeness score and the accuracy score are merged. The application terminal comprehensive quality score is calculated in combination with the scenario sensitivity coefficient. Step 5: Based on the feedback from the overall quality score of the incoming goods and the overall quality score of the application, iteratively optimize the quality control parameters.

[0006] Optionally, in step 1, the carbon emission factor ontology includes material list class, process boundary class, indicator system class, and traceability information class. Semantic mapping relationship is established through the material process indicator ternary association model. The format automatic mapping adopts the semantic similarity matching method to align the original data fields with the ontology elements. Automatic mapping is completed when the semantic similarity meets the preset threshold; otherwise, manual assistance or ontology expansion process is triggered.

[0007] Optionally, in step 2, the calculation logic for the basic scores of the four dimensions is as follows: The consistency baseline score is evaluated by integrating format consistency and semantic consistency. The exchangeability baseline score assesses the compatibility of the data format with mainstream carbon accounting tools and the resolvability of the data itself. The basic score for completeness is evaluated by combining the rate of missing necessary information with the rate of completeness of optional information; The basic score for transparency is assessed based on the completeness of the data traceability chain and the traceability of the traceability information.

[0008] Optionally, in step 2, the method for constructing the dynamic coupling weight matrix is ​​as follows: first, calculate the initial coupling weight based on the mutual information entropy between dimensions to characterize the correlation strength between dimensions; then, adjust the initial coupling weight according to the needs of specific application scenarios; and finally, perform normalization processing to obtain the final dynamic coupling weight.

[0009] Optionally, in step 2, the collaborative correction coefficient is calculated based on the degree of difference in the base scores of each dimension; the greater the difference, the higher the correction strength. The formula for calculating the comprehensive attenuation factor is: ;in This is a time decay factor calculated based on the time difference between the year the data was collected and the current year. This is a geographical attenuation factor calculated based on the geographical differences between the data collection area and the application area.

[0010] Optionally, in step 2, the overall quality score of the incoming goods is calculated. The calculation formula is: ;in The scores are adjusted based on four dimensions: consistency, commutability, integrity, and transparency. For each dimension's basic weights determined through the entropy weight method, F is a comprehensive attenuation factor ranging from 0 to 1, and S is a scenario adaptation factor ranging from 0 to 1, determined based on industry adaptation, lifecycle stage adaptation, and accounting purpose adaptation.

[0011] Optionally, in step 3, the data source types include four categories: enterprise metrological data, industry statistical data, literature report data, and data from unknown sources; the differentiated qualification threshold is the product of the basic qualification threshold of 0.7 and the verification strength coefficient, with the verification strength coefficients set to 1.0, 0.85, 0.7, and 0.55 for the four types of data, respectively; the quality rating is a five-star rating, and the rating method is as follows: like It is five stars, if It is a four-star rating, if For Samsung, if For two stars, if It is rated as one star.

[0012] Optionally, in step 4, the representative score R is calculated by evaluating five categories of indicators: technical relevance, data source reliability, time relevance, geographical relevance, and scenario matching degree, and by using a combination of subjective weights, objective weights, and scenario weights; the accuracy score A is calculated by Monte Carlo simulation, and the initial uncertainty of the data is corrected by using the comprehensive quality score of the data entering the database during the simulation process.

[0013] Optionally, in step 4, the overall quality score of the application is calculated. ; in , , The weights for the entry score, normalized representativeness score, and accuracy score are 0.4, 0.3, and 0.3, respectively. The sensitivity coefficient is a weighted sum of the parameter sensitivity coefficient and the hypothesis sensitivity coefficient, with weights of 0.6 and 0.4, respectively. The sensitivity coefficient is 0.2.

[0014] Optionally, in step 5, the closed-loop iterative optimization method is: periodically update the dynamic coupling weight, decay factor, and scene weight based on user feedback and application adaptation results; if the application's overall quality score... Data that is below the set first score threshold and cannot be corrected is discarded. The overall quality score of the application is... After supplementing the data (data greater than or equal to the first scoring threshold but less than the second scoring threshold), the overall quality score for data entry will be recalculated. Overall quality score of application The iteration effect is achieved through Average improvement rate The accuracy of the compatibility rate and the rate of decrease in user feedback were verified.

[0015] In summary, the present invention has at least one of the following beneficial technical effects: This invention provides a method for controlling the data quality of a carbon emission factor database. By constructing a domain ontology, it achieves standardized processing of multi-source heterogeneous data, solving the problem of inconsistent data formats and semantics, and laying a solid foundation for quality assessment. The construction of a dynamically coupled weight matrix fully considers the correlation strength of each quality dimension and scenario requirements. The collaborative correction coefficient effectively eliminates the assessment bias caused by dimensional differences, and the comprehensive attenuation factor achieves accurate adaptation of the data's spatiotemporal validity, enabling the overall quality score of the database to truly reflect the inherent quality of the data and its potential for practical application.

[0016] A differentiated data entry mechanism sets reasonable thresholds and review processes based on the characteristics of the data source, ensuring both data quality and entry efficiency. The application-side comprehensive quality score integrates entry quality, representativeness, and accuracy, combined with a scenario sensitivity coefficient, achieving a deep binding between quality assessment and specific application scenarios, thus improving data application adaptability. A closed-loop iterative optimization mechanism continuously adjusts quality control parameters based on application feedback, driving dynamic upgrades to data quality and control methods. This significantly improves the reliability, adaptability, and long-term availability of the carbon emission factor database, providing high-quality data support for carbon accounting and related decision-making. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a method for controlling the data quality of a carbon emission factor database according to the present invention. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to the accompanying drawings.

[0019] This invention discloses a method for controlling the data quality of a carbon emission factor database.

[0020] Reference Figure 1 Example 1: A method for controlling the data quality of a carbon emission factor database, comprising the following steps: Step 1: Construct an ontology for carbon emission factors, perform automatic format mapping and semantic alignment on multi-source heterogeneous data, and clean up missing and outlier values ​​to generate standardized data; Step 2: Calculate the base scores for the four dimensions of consistency, exchangeability, integrity, and transparency; construct a dynamic coupling weight matrix; correct the base scores for each dimension based on the collaborative correction coefficient to obtain the corrected scores; and calculate the overall quality score for data entry by combining the comprehensive attenuation factor and the scenario adaptation factor. Step 3: Set differentiated qualification thresholds according to the data source type, conduct quality rating based on the comprehensive quality score of the incoming goods, execute the hierarchical incoming goods review process, and generate complete data information, which includes the comprehensive quality score of the incoming goods and rating information. Step 4: When the data information is called by the application terminal in a specific application scenario, based on the comprehensive quality score of the data information and the characteristic parameters of the specific application scenario, the representativeness of the evaluation data is expanded and a representativeness score is obtained. The accuracy score is calculated, and the comprehensive quality score of the data information, the normalized representativeness score and the accuracy score are merged. The application terminal comprehensive quality score is calculated in combination with the scenario sensitivity coefficient. Step 5: Based on the feedback from the overall quality score of the incoming goods and the overall quality score of the application, iteratively optimize the quality control parameters.

[0021] In Example 2, in step 1, the carbon emission factor ontology includes a material list class, a process boundary class, an indicator system class, and a traceability information class. A semantic mapping relationship is established through a ternary association model of material process indicators. The automatic format mapping adopts a semantic similarity matching method to align the original data fields with the ontology elements. Automatic mapping is completed when the semantic similarity meets the preset threshold; otherwise, manual assistance or ontology expansion process is triggered.

[0022] Example 3, in step 2, the calculation logic for the basic scores of the four dimensions is as follows: The consistency baseline score is evaluated by integrating format consistency and semantic consistency. The exchangeability baseline score assesses the compatibility of the data format with mainstream carbon accounting tools and the resolvability of the data itself. The basic score for completeness is evaluated by combining the rate of missing necessary information with the rate of completeness of optional information; The basic score for transparency is assessed based on the completeness of the data traceability chain and the traceability of the traceability information.

[0023] By adopting the above technical solutions, the primary challenge in quality control of multi-source heterogeneous data is addressing the inconsistency in format and semantics. The construction of an ontology for carbon emission factors essentially establishes a unified semantic framework, integrating scattered material information, process boundaries, indicator specifications, and traceability elements. A ternary correlation model of material process indicators is used to build mapping relationships between these elements, providing a unified alignment benchmark for data from different sources. The application of semantic similarity matching methods, based on the characteristics of ontology elements, automatically maps raw data fields to standard elements, reducing manual intervention while ensuring mapping accuracy. The setting of preset thresholds balances automatic processing efficiency with mapping accuracy; exceeding the threshold triggers manual assistance or ontology expansion to ensure no omissions in the adaptation. The core principle of missing and outlier cleaning is to remove data noise, preventing invalid or erroneous data from interfering with subsequent quality assessments and providing reliable basic data for quality quantification.

[0024] The core characteristics of data quality are reflected in four dimensions: consistency, interoperability, integrity, and transparency. Each dimension corresponds to different levels of data quality requirements. Consistency assessment focuses on the uniformity of data format and semantics, ensuring data interoperability in different scenarios; interoperability assessment ensures data is compatible with mainstream accounting tools, improving data reusability; integrity assessment covers necessary and optional information, ensuring data meets basic application requirements while achieving sufficient richness; transparency assessment, through traceability and verifiability, ensures data credibility and verifiability. The basic scores of these four dimensions provide a preliminary quantification of data quality.

[0025] The core of constructing a dynamically coupled weight matrix is ​​to overcome the rigid limitations of fixed weights. Mutual information entropy is used to capture the inherent correlation strength between dimensions, avoiding biases caused by isolated evaluations. Weights are adjusted according to scenario requirements because the importance of each dimension varies across different application scenarios, making the weight allocation more aligned with actual usage needs. Normalization ensures the rationality and computability of the weights. The principle of collaborative correction considers the potential for extreme differences in scores across dimensions; the greater the difference, the more significant the interference with the total score. Adjusting the correction strength based on the degree of difference balances the impact of each dimension, allowing the corrected score to better reflect the overall quality of the data.

[0026] The introduction of a comprehensive attenuation factor is based on the objective law that the actual utility of data quality changes with time and geographical differences. The longer the time, the weaker the timeliness of the data; the greater the geographical differences, the lower the data's adaptability. By multiplying the time attenuation factor and the geographical attenuation factor, this degree of utility decline is quantified, making the data entry score more closely match the actual application value of the data. The scenario adaptation factor, on the other hand, incorporates the adaptation requirements of the application scenario in advance on top of the data's inherent quality. This ensures that the data entry score not only reflects the inherent quality of the data but also predicts its basic adaptability potential in different scenarios, building a bridge for subsequent application-side evaluation.

[0027] Data from different sources inherently differs in reliability. Enterprise measurement data, based on actual measurements, has the highest reliability, while data from unknown sources lacks clear traceability and has the lowest reliability. Setting differentiated acceptance thresholds based on this difference allows for rigorous control over high-reliability data, ensuring core database quality, while also setting reasonable standards for low-reliability data, avoiding data resource waste due to over-screening. The principle of quality rating is to intuitively quantify data quality, facilitating rapid identification of data quality levels by subsequent applications. The tiered entry review process matches different levels of rigor to the data source and quality level, improving entry efficiency while ensuring quality. Generating complete data information is crucial for retaining key data quality data, providing a foundation for subsequent application evaluation, data traceability, and iterative optimization.

[0028] The overall quality score for data entry is a quality assessment result under general scenarios. However, specific scenario characteristics exist at the application level, thus requiring further optimization of the assessment. The extended assessment of representative scores, based on the characteristic parameters of the application scenario, supplements indicators such as scenario matching degree, making the assessment dimensions more aligned with specific application needs. The combined weighting method combines subjective judgment, objective data characteristics, and scenario requirements to ensure the comprehensiveness and rationality of the representative score. The calculation of the accuracy score, combined with the overall quality score for data entry, corrects for initial uncertainty because the data entry score already reflects the inherent quality of the data; higher quality data has lower uncertainty. Monte Carlo simulation can quantify this accuracy, making the assessment results more convincing.

[0029] The core of the integrated calculation of application-side comprehensive quality scores is to combine general data quality with scenario-specific requirements. The overall quality score upon data entry ensures basic data reliability, the normalized representativeness score ensures the data's suitability for the scenario, and the accuracy score ensures the precision of the application results. These three scores are integrated with appropriate weights, and then the impact of parameters and assumption sensitivity on application quality is corrected through a scenario sensitivity coefficient. Ultimately, a quality assessment result tailored to the specific application scenario is obtained, resolving the disconnect between general quality scores and scenario requirements.

[0030] Data quality control is not a static process; feedback from the application end can directly reflect the rationality of the initial quality control parameters. Based on feedback from the overall quality score of the database and the overall quality score of the application end, the dynamic coupling weights, attenuation factors, and scenario weights are updated regularly. This process reverses the actual application effects on the quality control system, making the control parameters more aligned with data characteristics and application needs. Data with different application quality scores is processed by removing or supplementing information to continuously optimize the quality of the database inventory and avoid low-quality data consuming resources. Three verification indicators for the iteration effect quantify the optimization results from three dimensions: improvement in inherent data quality, improvement in application adaptation, and optimization of user experience, forming a closed loop to ensure continuous upgrading of data quality and control methods to adapt to the ever-changing needs of carbon accounting and carbon management.

[0031] In step 2, the calculation logic and specific acquisition methods for the basic scores of the four dimensions are as follows: Consistency Baseline Score The evaluation combines format consistency and semantic consistency, with the specific calculation method as follows: The basic consistency score equals the format consistency score multiplied by 0.5 plus the semantic consistency score multiplied by 0.5. The format consistency score is equal to the number of standardized field matches divided by the total number of fields; the semantic consistency score is equal to the number of semantically matched entries divided by the total number of entries, with values ​​ranging from 0 to 1.

[0032] Exchangeability Basic Score The compatibility of mainstream carbon accounting tools and the parsability of the data itself are evaluated based on the data format. Specific calculation methods are as follows: The basic score for commutativity equals the tool compatibility score multiplied by 0.6 plus the resolvability score multiplied by 0.4. The tool compatibility score is equal to the number of compatible mainstream carbon accounting tools / the total number of mainstream carbon accounting tools; the parsable score is equal to the number of data entries without formatting errors / the total number of data entries, with values ​​ranging from 0 to 1.

[0033] Integrity Baseline Score The evaluation combines the rate of missing necessary information with the rate of complete optional information. The specific calculation method is as follows: The baseline completeness score equals (1 minus the required information missing rate) multiplied by 0.7, plus the optional information completeness rate multiplied by 0.3. Among them, the missing rate of necessary information is equal to the number of missing fields of necessary information / the total number of fields of necessary information; the completeness rate of optional information is equal to the number of fields filled in optional information / the total number of fields of optional information, and the value of both is 0-1.

[0034] Transparency Baseline Score An assessment is conducted based on the integrity of the data traceability chain and the traceability of the traceability information. The specific calculation method is as follows: The basic score for transparency equals the traceability chain integrity score multiplied by 0.6, plus the traceability score multiplied by 0.4. Among them, the traceability chain integrity score is equal to the number of complete traceability chain entries / the total number of entries; the traceability score is equal to the number of locatable traceability information entries / the total number of entries, and the values ​​are both 0-1.

[0035] In Example 4, step 2, the method for constructing the dynamic coupling weight matrix is ​​as follows: first, the initial coupling weight is calculated based on the mutual information entropy between dimensions to characterize the correlation strength between dimensions; then, the initial coupling weight is adjusted according to the needs of specific application scenarios; and finally, normalization is performed to obtain the final dynamic coupling weight.

[0036] In Example 5, step 2, the collaborative correction coefficient is calculated based on the degree of difference in the basic scores of each dimension; the greater the difference, the higher the correction strength. The formula for calculating the comprehensive attenuation factor is: ;in The time decay factor is calculated based on the time difference between the data collection year and the current year. The calculation formula is as follows: , This represents the difference between the current year and the year the data was collected. This is the time decay coefficient, with a value of 0.05; This is a geographic attenuation factor calculated based on the geographic differences between the data collection area and the application area. The specific calculation method is as follows: Geographical distance is quantified according to differences in administrative regional levels, and the calculation formula is: In the formula: d is the geographical difference coefficient: d=0 for the same county-level region, d=0.2 for different counties in the same city, d=0.4 for different cities in the same province, d=0.6 for different provinces, and d=0.8 for cross-provincial regions; The geographical attenuation coefficient is set to 0.5; if the calculated result is less than 0, it is set to 0, and if it is greater than 1, it is set to 1.

[0037] Example 6, in step 2, the overall quality score of the incoming goods. The calculation formula is: ;in The scores are adjusted based on four dimensions: consistency, commutability, integrity, and transparency. For each dimension's basic weights determined through the entropy weight method, F is a comprehensive attenuation factor ranging from 0 to 1, and S is a scenario adaptation factor ranging from 0 to 1, determined based on industry adaptation, lifecycle stage adaptation, and accounting purpose adaptation.

[0038] By adopting the above technical solutions, the four core dimensions of data quality are not isolated but intrinsically related, and the strength of this relationship directly affects the accuracy of the overall quality assessment. Mutual information entropy, as an effective means of quantifying the degree of correlation between variables, can objectively capture the mutual influence strength between consistency, exchangeability, integrity, and transparency. The initial coupling weights calculated based on this are essentially a precise characterization of the inherent relationships between each dimension, avoiding the limitations of traditional fixed weights that ignore the interactions between dimensions.

[0039] Different application scenarios have varying requirements for data quality across different dimensions. For example, industrial carbon accounting scenarios may place greater emphasis on consistency dimensions related to data integrity and accuracy, while cross-regional carbon trading scenarios focus more on transparency dimensions related to data exchangeability and geographical adaptability. Therefore, adjustments are made to the initial coupled weights based on specific scenario requirements, ensuring that the weight allocation aligns with actual application needs and avoiding a disconnect between evaluation standards and usage scenarios. The core purpose of normalization is to guarantee the rationality and computability of the weight system, ensuring that the sum of the weights for each dimension satisfies the mathematical logic of quantitative evaluation, and guaranteeing the standardization and accuracy of subsequent comprehensive score calculations.

[0040] In multi-dimensional quality assessment, the base scores of each dimension may exhibit extreme differences. If these scores are directly used for comprehensive calculation, the extreme performance of a particular dimension may excessively dominate the total score, failing to accurately reflect the overall quality level of the data. The collaborative correction coefficient is designed based on the degree of difference in scores between dimensions; the greater the difference, the stronger the correction. Its core principle is to dynamically adjust and balance the contribution of each dimension to the total score, weakening the interference of extreme values, so that the corrected score better reflects the comprehensive balance of data quality, and avoiding the overemphasis on the shortcomings or advantages of a single dimension.

[0041] The practical application value of data diminishes over time and with changes in geographical scope; this is an inherent attribute of carbon emission factor data. The longer the time frame, the more likely changes in technology, energy structure, and other factors will reduce the data's timeliness. Greater geographical differences, such as variations in regional resource endowments and industrial characteristics, will reduce the data's adaptability. The comprehensive attenuation factor, designed by multiplying a time attenuation factor and a geographical attenuation factor, comprehensively quantifies this dual attenuation effect: the time attenuation factor... By calculating the time difference between the data collection year and the current year, the timeliness loss of data can be intuitively reflected; geographical attenuation factor. Based on the geographical differences between the data collection area and the application area, the loss of data geographic adaptability is quantified. The comprehensive attenuation factor formed by the product of the two factors can comprehensively measure the degree of decline in data quality and utility in the spatiotemporal dimensions, making the overall quality score of the data more consistent with its actual application value.

[0042] The core of the formula for calculating the overall quality score of data entry is to achieve comprehensive and accurate quantification of data quality across multiple dimensions and factors. The combination logic of each parameter revolves around objectively reflecting the inherent quality of the data, adapting to spatiotemporal effects, and predicting the potential of scenarios.

[0043] Corrected score This is a dimensional quality quantification result after collaborative correction, to balance the interference of differences between dimensions, and to truly reflect the actual quality level of each dimension; dimensional base weights. The entropy weight method is optimized and determined. Its core advantage is that it automatically assigns weights based on the distribution characteristics of the data itself, avoiding the bias of subjective human setting and making the importance allocation of each dimension more in line with the objective characteristics of the data.

[0044] The comprehensive attenuation factor F incorporates the spatiotemporal attenuation effect, while the scenario adaptation factor S takes into account the degree of data's compatibility with the industry, life cycle stage, and accounting purpose. The two factors are weighted and multiplied by the dimensional scores, which essentially adds the effects of spatiotemporal effectiveness and scenario adaptation potential to the inherent quality of the data.

[0045] The entire formula organically integrates the core influencing factors through multiplicative logic, ultimately yielding a comprehensive quality score for inventory entry. It comprehensively covers the core dimensions of data quality, fully considers the impact of spatiotemporal decay on data value, and predicts the basic adaptability of data in different scenarios, providing accurate and comprehensive quantitative basis for subsequent differentiated data entry control and application-side scenario-based evaluation.

[0046] In Example 7, step 3, the data source types include four categories: enterprise metrological data, industry statistical data, literature report data, and data from unknown sources; the differentiated qualification threshold is the product of the basic qualification threshold of 0.7 and the verification strength coefficient, and the verification strength coefficient is set to 1.0, 0.85, 0.7, and 0.55 for the four types of data, respectively; the quality rating is a five-star rating, and the rating method is as follows: like It is five stars, if It is a four-star rating, if For Samsung, if For two stars, if It is rated as one star.

[0047] By adopting the above technical solutions, the generation scenarios, collection methods, and traceability capabilities of carbon emission factor data differ fundamentally, directly determining the inherent reliability level of the data. Enterprise metrological data originates from real-time monitoring or precise measurement in actual production processes; the data generation process is standardized and controllable, the traceability chain is complete, and its reliability is at the highest level. Industry statistical data is compiled and organized by professional institutions or competent authorities based on data from multiple channels, undergoing systematic review and verification, and its reliability is second highest. Literature report data relies on academic research or special surveys; although it has undergone some academic verification, its reliability is at a medium level due to limitations such as research scope and experimental conditions. Data from unknown sources lacks a clear generation background, collection methods, and traceability information, making it impossible to verify the rigor of its generation process, and thus its reliability is the lowest. Dividing the data into four categories is fundamentally about accurately identifying the differences in the quality foundation of data from different sources, providing a basis for subsequently developing targeted quality control standards, and avoiding the over-review of high-quality data or the disorderly influx of low-quality data due to a single standard.

[0048] The basic qualification threshold serves as the baseline standard for ensuring the overall quality of the carbon emission factor database, while the verification intensity coefficient is set to closely align with the reliability differences of data from different sources. For the most reliable enterprise measurement data, the highest verification intensity coefficient is set to ensure its qualification threshold remains consistent with the basic qualification threshold, guaranteeing the purity of core high-quality data through stringent standards. The verification intensity coefficients for industry statistical data and literature reports decrease sequentially, with corresponding qualification thresholds appropriately relaxed. This avoids the loss of valuable, moderately reliable data due to overly stringent standards while preventing low-quality data from entering the database through threshold constraints. The verification intensity coefficient for data from unknown sources is the lowest, with a correspondingly lower qualification threshold. While adhering to the quality baseline, this allows potentially valuable data with missing traceability to be included in the database, achieving a balance between data resource utilization and quality control. This design logic ensures a precise match between quality standards and the inherent reliability of the data, guaranteeing the high quality of the database's core data while maximizing the application value of various data types.

[0049] While the overall quality score for data entry has quantified data quality, a clear and concise rating system is needed to meet the requirements of refined database management and efficient application. The five-star rating system categorizes data based on the overall quality score range, with higher scores corresponding to higher ratings, directly reflecting the degree of data quality. High-rated data indicates balanced and excellent performance across dimensions such as consistency, exchangeability, integrity, and transparency, demonstrating strong temporal and spatial adaptability and scenario potential. It can be prioritized for applications with stringent data quality requirements, such as carbon accounting and carbon trading. Low-rated data indicates weaknesses in certain quality dimensions and should be used cautiously, considering the tolerance margin of specific application scenarios. This rating system simplifies the data quality identification process, allowing users to quickly select suitable data. It also provides a clear basis for subsequent data review, upgrades, and differentiated management, improving database management efficiency and ease of use.

[0050] In Example 8, in step 4, the representative score R is calculated by evaluating five categories of indicators: technical relevance, data source reliability, time relevance, geographical relevance, and scene matching degree, and by using a weighting method that combines subjective weight, objective weight, and scene weight; the accuracy score A is calculated by Monte Carlo simulation, and the initial uncertainty of the data is corrected by using the comprehensive quality score of the data entering the database during the simulation process.

[0051] Example 9, in step 4, the overall quality score of the application terminal. ; in , , The weights for the entry score, normalized representativeness score, and accuracy score are 0.4, 0.3, and 0.3, respectively. The sensitivity coefficient is a weighted sum of the parameter sensitivity coefficient and the hypothesis sensitivity coefficient, with weights of 0.6 and 0.4, respectively. The sensitivity coefficient is 0.2.

[0052] Sscn is a weighted sum of the parameter sensitivity coefficient and the assumed sensitivity coefficient, with weights of 0.6 and 0.4, respectively. The parameter sensitivity coefficient is obtained by using a first-order local sensitivity analysis method to calculate the relative rate of change of the results when the core parameters (carbon emission factor, activity level data) fluctuate by ±10%. This rate of change is the parameter sensitivity coefficient. , The calculation formula is as follows: The value ranges from 0 to 1. This is the change in the accounting results. This is the original accounting result. These are parameter variation values. It is the original parameter value.

[0053] Assumption sensitivity coefficient acquisition method: Based on three core assumptions—accounting boundary, emission type, and calculation model—three types of assumptions are set: conservative, baseline, and relaxed. The relative dispersion of the accounting results under different assumptions is calculated, which is the assumption sensitivity coefficient. , The calculation formula is as follows: The value range is 0-1. It is the maximum value of the calculation result. It is the minimum value of the calculation result. It is the baseline value for the accounting results.

[0054] Sensitivity coefficient The final calculation formula is: ; The value range is 0-1.

[0055] By adopting the above technical solutions, the core of data adaptability in applications depends on its degree of matching with specific scenarios, and this degree of matching needs to be comprehensively characterized by multi-dimensional indicators. Technical relevance directly relates to the fit between the production technology corresponding to the data and the application scenario technology, which is the foundation for whether the data can effectively support scenario applications; the reliability of the data source determines the credibility of the data itself, which is a prerequisite for ensuring application effectiveness; temporal relevance and geographical relevance correspond to the timeliness and regional adaptability of the data, respectively, avoiding application deviations due to spatial and temporal differences; scenario matching focuses on the alignment between the data and the core needs of the application scenario, making up for the shortcomings of traditional representative assessments that are disconnected from specific scenarios.

[0056] The core design principle of the combined weighting method is to balance the advantages and limitations of different weights. Subjective weights, based on the experience of industry experts, can fully consider implicit needs and industry characteristics in practice, avoiding the one-sidedness of purely data-driven approaches. Objective weights are calculated based on the distribution characteristics and correlation patterns of the data itself, eliminating the interference of human subjective bias and ensuring the objectivity of weight allocation. Scenario weights are closely related to the core needs of specific application scenarios, making the weight allocation more in line with actual use cases and improving the suitability of representative scores to application requirements.

[0057] Monte Carlo simulation is a mature method for quantifying data uncertainty and calculating accuracy. Its core advantage lies in its ability to objectively reflect the fluctuation range of data application results by simulating data distribution through extensive random sampling. However, the initial uncertainty setting of data often lacks a connection with the inherent quality of the data, leading to inaccurate accuracy assessments. The comprehensive quality score for data entry fully quantifies the quality level of data in dimensions such as consistency, exchangeability, completeness, and transparency. Higher quality data generally indicates stronger standardization in its original collection and processing, and thus lower inherent uncertainty. Therefore, using the comprehensive quality score to correct initial uncertainty essentially establishes a logical connection between data quality and uncertainty, ensuring that accuracy assessment is no longer isolated but rather resonates with the overall quality control process. This allows the accuracy score to more accurately reflect the precision of data application results, enhancing the rationality and credibility of the assessment.

[0058] The overall quality score upon data entry is the basic quality quantification result after the data has undergone full-process control. It is the core guarantee of application quality and is therefore given the highest weight. The normalized representative score directly reflects the degree of data suitability to specific scenarios, while the accuracy score reflects the precision of data application results. Together, they determine the actual application value of data in the scenario and are therefore given considerable weight. The weighted sum of the three scores achieves a comprehensive integration of the core data quality dimensions and ensures the completeness of the evaluation.

[0059] The parameter sensitivity coefficient and hypothesis sensitivity coefficient quantify the impact of core parameter fluctuations and changes in modeling assumptions on application results, respectively. The sensitivity coefficient, obtained by weighting these two coefficients with appropriate weights, comprehensively reflects the combined impact of various sensitive factors in the scenario. The sensitivity effect coefficient is used to adjust the intensity of the influence of sensitive factors, preventing excessive sensitivity from dominating the total score and balancing the relationship between core quality dimensions and sensitive factors. The denominator incorporates the negative impact of sensitive factors into the calculation through a 1 plus sensitivity effect structure. This ensures that the overall application quality score not only reflects the quality and suitability of the data itself but also the risks that sensitive factors may bring in the scenario, making the score more practically instructive and providing a comprehensive and accurate decision-making basis for application data selection.

[0060] In Example 10, step 5, the closed-loop iterative optimization method is: periodically update the dynamic coupling weight, decay factor, and scene weight based on user feedback and application adaptation results; if the application's overall quality score... Data that is below the set first score threshold and cannot be corrected is discarded. The overall quality score of the application is... After supplementing the data (data greater than or equal to the first scoring threshold but less than the second scoring threshold), the overall quality score for data entry will be recalculated. Overall quality score of application The iteration effect is achieved through Average improvement rate The accuracy of the compatibility rate and the rate of decrease in user feedback were verified.

[0061] By adopting the above technical solution, the first scoring threshold is 0.55, and the second scoring threshold is 0.7. The overall quality score of the application end intuitively reflects the actual value of the data in a specific scenario. Differentiated processing based on the score range is the core of optimizing the database inventory and data quality to achieve efficient resource utilization. The first scoring threshold is set as the bottom line for the application value of data. Data below this threshold that cannot be corrected has quality shortcomings that exceed the scope of optimization. Continuing to retain it will affect the overall availability of the database and occupy storage and computing resources. Removing it can purify the database environment. Data between the first and second scoring thresholds, although it has quality shortcomings, has the potential to be supplemented and improved. By supplementing key information to make up for the shortcomings, the overall quality score of the database and application end is recalculated, which can make this type of data realize its application value, avoid the waste of potential data, and achieve the dual goals of improving data quality and saving resources.

[0062] The following specific embodiments illustrate the implementation principle of the present invention: Taking the quality control of carbon emission factor data in the thermal power industry of a certain province as an example, this paper details the actual application process of this technical solution.

[0063] Step 1: Preprocessing and adaptation of multi-source heterogeneous data; An ontology for carbon emission factors was constructed, encompassing material lists, process boundaries, indicator systems, and source information. Semantic mapping relationships were established using a ternary association model of material process indicators. Multi-source heterogeneous data were collected and standardized; the results are shown in Table 1. Table 1

[0064] Step 2: Multi-dimensional dynamic coupling quality verification; The base score and the revised score are shown in Table 2: Table 2

[0065] Key parameters and overall quality score of incoming goods: The basic weights for each dimension were determined using the entropy weighting method as follows: consistency 0.32, commutativity 0.24, integrity 0.26, and transparency 0.18. Other key parameters are... The calculation results are shown in Table 3: Table 3

[0066] Step 3, differentiated tiered warehousing control; The differentiated pass thresholds and quality rating results are shown in Table 4: Table 4

[0067] Enterprise measurement data is directly entered into the database after automatic system verification; industry statistical data is directly entered into the database after automatic system verification; literature and report data is entered into the database after being reviewed and approved by one power industry expert; data from unknown sources is entered into the database after being cross-reviewed by two experts and marked with a high uncertainty label. All data generates a data entry system containing... Complete files containing information such as ratings.

[0068] Step 4, Dynamic quality assessment of the application; A carbon trading institution used data for regional carbon trading accounting; the relevant assessment results are shown in Table 5. Table 5

[0069] Calculation in progress Equal to 0.4 Equal to 0.3 λ equals 0.3, λ equals 0.2. It is a weighted sum of the parameter sensitivity coefficient and the hypothesis sensitivity coefficient.

[0070] Step 5: Closed-loop iterative optimization; Data processing results: The first scoring threshold is 0.55, and the second scoring threshold is 0.7. (Data from unknown source) A score of 0.54 is below the first score threshold and cannot be corrected; therefore, it is discarded. (Literature report data) The value is 0.70, which falls between the two thresholds. After supplementing the experimental conditions, sample range, and other information, the result was recalculated as follows: Data type: Document report data, recalculated , ; Iteration effect verification: After iteration, the data in the database The average improvement rate was 6%. The high-quality adaptation rate was 35%, and the user feedback rate decreased by 22%, meeting the optimization goals.

[0071] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for controlling the data quality of a carbon emission factor database, characterized in that, Includes the following steps: Step 1: Construct an ontology for carbon emission factors, perform automatic format mapping and semantic alignment on multi-source heterogeneous data, and clean up missing and outlier values ​​to generate standardized data; Step 2: Calculate the base scores for the four dimensions: consistency, commutativity, integrity, and transparency. Construct a dynamic coupling weight matrix; The base scores for each dimension are corrected based on the collaborative correction coefficient to obtain the corrected scores. The overall quality score for data entry is calculated by combining the comprehensive attenuation factor and the scenario adaptation factor. Step 3: Set differentiated qualification thresholds according to the data source type, conduct quality rating based on the comprehensive quality score of the incoming goods, execute the hierarchical incoming goods review process, and generate complete data information, which includes the comprehensive quality score of the incoming goods and rating information. Step 4: When the data information is called by the application terminal in a specific application scenario, based on the comprehensive quality score of the data information and the characteristic parameters of the specific application scenario, the representativeness of the evaluation data is expanded and a representativeness score is obtained. The accuracy score is calculated, and the comprehensive quality score of the data information, the normalized representativeness score and the accuracy score are merged. The application terminal comprehensive quality score is calculated in combination with the scenario sensitivity coefficient. Step 5: Based on the feedback from the overall quality score of the incoming goods and the overall quality score of the application, iteratively optimize the quality control parameters.

2. The carbon emission factor database data quality control method according to claim 1, characterized in that, In step 1, the carbon emission factor ontology includes material list, process boundary, indicator system, and traceability information. Semantic mapping relationships are established through a ternary association model of material process indicators. Automatic format mapping uses a semantic similarity matching method to align the original data fields with ontology elements. Automatic mapping is completed when the semantic similarity meets a preset threshold; otherwise, manual assistance or ontology expansion processes are triggered.

3. The carbon emission factor database data quality control method according to claim 2, characterized in that, In step 2, the calculation logic for the basic scores of the four dimensions is as follows: The consistency baseline score is evaluated by integrating format consistency and semantic consistency. The exchangeability baseline score assesses the compatibility of the data format with mainstream carbon accounting tools and the resolvability of the data itself. The basic score for completeness is evaluated by combining the rate of missing necessary information with the rate of completeness of optional information; The basic score for transparency is assessed based on the completeness of the data traceability chain and the traceability of the traceability information.

4. The carbon emission factor database data quality control method according to claim 3, characterized in that, In step 2, the method for constructing the dynamic coupling weight matrix is ​​as follows: first, the initial coupling weight is calculated based on the mutual information entropy between dimensions to characterize the correlation strength between dimensions; then, the initial coupling weight is adjusted according to the needs of specific application scenarios; and finally, normalization is performed to obtain the final dynamic coupling weight.

5. The carbon emission factor database data quality control method according to claim 4, characterized in that, In step 2, the collaborative correction coefficient is calculated based on the degree of difference in the base scores of each dimension; the greater the difference, the higher the correction strength. The formula for calculating the comprehensive attenuation factor is: ;in This is a time decay factor calculated based on the time difference between the year the data was collected and the current year. This is a geographical attenuation factor calculated based on the geographical differences between the data collection area and the application area.

6. The carbon emission factor database data quality control method according to claim 5, characterized in that, In step 2, the overall quality score of the incoming goods is calculated. The calculation formula is: ;in The scores are adjusted based on four dimensions: consistency, commutability, integrity, and transparency. For each dimension's basic weights determined through the entropy weight method, F is a comprehensive attenuation factor ranging from 0 to 1, and S is a scenario adaptation factor ranging from 0 to 1, determined based on industry adaptation, lifecycle stage adaptation, and accounting purpose adaptation.

7. The carbon emission factor database data quality control method according to claim 6, characterized in that, In step 3, the data source types include four categories: enterprise metrological data, industry statistical data, literature report data, and data from unknown sources; the differentiated qualification threshold is the product of the basic qualification threshold of 0.7 and the verification strength coefficient, with the verification strength coefficients set to 1.0, 0.85, 0.7, and 0.55 for the four types of data, respectively; the quality rating is a five-star rating, and the rating method is as follows: like It is five stars, if It is a four-star rating, if For Samsung, if For two stars, if It is rated as one star.

8. The carbon emission factor database data quality control method according to claim 7, characterized in that, In step 4, the representative score R is calculated by evaluating five categories of indicators: technical relevance, data source reliability, time relevance, geographical relevance, and scenario matching degree, and by using a combination of subjective weights, objective weights, and scenario weights. The accuracy score A is calculated by Monte Carlo simulation, and the initial uncertainty of the data is corrected by using the comprehensive quality score of the data entering the database during the simulation process.

9. The carbon emission factor database data quality control method according to claim 8, characterized in that, In step 4, the overall quality score of the application is calculated. ; in , , The weights for the entry score, normalized representativeness score, and accuracy score are 0.4, 0.3, and 0.3, respectively. The sensitivity coefficient is a weighted sum of the parameter sensitivity coefficient and the hypothesis sensitivity coefficient, with weights of 0.6 and 0.4, respectively. The sensitivity coefficient is 0.

2.

10. The carbon emission factor database data quality control method according to claim 9, characterized in that, In step 5, the closed-loop iterative optimization method is as follows: periodically update the dynamic coupling weights, decay factors, and scene weights based on user feedback and application adaptation results; if the overall quality score of the application... Data that is below the set first score threshold and cannot be corrected is discarded. The overall quality score of the application is... After supplementing the data (data greater than or equal to the first scoring threshold but less than the second scoring threshold), the overall quality score for data entry will be recalculated. Overall quality score of application The iteration effect is achieved through Average improvement rate The accuracy of the compatibility rate and the rate of decrease in user feedback were verified.