A comprehensive evaluation method and system for carbon emission data quality from stationary sources

By employing a dual-source data collaboration and multi-dimensional verification approach, the issues of data accuracy and coverage in stationary source carbon dioxide emission monitoring have been resolved, enabling efficient quality evaluation and supervision of carbon emission data, and improving data quality and regulatory efficiency.

CN121434199BActive Publication Date: 2026-04-03BEIJING YUANSHENG ENERGY TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing stationary source carbon dioxide emission monitoring technologies have bottlenecks in terms of data accuracy, coverage, and regulatory efficiency. Material monitoring methods are subject to default values ​​and human factors, while flue gas monitoring methods are costly and cannot cover dispersed sources or identify carbon emissions from alternative fuels.

Method used

By collaborating on dual-source data, the consistency of correlation coefficient distribution is constructed using material monitoring and flue gas monitoring methods. Combined with multi-dimensional testing and risk scoring, a comprehensive evaluation method for the quality of stationary source carbon emission data is constructed, including effective data calculation, risk scoring, and invalid data identification, and outputting four-dimensional evaluation results.

Benefits of technology

It improved the accuracy and coverage of carbon emission data, reduced regulatory costs, and enhanced the data quality for the carbon market and corporate internal controls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434199B_ABST
    Figure CN121434199B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for comprehensive evaluation of the quality of stationary source carbon emission data, relating to the field of carbon emission monitoring technology. The method includes: S1 collecting dual-source monitoring and production data, identifying invalid data, calculating and aggregating emissions, classifying operating conditions, constructing correlation coefficients, and establishing a basic dataset for different operating conditions through coefficient of variation testing; S2 calculating correlation coefficients for valid data to be tested, obtaining a comprehensive risk score for valid data through multi-dimensional testing and scoring, and combining industry weights; S3 calculating the proportion of invalid data as the invalid data risk score; S4 performing a two-dimensional risk rating on valid and invalid data, and outputting a four-dimensional result. The system implements this method through four modules: data input, processing, risk assessment, and result output, solving the technical problems of existing data quality evaluation methods that only focus on valid data, have a single testing method, and lack a complete dual-source collaborative quality control and all-dimensional quality management scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon emission data monitoring technology, and in particular to a method and system for comprehensive evaluation of the quality of carbon emission data from stationary sources. Background Technology

[0002] Currently, the accounting and monitoring of carbon dioxide emissions from stationary sources mainly rely on two types of technical methods, both of which have been practically applied in the industry. However, due to limitations in technical principles and implementation conditions, both methods have significant shortcomings, resulting in emission data quality that fails to meet the needs of carbon market regulation and corporate internal control. The specific technical status and problems are as follows:

[0003] Material monitoring is the primary accounting method for the power generation industry in my country's carbon emissions trading market. It operates according to the "Guidelines for Enterprise Greenhouse Gas Emission Accounting and Reporting (Power Generation Facilities)": It collects parameters such as fossil fuel consumption, lower heating value of fuel, and carbon oxidation rate (some parameters use industry default values) and calculates stationary source carbon dioxide emissions based on material balance logic. Currently, the material monitoring method has several problems: First, the use of default values ​​(lower heating value of fossil fuels, carbon oxidation rate, etc.) in key emission calculations can lead to discrepancies with actual emissions. Second, some steps involve human factors, such as coal sampling and sample preparation, increasing the potential for falsification by enterprises and introducing data uncertainty. Third, the numerous verification steps, including data sampling, testing, verification, and equipment calibration, require significant investment of manpower and resources for supervision, resulting in high regulatory costs.

[0004] Furthermore, Europe and the United States generally adopt the flue gas monitoring method, which involves installing Continuous Emission Monitoring Systems (CEMS) at fixed pollution sources. This allows for real-time, continuous data collection of parameters such as flue gas flow rate and carbon dioxide concentration, directly calculating carbon dioxide emissions per unit time. This enables real-time dynamic monitoring of emission data and is valuable in scenarios with high data timeliness requirements. However, flue gas monitoring has several drawbacks. First, it cannot monitor dispersed sources because it requires installing fixed monitoring equipment at each emission source, which is costly and technically challenging, making it difficult to cover a wide range of numerous dispersed emission sources. Second, it carries a high risk of systematic errors. Currently, there is a lack of installation and calibration standards for flue gas monitoring, and improper installation and calibration can lead to systematic errors. Third, flue gas monitoring cannot identify carbon emissions from alternative fuels and feedstocks because the measured carbon emissions include those from alternative fuels and feedstocks, which are not included in the carbon market's accounting scope.

[0005] In summary, existing stationary source carbon dioxide emission monitoring and evaluation technologies face significant bottlenecks in terms of data accuracy, coverage, regulatory efficiency, and evaluation completeness. There is an urgent need to develop a technical solution based on dual-source data collaboration, multi-dimensional verification, and full-process quality control to systematically address data quality issues. Summary of the Invention

[0006] The purpose of this invention is to provide a comprehensive evaluation method and system for the quality of carbon emission data from stationary sources. By utilizing the consistency of the correlation coefficient distribution constructed from data from material monitoring and online monitoring methods, the method performs quality diagnosis on carbon emission data in the application stage to improve the quality of carbon market and other stationary source carbon dioxide data.

[0007] To achieve the above objectives, this invention provides a method for comprehensive evaluation of the quality of carbon emission data from stationary sources, comprising the following steps:

[0008] S1. Effective data calculation and basic dataset construction: Collect dual-source monitoring data and production operation data obtained by material monitoring method and flue gas monitoring method respectively. After invalid data identification, effective time period carbon emission accounting and aggregation, and production conditions classification, construct correlation coefficients based on minute-level carbon emissions, and construct a basic dataset of correlation coefficients for different working conditions through the coefficient of variation test.

[0009] S2. Valid Data Risk Score: For the valid data to be tested, calculate its correlation coefficients for different work conditions, use multi-dimensional testing methods to compare and score the basic dataset, and combine industry weights to calculate the valid data risk score.

[0010] S3. Invalid Data Risk Score: Calculate the percentage of invalid data time periods and use the percentage of invalid time periods as the invalid data risk score;

[0011] S4. Dual-dimensional risk rating: Risk levels are divided according to the risk scores of valid data and invalid data, and the final output includes a four-dimensional evaluation result containing the proportion of invalid data, the level of invalid data, the score of valid data, and the level of valid data.

[0012] Preferably, the specific steps for effective data calculation and basic dataset construction in S1 are as follows:

[0013] S1.1 Synchronously collect dual-source monitoring data and production operation data within a set period, wherein the dual-source monitoring data is collected at a frequency of no less than once per minute; the production operation data includes unit load, fuel type and process parameters of the fixed source;

[0014] S1.2. The scope of invalid data is defined as data collected when emission data is missing or monitoring equipment is under abnormal conditions such as calibration or maintenance. An intelligent identification algorithm is used to simultaneously identify data from both material monitoring and flue gas monitoring methods. When either source data within a given time period is invalid, the corresponding time period is marked as an invalid data interval. Only for time periods where both source monitoring data are valid, minute-level carbon emissions are calculated, and based on the time dimension, valid minute-level emission data are aggregated into lower-frequency emission data pairs.

[0015] S1.3. Using the production operation data obtained in step S1.1 as feature input, the production conditions are automatically divided and identified through an intelligent classification algorithm. The low-frequency emission data pairs obtained in step S1.2 are associated and matched with the corresponding operating condition categories to complete the operating condition category labeling of the data pairs.

[0016] S1.4 Constructing a correlation coefficient for carbon emissions based on material monitoring and flue gas monitoring methods. , To reflect the stable quantitative relationship between dual-source data, the indicators should exhibit consistent distribution characteristics under the same operating conditions;

[0017] For the dataset consisting of correlation coefficients under each working condition, calculate its coefficient of variation. The formula is:

[0018]

[0019] in, Indicates the first The coefficient of variation for each working condition;

[0020] Indicates the first The average correlation coefficient of each working condition is calculated using the following formula:

[0021]

[0022] Indicates the first The standard deviation of the correlation coefficient for each working condition is calculated using the following formula:

[0023]

[0024] in, Indicates the first The number of correlation coefficients in each working condition; Indicates belonging to the first The set of subscripts for the correlation coefficients of each working condition; Indicates the first Correlation coefficients for each time interval;

[0025] If the coefficient of variation for the kth working condition is less than or equal to the specified threshold, the basic dataset of the correlation coefficient for that working condition is considered to have been successfully constructed. If the coefficient of variation is greater than the specified threshold, the valid data for that working condition is supplemented, and steps S1.1-S1.4 are re-executed until the coefficient of variation meets the specified threshold.

[0026] Preferably, the specific steps for S2 valid data risk scoring are as follows:

[0027] S2.1 Collect and verify the carbon dioxide emission data of the stationary source to be tested, process it according to the process of steps S1.1-S1.3, obtain the valid data to be tested and the working condition to which the valid data to be tested belongs, and calculate the working condition correlation coefficient of the valid data to be tested.

[0028] S2.2. Using mean test, variance test, nonparametric test, and other distribution characteristic test methods, compare the difference between the correlation coefficient dataset of the valid data to be tested and the basic dataset of correlation coefficients of the corresponding working conditions; set a significance level. If the p-value obtained by any test method is less than the preset significance level, it indicates that the data under this method has a high suspicion of being falsified, and the test method scores 1 point. If the p-value is greater than the preset significance level, the test method scores 0 points, indicating that the data passes the test by this method.

[0029] S2.3. Based on the industry characteristics of the fixed source, the weights of the testing methods are set, and the scores of each testing method are summarized using a weighted calculation method to obtain the comprehensive risk score of the valid data to be tested under each working condition. The calculation formula is as follows:

[0030]

[0031] in, Indicates the first Comprehensive risk assessment of data for each working condition; Indicates the adopted first Each test algorithm weight; Indicates the first The scoring results of each test method.

[0032] Preferably, in S2.2, the significance level is set to 0.05.

[0033] Preferably, the specific steps of S3, invalid data risk scoring, include:

[0034] The percentage of invalid data time periods is calculated by determining the proportion of invalid data time periods to the total duration of invalid data during the normal operation of station equipment. The calculation formula is:

[0035]

[0036] in, This indicates the total duration of invalid data, in hours (h). This indicates the total duration of the data period requiring verification, in hours (h).

[0037] Preferably, the data period to be verified in step S3 refers to the period during which the enterprise's production equipment is operating normally, excluding downtime caused by production stoppages.

[0038] Preferably, the specific steps of S4, dual-dimensional risk rating include:

[0039] S4.1, Based on the percentage of invalid data time periods obtained in step S3.1, ... For low risk, Medium risk The data is classified as high-risk, and its quality risk level is determined accordingly.

[0040] S4.2, The comprehensive risk score obtained in step S2.3, is used as... For low risk, For medium risk, For high-risk data, classify the effective data quality risk level;

[0041] S4.3 Output a four-dimensional carbon emission data quality evaluation result that includes the percentage of invalid data time periods, the quality risk level of invalid data, the comprehensive risk score of valid data, and the quality risk level of valid data.

[0042] A comprehensive scoring system for the quality of carbon emission data from stationary sources, used to implement a comprehensive evaluation method for the quality of carbon emission data from stationary sources, includes four functional modules that work in sequence and collaboratively:

[0043] (1) Data input layer: used to receive raw data related to carbon emissions from stationary sources, providing a foundation for subsequent data processing;

[0044] (2) Data processing layer: used to identify invalid data, calculate emissions, classify operating conditions, calculate correlation coefficients and construct basic datasets for input data, and complete data preprocessing and quality control;

[0045] (3) Risk assessment layer: used for preprocessing, multi-dimensional testing, risk scoring and risk level classification of the data to be tested, so as to realize data quality risk assessment;

[0046] (4) Results output layer: used to integrate data quality evaluation results and output them in a preset format for easy viewing and use by users.

[0047] Preferably, the data input layer includes:

[0048] Dual-source monitoring data receiving module: used to receive raw data from material monitoring and flue gas monitoring methods;

[0049] Production operation data receiving module: used to receive production operation data from fixed sources and provide feature inputs for operating condition classification;

[0050] The data processing layer includes:

[0051] Invalid Data Intelligent Identification Module: Identifies and marks invalid data using intelligent identification algorithms;

[0052] Emissions calculation module: Calculates minute-level carbon emissions during the effective period of dual sources according to the formula, and aggregates the effective minute-level emission data into low-frequency emission data pairs;

[0053] Operating condition classification module: Classifies production operating conditions and labels data pairs using clustering models or operating condition pattern recognition algorithms;

[0054] Correlation coefficient calculation module: Calculates correlation coefficients for different working conditions according to formulas;

[0055] Basic dataset quality control module: Calculates and tests the coefficient of variation, and outputs a qualified basic dataset of correlation coefficients;

[0056] Preferably, the risk assessment layer includes:

[0057] Data preprocessing module: processes the data to be inspected and matches it with the operating conditions;

[0058] Valid data multi-dimensional testing module: Performs mean test, variance test, nonparametric test and other distribution characteristic tests and scores them;

[0059] Effective Data Risk Score Calculation Module: Calculates the comprehensive risk score of effective data according to the formula;

[0060] Invalid data percentage calculation module: Calculates the percentage of invalid data within a time period according to a formula;

[0061] Risk rating module: Classifies the quality risk level of invalid data and valid data;

[0062] The result output layer includes:

[0063] Results processing module: Used to integrate four-dimensional evaluation results including the percentage of invalid data time periods, the quality risk level of invalid data, the comprehensive risk score of valid data, and the quality risk level of valid data;

[0064] Report generation module: Used to automatically generate carbon emission data quality assessment reports that include detailed formula calculations, correlation coefficient distribution charts, and risk level summary tables;

[0065] Visualization module: Used to display the distribution trend of correlation coefficients and the proportion of risk levels in the form of histograms and pie charts, so that users can intuitively view the data quality.

[0066] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0067] This application achieves a breakthrough through technological improvements in dual-source synchronous data acquisition, intelligent invalid identification, and condition-based quality control:

[0068] First, based on the carbon emission accounting guidelines and flue gas monitoring industry standards, raw data from both material monitoring and flue gas monitoring methods are collected simultaneously to construct a dual-source data pair, forming a basis for cross-verification from a monitoring perspective.

[0069] Secondly, an invalid data identification module is built through intelligent algorithms to define the collected data in scenarios such as data missing, equipment calibration, and maintenance status as invalid data. When any dimension of the dual-source data is invalid, the corresponding time interval is marked as invalid, thus eliminating interference items from the data screening process.

[0070] Third, using production operation data as feature input, the system automatically divides production operating conditions through clustering models and operating condition pattern recognition algorithms. The aggregated emission data pairs are associated and labeled with the corresponding operating conditions, and the stability of the sub-operating condition association coefficient dataset is ensured through the coefficient of variation test, thus eliminating the interference of mixed operating conditions from the data classification stage.

[0071] Fourth, for the valid data to be tested, first calculate the correlation coefficient, then conduct risk testing through multiple testing algorithms under different working conditions, and finally determine the weight of each testing algorithm based on industry characteristics, and obtain the comprehensive risk score of the valid data under different working conditions through weighted calculation.

[0072] Fifth, risk rating is conducted from two dimensions: invalid data and valid data. For invalid data, the proportion of total invalid data duration during normal equipment operation is calculated to determine the risk level based on certain standards. For valid data, a comprehensive risk score for valid data under different operating conditions is used to determine the risk level based on certain standards. The final output is a four-dimensional evaluation result including the proportion of invalid data time periods, invalid data risk level, valid data risk score, and valid data risk level.

[0073] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0075] Figure 1 This is a flowchart illustrating an embodiment of the method for comprehensive evaluation of carbon emission data quality from stationary sources according to the present invention. Detailed Implementation

[0076] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0077] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0078] Example

[0079] like Figure 1 As shown, a comprehensive evaluation method for the quality of carbon emission data from stationary sources includes the following steps:

[0080] S1. Valid Data Calculation and Basic Dataset Construction: Based on current carbon emission accounting guidelines and industry standards for online monitoring methods, dual-source monitoring data and production operation data obtained through material monitoring and flue gas monitoring methods are collected respectively. After invalid data identification, valid time period carbon emission calculation and aggregation, and production condition classification, correlation coefficients are constructed based on minute-level carbon emissions. A basic dataset of correlation coefficients by production condition is then constructed through coefficient of variation testing. The specific steps are as follows:

[0081] S1.1 Synchronously collect dual-source monitoring data and production operation data within a set period. The collection frequency of dual-source monitoring data shall not be less than 1 time / min. Production operation data includes unit load, fuel type and process parameters of fixed source.

[0082] Current carbon emission accounting guidelines and industry standards for online monitoring methods include:

[0083] GB 17167 General Rules for the Configuration and Management of Energy Measuring Instruments in Energy-Using Units

[0084] GB / T 7721 Continuous Accumulation Automatic Weighing Instruments (Electronic Belt Scales)

[0085] GB / T 21923 General Rules for Inspection of Solid Biomass Fuels

[0086] GB / T 30727 Determination of calorific value of solid biomass fuels

[0087] GB / T 28734 Method for Determination of Carbon and Hydrogen in Solid Biomass Fuels

[0088] GB / T 28017 Pressure-resistant metering coal feeder

[0089] HJ 75 Technical Specification for Continuous Monitoring of Flue Gas Emissions (SO2, NOx, Particulate Matter) from Stationary Sources

[0090] HJ 76 Technical Requirements and Testing Methods for Continuous Monitoring of Flue Gas Emissions (SO2, NOx, Particulate Matter) from Stationary Sources

[0091] HJ / T 397 Technical Specification for Monitoring of Stationary Source Exhaust Gas

[0092] JJG 444 Standard Rail Scale Verification Procedure

[0093] CJ / T 96 General Test Methods for Chemical Properties of Municipal Solid Waste

[0094] CJ / T 313 Sampling and Analysis Methods for Municipal Solid Waste

[0095] Guidelines for Enterprise Greenhouse Gas Emissions Accounting and Reporting - Power Generation Facilities

[0096] S1.2. Invalid data is defined as data collected when emission data is missing or monitoring equipment is under abnormal conditions such as calibration or maintenance. An intelligent identification algorithm is used to simultaneously identify data from both material monitoring and flue gas monitoring methods. When either source data within a given time period is invalid, the corresponding time period is marked as invalid. Only for time periods where both source monitoring data are valid, minute-level carbon emissions are calculated, and based on the time dimension, valid minute-level emission data are aggregated into lower-frequency emission data pairs at the 15-minute or 30-minute level.

[0097] S1.3. Using the production operation data obtained in step S1.1 as feature input, the production conditions are automatically divided and identified through intelligent classification algorithms (random forest or k-clustering). The low-frequency emission data pairs of the corresponding time obtained in step S1.2 are associated and matched with the corresponding operating condition categories to complete the category labeling of the data pairs.

[0098] S1.4 Constructing a correlation coefficient for carbon emissions based on material monitoring and flue gas monitoring methods. . To reflect the stable quantitative relationship of dual-source data, the index should exhibit consistent distribution characteristics under the same operating conditions, as shown in the formula:

[0099]

[0100] in, This represents the correlation coefficient for the i-th time interval; Represents the carbon emissions from the material monitoring method for the i-th time interval; This represents the carbon emissions from the flue gas monitoring method during the i-th time interval.

[0101] Alternatively, the correlation coefficient formula is:

[0102]

[0103] in, This represents the correlation coefficient for the i-th time interval; Represents the carbon emissions from the material monitoring method for the i-th time interval; This represents the carbon emissions from the flue gas monitoring method during the i-th time interval.

[0104] For the dataset consisting of correlation coefficients under each working condition, calculate its coefficient of variation. The formula is:

[0105]

[0106] in, Indicates the first The coefficient of variation for each working condition;

[0107] Indicates the first The average correlation coefficient of each working condition is calculated using the following formula:

[0108]

[0109] Indicates the first The standard deviation of the correlation coefficient for each working condition is calculated using the following formula:

[0110]

[0111] in, Indicates the first The number of correlation coefficients in each working condition; Indicates belonging to the first The set of subscripts for the correlation coefficients of each working condition; Indicates the first Correlation coefficients for each time interval;

[0112] If the coefficient of variation for the kth working condition is ≤0.2, the basic dataset for the correlation coefficient of that working condition is considered to have been successfully constructed; if the coefficient of variation is >0.2, supplement the valid data for that working condition, and repeat steps S1.1-S1.4 until the coefficient of variation meets the specified threshold.

[0113] S2. Valid Data Risk Scoring: For the valid data to be tested, calculate its correlation coefficients for different work conditions, compare it with the basic dataset using a multi-dimensional testing method and score it, and calculate the comprehensive risk score by combining industry weights; the specific steps are as follows:

[0114] S2.1 Collect the carbon dioxide emission data of the stationary source to be tested, process it according to the process of steps S1.1-S1.3, obtain the valid data to be tested and the working condition to which the valid data to be tested belongs, and calculate the working condition correlation coefficient of the valid data to be tested.

[0115] S2.2. Using mean test, variance test, nonparametric test, and other distribution characteristic test methods, compare the differences between the correlation coefficient dataset of the valid data to be tested and the basic correlation coefficient dataset of the corresponding working conditions; specifically:

[0116] The mean-based test algorithm calculates the variance (or standard deviation) of the dataset to be tested and the corresponding baseline dataset, and constructs a test statistic based on sample size and data distribution characteristics. This statistic determines whether there is a significant difference in the population variance of the two datasets due to factors other than random fluctuations, thus verifying whether the dispersion of the two populations is fundamentally different. When the obtained p-value is less than the pre-set significance level, the variance test score of the dataset to be tested is recorded as 1. When the p-value is greater than the significance level, the score is 0.

[0117] Based on the variance test algorithm: By calculating the variance (or standard deviation) of the dataset to be tested and the corresponding baseline dataset, and combining the sample size and data distribution characteristics, a test statistic is constructed to determine whether there is a significant difference in the population variance of the two datasets caused by factors other than random fluctuations, thereby verifying whether the dispersion of the two populations is fundamentally different. When the obtained p-value is less than the pre-set significance level, the variance test score of the dataset to be tested is recorded as 1. When the p-value is greater than the significance level, the score is 0.

[0118] Based on a nonparametric test algorithm: By analyzing nonparametric features such as rank (or sorting position) and frequency distribution of the dataset to be tested and the corresponding baseline dataset, the algorithm determines whether there is a fundamental difference in the overall distribution represented by the two datasets, thereby verifying the data differences. When the obtained p-value is less than the pre-set significance level, the nonparametric test score of the dataset to be tested is recorded as 1. When the p-value is greater than the significance level, the score is 0.

[0119] Based on artificial intelligence algorithms (random forest or support vector machine), the baseline dataset of association coefficients is divided into two classes. The first class contains the true association coefficients, with an objective function of 0. The second class contains the modified association coefficients, formed by multiplying the original association coefficients by a random coefficient, with an objective function of 1. An artificial intelligence algorithm is used to learn and generate a testing model from the two classes of data and their objective functions. This testing model is then applied to determine the category of the data to be tested. When more than 5% of the data points are classified as 1, the final score is recorded as 1; otherwise, the final score is recorded as 0. The formula for calculating the modified association coefficient for the second class is as follows:

[0120]

[0121] in, Indicates the first i The correlation coefficient after the time interval is modified Indicates the first i The correlation coefficient for each time interval. This represents a non-zero random modification coefficient, ranging from 0% to 3.5%.

[0122] Based on other testing algorithms: The correlation coefficient between the material monitoring method and the online monitoring method also exhibits other distributional characteristics. Other testing methods are applied to determine whether there is an essential difference in the distributional characteristics represented by the dataset to be tested and the corresponding basic dataset. When the obtained p-value is less than the pre-set significance level, the other test scores for the dataset to be tested are recorded as 1. When the p-value is greater than the significance level, the score is 0.

[0123] S2.3. Based on the industry characteristics of the fixed source, assign weights to each of the above testing methods. Summarize the scores of each testing method using a weighted calculation method to obtain the comprehensive risk score for the valid data to be tested under each working condition. The calculation formula is as follows:

[0124]

[0125] in, Indicates the first Comprehensive risk assessment of data for each working condition; Indicates the adopted first Each test algorithm weight; Indicates the first The scoring results of each test method.

[0126] S3. Invalid Data Risk Scoring: Calculate the percentage of invalid data periods and use this percentage as the risk score. Specific steps include:

[0127] Calculate the proportion of invalid data corresponding to the total normal operating time of the station equipment during the normal operation period. The normal operating period refers to the time during which the stationary source operates continuously according to its design conditions, excluding downtime caused by equipment failure or planned maintenance.

[0128] Percentage of invalid data periods The calculation formula is:

[0129]

[0130] in, This indicates the total duration of invalid data, in hours (h). This indicates the total duration of the data period requiring verification, in hours (h).

[0131] S4. Two-Dimensional Risk Rating: Risk levels are assigned based on the risk scores of valid and invalid data, resulting in a four-dimensional evaluation that includes the percentage of invalid data, the level of invalid data, the score of valid data, and the level of valid data. Specific steps include:

[0132] S4.1, Based on the percentage of invalid data time periods obtained in step S3.1, ... For low risk, Medium risk For high-risk data, invalid data quality risk levels are categorized and presented in a table below:

[0133] Table 1 Classification of Invalid Data Quality Risk Levels

[0134]

[0135] S4.2, The comprehensive risk score obtained in step S2.3, is used as... Low risk, 0.35 For medium risk, For high-risk data, the effective data quality risk levels are classified and represented in the table below:

[0136] Table 2 Classification of Valid Data Quality Risk Levels

[0137]

[0138] S4.3 Output the four-dimensional carbon emission data quality assessment results, including the percentage of invalid data periods, the quality risk level of invalid data, the comprehensive risk score of valid data, and the quality risk level of valid data. Presented in a table below:

[0139] Table 3 Summary of Carbon Emission Data Quality Assessment Results

[0140]

[0141] A comprehensive scoring system for the quality of stationary source carbon emission data, used to implement a comprehensive evaluation method for the quality of stationary source carbon emission data, includes:

[0142] (1) Data input layer, including:

[0143] Dual-source monitoring data receiving module: used to receive raw data from material monitoring and flue gas monitoring methods;

[0144] Production operation data receiving module: used to receive production operation data from fixed sources and provide feature inputs for operating condition classification;

[0145] (2) Data processing layer, including:

[0146] Invalid Data Intelligent Identification Module: Identifies and marks invalid data using intelligent identification algorithms;

[0147] Emissions calculation module: Calculates minute-level carbon emissions during the effective period of dual sources according to the formula, and aggregates the effective minute-level emission data into low-frequency emission data pairs;

[0148] Operating condition classification module: Classifies production operating conditions and labels data pairs using clustering models or operating condition pattern recognition algorithms;

[0149] Correlation coefficient calculation module: Calculates correlation coefficients for different working conditions according to formulas;

[0150] Basic dataset quality control module: Calculates and tests the coefficient of variation, and outputs a qualified basic dataset of correlation coefficients;

[0151] (3) Risk assessment layer:

[0152] Data preprocessing module: processes the data to be inspected and matches it with the operating conditions;

[0153] Valid data multi-dimensional testing module: Performs mean test, variance test, nonparametric test and other distribution characteristic tests and scores them;

[0154] Effective Data Risk Score Calculation Module: Calculates the comprehensive risk score of effective data according to the formula;

[0155] Invalid data percentage calculation module: Calculates the percentage of invalid data within a time period according to a formula;

[0156] Risk rating module: Classifies the quality risk level of invalid data and valid data;

[0157] (4) Result output layer, including:

[0158] Results processing module: Used to integrate four-dimensional evaluation results including the percentage of invalid data time periods, the quality risk level of invalid data, the comprehensive risk score of valid data, and the quality risk level of valid data;

[0159] Report generation module: Used to automatically generate carbon emission data quality assessment reports that include detailed formula calculations, correlation coefficient distribution charts, and risk level summary tables;

[0160] Visualization module: Used to display the distribution trend of correlation coefficients and the proportion of risk levels in the form of histograms and pie charts, so that users can intuitively view the data quality.

[0161] The remaining technical features in the above embodiments can be flexibly selected by those skilled in the art to meet different specific practical needs according to actual circumstances. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims. In the above description, numerous specific details have been set forth to provide a thorough understanding of the present invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to implement the present invention. In other instances, to avoid obscuring the present invention, well-known techniques, such as specific construction details, operating conditions, and other technical conditions, have not been specifically described.

[0162] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for comprehensive evaluation of the quality of carbon emission data from stationary sources, characterized in that, The steps are as follows: S1. Effective data calculation and basic dataset construction, the specific steps are as follows: S1.1 Synchronously collect dual-source monitoring data and production operation data within a set period, wherein the collection frequency of the dual-source monitoring data is not less than the preset frequency; S1.

2. The scope of invalid data is defined as data collected under abnormal conditions. An intelligent identification algorithm is used to simultaneously identify data from both material monitoring and flue gas monitoring methods. When either source data within a given time period is invalid, the corresponding time period is marked as an invalid data interval. Only for time periods where both source monitoring data are valid, minute-level carbon emissions are calculated, and based on the time dimension, valid minute-level emission data are aggregated into low-frequency emission data pairs. S1.

3. Using the production operation data obtained in step S1.1 as feature input, the production conditions are automatically divided and identified through an intelligent classification algorithm. The low-frequency emission data pairs obtained in step S1.2 are associated and matched with the corresponding operating condition categories to complete the operating condition category labeling of the data pairs. S1.4 Constructing a correlation coefficient for carbon emissions based on material monitoring and flue gas monitoring methods. The formula is as follows: (This implies that the distribution characteristics are consistent under the same working conditions.) ; in, This represents the correlation coefficient for the i-th time interval; Represents the carbon emissions from the material monitoring method for the i-th time interval; This represents the carbon emissions from the flue gas monitoring method during the i-th time interval. For the dataset consisting of correlation coefficients under each working condition, calculate its coefficient of variation. The formula is: ; in, Indicates the first The coefficient of variation for each working condition Indicates the first The average correlation coefficient of each working condition Indicates the first Standard deviation of the correlation coefficient for each working condition; If the coefficient of variation for the kth working condition is ≤0.2, the basic dataset for the correlation coefficient of that working condition is considered to have been successfully constructed; if the coefficient of variation is >0.2, supplement the valid data for that working condition, and repeat steps S1.1-S1.4 until the coefficient of variation meets the specified threshold. S2. Valid Data Risk Score: For the valid data to be tested, calculate its correlation coefficients for different work conditions, compare it with the basic dataset using a multi-dimensional testing method and score it, and calculate the valid data risk score by combining industry weights; the specific steps are as follows: S2.1 Collect and verify the carbon dioxide emission data of the stationary source to be tested, process it according to the process of steps S1.1-S1.3, obtain the valid data to be tested and the working condition to which the valid data to be tested belongs, and calculate the working condition correlation coefficient of the valid data to be tested. S2.

2. Using mean test, variance test, nonparametric test, and other distribution characteristic test methods, compare the difference between the dataset of correlation coefficients of the working conditions to be tested and the basic dataset of correlation coefficients of the corresponding working conditions. Set a preset significance level. If the difference obtained by any test method is less than the preset significance level, the test method scores 1 point, indicating that the data under the method is suspected of being falsified. If the difference is greater than the preset significance level, the test method scores 0 points, indicating that the data passes the test by the method. S2.

3. Based on the industry characteristics of the fixed source, the weights of the testing methods are set, and the scores of each testing method are summarized using a weighted calculation method to obtain the comprehensive risk score of the valid data to be tested under each working condition. The calculation formula is as follows: ; in, Indicates the first Comprehensive risk assessment of data for each working condition; Indicates the adopted first Each test algorithm weight; Indicates the first The scoring results of each test method; S3. Invalid Data Risk Scoring: Calculate the percentage of invalid data time periods and use this percentage as the invalid data risk score; specific steps include: The proportion of invalid data during the normal operation of fixed source equipment is calculated as the percentage of the total normal operation time of the station equipment. The calculation formula is: ; in, Shows the total duration corresponding to invalid data, in units of ; The total duration of the data period to be verified is expressed in hours (h). The data period to be verified refers to the period during which the enterprise's production equipment is operating normally, excluding downtime caused by production stoppages. S4. Two-Dimensional Risk Rating: Risk levels are divided based on both valid and invalid data risk scores, ultimately outputting a four-dimensional evaluation result including the proportion of invalid data, invalid data level, valid data score, and valid data level. Specific steps include: S4.1, Based on the percentage of invalid data time periods obtained in step S3, ... For low risk, Medium risk For high-risk data, invalid data quality risk levels are categorized; among them, ; S4.2, The comprehensive risk score obtained in step S2.3, is used as... For low risk, For medium risk, For high-risk data, the effective data quality risk level is classified; among them, , ; S4.3 Output a four-dimensional carbon emission data quality evaluation result that includes the percentage of invalid data time periods, the quality risk level of invalid data, the comprehensive risk score of valid data, and the quality risk level of valid data.

2. The method for comprehensive evaluation of the quality of stationary source carbon emission data according to claim 1, characterized in that: In S2.2, the significance level is set to 0.

05.

3. A comprehensive scoring system for the quality of stationary source carbon emission data, used to implement the comprehensive evaluation method for the quality of stationary source carbon emission data as described in any one of claims 1-2, characterized in that, It includes four functional modules that work in sequence and in coordination: (1) Data input layer: used to receive raw data related to carbon emissions from stationary sources, providing a foundation for subsequent data processing; (2) Data processing layer: used to identify invalid data, calculate emissions, classify operating conditions, calculate correlation coefficients and construct basic datasets for input data, and complete data preprocessing and quality control; (3) Risk assessment layer: used for preprocessing, multi-dimensional testing, risk scoring and risk level classification of the data to be tested, so as to realize data quality risk assessment; (4) Results output layer: used to integrate data quality evaluation results and output them in a preset format for easy viewing and use by users.

4. The stationary source carbon emission data quality comprehensive scoring system according to claim 3, characterized in that: The data input layer includes: Dual-source monitoring data receiving module: used to receive raw monitoring data corresponding to material monitoring method and flue gas monitoring method; Production operation data receiving module: used to receive production operation-related data from fixed sources and provide feature inputs for operating condition classification; The data processing layer includes: Invalid data identification module: Identifies and marks invalid data using data identification algorithms; Emissions calculation module: Calculates minute-level carbon emissions for the effective period of dual sources according to the corresponding formula, and aggregates the effective minute-level emission data into low-frequency emission data pairs; Operating condition classification module: Classifies production operating conditions and labels data pairs using an operating condition classification algorithm; Correlation coefficient calculation module: Calculates correlation coefficients for different working conditions according to the corresponding formula; Basic dataset quality control module: Calculates and tests the coefficient of variation, and outputs a qualified basic dataset of correlation coefficients.

5. A comprehensive scoring system for the quality of stationary source carbon emission data according to claim 3, characterized in that: The risk assessment layer includes: Data preprocessing module: processes the data to be inspected and matches it with the corresponding working conditions; Valid data multi-dimensional testing module: performs mean test, variance test, nonparametric test and other distribution characteristic test and scores; Effective Data Risk Score Calculation Module: Calculates the comprehensive risk score of effective data according to the corresponding formula; Invalid data percentage calculation module: Calculates the percentage of invalid data within a time period according to the corresponding formula; Risk rating module: Classifies the quality risk level of invalid data and valid data; The result output layer includes: Results processing module: Used to integrate four-dimensional evaluation results including the percentage of invalid data time periods, the quality risk level of invalid data, the comprehensive risk score of valid data, and the quality risk level of valid data; Report generation module: Used to automatically generate carbon emission data quality assessment reports that include detailed formula calculations, correlation coefficient distribution charts, and risk level summary tables; Visualization module: Displays the distribution trend of correlation coefficients and the proportion of risk levels in a visual way, making it easy for users to intuitively view the data quality.

Citation Information

Patent Citations

  • Multi-source data carbon emission evaluation method and device

    CN116862253A

  • Calculation method for project carbon emission accounting data quality

    CN118797232A