A method and system for dynamically calibrating fuel end carbon emission factors of a control and discharge enterprise

By establishing coal type classification and identification models and prediction models, combined with three-level data verification rules, the problem of insufficient accuracy of carbon emission factors was solved, and the accuracy and reliability of carbon emission accounting data were achieved, supporting the management of the carbon market and emission reduction measures.

CN120125254BActive Publication Date: 2025-12-26SOUTH CHINA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510197269.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-12-26
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In existing technologies, the accuracy and precision of carbon emission factors are poor, resulting in inaccurate carbon accounting results, causing economic losses to enterprises, and posing a risk of data falsification.

Method used

By establishing classification and identification models for different coal types, and using particle swarm optimization support vector machine combined with backward sequence selection algorithm, a prediction model for coal combustion carbon emission factor, elemental carbon content and lower heating value is established. Verification thresholds are set, and a three-level data verification rule is used to automatically verify the enterprise's coal combustion carbon emission factor data.

Benefits of technology

It enables accurate verification of corporate coal-fired carbon emission factor data, timely identification of abnormal data, ensures the quality of carbon emission accounting data, provides basic data support for the carbon market, helps set reasonable carbon emission limits and pricing, and promotes emission reduction actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125254B_ABST
    Figure CN120125254B_ABST
Patent Text Reader

Abstract

The application discloses a kind of control enterprise fuel end carbon emission factor dynamic calibration method and system, by establishing different coal classification identification model, the prediction model of coal combustion carbon emission factor CEF, element carbon content C, low calorific value NCV is established, according to the error absolute value of prediction result further analysis the prediction model, the calibration threshold of three is determined respectively, and the rationality of coal-fired carbon emission factor data of enterprise is carried out automatic calibration using three-level data calibration rule, can identify abnormal data in time, effectively assist third party verification organization to calibrate the accuracy and rationality of enterprise fuel end carbon emission factor result, to effectively guarantee carbon emission accounting data quality.In addition, it also provides basic data support and guidance for the design and operation of carbon market, which helps market managers to master the carbon emission factor trend of different activities and industries, so as to set reasonable carbon emission limit and pricing, effectively guide and promote emission reduction action.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of low-carbon energy saving and environmental protection, and particularly relates to a fuel end carbon emission factor dynamic calibration method and system for emission control enterprises. BACKGROUND

[0002] In the carbon market, the carbon emission factor, a key numerical indicator, provides a basis for the quantification and evaluation of carbon emission levels. Accurate detection of the carbon emission factor is the basis for calculating the carbon emission of enterprises, providing basic data support and guidance for the design and operation of the carbon market. The prediction of the carbon emission factor helps market managers to grasp the trend of the carbon emission factor of different activities and industries, so as to set reasonable carbon emission limits and prices, effectively guide and promote emission reduction actions. Combined with the carbon market enterprise data reporting, the construction of the greenhouse gas emission factor database can provide technical parameters for carbon accounting at different levels, reduce the cost of carbon accounting and improve the accuracy of accounting.

[0003] At present, most of the carbon accounting work refers to the accounting method and guidelines of the Intergovernmental Panel on Climate Change (IPCC), but the actual conditions in different places cause poor pertinence and accuracy. For carbon emission, most of them are calculated according to the emission factor. For fossil fuel combustion emissions, the emission factor can use the self-test value or the default value provided by the standard guide. The carbon emission data calculated by the self-test value provided by the enterprise cannot be calibrated, so its accuracy cannot be guaranteed; the carbon content of fuel unit heat value is different in different regions, and there is a big difference in the degree of fuel combustion. If the default value is used for accounting, the result will be less accurate, and it will also cause significant economic losses to power generation enterprises with emission of millions of tons. The national scale of industry carbon accounting standards is not unified, which leads to the existence of elastic space for data, and provides soil for data fraud.

[0004] Therefore, there is an urgent need for an accurate and efficient carbon emission factor dynamic calibration method to effectively guarantee the quality of carbon emission accounting data. SUMMARY

[0005] One of the purposes of the present application is to provide a fuel end carbon emission factor dynamic calibration method for emission control enterprises, which establishes different coal classification identification models, establishes prediction models of coal combustion carbon emission factor CEF, elemental carbon content C and low calorific value NCV, further analyzes the absolute error value of the prediction model according to the prediction result, respectively determines the calibration threshold of the three, and adopts three-level data calibration rules to automatically calibrate the rationality of the coal combustion carbon emission factor data of the enterprise, which can timely identify abnormal data, effectively assist the third-party verification agency to calibrate the accuracy and rationality of the fuel end carbon emission factor result of the enterprise, and thus effectively guarantee the quality of carbon emission accounting data.

[0006] To achieve the above-mentioned purposes, the technical solutions adopted by the present application are as follows:

[0007] A control and discharge enterprise fuel end carbon emission factor dynamic calibration method, comprising the following steps:

[0008] Step S1. Establish a different coal classification identification model: first, sample collection and processing, collect coal quality analysis parameter data of different coal types (including total water (Mt), industrial analysis (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), total sulfur content (St), etc.), divide the coal quality analysis parameter data into training set and test set, set different labels for different categories of coal to distinguish (for example, the label of bituminous coal is set to 1, the label of lignite is set to 2, and so on), the coal quality analysis data of the training set of coal samples should be the data of standard coal, or the data of coal samples detected by CMA recognized or CNAS recognized detection institutions / laboratories; according to the training set, the K nearest neighbor (K-Nearest Neighbor, KNN) algorithm is used to establish a coal classification identification model, then the coal type is judged, and finally the performance of the established coal classification identification model is evaluated using the test set samples;

[0009] Step S2. Establish a prediction model for the carbon emission factor CEF of coal combustion, the elemental carbon content C, and the low calorific value NCV (i.e. low heat value): adopt particle swarm optimization-support vector regression (PSO-SVR) combined with sequential backward selection (SBS) algorithm to establish PSO-SVR prediction models for the carbon emission factor CEF of coal combustion, the elemental carbon content C, and the low calorific value NCV of different coal types (such as bituminous coal, lignite, anthracite, etc.);

[0010] Step S3. Set the calibration threshold: according to the prediction results of the established PSO-SVR prediction model for the carbon emission factor CEF of coal combustion, the elemental carbon content C, and the low calorific value NCV, further analyze the absolute error value of the prediction model, and determine the calibration threshold of each, as the basis for determining the reasonableness of carbon emission data in the data calibration process;

[0011] Step S4. Set the three-level data checking rules: the third-party checking institution can calculate the error between the predicted value calculated according to the coal combustion carbon emission factor prediction model and the measured calculated value, the error between the predicted value calculated according to the element carbon content C prediction model and the measured calculated value, and the error between the predicted value calculated according to the low calorific value NCV prediction model and the measured calculated value according to the enterprise reported data, and compare them with the corresponding carbon emission factor CEF, element carbon content C, and low calorific value NCV prediction model checking threshold in step S3, so as to judge whether the enterprise reported carbon emission factor data is abnormal, if it is judged to be abnormal, further identify the source of the abnormal data, and replace it with the corresponding default value for recalculation. The three-level data checking rules can automatically check the rationality of the enterprise's coal combustion carbon emission factor data, identify abnormal data in time, and grasp the data quality of carbon emission calculation according to the same.

[0012] Preferably, the K-nearest neighbor algorithm is used to establish the coal type classification and identification model in step S1, and the method is as follows: given the training set data T and the K value;

[0013] The training set data T is represented as:

[0014] T = {(x1, y1), (x2, y2),..., (xN, yN)} (i = 1, 2,..., N); i ,y i )} (i = 1, 2,..., N);

[0015] In the formula, x i is the coal quality analysis parameter data (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), and total sulfur content (St) of the coal sample, which is used as the model input variable; y i is the category of the coal type training set sample, which has m categories, and N is the number of samples;

[0016] In the K-nearest neighbor algorithm, the value of the number of nearest neighbors (K) will affect the accuracy of the final result. If the K value is too large, overfitting will occur, and if the K value is too small, underfitting will occur. Therefore, the K-fold cross-validation method is used to find the most suitable K value. The steps of using the K-fold cross-validation method to determine the optimal K value in the K-nearest neighbor algorithm are as follows:

[0017] a1. Set the range of K value: first determine a range of K value, usually starting from 1, and the upper limit can be set to 20 or a smaller number;

[0018] a2. Divide the data set: divide the entire data set into K subsets, and each subset is used as the test set in turn during cross-validation, and the remaining K-1 subsets are used as the training set;

[0019] a3. Model training and validation: For each K value, repeat the following steps K times: train a KNN model using data from K-1 subsets, and use the remaining one subset as the test set to evaluate the model performance;

[0020] a4. Calculate the average accuracy: For each K value, calculate its average accuracy in K times of cross-validation;

[0021] a5. Determine the optimal K value: Compare the average accuracy of different K values, and select the K value with the highest accuracy as the optimal K value.

[0022] Preferably, the method for judging the category of the unknown coal sample in step S1 is as follows:

[0023] b1. Calculate the distance between each training set sample and the unknown coal sample, and the calculation formula is:

[0024]

[0025] where x i is the coal quality analysis parameter data of the i-th training set sample; x j is the coal quality analysis parameter data of the j-th unknown coal sample; x i ,x j ∈R n ; l is the ordinal number of the coal quality analysis parameter, from 1 to n; n is the total number of coal quality analysis parameters; p is a variable parameter, p≥1, and according to the different parameters p, different types of distances can be represented, when p=1 Manhattan distance, p=2 Euclidean distance; if the feature space is multi-dimensional and each dimension is independent, Manhattan distance is selected; while when the feature space is continuous and each dimension is related, and the straight-line distance between data points needs to be reflected, Euclidean distance is selected;

[0026] b2. Count the data points M k (x) of the K nearest neighbor training set samples to the data x of the unknown coal sample;

[0027] b3. The category with the most votes in the data points M k (x) of the K nearest neighbor training set samples is taken as the final output category of the unknown coal sample, and the calculation formula is:

[0028]

[0029] where f is the judgment indicator function, y i is the actual coal type training set sample category, c j is the coal type classification and recognition model predicted coal type category, when y i =c jf = 1 if t = 1, otherwise f = 0.

[0030] Preferably, the step S1 further comprises a model performance evaluation step: selecting the classification accuracy (Accuracy) as the evaluation index of the coal classification identification model, to evaluate the performance of the coal classification identification model. Generally speaking, the higher the classification accuracy, the better the recognition effect of the classification model, and the calculation formula is as follows:

[0031]

[0032] In the formula, Accuracy represents the classification accuracy of the sample coal, T is the number of correct classification, and F is the number of incorrect classification.

[0033] Preferably, in the step S2, the establishment of the prediction model of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV comprises the following steps:

[0034] c1. Data preprocessing: first, give the training sample set q:

[0035] q = {(x1, y1), (x2, y2),..., (x i , y i ),..., (x n , y n )}

[0036] In the formula, xi is the input feature vector, which is represented as coal quality analysis parameter data; yi is the output feature vector, which is represented as the coal combustion carbon emission factor or the elemental carbon content or the low calorific value;

[0037] Then, normalize all data of the training set of the coal sample according to the following formula:

[0038]

[0039] In the formula, X ik is the data normalized value; x ik is the original data value; x k min is the minimum value of the original data; x k max is the maximum value of the original data;

[0040] c2. Establishing an initial full characteristic model: first, a group of coal samples with known coal quality analysis parameter index values (including total water (Mt), proximate analysis (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), total sulfur content (St), etc.) are used for calibration to establish preliminary SVR models (support vector regression models) of the full characteristics of coal combustion carbon emission factors CEF, element carbon content C, and low calorific value NCV for different coal types (such as bituminous coal, lignite, anthracite, and mixed coal, etc.). The coal quality analysis data of the coal samples used for calibration should be standard coal data, or coal sample data detected by CMA-recognized or CNAS-recognized detection institutions / laboratories; the preliminary SVR model maps the nonlinear feature vector in the data sample from a low-dimensional space to a high-dimensional feature space, and then fits a linear regression function in the feature space:

[0041]

[0042] where ω is the weight vector; b is the bias constant; is a nonlinear mapping from the input space to the high-dimensional feature space;

[0043] The preliminary SVR model defines a loss function ε, which optimizes the model parameters by finding the minimum value of the function, to find the best parameter value of the regression function, as shown in the following formula:

[0044]

[0045] The constraint condition is:

[0046]

[0047] where c is the penalty factor; ε is the loss function; b is the bias constant; ξ i , ξ' i are the slack variables; ω is the weight vector; x i is the input feature vector, and y i is the output feature vector;

[0048] After operation, the linear regression function is as shown in the following formula:

[0049]

[0050] where βi, αi are the Lagrange multipliers, K(x g x i ) is the radial basis kernel function of the support vector machine, which has good generalization performance and involves fewer parameters, having certain advantages in parameter optimization. The radial basis kernel function is as shown in the following formula: ​

[0051]

[0052] wherein g is a kernel function parameter, representing the mean square deviation of the Gaussian function, i.e. the width of the function in the independent variable direction; when g is small, the fitting performance of the function to the data is good; when g is large, the generalization ability of the function is strong; therefore, the determination of the optimal parameter directly affects the prediction model of the SVR coal combustion carbon emission factor of the radial basis kernel function type, in order to further improve the prediction performance of the SVR model, an effective optimization algorithm should be found to optimize the parameters of the SVR model;

[0053] c3. Model parameter optimization: for the SVR (support vector machine) of the radial basis kernel function, the penalty factor c and the kernel parameter g are the main parameters affecting the performance of the SVR, and the accuracy and reliability of the SVR prediction result depend on the best choice of (c, g); therefore, the particle swarm optimization algorithm (particle swarm optimization, PSO) is used to optimize the selection of the penalty factor c and the kernel parameter g of the SVR of the radial basis kernel function, and the PSO-SVR model for predicting the coal combustion carbon emission factor, the elemental carbon content and the low calorific value is constructed;

[0054] c4. Feature selection: based on the SBS algorithm, combined with PSO-SVR, a feature is removed from the current feature set, and a new model is established, and the model performance is calculated (the root mean square error of prediction (The root mean square error of prediction, RMSE) is used as the model performance evaluation index); each parameter in the full feature set is traversed, and one parameter is removed from the full feature set in each round, and a new model is established, and the model performance is calculated; then the performance difference between the new model and the original model is compared, if the model performance does not decrease significantly, the feature is removed; otherwise, the feature is retained; by repeatedly repeating the above steps, irrelevant features are gradually removed, until the performance of the model no longer improves, thereby selecting the most relevant feature subset of the carbon emission factor, and using the finally selected feature subset to construct the PSO-SVR model, i.e. the optimal model of the coal combustion carbon emission factor, the elemental carbon content and the low calorific value is obtained; by reducing irrelevant features, the generalization error of the model can be reduced, and the prediction ability of the model can be improved.

[0055] c5. Model performance evaluation: the above feature selection uses the root mean square error (RMSE) to evaluate the performance of the model, and under normal circumstances, the smaller the RMSE, the higher the prediction accuracy of the model, and the calculation formula is:

[0056]

[0057] y i , respectively, n represents the sample number of the coal sample, and i is the number of independent variables.

[0058] Preferably, the mean absolute error (MAE) is used as the basis for setting the verification threshold in the step S3; the smaller the mean absolute error (MAE), the higher the prediction accuracy of the model, and the calculation formula is as follows:

[0059]

[0060] wherein y i , respectively, n represents the sample number of the coal sample, and i is the number of independent variables.

[0061] Preferably, the three-level data verification process of the step S4 includes the following steps.

[0062] S41. First-level data verification: according to the coal quality data reported by the enterprise, the measured low calorific value NCV and the element carbon content C are used to calculate the measured and calculated value of the coal combustion carbon emission factor CEF, and then the prediction model of the coal combustion carbon emission factor CEF established in the step S3 is used to obtain the model prediction value of the carbon emission factor CEF, the error between the model prediction value and the measured and calculated value of the carbon emission factor CEF is calculated, if the error is less than the preset threshold, the verification is passed, the carbon emission factor CEF data is automatically added to the database system and dynamically updated, if the error is greater than the preset threshold, the second-level data verification is entered;

[0063] S42. Second-level data verification: according to the coal quality data reported by the enterprise, the prediction model of the element carbon content C established in the step S3 is used to obtain the model prediction value of the element carbon content, and the error between the model prediction value and the measured value of the element carbon content is calculated, if the error is greater than the preset threshold, it is indicated that the abnormal source of the carbon emission factor is the element carbon content, the element carbon content is marked as abnormal, and the enterprise manager is notified that the measured data of the element carbon content is doubtful, and the default value of the unit heat value carbon content in the guide is used to replace the element carbon content to calculate the carbon emission factor; if the error is less than the preset threshold, the third-level data verification is entered;

[0064] S43. Third-level data checking: according to the coal quality data reported by the enterprise, the model predicted value of the low calorific value is obtained by using the low calorific value pre-NCV prediction model established in step S3, the error between the model predicted value and the measured value of the low calorific value is calculated, if the error is greater than the preset threshold, it indicates that the abnormal source of the carbon emission factor is the low calorific value, the low calorific value is marked as abnormal, and the enterprise responsible person is notified that the low calorific value measured data is doubtful, and the guide low calorific value default value is used to replace the calculation of the carbon emission factor; if the error is less than the preset threshold, the enterprise responsible person is notified, the detailed report of the abnormal coal quality data is provided or the on-site verification is carried out, and further analysis is carried out to determine the root cause of the abnormal carbon emission factor data.

[0065] Preferably, step S5: model updating is further included; in order to ensure the prediction accuracy of the prediction model of the coal-fired combustion carbon emission factor, the elemental carbon content and the low calorific value, the model needs to be updated regularly, including collecting new coal quality data, retraining the model and the like, so as to adapt to the changes of coal types and combustion efficiency.

[0066] Preferably, the step of searching for the optimal SVR parameter by PSO algorithm in step c3 is as follows:

[0067] c31. initialize the PSO-SVR model parameters, including: the number of iterations, the number of particles, the learning factors c1 and c2, the search range of the penalty factor c and the kernel parameter g, and assume that the position of the i-th particle in the D-dimensional space in the particle swarm is x i =(x i1 ,x i2 ,···,x ij ,···,x id ), the particle velocity is v i =(v i1 ,v i2 ,···,v ij ,···,v id ), and the single iteration displacement of the particle in the search process is determined;

[0068] c32. learn the training samples by using SVR, take the root mean square error (RMSE) of the prediction results of the training samples of the coal as the fitness function, update the velocity and position of the population particles; compare the current fitness value of the particle with the optimal fitness value of the particle, if the current value is better, the current position is taken as the optimal position of the particle;

[0069] c33. update the velocity and position of the particle by using the following two formulas:

[0070]

[0071] Wherein, k is the current iteration number, c1, c2 are non-negative learning factors, r1 and r2 are random numbers distributed in the interval (0, 1), Vi(k) and xi(k) respectively represent the velocity and position of particle i at the kth iteration in d-dimensional space, is the global extreme value after k iterations, Vi(k+1) and xi(k+1) respectively represent the velocity and position of particle i at the (k+1)th iteration in d-dimensional space; ω is an inertia factor, representing the global search ability, which is defined as:

[0072] ω = ω min +(iter max -iter)·(ω max -ω min ) / iter max

[0073] Wherein: ω max and ω min are the maximum and minimum weight factors respectively, iter is the current iteration number, and iter max is the maximum iteration number.

[0074] c34. Determine whether the set maximum iteration number or accuracy target and other optimization termination conditions are met, if yes, the optimal solution is obtained, and if not, go to step c32;

[0075] Finally, the optimal position vector (c, g) is output, that is, the optimal PSO-SVR coal combustion carbon emission factor CEF, element carbon content C and low calorific value NCV model parameters are obtained: penalty factor (bestc) and kernel function parameter (bestg).

[0076] The second object of the present application is to provide a dynamic calibration system for controlling the carbon emission factor of fuel end of an enterprise, which is formed by the above method, and the system comprises:

[0077] A classification model generation module, a prediction model generation module, a calibration threshold determination module and a calibration module.

[0078] The classification model generation module is used for sample collection and processing, collecting coal quality analysis parameter data of different coal types, dividing the coal quality analysis parameter data into a training set and a test set, setting different labels for different types of coal to distinguish them, selecting a K nearest neighbor algorithm to establish a coal type classification and identification model according to the training set, then judging the coal type, and finally using the test set sample to evaluate the performance of the established coal type classification and identification model.

[0079] The prediction model generation module is used to establish prediction models of the coal-fired carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV of different coal types by using a particle swarm optimization support vector machine combined with a backtracking sequential selection algorithm.

[0080] The check threshold determination module is used to further analyze the error absolute values of the prediction models according to the prediction results of the established prediction models of the coal-fired carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV, and determine the check thresholds of the three, which are used as the judgment basis for the reasonableness of the carbon emission data in the data checking process.

[0081] The checking module is used for the third-party checking institution to calculate the errors between the predicted values and the measured values calculated according to the prediction models of the coal-fired carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV in sequence according to the data reported by the enterprise, and compare the errors with the check thresholds of the corresponding carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV prediction models in the check threshold determination module, so as to judge whether the carbon emission factor data reported by the enterprise is abnormal, and if the data is abnormal, further identify the source of the abnormal data and replace it with a corresponding default value for recalculation.

[0082] Beneficial effects:

[0083] The fuel end carbon emission factor dynamic checking method and system for the control and emission enterprise can timely identify abnormal data and effectively assist the third-party checking institution in checking the accuracy and reasonableness of the fuel end carbon emission factor results of the enterprise, so as to effectively guarantee the quality of the carbon emission accounting data. In addition, the method and system provide basic data support and guidance for the design and operation of the carbon market, which helps the market managers to master the change trend of the carbon emission factor of different activities and industries, so as to set a reasonable carbon emission limit and price and effectively guide and promote the emission reduction action. BRIEF DESCRIPTION OF DRAWINGS

[0084] Figure 1 The figure shows a fuel end carbon emission factor dynamic checking method and flowchart of the method for the control and emission enterprise. DETAILED DESCRIPTION

[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, specific implementation manners of the present application will be described below with reference to the drawings. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without any creative effort, and other embodiments can also be obtained.

[0086] The technical solutions of the present application will be described in detail below with specific embodiments.

[0087] Reference Figure 1 A dynamic calibration method for controlling and discharging fuel end carbon emission factor of an enterprise, and the calculation formula of the carbon dioxide emission factor of coal-fired fuel is:

[0088]

[0089] In the formula, CEF is the carbon dioxide emission factor of coal-fired fuel, the unit is tCO2 / GJ; C is the elemental carbon content of coal-fired fuel, the unit is tC / t; NCV is the low calorific value of coal-fired fuel, i.e. low heat value, the unit is GJ / t; OF is the oxidation rate of coal-fired fuel, the unit is expressed in %, and the carbon oxidation rate of coal-fired fuel specified in the standard guide is 99%.

[0090] Specifically, the dynamic calibration method for controlling and discharging fuel end carbon emission factor of an enterprise includes the following steps:

[0091] Step S1. Establish a different coal classification and identification model: first, sample collection and processing, collect coal quality analysis parameter data of different coal-fired coal types (including total water (Mt), industrial analysis (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), total sulfur content (St) and the like), divide the coal quality analysis parameter data into training set and test set, set different labels for different categories of coal types to distinguish (for example, the label of bituminous coal is set to 1, the label of lignite is set to 2, and so on), the coal quality analysis data of the coal sample in the training set needs to be the data of standard coal, or the data of coal sample detected by a CMA-recognized or CNAS-recognized detection institution / laboratory; according to the training set, a K-nearest neighbor (K-Nearest Neighbor, KNN for short) algorithm is used to establish a coal classification and identification model, then the coal type is judged, and finally the performance of the established coal classification and identification model is evaluated using the test set sample;

[0092] Step S2. Establishing a prediction model of the coal combustion carbon emission factor CEF, the elemental carbon content C, and the low calorific value NCV (i.e. low heat value): a particle swarm optimization-support vector regression (PSO-SVR) combined with a Sequential Backward Selection (SBS) algorithm is used to establish a PSO-SVR prediction model of the coal combustion carbon emission factor CEF, the elemental carbon content C, and the low calorific value NCV for different coal types (e.g. bituminous coal, lignite, anthracite, etc.);

[0093] Step S3. Setting a checking threshold: according to the prediction results of the established PSO-SVR prediction model of the coal combustion carbon emission factor CEF, the elemental carbon content C, and the low calorific value NCV, the absolute error of the prediction model is further analyzed, and the checking threshold of each is determined as a basis for determining the reasonableness of the carbon emission data in the data checking process;

[0094] Step S4. Setting a three-level data checking rule: a third-party checking institution can calculate the error between the predicted value and the measured value calculated according to the coal combustion carbon emission factor CEF prediction model, the error between the predicted value and the measured value calculated according to the elemental carbon content C prediction model, and the error between the predicted value and the measured value calculated according to the low calorific value NCV prediction model according to the data reported by the enterprise, and compare them with the corresponding carbon emission factor CEF, elemental carbon content C, and low calorific value NCV prediction model checking threshold in the checking threshold determination module, so as to determine whether the carbon emission factor data reported by the enterprise is abnormal, and if it is determined to be abnormal, further identify the source of the abnormal data, and replace it with the corresponding default value for recalculation. The three-level data checking rule can automatically check the reasonableness of the coal combustion carbon emission factor data of the enterprise, identify abnormal data in time, and determine the quality of the carbon emission accounting data based on this.

[0095] The method for selecting the K-nearest neighbor algorithm to establish the coal type classification and identification model in step S1 is as follows: a training set data T is given and a K value is set;

[0096] The training set data T is represented as:

[0097] T={(x1,y1),(x2,y2),…(x i ,y i )}(i=1,2,…,N);

[0098] In the formula, x iThe coal quality analysis parameter data (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)) of the coal sample, calorific value (Q), hydrogen content (H), and total sulfur content (St) are used as model input variables; y i The category of the coal sample in the training set, which has m categories, and N is the number of samples;

[0099] In the K-nearest neighbor algorithm, the value of the number of nearest neighbors (K) affects the accuracy of the final result. If the value of K is too large, overfitting will occur, and if the value of K is too small, underfitting will occur. Therefore, the K-fold cross-validation method is used to find the most suitable K value. The steps for determining the optimal K value in the K-nearest neighbor algorithm using the K-fold cross-validation method are as follows:

[0100] a1. Set the range of K values: First, determine a range of K values, usually starting from 1, and the upper limit can be set to 20 or a smaller number;

[0101] a2. Divide the data set: Divide the entire data set into K subsets evenly, and each subset is used as the test set in turn during cross-validation, and the remaining K-1 subsets are used as the training set;

[0102] a3. Model training and validation: For each K value, repeat the following steps K times: use the data of K-1 subsets to train the KNN model, and use the remaining one subset as the test set to evaluate the model performance;

[0103] a4. Calculate the average accuracy: For each K value, calculate the average accuracy in K cross-validation;

[0104] a5. Determine the optimal K value: Compare the average accuracy of different K values, and select the K value with the highest accuracy as the optimal K value.

[0105] Preferably, the method for determining the category of the unknown coal sample in step S1 is as follows:

[0106] b1. Calculate the distance between each training set sample and the unknown coal sample, and the calculation formula is:

[0107]

[0108] where x i is the coal quality analysis parameter data of the i-th training set sample; x j is the coal quality analysis parameter data of the j-th unknown coal sample; x i , x j ∈R n; l is the ordinal number of the coal quality analysis parameter, from 1 to n; n is the total number of the coal quality analysis parameters; p is a variable parameter, p≥1, and according to the different parameters p, different types of distances can be represented, when p=1, Manhattan distance, p=2, Euclidean distance; if the feature space is multi-dimensional and each dimension is independent of each other, Manhattan distance is selected; when the feature space is continuous and each dimension is related to each other, and the straight-line distance between data points needs to be reflected, Euclidean distance is selected;

[0109] b2. Statistics of the data points M of the K closest coal training set samples to the data x of the unknown coal sample k (x);

[0110] b3. Voting the most in the data points M k (x) of the K closest neighbor training set samples as the final output category of the unknown coal sample, and the calculation formula is:

[0111]

[0112] In the formula, f is a judgment indication function, y i is the actual coal training set sample category, c j is the coal type category predicted by the coal type classification and identification model, when y i =c j , f=1, otherwise f=0.

[0113] Preferably, the step S1 further comprises a model performance evaluation step: selecting the classification accuracy (Accuracy) as the evaluation index of the coal type classification and identification model, to evaluate the performance of the coal type classification and identification model, in general, the higher the classification accuracy, the better the recognition effect of the classification model, and the calculation formula is as follows:

[0114]

[0115] In the formula, Accuracy represents the classification accuracy of the sample coal type, T is the number of correct classification, and F is the number of wrong classification.

[0116] In the step S2, the establishment of the prediction model of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV includes the following steps:

[0117] c1. Data preprocessing: first, give the training sample set q:

[0118] q= {(x1,y1),( x2,y2),...,(x i ,y i ),...,( x n ,y n )}

[0119] wherein xi is an input feature vector, denoted as coal quality analysis parameter data; and yi is an output feature vector, denoted as a coal-fired carbon emission factor or elemental carbon content or low calorific value;

[0120] Next, normalization processing is performed on all data of the training set of the coal sample according to the following formula:

[0121]

[0122] wherein X ik is a data normalization value; x ik is an original data value; x k min is a minimum value of the original data; x k max is a maximum value of the original data;

[0123] c2. Establishing an initial full feature model: first, a set of coal samples with known coal quality analysis parameter index values (including total water (Mt), industrial analysis (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), total sulfur content (St), etc.) are used for calibration, and preliminary SVR models (support vector regression models) of the full features of the coal-fired carbon emission factor CEF, elemental carbon content C, and low calorific value NCV of different coal types (such as bituminous coal, lignite, anthracite, and mixed coal, etc.) are established, respectively. The coal quality analysis data of the coal samples used for calibration need to be standard coal data, or coal sample data detected by a CMA-recognized or CNAS-recognized detection institution / laboratory. The preliminary SVR model maps the nonlinear feature vector in the data sample from a low-dimensional space to a high-dimensional feature space through a kernel function and then fits a linear regression function in the feature space:

[0124]

[0125] wherein ω is a weight vector; and b is a bias constant; is a nonlinear mapping from the input space to the high-dimensional feature space;

[0126] The preliminary SVR model defines a loss function ε, and the model is optimized by the minimum value of the function to find the regression function with the best parameter value, as shown in the following formula:

[0127]

[0128] The constraint condition is:

[0129]

[0130] where c is a penalty factor; ε is a loss function; b is a bias constant; ξ i , ξ' i are slack variables; ω is a weight vector; x i is an input feature vector, y i is an output feature vector;

[0131] After operation, the regression function can be obtained as shown in the following formula:

[0132]

[0133] where β i , α i is a Lagrange multiplier, K(x g x i ) is a radial basis kernel function of the support vector machine, wherein the radial basis kernel function has good generalization performance and involves fewer parameters, and has certain advantages in parameter optimization, and the radial basis kernel function is shown in the following formula:

[0134]

[0135] where g is a kernel function parameter, representing the mean square deviation of the Gaussian function, i.e., the width of the function in the independent variable direction; when g is small, the fitting performance of the function to the data is good; when g is large, the generalization ability of the function is strong; therefore, the determination of the optimal parameter directly affects the SVR prediction model of the radial basis kernel function type of the coal-fired combustion carbon emission factor, and in order to further improve the prediction performance of the SVR model, an effective optimization algorithm should be found to optimize the parameters of the SVR model;

[0136] c3. Model parameter optimization: for the radial basis kernel function of the SVR (support vector machine), the penalty factor c and the kernel parameter g are the main parameters affecting the performance of the SVR, and the accuracy and reliability of the SVR prediction result depend on the best choice of (c, g); therefore, the particle swarm optimization algorithm (particle swarm optimization, PSO) is used to optimize the penalty factor c and the kernel parameter g of the radial basis kernel function of the SVR, and the PSO-SVR model for predicting the coal-fired combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV is constructed;

[0137] The steps of the PSO algorithm in step c3 for iteratively finding the best SVR parameters are as follows:

[0138] c31. Initialize the PSO-SVR model parameters, including: the number of iterations, the total number of particles, the learning factors c1 and c2, the search range of the penalty factor c and the kernel parameter g, and assume that the position of the i-th particle in the D-dimensional space in the particle swarm is represented as x i = (x i1x i2 ,···,x ij ,···,x id ), particle velocity is v i =(v i1 ,v i2 ,···,v ij ,···,v id ), determine the single iteration displacement of the particle in the search process;

[0139] c32. Learning the training samples by using SVR, taking the root mean square error (RMSE) of the prediction results of the training samples of the coal-fired power plant as the fitness function, updating the velocity and position of the population particles; comparing the current fitness value of the particle with the optimal fitness value of the particle, if the current value is better, then the current position is taken as the optimal position of the particle;

[0140] c33. Updating the velocity and position of the particle by using the following two formulas:

[0141]

[0142] In the formula, k is the current iteration number, c1 and c2 are non-negative learning factors, r1 and r2 are random numbers distributed in the interval (0, 1), vi k represents the velocity of the particle i at the kth iteration in the d-dimensional space, pi k represents the individual extreme value, and xi k represents the position, is the global extreme value after k iterations, respectively represent the velocity and position of the particle i at the (k+1)th iteration in the d-dimensional space; ω is an inertia factor, representing the global search ability, which is defined as:

[0143] ω = ω min +(iter max -iter)·(ω max -ω min ) / iter max

[0144] In the formula, ω max and ω min are the maximum and minimum weight factors, respectively, iter is the current iteration number, and iter max is the maximum iteration number.

[0145] c34. Judging whether the optimization termination conditions such as the set maximum iteration number or the accuracy target are met, if yes, then the optimal solution is obtained, and if not, then going to step c32;

[0146] ​The final output optimal position vector (c, g) is the optimal PSO-SVR coal combustion carbon emission factor CEF, element carbon content C, and low calorific value NCV model parameters: penalty factor (bestc) and kernel function parameter (bestg).

[0147] c4. Feature selection: based on the SBS algorithm, combined with PSO-SVR, a feature is removed from the current feature set, and a new model is established to calculate the model performance (the root mean square error of prediction (RMSE) is used as the model performance evaluation index); each parameter in the full feature set is traversed, and one parameter is removed from the full feature set in each round, and a new model is established to calculate the model performance; then the performance difference between the new model and the original model is compared, if the model performance does not decrease significantly, the feature is removed; otherwise, the feature is retained; by repeatedly repeating the above steps, irrelevant features are gradually removed, until the performance of the model no longer improves, thereby selecting the most relevant feature subset of the carbon emission factor, and using the finally selected feature subset to construct the PSO-SVR model, i.e. the optimal model of the coal combustion carbon emission factor, the element carbon content, and the low calorific value; reducing irrelevant features can reduce the generalization error of the model and improve the prediction ability of the model.

[0148] c5. Model performance evaluation: the above feature selection uses the root mean square error (RMSE) to evaluate the performance of the model, and under normal circumstances, the smaller the RMSE, the higher the prediction accuracy of the model, and the calculation formula is:

[0149]

[0150] In the formula, y i , and y i are the reference value and the model predicted value of the coal quality parameter respectively, n represents the sample number of the coal sample, and i is the number of independent variables.

[0151] The mean absolute error (MAE) is used as the basis for setting the verification threshold in step S3; the smaller the mean absolute error (MAE), the higher the prediction accuracy of the model, and the calculation formula is:

[0152]

[0153] In the formula, y i , and y i are the reference value and the model predicted value of the coal quality parameter respectively, n represents the sample number of the coal sample, and i is the number of independent variables.

[0154] The three-level data verification process of step S4 includes the following steps:

[0155] S41. Primary data verification: According to the coal quality data reported by the enterprise, the measured and calculated carbon emission factor CEF of coal-fired combustion is obtained by using the measured low calorific value NCV and elemental carbon content C, and then the model prediction value of the carbon emission factor CEF is obtained according to the carbon emission factor CEF prediction model established in step S3. The error between the model prediction value and the measured and calculated value of the carbon emission factor CEF is calculated. If the error is less than the preset threshold, the verification is passed, the carbon emission factor CEF data is automatically added to the database system and dynamically updated. If the error is greater than the preset threshold, it enters the secondary data verification;

[0156] S42. Secondary data verification: According to the coal quality data reported by the enterprise, the model prediction value of the elemental carbon content is obtained by using the elemental carbon content C prediction model established in step S3. The error between the model prediction value and the measured value of the elemental carbon content is calculated. If the error is greater than the preset threshold, it indicates that the abnormal source of the carbon emission factor is the elemental carbon content, and the elemental carbon content is marked as abnormal and the enterprise manager is notified. The measured data of the elemental carbon content is doubtful, and the unit heat value carbon content default value is used to replace it to calculate the carbon emission factor. If the error is less than the preset threshold, it enters the tertiary data verification;

[0157] S43. Tertiary data verification: According to the coal quality data reported by the enterprise, the model prediction value of the low calorific value is obtained by using the low calorific value NCV prediction model established in step S3. The error between the model prediction value and the measured value of the low calorific value is calculated. If the error is greater than the preset threshold, it indicates that the abnormal source of the carbon emission factor is the low calorific value, and the low calorific value is marked as abnormal and the enterprise manager is notified. The measured data of the low calorific value is doubtful, and the guide low low calorific value default value is used to replace it to calculate the carbon emission factor. If the error is less than the preset threshold, the enterprise manager is notified, a detailed report of the abnormal coal quality data is provided or on-site verification is carried out for further analysis to determine the root cause of the abnormal carbon emission factor data.

[0158] It also includes step S5: model updating; In order to ensure the prediction accuracy of the prediction model of the carbon emission factor of coal-fired combustion, the elemental carbon content and the low calorific value, the model needs to be updated regularly, including collecting new coal quality data, retraining the model, etc. to adapt to the changes of coal types and combustion efficiency.

[0159] The specific implementation cases of the above-mentioned dynamic verification method of the fuel end carbon emission factor of the control emission enterprise are as follows:

[0160] 1. Establish a different coal type classification and identification model:

[0161] (1) Sample collection and processing: In this embodiment, 557 coal samples (including 461 bituminous coal samples and 96 lignite samples) are collected for coal quality analysis parameters (including total water (Mt), industrial analysis (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), total sulfur content (St), etc.), which are all coal sample data detected by a detection institution certified by CMA. According to the proportion of the training set (80%) and the test set (20%), 446 samples are randomly selected as the training set samples for the coal type classification and identification model, and the remaining 111 samples, including 92 bituminous coal samples and 19 lignite samples, are used as the test set samples to test the performance of the model and verify the effect of classification and identification. Different labels are set for different types of coal quality for differentiation, and the same labels are set for the training set and test set samples of the same type. In this embodiment, the label of bituminous coal is set to 1, and the label of lignite is set to 2. Then the data is standardized to ensure that different features are on the same scale.

[0162] (2) Coal type classification and identification model establishment: K nearest neighbor algorithm (KNN) is selected to establish the coal type classification and identification model, i.e. KNN model, and coal quality analysis parameter data (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), heat value (Q), hydrogen content (H), and total sulfur content (St)) are used as input variables. The ten-fold cross-validation method is used to find the most suitable number of nearest neighbors K as 1, and the parameters of the KNN model are finally set as K value 1 and distance function Euclidean distance. The modeling of 446 training set samples is performed, and the obtained model has good effect, as shown in Table 1. The identification accuracy rates of bituminous coal and lignite are 99.2% and 97.4%, respectively. Among the 369 bituminous coals, only 3 are misjudged, and 366 are correct. Among the 77 lignites, only 2 are misjudged, and 75 are correct. The overall average identification accuracy is 98.2%, indicating that the coal type classification and identification model has good effect.

[0163] Table 1 Prediction results of KNN coal type classification and identification model

[0164]

[0165] (3) Model testing: The performance of the coal type classification and identification model is evaluated using 111 test set samples, and the results are shown in Table 2. The identification accuracy rates of bituminous coal and lignite are 98.9% and 94.7%, respectively. Among the 92 bituminous coals, only 1 is misjudged, and 92 are correct. Among the 19 lignites, only 1 is misjudged, and 18 are correct. The overall average identification accuracy is 98.2%, indicating that the coal type classification and identification model has good effect.

[0166] Table 2 Prediction results of KNN coal type classification and identification model

[0167]

[0168] In actual application, when the coal type of an enterprise is unknown or uncertain, the coal type can be automatically judged by the above-mentioned coal type classification and identification model according to the coal quality data reported by the enterprise.

[0169] 2. A prediction model of a coal combustion carbon emission factor (CEF), an element carbon content (C) and a low calorific value (NCV) is established.

[0170] (1) Data preprocessing: first, the data of 461 coal samples detected by a CMA-recognized detection institution are normalized according to the following formula:

[0171]

[0172] In the formula, X is a data normalized value; x is an original data value; x is a minimum value of the original data; and x is a maximum value of the original data. ik ik kmin k max

[0173] (2) An initial full-feature model is established.

[0174] First, 461 coal samples with known coal quality analysis parameter index values (including total water (Mt), industrial analysis (moisture (M), ash (A), volatile matter (V), fixed carbon (FC)), calorific value (Q), hydrogen element content (H), total sulfur element content (St) and the like) are trained to respectively establish preliminary SVR models of the full features of the coal combustion carbon emission factor CEF, the element carbon content C and the low calorific value NCV of different coal types (taking bituminous coal and lignite as examples in this embodiment).

[0175] (3) Model parameter optimization:

[0176] ​​​​The PSO algorithm is used for optimizing and selecting the penalty factor c and the kernel parameter g in the preliminary SVR model parameters of the coal combustion carbon emission factor CEF, the element carbon content C and the low calorific value NCV, and the PSO-SVR models for predicting the coal combustion carbon emission factor CEF, the element carbon content C and the low calorific value NCV are respectively constructed; the PSO particle swarm parameters are initialized, the total number of particles is set to 20, the iteration number is set to 100, the learning factor c1 is set to 1.5, the learning factor c2 is set to 1.7, c is in the range of 1 to 100, g is in the range of 0.1 to 100, and each particle position and speed is randomly initialized; then the fitness function (in the embodiment, the fitness function is the root mean square error (RMSE) of the prediction result of the coal training sample) of each particle is calculated, the iteration optimization of the fitness of all particles is performed, the global optimal fitness individual is obtained, the particle position and speed are updated, and the latest particle position and speed are calculated through the following formula:

[0177] v ij =w·v ij +c1r1(p ij -x ij )+c2r2(g i -x ij )

[0178] x ij =x ij +v ij

[0179] The optimization iteration is performed in the local and global aspects, the particle with the optimal fitness is obtained, the optimal position vector (c, g) corresponding to the optimal SVR model penalty factor (best c) and the kernel function parameter (best g) is obtained, and finally the optimal parameters of the coal combustion carbon emission factor CEF, the element carbon content C and the low calorific value NCV model are obtained as shown in Table 3.

[0180] Table 3 Optimal parameters of the PSO-SVR model

[0181]

[0182]

[0183] (4) Feature selection:

[0184] The SBS algorithm is used to remove a feature from the current feature set, a new model is established, the model performance (in this embodiment, the model performance evaluation index is the root mean square error (RMSE)) is calculated, and the performance difference between the new model and the original model is compared. If the model performance does not decrease significantly, the feature is removed; otherwise, the feature is retained. By repeatedly performing the above steps, the features are gradually removed

[0185] Except for irrelevant features, until the performance of the model no longer improves. Finally, the optimal input variables of the carbon emission factor, elemental carbon content, and low calorific value of bituminous coal and lignite combustion are shown in Table 4.

[0186] Table 4 Optimal input variables of PSO-SVR model

[0187]

[0188] (5) Optimal model: using the final selected optimal input variables and the best model parameters to build the PSO-SVR prediction model, the optimal model of the carbon emission factor, elemental carbon content, and low calorific value of bituminous coal and lignite combustion is obtained, and the RMSE results of the model are shown in Table 5.

[0189] Table 5 RMSE results of PSO-SVR model

[0190] Parameters Bituminous coal Lignite CEF (tCO2 / GJ) 0.1031 0.1092 C(%) 0.5135 0.4364 NCV (GJ / t) 0.3602 0.2259

[0191] 3, Set the threshold for verification

[0192] According to the prediction results obtained by the built PSO-SVR prediction model of the carbon emission factor, elemental carbon content, and low calorific value of coal combustion, the absolute error value of the PSO-SVR prediction model is further determined, and the threshold for verification of the PSO-SVR prediction model of the carbon emission factor, elemental carbon content, and low calorific value of coal combustion is determined as the basis for determining the reasonableness of the carbon emission data in the data verification process.

[0193] (1) Absolute error value of carbon emission factor

[0194] As shown in Table 6, for bituminous coal, in the 369 calibration sets, the absolute error value of the model prediction value and the measured and calculated value of the carbon emission factor is ≤MAE (0.07tCO2 / GJ) in 246, accounting for 66.67%; ≤2MAE (0.14tCO2 / GJ) in 322, accounting for 87.26%; ≤3MAE (0.21tCO2 / GJ) in 348, accounting for 94.31%; ≤4MAE (0.28tCO2 / GJ) in 361, accounting for 97.83%. For lignite, in the 77 calibration sets, the absolute error value of the model prediction value and the measured and calculated value of the carbon emission factor is ≤MAE (0.08tCO2 / GJ) in 42, accounting for 54.55%; ≤2MAE (0.16tCO2 / GJ) in 68, accounting for 88.31%; ≤3MAE (0.24tCO2 / GJ) in 75, accounting for 97.40%; ≤4MAE (0.32tCO2 / GJ) in 77, accounting for 100%.

[0195] (2) Absolute error value of elemental carbon content

[0196] As shown in Table 6, for bituminous coal, in 369 calibration sets, the absolute error of the model predicted value and the measured value of the elemental carbon content was ≤MAE (0.36%) in 245, accounting for 66.40%; ≤2MAE (0.72%) in 323, accounting for 87.53%; ≤3MAE (1.08%) in 350, accounting for 94.85%; ≤4MAE (1.44%) in 359, accounting for 97.29%. For lignite, in 77 calibration sets, the absolute error of the model predicted value and the measured value of the elemental carbon content was ≤MAE (0.33%) in 45, accounting for 58.44%; ≤2MAE (0.66%) in 62, accounting for 80.52%; ≤3MAE (0.99%) in 75, accounting for 97.40%; ≤4MAE (1.32%) in 77, accounting for 100%.

[0197] (3) Absolute error of low calorific value

[0198] As shown in Table 6, for bituminous coal, in 369 calibration sets, the absolute error of the model predicted value and the measured value of the low calorific value was ≤MAE (0.24 GJ / t) in 247, accounting for 66.94%; ≤2MAE (0.48 GJ / t) in 327, accounting for 88.62%; ≤3MAE (0.72 GJ / t) in 356, accounting for 96.48%; ≤4MAE (0.96 GJ / t) in 361, accounting for 97.83%. For lignite, in 77 calibration sets, the absolute error of the model predicted value and the measured value of the low calorific value was ≤MAE (0.16 GJ / t) in 45, accounting for 58.44%; ≤2MAE (0.32 GJ / t) in 68, accounting for 88.31%; ≤3MAE (0.48 GJ / t) in 74, accounting for 96.10%; ≤4MAE (0.64 GJ / t) in 75, accounting for 97.40%.

[0199] In summary, the third-party verification agency can select the calibration threshold of the carbon emission factor, the low calorific value and the elemental carbon content between 2RMSE~3RMSE according to the verification requirements, as the basis for determining the reasonableness of the carbon emission data in the data calibration process.

[0200] Table 6 Model absolute error ratio results

[0201]

[0202]

[0203] 4. Data calibration

[0204] In order to ensure the rationality and accuracy of the coal-fired carbon emission factor data reported by enterprises, the third-party verification agency can automatically check the rationality of the coal-fired carbon emission factor data of enterprises according to the data reported by enterprises, can identify abnormal data in time, and can grasp the data quality of carbon emission calculation according to the same.

[0205] (1) First-level data checking: ten test samples are used in this embodiment, including five bituminous coals (samples with serial numbers T1 to T5) and five lignite coals (samples with serial numbers T6 to T10). First, the first-level checking of the coal-fired carbon emission factor data reported by enterprises is performed. According to the coal quality data reported by enterprises, the measured low calorific value NCV and the elemental carbon content C are used to calculate the CEF measured value, and then the CEF prediction model of coal combustion carbon emission factor built above is used to obtain the prediction value of CEF, and the error between the model prediction value and the measured calculation value of CEF is calculated. If the error is less than the preset threshold value, the verification is passed, the CEF data is automatically added to the database system, and is dynamically updated. If the error is greater than the preset threshold value, it enters the second-level checking. The threshold value of CEF in this embodiment is MAE (0.07 tCO2 / GJ for bituminous coal and 0.08 tCO2 / GJ for lignite), and the first-level data checking result is shown in Table 7.

[0206] Table 7 First-level data checking result

[0207]

[0208] According to Table 7, the error of five samples in the ten test set samples is less than the preset threshold value, so the checking is passed, and the error of the remaining five samples is greater than the preset threshold value, so it needs to enter the second-level data checking.

[0209] (2) Second-level data checking: the elemental carbon content prediction model built above is used to predict the elemental carbon content of the five test set samples entering the second-level data checking, to obtain the model prediction value of the elemental carbon content; the error between the model prediction value and the measured value of the elemental carbon content is calculated. If the error is greater than the preset threshold value, it indicates that the abnormal source of the carbon emission factor is the elemental carbon content, the elemental carbon content is marked as abnormal, and the enterprise manager is notified that the measured data of the elemental carbon content is suspicious and needs to be replaced by the default value of the elemental carbon content in the guide to calculate the carbon emission; if the error is less than the preset threshold value, it enters the third-level data checking.

[0210] The threshold value of the elemental carbon content in this case is MAE (0.36% for bituminous coal and 0.33% for lignite), and the second-level data checking result is shown in Table 8.

[0211] Table 8 Second-level data checking result

[0212]

[0213] According to Table 8, the errors of three of the five test set samples are greater than the preset threshold, so it is indicated that the abnormal source of the carbon emission factor is the element carbon content, the element carbon content is marked as abnormal, and the person in charge of the enterprise is notified; the errors of the remaining two samples (T5 and T10) are less than the preset threshold, and then the three-level checking is entered.

[0214] (3) Three-level data checking: the low calorific value NCV prediction model is used to predict the low calorific value of the two test set samples entering the three-level checking, and the low calorific value prediction value is obtained. The error between the model prediction value and the measured value of the low calorific value is calculated. If the error is greater than the preset threshold, it indicates that the abnormal source of the carbon emission factor is the low calorific value, the low calorific value is marked as abnormal, and the person in charge of the enterprise is notified that the low calorific value measured data is questionable and needs to be replaced by the guide low calorific value default value for carbon emission calculation; if the error is less than the preset threshold, the person in charge of the enterprise is notified, a detailed report of abnormal coal quality data is provided or on-site verification is carried out, and further in-depth analysis is carried out to determine the root cause of the abnormal carbon emission factor data.

[0215] The threshold value of the low calorific value in this case is MAE (0.24 GJ / t for bituminous coal and 0.16 GJ / t for lignite), and the three-level data checking result is shown in Table 9.

[0216] Table 9 Three-level data checking result

[0217]

[0218] According to Table 9, the error of T5 is less than the preset threshold, so the person in charge of the enterprise is notified, a detailed report of abnormal coal quality data is provided or on-site verification is carried out, and further in-depth analysis is carried out to determine the root cause of the abnormal carbon emission factor data; the error of T10 is greater than the preset threshold, indicating that the abnormal source of the carbon emission factor is the low calorific value, and the low calorific value is marked as abnormal and the person in charge of the enterprise is notified.

[0219] Example 2

[0220] The embodiment discloses a dynamic checking system for fuel end carbon emission factor of a control emission enterprise, which is formed by the method described in the above embodiment, and the system comprises:

[0221] a classification model generation module, a prediction model generation module, a checking threshold determination module, and a checking module;

[0222] The classification model generation module is used for sample collection and processing, collecting coal quality analysis parameter data of different coal types, dividing the coal quality analysis parameter data into a training set and a test set, setting different labels for different types of coal to distinguish them, selecting a K nearest neighbor algorithm to establish a coal type classification and identification model according to the training set, then performing coal type classification and identification, and finally using the test set sample to evaluate the performance of the established coal type classification and identification model.

[0223] The prediction model generation module is used for combining a particle swarm optimization support vector machine with a backtracking sequential selection algorithm to respectively establish prediction models of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV of different coal types.

[0224] The threshold value determination module is used for further analyzing the absolute error of the prediction model according to the prediction results of the established prediction models of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV, respectively determining the threshold values of the three as the basis for determining the reasonableness of the carbon emission data in the data checking process.

[0225] The checking module is used for a third-party checking institution to calculate the error between the predicted value and the measured value calculated according to the coal combustion carbon emission factor CEF prediction model, the error between the predicted value and the measured value calculated according to the elemental carbon content C prediction model, and the error between the predicted value and the measured value calculated according to the low calorific value NCV prediction model according to the data reported by the enterprise, and respectively compare them with the threshold values of the carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV prediction model in the threshold value determination module, so as to determine whether the carbon emission factor data reported by the enterprise is abnormal, and if it is determined to be abnormal, further identify the source of the abnormal data, and replace it with the corresponding default value for recalculation.

[0226] The above describes in detail the embodiments of the fuel end carbon emission factor dynamic checking method and method for controlling the carbon emission of an enterprise. This paper applies specific examples to describe the principles and implementation methods of the present application. The above description of the embodiments is only used to help understand the core idea of the present application. It should be pointed out that for ordinary skilled persons in the technical field, without departing from the principles of the present application, the present application can be improved and modified, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for dynamically calibrating the fuel end carbon emission factor of a controlled emission enterprise, characterized in that, Comprising the following steps: Step S1. First, sample collection and processing are carried out, coal quality analysis parameter data of different coal types are collected, the coal quality analysis parameter data is divided into a training set and a test set, different labels are set for different types of coal to distinguish them, according to the training set, a coal type classification and identification model is established by selecting a K nearest neighbor algorithm, then the coal type is judged, and finally the performance of the established coal type classification and identification model is evaluated using the test set samples; Step S2. A particle swarm optimization support vector machine combined with a backtracking sequential selection algorithm is used to establish PSO-SVR prediction models of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV of different coal types; Step S3. According to the prediction results obtained by the prediction model established in step S2, the absolute error of the prediction model is further analyzed, and the calibration threshold of the three prediction models is determined respectively; Step S4. A three-level data calibration rule is set: a third-party verification institution can calculate the error between the predicted values calculated according to the coal combustion carbon emission factor CEF prediction model, the elemental carbon content C prediction model and the low calorific value NCV prediction model and the respective measured calculated values according to the enterprise reported data, and compare them with the corresponding calibration thresholds in step S3, to determine whether the enterprise reported carbon emission factor data is abnormal, if it is determined to be abnormal, the source of the abnormal data is further identified, and the corresponding default value is used to replace it for recalculation; The three-level data calibration process comprises the following steps: S41. First-level data calibration: according to the coal quality data reported by the enterprise, the measured low calorific value NCV and the elemental carbon content C are used to calculate the measured calculated value of the coal combustion carbon emission factor CEF, then the model predicted value of the carbon emission factor CEF is obtained according to the coal combustion carbon emission factor CEF prediction model established in step S3, the error between the model predicted value and the measured calculated value of the carbon emission factor CEF is calculated, if the error is less than the preset threshold, the verification is passed, the carbon emission factor CEF data is automatically added to the database system and dynamically updated, if the error is greater than the preset threshold, the second-level data calibration is entered.

2. The method of claim 1, wherein the method further comprises: The method for selecting a K nearest neighbor algorithm to establish a coal type classification and identification model in step S1 is as follows: given a training set data T and a K value; The training set data T is represented as: T = {(X1, Y1), (X2, Y2),... (X m , Y m )... (X N , Y N )} (m = 1, 2,..., N) In the formula, X m is the coal quality analysis parameter data of the coal sample as a model input variable; Y m is the category of the coal training set sample, and there are w categories in total, and N is the number of samples; The steps for determining the optimal K value in the K nearest neighbor algorithm using K-fold cross-validation method are as follows: a1. Set the range of K value: start from 1, and set the upper limit to 20; a2. Divide the data set: evenly divide the entire data set into K subsets, each subset is used as the test set in turn during cross-validation, and the remaining K-1 subsets are used as the training set; a3. Model training and verification: for each K value, repeat the following steps K times: use the data of K-1 subsets to train the coal type classification and identification model, and use the remaining one subset as the test set to evaluate the model performance; a4. Calculate the average accuracy: for each K value, calculate the average accuracy in K cross-validation; a5. Determine the optimal K value: compare the average accuracy corresponding to different K values, and select the K value with the highest accuracy as the optimal K value.

3. The method of claim 2, wherein the method further comprises: The method for judging the category of the unknown coal sample in the step S1 is as follows: b1. Calculate the distance between each training set sample and the unknown coal sample, and the calculation formula is: In the formula, X m is the coal quality analysis parameter data of the mth training set sample; X j is the coal quality analysis parameter data of the jth unknown coal sample, l is the ordinal number of the coal quality analysis parameter, and takes a value from 1 to z; z is the total number of the coal quality analysis parameters; p is a variable parameter, and p≥1; p represents different types of distances according to the different parameters p; b2. Statistics out of the data X of the unknown coal sample K coal training set sample data points M closest to the distance of the data k (X); b3. Data points M of K nearest neighbor training set samples k (x) the class with the most votes in (x) as the final output class for the unknown coal sample, calculated by: In the formula, f is a judgment indication function, Y m is the category of the actual coal training set sample, c d is the coal category predicted by the coal category identification model, and f = 1 when Y m = c d , otherwise f = 0.

4. The method of claim 1, wherein the method further comprises: The step S1 further includes a model performance evaluation step: the classification accuracy is selected as the evaluation index of the coal category classification and identification model, and is used to evaluate the performance of the coal category classification and identification model, and the calculation formula is as follows: In the formula, Accuracy represents the classification accuracy of the sample coal category, T is the number of correct classification, and F is the number of incorrect classification.

5. The method of claim 1, wherein the method further comprises: In the step S2, the establishment of the prediction model of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV includes the following steps: c1. Data preprocessing: first, the training sample set q is given: q = {(xl,yl), (x2,y2),..., (xn,yn)} i i n n ​​​​ In the formula, x i is an input feature vector, denoted as coal quality analysis parameter data; y i is an output feature vector, denoted as a carbon emission factor of coal combustion or elemental carbon content or low calorific value. Then, the training set data of the coal sample is normalized according to the following formula: where X ik is the data normalized value; x ik is the original data value; x k min is the original data minimum value; x k max is the original data maximum value; c2. Establishing initial full feature model: first, a group of coal samples with known coal quality analysis parameter index values are used for calibration, and preliminary SVR models of full features of coal combustion carbon emission factor CEF, elemental carbon content C, and low calorific value NCV of different coal types are established, the preliminary SVR model is mapped from a low-dimensional space to a high-dimensional feature space through a kernel function The nonlinear feature vector in the data sample is mapped from a low-dimensional space to a high-dimensional feature space, and then a linear regression function is fitted in the feature space: In the formula, ω is a weight vector; b is a bias constant; is a nonlinear mapping from the input space to the high-dimensional feature space; The loss function ε is defined in the preliminary SVR model, and the regression function with the optimal parameter value is found by optimizing the parameters of the model through the minimum value of the function, as shown in the following formula: The constraint condition is: where c is a penalty factor; ε is a loss function; b is a bias constant; ξ i ’ i is a slack variable; ω is a weight vector; x i is an input feature vector, y i is an output feature vector;​ After operation, the linear regression function is as follows: where β i , α i are Lagrange multipliers, K(x x i ) is a radial basis kernel function of the support vector machine, and the radial basis kernel function is as follows: Wherein, g is the kernel function parameter, representing the mean square deviation of the Gaussian function; c3. Model parameter optimization: the parameters of the preliminary SVR model are optimized, and for the SVR of the radial basis kernel function, the particle swarm optimization algorithm is used to optimize and select the penalty factor C and the kernel parameter g of the SVR of the radial basis kernel function, and the PSO-SVR models for predicting the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV are constructed; c4. Feature selection: based on the SBS algorithm, the PSO-SVR is combined to remove a feature from the current feature set, and a new model is established to calculate the model performance; then the performance difference between the new model and the original model is compared, if the model performance does not decrease significantly, the feature is removed; otherwise, the feature is retained; by repeatedly repeating the above steps, irrelevant features are gradually removed, until the performance of the model no longer improves, so that the most relevant feature subset of the carbon emission factor, the elemental carbon content and the low calorific value is selected, and the PSO-SVR model is constructed using the finally selected feature subset, and the optimal models of the coal combustion carbon emission factor, the elemental carbon content and the low calorific value are obtained; c5. Model performance evaluation: the root mean square error RMSE is used to evaluate the performance of the model in the above feature selection, and the calculation formula is as follows: In the formula, y i , are the reference value and the model predicted value of the coal quality parameter, respectively.

6. The method of claim 1, wherein the method further comprises: In the step S3, the mean absolute error MAE is used as the basis for setting the check threshold; the smaller the mean absolute error MAE, the higher the prediction accuracy of the model, and the calculation formula is as follows: where y i , are the reference and model predicted values of the coal quality parameters, respectively, n represents the number of samples of the coal sample, and i is the i-th sample.

7. The method of claim 1, wherein the method further comprises: The three-level data checking process of the step S4 further includes the following steps: S42. Secondary data checking: According to the coal quality data reported by the enterprise, the model predicted value of elemental carbon content is obtained by using the elemental carbon content prediction model established in step S3, the error between the model predicted value and the measured value of elemental carbon content is calculated, if the error is greater than the preset threshold, it indicates that the abnormal source of the carbon emission factor is the elemental carbon content, mark the elemental carbon content as abnormal, and notify the enterprise responsible person that the measured data of the elemental carbon content is questionable, and the default value of the carbon content per unit heat value in the guide is used to replace it to calculate the carbon emission factor; if the error is less than the preset threshold, then enter the tertiary data checking; S43. Tertiary data checking: According to the coal quality data reported by the enterprise, the model predicted value of low calorific value is obtained by using the low calorific value prediction model ncv prediction model established in step S3, the error between the model predicted value and the measured value of low calorific value is calculated, if the error is greater than the preset threshold, it indicates that the abnormal source of the carbon emission factor is the low calorific value, mark the low calorific value as abnormal, and notify the enterprise responsible person that the measured data of the low calorific value is questionable, and the default value of the low calorific value in the guide is used to replace it to calculate the carbon emission factor; if the error is less than the preset threshold, then notify the enterprise responsible person, provide a detailed report of the abnormal coal quality data or conduct on-site verification, and further analyze to determine the root cause of the abnormal carbon emission factor data.

8. The method of claim 1, wherein the method further comprises: It also includes step S5: model updating: In order to ensure the prediction accuracy of the prediction models of the coal-fired combustion carbon emission factor, elemental carbon content and low calorific value, the model needs to be updated regularly, including collecting new coal quality data and retraining the model to adapt to the changes of coal types and combustion efficiency.

9. The method of claim 5, wherein the method further comprises: The steps of the PSO algorithm in step c3 for finding the optimal SVR parameters are as follows: c31. Initialize the PSO-SVR model parameters, including: the number of iterations, the total number of particles, the search range of learning factors c1and c2, the penalty factor c and the kernel parameter g, and assume that the position of the Ith particle in the particle swarm in the D-dimensional space is represented as x I = (x I1 , x I2 , ···, x IJ , ···, x Id ), the particle velocity is v I = (v I1 , v I2 , ···, v IJ , ···, v Id ), and the single iteration displacement of the particle in the search process is determined; c32. Using SVR to learn the training samples, taking the root mean square error RMSE of the prediction results of the training samples of the coal as the fitness function, updating the speed and position of the population particles; compare the current fitness value of the particle with the optimal fitness value of the particle, if the current value is better, then the current position is taken as the optimal position of the particle; c33. Update the speed and position of the particle with the following two formulas: where k is the current iteration number, c1, c2 are non-negative learning factors, and r1 and r2 are random numbers distributed in the interval (0, 1), respectively represent the velocity, individual extremum and position of particle I at the kth iteration in d-dimensional space, is the global extremum after k iterations, respectively represent the velocity and position of particle I at the (k+1)th iteration in d-dimensional space; ω is an inertia factor, representing the global search ability, which is defined as: ω = ω min + (iter max - iter) · (ω max - ω min ) / iter max where ω max and ω min are the maximum and minimum weight factors, respectively, iter is the current iteration number, and iter max is the maximum number of iterations. c34. Determine whether the set maximum iteration number or accuracy target optimization termination condition is met, if yes, then find the optimal solution, if not, then go to step c32; Finally output the optimal position vector (c, g), that is, obtain the optimal PSO-SVR coal-fired combustion carbon emission factor CEF, elemental carbon content C, and low calorific value NCV model parameters: penalty factor and kernel function parameter.

10. A dynamic verification system for carbon emission factors at the fuel end of controlled emission enterprises, characterized in that, It forms the system by the method of any one of claims 1-9, the system comprising: a classification model generation module, a prediction model generation module, a checking threshold determination module, and a checking module; The classification model generation module is used for sample collection and processing, collecting coal quality analysis parameter data of different coal types, dividing the coal quality analysis parameter data into a training set and a test set, setting different labels for different types of coal to distinguish them, selecting a K nearest neighbor algorithm to establish a coal type classification and identification model according to the training set, then performing coal type classification, and finally using the test set samples to evaluate the performance of the established coal type classification and identification model; The prediction model generation module is used for establishing prediction models of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV of different coal types by using a particle swarm optimization support vector machine combined with a backtracking sequential selection algorithm; The threshold determination module is used for further analyzing the absolute error of the prediction model according to the prediction results of the established prediction models of the coal combustion carbon emission factor CEF, the elemental carbon content C and the low calorific value NCV, and determining the threshold values of the three, which are used as the judgment basis for reasonable carbon emission data in the data checking process; The checking module is used for the third-party checking institution to calculate the errors between the predicted values and the measured values calculated according to the coal combustion carbon emission factor CEF prediction model, the errors between the predicted values and the measured values calculated according to the elemental carbon content C prediction model, and the errors between the predicted values and the measured values calculated according to the low calorific value NCV prediction model according to the data reported by the enterprise, and compare them with the threshold values of the corresponding carbon emission factors CEF, elemental carbon content C and low calorific value NCV prediction models in the threshold determination module, so as to judge whether the carbon emission factor data reported by the enterprise is abnormal, if it is judged to be abnormal, further identify the source of the abnormal data, and replace it with the corresponding default value for recalculation.

Citation Information

Patent Citations

  • Method for calculating dynamic carbon emission factors of multi-coal-electricity regional power grid

    CN118134693A

  • Metering-based carbon emission accounting system

    CN119443533A