CO2 emission prediction method based on PCA-WOA-GPR neural network model

By using the PCA-WOA-GPR neural network model and optimizing the Gaussian process regression model with principal component analysis and whale optimization algorithm, the problem of predicting CO2 emissions with high-dimensional features was solved, and higher accuracy prediction of CO2 emissions from coal-fired power plants was achieved.

CN120996364APending Publication Date: 2025-11-21SHENYANG INST OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511149847.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively predict CO2 emissions based on high-dimensional features, and the lack of objectivity and deep correlation in feature selection leads to low prediction accuracy.

Method used

A method for predicting CO2 emissions was constructed by using a PCA-WOA-GPR neural network model, extracting comprehensive indicators through principal component analysis, and optimizing the hyperparameters of the Gaussian process regression model using the whale optimization algorithm.

Benefits of technology

It reduces the input dimensionality, improves prediction accuracy and performance, and enables more accurate prediction of CO2 emissions from coal-fired power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996364A_ABST
    Figure CN120996364A_ABST
Patent Text Reader

Abstract

The invention relates to CO2 emission prediction, in particular to a CO2 emission prediction method based on a PCA-WOA-GPR neural network model. The method comprises the steps that S1, influence factor data influencing carbon dioxide emission in a coal-fired power plant are collected, and a basic database is formed; s2, selecting a predictive factor through correlation analysis, standardizing an original input sample, and performing feature extraction on predictive factor data by adopting a principal component analysis method to generate a comprehensive index; s3, inputting the comprehensive influence factors obtained after principal component analysis in the step S2 into a WOA-GPR network model; dividing a training set and a test set, and training and testing the training set and the test set to obtain corresponding carbon dioxide emissions. According to the method, the input dimension is reduced, the correlation of training samples is reduced, and higher precision and better performance are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to CO2 emission prediction, in particular to a CO2 emission prediction method based on a PCA-WOA-GPR neural network model. BACKGROUND

[0002] The low-carbon transformation of the power industry is a key industry for deep emission reduction to cope with climate change, and CO2 emission reduction in the power industry is of great significance to achieve long-term emission reduction targets. Controlling the CO2 emissions of the power industry must first account for the CO2 emissions in the power production process, which is conducive to the scientific and rational formulation of CO2 emission reduction targets by coal-fired power plants.

[0003] From the perspective of generation mechanism, CO2 emission concentration has correlation with numerous process variables, and there is also difference under different working conditions, which makes the high-dimensional characteristics of modeling data the primary problem faced by CO2 prediction modeling. A large amount of calculation is required to process high-dimensional characteristics, and feature selection is an effective strategy to reduce dimension, delete irrelevant data and improve learning accuracy. For complex industrial processes, only experts with long-term operating experience can select more relevant process variables based on their rich prior knowledge, but there is also a certain subjectivity and randomness. At present, research has not solved how to determine the threshold for feature selection according to data characteristics, and has not considered the deep relevance between features. Therefore, it is very necessary to design a CO2 emission prediction method for coal-fired power plants based on a PCA-WOA-GPR network model. SUMMARY

[0004] The application is aimed at the defects in the prior art and provides a CO2 emission prediction method based on a PCA-WOA-GPR neural network model.

[0005] To achieve the above-mentioned purpose, the application adopts the following technical scheme, a CO2 emission prediction method based on a PCA-WOA-GPR neural network model, comprising:

[0006] S1: Collecting the influence factor data of CO2 emission in a coal-fired power plant to form a basic database;

[0007] S2: Selecting a prediction factor through correlation analysis, standardizing the original input sample, extracting features from the prediction factor data by using a principal component analysis method, and generating a comprehensive index;

[0008] S3: Inputting the comprehensive influence factor obtained after the principal component analysis in S2 into a WOA-GPR network model; dividing a training set and a test set, and obtaining the corresponding CO2 emission through training and testing of the training set and the test set.

[0009] Further, in S1, the influence factor data of the carbon dioxide emission is 20-dimensional variable data, which is specifically divided into two categories: coal-fired power plant operation parameters and coal quality characteristic parameters;

[0010] The coal-fired power plant operation parameters include: unit load (Load), main steam temperature (T A ), main steam pressure (P A ), furnace outlet flue gas temperature (θ″ f ), excess air coefficient (α), primary air volume (L1), secondary air volume (L2), primary air temperature (T a ), secondary air temperature (T S ), flue gas temperature (T py ), desulfurizer consumption (W i ), desulfurization efficiency (η), flue gas oxygen content (O2), coal consumption for power supply (b g ), auxiliary power rate (W d );

[0011] The coal quality characteristic parameters include: low calorific value (Q net,ar ), fixed carbon (FC ar ), total moisture (M ar ), ash content (A ar ), volatile matter (V ar ).

[0012] Further, S2 is specifically divided into:

[0013] S2.1, correlation analysis is performed on the 20-dimensional variable data to determine whether it is suitable for principal component analysis;

[0014] S2.2, standardization processing is performed on the variable data passing through the correlation analysis;

[0015] S2.3, a correlation coefficient matrix is calculated based on the standardized data, and principal component analysis is performed;

[0016] S2.4, a comprehensive index is extracted according to the principal component variance contribution rate, and the input feature after dimension reduction is generated.

[0017] Further, S2.1 specifically includes: (Kaiser-Meyer-Olkin) KMO test and Bartlett's Sphericity Test are performed on the 20-dimensional variable data to determine whether the data is suitable for principal component analysis; wherein,

[0018] In the KMO test, the ratio of the partial correlation coefficient and the simple correlation coefficient between any two variables in the 20-dimensional variable data is evaluated by calculating the sampling adequacy statistic; when the KMO value is greater than a preset threshold (such as 0.6), it indicates that the overall correlation between variables is suitable for principal component analysis;

[0019] In the Bartlett's Sphericity Test, by testing whether the correlation coefficient matrix of the 20-dimensional variable data is significantly different from the unit matrix, when the significance level (P value) of the test is less than the preset threshold (0.05), the null hypothesis that the correlation coefficient matrix is the unit matrix is rejected, indicating that there is significant correlation in the data.

[0020] If both the KMO value is greater than the preset threshold and the Bartlett's test P value is less than the preset threshold, it is determined that the 20-dimensional variable data is suitable for principal component analysis.

[0021] Further, S2.2 specifically includes: for the 20-dimensional variable data determined to be suitable for principal component analysis by KMO and Bartlett's test in S2.1, performing standardization processing: for each dimension variable, the mean and standard deviation of all sample values are calculated respectively; the original value of each sample is subtracted from the mean value of the variable, and then divided by the standard deviation to obtain the standardized data; the processed data is used as the input for the calculation of the correlation coefficient matrix in S2.3.

[0022] Further, S2.3 specifically includes:

[0023] S2.3.1, calculating the sample correlation coefficient matrix r of the standardized 20-dimensional variable data, the expression is:

[0024]

[0025] Wherein, n is the number of samples, x ti and x tj are the standardized data of the i-th dimension and the j-th dimension of the t-th sample; r ij represents the correlation coefficient of the i-th variable and the j-th variable, p=20 is the total number of variables; i, j=1, 2, …20;

[0026] S2.3.2, the eigenvalues (λ1, λ2, …, λ p ) and the corresponding eigenvectors of the correlation coefficient matrix r are obtained as:

[0027] a i =(a i1 ,a i2 ,…a ip ), i=1, 2, …, p

[0028] S2.3.3, the eigenvalues and eigenvectors are used as the output of the principal component analysis, which is used for subsequent comprehensive index extraction.

[0029] Further, S2.4 specifically includes: calculating the contribution rate of each principal component according to the eigenvalues obtained in S2.3, the contribution rate being the proportion of the variance of the principal component to the total sum of the variances of all principal components, and the calculation formula being:

[0030]

[0031] The first k principal components with cumulative contribution rate of more than 85% are selected as the comprehensive index, and the standardized original data is projected to the selected principal component direction to generate the reduced comprehensive index.

[0032] Further, S3 is specifically divided into: (that is, the construction and optimization of the WOA-GPR network model specifically includes:)

[0033] S3.1, Gaussian process regression model construction;

[0034] Based on the k principal component comprehensive indexes obtained in S2.4 as input features, a Gaussian process regression model is constructed; wherein the covariance function of the model adopts a square exponential kernel function, which contains two to-be-optimized hyperparameters: variance parameter σ 2 The output amplitude is controlled, and the scale parameter θ controls the correlation strength between features;

[0035] S3.2, parameter optimization based on whale optimization algorithm;

[0036] The σ 2 and θ determined in S3.1 are taken as optimization objectives, and the whale optimization algorithm is used for parameter optimization, and the specific process is as follows:

[0037] (1) initialization stage: set the population size to N, and each individual represents a set of (σ 2 , θ) parameter combinations;

[0038] (2) take the root mean square error on the training set as the fitness function to evaluate the prediction performance of the current parameter combination;

[0039] (3) according to the fitness evaluation result, the following strategies are used to update the parameters:

[0040] When the fitness improves, the surrounding strategy is executed to finely search in the neighborhood of the current optimal parameters;

[0041] When the fitness stagnates, the random search strategy is executed to expand the parameter search range;

[0042] Periodically execute the spiral search strategy to maintain parameter diversity;

[0043] (4) when the optimal fitness value continuously iterates no longer improves or reaches the maximum iteration number, output the optimal (σ 2 , θ) parameter combination;

[0044] S3.3, the optimal hyperparameters obtained by S3.2 are substituted into the Gaussian process regression model, and the model training is completed using the training set data; finally, the test set data is input into the trained model, and the predicted value of CO2 emission is output, thereby completing the quantitative evaluation of the carbon emission of the coal-fired power plant.

[0045] Compared with the prior art, the present application has the following advantages.

[0046] The present application has good prediction effect for CO2 emission prediction, and compared with the traditional prediction network, the model reduces the input dimension, reduces the correlation of the training samples, has higher precision and better performance, and can be used as an effective method for CO2 emission prediction. With further research, the CO2 emission prediction method of the coal-fired power plant based on the PCA-WOA-GPR neural network model will be more widely applied. BRIEF DESCRIPTION OF DRAWINGS

[0047] The present application will be further described below in combination with the drawings and specific embodiments. The protection scope of the present application is not limited to the following content.

[0048] Figure 1 is the flow chart of the CO2 emission prediction method based on the PCA-WOA-GPR neural network model.

[0049] Figure 2 is the specific flow chart of the modeling process.

[0050] Figure 3 is the original feature WOA-GPR convergence process.

[0051] Figure 4 is the principal component feature WOA-GPR convergence process.

[0052] Figure 5 Comparison chart of original feature WOA-GPR prediction results.

[0053] Figure 6 Comparison chart of principal component feature WOA-GPR prediction results. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present disclosure will be described clearly and completely in combination with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present disclosure.

[0055] The terminology used in the disclosure herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in the description of the disclosure and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0056] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting."

[0057] For ease of understanding, the embodiments of the disclosure will be described in detail.

[0058] As Figure 1 shown, the coal-fired power plant CO2 emission prediction method based on the PCA-WOA-GPR network model includes the following steps:

[0059] Step 1, collect the influencing factor data of carbon dioxide emission in the coal-fired power plant to form a basic database;

[0060] Step 2, select the prediction factor through correlation analysis, then standardize the original input sample, use the principal component analysis method to extract the features of the prediction factor data, and obtain several comprehensive indexes.

[0061] Step 3, construct a prediction model of carbon dioxide emission of coal-fired power plant based on WOA-GPR network, and import the information data obtained after principal component analysis into the WOA-GPR network model, and divide the training set and test set, and obtain the corresponding carbon dioxide emission by training and testing the training set and test set.

[0062] One embodiment, the selected 20-dimensional variables affecting CO2 emission include: unit load (Load), main steam temperature (T A ), main steam pressure (P A ), low heat release (Q net,ar ), fixed carbon (FC ar ), total moisture (M ar ), ash content (A ar ), volatile matter (V ar ), furnace outlet flue gas temperature (θ″ f ), excess air coefficient (α), primary air volume (L1), secondary air volume (L2), primary air temperature (T a ) secondary air temperature (T S ), flue gas temperature (T py ), desulfurizer consumption (W i ), desulfurization efficiency (η), oxygen content in flue gas (O2), power supply coal consumption (b g ), auxiliary power rate (Wd ).

[0063] Correlation analysis. First, the sample data is KMO and Bartlett test, indicating that there is a multicollinearity between the independent variables, suitable for principal component analysis. KMO value for checking the correlation between variables, the closer to 1 KMO value, the stronger the correlation between variables, factor analysis is better; Bartlett test for determining whether the correlation matrix is a unit matrix, that is, whether the variables have a strong correlation, P value less than or greater than 0.5, subject to test, suitable for factor analysis.

[0064] Standardization of the original input sample. Because the carbon emissions of each influencing factor (i.e. the original input sample) has different dimensions and orders of magnitude, the difference between the data is large, so the original input sample is standardized before principal component analysis.

[0065] One embodiment is:

[0066]

[0067] Calculate the sample correlation coefficient matrix:

[0068]

[0069] For convenience, assume that the original data is still represented by X after standardization, then the correlation coefficient of the standardized data is:

[0070]

[0071] The eigenvalues (λ1, λ2, …, λp) of the correlation coefficient matrix R and the corresponding eigenvectors are: p

[0072] a i = (a i1 , a i2 , … a ip ), i = 1, 2, …, p (5)

[0073] Extract features to obtain several comprehensive indexes. Principal component analysis can obtain p principal components. However, since the variance of each principal component is decreasing, the amount of information contained is also decreasing, and in actual analysis, the first k principal components are usually selected according to the size of the cumulative contribution rate of each principal component, where the contribution rate refers to the proportion of the variance of a principal component to the total variance, which is actually the proportion of a characteristic value to the total characteristic value, that is:

[0074]

[0075] ​The greater the contribution rate, the stronger the information of the original variable contained in the principal component. The selection of the number of principal components k is mainly determined by the cumulative contribution rate of the principal components, that is, the cumulative contribution rate is generally required to reach more than 85%, so as to ensure that the comprehensive variable can include most of the information of the original variable.

[0076] An embodiment, as shown in Figure 2 , the establishment of whale shark algorithm optimization Gaussian process regression prediction model, including the following steps:

[0077] Gaussian process regression takes Gaussian process as model prior, which is expressed by the determination of mean function and covariance function, formula (7), wherein the mean function represents the proxy model of the target function, and the covariance function represents the uncertainty of the model:

[0078]

[0079] wherein m(x)=E[f(x)] is the mean function; is the covariance function;

[0080] The prior mean function is expressed as a linear combination of a set of basis functions:

[0081] m(x)=s T (x)β (8)

[0082] wherein s T (x)=[s1(x),s2(x),…,s p (x)] is a vector composed of basis functions; β is a regression parameter vector of p×1; the prior mean function takes 0;

[0083] The covariance function determines the mutual dependence between samples, which is expressed by a kernel function, and the expression is:

[0084]

[0085] wherein σ 2 is the variance parameter; is the correlation kernel, and θ is the parameter vector;

[0086] The square exponential covariance function is selected as the covariance function, and the formula is as follows:

[0087]

[0088] wherein θ is the scale parameter, representing the dependence degree of the distance between observation points; is the kernel function selected by the specific covariance function, which is a specific function of formula (9).

[0089] ​Based on the definition of Gaussian process prior, let the prior mean function be 0 and the prior covariance function be Then the training data where j = 1, 2, 3, …, k, k is the number of principal components, The probability distribution of the prediction data x = [x1, x2, …, x k ] is respectively expressed as:

[0090] F(X A ) ~ N(0, K + Δ) (11)

[0091] F(x) ~ N(0, k t (x, x)) (12)

[0092] where,

[0093] Based on formulas (11) and (12), the joint Gaussian distribution of the two is expressed as:

[0094]

[0095] where, k t (x) = k t (x, x), k t T (x) = [k t (x, x A m )] m=1,...,N ;

[0096] From the marginal distribution property of the joint Gaussian distribution, it is obtained that the prediction data x is subject to the distribution:

[0097]

[0098] where:

[0099]

[0100] Another embodiment, from formula (11), there are other necessary parameters and hyperparameters in the model, which are obtained by whale optimization algorithm, and the specific process is as follows:

[0101] Initialize the population size N of parameters, the spatial dimension dim, the number of iterations max_iter, the initial population of whales x i (i = 1, 2, …, N) and the initial position and fitness of the optimal whale.

[0102] Enter the iteration, first to the whole population set boundary condition processing, then to the population of each individual evolution processing, coefficient vector and p. The uniformly distributed decision random number is generated.

[0103] When the decision coefficient p < 0.5 and >= 1, the whale population is based on the random whale position solution, and the random walk foraging strategy is executed according to the following formula.

[0104]

[0105] When the decision coefficient p < 0.5 and < 1, the whale population is based on the current optimal solution, and the shrinking encircling prey strategy is executed according to the following formula.

[0106]

[0107] When the decision coefficient p >= 0.5, the whale population is based on the current optimal solution, and the spiral bubble net predation strategy is executed according to the following formula.

[0108]

[0109] Determine whether the end condition is met, if it is met, record the current optimal solution and its target value, if it is not met, return to step 2 for the next iteration. Finally, the optimal solution of the hyperparameter is obtained. Based on the principal component data obtained in step 2, the CO2 emission model F(·) of the coal-fired power plant obtained in step 3 is substituted, thereby obtaining N groups of CO2 emission data of the target coal-fired power plant, and realizing the prediction of CO2 emission.

[0110] The working mechanism and experimental analysis of the application are as follows:

[0111] I. From the perspective of data analysis, the application calculates the CO2 emission of the power plant through real-time operation data collection of the power plant. The coal combustion process, desulfurization process and purchased electricity process in the coal-fired power plant power production process are taken as the calculation emission boundary, and it is considered that the CO2 emission in the coal-fired power plant power production process is composed of two parts, which are the CO2 emission generated by combustion and the CO2 emission generated in the desulfurization process.

[0112] The CO2 emission generated by combustion is calculated by the coal consumption, received base carbon content of coal, ash content in coal and boiler thermal efficiency. The CO2 emission generated in the desulfurization process is calculated by the coal consumption, received base sulfur content of coal and boiler thermal efficiency. The application collects related parameters such as coal and boiler thermal efficiency under the field operation from 121MW to 319MW load for CO2 emission calculation.

[0113] II. Whale algorithm population size and iteration number value.

[0114] Population size refers to the number of whale individuals initialized in the whale optimization algorithm. Each whale individual represents a potential optimal solution to the optimization problem. A larger population size means that the algorithm has more "search agents" in the search space, which can cover a wider area and help find the global optimal solution, reducing the likelihood of getting stuck in a local optimum, but at the same time, it will increase the amount of calculation and calculation time.

[0115] The number of iterations refers to the number of times the whale optimization algorithm updates the position of the whale individual from the initial state according to its set rules and repeatedly executes the optimization process. The algorithm will evaluate the quality of each whale individual according to the fitness function in each iteration, and constantly update the optimal solution. When the maximum number of iterations is reached, the algorithm stops running and outputs the optimal solution found. The more iterations, the more opportunities the algorithm has to search and optimize, and it may find a better solution, but it will also consume more computing resources and time. If the number of iterations is too small, the algorithm may not have fully searched for the optimal solution before ending, resulting in an unsatisfactory result.

[0116] The present application sets the number of iterations to 10. During the algorithm running process, it is found that when the number of iterations increases from 10 to 20, the MSE value gradually decreases from 0.000503 to 0.000474, indicating that the algorithm is effectively optimizing the model. However, when the number of iterations increases from 20 to 50, the MSE value remains at 0.000474, indicating that the algorithm has converged to a stable solution, and subsequent iterations have not found a better parameter. Therefore, the number of iterations is selected as 20.

[0117] During the adjustment of the population size from 10 to 20, it is found that the convergence speed of the algorithm does not change significantly, and the optimal solution also does not change, so it is believed that the population size of 10 already meets the calculation requirements.

[0118] III. Technical effect analysis

[0119] For the enhancement effect of WOA-GPR algorithm on principal component PCA data processing, the present application carries out prediction with PCA data processing (such as Figure 4 ) and prediction without PCA data processing (such as Figure 3 ) under the same population size (10) and maximum number of iterations (50), and the results are as follows:

[0120] By comparing the two graphs, it can be seen that the WOA-GPR algorithm without using principal component analysis for data dimensionality reduction has a long-term MSE value of about 26400, indicating that the root mean square of the deviation between the predicted value and the actual value is high, the prediction error is large, and the model's fitting and prediction effect on the original data is extremely poor, and it cannot effectively capture the data regularity. While Figure 4The MSE in the model rapidly decreased from close to 2.0 to near 0, eventually stabilizing at an extremely low level. This indicates that the model's prediction error was significantly compressed, the predicted values ​​were highly close to the true values, and the accuracy of the fitted prediction was significantly improved.

[0121] pass Figure 3 , Figure 4 The comparison revealed that principal component analysis (PCA) performs exceptionally well in addressing the challenges of high correlation, large data dimensionality, and poor regression prediction accuracy in CO2 emission prediction data from coal-fired power plants. PCA effectively reduces data dimensionality and information overlap, thereby improving modeling accuracy and saving modeling time.

[0122] Figure 5 and Figure 6 The images show the WOA-GPR prediction results without principal component analysis and the WOA-GPR prediction results after using principal component analysis, respectively. This also verifies the effectiveness of principal component analysis in improving the prediction accuracy of the algorithm.

[0123] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "preferred embodiment," "detailed description," or "preferred embodiment," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0124] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Therefore, these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for predicting CO2 emissions based on a PCA-WOA-GPR neural network model, characterized in that, include: S1. Collect data on factors affecting carbon dioxide emissions from coal-fired power plants to form a basic database; S2. Select predictor factors through correlation analysis, standardize the original input samples, extract features from the predictor factor data using principal component analysis, and generate comprehensive indicators. S3. Input the comprehensive impact factor obtained from principal component analysis in S2 into the WOA-GPR network model; divide the model into training and testing sets, and obtain the corresponding carbon dioxide emissions by training and testing the training and testing sets.

2. The method for predicting CO2 emissions from coal-fired power plants based on the PCA-WOA-GPR network model according to claim 1, characterized in that, In S1, the data on factors affecting carbon dioxide emissions are 20-dimensional variable data, specifically divided into two categories: operating parameters of coal-fired power plants and coal quality characteristic parameters; The operating parameters of a coal-fired power plant include: unit load, main steam temperature, main steam pressure, flue gas temperature at furnace outlet, excess air coefficient, primary air volume, secondary air volume, primary air temperature, secondary air temperature, flue gas temperature, desulfurizing agent consumption, desulfurization efficiency, oxygen content in flue gas, coal consumption for power supply, and plant power consumption rate. Coal quality characteristics include: lower heating value, fixed carbon, total moisture, ash content, and volatile matter.

3. The method for predicting CO2 emissions from coal-fired power plants using the PCA-WOA-GPR network model according to claim 2, characterized in that, S2 is specifically divided into: S2.1 Perform correlation analysis on the 20-dimensional variable data to determine whether principal component analysis is suitable; S2.2 Standardize the variable data obtained through correlation analysis; S2.3 Calculate the correlation coefficient matrix based on the standardized data and perform principal component analysis; S2.4 Extract comprehensive indicators based on the principal component variance contribution rate to generate dimensionality-reduced input features.

4. The method for predicting CO2 emissions from coal-fired power plants using the PCA-WOA-GPR network model according to claim 2, characterized in that, S2.1 specifically includes: performing the KMO test and Bartlett's test of sphericity on the 20-dimensional variable data to determine whether the data is suitable for principal component analysis; wherein, In the KMO test, the ratio of the partial correlation coefficient to the simple correlation coefficient between any two variables in the 20-dimensional variable data is evaluated by calculating the sampling adequacy statistic. When the KMO value is greater than the preset threshold, it indicates that the overall correlation between the variables is suitable for principal component analysis. In the Bartlett test for sphericity, the correlation coefficient matrix of the 20-dimensional variable data is tested to see if there is a significant difference from the identity matrix. When the significance level of the test is less than a preset threshold, the null hypothesis that "the correlation coefficient matrix is ​​an identity matrix" is rejected, indicating that there is a significant correlation in the data. If both the KMO value and the Bartlett test P value are greater than the preset threshold, then the 20-dimensional variable data are deemed suitable for principal component analysis.

5. The method for predicting CO2 emissions from coal-fired power plants using the PCA-WOA-GPR network model according to claim 2, characterized in that, S2.2 specifically includes: standardizing the 20-dimensional variable data that were determined to be suitable for principal component analysis by the KMO and Bartlett tests in S2.1; for each variable, calculating the mean and standard deviation of all sample values; subtracting the mean of the variable from the original value of each sample and then dividing by the standard deviation to obtain the standardized data; the processed data is used as the input for calculating the correlation coefficient matrix in S2.

3.

6. The method for predicting CO2 emissions from coal-fired power plants using the PCA-WOA-GPR network model according to claim 2, characterized in that, S2.3 specifically includes: S2.3.1 Calculate the sample correlation coefficient matrix r of the standardized 20-dimensional variable data. Its expression is: Where n is the number of samples, x ti and x tj These are the standardized data of the i-th and j-th dimensions of the t-th sample, respectively; r ij Let represent the correlation coefficient between the i-th and j-th variables, where p = 20 is the total number of variables; i, j = 1, 2, ..., 20; S2.3.2 Calculate the eigenvalues ​​(λ1, λ2, ..., λ) of the correlation coefficient matrix r. p The corresponding eigenvectors are: a i =(a i1 ,a i2 ,…a ip ),i=1,2,…,p S2.3.

3. Use the eigenvalues ​​and eigenvectors as the output of principal component analysis for subsequent comprehensive index extraction.

7. The method for predicting CO2 emissions from coal-fired power plants using the PCA-WOA-GPR network model according to claim 2, characterized in that, S2.4 specifically includes: calculating the contribution rate of each principal component based on the eigenvalues ​​obtained in S2.

3. The contribution rate is the proportion of the variance of that principal component to the total variance of all principal components, and its calculation formula is as follows: The top k principal components with a cumulative contribution rate of over 85% are selected as comprehensive indicators. The standardized original data are then projected onto the selected principal component directions to generate the dimensionality-reduced comprehensive indicators.

8. The method for predicting CO2 emissions from coal-fired power plants using the PCA-WOA-GPR network model according to claim 1, characterized in that, S3 is specifically divided into: S3.1 Construction of Gaussian process regression model; Based on the k principal component comprehensive indices obtained from S2.4 as input features, a Gaussian process regression model is constructed. The model's covariance function employs a squared exponential kernel function, which includes two hyperparameters to be optimized: the variance parameter σ. 2 The output amplitude is controlled, and the scale parameter θ controls the correlation strength between features. S3.2 Parameter optimization based on whale optimization algorithm; The σ determined in S3.1 2 Using θ as the optimization objective, the whale optimization algorithm is used to optimize the parameters. The specific process is as follows: (1) Initialization phase: Set the population size to N, and each individual represents a group (σ). 2 ,θ) parameter combinations; (2) Use the root mean square error on the training set as the fitness function to evaluate the prediction performance of the current parameter combination; (3) Based on the fitness assessment results, the following strategies are used to update the parameters: When the fitness is improved, the encirclement strategy is executed to perform a fine search within the neighborhood of the current optimal parameters. When fitness stagnates, a random search strategy is executed to expand the range of parameters to be searched. Regularly execute the spiral search strategy to maintain parameter diversity; (4) When the optimal fitness value no longer improves with continuous iterations or reaches the maximum number of iterations, output the optimal (σ) value. 2 ,θ) parameter combinations; S3.3 Substitute the optimal hyperparameters obtained from S3.2 into the Gaussian process regression model and use the training set data to complete the model training; finally, input the test set data into the trained model and output the predicted value of CO2 emissions to complete the quantitative assessment of carbon emissions from coal-fired power plants.