Carbon emission prediction method, system, equipment and medium
By employing adaptive model selection and dynamic feature filtering methods, a high-precision carbon emission prediction model is constructed. This solves the problems of difficulty in unifying carbon emission data in the thermal power industry and the low accuracy of traditional models, enabling accurate prediction of carbon emissions from coal-fired power plants and supporting the green and low-carbon development of enterprises.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-10
AI Technical Summary
Existing methods for carbon emission accounting and forecasting in the thermal power industry suffer from problems such as untimely data collection, statistical lag, and inconsistent accounting boundaries and methods. This makes it difficult to effectively compare, verify, and unify data, and fails to accurately capture carbon emission characteristics with high spatiotemporal resolution. Furthermore, traditional models struggle to capture the complex nonlinear relationships of carbon emissions from coal-fired power plants, resulting in low prediction accuracy.
An adaptive model selection mechanism is adopted, combined with dynamic training cycles and multi-level feature screening methods. By selecting a suitable basic prediction model and performing hyperparameter tuning, a multi-source data system covering coal quality, combustion, equipment and environment is constructed to build a high-precision carbon emission prediction model. The model parameters are optimized through a closed-loop feedback mechanism of multi-index evaluation.
It enables accurate and reliable prediction of carbon emissions from coal-fired power plants, provides stable and reliable carbon emission prediction results, supports enterprises in precise carbon management and energy-saving and carbon-reduction decisions, and promotes green and low-carbon development.
Smart Images

Figure CN121835977A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of coal power carbon emission, and particularly relates to a carbon emission prediction method, system, device and medium. BACKGROUND
[0002] With the increasingly serious global climate change problem, reducing carbon emissions and realizing green and low-carbon development have become the consensus of the international community. As an important part of power production, the thermal power industry is a key field of energy consumption and carbon emission. Thermal power generation generally refers to a mode of using heat energy generated by burning fuels such as coal, oil and natural gas to heat water, so that the water becomes high-temperature and high-pressure steam, and then the steam drives a generator to generate electricity. Thermal power generation includes coal-fired power generation, gas-fired power generation, oil-fired power generation, waste heat / power / air power generation, garbage incineration power generation and biomass power generation. For a long time, coal power carbon emission has been one of the main carbon emission sources in the energy field. As a main component of thermal power, coal power will also be affected by these overall trends, and its carbon emission proportion may further decrease with the development of clean energy. However, the accounting and prediction methods currently used by the thermal power industry mainly rely on the IPCC inventory preparation method, the national and provincial greenhouse gas inventory preparation guidelines and the like. There are disadvantages such as untimely data collection, lagging statistical accounting information, non-uniform accounting boundaries and methods. The carbon emission sources of thermal power enterprises are complex, covering point sources, line sources and area sources, and have significant spatial differences. Moreover, the multi-energy consumption data and emission factors have great differences in terms of collection accuracy, data sources, accounting system and spatio-temporal characteristics, which makes it difficult to effectively compare, verify and unify the data. At the same time, the existing data sources are complicated, and the spatio-temporal scale is limited, which cannot accurately capture the carbon emission characteristics of high spatio-temporal resolution. It is difficult to accurately account for the carbon emission of the thermal power industry. SUMMARY
[0003] To solve the above technical problems, the present application provides the following technical solutions. In a first aspect, the present application provides a carbon emission prediction method, comprising: based on coal power operation data, adaptively selecting a basic prediction model according to the data characteristics thereof, and performing hyperparameter tuning on the basic prediction model. Performing dynamic cycle selection and multi-level feature screening on the training data set to determine target training samples and a target feature subset that adapt to the current data characteristics. Training the tuned prediction model based on the target training samples and the feature subset to obtain a target prediction model. Using the target prediction model to predict carbon emission and outputting a prediction result.
[0004] As a preferred scheme of the carbon emission prediction method of the present application, wherein adaptively selecting a basic prediction model according to the data characteristics thereof comprises: selecting a basic prediction model from a preset model set according to linear and nonlinear relationships between data characteristics; when it is judged that the relationship between data characteristics conforms to a linear hypothesis, a linear regression model is selected; when it is judged that the data has multicollinearity, an elastic regression model is selected; when it is judged that the data is a high-dimensional nonlinear data set, an XGBoost model is selected.
[0005] As a preferred scheme of the carbon emission prediction method, wherein: the hyperparameter tuning of the XGBoost model adopts a grid search combined with hierarchical Bayesian optimization; The hierarchical Bayesian optimization divides the hyperparameter search into a global layer and a local layer; In the global layer, the number of trees parameter is searched in an interval; In the local layer, the depth and learning rate parameters of the trees are finely searched within the number of trees interval determined in the global layer.
[0006] As a preferred scheme of the carbon emission prediction method, wherein: dynamic period selection is performed on the training data set, including, A rolling window division method is adopted to take different period data sets forward from the current time as the end point; The intercepted data sets are divided into a training set and a test set at a fixed ratio; The mean absolute percentage error of the test set corresponding to each period is calculated; The period that makes the MAPE minimum is selected as the current dynamic training set period.
[0007] As a preferred scheme of the carbon emission prediction method, wherein: the multi-level feature screening adopts a framework of fusion filtering, wrapping and embedding methods; The execution flow of the framework includes, Preliminary rough screening is performed by adopting Pearson correlation analysis to eliminate features irrelevant to the carbon emission target variable; The feature set after rough screening is simultaneously or sequentially subjected to fine screening by adopting a recursive feature elimination method and a Lasso regression method; The prediction performance of different feature subsets is evaluated by cross-validation to determine the target feature subset.
[0008] As a preferred scheme of the carbon emission prediction method, wherein: the method further includes comparing the prediction result of the target prediction model with the actual carbon emission data to optimize the model parameters.
[0009] As a preferred scheme of the carbon emission prediction method of the present application, wherein: the period includes a short period of 7 days or 14 days, a medium period of 1 to 3 months, and a long period of 6 months or more.
[0010] In a second aspect, the present application provides a carbon emission prediction system, comprising: a selection module for selecting a basic prediction model based on coal power operation data according to the data characteristics thereof, and performing hyperparameter tuning on the basic prediction model; A determination module for performing dynamic period selection and multi-level feature screening on the training data set to determine target training samples and a target feature subset that adapt to the current data characteristics; A training module for training the tuned prediction model based on the target training samples and the feature subset to obtain a target prediction model; A prediction module for predicting carbon emissions using the target prediction model and outputting the prediction results.
[0011] In a third aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the above method when executing the computer program.
[0012] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the above method.
[0013] Compared with the prior art, the present application has the following advantages: by constructing a multi-source data system covering coal quality, combustion, equipment and environment, and using an adaptive model selection mechanism and a hierarchical optimization strategy, combined with a dynamic training period and a multi-level feature screening method, a high-precision prediction model capable of accurately capturing the nonlinear characteristics of coal power carbon emissions is constructed. At the same time, by introducing a closed-loop feedback mechanism based on multi-index evaluation, the continuous optimization of model parameters and the self-improvement of prediction performance are realized, thereby providing stable and reliable carbon emission prediction results for coal power enterprises, effectively supporting production optimization and emission reduction decisions, and promoting green and low-carbon development. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0015] Figure 1 The flowchart of the carbon emission prediction method. DETAILED DESCRIPTION
[0016] In order to make the above objectives, characteristics and advantages of the present application more apparent, more comprehensible, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should belong to the protection scope of the present application.
[0017] Embodiment 1, refer to Figure 1 For the first embodiment of the present application, the embodiment provides a carbon emission prediction method, comprising: S100: based on coal power operation data, a basic prediction model is adaptively selected according to the data characteristics it presents, and the hyperparameters of the basic prediction model are optimized; S200: dynamic cycle selection and multi-level feature screening are performed on the training data set to determine the target training sample and the target feature subset that adapt to the current data characteristics; S300: based on the target training sample and the feature subset, the optimized prediction model is trained to obtain a target prediction model; S400: the target prediction model is used to predict carbon emissions, and the prediction result is output.
[0018] It should be noted that the carbon emission is affected by the coupling of coal quality composition, combustion process and other factors, and the traditional linear model is difficult to capture the complex nonlinear relationship, resulting in low prediction accuracy. At the same time, the data also has problems such as lagging collection, different sources, and uneven quality, and the dynamic trend of carbon emission has both long-term regularity and short-term fluctuation. The model trained with fixed training cycle and feature set cannot adapt to the dynamic change characteristics of the actual operation data of the enterprise.
[0019] Therefore, in view of the above problems, the method through the steps of S100-S400 first intelligently selects and optimizes the most suitable prediction model according to the inherent characteristics of the data, effectively improving the fitting ability of the model to complex nonlinear relationships. Then, through dynamic cycle selection and multi-level feature screening, the optimal training data and core feature set that best reflect the current running state are automatically determined. Finally, a high-precision target prediction model is constructed, realizing accurate and reliable prediction of the future carbon emissions of coal power enterprises, and providing solid data support for precise carbon management and energy saving and carbon reduction decision-making of enterprises.
[0020] Embodiment 2, refer to Figure 1 For an embodiment of the present application, based on the above embodiment, a carbon emission prediction method is provided.
[0021] In the embodiments of the present application, in step S100, based on the coal power operation data, a basic prediction model is adaptively selected according to the data characteristics presented by the coal power operation data, and hyperparameter tuning is performed on the basic prediction model, including the following steps A1-A2: It can be understood that the coal power operation data includes coal quality composition data, combustion process parameter data, combustion equipment state data and environmental condition data, wherein the coal quality composition data includes at least one of received base element carbon content, air-dry base element carbon content, dry base element carbon content, received base low calorific value and unit calorific value carbon content; the combustion process parameter data includes at least one of coal consumption, unit power generation carbon emission and instantaneous carbon emission intensity; the combustion equipment state data includes at least one of unit type, unit use time, boiler combustion efficiency and equipment failure rate; and the environmental condition data includes at least one of environmental temperature, precipitation, wind speed and power grid emission factor.
[0022] It should be noted that before the basic prediction model is adaptively selected based on the coal power operation data according to the data characteristics presented by the coal power operation data, the above-mentioned coal power operation data is pre-processed, wherein the pre-processing includes processing missing values in the data by interpolation or mean substitution; processing abnormal values in the data by identifying and eliminating or correcting, and normalizing the data to ensure the consistency of the data format.
[0023] In an optional embodiment, the coal power operation data can also be obtained by a distributed sensing network based on Internet of Things and edge computing, that is, a plurality of types of micro intelligent sensors are deployed at key process nodes such as coal conveying belts, coal mills, boiler furnaces and flues to form an Internet of Things sensing layer; the high-frequency vibration, infrared thermal imaging, sound spectrum and micro-environmental air pressure data collected by these sensors are subjected to preliminary filtering, compression and feature extraction by edge computing nodes deployed near the equipment, for example, by analyzing vibration spectrum characteristics to indirectly calculate coal particle size and flow rate, or by real-time inversion of the in-furnace combustion temperature field distribution through infrared thermal imaging data; the feature data pre-processed on the edge side is uploaded to the central data platform through industrial Ethernet or 5G network, and is fused with the conventional parameters recorded by the DCS system, thereby constructing an enhanced coal power operation data set with higher spatiotemporal resolution and better reflecting the internal microstate of the equipment.
[0024] In another alternative embodiment, the coal power operation data can also be obtained by tracing the mapping of supply chain information and in-plant digital twin, that is, first, a digital supply chain tracing system from coal suppliers to power plants is established to obtain the accurate mine source, washing and sorting records and detailed coal quality test sheets issued by third parties for each batch of coal entering the plant; then, a dynamic digital twin model covering the whole process of "coal transportation-pulverization-combustion-flue gas" is built on the power plant side, the traced coal quality information of the coal entering the furnace is input as the initial boundary condition into the model, and the theoretical combustion characteristics, pollutant generation potential and carbon emission benchmark value of the current coal quality under the specific equipment parameters and operating conditions are deduced through real-time simulation calculation; these deduced data and actual operation monitoring data are combined to serve as inputs for the prediction model, thereby realizing the whole-process data correlation and mapping from the fuel source characteristics to the end emission.
[0025] A1: selecting a basic prediction model from a preset model set according to the linear and nonlinear relationship between data characteristics; wherein when it is judged that the relationship between data characteristics conforms to the linear assumption, a linear regression model is selected; When it is judged that the data has multicollinearity, an elastic regression model is selected, wherein the elastic regression model is an extension of the linear regression; It can be understood that many parameters in coal power data (such as carbon content of different bases and various efficiency indicators) can be highly correlated, and the elastic regression model can handle such problems to prevent unstable model coefficients.
[0026] When it is judged that the data is a high-dimensional nonlinear data set, an XGBoost model is selected.
[0027] It can be understood that such a data set as coal power data involves multi-dimensional complex interactions of combustion chemistry, equipment physics and environmental variables, and nonlinearity is its main feature. Therefore, in actual application, the system will most likely go to this branch.
[0028] Preferably, through the selection mechanism of this step, the system can adapt to different power plants, different units and even different time periods of data state, which can ensure the rationality of the use of the model basis.
[0029] A2: the hyperparameter tuning of the XGBoost model adopts the way of grid search combined with hierarchical Bayesian optimization; Hierarchical Bayesian optimization (i.e. BO method based on Gaussian process) divides the hyperparameter search into global layer and local layer; In the global layer, the number of trees parameter is searched in an interval; It can be understood that the number of trees parameter is searched in an interval in order to quickly locate a better number interval (for example, it is found that the model performs well when the number of trees is between 100-300), rather than determining a single optimal point.
[0030] In the local layer, the depth of the tree and the learning rate parameter are finely searched within the number interval of the tree determined in the global layer.
[0031] It should be noted that after a certain candidate value of the number of trees is determined in the global layer, the fine search of the two closely related parameters of the depth of the tree and the learning rate is immediately performed in the local layer. This process is dynamic and adaptive, and the Bayesian optimizer intelligently suggests the next combination of parameters to be tried according to the previous evaluation results.
[0032] Preferably, the traditional grid search is performed in the full space of all parameters, which has a very high calculation cost. However, the present method greatly reduces unnecessary calculation waste by decomposing the search space, and can find a better combination of parameters with fewer attempts. Further, by finding a set of coordinated optimal hyperparameters, the final XGBoost model can neither underfit nor overfit, thereby having stronger generalization ability and stability.
[0033] In an optional embodiment, the adaptive selection of the base prediction model based on the coal power operation data according to the data characteristics presented by the data in step S100 can also be achieved by a dynamic selection method based on time series feature enhancement and model performance pre-evaluation, that is, first, the coal power operation data is subjected to deep time series feature engineering, and multi-dimensional time series features including volatility, trend intensity, and periodicity significance are extracted; then the time series features and static statistical features are input into a lightweight meta-classifier, which has been trained based on historical project data and can quickly predict the potential performance ranking of different base prediction models (such as LightGBM, time series convolution network, etc.) on the current data set according to the overall feature profile of the input data; finally, the system selects the model with the highest pre-evaluation performance as the base prediction model, realizing intelligent model selection under data driving.
[0034] In another optional embodiment, the adaptive selection of the base prediction model based on the coal power operation data according to the data characteristics presented by the data in step S100 can also be achieved by constructing a heterogeneous model pool and dynamically switching based on online learning performance, that is, the system initializes a heterogeneous model pool containing tree models, neural networks, and models optimized for time series data (such as Prophet); in the early stage of model training, all models in the pool are allowed to perform preliminary training in parallel, and their core performance indicators (such as MAPE) are continuously monitored on a fixed validation window; after a preset evaluation period, the system automatically locks and switches to the model with the best performance on the current validation set as the base prediction model for subsequent training and optimization, thereby realizing the best matching of model selection and real-time data characteristics.
[0035] In the embodiments of the present application, the dynamic period selection and multi-level feature screening are performed on the training data set in step S200 to determine the target training sample and target feature subset that adapt to the current data characteristics, including the following steps B1-B2: B1: performing dynamic period selection on the training data set, including, Using the rolling window division method, the data set of different periods is intercepted forward from the current time as the end point; The intercepted data set is divided into a training set and a test set according to a fixed ratio, for example, a fixed ratio of 10% length; The mean absolute percentage error of the test set corresponding to each period is calculated; The period that minimizes the MAPE is selected as the current dynamic training set period.
[0036] It can be understood that the carbon emissions of the coal power system are simultaneously affected by short-term operating condition fluctuations and long-term seasonal and maintenance rules, and the use of a fixed training period cannot capture both modes, so the dynamic selection used in this step is to determine how long the historical data should be used to train the model to most effectively predict the future.
[0037] It should be noted that the period includes short periods of 7 days or 14 days, medium periods of 1 to 3 months, and long periods of 6 months or more, wherein the short period can quickly reflect recent carbon emission changes and is suitable for capturing short-term fluctuations, but lacks the ability to identify long-term trends; the medium period can balance data timeliness and quantity and capture seasonal changes; the long period helps to identify long-term trends and periodic rules and improve the stability of the model, but is not sensitive to recent changes. It should be emphasized that the test set index of each period is recalculated every 3 days and the training set period is updated to ensure that the model adapts to the latest data characteristics.
[0038] Preferably, through this step, not only can the prediction performance decline caused by the fixed use of long periods (insensitive to recent changes) or short periods (unable to learn long-term rules) be avoided, but also the model can take into account both long-term trends and short-term fluctuations, significantly improving the universality and robustness of the model.
[0039] B2: multi-level feature screening, using a framework that integrates filtering, wrapping and embedding methods; The execution process of the framework includes, Preliminary rough screening is performed using Pearson correlation analysis to eliminate features that are not associated with the carbon emission target variable; The feature set after rough screening is simultaneously or sequentially subjected to recursive feature elimination method and Lasso regression method for fine screening; The prediction performance of different feature subsets is evaluated through cross-validation to determine the target feature subset.
[0040] It can be understood that the coal electricity data has high feature dimension, and there are a large number of redundant and irrelevant features. Through screening, the model complexity can be reduced, overfitting can be prevented, the training speed can be improved, and the stability and interpretability of the model can be enhanced.
[0041] It should be noted that this step first filters out the features that are not related to carbon emissions by calculating the linear correlation between each feature and the carbon emission target variable. From the remaining features, a core feature set with the highest prediction power is further selected. Specifically, the recursive feature elimination method is used to evaluate the feature importance using the prediction model (such as XGBoost) itself, iteratively remove the least important features, and further use the Lasso regression method to automatically complete feature selection by introducing an L1 regularization penalty term in the loss function during model training, and compress the coefficients of unimportant features to zero. Finally, the performance of the candidate feature subsets obtained after fine screening is evaluated using cross-validation under the actual prediction model, and the feature set with the best performance in cross-validation is selected as the target feature subset.
[0042] Preferably, the target feature subset obtained through this step eliminates redundancy and noise, effectively extracts potential feature interaction patterns and time series dependencies, and further improves the stability and accuracy of the prediction model.
[0043] In the embodiments of the present application, the target prediction model is obtained by training the optimized prediction model based on the target training samples and the feature subset in step S300, including steps C1-C3: C1: Limit the target training samples determined by the dynamic cycle selection to the feature dimension corresponding to the target feature subset obtained by multi-level feature screening.
[0044] It can be understood that this step is to ensure that the data input into the trainer is pure and consistent.
[0045] C2: Train the prediction model optimized by the hyperparameters using the data of the target training samples on the target feature subset.
[0046] C3: Through the training process, the model learns the final mapping relationship between the features in the target feature subset and the carbon emissions, and generates a target prediction model for actual prediction.
[0047] It should be noted that the training process is essentially a mathematical optimization process, that is, by adjusting the tens of thousands of parameters inside the model (such as the structure of all decision trees and leaf weights in XGBoost), the error between the output (predicted value) of the model and the true value is minimized. When the training converges, the internal parameters of the model are fixed, which collectively encode the mapping relationship between the features and the target variable. It can be understood that the mapping relationship can enable the model to accurately calculate the carbon emissions when facing new, never-before-seen data (but belonging to the same feature subset).
[0048] In an alternative embodiment, the target prediction model obtained in step S300 can also be obtained by integrating learning-based multi-model weighted fusion, that is, multiple base models such as XGBoost, LightGBM and random forest are trained in parallel on the target training sample and the feature subset, and then the average absolute percentage error (MAPE) of each model on the validation set is used to dynamically allocate weights, and finally the prediction results of these base models are fused by weighted voting to form a stronger prediction model with more stable and stronger generalization ability as the final target prediction model.
[0049] In another alternative embodiment, the target prediction model obtained in step S300 can also be obtained by introducing a model fine-tuning method of transfer learning, that is, a pre-trained base model architecture on a large general industrial time series dataset is first loaded, and then the pre-trained model is fine-tuned using the target training sample and the feature subset of the coal-fired power plant, only updating the last few fully connected layer parameters of the model, so that it quickly adapts to the data distribution of the current specific enterprise while retaining the ability to extract general time series features, thereby efficiently constructing a high-precision target prediction model.
[0050] In the embodiments of the present application, the target prediction model is used to predict carbon emissions in step S400, and the prediction result is output, including the following steps D1-D3: D1: The coal-fired power plant operation data corresponding to the to-be-predicted period is input into the target prediction model after being subjected to data preprocessing and feature selection operations consistent with the training phase.
[0051] Preferably, this step eliminates errors introduced by inconsistent data formats, scales or feature sets, ensuring that the model can reason in its familiar data space and output reliable prediction results.
[0052] D2: Obtain the carbon emission prediction value of the future specified period output by the target prediction model.
[0053] It should be noted that the data standardized by step D1 is sent to the target prediction model. The model calculates according to the complex mapping relationship learned by S300, and outputs one or more future carbon emission prediction values.
[0054] Preferably, the specific prediction value output by this step can enable coal-fired power enterprises to proactively understand the carbon emission trend rather than post hoc statistics.
[0055] D3: Compare the prediction results of the target prediction model with the actual carbon emission data to optimize the model parameters.
[0056] It should be noted that when the predicted future period becomes the present and actual carbon emission data is generated, the system compares the prediction value of step D2 with the actual measured value, and uses mean square error MSE (used to measure the average of the square of the difference between the predicted value and the actual value), root mean square error RMSE (used to represent the standard deviation of the error between the predicted value and the actual value, with the same unit as the original data), mean absolute error MAE (used to measure the average absolute value of the difference between the predicted value and the actual value), coefficient of determination R 2 (used to measure the goodness of fit of the model, indicating the proportion of the variance explained by the model to the total variance), mean absolute percentage error MAPE (used to measure the average percentage of the difference between the predicted value and the actual value) to quantify the performance of this prediction, and according to the evaluation results, trigger the optimization mechanism. Optimization mechanism includes re-execution or fine-tuning of hyperparameters, and re-evaluation and adjustment of training set period or feature subset.
[0057] The errors of each model are compared as shown in Table 1 below: Table 1 Error comparison table of each model It can be understood that RMSE, MAE and MAPE are error indicators, which measure the difference between the predicted value and the true value. The smaller the value, the closer the predicted value is to the true value, and the better the model performance. R 2 is used to measure how much proportion of the variance of the dependent variable is explained by the model, and its value range is [0, 1], so the larger the value, the better, R 2 closer to 1, the better the model fitting effect, the stronger the model's ability to explain data, and the better the performance, R 2 closer to 0, indicating that the model has little explanation ability.
[0058] As can be seen from Table 1 above, the above error comparison results show that the coal-fired power data is more complex, and the linear model performs worse, while the XGBoost model is better than the linear model.
[0059] In an alternative embodiment, the prediction of carbon emissions in step S400 can also be achieved by a dynamic rolling prediction method based on real-time data assimilation, that is, after obtaining the initial prediction result, the system continuously collects the latest actual operation data and inputs it into the assimilation algorithm based on Kalman filtering or particle filtering to dynamically correct the internal state or bias of the target prediction model, thereby achieving rolling update and accurate tracking of the carbon emission trajectory in the next few hours, significantly improving the timeliness and accuracy of short-term prediction.
[0060] In another alternative embodiment, the prediction of carbon emissions in step S400 can also be achieved by a probabilistic interval prediction method combined with scenario generation technology, that is, by using the feature distribution under different operating conditions in historical data, a plurality of possible future operating scenarios are constructed by a conditional variational autoencoder generation model, and these scenario data are input into the target prediction model, and finally a carbon emission prediction interval with a probability density distribution is output, providing a decision basis containing uncertainty information for risk control.
[0061] In summary, the carbon emission prediction method has the following advantages: by constructing a multi-source data system covering coal quality, combustion, equipment and environment, and using an adaptive model selection mechanism and a hierarchical optimization strategy, combined with a dynamic training period and a multi-level feature selection method, a high-precision prediction model capable of accurately capturing the nonlinear characteristics of coal-fired power plant carbon emissions is constructed; at the same time, by introducing a closed-loop feedback mechanism based on multi-index evaluation, the continuous optimization of model parameters and the self-improvement of prediction performance are realized, thereby providing stable and reliable carbon emission prediction results for coal-fired power plants, effectively supporting production optimization and emission reduction decisions, and promoting green and low-carbon development.
[0062] Embodiment 3, the above is a schematic scheme of a carbon emission prediction method. It should be noted that the technical scheme of the carbon emission prediction system belongs to the same concept as the technical scheme of the carbon emission prediction method described above. The technical scheme of the carbon emission prediction system in this embodiment is not described in detail, and the description of the technical scheme of the carbon emission prediction method can be referred to.
[0063] The embodiment also provides a carbon emission prediction system, comprising: The selection module is configured to select a basic prediction model based on coal-fired power plant operation data according to the data characteristics thereof, and to perform hyperparameter tuning on the basic prediction model; The determination module is configured to perform dynamic period selection and multi-level feature selection on the training data set to determine target training samples and a target feature subset that adapt to the current data characteristics; The training module is configured to train the tuned prediction model based on the target training samples and the feature subset to obtain a target prediction model; The prediction module is configured to predict the carbon emission by using a target prediction model and output a prediction result.
[0064] The embodiment further provides an electronic device suitable for the carbon emission prediction, comprising a memory and a processor; the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to realize the carbon emission prediction method proposed in the above embodiment.
[0065] The embodiment further provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to realize the carbon emission prediction method proposed in the above embodiment.
[0066] The storage medium proposed in the embodiment and the carbon emission prediction method proposed in the above embodiment belong to the same inventive concept, and the technical details not described in the embodiment can be referred to the above embodiment, and the embodiment has the same beneficial effects as the above embodiment.
[0067] From the above description about the embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and necessary universal hardware, and of course can also be realized by hardware. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a floppy disk, a ROM, a RAM, a FLASH, a hard disk or an optical disk, and includes a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute the methods of various embodiments of the present application.
[0068] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application, and although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and all of them should be covered in the scope of the claims of the present application.
Claims
1. A method of carbon emissions prediction, characterized by: The method comprises the following steps: Adaptive selection of a basic prediction model according to the data characteristics presented by the coal power operation data, and hyperparameter tuning of the basic prediction model; Dynamic cycle selection and multi-level feature screening are performed on the training data set to determine the target training sample and the target feature subset that adapt to the current data characteristics; The tuned prediction model is trained based on the target training sample and the feature subset to obtain a target prediction model; The target prediction model is used to predict carbon emissions, and the prediction result is output.
2. The carbon emission prediction method of claim 1, wherein: The adaptive selection of the basic prediction model according to the data characteristics presented by the coal power operation data comprises the following steps: According to the linear and nonlinear relationship between data characteristics, a basic prediction model is selected from a preset model set; when it is judged that the relationship between data characteristics conforms to the linear hypothesis, a linear regression model is selected; When it is judged that the data has multicollinearity, an elastic regression model is selected; When it is judged that the data is a high-dimensional nonlinear data set, an XGBoost model is selected.
3. A carbon emission prediction method as claimed in claim 2, wherein: The hyperparameter tuning of the XGBoost model adopts a grid search combined with hierarchical Bayesian optimization; The hierarchical Bayesian optimization divides the hyperparameter search into a global layer and a local layer; In the global layer, the number of trees is searched in an interval; In the local layer, the depth and learning rate parameters of the trees are finely searched within the number interval of the trees determined in the global layer.
4. A carbon emission prediction method as claimed in claim 3, wherein: The dynamic cycle selection of the training data set comprises the following steps: A rolling window division method is used to take different period data sets forward from the current time as the end point; The taken data sets are divided into a training set and a test set at a fixed ratio; The mean absolute percentage error (MAPE) of the test set corresponding to each period is calculated; The period that makes the MAPE minimum is selected as the current dynamic training set period.
5. A carbon emission prediction method as claimed in claim 4, characterized by: The multi-level feature screening adopts a framework that fuses filtering, wrapping, and embedding methods; The execution process of the framework comprises the following steps: Preliminary coarse screening is performed by using Pearson correlation analysis to eliminate features irrelevant to the carbon emission target variable; Recursive feature elimination and Lasso regression methods are used to perform fine screening on the coarse screened feature set synchronously or sequentially; The prediction performance of different feature subsets is evaluated through cross-validation to determine the target feature subset.
6. A method of carbon emission prediction according to any one of claims 1 to 5 wherein: The method further comprises comparing the prediction result of the target prediction model with the actual carbon emission data to optimize the model parameters.
7. The carbon emission prediction method of claim 4, wherein: The period includes short periods of 7 days or 14 days, medium periods of 1 to 3 months, and long periods of 6 months or more.
8. A carbon emission prediction system applying the method according to any one of claims 1 to 7, characterized in that, The method comprises the following steps: A selection module is configured to adaptively select a basic prediction model according to the data characteristics presented by the coal power operation data, and to tune the hyperparameters of the basic prediction model; A determination module is configured to perform dynamic cycle selection and multi-level feature screening on the training data set to determine the target training sample and the target feature subset that adapt to the current data characteristics; A training module is configured to train the tuned prediction model based on the target training sample and the feature subset to obtain a target prediction model; A prediction module is configured to use the target prediction model to predict carbon emissions and output the prediction result. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.