Carbon emission prediction optimization system

By introducing quantitative evaluation and similarity analysis of uncertainty factor data sets, the carbon emission prediction model is optimized, which solves the problem of inaccurate predictions of traditional methods in complex environments and achieves more accurate and stable carbon emission predictions.

CN120688698APending Publication Date: 2025-09-23DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510939771.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing carbon emission forecasting methods are unable to accurately capture the impact of uncertain factors in complex and changing environments, resulting in insufficient stability and large errors in the forecast results.

Method used

A historical uncertainty factor dataset is introduced, and the impact intensity, change fluctuation, development trend and stability reliability index are calculated through a quantitative evaluation mechanism to generate a comprehensive score. It is then integrated with the benchmark carbon emission characteristics into a unified feature set to construct a carbon emission prediction model. The similarity analysis is combined to screen uncertainty factors and optimize the model input.

Benefits of technology

It significantly improves the accuracy of carbon emission prediction and its ability to adapt to complex environments, and improves the generalization ability and prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688698A_ABST
    Figure CN120688698A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon emission prediction optimization system, which relates to the technical field of carbon emission prediction, and realizes full-process optimization design from model training to a prediction stage by effectively combining structured modeling of a historical uncertainty factor data set and a screening mechanism of an uncertainty factor data set in to-be-predicted data. The accuracy, scientificity and practicability of a carbon emission prediction result are effectively improved, and the method has good engineering application prospects and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of carbon emission prediction, and in particular to a carbon emission prediction optimization system. Background Art

[0002] As global climate change becomes increasingly severe, carbon emission forecasting, as a crucial basis for formulating emission reduction policies and promoting low-carbon development, has become a core topic of common concern in academia and industry. Currently, mainstream carbon emission forecasting methods primarily rely on historical baseline carbon emission data. These models construct forecasting models by extracting conventional variables such as time series trends, seasonal characteristics, and changes in energy structure. These models can, to a certain extent, reflect the basic evolution of carbon emissions. However, in the face of complex and ever-changing real-world environments, they often struggle to accurately capture the impact of fluctuations caused by uncertainties. For example, economic policy adjustments, sudden natural disasters, and changes in industrial structure can all significantly impact carbon emissions. However, due to their random and nonlinear nature, traditional methods typically fail to incorporate these factors into model training and lack systematic quantitative modeling. This approach not only limits the model's ability to learn from uncertainty information but also results in insufficient stability and large errors in forecast results when faced with new scenarios. Therefore, existing technologies urgently need technical solutions for carbon emission forecasting optimization systems. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides a carbon emission prediction and optimization system, which specifically includes the following modules:

[0004] Historical data collection module: used to collect historical carbon emission data and pre-process the historical carbon emission data, wherein the historical carbon emission data includes a historical baseline carbon emission data set and a historical uncertainty factor data set;

[0005] Feature set acquisition module: connected to the historical data collection module, used to extract features from the pre-processed historical baseline carbon emission dataset and historical uncertainty factor dataset, and integrate the extracted features into a feature set;

[0006] Historical data quantification unit: used to perform quantification processing on the historical uncertainty factor data set and obtain a comprehensive score for each uncertainty factor data item in the historical uncertainty factor data set based on the quantification result;

[0007] Index calculation interface: used to calculate the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertainty factor data item in the historical uncertainty factor data set;

[0008] Weight analysis interface: used to calculate the weight of each index based on the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertain factor data item;

[0009] Standardization processing sub-interface: used to perform standardization on each index of each uncertainty factor data item in the historical uncertainty factor dataset;

[0010] Global mean and standard deviation calculation sub-interface: used to calculate the global mean and standard deviation of each index of each uncertainty factor data item in the entire historical uncertainty factor data set based on the standardization processing results;

[0011] Deviation calculation sub-interface: used to calculate the deviation of each index relative to the global mean for each uncertainty factor data item;

[0012] Weight calculation sub-interface: used to normalize the deviation of each index of each uncertainty factor data item relative to the global mean to obtain the weight of each index of each uncertainty factor data item;

[0013] Comprehensive score calculation interface: used to calculate the comprehensive score of each uncertain factor data item based on each index and the weight of each index of each uncertain factor data item using a weighted summation method;

[0014] Feature set integration unit: used to integrate the comprehensive score of each uncertainty factor data item in the historical uncertainty factor dataset and the features extracted from the historical benchmark carbon emission dataset into a feature set;

[0015] Carbon emission prediction model training module: connected to the feature set acquisition module, used to build a carbon emission prediction model, and integrate historical carbon emission data with the feature set into a training set, and use the training set to train the carbon emission prediction model to obtain a trained carbon emission prediction model;

[0016] The data analysis module to be predicted is connected to the carbon emission prediction model training module to obtain the carbon emission data to be predicted, analyze the carbon emission data to be predicted, and obtain the final carbon emission data to be predicted;

[0017] Data quantification unit to be predicted: used to calculate the comprehensive score of each uncertainty factor data item in the uncertainty factor data set of the carbon emission data to be predicted;

[0018] Similarity value analysis unit: used to analyze the similarity between the comprehensive scores of each uncertainty factor data item in the uncertainty factor data set, and obtain the similarity value between the comprehensive scores of each uncertainty factor data item;

[0019] One-dimensional feature vector construction interface: used to construct a one-dimensional feature vector based on the comprehensive scores of all uncertain factor data items, where each element in the one-dimensional feature vector represents the comprehensive score of an uncertain factor data item;

[0020] Similarity value calculation interface: used to calculate the numerical proximity of the comprehensive score between each uncertain factor data item using a pairwise comparison method, and obtain the similarity value of the comprehensive score between each uncertain factor data item;

[0021] A screening threshold setting unit is used to calculate a first average similarity value between the comprehensive scores of all uncertainty factor data items, and set a screening threshold value according to the first average similarity value;

[0022] Uncertainty factor data item retaining unit: used to screen each uncertainty factor data item according to the screening threshold value, and determine the retained uncertainty factor data item based on the screening result;

[0023] Second average similarity value calculation interface: used to calculate the second average similarity value between the current uncertain factor data item and all other uncertain factor data items based on the similarity value between each uncertain factor data item;

[0024] Screening interface: used to compare the second average similarity value of each uncertainty factor data item with the screening threshold;

[0025] Screening result acquisition interface: used to determine whether to retain the current uncertain factor data item if the second average similarity value of the current uncertain factor data item is less than the screening threshold; if the second average similarity value of the current uncertain factor data item is greater than or equal to the screening threshold, then eliminate the current uncertain factor data item;

[0026] Final uncertainty factor data set acquisition unit: used for aggregating each uncertainty factor data item determined to be retained to generate a final uncertainty factor data set;

[0027] Final data acquisition unit to be predicted: used to integrate the final uncertainty factor data set with the benchmark carbon emission data set to obtain the final carbon emission data to be predicted;

[0028] Prediction result acquisition module: connected to the data analysis module to be predicted, used to input the final carbon emission data to be predicted into the trained carbon emission prediction model and output the carbon emission prediction result.

[0029] The embodiments of the present invention have the following technical effects:

[0030] The present invention provides a carbon emission prediction method, the core innovation of which is that it introduces historical uncertainty factor data sets into the training process of carbon emission prediction models for the first time, and by constructing a complete set of quantitative evaluation mechanisms, it converts the uncertainty factor information that was originally difficult to structure into numerical features that can be recognized and learned by the model. In traditional carbon emission prediction models, the features used for training are often limited to conventional features such as trends and periodicity extracted from historical benchmark carbon emission data. There is a lack of effective modeling of external factors that affect carbon emissions but have uncertainty, resulting in insufficient generalization ability and limited prediction accuracy of the model when facing complex real-life scenarios. To this end, the present invention proposes to evaluate each uncertainty factor data item in the historical uncertainty factor data set based on four key indicators - including impact intensity index, change volatility index, development trend index and stability reliability index, and dynamically calculate weights based on the deviation of each indicator relative to the global distribution, and finally generate a comprehensive score for the data item through weighted fusion. The comprehensive score serves as a quantitative expression of the uncertainty factor. , which can accurately reflect the degree of its impact on carbon emissions in numerical form, and be integrated with the benchmark carbon emission features into a unified feature set for training the carbon emission prediction model. This mechanism not only significantly improves the dimension and information integrity of the training features, but also enables the model to have the ability to adapt to the uncertainty environment; in addition, in the prediction stage, the present invention further introduces a screening mechanism based on similarity analysis, by performing a comprehensive score calculation on the uncertainty factor data set in the current data to be predicted, and constructing a similarity matrix, setting a screening threshold and eliminating atypical factors that deviate greatly from the overall trend, retaining the most representative uncertainty factor data items, thereby improving the consistency and quality of the input data; in summary, the present invention effectively combines the structured modeling of the historical uncertainty factor data set with the screening mechanism of the uncertainty factor data set in the data to be predicted, thereby realizing the full-process optimization design from model training to prediction stage, effectively improving the accuracy, scientificity and practicality of the carbon emission prediction results, and has good engineering application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 It is a framework diagram of the carbon emission prediction and optimization system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0033] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0034] Example 1: Figure 1 As shown, the carbon emission prediction and optimization system provided by the present invention includes the following modules:

[0035] Historical data collection module: used to collect historical carbon emission data and pre-process the historical carbon emission data, wherein the historical carbon emission data includes a historical baseline carbon emission data set and a historical uncertainty factor data set;

[0036] It is worth noting that the historical benchmark carbon emissions dataset mainly includes data reflecting the long-term trend, cyclical laws and structural characteristics of carbon emissions, specifically including the following categories: 1. Time series carbon emissions, that is, by recording the total carbon emissions in each time period such as year, quarter, and month, usually obtained from the annual report released by the environmental department; 2. Energy consumption structure data, including the proportion of various energy sources such as coal, oil, natural gas, and renewable energy and their corresponding carbon emission coefficients. The data source is obtained through the Energy Bureau; 3. Industrial structure data, reflecting the contribution of different industries such as industry, agriculture, and services to carbon emissions, specifically obtained through economic census reports, industry research literature, and information released by local development and reform commissions; 4. Seasonal factor data, used to capture the trend of carbon emissions changing with seasons For example, the increased carbon emissions caused by peak electricity consumption during the winter heating season are specifically obtained through historical temperature and energy consumption data from meteorological departments. To ensure the effectiveness and stability of subsequent model training, this data must undergo a series of standardization and cleaning operations before being used for modeling. Specifically, missing value processing is required. Missing items in the dataset are filled using methods such as linear interpolation and mean imputation to avoid affecting the subsequent modeling process. Next, outlier detection and correction are performed. Statistical methods such as box plots and Z-scores are used to identify abnormal data points. Obviously incorrect or unreasonable values, such as negative carbon emissions, are corrected or eliminated based on contextual information. Next, data standardization is required. Because data of different dimensions may have orders of magnitude differences, the dimensions must be unified. Z-score standardization is used for continuous variables to ensure that the data follows a distribution with a mean of 0 and a standard deviation of 1. One-hot encoding is performed for categorical variables.

[0037] Feature set acquisition module: connected to the historical data collection module, used to extract features from the pre-processed historical baseline carbon emission dataset and historical uncertainty factor dataset, and integrate the extracted features into a feature set;

[0038] It is worth noting that after completing the preprocessing of the historical benchmark carbon emission dataset, the historical benchmark carbon emission dataset has a well-structured and standardized form. On this basis, the following key features can be extracted from it, including time trend characteristics, such as the average carbon emissions at different granularities such as annual, quarterly, monthly, and weekly; periodic characteristics, such as seasonal fluctuation components, specifically the periodic terms extracted through Fourier transform; energy structure-related characteristics, such as the carbon emission coefficients of different energy types and their average values; the above characteristics together constitute the basic feature set in the training of the carbon emission prediction model, reflecting the basic laws and structural characteristics of carbon emissions.

[0039] Historical data quantification unit: used to perform quantification processing on the historical uncertainty factor data set and obtain a comprehensive score for each uncertainty factor data item in the historical uncertainty factor data set based on the quantification result;

[0040] Index calculation interface: used to calculate the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertainty factor data item in the historical uncertainty factor data set;

[0041] It is worth noting that the calculation process of the impact intensity index of each uncertain factor data item is as follows: First, the actual carbon emissions C when each uncertain factor data item is considered should be obtained. t , the actual carbon emissions C t The data is obtained through environmental monitoring systems, government statistical platforms, or corporate carbon accounts. The time granularity is usually monthly or quarterly, such as the total monthly carbon emissions of a region during the implementation of a policy. In addition, the baseline carbon emissions without considering the uncertainty factor need to be obtained. The method of obtaining the carbon emissions is to take the average carbon emissions in the same period of several years, such as 5 years, as a reference, and then, based on C t and Further calculation to obtain the standardized impact value in It is expressed as the change in carbon emissions caused by the uncertain factor data item in the tth time period; then, based on the above data, the impact intensity index of each uncertain factor data item is calculated. The specific calculation formula is: Among them, I irepresents the impact intensity index of the i-th uncertainty factor data item; T represents the total number of time periods in which the i-th uncertainty factor data item acts; Represents the change in carbon emissions caused by the i-th uncertainty factor data item in the k-th time period; represents the baseline carbon emissions in the kth time period without considering the i-th uncertainty factor data item;

[0042] The purpose of this calculation is to measure the overall disturbance intensity of a certain uncertainty factor data item on carbon emissions during its action cycle; and by introducing the "time accumulation effect" into the above calculation formula, it is possible to effectively capture those uncertainty factor data items that have a small initial impact but gradually increase in intensity; the practical significance of this calculation is that, for example, if a policy has a limited initial impact, that is, a certain uncertainty factor data item, but as its implementation deepens, its disturbance gradually increases, this can be accurately portrayed through the above calculation process, thereby improving the subsequent model's ability to predict future trends. Specifically, the above calculation process helps to identify uncertainty factor data items with deep disturbance mechanisms on carbon emissions, thereby improving the subsequent model's ability to perceive complex external disturbance factors;

[0043] The calculation process of the change fluctuation index of each uncertainty factor data item is as follows: First, based on the obtained standardized impact value The impact value reflects the degree of influence of the current uncertainty factor data item on carbon emissions in different time periods, and then calculates its average value in all time periods. Then calculate the change fluctuation index of each uncertain factor data item. The specific calculation formula is:

[0044]

[0045] Among them, V i Represents the change fluctuation index of the i-th uncertainty factor data item;

[0046] The purpose of this calculation is to reflect the stability of the impact of current uncertainties on carbon emissions. A larger fluctuation index indicates greater differences in the impact of the uncertainties over time, leading to higher prediction difficulty. This helps identify "atypical" uncertainty factor data items. Its practical significance lies in helping the model filter out factors that fluctuate violently and are difficult to model, thereby improving the consistency of subsequent model input data and enhancing the stability of prediction results.

[0047] The calculation process of the development trend index of each uncertain factor data item is as follows: first calculate the difference between the standardized impact values ​​of adjacent time periods, that is, It is used to determine whether the impact of each uncertainty factor data item on carbon emissions shows an upward or downward trend; then, based on the difference in standardized impact values, the development trend index of each uncertainty factor data item is calculated. The specific calculation formula is:

[0048]

[0049] Among them, D i Represents the development trend index of the i-th uncertainty factor data item;

[0050] The purpose of this calculation is to assess whether the impact of current uncertainty factor data items on carbon emissions is showing an increasing or decreasing trend, thereby identifying uncertainty factor data items with persistent or phased characteristics, which will help subsequent models identify long-term influencing factors, thereby improving the subsequent models' ability to predict future scenarios, that is, enhancing the subsequent models' ability to adapt to complex uncertainty factors;

[0051] The calculation process of the stability reliability index of each uncertainty factor data item is as follows: first, obtain the standardized influence value in all time periods, and based on the obtained standardized influence values ​​in all time periods, call the maximum value of the standardized influence value and minimum value Then we need to call the average of the standardized impact values ​​in all time periods Then the stability reliability index of each uncertainty factor data item is calculated. The specific calculation formula is:

[0052]

[0053] Among them, S i Represents the stability reliability index of the i-th uncertainty factor data item;

[0054] The purpose of this calculation is to measure the consistency and predictability of the current uncertainty factor over multiple sample periods; and by dividing the range by the mean to obtain the relative fluctuation ratio, the index can be applied to uncertainty factor data items of different dimensions. At the same time, the higher the value, the more stable the factor behavior. The practical significance of this calculation is to screen out uncertainty factors with strong behavioral regularity and small interference, thereby improving model training efficiency and prediction accuracy.

[0055] Weight analysis interface: used to calculate the weight of each index based on the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertain factor data item;

[0056] Standardization processing sub-interface: used to perform standardization on each index of each uncertainty factor data item in the historical uncertainty factor dataset;

[0057] Global mean and standard deviation calculation sub-interface: used to calculate the global mean and standard deviation of each index of each uncertainty factor data item in the entire historical uncertainty factor data set based on the standardization processing results;

[0058] Deviation calculation sub-interface: used to calculate the deviation of each index relative to the global mean for each uncertainty factor data item;

[0059] It is worth noting that for each uncertainty factor data item, the deviation of each index relative to the global mean is calculated, which is mainly used to measure whether the current uncertainty factor data item has significant characteristics in each index. The specific calculation formula is as follows: Where d i Represents the deviation of the i-th index relative to the global mean; Represents the i-th standardized exponential value of the current uncertainty factor data item; μ i Represents the global mean of the i-th index of the current uncertainty factor data item in the entire historical uncertainty factor data set; σ i Represents the global standard deviation of the i-th index of the current uncertainty factor data item in the entire historical uncertainty factor data set;

[0060] Among them, the larger the deviation, the more "special" the performance of the current uncertainty factor data item on the index is, and the more important it is to carbon emission prediction.

[0061] Weight calculation sub-interface: used to normalize the deviation of each index of each uncertainty factor data item relative to the global mean to obtain the weight of each index of each uncertainty factor data item;

[0062] It is worth noting that the calculation formula for the weight of each index of each uncertainty factor data item is:

[0063]

[0064] Where w i Represents the weight of the i-th index of the current uncertainty factor data item;

[0065] This weight allocation mechanism is entirely determined by the index performance of the uncertainty factor data item itself, thus achieving an objective weighting mechanism that is closely related to the characteristics of the uncertainty factor data item itself, which can significantly improve the calculation accuracy and scientificity of the final comprehensive score;

[0066] It is worth further explaining that by comparing the degree of deviation between each uncertain factor data item in a certain index and the average value of its category and using it as the source of weight, the subjective bias caused by artificially setting weights can be effectively avoided. This data-driven weight generation mechanism is more in line with the trend of modern machine learning and modeling optimization.

[0067] Comprehensive score calculation interface: used to calculate the comprehensive score of each uncertain factor data item based on each index and the weight of each index of each uncertain factor data item using a weighted summation method;

[0068] It is worth noting that the above-mentioned four indices are introduced to characterize the impact of uncertainty factors on carbon emissions from different dimensions. However, these indices are independent and difficult to use directly for modeling or ranking comparison. Therefore, a weighted summation method is adopted to integrate information from multiple dimensions into a unified numerical indicator, forming a complete comprehensive evaluation system for uncertainty factors. This not only improves the structured degree of data, but also provides standardized input features for subsequent model training, enhancing the operability and scalability of the system.

[0069] It is worth further explaining that the final generated comprehensive score is incorporated into the input feature set of the carbon emission prediction model as a high-order feature to assist the model in understanding the impact trend of external uncertainty factors on carbon emissions. Compared with the traditional method of ignoring uncertainty factors, the present invention, through this comprehensive scoring mechanism, achieves a better portrayal of the potential impact of uncertainty factors on carbon emissions, improves the model's ability to perceive complex external disturbance factors, and enhances the accuracy and credibility of prediction results. Therefore, the comprehensive score is not only an effective means of quantifying uncertainty factors, but also one of the key supports for promoting the performance improvement of the entire prediction model.

[0070] It is worth further explaining that since the weight of each index is obtained by normalizing its deviation from the global mean, the weight itself has a certain degree of dynamics. This means that when faced with different types of historical uncertainty factor data sets or carbon emission backgrounds in different regions, the comprehensive score generation mechanism can automatically adapt to changes in data distribution and maintain the scientific nature and consistency of the scoring system. This flexibility makes the present invention not only suitable for the current data environment, but also has good migration and generalization capabilities, and can maintain stable performance in various carbon emission prediction tasks.

[0071] Feature set integration unit: used to integrate the comprehensive score of each uncertainty factor data item in the historical uncertainty factor dataset and the features extracted from the historical benchmark carbon emission dataset into a feature set;

[0072] It is worth noting that since the comprehensive score comes from different uncertainty factor data items and their original value ranges are prone to large differences, the standardization of each index mentioned above can make the comprehensive score obey a distribution with a mean of 0 and a standard deviation of 1, thereby avoiding the impact of scale differences on model learning accuracy. On this basis, it is still necessary to perform outlier detection and correction on the comprehensive score of each uncertainty factor data item, that is, use the IQR method or Z-score method to identify outliers in the comprehensive score. For data points that deviate far from the overall distribution, use adjacent values ​​to replace them or set upper and lower limits. Finally, the processed comprehensive score is used as a new numerical feature and is combined with the features extracted from the historical benchmark carbon emissions dataset for feature splicing, that is, integration, to form a complete training feature set.

[0073] It should be noted that the comprehensive score of each uncertainty factor data item in the historical uncertainty factor dataset not only reflects the potential impact of the uncertainty factor on carbon emissions, but also has good interpretability and operability, and can be directly identified and learned by machine learning models. In practical applications, these comprehensive scores will serve as important input features in the training stage of the carbon emission prediction model, and will participate in model training together with historical benchmark carbon emission features. In this way, the model can not only learn the historical laws of carbon emissions, but also perceive and adapt to the changing trends of external uncertainty factors, thereby significantly improving the accuracy and robustness of the prediction results.

[0074] Carbon emission prediction model training module: connected to the feature set acquisition module, used to build a carbon emission prediction model, and integrate historical carbon emission data with the feature set into a training set, and use the training set to train the carbon emission prediction model to obtain a trained carbon emission prediction model;

[0075] It is worth noting that after completing the collection, preprocessing, feature extraction, and comprehensive score calculation of historical carbon emission data, including historical baseline carbon emission data sets and uncertainty factor data sets, the next step is to build a carbon emission prediction model based on these data. Among them, the carbon emission prediction model proposed in the present invention is a time series prediction model driven by multiple features and integrating uncertainty factor information. Its training process not only relies on the historical carbon emission data itself, but also deeply integrates the structured features generated by the uncertainty factor quantitative evaluation system, thereby significantly improving the model's predictive ability and robustness for future carbon emission trends.

[0076] The composition of the training set required in the carbon emission prediction model training process includes two parts: first, the historical benchmark carbon emission data set, including basic data such as time series carbon emissions and energy structure, which has undergone pre-processing operations such as standardization, missing value filling, and outlier correction, and is the core basis for the model to learn the basic laws of carbon emissions; second, the feature set, including time trend characteristics, periodic characteristics, energy structure-related characteristics, regional characteristics, etc. extracted from the historical benchmark carbon emission data set, and also includes a comprehensive score generated after quantitative evaluation of uncertainty factor data items. Among them, the four indexes of impact intensity, change fluctuation, development trend, and stability and reliability generated after quantitative evaluation of uncertainty factor data items can also be used as part of the feature set, because these features can together constitute the key input for the model to understand external disturbance factors; it is worth emphasizing that: the present invention innovatively incorporates the quantified results of uncertainty factors as important features into the training set, so that the model can not only learn the historical laws of carbon emissions, but also perceive and adapt to uncertainty disturbances in the external environment, thereby improving prediction accuracy and generalization ability;

[0077] The input to the carbon emission prediction model is a multidimensional feature vector, with each set of inputs corresponding to a time period such as a month or quarter. These include, but are not limited to: timestamps such as year and month; historical carbon emissions; energy structure characteristics such as the proportion of coal, natural gas, and renewable energy; industrial structure characteristics such as the proportion of carbon emissions from industry, services, and agriculture; seasonal characteristics such as seasonal dummy variables and temperature change rates; and quantitative characteristics of uncertainty factors such as a comprehensive score, which can include four corresponding indices. In addition, all input features have been standardized and normalized to ensure the stability of model training.

[0078] The output of the model is a numerical variable, such as the expected carbon emissions in the next month or quarter, that is, the expected carbon emissions in the next time period;

[0079] Furthermore, the training label of the carbon emission prediction model is the actual carbon emissions in the next time period, that is, given all the features of the current time period, the carbon emissions value in the next time period is predicted;

[0080] Furthermore, regarding the selection of loss function, the present invention preferably adopts mean square error as the main loss function; and regarding model evaluation indicators, the present invention adopts the following multiple evaluation indicators, such as mean absolute error, mean absolute percentage error, and coefficient of determination. These indicators will be evaluated separately on the validation set and the test set for model tuning and performance comparison;

[0081] Therefore, based on the above, the carbon emission prediction model training mechanism in the present invention is not only a powerful supplement to the existing technology, but also a key technical support for promoting carbon emission prediction from "experience-driven" to "data + knowledge dual-driven".

[0082] The data analysis module to be predicted is connected to the carbon emission prediction model training module to obtain the carbon emission data to be predicted, analyze the carbon emission data to be predicted, and obtain the final carbon emission data to be predicted;

[0083] Data quantification unit to be predicted: used to calculate the comprehensive score of each uncertainty factor data item in the uncertainty factor data set of the carbon emission data to be predicted;

[0084] Similarity value analysis unit: used to analyze the similarity between the comprehensive scores of each uncertainty factor data item in the uncertainty factor data set, and obtain the similarity value between the comprehensive scores of each uncertainty factor data item;

[0085] One-dimensional feature vector construction interface: used to construct a one-dimensional feature vector based on the comprehensive scores of all uncertain factor data items, where each element in the one-dimensional feature vector represents the comprehensive score of an uncertain factor data item;

[0086] Similarity value calculation interface: used to calculate the numerical proximity of the comprehensive score between each uncertain factor data item using a pairwise comparison method, and obtain the similarity value of the comprehensive score between each uncertain factor data item;

[0087] It is worth noting that the core technology in the stage from the one-dimensional feature vector construction interface to the similarity value calculation interface is to provide a specific and feasible similarity calculation method to support the subsequent screening decision of uncertainty factors. Due to the wide variety of uncertainty factors and their different manifestations, if there is a lack of a unified quantitative standard, it is difficult to determine which factors have high duplication or similarity, which affects the scientific nature of the screening results. By constructing a one-dimensional feature vector and performing pairwise comparisons, the degree of difference between the factors can be intuitively reflected, providing a reliable data basis for the subsequent calculation of the average similarity and setting of the screening threshold. At the same time, this method has good scalability and engineering feasibility, is suitable for data sets of different sizes, and is easy to deploy and implement in actual systems; therefore, the core technology in the stage from the one-dimensional feature vector construction interface to the similarity value calculation interface provides key technical support for the uncertainty factor screening mechanism, enhances the model's ability to identify complex disturbance factors in the prediction stage, improves the system's intelligence level and application adaptability, and is one of the important technical means to promote the carbon emission prediction model from experience-driven to data-driven.

[0088] A screening threshold setting unit is used to calculate a first average similarity value between the comprehensive scores of all uncertainty factor data items, and set a screening threshold value according to the first average similarity value;

[0089] Uncertainty factor data item retaining unit: used to screen each uncertainty factor data item according to the screening threshold value, and determine the retained uncertainty factor data item based on the screening result;

[0090] Second average similarity value calculation interface: used to calculate the second average similarity value between the current uncertain factor data item and all other uncertain factor data items based on the similarity value between each uncertain factor data item;

[0091] Screening interface: used to compare the second average similarity value of each uncertainty factor data item with the screening threshold;

[0092] Screening result acquisition interface: used to determine whether to retain the current uncertain factor data item if the second average similarity value of the current uncertain factor data item is less than the screening threshold; if the second average similarity value of the current uncertain factor data item is greater than or equal to the screening threshold, then eliminate the current uncertain factor data item;

[0093] It is worth noting that the core technology purpose of the stage from the second average similarity value calculation interface to the screening result acquisition interface is to provide a specific screening judgment rule, which realizes the automation and standardization of the uncertainty factor screening process. In actual operation, there are often deviations in setting the screening conditions based solely on subjective judgment or experience. By introducing the statistical indicator of the second average similarity, the screening process has stronger objectivity and repeatability. When the overall similarity between a certain uncertainty factor and other factors is high, it indicates that its information contribution is limited and should be eliminated; otherwise, it should be retained. This automated screening mechanism not only improves the efficiency of the model operation, but also significantly reduces the uncertainty caused by human intervention, and enhances the stability and credibility of the model prediction results. Therefore, the core technology of the stage from the second average similarity value calculation interface to the screening result acquisition interface strengthens the model's input quality control capability in the prediction stage, is an indispensable part of achieving high-precision carbon emission prediction, and fully reflects the systematic design and technological advancement of the present invention in uncertainty factor processing.

[0094] Final uncertainty factor data set acquisition unit: used for aggregating each uncertainty factor data item determined to be retained to generate a final uncertainty factor data set;

[0095] Final data acquisition unit to be predicted: used to integrate the final uncertainty factor data set with the benchmark carbon emission data set to obtain the final carbon emission data to be predicted;

[0096] It is worth noting that in the stage from the quantification unit of the data to be predicted to the final acquisition unit of the data to be predicted, the core role of this technology is to introduce an uncertainty factor screening mechanism based on comprehensive scoring, so as to optimize the structure and quality of the model input features. In actual application, faced with a large number of uncertainty factors that will appear in the future, if they are not screened and all included in the model, noise information will inevitably be introduced, affecting the model's prediction accuracy and stability. Therefore, by constructing a comprehensive scoring system and performing similarity analysis, redundant or inefficient factors can be effectively identified, and truly representative key factors can be retained, thereby improving the model's prediction efficiency and accuracy. In addition, the screening mechanism has good dynamic adaptability and can continuously update the factor set over time. It is suitable for carbon emission prediction tasks under complex and changeable scenarios such as policy adjustments and climate change. Therefore, the core technology of the stage from the quantification unit of the data to be predicted to the final acquisition unit of the data to be predicted provides important support for connecting the training stage and the prediction stage, and is a key link in achieving the "efficient, accurate, and explainable" carbon emission prediction goals.

[0097] Prediction result acquisition module: connected to the data analysis module to be predicted, used to input the final carbon emission data to be predicted into the trained carbon emission prediction model and output the carbon emission prediction result;

[0098] It is worth noting that the final carbon emission data to be predicted specifically includes the quantitative characteristics of the baseline carbon emission data and the uncertainty factor data items, and the carbon emission prediction result finally output by the trained carbon emission prediction model is a numerical variable, such as an enterprise's carbon emissions are expected to be 500 tons next month; the final output can help enterprises optimize production plans and control carbon footprints.

[0099] It should be noted that the terms used in the present invention are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.

[0100] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. Carbon emission prediction and optimization system, characterized by: Includes the following modules: Historical data collection module: used to collect historical carbon emission data and pre-process the historical carbon emission data, wherein the historical carbon emission data includes a historical baseline carbon emission data set and a historical uncertainty factor data set; Feature set acquisition module: connected to the historical data collection module, used to extract features from the pre-processed historical baseline carbon emission dataset and historical uncertainty factor dataset, and integrate the extracted features into a feature set; Carbon emission prediction model training module: connected to the feature set acquisition module, used to build a carbon emission prediction model, and integrate historical carbon emission data with the feature set into a training set, and use the training set to train the carbon emission prediction model to obtain a trained carbon emission prediction model; The data analysis module to be predicted is connected to the carbon emission prediction model training module to obtain the carbon emission data to be predicted, analyze the carbon emission data to be predicted, and obtain the final carbon emission data to be predicted; Prediction result acquisition module: connected to the data analysis module to be predicted, used to input the final carbon emission data to be predicted into the trained carbon emission prediction model and output the carbon emission prediction result.

2. The carbon emission prediction and optimization system according to claim 1, characterized in that: The method extracts features from the pre-processed historical baseline carbon emission dataset and the historical uncertainty factor dataset, and integrates the extracted features into a feature set, including: Historical data quantification unit: used to perform quantification processing on the historical uncertainty factor data set and obtain a comprehensive score for each uncertainty factor data item in the historical uncertainty factor data set based on the quantification result; Feature set integration unit: used to integrate the comprehensive score of each uncertainty factor data item in the historical uncertainty factor dataset and the features extracted from the historical benchmark carbon emission dataset into a feature set.

3. The carbon emission prediction and optimization system according to claim 2, characterized in that: The quantification process is performed on the historical uncertainty factor dataset, and a comprehensive score of each uncertainty factor data item in the historical uncertainty factor dataset is obtained based on the quantification result, including: Index calculation interface: used to calculate the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertainty factor data item in the historical uncertainty factor data set; Weight analysis interface: used to calculate the weight of each index based on the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertain factor data item; Comprehensive score calculation interface: used to calculate the comprehensive score of each uncertain factor data item based on each index of each uncertain factor data item and the weight of each index, using a weighted summation method.

4. The carbon emission prediction and optimization system according to claim 3, characterized in that: The weight of each index is calculated based on the impact intensity index, change fluctuation index, development trend index and stability reliability index of each uncertain factor data item, including: Standardization processing sub-interface: used to perform standardization on each index of each uncertainty factor data item in the historical uncertainty factor dataset; Global mean and standard deviation calculation sub-interface: used to calculate the global mean and standard deviation of each index of each uncertainty factor data item in the entire historical uncertainty factor data set based on the standardization processing results; Deviation calculation sub-interface: used to calculate the deviation of each index relative to the global mean for each uncertainty factor data item; Weight calculation sub-interface: used to normalize the deviation of each index of each uncertainty factor data item relative to the global mean to obtain the weight of each index of each uncertainty factor data item.

5. The carbon emission prediction and optimization system according to claim 1, characterized in that: The analysis of the carbon emission data to be predicted to obtain the final carbon emission data to be predicted includes: Data quantification unit to be predicted: used to calculate the comprehensive score of each uncertainty factor data item in the uncertainty factor data set of the carbon emission data to be predicted; Similarity value analysis unit: used to analyze the similarity between the comprehensive scores of each uncertainty factor data item in the uncertainty factor data set, and obtain the similarity value between the comprehensive scores of each uncertainty factor data item; A screening threshold setting unit is used to calculate a first average similarity value between the comprehensive scores of all uncertainty factor data items, and set a screening threshold value according to the first average similarity value; Uncertainty factor data item retaining unit: used to screen each uncertainty factor data item according to the screening threshold value, and determine the retained uncertainty factor data item based on the screening result; Final uncertainty factor data set acquisition unit: used for aggregating each uncertainty factor data item determined to be retained to generate a final uncertainty factor data set; Final data acquisition unit to be predicted: used to integrate the final uncertainty factor data set with the benchmark carbon emission data set to obtain the final carbon emission data to be predicted.

6. The carbon emission prediction and optimization system according to claim 5, characterized in that: The analyzing the similarity between the comprehensive scores of each uncertainty factor data item in the uncertainty factor data set to obtain the similarity value between the comprehensive scores of each uncertainty factor data item includes: One-dimensional feature vector construction interface: used to construct a one-dimensional feature vector based on the comprehensive scores of all uncertain factor data items, where each element in the one-dimensional feature vector represents the comprehensive score of an uncertain factor data item; Similarity value calculation interface: used to calculate the numerical proximity of the comprehensive score between each uncertain factor data item using a pairwise comparison method, and obtain the similarity value of the comprehensive score between each uncertain factor data item.

7. The carbon emission prediction and optimization system according to claim 5, characterized in that: The analyzing the similarity between the comprehensive scores of each uncertainty factor data item in the uncertainty factor data set to obtain the similarity value between the comprehensive scores of each uncertainty factor data item includes: Second average similarity value calculation interface: used to calculate the second average similarity value between the current uncertain factor data item and all other uncertain factor data items based on the similarity value between each uncertain factor data item; Screening interface: used to compare the second average similarity value of each uncertainty factor data item with the screening threshold; Screening result acquisition interface: used to determine whether to retain the current uncertain factor data item if the second average similarity value of the current uncertain factor data item is less than the screening threshold; if the second average similarity value of the current uncertain factor data item is greater than or equal to the screening threshold, then eliminate the current uncertain factor data item.