Enterprise electricity and carbon calculation method based on clustering and integration algorithm
The enterprise carbon emissions calculation method based on clustering and integration algorithms solves the problems of high equipment cost, high data collection requirements and slow data update in existing technologies, realizes accurate and timely carbon emissions monitoring and management, and is suitable for multiple types of enterprises.
Patent Information
- Application Number
- CN202510730455.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
AI Technical Summary
Existing corporate carbon emissions calculation methods have problems such as high equipment costs, high data collection requirements, and slow data updates, making it difficult to achieve accurate and timely carbon emissions monitoring and management.
Using a method based on clustering and integration algorithms, through data preprocessing, PCA dimensionality reduction, cluster analysis and XGBoost integrated learning model, a calculation model for energy consumption and output of enterprise groups is constructed to calculate the carbon emission factors of enterprises.
It improves the accuracy and timeliness of carbon emission calculations, reduces the difficulty and cost of data collection, and is suitable for multiple types of enterprises to meet the carbon emission calculation needs of different enterprise groups.
Smart Images

Figure CN120671970A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of carbon emission calculation, and specifically relates to a method for calculating enterprise electricity carbon based on clustering and integration algorithms. Background Art
[0002] As the pressure of global climate change intensifies, countries around the world are enacting and implementing stricter carbon emission regulations and standards to effectively constrain corporate carbon emissions. These measures not only provide the necessary data foundation for the operation of carbon trading markets but also help companies achieve compliance. Consequently, there is an urgent need for regulators and businesses to strengthen the monitoring and calculation of carbon emissions.
[0003] Currently, there are three methods for calculating carbon emissions for enterprises, namely direct measurement method, material balance method and carbon emission factor method.
[0004] 1. Direct measurement method
[0005] The direct measurement method uses a gas analyzer to directly measure the gas concentration and flow rate of the greenhouse gas emission source of the measured object, calculate the carbon emissions and compare them with known standards to obtain accurate measurement results. This method has the following advantages:
[0006] 1) Accurate Data: Direct measurement provides objective and unbiased data because it is based on direct observation or measurement and is not affected by the researcher's subjective interpretation.
[0007] 2) Intuitive results: By directly comparing with known standards, more accurate measurement results can be obtained, and the results are intuitive and easy to understand.
[0008] It also has the following disadvantages:
[0009] 1) High cost: Direct measurement requires expensive equipment and instruments, as well as modifications to existing equipment and devices.
[0010] 2) Reactivity bias: Participants may change their behavior because they know they are being observed or measured, which leads to reactivity bias.
[0011] 2. Material Balance Method
[0012] The material balance method is based on the principle of conservation of mass. It calculates carbon emissions by calculating the input, output, and loss of materials during the production process. Specifically, it requires detailed records and measurements of the quality and composition of the incoming and outgoing materials of an enterprise or facility, and then calculates carbon emissions based on this data. This method has the following advantages:
[0013] 1) Accurate calculation: The material balance method can fully consider the transformation and loss of materials in the production process, so the calculation results of this method are relatively accurate.
[0014] 2) Wide applicability: It is applicable to all types of enterprises and facilities. As long as the input, output and loss of materials can be accurately recorded and measured, carbon emissions can be calculated.
[0015] It also has the following disadvantages:
[0016] 1) Large workload: It is necessary to collect detailed industrial production process data and fully understand the production process, chemical reactions, side reactions and management.
[0017] 2) High management requirements: The company has high requirements on production technology and management level, and needs to ensure accurate measurement and recording of materials.
[0018] 3. Carbon Emission Factor Method
[0019] The carbon emission factor method estimates carbon emissions by multiplying activity data by an emission factor. Activity data refers to the amount of activity related to carbon emissions (such as energy consumption and product output), while the emission factor refers to the amount of carbon emissions generated per unit of activity. This method has the following advantages:
[0020] 1) Simple and clear: The carbon emission factor method is simple, clear and easy to understand, and is applicable to all types of enterprises and facilities.
[0021] 2) Widely used: It is the first carbon emission estimation method proposed by the IPCC and is also the most widely used method at present.
[0022] It also has the following disadvantages:
[0023] 1) Large uncertainty: The carbon emission factor is affected by many factors such as technology level, production conditions, energy utilization and process, so there is a certain degree of uncertainty.
[0024] 2) Slow data update: Since the determination of emission factors requires a large amount of experimental data and statistical analysis, their update speed may be relatively slow and cannot reflect the latest technology and production conditions in a timely manner.
[0025] In summary, the current methods for calculating corporate carbon emissions have the following shortcomings due to the limitations of the technology itself and the complexity of practical application:
[0026] 1. The cost of equipment and facility modification for direct measurement limits its widespread application;
[0027] 2. The material balance method has high requirements for enterprise management and data collection, making it difficult for many enterprises to implement;
[0028] 3. The carbon emission factor method requires statistical data, which results in slow data updates and makes it impossible to timely monitor corporate carbon emissions and provide process management monitoring for regulatory agencies. Summary of the Invention
[0029] The purpose of the present invention is to provide a method for calculating enterprise electricity carbon based on clustering and integration algorithms, which solves the limitations of existing carbon emission calculation methods.
[0030] The present invention is achieved through the following technical solutions:
[0031] A method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms includes the following steps:
[0032] S1. Collect enterprise data;
[0033] S2. Preprocess the enterprise data to obtain the original data;
[0034] S3. Preliminarily classify the original data into corresponding industry and enterprise groups, and add industry classification data to the original data to form industry and enterprise group data;
[0035] S4. Perform PCA dimensionality reduction on the industry enterprise group data to obtain dimensionality-reduced data;
[0036] S5. Use clustering algorithm to perform cluster analysis on the dimensionality reduction data to obtain classified data;
[0037] S6. Input the classified data into a pre-built enterprise group energy consumption and output calculation model to calculate the energy consumption of the enterprise;
[0038] S7. Calculate the enterprise’s carbon emission factor based on the enterprise’s energy consumption.
[0039] Furthermore, in S1, enterprise data includes enterprise basic information, enterprise electricity data, enterprise energy consumption data and enterprise production data.
[0040] Furthermore, in S2, data preprocessing specifically includes: setting a time period to be calculated, extracting enterprise data in the time period, and using the complete enterprise data as the original data.
[0041] Furthermore, in S4, PCA dimensionality reduction is performed on the industry enterprise group data to obtain dimensionality-reduced data, which specifically includes the following steps:
[0042] Standardize the data of industry and enterprise groups to obtain standardized data;
[0043] Using standardized data, construct a standardized data matrix;
[0044] Calculate the covariance matrix of the standardized data matrix. The covariance matrix is used to describe the linear relationship between different features.
[0045] Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and corresponding eigenvectors;
[0046] Sort the eigenvectors according to the size of the eigenvalues of the data set of the industry enterprise group, and select a preset number of eigenvalues with the highest order to construct the eigenvector space;
[0047] The dataset of the industry enterprise group is projected into the selected feature vector space to obtain the dimension-reduced data.
[0048] Furthermore, S5 specifically includes the following steps:
[0049] S5.1. Divide the dimensionality-reduced data into three clusters; randomly select one data point in each cluster as the initial centroid, thus forming three initial centroids;
[0050] S5.2. For each data point, calculate its distance to all initial centroids and assign each data point to the cluster corresponding to the nearest centroid.
[0051] S5.3. Calculate the average distance between all data points in each cluster and the initial centroid of the cluster, and update the centroid;
[0052] Repeat S5.2 and S5.3 until the centroid no longer changes significantly or the maximum number of iterations is reached, and the enterprise cluster classification characteristics of each enterprise are obtained. Combined with the standardized data matrix, the industry enterprise cluster dataset is output.
[0053] Furthermore, in S6, the construction process of the pre-built enterprise group energy consumption and output calculation model is as follows:
[0054] The industry enterprise group dataset is divided into a training set and a test set. Energy consumption is used as the dependent variable, and electricity and enterprise category are used as independent variables. Based on the XGBoost ensemble learning model, the grid search method is used for parameter selection. The optimal parameters are obtained through training, and a qualified XGBoost ensemble learning model is obtained as the energy consumption and output calculation model of the enterprise group.
[0055] Furthermore, in S7, the energy carbon emission factor is used to calculate the carbon emission factor of the enterprise. The specific calculation formula is as follows:
[0056]
[0057] Among them, E is the carbon emission factor of the enterprise; AC i is the consumption of the i-th energy; EF iis the carbon emission factor of the i-th energy source; i = 1, 2, 3...n, where n is the number of energy types.
[0058] The present invention also discloses an enterprise electric carbon calculation system based on clustering and integration algorithms, comprising:
[0059] Data collection module, used to collect enterprise data;
[0060] The preprocessing module is used to preprocess enterprise data to obtain original data;
[0061] The preliminary classification module is used to preliminarily classify the original data into corresponding industry and enterprise groups, and add industry classification data to the original data to form industry and enterprise group data;
[0062] Dimensionality reduction module, used to perform PCA dimensionality reduction on industry enterprise group data to obtain dimensionality-reduced data;
[0063] Clustering module, used to perform cluster analysis on dimensionality-reduced data using clustering algorithms to obtain classified data;
[0064] The first calculation module is used to input the classified data into a pre-built enterprise group energy consumption and output calculation model to calculate the energy consumption of the enterprise;
[0065] The second calculation module is used to calculate the enterprise's carbon emission factor based on the enterprise's energy consumption.
[0066] The present invention also discloses a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the online beat tracking method are implemented.
[0067] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the online beat tracking method are implemented.
[0068] Compared with the prior art, the present invention has the following beneficial technical effects:
[0069] The present invention discloses a method for calculating enterprise electric carbon based on clustering and integration algorithms. The method performs data preprocessing on enterprise data, which can remove noise, fill missing values, process abnormal data, etc., and improve data quality. The preprocessed data can more accurately reflect the actual situation of the enterprise, reduce the impact of data errors on subsequent calculations, lay the foundation for building an accurate electric carbon calculation model, and help improve the stability and generalization ability of the model. Classifying the original data into corresponding industry enterprise groups and adding industry classification data can help consider the characteristics and differences of different industries. Through this classification, electric carbon calculations can be performed on enterprises in different industries in a more targeted manner, improving the accuracy and rationality of the calculation results. PCA dimensionality reduction is performed on the industry enterprise group data. PCA dimensionality reduction can remove redundant information in the data, reduce data dimensions, and reduce calculation amount and storage space; clustering analysis is performed on the dimensionality reduction data using a clustering algorithm, which can classify enterprises with similar characteristics into one category, which helps to discover potential relationships and patterns between enterprises; the energy consumption of the enterprise is calculated using a pre-constructed enterprise group energy consumption and output calculation model. The model can accurately calculate the energy consumption of the enterprise based on factors such as the enterprise's production characteristics, industry classification, and historical data. The final calculation results in the enterprise carbon emission factor, which provides a key indicator for the enterprise's carbon emission accounting and management. The carbon emission factor is an important basis for measuring the carbon emission level of an enterprise. By accurately calculating this factor, the enterprise can understand its own carbon emission status and formulate corresponding emission reduction targets and measures. In summary, the present invention has made significant improvements in the enterprise carbon emission accounting method. Through the integration of clustering and integration algorithms, it not only improves the computational applicability, but also reduces the difficulty and cost of data collection, increases the timeliness of data, and meets the carbon emission calculation needs of different enterprise groups. It has good prospects for promotion and application.
[0070] Furthermore, enterprise data includes basic enterprise information, enterprise electricity data, enterprise energy consumption data and enterprise production data, which can comprehensively obtain information related to enterprise electricity carbon calculation and provide a basis for subsequent analysis and calculation. Accurate and comprehensive data collection helps to improve the accuracy and reliability of electricity carbon calculation.
[0071] Furthermore, a pre-built enterprise cluster energy consumption and output calculation model was constructed based on the XGBoost ensemble learning model. The XGBoost ensemble algorithm effectively utilizes electricity data and enterprise-related feature data. Based on the principle of gradient boosting, it integrates multiple decision trees to predict carbon emissions. Grid search is used to optimize algorithm parameters, giving the model greater generalization capabilities. By selecting optimal parameters and optimizing the ensemble algorithm, this method ensures accurate and robust carbon emissions calculations for complex enterprise clusters, reducing the risk of model overfitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1This is a flow chart of a method for calculating enterprise electricity carbon based on clustering and integration algorithms of the present invention;
[0073] Figure 2 This is a schematic diagram of an enterprise electric carbon calculation system based on clustering and integration algorithms of the present invention;
[0074] Figure 3 This is a schematic diagram of industry enterprise group classification based on Kmeans;
[0075] Figure 4 Schematic diagram of using the XGBoost algorithm to build an enterprise group energy consumption and output calculation model. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following is a further detailed description with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. That is, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments.
[0077] The components described and illustrated in the drawings and embodiments of the present invention may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the following drawings is not intended to limit the scope of the claimed invention, but merely represents a selected embodiment of the present invention. All other embodiments derived by those skilled in the art based on the drawings and embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.
[0078] It should be noted that the terms "comprises", "includes" or any other variations are intended to cover non-exclusive inclusion, so that a process, element, method, article or apparatus that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to the process, element, method, article or apparatus.
[0079] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0080] The present invention proposes a method for calculating enterprise electricity carbon based on clustering and integration algorithms. It collects relevant data such as basic enterprise information, electricity, energy consumption and product characteristics, uses aggregation algorithms to classify enterprise groups, and then uses integration algorithms to calculate the enterprise's carbon emissions using electricity data.
[0081] Example 1
[0082] The present invention discloses a method for calculating enterprise electricity carbon based on clustering and integration algorithms, comprising the following steps:
[0083] S1. Collect enterprise data;
[0084] S2. Preprocess the enterprise data to obtain the original data;
[0085] S3. Preliminarily classify the original data into corresponding industry and enterprise groups, add industry classification data to the original data, and form industry and enterprise group data;
[0086] S4. Perform PCA (Principal Component Analysis) dimensionality reduction on the industry enterprise group data to obtain reduced-dimensionality data;
[0087] S5. Use clustering algorithm to perform cluster analysis on the dimensionality reduction data to obtain classified data;
[0088] S6. Input the data into a qualified enterprise group energy consumption and output calculation model to calculate the energy consumption of the enterprise;
[0089] S7. Calculate the enterprise’s carbon emission factor based on the enterprise’s energy consumption.
[0090] This method uses a clustering algorithm to classify companies, dividing them into multiple clusters despite having different energy structures and production processes but belonging to the same industry. This approach addresses the poor applicability of traditional calculation methods. Based on historical enterprise data and an integrated algorithm, this method constructs a carbon emissions calculation model based on enterprise clusters. This approach is widely applicable to a wide range of enterprises, reducing calculation errors caused by enterprise diversity and ensuring the method's versatility and accuracy.
[0091] Focusing on electricity data and integrating it with highly relevant data such as company type, energy consumption, and product output, this method addresses the traditional carbon emissions accounting problem of relying on lagged statistical data. The high frequency and availability of electricity data enable the calculation method to promptly reflect changes in corporate carbon emissions, enabling real-time monitoring of carbon emissions and improving the timeliness of carbon emissions accounting.
[0092] Example 2
[0093] Based on Example 1, S1-S3 are mainly introduced.
[0094] The company's basic information, electricity consumption, energy consumption, production and other related data are collected from the company's ledgers or the production and operation data reported to the competent authorities. Specific data details are shown in Table 1 below.
[0095] Table 1 Enterprise data collection table
[0096]
[0097]
[0098]
[0099] The data preprocessing in S2 is as follows:
[0100] Set the time period to be calculated, extract the enterprise data for that time period, and use the complete enterprise data without any missing data as the original data.
[0101] For example, if you need to calculate the data for the period from January 2024 to December 2024, then set the time period to January 2024 to December 2024, extract the data for this time period, and retain the complete enterprise data as the original data.
[0102] The process of preliminary classification of raw data is as follows:
[0103] This paper uses the subcategories in the "National Economic Industry Classification GB / T 4754-2017" to divide the enterprise industry based on the basic information of the enterprise to form industry enterprise groups.
[0104] Industry classification data will be added to enterprise data, and the industry classification data includes industry classification code data and industry classification name data.
[0105] For example, the main business activity of a certain enterprise is the production of children's clothing. After inquiry, it is found that its industry code is 1810 (woven clothing manufacturing), which belongs to the subcategory 1810 under the middle category 181 (woven clothing manufacturing) in the major category 18 (textile clothing and apparel industry) under the manufacturing category.
[0106] Enterprises are categorized and aggregated according to their corresponding sub-category industry codes, and enterprises with the same sub-category code are grouped together as an industry enterprise group. For example, all enterprises with industry code 1810 can form a woven garment manufacturing enterprise group.
[0107] Example 3
[0108] Based on Example 1, S4 is introduced.
[0109] PCA dimensionality reduction is a commonly used dimensionality reduction technique. The core idea is to transform the original data into a new coordinate system through linear transformation, so that the variance of the data in the new coordinate system is maximized. The directions of maximum variance are the directions of the principal components. The goal is to find a set of mutually orthogonal coordinate axes so that the projection of the data onto these axes can retain the original data information as much as possible while removing correlation and noise in the data.
[0110] The present invention chooses to reduce the data dimension to less than 5 dimensions, which can significantly reduce the amount of computation and improve the efficiency of subsequent data analysis and model training. At the same time, the lower-dimensional model is relatively simple and less prone to overfitting, which can improve the model's generalization ability while ensuring model accuracy.
[0111] Using PCA to reduce dimensionality, selecting fewer than five principal components typically preserves the most critical information in the original data. Generally speaking, the first few principal components tend to capture the main trends and characteristics of the data. Reducing the dimensionality to fewer than five dimensions effectively compresses and simplifies the data while minimizing information loss.
[0112] PCA dimensionality reduction is used to reduce the multi-dimensional features of industry enterprise group data to 5 dimensions. The specific method is as follows:
[0113] 1) Standardized data
[0114] Since PCA dimensionality reduction is sensitive to the scale of the data, the industry enterprise group data is first standardized so that the mean of each feature is 0 and the standard deviation is 1. This is achieved through the following formula:
[0115]
[0116] Among them, x i is the data item value in the model training data set, μ is the mean, σ is the standard deviation, z i are the values after normalization.
[0117] 2) Calculate the covariance matrix
[0118] Using standardized data, a standardized data matrix is constructed.
[0119] Assume there are n companies, one company corresponds to one sample, and one sample has p features. Then the standardized data matrix consists of n samples and p features. Each row of the standardized data matrix represents the p features of a sample, and each column represents the pth feature of n samples. For example, x 1,p Represents the pth feature of the first sample. The specific expression is as follows:
[0120]
[0121] The covariance matrix C is calculated based on the standardized data matrix X to describe the linear relationship between different features. The covariance matrix C is expressed as:
[0122]
[0123] Where X is the normalized data matrix, X T is the transposed matrix of the normalized data.
[0124] 3) Eigenvalue decomposition
[0125] Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues and corresponding eigenvectors.
[0126] The eigenvalue represents the variance of the data in the direction of the eigenvector, and the eigenvector represents the direction of the new coordinate axis.
[0127] The expression of eigenvalue decomposition is: Cv=λv
[0128] Where λ is the eigenvalue of the matrix C, and v is the eigenvector of the matrix C corresponding to the eigenvalue λ.
[0129] 4) Select the principal components:
[0130] Assume that the dataset X of the industry enterprise group has 20 eigenvalues (λ1,λ2,...,λ20). Sort the eigenvectors according to the size of the eigenvalues (λ1,λ2,...,λ20), and select the eigenvectors (v1,v2,v3,v4,v5) corresponding to the first 5 eigenvalues (λ1,λ2,λ3,λ4,λ5) to form the basis values required for dimensionality reduction.
[0131] 5) Data conversion:
[0132] Project the normalized data matrix X onto the selected eigenvector space to obtain the reduced-dimensional data matrix, and then obtain the reduced-dimensional data. The expression is:
[0133] Z=XW=X[v1,v2,v3,v4,v5]
[0134] Among them, W is the matrix composed of the first 5 selected eigenvectors, and Z is the data matrix after dimensionality reduction.
[0135] Example 4
[0136] Based on Example 3, this example mainly introduces the process of constructing a calculation model for energy consumption and output of an enterprise group.
[0137] Similar to Example 2, a large amount of enterprise data needs to be collected in the early stage. Through data preprocessing, complete enterprise data with no missing data is extracted to create a model training data set.
[0138] The model training data set is preliminarily classified according to the industry to which the enterprises belong, forming enterprise clusters in different industries; then PCA dimensionality reduction is performed on the industry enterprise cluster data of different industries to reduce the data dimension and obtain reduced dimensionality data; then the Kmeans algorithm is used to perform cluster analysis on the reduced dimensionality data, and the enterprises are divided into three categories according to the spatial relationship characteristics between the data to realize the classification of enterprise clusters.
[0139] Cluster analysis is as follows: Figure 3 As shown in the figure, based on the electricity and energy consumption characteristics of enterprises in the industry enterprise group, the KMeans algorithm is used to cluster the enterprises in the industry, and an industry enterprise group with similar "electricity-energy" characteristics is constructed to achieve further classification of the industry enterprise group.
[0140] Specifically, the industry enterprise group is divided into three types of enterprises as follows:
[0141] 1) Determine parameters
[0142] In this embodiment, the industry enterprise groups are determined to be 3 clusters.
[0143] 2) Initialize the centroid
[0144] One data point is randomly selected in each cluster as the initial centroid, thus forming three initial centroids.
[0145] 3) Assign data points
[0146] For each enterprise data point, calculate its distance from all initial centroids and assign each data point to the cluster corresponding to the nearest centroid. This step can be expressed as follows:
[0147]
[0148] Among them, g i is a data point, c j is the initial centroid of the jth cluster, j = 1, 2, 3; Cluster(g i ) is the minimum distance from each enterprise data point to the initial centroid.
[0149] 4) Update the centroid
[0150] Calculate the average distance of all data points in each cluster to the initial centroid of the cluster and update the centroid. This step can be expressed as follows:
[0151]
[0152] in, is the distance from all data points in the jth cluster to the initial centroid of the cluster, n j is the number of data points in the jth cluster.
[0153] 5) Iteration
[0154] Repeat 3) and 4) until the centroid no longer changes significantly or the maximum number of iterations is reached, and the enterprise cluster classification characteristics of each enterprise are obtained. Combined with the standardized data matrix X, the industry enterprise cluster dataset X is output. label , as shown below.
[0155]
[0156] The industry enterprise group dataset X label The K-FOLD cross-validation method with a K value of 5 was used to divide the dataset into 5 training sets and test sets. Energy consumption was used as the dependent variable, and electricity and enterprise category were used as independent variables. Based on the XGBoost (eXtreme Gradient Boosting) ensemble learning model, the grid search method was used for parameter selection. The XGBoost ensemble learning model with the optimal parameters was obtained through training to realize the calculation of enterprise energy consumption and output.
[0157] The details are as follows:
[0158] (1) Using the K-FOLD cross-validation method with a K value of 5, the industry enterprise group dataset X label The dataset is randomly divided into 5 equal or nearly equal subsets. In each iteration, 4 of them are used as training sets and the remaining 1 is used as a test set. This ensures that each subset has a chance to be used as a test set to reduce the risk of overfitting and evaluate the generalization ability of the model.
[0159] (2) Combined with grid search, the parameters of the XGBoost ensemble learning model are optimized as follows:
[0160] 1) Combining the grid search principle, multiple parameter combinations (such as the maximum depth of the tree, learning rate, number of trees, subsampling ratio, feature sampling ratio and minimum loss reduction value of split nodes) are set for traversal.
[0161] 2) Initialize the XGBoost ensemble learning model
[0162] XGBoost starts with a simple initial model, the initial prediction value is the industry enterprise group dataset X label The mean of the dependent variable for all samples in . Assume that the initial forecast is but:
[0163]
[0164] Among them, y i For the industry enterprise group dataset X label The true value of all sample dependent variables in , m is the industry enterprise group data set X label The number of samples.
[0165] 3) For each iteration, the residual of each sample is calculated, that is, the difference between the actual value and the current model's predicted value. These residuals will be used to train a new decision tree.
[0166]
[0167] Among them, R i (t) is the residual of round t, Y i is the true value of the sample, is the predicted value of the previous iteration.
[0168] 4) XGBoost fits a new decision tree by optimizing the objective function, which aims to minimize the residual error of the current model. The objective function is defined as:
[0169]
[0170] in, Is the loss function, commonly used square loss
[0171] Ω(F t ) is a regularization term that limits the model complexity to avoid overfitting.
[0172] 5) Update the predicted value. The predicted value of the new model is equal to the previous round of predicted value plus the predicted value of the new tree. The update formula is:
[0173]
[0174] Where η is the learning rate, which controls the contribution of each tree to the final prediction. XGBoost achieves the optimal model performance by adjusting the learning rate and the number of trees.
[0175] 6) Introduce the regularization term Ω(f t ), constraining the complexity of each tree. This reduces the risk of overfitting, enables the model to generalize well on high-dimensional data, and completes model construction.
[0176] (3) Model training and optimal parameter selection: The model is trained using the optimal parameters in each K-Fold grouping, and the model performance is verified in the test set. After completing all group training, the results of each round of verification are averaged to evaluate the generalization ability and prediction effect of the model.
[0177] like Figure 4 As shown in the figure, the XGBoost algorithm is used to build an energy consumption and output calculation model for the enterprise group in the industry, and a comprehensive energy consumption calculation is performed on the target enterprise. The XGBoost algorithm generates multiple sample sets based on the data sampling of all enterprises in the enterprise group, and iteratively constructs a weak regression tree for each sample set. Finally, the calculation result of the comprehensive target enterprise on all weak regression trees is the enterprise's energy consumption.
[0178] Example 5
[0179] The energy consumption of each enterprise is calculated by the enterprise group energy consumption and output calculation model obtained in Example 4. The carbon emission factor of each enterprise is calculated by combining the energy carbon emission factor. The calculation formula is as follows:
[0180]
[0181] Among them, E is the carbon emission factor of the enterprise; AC i is the consumption of type i energy (including electricity consumption); EF i is the carbon emission factor of the i-th energy source (including the carbon emission factor of electricity); i = 1, 2, 3...n, where n is the number of energy types.
[0182] This method is an extension of the existing enterprise carbon emission calculation method. First, the collected enterprises are classified using the national economic industry classification standard. Secondly, the classified enterprises are classified using a clustering algorithm according to their energy consumption, electricity consumption, output and related characteristics to form enterprise groups. Based on the enterprise groups, according to the correlation between electricity data and energy consumption and output, an energy consumption and output calculation model for the enterprise group is constructed based on an integrated learning algorithm. Finally, the energy, product and electricity carbon emission factors are used to realize the calculation of enterprise carbon emissions.
[0183] Example 6
[0184] like Figure 2 As shown, the present invention discloses an enterprise electric carbon calculation system based on clustering and integration algorithms, comprising:
[0185] Data collection module, used to collect enterprise data;
[0186] The preprocessing module is used to preprocess enterprise data to obtain original data;
[0187] The preliminary classification module is used to preliminarily classify the original data into corresponding industry and enterprise groups, and add industry classification data to the original data to form industry and enterprise group data;
[0188] Dimensionality reduction module, used to perform PCA dimensionality reduction on industry enterprise group data to obtain dimensionality-reduced data;
[0189] Clustering module, used to perform cluster analysis on dimensionality-reduced data using clustering algorithms to obtain classified data;
[0190] The first calculation module is used to input the classified data into a pre-built enterprise group energy consumption and output calculation model to calculate the energy consumption of the enterprise;
[0191] The second calculation module is used to calculate the enterprise's carbon emission factor based on the enterprise's energy consumption.
[0192] Example 7
[0193] The present invention also discloses a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the online beat tracking method are implemented. The memory may include internal memory, such as a high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device. The processor, network interface, and memory are interconnected via an internal bus. The internal bus may be an industrial standard architecture bus, a peripheral component interconnect standard bus, an extended industrial standard architecture bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0194] Example 8
[0195] The present invention also discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer-readable storage medium implements the steps of the online beat tracking method. Specifically, the computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include random access memory and / or cache memory, etc. The non-volatile memory may include read-only memory, hard disk, flash memory, optical disk, magnetic disk, etc.
[0196] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) containing computer-usable program code.
[0197] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0198] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0199] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for calculating enterprise electricity carbon based on clustering and integration algorithms, characterized by: The following steps are involved: S1. Collect enterprise data; S2. Preprocess the enterprise data to obtain the original data; S3. Preliminarily classify the original data into corresponding industry and enterprise groups, and add industry classification data to the original data to form industry and enterprise group data; S4. Perform PCA dimensionality reduction on the industry enterprise group data to obtain dimensionality-reduced data; S5. Use clustering algorithm to perform cluster analysis on the dimensionality reduction data to obtain classified data; S6. Input the classified data into a pre-built enterprise group energy consumption and output calculation model to calculate the energy consumption of the enterprise; S7. Calculate the enterprise’s carbon emission factor based on the enterprise’s energy consumption.
2. The method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms according to claim 1 is characterized in that: In S1, enterprise data includes basic enterprise information, enterprise electricity data, enterprise energy consumption data and enterprise production data.
3. The method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms according to claim 1 is characterized in that: In S2, data preprocessing specifically includes: setting the time period to be calculated, extracting the enterprise data in the time period, and taking the complete enterprise data as the original data.
4. The method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms according to claim 1 is characterized in that: In S4, PCA dimensionality reduction is performed on the industry enterprise group data to obtain reduced dimensionality data, which specifically includes the following steps: Standardize the data of industry and enterprise groups to obtain standardized data; Using standardized data, construct a standardized data matrix; Calculate the covariance matrix of the standardized data matrix. The covariance matrix is used to describe the linear relationship between different features. Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and corresponding eigenvectors; Sort the eigenvectors according to the size of the eigenvalues of the data set of the industry enterprise group, and select a preset number of eigenvalues with the highest order to construct the eigenvector space; The dataset of the industry enterprise group is projected into the selected feature vector space to obtain the dimension-reduced data.
5. The method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms according to claim 1 is characterized in that: S5 specifically includes the following steps: S5.
1. Divide the dimensionality-reduced data into three clusters; randomly select one data point in each cluster as the initial centroid, thus forming three initial centroids; S5.
2. For each data point, calculate its distance to all initial centroids and assign each data point to the cluster corresponding to the nearest centroid. S5.
3. Calculate the average distance between all data points in each cluster and the initial centroid of the cluster, and update the centroid; Repeat S5.2 and S5.3 until the centroid no longer changes significantly or the maximum number of iterations is reached, and the enterprise cluster classification characteristics of each enterprise are obtained. Combined with the standardized data matrix, the industry enterprise cluster dataset is output.
6. The method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms according to claim 1 is characterized in that: In S6, the construction process of the pre-built enterprise group energy consumption and output calculation model is as follows: The industry enterprise group dataset is divided into a training set and a test set. Energy consumption is used as the dependent variable, and electricity and enterprise category are used as independent variables. Based on the XGBoost ensemble learning model, the grid search method is used for parameter selection. The optimal parameters are obtained through training, and a qualified XGBoost ensemble learning model is obtained as the energy consumption and output calculation model of the enterprise group.
7. The method for calculating enterprise electricity carbon emissions based on clustering and integration algorithms according to claim 1 is characterized in that: In S7, the energy carbon emission factor is used to calculate the enterprise's carbon emission factor. The specific calculation formula is as follows: Among them, E is the carbon emission factor of the enterprise; AC i is the consumption of the i-th energy; EF i is the carbon emission factor of the i-th energy source; i = 1, 2, 3...n, where n is the number of energy types.
8. An enterprise electricity carbon calculation system based on clustering and integration algorithm, characterized by: include: Data collection module, used to collect enterprise data; The preprocessing module is used to preprocess enterprise data to obtain original data; The preliminary classification module is used to preliminarily classify the original data into corresponding industry and enterprise groups, and add industry classification data to the original data to form industry and enterprise group data; Dimensionality reduction module, used to perform PCA dimensionality reduction on industry enterprise group data to obtain dimensionality-reduced data; Clustering module, used to perform cluster analysis on dimensionality-reduced data using clustering algorithms to obtain classified data; The first calculation module is used to input the classified data into a pre-built enterprise group energy consumption and output calculation model to calculate the energy consumption of the enterprise; The second calculation module is used to calculate the enterprise's carbon emission factor based on the enterprise's energy consumption.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the online beat tracking method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the online beat tracking method according to any one of claims 1 to 7 are implemented.