Multi-factor Carbon Emission Accounting Method and Device

By adopting spatial geographic weighting and ensemble learning methods in carbon emission accounting, each carbon emission driver is independently modeled and a multi-factor linear weighted expression is obtained through ensemble learning, which solves the problem of spatial heterogeneity not reflecting and the limitations of the research object expansion in traditional methods, and achieves a more flexible and accurate research on carbon emission drivers.

CN114662282BActive Publication Date: 2025-06-20SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210187457.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-06-20
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

Traditional carbon emission accounting and driving factor research methods have problems such as spatial heterogeneity and limitations in the expansion of research objects, making it difficult to flexibly expand the research objects, and it is assumed that the same factors have no difference in the impact of carbon emissions between different regions or enterprises.

Method used

A multifactorial carbon emission accounting method based on spatial geographic weighted and integrated learning is adopted, and each carbon emission driver is independently modeled through the combination of geographic weighted regression model (GWR) and logistic regression model (LR), and a multifactorial linear weighted expression is obtained through ensemble learning to quantify the contribution of each driver to carbon emissions.

Benefits of technology

It alleviates the multicollinear interference in multivariate modeling, improves the ability of the regression model to reflect the correlation between variables and carbon emissions, expands the research scope of carbon emission drivers, enhances the flexibility of carbon emission factor decomposition methods, and more accurately reflects spatial heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114662282B_ABST
    Figure CN114662282B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-factor carbon emission accounting method, including: selecting multiple driving factors that affect carbon emissions; independently modeling each driving factor as a base model for ensemble learning, performing ensemble learning on the base models of each driving factor to obtain a multi-factor linear weighted expression; and quantifying the contribution degree of each driving factor to carbon emissions according to the multi-factor linear weighted expression. This method better measures the different degrees of influence on carbon emissions caused by the internal differences such as economy, culture, and development level of each region; the research method of "independent modeling - ensemble learning" gets rid of the limitation of the need to construct a rigorous identity at the beginning stage in the general carbon emission factor decomposition process, alleviates the multicollinearity interference existing in the traditional multi-variable regression method, expands the research object range of the traditional factor decomposition method, and provides a more flexible way to conduct research and accounting on the driving factors of carbon emissions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research field of "carbon peak - carbon neutrality" (hereinafter referred to as "dual carbon"), and particularly relates to a multi - factor carbon emission accounting method based on geospatial weighted and ensemble learning. Background Technique

[0002] The fundamental to achieving the "dual carbon" goal lies in reducing carbon emissions. To achieve carbon emission reduction, the first task is to calculate carbon emissions following scientific and rigorous measurement methods. Currently, the carbon emission accounting methods are mainly divided into the direct method and the indirect method. The direct method mainly uses the default values of carbon emission sources and their emission coefficients in the IPCC guidelines as the basis for calculating carbon emissions at the national or regional boundaries. The indirect method decomposes the total carbon emissions of a region into multiple factors that affect carbon emissions in the region and quantifies the contribution degree of each factor to the total carbon emissions. As a comprehensive research method that reasonably considers the regional humanities, economy, and development level, the indirect method is more in line with the complex basic national conditions and provides an important scientific basis for formulating carbon emission reduction routes and achieving the "dual carbon" goal.

[0003] The existing idea of carbon emission factor decomposition is to construct an identity between carbon emissions and multiple driving factors and generate a regression model through parameter fitting. Commonly used identities include the Kaya identity and the STIRPAT identity, and their general content is to decompose carbon emissions into the sum or product of several factors (such as key indicators like economy, environment, technology, population, etc.). Commonly used index decomposition methods include the Logarithmic Mean Divisia Index (LMDI), which can decompose the research object without residuals. Traditional factor decomposition methods perform logarithmic transformation and index decomposition based on the identity to quantify the influence degree of different factors on carbon emissions and then identify the key driving factors. Due to the characteristics of simple form and reasonable explanation, traditional factor decomposition methods have been widely used in related research fields. However, there are two deficiencies in the specific applications in the current "dual carbon" research field: one is that it does not reflect the spatial effect. For a country with a large land area, due to its vast territory, there are differences in the development levels of humanities, economy, etc. in different regions, resulting in differences in the influence degree of the same carbon emission factors on carbon emissions in different regions, that is, spatial heterogeneity. Currently, most studies directly assume that the cross - sectional units of panel data are homogeneous, that is, there are no spatial differences in the economic behaviors between regions or enterprises, which does not conform to the actual situation. The other is the limitation in the expansion of the research object. Each driving factor in the identity is given a specific meaning, such as per capita GDP, unit energy consumption, etc. It requires a strong logical relationship between the driving factors and is difficult to flexibly expand the research object.

[0004] To address the above challenges, recent studies have attempted to introduce spatial statistical analysis and use spatial weight matrices to reflect geographical heterogeneity, such as the economic geographical structures of developed and underdeveloped regions, core and peripheral regions, etc. The Geographically Weighted Regression (GWR) model incorporates geographical location information as regression parameters through a spatial weight matrix, extending the ordinary linear regression model. It constructs the relationship between driving factors and the total carbon emissions using local fitting, enabling the regression parameters of a specific region to vary with the local geographical location in space. Although the GWR model effectively addresses the issue of spatial heterogeneity, its local fitting of multiple independent variables uses the same bandwidth and cannot reflect the differences among different factors during the regression process. To address the deficiencies of GWR, the Multiscale Geographically Weighted Regression (MGWR) model allows each independent variable to have a different level of spatial smoothing, giving each variable its own statistical standard, reducing estimation bias, and making the regression results more reliable. The bandwidth of each independent variable can reflect the spatial scale of the spatial process of each, and the multi-bandwidth method produces a spatial process model that is closer to the truth and more useful.

[0005] However, the problem of multicollinearity inevitably exists in the fitting process of the multi-variable regression model, that is, if there are highly linearly correlated variables among multiple independent variables, it will lead to a distorted regression effect of the model. In addition, neither GWR nor MGWR considers the spatial heterogeneity of carbon emission driving factors and does not intuitively reflect the contribution of each driving factor to the total carbon emissions. Summary of the Invention

[0006] The purpose of the present invention is to provide a multi-factor carbon emission accounting method and device.

[0007] The technical problem to be solved by the present invention is as follows: In traditional carbon emission accounting and driving factor research methods, first, an identity equation (such as Kaya, STIRPAT, etc.) needs to be constructed between carbon emissions and multiple driving factors, expressing carbon emissions in the form of a product of multiple driving factors. Then, the research time point and the reference time point are substituted into the formula and subtracted to obtain an expression for the increment of carbon emissions. A logarithmic transformation is performed on the increment expression to convert the original factor multiplication into the form of factor logarithm addition. The formula after logarithmic transformation reflects the relationship between the change amplitude of carbon emissions and the change amplitude of driving factors. Finally, the change in the total carbon emissions caused by each factor is calculated through index decomposition (such as LMDI, GDIM, etc.). This method mainly has two problems: First, when constructing the identity equation, the change in carbon emissions is decomposed into the product of several factors, and the specific economic meanings assigned to each factor (such as per capita GDP, unit energy consumption, etc.) depend on their logical relationships with each other, making it difficult to flexibly expand the research object. Second, the carbon emission accounting method based on index decomposition directly assumes that the impact of the same factor on carbon emissions is the same among different regions or enterprises. In fact, the geographical space lacks homogeneity, with economic geographical structures such as developed and underdeveloped regions, core and peripheral regions, etc. Assuming that there are heterogeneous differences in carbon emission driving factors among regions is more in line with reality.

[0008] The present invention is oriented to the "dual carbon" research field and designs and implements a carbon emission accounting method based on spatial geographical weighting and ensemble learning. This method separately models multiple predefined carbon emission driving factors in a spatially geographically weighted manner, constructs a logistic regression model (LR model) based on the geographically weighted regression model (GWR model), and obtains the final multi-factor linear weighted expression through ensemble learning. Among them, the GWR model reflects the spatial heterogeneity of carbon emission driving factors by introducing a spatial weight matrix, making the regression process closer to the real situation. Secondly, separate modeling of each research variable not only alleviates the multiple collinearity interference existing in multi-variable modeling but also improves the ability of the regression model to reflect the correlation between the research variable and carbon emissions. Finally, taking the GWR model as the base model, a linear weighted expression of all base models is obtained through logistic regression. The weight value of each base model is used as the influence degree of the research variable on carbon emissions, intuitively quantifying the contribution of each driving factor to carbon emissions, expanding the research scope of carbon emission driving factors, and improving the flexibility of the carbon emission factor decomposition method.

[0009] For the above purpose, the present invention adopts the following technical solutions:

[0010] A multi-factor carbon emission accounting method, comprising:

[0011] Select multiple driving factors that affect carbon emissions;

[0012] Independently model each driving factor as the base model of the ensemble learning, and perform ensemble learning on the base models of each driving factor to obtain a multi-factor linear weighted expression;

[0013] According to the multi-factor linear weighted expression, quantify the contribution degree of each driving factor to carbon emissions.

[0014] As a preferred implementation method, during the modeling process, take the driving factors as research variables, traverse the custom research variables, use the current research variable as the independent variable, and the total carbon emissions as the dependent variable.

[0015] As a preferred implementation method, separately construct a geographically weighted regression model for each driving factor as the base model of the ensemble learning, and select the best local fitting bandwidth according to each driving factor to perform regression on carbon emissions and the current driving factor in different geographical spaces.

[0016] As a preferred implementation method, perform k-fold cross-validation on each base model to generate single-dimensional features and combine them into the training dataset and test dataset required for the ensemble learning stage.

[0017] As a preferred implementation method, enter the ensemble learning stage, construct a logistic regression model as the combined model of the ensemble learning, and use the ensemble learning training dataset and test dataset composed of the generated features of all base models for training and validation; after the combined model training is completed, obtain a multi-factor linear weighted expression, and use the independent variable coefficients therein as the quantification values of the contribution degrees of the corresponding carbon emission driving factors in the total carbon emissions.

[0018] As a preferred implementation method, the k-fold cross-validation includes:

[0019] Take 1 subset from the training set evenly divided into k parts as the validation set, and the remaining k - 1 subsets as the training set for this round to train the model;

[0020] After the training is completed, first make a prediction on the validation set and then make a prediction on the test set, and cycle k rounds in sequence;

[0021] At this time, after k-fold cross-validation, obtain the predicted values of k validation sets and the predicted values of k test sets. Vertically combine the k predicted values of the current base model's validation sets as a feature dimension in the ensemble learning training set, and its sample quantity is the same as the original training set;

[0022] Take the mean of the k predicted values of the current base model's test sets as a feature dimension in the ensemble learning test set, and its sample quantity is the same as the original test set, and the k-fold cross-validation process ends.

[0023] As a preferred implementation, the number of independent variables in the constructed logistic regression model is consistent with the number of research variables; each feature dimension in the dataset is generated by a base model representing different driving factors, and the prediction object is the actual total carbon emissions. Then, the independent variables of the logistic regression model are made to correspond one-to-one with the driving factors. Finally, the linear weighted expression of all the base models is learned, and the coefficients of the independent variables in the expression are used as the weights of the research variables in the total carbon emissions, so as to quantify the influence degree of each driving factor on the total carbon emissions.

[0024] A multi-factor carbon emission accounting method, comprising:

[0025] The multi-factor independent modeling stage, including constructing an independent geographically weighted regression model for a single carbon emission driving factor as the base model, and also including the process of cross-validating the base model to generate an integrated learning dataset;

[0026] The integrated learning stage, including the process of constructing a logistic regression model for multiple base models as a combined model and learning the linear weighted expressions of multiple base models, and also including using the coefficients of the independent variables in the linear weighted expression as the quantification values of the contribution degrees of the corresponding carbon emission driving factors to the total carbon emissions.

[0027] A multi-factor carbon emission accounting device for executing the above method.

[0028] A method for generating quantification values of the contribution degrees of carbon emission driving factors to carbon emissions, comprising:

[0029] Providing the above device;

[0030] Providing multiple driving factors affecting carbon emissions;

[0031] The device generates quantification values of the contribution degrees of each driving factor to the total carbon emissions according to the above method.

[0032] Compared with the existing technology, the present invention has the following advantages and effects:

[0033] 1. Aiming at the limitation that the traditional carbon emission driving factor research method based on index decomposition overly relies on constructing identity equations, the present invention proposes a research method of "independent modeling - integrated learning", which alleviates the multicollinearity interference existing in multi-variable modeling, improves the ability of the regression model to reflect the correlation between research variables and carbon emissions, expands the range of research variables for factor decomposition, and enhances the flexibility of the carbon emission driving factor decomposition method.

[0034] 2. In view of the unreasonableness of the traditional carbon emission accounting method that directly assumes that the impact of the same factors on carbon emissions has no difference between different regions, the present invention uses a spatially geographically weighted method to model the carbon emission driving factors, and reflects the spatial heterogeneity of the carbon emission driving factors by introducing a spatial weight matrix, making the regression process closer to the real situation.

[0035] 3. The present invention proposes a method for quantifying the contribution degree of a single factor to the total carbon emissions based on ensemble learning. The GWR model is selected as the base model, and the LR model is selected as the combined model to obtain the linear weighted expression of all base models. The weights of the base models are used as the influence degree of the corresponding research variables on carbon emissions, intuitively quantifying the contribution of each factor to the total carbon emissions, which is of great significance for the related research on carbon emission driving factors.

[0036] 4. In the novel carbon emission accounting method proposed by the present invention, the GWR model is introduced to consider the spatial heterogeneity of carbon emission driving factors, better measuring the different influence degrees on carbon emissions caused by the internal differences such as economy, culture, and development degree of each region; the research method of "independent modeling - ensemble learning" gets rid of the limitation that a rigorous identity needs to be constructed at the beginning stage in the general carbon emission factor decomposition process, alleviating both the multiple collinearity interference existing in the traditional multi-variable regression method and expanding the research object range of the traditional factor decomposition method, providing a more flexible way to conduct research and accounting on carbon emission driving factors. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the structural diagram of the carbon emission accounting model based on geographically weighted regression and ensemble learning;

[0038] Figure 2 is the flowchart of the carbon emission accounting method based on geographically weighted regression and ensemble learning;

[0039] Figure 3 is the schematic diagram of the cross-validation training process of the carbon emission accounting model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To better understand the purpose, structure, features, and effects of the present invention, the present invention will be further described below in conjunction with the drawings and specific embodiments. It should be noted that the described embodiments are part of the embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0041] The present disclosure provides a multi-factor carbon emission accounting method based on geographically weighted regression and ensemble learning. Through the research method of "independent modeling - ensemble learning", the carbon emissions are decomposed into multiple driving factors, and the final multi-factor linear weighted expression is obtained through ensemble learning to quantify the contribution degree of each driving factor to carbon emissions. First, multiple predefined carbon emission driving factors are independently modeled in a spatially geographically weighted manner, and the optimal bandwidth is selected according to the driving factors to alleviate the multicollinearity interference in multi-variable modeling, and at the same time improve the ability of the regression model to reflect the correlation between the research variables and carbon emissions; secondly, the GWR model reflects the spatial heterogeneity of the carbon emission driving factors by introducing a spatial weight matrix, and makes the regression process closer to the real situation through local fitting; finally, ensemble learning is performed on the base models of each driving factor to obtain the multi-factor linear weighted expression of carbon emissions. Compared with the general carbon emission factor decomposition research method that needs to construct a rigorous identity at the beginning stage, it gets rid of the limitation of maintaining a strong logical relationship between multiple driving factors.

[0042] The method of the present disclosure first predefines n driving factors affecting carbon emissions according to the research content and direction as research variables, such as regional per capita GDP, regional gross domestic product, regional proportion of the tertiary industry, etc. There is no need to emphasize the internal logical relationship between the research variables, that is, there is no need to construct an identity between carbon emissions and research variables at the beginning stage of the research.

[0043] A data set is provided, and the data set is divided into a training set and a test set according to a certain ratio. The division ratio usually adopts 8:2 or 7:3. The training set is further evenly divided into k parts, usually making k = 5. For each base model in the ensemble learning, the divided training set and test set will be used as the data source for its k-fold cross-validation training process.

[0044] Traverse the predefined n research variables, and separately construct a GWR model for each research variable. Using the current research variable as the independent variable and carbon emissions as the dependent variable, select the optimal bandwidth for local fitting according to the research variable, and perform k-fold cross-validation training.

[0045] The k-fold cross-validation process for the GWR model of each research variable is as follows: Take 1 subset from the k evenly divided training sets as the validation set, and the remaining k - 1 subsets as the current training set, and train the model. After training, make a prediction on the validation set first, and then make a prediction on the test set. Cycle k rounds in sequence.

[0046] After the GWR model of the current research variable is trained, it serves as a base model in the ensemble learning. At this time, the predicted values of the k validation sets and the predicted values of the k test sets are obtained through k-fold cross-validation.

[0047] Vertically combine the prediction values of the k validation sets of the current base model as a feature dimension in the ensemble learning training set, and the number of samples is the same as that of the original training set. Take the mean of the prediction values of the k test sets of the current base model as a feature dimension in the ensemble learning test set, and the number of samples is the same as that of the original test set.

[0048] After all the base models of the n carbon emission driving factors are trained, the entire training set and test set of the ensemble learning have been obtained at this time, and the number of samples is the same as that of the original training set and the original test set; the number of feature dimensions in the dataset is the same as the number of predefined driving factors.

[0049] Construct an LR model on the ensemble learning training set as the combined model to obtain the linear weighted expression of the outputs of all base models, and verify the combined model on the ensemble learning test set. If necessary, other linear regression models can also be selected as the combined model, such as decision trees.

[0050] Take the coefficients of the independent variables in the linear weighted expression, and the weights of each independent variable as the quantified values of the contributions of the corresponding carbon emission driving factors to carbon emissions.

[0051] The present disclosure also provides a multi-factor carbon emission accounting device based on geographically weighted regression and ensemble learning for performing the above method. The device can be a computer. Through this device, in the case of providing multiple driving factors affecting carbon emissions, the quantified values of the contributions of each driving factor to the total carbon emissions can finally be obtained.

[0052] The above detailed description is only for the description of the preferred embodiments of the present invention, and does not limit the patent scope of the present invention. Therefore, all equivalent technical changes made by using the content of this creation are included in the patent scope of this creation.

Claims

1. A multi-factor carbon emission accounting method, characterized in that, Including: Selecting multiple driving factors that affect carbon emissions; Individually constructing a geographically weighted regression model for each driving factor among the multiple driving factors as the base model for ensemble learning; in the modeling process, taking the driving factors as research variables, traversing the custom research variables, using the current research variable as the independent variable, and the total carbon emissions as the dependent variable; selecting the best local fitting bandwidth according to each driving factor, and performing regression on carbon emissions and the current driving factor in different geographical spaces; Performing k-fold cross-validation on each base model to generate one-dimensional features, and combining them into the training dataset and test dataset required for the ensemble learning stage; Constructing a logistic regression model for the base models of multiple driving factors as the combined model for ensemble learning, and using the ensemble learning training dataset and test dataset combined by the generated features of all base models for training and validation; Combining After the model training is completed, obtaining the linear weighted expressions of all base models; According to the linear weighted expressions of all base models, taking the coefficient of the independent variable in the linear weighted expression as the quantification value of the contribution degree of the corresponding carbon emission driving factor to the total carbon emissions, so as to quantify the contribution degree of each driving factor to the carbon emissions.

2. The method according to claim 1, characterized in that, Performing k-fold cross-validation includes: Taking 1 subset from the training set evenly divided into k parts as the validation set, and the remaining k - 1 subsets as the training set for this round, and training the model; After the training is completed, making a prediction on the validation set first, and then making a prediction on the test set, and cycling k rounds in sequence; At this time, after k-fold cross-validation, obtaining the predicted values of k validation sets and the predicted values of k test sets, vertically combining the k predicted values of the validation sets of the current base model as a feature dimension in the ensemble learning training set, and the number of samples is the same as the original training set; Taking the average of the k predicted values of the test sets of the current base model as a feature dimension in the test set of the ensemble learning, and the number of samples is the same as the original test set, and the k-fold cross-validation process ends.

3. The method according to claim 1, characterized in that, The number of independent variables for constructing the logistic regression model is the same as the number of research variables; each feature dimension in the dataset is generated by the base model representing different driving factors, and the prediction object is the real total carbon emissions, then making the independent variables of the logistic regression model correspond one by one with the driving factors, finally learning the linear weighted expressions of all base models, and taking the coefficient of the independent variable in the expression as the weight of the research variable in the total carbon emissions, so as to quantify the influence degree of each driving factor on the total carbon emissions.

4. A multi-factor carbon emission accounting device for performing the method according to any one of claims 1-3.

5. A method for generating a quantification value of the contribution degree of carbon emission driving factors to carbon emissions, characterized in that, Including: Providing the device described in claim 4; Providing multiple driving factors that affect carbon emissions; The device generates the quantification value of the contribution degree of each driving factor to the total carbon emissions according to the method described in any one of claims 1 - 3.