Building carbon emission evaluation method based on improved LightGBM and carbon emission efficiency ratio

By improving the LightGBM model and carbon emission efficiency ratio method, the problems of ignoring inherent attributes and having a single assessment dimension in building carbon emission assessment are solved. Multi-dimensional dynamic modeling and transparent identification of driving factors are achieved, providing a fair and credible carbon emission assessment and supporting precise energy-saving retrofitting.

CN121920884APending Publication Date: 2026-04-24GUANGZHOU YUANZHENG INTELLIGENCE TECH
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU YUANZHENG INTELLIGENCE TECH
Filing Date
2025-12-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing building carbon emission assessment methods fail to fully consider inherent attributes such as building size, function type, and number of energy users. The assessment dimensions are too limited, lacking the quantification of multi-dimensional influencing factors. They cannot dynamically predict reasonable emission levels, and traditional black-box models are insufficient to provide precise directions for energy-saving renovations.

Method used

By adopting an improved LightGBM model combined with carbon emission efficiency ratio, and by acquiring energy consumption data and influencing factors, an influencing factor sequence is constructed. The SHAP value model is used for interpretability analysis to identify key features, calculate carbon emission efficiency ratio, and generate a carbon emission efficiency rating table, thereby achieving multi-dimensional dynamic modeling and transparency of key driving factors.

Benefits of technology

It achieves accurate time-series prediction of building carbon emissions, significantly suppresses overfitting, provides a fair evaluation benchmark, makes key driving factors transparent, ensures the fairness and credibility of evaluation results, and supports precise energy-saving retrofitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920884A_ABST
    Figure CN121920884A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of carbon emission evaluation, in particular to a building carbon emission evaluation method based on an improved LightGBM and carbon emission efficiency ratio. The method comprises the following steps: acquiring energy consumption data and related influence factor data of different types of public buildings in a target area in an operation stage; carbon accounting is carried out based on the energy consumption data, and the annual total carbon emission amount of each building is obtained; constructing an influence factor sequence based on the related influence factor data; dividing the influence factor sequence and the corresponding annual total carbon emission into a training set and a test set based on a time sequence; and training the improved LightGBM model by using the training set to obtain a carbon emission prediction model. According to the method, accurate prediction and interpretable efficiency rating of the operation carbon emission of the public building under the multi-dimensional heterogeneous condition are achieved, the fairness defect of traditional unit area sorting is thoroughly eliminated, and a targeted energy-saving transformation basis is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon emission assessment technology, and in particular to a building carbon emission assessment method based on an improved LightGBM and carbon emission efficiency ratio. Background Technology

[0002] Currently, regional building carbon emission assessments still generally use a simple ranking or quota method based on carbon emissions per unit area. The core drawbacks of this approach are twofold: firstly, it fails to consider inherent attributes such as building size, function type, and number of energy users, directly judging quality based on absolute values, which is clearly unfair; secondly, the assessment dimensions are singular, relying primarily on building area or total energy consumption, lacking a systematic quantification of multi-dimensional influencing factors such as operational characteristics and environmental factors, leading to one-sided evaluation results; furthermore, existing methods are mostly static historical statistics, unable to dynamically predict reasonable emission levels based on given conditions, making it difficult to scientifically measure the true management efficiency of buildings; and finally, traditional black-box models fail to adequately explain key driving factors, failing to provide a clear direction for subsequent precise energy-saving renovations. Summary of the Invention

[0003] Therefore, it is necessary for the present invention to provide a building carbon emission assessment method based on an improved LightGBM and carbon emission efficiency ratio to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, a building carbon emission assessment method based on an improved LightGBM and carbon emission efficiency ratio includes the following steps: Step S1: Obtain energy consumption data and related influencing factor data of different types of public buildings in the target area during the operation phase; perform carbon accounting based on energy consumption data to obtain the total annual carbon emissions of each building; Step S2: Construct an impact factor sequence based on relevant influencing factor data; Step S3: Divide the influencing factor sequence and its corresponding annual total carbon emissions into training set and test set based on time order; use the training set to train and improve the LightGBM model to obtain the carbon emission prediction model; Step S4: Based on the preset SHAP value model, perform interpretability analysis on the carbon emission prediction model to identify key features affecting carbon emission prediction; Step S5: Calculate the expected carbon emissions of each building under key characteristics using the carbon emission prediction model, and calculate the carbon emission efficiency ratio of each building; Step S6: Generate a carbon emission efficiency rating table based on the carbon emission efficiency ratio; and assess the operational carbon emission efficiency of each building in the region according to the carbon emission efficiency rating table.

[0005] The beneficial effects of this invention are as follows: On the one hand, by using multi-dimensional dynamic modeling and improving the three structural optimizations of LightGBM, inherent and dynamic attributes such as building scale, functional type, number of energy users, operating environment interference, and regional electricity carbon emission factors are incorporated into a unified feature space, which significantly suppresses overfitting and achieves accurate time-series prediction of reasonable emission levels. This fundamentally solves the fairness defects caused by the neglect of building heterogeneity in traditional static ranking. Specifically, the adaptive learning rate, path smoothing regularization, and adaptive gain threshold work together to ensure that the model maintains stable extrapolation even in scenarios with limited sample size or long time spans, ensuring that a fair benchmark for the same conditions and emissions can be quantified.

[0006] On the other hand, by introducing SHAP interpretability analysis, the marginal contribution of each feature to carbon emission prediction is quantified, opening up the original black box and making key driving factors transparent; it can be seen intuitively that a building's expected emissions increase due to excessive per capita area or high regional power grid carbon factor.

[0007] On the other hand, by redefining the evaluation scale with the carbon emission efficiency ratio of actual emissions to expected emissions, and combining quantile transformation and statistical validity testing, a rating system that takes into account fairness, comparability and credibility is constructed. Regardless of the size of the building or the difference in function, an efficiency ratio greater than 1 clearly exposes management or technical shortcomings, while an efficiency ratio less than 1 confirms excellent practices. The significance of the Pearson correlation coefficient is then verified to ensure that the rating results are statistically consistent with the traditional unit area ranking. Attached Figure Description

[0008] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings: Figure 1 A schematic flowchart of the steps of a building carbon emission assessment method based on an improved LightGBM and carbon emission efficiency ratio is shown in one embodiment.

[0009] Figure 2 A detailed flowchart of step S11 of one embodiment is shown.

[0010] Figure 3 A graph showing the adaptive minimum gain threshold variation of one embodiment is provided. Detailed Implementation

[0011] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0012] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0013] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0014] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a building carbon emission assessment method based on an improved LightGBM and carbon emission efficiency ratio, comprising the following steps: Step S1: Obtain energy consumption data and related influencing factor data of different types of public buildings in the target area during the operation phase; perform carbon accounting based on energy consumption data to obtain the total annual carbon emissions of each building; Step S2: Construct an impact factor sequence based on relevant influencing factor data; Step S3: Divide the influencing factor sequence and its corresponding annual total carbon emissions into training set and test set based on time order; use the training set to train and improve the LightGBM model to obtain the carbon emission prediction model; Step S4: Based on the preset SHAP value model, perform interpretability analysis on the carbon emission prediction model to identify key features affecting carbon emission prediction; Step S5: Calculate the expected carbon emissions of each building under key characteristics using the carbon emission prediction model, and calculate the carbon emission efficiency ratio of each building; Step S6: Generate a carbon emission efficiency rating table based on the carbon emission efficiency ratio; and assess the operational carbon emission efficiency of each building in the region according to the carbon emission efficiency rating table.

[0015] Preferably, in step S1, after obtaining the total annual carbon emissions of each building, the method further includes: Step S111: Record the energy consumption data and related influencing factor data as the initial dataset; Step S112: Remove the sample records with missing values ​​from the initial dataset to construct an intermediate dataset; In some embodiments, the initial dataset refers to a collection of original samples collected within the target area, containing energy consumption data and related influencing factor data of various types of public buildings during their operation. This dataset may contain missing values ​​in certain key feature fields in some sample records due to sensor malfunctions, omissions in manual data entry, or system communication interruptions.

[0016] Specifically, each sample record is iterated through to identify whether it has missing values ​​in any key feature field. If any sample record has a missing field, the entire record is removed.

[0017] Step S113: For the numerical feature fields in the intermediate dataset, calculate the first quartile and the third quartile of the numerical distribution of each feature field, and determine its upper and lower bounds based on the first quartile and the third quartile. Specifically, for a numerical feature field (such as annual carbon emissions or energy consumption per unit area), the values ​​of this field in all samples are sorted, and its first quartile (Q1) and third quartile (Q3) are calculated. Further, based on the interquartile range (IQR = Q3 − Q1), the reasonable numerical range boundaries of this field are determined: lower bound = Q1 − 1.5 × IQR; upper bound = Q3 + 1.5 × IQR; this range is considered the statistically normal fluctuation range of this feature field. Sample values ​​exceeding this range are initially marked as potential outliers.

[0018] Step S114: Identify sample values ​​in the intermediate dataset that exceed the range defined by the upper and lower bounds as outliers, and replace and fill out the outliers using the median of the feature field in similar building samples.

[0019] In some embodiments, in order to improve data quality and avoid outliers from causing bias in model training, the identified outliers are corrected rather than deleted directly.

[0020] Specifically, for a value marked as abnormal (such as an abnormally high annual gas consumption of an office building), the functional type of the building (such as an office building) is determined, and the median of the feature field is found in the subset of functional types. The original abnormal value is then replaced with the median.

[0021] Preferably, after constructing the impact factor sequence, step S2 further includes feature engineering of the impact factor sequence, wherein the feature engineering includes: Construct interactive features, which include at least the per capita area, the product of the operating environment disturbance index and the building area, the product of the operating environment disturbance index and the number of energy users, and the product of the regional electricity carbon emission factor and the building area. In some embodiments, three additional product-based interactive features are calculated and added within the same building sample: per capita area × operating environment disturbance index × building area, which simultaneously characterizes the superposition effect of spatial congestion and external disturbance intensity on the building scale dimension. The higher the value, the more the building faces the dual pressure of per capita space shortage and climate fluctuations despite its large size; operating environment disturbance index × number of energy users, which reflects the synergistic impact of external disturbances and internal population density. When a large transportation hub with more than 10,000 energy users encounters extreme high temperatures (disturbance index 1.5), the product term can reach 15,000, and the model marks high-sensitivity samples accordingly; regional electricity carbon emission factor × building area, which directly links the cleanliness of the power grid with the building scale.

[0022] Construct lagging characteristics, which include at least the operating environment disturbance index of the previous year, the operating environment disturbance index of the previous two years, the regional power carbon emission factor of the previous year, and the regional power carbon emission factor of the previous two years.

[0023] In some embodiments, four types of lag features are further extracted for each sample and directly added to the modeling dataset. These include: the operating environment disturbance index of the previous year, used to characterize the delayed impact of external disturbances (such as extreme temperatures and equipment start-up and shutdown fluctuations) in the previous year on the energy consumption inertia of the current year; the operating environment disturbance index of the previous two years, used to identify whether the building has triggered equipment updates or management strategy adjustments under high disturbance scenarios for two consecutive years; the regional power carbon emission factor of the previous year, reflecting the delayed manifestation of indirect emission benefits or pressures brought about by the transformation of the regional power grid structure (such as the increase in the proportion of new energy grid connection) in the current year; and the regional power carbon emission factor of the previous two years, used to capture the continuous trend of regional decarbonization or recarbonization, and help the model distinguish between trend-based emission reduction and passive emission rebound.

[0024] Preferably, in step S3, the improved LightGBM model is constructed by integrating the standard LightGBM model with three structural improvements. The three structural improvements jointly determine their hyperparameters using a Bayesian optimization framework and time-series cross-validation. The first structural improvement among the three improvements is adaptive learning rate adjustment, including: Set the initial learning rate at the beginning of training. ; During the training process, the first The training rounds decay according to an exponential decay strategy, where the specific calculation formula for the exponential decay strategy is as follows: ; in, For the first Actual learning rate per round As the attenuation factor, The attenuation frequency; Attenuation factor Along with the decay frequency k, it is used as a key hyperparameter and is automatically optimized together with other model parameters in the Bayesian optimization framework.

[0025] In this embodiment of the invention, a uniform initial learning rate is set for all decision trees before training begins. (For example, 0.1). This value ensures that the first round of gradient boosting has a sufficiently large step size to quickly cover the main descent direction of the loss function.

[0026] As the iteration progresses, the learning rate is dynamically reduced according to the round t. ∈(0,1) is the decay factor, which controls the reduction factor per k steps; k is the decay frequency, which means that a proportional reduction is triggered once every k rounds.

[0027] For example, take =0.1、 =0.8, k=50, then in round 0 =0.1, Round 50 =0.1× =0.08, Round 100 =0.1× =0.064, achieving a smooth transition from fast to slow.

[0028] To avoid subjective bias caused by manual parameter adjustment, Along with k, other LightGBM hyperparameters such as max_depth, num_leaves, and min_child_samples are incorporated into the Bayesian optimization framework; the objective function is set as the mean negative mean squared error (−MSE) under time-series cross-validation. The optimizer uses a TPE or GP surrogate model in a predefined search space (e.g., ∈[0.6,0.95], Automatic exploration within [30, 200] will converge to the optimal value after ≤100 iterations. , )combination.

[0029] Preferably, the second structural improvement among the three structural improvements includes: During model training, feature selection and node splitting are performed iteratively to construct a model decision tree set consisting of multiple decision trees; Most importantly, the improved LightGBM employs a gradient boosting mechanism for iterative training during the training phase: In each training round, a new decision tree is constructed based on the prediction residuals of the current model; For the construction of each decision tree, the following sub-steps are executed sequentially: Step S021: For the current node to be split, traverse all features and calculate the information gain corresponding to all candidate split points for each feature; Step S022: Select the feature with the largest information gain and its splitting threshold from all candidate splitting points as the splitting basis for the current node, and divide the node sample into left child node and right child node according to the splitting threshold; Step S023: Repeat steps S021 to S022 for the generated left and right child nodes respectively until a preset stopping condition is met; wherein, the stopping condition includes at least one of the following: reaching the maximum depth of the tree, or the number of samples in the node is lower than the minimum number of samples threshold; Step S024: Determine the nodes that meet the stopping conditions as leaf nodes, and calculate the output value of the leaf node based on the mean or median of the residuals of the samples falling into the leaf node. Step S025: As the number of training rounds increases, the decision trees generated in each round are sequentially superimposed on the model to form a decision tree ensemble model composed of multiple decision trees.

[0030] For example, in the task of predicting carbon emissions from public buildings, the boosting rounds are set to 800 rounds. The first tree splits with the per capita area as the root node, the 50th tree mainly selects the regional electricity carbon emission factor, and the 200th tree frequently uses the previous year's operating environment disturbance index as the splitting feature, finally obtaining a decision tree set consisting of 800 trees.

[0031] The splitting feature selection frequency and splitting threshold variance of all nodes are calculated based on the model decision tree set. Of particular importance is that after the improved LightGBM completes all boosting rounds, and the model training is finished, centralized post-processing is performed on all decision trees in the decision tree ensemble. The post-processing includes the following steps: S031: Traverse each decision tree in the decision tree ensemble, record the splitting feature name and splitting threshold used by each non-leaf node, and form a feature-threshold raw data list; S032: Based on the feature-threshold raw data list, count the cumulative number of times each feature is selected as a splitting feature, calculate its feature selection frequency, and identify the key feature set and secondary feature set accordingly. S033: Based on the feature-threshold raw data list, calculate the distribution variance of the splitting threshold corresponding to each feature, and identify robust and sensitive features for splitting decisions based on the variance; S034: Calculate the path divergence coefficient for each feature based on the feature selection frequency and threshold distribution variance, and generate a path smoothing regularization term after aggregation. S035: Generate an optimization report: Based on the analysis results of steps S032 to S034, generate a structured report that includes a list of key features, warnings for sensitive features, stability scores, and suggestions for hyperparameter adjustment.

[0032] For example, in a typical scenario with 800 trees and a total of 12,600 non-leaf nodes, the frequency of splitting feature selection was statistically analyzed. Specifically, the probability of each feature being selected as a splitting condition was obtained by summing the occurrences of each feature dimension and dividing by the total number of nodes. The results showed that the operating environment interference index appeared 3,024 times, with a selection frequency of 24%, ranking first; while the regional electricity carbon emission factor from the previous year appeared 1,512 times, with a frequency of 12%, ranking third.

[0033] Furthermore, the splitting threshold variance is calculated on the same feature-threshold list: the mean and variance of all occurrence thresholds for the same feature are calculated. The larger the variance, the more sensitive the model is to the splitting position of that feature, and the more easily it drifts with sample perturbation.

[0034] For example, the threshold variance of the building area feature is as high as 1842. It is far higher than the variance of per capita area of ​​95. The data suggests that the building area split point fluctuates wildly across different subsamples, indicating a risk of overfitting; while the regional electricity carbon emission factor variance is only 0.008. This indicates that its threshold is relatively robust.

[0035] Based on the splitting characteristics, select the frequency and splitting threshold variance to determine the path smoothing regularization term. ; In some embodiments, after calculating the frequency of splitting feature selection and the variance of splitting threshold, it is necessary to further integrate the two into a single index to quantify the instability of each splitting path. To this end, a path smoothing regularization term is introduced. The logic behind its construction is that paths with high frequency and large threshold drift are most sensitive to sample perturbations and should be subject to greater penalties.

[0036] Specifically, for each feature f, first select its frequency. With threshold variance Normalized to the [0,1] interval, and further weighted product form is used: This design ensures that the filter is applied only when the feature is frequently selected and the threshold fluctuates drastically. It makes a significant contribution and avoids bias from a single indicator.

[0037] For example, in the statistical results of 800 trees and 12,600 non-leaf nodes, the building area characteristic... =18%, =1842 After normalization, the corresponding component is 0.18 × 0.76 = 0.137; while the operating environment interference index, although... Up to 24%, but With a value of only 0.015, the normalized component is only 0.24 × 0.01 = 0.002.

[0038] Ultimately, this The value (0.28 in the example scenario) serves as the quantifiable penalty base for subsequent regularization loss, and is the path smoothing coefficient. The introduction of provides direct input.

[0039] Path smoothing regularization term Introducing the original loss function of the standard LightGBM model Construct the smoothed loss function ,in, This is the path smoothing coefficient; In some embodiments, upon obtaining Then, it is injected into the original loss function of standard LightGBM in a weighted form to form the smoothed loss: This is used to control the penalty intensity for unstable split paths. During the training phase, gradient backpropagation simultaneously optimizes two objectives: reducing residuals and suppressing high-frequency, high-drift paths. This automatically reduces the model's sensitivity to sample perturbations without pruning or reducing the number of trees.

[0040] For example, taking a public building carbon emission prediction task as an example, the current batch residual loss =0.18 (MSE), calculated in the previous step. =0.28, let =0.01, then: =0.18 + 0.01 × 0.28 = 0.1828; Although the absolute value only increases by 0.0028, it generates an additional value during gradient calculation. This forced the splitting algorithm to compromise between information gain and path stability. After 50 more training rounds, It dropped to 0.12, while The variance decreased to 0.15, indicating that while maintaining the decrease in residuals, the model successfully compressed the building area threshold variance by 38%, thus achieving explicit suppression of the overfitting path.

[0041] The path smoothing coefficient λ is incorporated as a key hyperparameter into the Bayesian optimization framework and jointly optimized with the key hyperparameter of adaptive learning rate adjustment and the minimum gain split control threshold to obtain a hyperparameter combination; wherein, the objective function of the joint optimization is the average negative mean square error obtained by time-sequential cross-validation on the training set. In some embodiments, Compared with the first improvement k and the minimum gain threshold in the third improvement They are collectively defined as a four-dimensional hyperparameter vector and incorporated into the same Bayesian optimization search space: ∈[1× 5× (Logarithmically uniform); k ∈ [0.60, 0.95] (uniform); k ∈ [30, 200] (integer); ∈[0.01,0.20] (uniform); the optimizer uses a TPE kernel, with 30 prior random samples and subsequent 70 iterative samples collected using the EI sampling function; the objective function is fixed as the mean negative mean square error (−MSE) of time-series cross-validation to ensure the generalization ability of the hyperparameters under the real time-series distribution.

[0042] For example, panel data from 1200 public buildings between 2015 and 2022 was divided into 80% segments by year and verified sequentially. Candidate combinations emerged in the 42nd iteration: =0.012、 =0.77, k=60, =0.08, and its −MSE on the training set is −0.118, which is better than the previous best −0.124. After optimization convergence, the final hyperparameter combination is locked: =0.012、 =0.77、 =60、 =0.08. Using this combination to retrain 800 trees, the validation set RMSE decreased by 9.3% compared to the standard LightGBM. It improves by 4.1%, and the variance of the splitting threshold of building area features decreases by 38%, achieving optimal simultaneous minimization of residuals and path stability.

[0043] Improve the LightGBM model by training and ensemble using hyperparameter combinations.

[0044] In some embodiments, after Bayesian optimization converges, the joint optimal hyperparameter combination is obtained: =0.012、 =0.77、 =60、 =0.08, and the previously determined initial learning rate. =0.10, number of trees 800, and other fixed parameters. Furthermore, this combination is written into the improved LightGBM configuration dictionary all at once, initiating the complete training process: data streams are fed in chronological order, and each iteration first... Dynamically scale the learning rate; when splitting a node, only if the information gain is greater than the adaptive learning rate. Furthermore, splitting is only performed when the path smoothing regularization gradient constraint is satisfied; calculation is performed immediately after each tree is completed. The increments are then added back to the leaf weights to suppress unstable paths while building the tree.

[0045] Preferably, the third structural improvement among the three structural improvements includes: During the decision tree node splitting process, the information gain corresponding to each candidate split point is calculated; In some embodiments, the sample subset of the current node to be split is sequentially scanned to determine the optimal splitting method. Specifically, for each numerical feature, it is first sorted in ascending order, and then the midpoint between adjacent values ​​is taken as a candidate splitting threshold; for categorical features, all possible binary splitting combinations are directly enumerated. Furthermore, for each candidate threshold, the reduction in mean squared error (MSE) before and after splitting is calculated, and this reduction is defined as the information gain of that splitting point. The larger the information gain, the more significant the decrease in residual after splitting according to this threshold, and the more obvious the improvement in model fitting effect.

[0046] For example, within a public building node, the sample set contains 128 annual energy consumption records. After sorting the features based on the previous year's operating environment disturbance index, 12 candidate split points are obtained. Among them, when the threshold is set to 1.25, the MSE of the left child node is 0.014t. The MSE of the right child node is 0.018t. After weighting, the overall MSE decreased by 0.018t compared to before the split. Therefore, the information gain of this point is 0.018, the highest among all candidate points, and it will be prioritized as the basis for splitting the current node. If this gain subsequently exceeds the adaptive minimum gain threshold, the splitting operation will be formally executed; otherwise, the node will be marked as a leaf node, and further partitioning will cease.

[0047] The minimum gain threshold is set as a key hyperparameter, where the minimum gain threshold is adaptively adjusted according to the number of samples in the current node and the tree depth; In some embodiments, instead of using a fixed constant as the splitting threshold, a minimum gain threshold is used. Designed as a bivariate function of sample size N + tree depth d. The specific function form is: ; in, The base gain is obtained by Bayesian optimization within the range of [0.01, 0.20]. , The sensitivity coefficient controls the sample sparsity and depth penalty weights, respectively. 1% of the total number of samples in the training set is used to dimensionless the exponential term.

[0048] For example, take the optimal hyperparameter =0.05、 =0.15、 =0.01, =120 (Total sample size: 12,000): Root node scenario: N=12000, d=0; =0.05×(1+0.15·e^(−100)+0)≈0.050; the threshold is close to the base value, allowing for sufficient splitting.

[0049] Intermediate node scenario: N=320, d=5; =0.05×(1+0.15·e^(−2.67)+0.05)≈0.056; The threshold is slightly raised, and a slight suppression is applied to overly deep paths.

[0050] Deep small sample size scenario: N=32, d=8; =0.05×(1+0.15·e^(−0.27)+0.08)≈0.068; the threshold increases by 36%, if the information gain is less than 0.068t If the node is not found, it is directly marked as a leaf node, effectively blocking overfitting branches.

[0051] The adjustment parameters for the minimum gain threshold are determined by combining a Bayesian optimization framework with time series cross-validation. In some embodiments, an adaptive function for minimizing the gain threshold is used. To truly align with the temporal distribution of building carbon emissions, , , The three-dimensional vector is used as a key hyperparameter, along with the learning rate decay factor in the first improvement. The attenuation frequency k and the path smoothing coefficient in the second improvement Incorporate them into the same Bayesian optimization search space: ∈[0.01,0.20] (continuous and uniform); ∈[0.05,0.30] (continuous and uniform); ∈[0.005,0.020] (continuous and uniform); the optimizer uses a TPE kernel, with 30 prior random samples and subsequent 70 iterative samples collected using the EI sampling function; the objective function is set to the mean negative mean square error (−MSE) of time-series cross-validation to ensure adaptiveness. Generalization ability on real-year sliding windows.

[0052] The information gain of each candidate split point is compared with the minimum gain threshold corresponding to the current node. The split operation of the node is only performed if the information gain of the candidate split point is greater than the minimum gain threshold. If the information gain of all candidate split points is not greater than the minimum gain threshold, the node is marked as a leaf node and further splitting is stopped.

[0053] In some embodiments, the adaptive function is invoked immediately after the information gain calculation of the candidate split points is completed. The minimum gain threshold for the current node is generated and compared with the maximum information gain. The comparison result determines whether the node continues to split or becomes a leaf node, thus achieving dynamic early stopping.

[0054] For example, the optimal parameters after Bayesian optimization are taken. =0.08、 =0.15、 =0.009: Large sample shallow nodes: N=1200, d=2 → ≈0.080; Maximum information gain = 0.118> The split was approved, and the nodes continued to be divided according to the threshold of 0.62.

[0055] Medium-sized sample mid-level nodes: N=160, d=5 → ≈0.067; Maximum information gain = 0.069> (Exceeding 3%), splits are still allowed, but the threshold has been raised by 11%.

[0056] Small sample deep nodes: N=24, d=9 → ≈0.071; Maximum information gain = 0.058 If none of the candidate points meet the criteria, the split is terminated immediately, the node is marked as a leaf node, and the output weight is the median of the sample residuals minus 0.09t. .

[0057] In one embodiment, see Figure 3 The horizontal axis uses a logarithmic scale to represent the number of samples N, and the vertical axis represents the minimum gain threshold. The graph shows four curves, corresponding to tree depths d=0, 3, 5, and 8, distinguished by different line types. The minimum gain threshold is determined by the formula... Perform adaptive adjustment, where =0.05、 =0.15、 =0.01、 =120. As can be seen from the figure, at the same tree depth, the minimum gain threshold gradually decreases and tends to stabilize as the number of samples increases; for the same number of samples, the greater the tree depth, the higher the minimum gain threshold. The figure marks three typical scenarios: the root node scenario (N=12000, d=0). ≈0.050, in the intermediate node scenario (N=320, d=5) ≈0.056, for deep small sample scenarios (N=32, d=8) The threshold value was approximately 0.068, which increased by 36%, effectively preventing the generation of overfitting branches.

[0058] Preferably, the formula for calculating the carbon emission efficiency ratio is: ; in, Let be the carbon emission efficiency ratio of the i-th building. Its actual annual carbon emissions, The carbon emission prediction model is based on the building's feature vector. The projected expected carbon emissions.

[0059] Preferably, generating a carbon emission efficiency rating table based on the carbon emission efficiency ratio includes: A standard score set is constructed by performing quantile transformation on the carbon emission efficiency ratios of all buildings. In some embodiments, to eliminate the influence of the right-skewed distribution of carbon emission efficiency ratio (CEER) on subsequent rating classification, after obtaining the CEER vectors of all 1200 buildings, a quantile transformation is performed on them to map them to standard normal quantiles, forming a standard rating set.

[0060] The specific steps are as follows: Sort the CEER values ​​from smallest to largest to obtain the order statistic. Calculate each Corresponding cumulative probability Through the inverse standard normal cumulative distribution function (This can be done by consulting a table or approximating the value), Transform into standard normal quantiles : = ; thus generated The standard score follows N(0,1) and its range falls within the interval [−3,3].

[0061] For example, a certain convention center has a CEER of 1.42, which corresponds to the following after sorting. =0.916, from the table we get =1.38; while another ultra-low emission office building has a CEER of 0.73, corresponding to =0.048, =−1.66. Through this transformation, the original right-skewed distribution is stretched into a symmetrical bell shape.

[0062] All buildings corresponding Composition vector Z={ , ,..., } is the standard rating set.

[0063] The lower quartile, median, and upper quartile of the transformed distribution in the standard score set were selected as the core statistical benchmarks. In some embodiments, for the standard score set Z={ that has undergone quantile transformation , ,..., Quartile extraction is performed on n=1200 to determine the core statistical benchmark for subsequent grading.

[0064] The specific steps are as follows: Arrange them in ascending order to obtain the order statistic. ≤ ≤…≤ Lower quartiles Position = 0.25 × (n + 1) = 300.25 → linear interpolation =−0.68; median Position = 0.5 × (n + 1) = 600.5 → Interpolation value =0.02; Upper quartile Position = 0.75 × (n + 1) = 900.75 → Interpolation value =0.71; put { =−0.68, =0.02, =0.71} is denoted as the core statistical benchmark.

[0065] The theoretical average efficiency line with a carbon emission efficiency ratio of 1 is used as the key baseline, and a carbon emission efficiency rating table is constructed based on the core statistical benchmark and the key baseline.

[0066] In some embodiments, the lower quartile, median, and upper quartile of the standard score set are obtained. =−0.68、 =0.02、 After that (when CEER = 0.71), the theoretical average efficiency line with a carbon emission efficiency ratio CEER = 1 is further introduced as a fairness benchmark. This line indicates that the actual emissions are exactly equal to the emissions expected by the model, corresponding to a standard score of 0, which is the critical value for determining whether the building operation efficiency meets the standard. Mapping this key benchmark line back to the original CEER scale together with the above three statistical benchmark points forms five-level rating cut-off points: CEER ≤ 0.83 (corresponding to z ≤ ) → Grade A (efficient); 0.83 < CEER ≤ 0.98 ( < z ≤ 0) → Grade B; 0.98 < CEER ≤ 1.02 (0 < z ≤ ) → Grade C (meeting the standard); 1.02 < CEER ≤ 1.20 ( < z ≤ ) → Grade D; [[ID=X]] CEER > 1.20 (z > ) → Grade E (high consumption); Preferably, after step S6, it further includes a step of statistically validating the rating results, specifically including: Based on the total annual carbon emissions and building area of each building, calculate the absolute value of carbon emissions per unit building area of each building, and sort all buildings according to the absolute value of carbon emissions per unit building area to determine the carbon emission efficiency grade sequence; In some embodiments, perform a standardized calculation of total carbon emissions ÷ building area for each building to obtain the absolute value of carbon emissions per unit building area (t / ), and further sort the full sample from small to large according to this index to form the carbon emission efficiency grade sequence in the traditional sense.

[0067] The specific steps are as follows: For building , the formula is: [[ID=X]]: where is the total annual carbon emissions (t ), is the building area ( ). Sort the ACE values of 1200 public buildings in ascending order to obtain the order statistics { , , …, }.

[0068] Use the equal-frequency quintile method: Buildings with ACE in the top 20% are designated as Grade A′; 20% - 40% are Grade B′; 40% - 60% are Grade C′; 60% - 80% are Grade D′; the last 20% are Grade E′; Exemplarily, a convention and exhibition center has an annual carbon emission of 9600 t Building area 50,000 ACE=0.192t / After sorting, it is ranked 1056th, at the 88th percentile, and is therefore classified as E′; while the other office building has an ACE of 0.068t. / It is ranked 180th, at the 15th percentile, and is classified as Grade A′.

[0069] Map the ranking results of the absolute value of carbon emissions per unit building area to the corresponding traditional ranking levels to determine the traditional ranking level sequence; In some embodiments, after sorting the absolute carbon emissions per unit building area (ACE) in ascending order, the continuous sorting needs to be mapped to discrete traditional sorting levels. The mapping uses the equal-frequency quintile method to ensure that the sample size is the same at each level.

[0070] The specific steps are as follows: Calculate the total sample size n=1200, and determine the sample size quota for each level = n / 5=240. Number the samples sequentially from ACE from smallest to largest 1→1200, and extract the corresponding quantiles: 1–240 → A′ level (high efficiency); 241–480 → B′ level; 481–720 → C′ level; 721–960 → D′ level; 961–1200 → E′ level (high consumption). Calculate the Pearson correlation coefficient between the carbon emission efficiency ranking series and the traditional ranking series; In some embodiments, the carbon emission efficiency ranking sequence is aligned with the traditional ranking sequence by building number to form paired samples, and then the Pearson correlation coefficient formula is applied: ; in, The CEER rating score for the i-th building is numerically converted (A=5, B=4, C=3, D=2, E=1). The traditional ACE rating for the same building is numerically converted (A′=5, B′=4, C′=3, D′=2, E′=1). For example, in a panel of 1200 public buildings, the paired sequences {X} and {Y} are obtained, and the following is calculated: =2.98, =2.98; covariance Standard deviation =1.41, =1.41; Substituting into the formula, we get the Pearson correlation coefficient r = 1.74 / (1.41 × 1.41) ≈ 0.88.

[0071] The significance of the Pearson correlation coefficient was tested to obtain its corresponding probability value; In some embodiments, after obtaining a Pearson correlation coefficient r≈0.88, a t-test is further performed to determine whether it is significantly different from zero, in order to quantify the statistical reliability of the correlation between the two rank sequences. The test statistic is calculated using the following formula: Where n=1200 is the building sample size, and the degrees of freedom are... .

[0072] Example calculation: ; Further consulting the t-distribution table or using the SciPy Stats two-tailed cumulative distribution function yields: .

[0073] The probability value is compared with a preset significance level. If the probability value is less than the preset significance level, the carbon emission rating result of the public building is deemed valid.

[0074] In some embodiments, after performing the significance test of the Pearson correlation coefficient, the two-tailed p≈ is obtained. Further, this probability value is compared with a preset significance level. =0.01 for single-point comparison We rejected the null hypothesis that there was no correlation and determined that the carbon emission efficiency ranking sequence is significantly positively correlated with the traditional ranking sequence.

[0075] Preferably, the method further includes quantitative verification of the prediction performance of the improved LightGBM model, specifically including: On the test set, carbon emission predictions were performed using the improved LightGBM model and the standard LightGBM model, and the coefficient of determination and root mean square error of the prediction results of the two models were calculated respectively. In some embodiments, after training with panel data of public buildings from 2015 to 2020, 240 buildings from 2021 to 2022 were independently archived as a test set in chronological order, and used to make predictions for both the improved LightGBM and standard LightGBM models. Both models used the same 105-dimensional feature input, and the number of trees was uniformly set to 800 to ensure fair comparison. After prediction, the deviation between the actual and predicted values ​​was calculated for each building, and the coefficient of determination and root mean square error were obtained accordingly.

[0076] Exemplary results show that the standard LightGBM has a determination coefficient of 0.794 and a root mean square error of 0.124t on the test set. The improved LightGBM achieved a determination coefficient of 0.845 and a root mean square error of 0.112t. Compared to the baseline, the coefficient of determination increased by 5.1 percentage points, the root mean square error decreased by 9.7%, and the residual distribution tightened significantly, indicating that the three structural improvements effectively enhanced the model's generalization ability and prediction accuracy for carbon emissions in unknown years.

[0077] Based on the coefficient of determination and root mean square error of the prediction results from the two models, the following judgments are made: When the improvement in the coefficient of determination of the improved LightGBM model on the test set reaches a preset first threshold or more, and the reduction in the root mean square error reaches a preset second threshold or more, the structural improvement of the improved LightGBM model is deemed effective; wherein the preset first threshold is less than the preset second threshold.

[0078] In some embodiments, a determination module is invoked to make a decision on the effectiveness of the structural improvement. A preset first threshold is set ( The threshold for raising the error is set to 3%, and the second threshold for lowering the RMSE is set to 8%. The first threshold is explicitly lower than the second threshold to emphasize that error reduction takes precedence over slight increases in goodness of fit.

[0079] Using the results from the previous step: The improved model is superior to the standard model. The PMI rose from 0.794 to 0.845, an increase of approximately 6.4% (0.845 − 0.794) / 0.794; the RMSE increased from 0.124t. Reduced to 0.112t The decrease was approximately 9.7% (0.124 − 0.112) / 0.124.

[0080] Since 6.4% > 3% and 9.7% > 8%, both conditions are met simultaneously, indicating that the structural improvement is effective and allowing the improved model to enter the subsequent online rating and policy application stages; if either indicator fails to reach the corresponding threshold, the improvement is deemed ineffective, triggering a rollback or readjustment process.

[0081] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0082] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A building carbon emission assessment method based on an improved LightGBM and carbon emission efficiency ratio, characterized in that, Includes the following steps: Step S1: Obtain energy consumption data and related influencing factor data of different types of public buildings in the target area during the operation phase; Carbon accounting is performed based on energy consumption data to obtain the total annual carbon emissions of each building; Step S2: Construct an impact factor sequence based on relevant influencing factor data; Step S3: Divide the influencing factor sequence and its corresponding annual total carbon emissions into training set and test set based on time order; use the training set to train and improve the LightGBM model to obtain the carbon emission prediction model; Step S4: Based on the preset SHAP value model, perform interpretability analysis on the carbon emission prediction model to identify key features affecting carbon emission prediction; Step S5: Calculate the expected carbon emissions of each building under key characteristics using the carbon emission prediction model, and calculate the carbon emission efficiency ratio of each building; Step S6: Generate a carbon emission efficiency rating table based on the carbon emission efficiency ratio; The carbon emission efficiency rating table is used to assess the operational carbon emission efficiency of each building in the region.

2. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio as described in claim 1, characterized in that, In step S1, after obtaining the total annual carbon emissions of each building, the following steps are also included: Step S111: Record the energy consumption data and related influencing factor data as the initial dataset; Step S112: Remove the sample records with missing values ​​from the initial dataset to construct an intermediate dataset; Step S112: For the numerical feature fields in the intermediate dataset, calculate the first quartile and the third quartile of the numerical distribution of each feature field, and determine its upper and lower bounds based on the first quartile and the third quartile. Step S114: Identify sample values ​​in the intermediate dataset that exceed the range defined by the upper and lower bounds as outliers, and replace and fill out the outliers with the median of the feature field in similar building samples.

3. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio as described in claim 1, characterized in that, Step S2, after constructing the impact factor sequence, also includes feature engineering of the impact factor sequence. Feature engineering includes: Construct interactive features, which include at least the per capita area, the product of the operating environment disturbance index and the building area, the product of the operating environment disturbance index and the number of energy users, and the product of the regional electricity carbon emission factor and the building area. Construct lagging characteristics, which include at least the operating environment disturbance index of the previous year, the operating environment disturbance index of the previous two years, the regional power carbon emission factor of the previous year, and the regional power carbon emission factor of the previous two years.

4. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 1, characterized in that, In step S3, the improved LightGBM model is constructed by integrating the standard LightGBM model with three structural improvements. These three structural improvements jointly determine their hyperparameters using a Bayesian optimization framework and time-series cross-validation. The first structural improvement among the three is adaptive learning rate adjustment, which includes: Set the initial learning rate at the beginning of training. ; During the training process, the first The training rounds decay according to an exponential decay strategy, where the specific calculation formula for the exponential decay strategy is as follows: ; in, For the first Actual learning rate per round As the attenuation factor, The attenuation frequency; Attenuation factor Along with the decay frequency k, it is used as a key hyperparameter and is automatically optimized together with other model parameters in the Bayesian optimization framework.

5. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 4, characterized in that, The second of the three structural improvements includes: During model training, feature selection and node splitting are performed iteratively to construct a model decision tree set consisting of multiple decision trees; The splitting feature selection frequency and splitting threshold variance of all nodes are calculated based on the model decision tree set. Based on the splitting characteristics, select the frequency and splitting threshold variance to determine the path smoothing regularization term. ; Path smoothing regularization term Introducing the original loss function of the standard LightGBM model Construct the smoothed loss function ,in, This is the path smoothing coefficient; The path smoothing coefficient λ is incorporated as a key hyperparameter into the Bayesian optimization framework and jointly optimized with the key hyperparameter of adaptive learning rate adjustment and the minimum gain split control threshold to obtain a hyperparameter combination; wherein, the objective function of the joint optimization is the average negative mean square error obtained by time-sequential cross-validation on the training set. Improve the LightGBM model by training and ensemble using hyperparameter combinations.

6. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 4, characterized in that, The third structural improvement among the three structural improvements includes: During the decision tree node splitting process, the information gain corresponding to each candidate split point is calculated; The minimum gain threshold is set as a key hyperparameter, where the minimum gain threshold is adaptively adjusted according to the number of samples in the current node and the tree depth; The adjustment parameters for the minimum gain threshold are determined by combining a Bayesian optimization framework with time series cross-validation. The information gain of each candidate split point is compared with the minimum gain threshold corresponding to the current node. The split operation of the node is only performed if the information gain of the candidate split point is greater than the minimum gain threshold. If the information gain of all candidate split points is not greater than the minimum gain threshold, the node is marked as a leaf node and further splitting is stopped.

7. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 1, characterized in that, The formula for calculating the carbon emission efficiency ratio is: ; in, Let be the carbon emission efficiency ratio of the i-th building. Its actual annual carbon emissions, The carbon emission prediction model is based on the building's feature vector. The projected expected carbon emissions.

8. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 1, characterized in that, The carbon emission efficiency rating table generated based on the carbon emission efficiency ratio includes: A standard score set is constructed by performing quantile transformation on the carbon emission efficiency ratios of all buildings. The lower quartile, median, and upper quartile of the transformed distribution in the standard score set were selected as the core statistical benchmarks. The theoretical average efficiency line with a carbon emission efficiency ratio of 1 is used as the key baseline, and a carbon emission efficiency rating table is constructed based on the core statistical benchmark and the key baseline.

9. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 1, characterized in that, Step S6 is followed by a step to verify the statistical validity of the rating results, which specifically includes: Based on the annual total carbon emissions and building area of ​​each building, the absolute value of carbon emissions per unit building area of ​​each building is calculated, and all buildings are ranked according to the absolute value of carbon emissions per unit building area to determine the carbon emission efficiency level sequence. Map the ranking results of the absolute value of carbon emissions per unit building area to the corresponding traditional ranking levels to determine the traditional ranking level sequence; Calculate the Pearson correlation coefficient between the carbon emission efficiency ranking series and the traditional ranking series; The significance of the Pearson correlation coefficient was tested to obtain its corresponding probability value; The probability value is compared with a preset significance level. If the probability value is less than the preset significance level, the carbon emission rating result of the public building is deemed valid.

10. The building carbon emission assessment method based on the improved LightGBM and carbon emission efficiency ratio according to claim 1, characterized in that, The method also includes quantitative verification of the prediction performance of the improved LightGBM model, specifically including: On the test set, carbon emission predictions were performed using the improved LightGBM model and the standard LightGBM model, and the coefficient of determination and root mean square error of the prediction results of the two models were calculated respectively. Based on the coefficient of determination and root mean square error of the prediction results from the two models, the following judgments are made: When the improvement in the coefficient of determination of the improved LightGBM model on the test set reaches a preset first threshold or more, and the reduction in the root mean square error reaches a preset second threshold or more, the structural improvement of the improved LightGBM model is deemed effective; wherein the preset first threshold is less than the preset second threshold.

Citation Information

Cited By

  • A carbon emission management method and system for environmentally friendly cables

    CN122155758A