Multi-dimensional dynamic carbon emission factor modeling method and system
Through the multi-dimensional dynamic carbon emission factor modeling method, multi-dimensional data is collected, key feature subsets are screened, association matrix is established and main factors are extracted, gradient mask feature network is constructed, dynamic fusion model results are combined with incremental learning and transfer learning, and the bottleneck of cross-scale data fusion in the existing technology is solved, realizing the accuracy of local anomaly emission events and the reliability of global evaluation.
Patent Information
- Application Number
- CN202510829467.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing carbon emission factor modeling methods cannot accurately characterize the contribution of local anomalies in the global assessment when cross-scale data fusion, resulting in a reduced sensitivity of monitoring models to identify key risk signals, forming a bottleneck for cross-scale data fusion.
Multi-dimensional monitoring data is collected, key feature subsets are screened through statistical tests, association matrix is established, cross-dimensional main factors are extracted, gradient mask feature importance network is constructed, special sub-models are fused with incremental learning and transfer learning, and dynamic optimization evaluation results are output through Monte Carlo simulation.
It improves the analysis ability of carbon emission factor modeling to analyze dynamic features across scale data, enhances the model's ability to identify local anomaly emission events, ensures the reliability and compliance of the evaluation results, and solves the problem of feature contribution distortion caused by linear aggregation in traditional methods.
Smart Images

Figure CN120337799A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon emission monitoring. More specifically, the present invention relates to a multi-dimensional dynamic carbon emission factor modeling method and system. Background Art
[0002] When integrating data at different scales, existing carbon emission factor modeling methods usually adopt a unified aggregation strategy to process device-level monitoring data and regional-level energy data. Such methods standardize and integrate data at different levels based on preset rules, and fail to effectively distinguish the differential impacts of key local fluctuations and overall trends.
[0003] During the cross-scale data fusion process of existing methods, due to the insufficient differential response of the standardization integration mechanism to micro-dynamic characteristics and macro-steady characteristics, the modeling results of carbon emission factors cannot accurately represent the contribution degree of local abnormal emission events to global evaluation. This distortion of the influence transmission between data levels will reduce the recognition sensitivity of the monitoring model to key risk signals and form a bottleneck in cross-scale data fusion. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a multi-dimensional dynamic carbon emission factor modeling method and system to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions: A multi-dimensional dynamic carbon emission factor modeling method, comprising: S1. Collect multi-dimensional monitoring data to construct a multi-dimensional data set, and screen a key feature subset significantly related to the carbon emission factor based on statistical tests; S2. Establish an association matrix between the key feature subset and the carbon emission factor, extract cross-dimensional main factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional main factor, and generate hierarchical weight parameters; S3. Construct a feature importance network based on gradient masking based on the hierarchical weight parameters, optimize the feature distribution of the multi-dimensional data set through feature channel scaling, and generate an initial prediction model; S4. Construct dedicated sub-models for multiple preset dimensions, and adopt a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor; S5. Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, and transfer historical knowledge to the dedicated sub-model through transfer learning and filter input data noise; S6. Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor according to the historical verification data, and output a dynamically optimized evaluation result.
[0006] In a preferred embodiment, multi-dimensional monitoring data is collected to construct a multi-dimensional data set, and a key feature subset significantly correlated with carbon emission factors is screened based on statistical tests, including: Monitoring indicators of fuel property dimension, regional characteristic dimension, time dynamic dimension, process technology dimension, equipment efficiency dimension and environmental parameter dimension are collected to construct a multi-dimensional data set; After performing standardized preprocessing on the multi-dimensional data set, the chi-square test or mutual information method is used to calculate the correlation between the features of each dimension and the carbon emission factors, and the features with a correlation higher than the preset significance threshold are retained to form a key feature subset.
[0007] In a preferred embodiment, the fuel property dimension includes fuel type, calorific value and carbon content indicators, the regional characteristic dimension includes regional energy structure and policy intensity indicators, and the time dynamic dimension includes seasonal energy consumption fluctuations and equipment operation cycle indicators.
[0008] In a preferred embodiment, an association matrix between the key feature subset and the carbon emission factors is established, cross-dimensional principal factors are extracted through principal component analysis, and the non-linear dependence strength of each cross-dimensional principal factor is quantified to generate hierarchical weight parameters, including: Calculate the Spearman rank correlation coefficient between the features of each dimension in the key feature subset and the carbon emission factors, and screen the features with the absolute value of the Spearman rank correlation coefficient greater than the preset correlation threshold to generate an initial association matrix; Perform eigenvalue decomposition on the initial association matrix, and extract the principal components with a cumulative contribution rate exceeding the preset contribution threshold as cross-dimensional principal factors; Use the kernel density estimation method to fit the joint probability distribution of the cross-dimensional principal factors and the carbon emission factors, and calculate their non-linear dependence strength based on the mutual information entropy formula; Calculate the weight distribution ratio based on the entropy weight method combined with the non-linear dependence strength of each cross-dimensional principal factor to generate hierarchical weight parameters.
[0009] In a preferred embodiment, a feature importance network based on gradient masking is constructed based on the hierarchical weight parameters, and the feature distribution of the multi-dimensional data set is optimized through feature channel scaling to generate an initial prediction model, including: Take the hierarchical weight parameters as the initial feature importance weights, and dynamically update the weight gradients of each feature channel through gradient backpropagation to generate a dynamic feature importance mask; Perform element-wise multiplication of the dynamic feature importance mask and the multi-dimensional data set to suppress low-importance feature noise and enhance key feature signals, and generate an optimized multi-dimensional feature distribution; Based on the optimized multi-dimensional feature distribution, select the gradient boosting tree model architecture, and use cross-validation and Bayesian optimization to adjust the model hyperparameters to generate an initial prediction model.
[0010] In a preferred embodiment, dedicated sub-models are constructed for multiple preset dimensions, and a dynamic weight allocation mechanism is adopted to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor, including: Based on the optimized multi-dimensional feature distribution, gradient boosting tree sub-models corresponding to the fuel property dimension, regional feature dimension, and process technology dimension are trained respectively; Calculate the dynamic weights according to the historical prediction errors of each gradient boosting tree sub-model, and perform weighted summation on the prediction results of the gradient boosting tree sub-models according to the dynamic weights; Normalize the weighted summation result and output the multi-scenario fusion carbon emission factor.
[0011] In a preferred embodiment, the initial prediction model is updated based on the incremental learning mechanism to generate historical verification data, and the historical knowledge is transferred to the dedicated sub-model by combining transfer learning to filter the input data noise, including: Input the newly added monitoring data into the initial prediction model, and incrementally update the model parameters through an online sequential extreme learning machine to generate an updated initial prediction model and historical verification data; Extract the weight distribution characteristics of the updated initial prediction model and adapt them to the weight space of the gradient boosting tree sub-model through a feature mapping function; Detect outliers in the newly added monitoring data based on the isolation forest algorithm, and input the data after removing the outliers into the gradient boosting tree sub-model.
[0012] In a preferred embodiment, the uncertainty interval of the multi-scenario fusion carbon emission factor is calibrated according to the historical verification data, and a dynamically optimized evaluation result is output, including: Generate the probability distribution of the multi-scenario fusion carbon emission factor based on Monte Carlo simulation, and calculate the confidence interval with the confidence level being a preset threshold; Adjust the upper and lower bound thresholds of the confidence interval in combination with preset rules to generate a dynamically optimized evaluation result.
[0013] In a preferred embodiment, the execution steps of Monte Carlo simulation include: randomly sampling from the historical verification data to generate a simulation data set, and iteratively calculating the distribution statistics of the multi-scenario fusion carbon emission factor; the preset rules include the statistical laws of historical data and external constraint conditions.
[0014] On the other hand, the present invention provides a multi-dimensional dynamic carbon emission factor modeling system, including: A multi-dimensional acquisition module: acquiring multi-dimensional monitoring data to construct a multi-dimensional data set, and screening a key feature subset significantly related to the carbon emission factor based on statistical tests; Cross - dimensional analysis module: Establish an association matrix between the key feature subset and carbon emission factors, extract cross - dimensional principal factors through principal component analysis, quantify the non - linear dependence strength of each cross - dimensional principal factor, and generate hierarchical weight parameters; Weight optimization module: Construct a feature importance network based on gradient masking based on hierarchical weight parameters, optimize the feature distribution of the multi - dimensional dataset through feature channel scaling, and generate an initial prediction model; Dynamic fusion module: Build dedicated sub - models for multiple preset dimensions, adopt a dynamic weight allocation mechanism to fuse the output results of the dedicated sub - models, and generate multi - scenario fusion carbon emission factors; Incremental migration module: Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, combine transfer learning to transfer historical knowledge to the dedicated sub - models and filter the input data noise; Dynamic calibration module: Calibrate the uncertainty interval of the multi - scenario fusion carbon emission factors according to the historical verification data, and output the dynamically optimized evaluation results.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Through the multi - dimensional dynamic feature fusion and hierarchical weight optimization mechanism, it effectively improves the analytical ability of carbon emission factor modeling for the dynamic features of cross - scale data. Based on the quantification of the non - linear dependence strength of the key feature subset and the gradient masking feature importance network, it accurately identifies the contribution degree of local abnormal emission events to the global evaluation, and solves the problem of distorted feature contribution degree caused by linear aggregation in traditional methods. Through dynamic weight allocation and multi - scenario sub - model collaborative fusion, it enhances the adaptability of the model to multi - dimensional differences such as regions and processes, and avoids the risk of scenario bias under static weight allocation. The combination of incremental learning and transfer learning continuously optimizes the model parameters based on historical verification data, ensures the real - time response ability of the modeling process to the dynamic evolution of data, and improves the quality of input data through noise filtering to ensure the robustness of the model.
[0016] 2. Through the uncertainty interval dynamic calibration mechanism, integrating Monte Carlo simulation and external constraint rules, it realizes the reliability and compliance of the evaluation results in multiple scenarios. The collaborative design of cross - dimensional principal factor extraction and feature channel scaling breaks through the technical bottleneck that it is difficult to be compatible with macroscopic steady - state features and microscopic dynamic features in traditional methods, enabling the model to still maintain high sensitivity and generalization performance in complex data environments. Through the data flow closed - loop and logical coupling, a complete technical chain is formed, systematically solving problems such as local signal omission, poor scenario adaptability, and inaccurate evaluation results in carbon emission monitoring. Description of the Drawings
[0017] Figure 1 It is a flowchart of a multi - dimensional dynamic carbon emission factor modeling method of the present invention; Figure 2This is a schematic structural diagram of a multi-dimensional dynamic carbon emission factor modeling system of the present invention. Detailed implementation manners
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] Embodiment 1: Figure 1 A multi-dimensional dynamic carbon emission factor modeling method of the present invention is given, including: S1. Collect multi-dimensional monitoring data to construct a multi-dimensional data set, and screen a key feature subset significantly related to the carbon emission factor based on statistical tests; S2. Establish an association matrix between the key feature subset and the carbon emission factor, extract cross-dimensional main factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional main factor, and generate hierarchical weight parameters; S3. Construct a feature importance network based on gradient masks based on the hierarchical weight parameters, optimize the feature distribution of the multi-dimensional data set through feature channel scaling, and generate an initial prediction model; S4. Construct dedicated sub-models for multiple preset dimensions, and adopt a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor; S5. Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, and transfer historical knowledge to the dedicated sub-model and filter input data noise in combination with transfer learning; S6. Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor according to the historical verification data, and output a dynamically optimized evaluation result.
[0020] S1. Collect multi-dimensional monitoring data to construct a multi-dimensional data set, and screen a key feature subset significantly related to the carbon emission factor based on statistical tests. The specific implementation steps are as follows: Collect monitoring indicators for the fuel property dimension. The fuel property dimension includes fuel type, calorific value, and carbon content indicators. The fuel type is obtained from the fuel classification list provided by the energy supplier. The fuel classification list is divided according to the fuel classification standard issued by the International Energy Agency and includes three major categories: fossil fuels, biomass fuels, and synthetic fuels. Each category is further subdivided into specific fuel names. For example, coal is subdivided into anthracite and lignite, and natural gas is subdivided into liquefied natural gas and pipeline natural gas. The calorific value is obtained through laboratory testing or the standard calorific value table provided by the fuel supplier. In laboratory testing, the oxygen bomb calorimetry method is used. The fuel sample is completely burned in an oxygen bomb calorimeter, and the released heat value is measured, with the unit of megajoules per kilogram (MJ / kg). The standard calorific value table refers to the fuel calorific value database issued by the International Organization for Standardization and indexes and matches the corresponding calorific value according to the fuel type. The carbon content is determined by an elemental analyzer or by referring to the standard value in the international carbon content database. When using an elemental analyzer to determine, after the fuel sample is burned, the carbon dioxide emission is analyzed by gas chromatography and converted into the carbon emission intensity per unit calorific value, with the unit of kilograms of carbon dioxide equivalent per megajoule (kg CO2e / MJ).
[0021] Collect monitoring indicators for the regional characteristic dimension. The regional characteristic dimension includes the regional energy structure and policy intensity indicators. The regional energy structure is obtained from the annual energy consumption statistics report issued by the local government. The proportion data of coal power, hydropower, and wind power in the statistics report is calculated by dividing the consumption of each energy type by the total energy consumption. The calculation formula is that the proportion of an energy is equal to the consumption of that energy divided by the total energy consumption, and the result is expressed as a percentage (%). The policy intensity indicator is quantified through the carbon emission supervision policy documents publicly released by the environmental protection department. The carbon tax rate in the policy document is directly extracted, with the unit of yuan per ton of carbon dioxide equivalent (yuan / t CO2e). The completion rate of the emission reduction target is calculated by the ratio of the actual emission reduction amount to the target emission reduction amount, and the result is expressed as a percentage (%).
[0022] Collect monitoring indicators for the time dynamic dimension. The time dynamic dimension includes seasonal energy consumption fluctuations and equipment operation cycle indicators. Seasonal energy consumption fluctuations are extracted from historical energy consumption records. The standard deviation of monthly energy consumption is statistically calculated on a quarterly basis. The standard deviation calculation formula is the square root of the mean of the sum of the squared deviations of monthly energy consumption from the quarterly average energy consumption, which is used to characterize the amplitude of energy consumption fluctuations, with the unit of megawatt-hours (MWh). The equipment operation cycle indicator is collected through equipment sensors and includes the continuous operation duration of the equipment and the shutdown and maintenance interval. The continuous operation duration is the cumulative number of hours (h) from the start to the shutdown of the equipment, and the shutdown and maintenance interval is the number of days (d) between adjacent two maintenance operations.
[0023] Collect monitoring indicators for the process technology dimension. The process technology dimension includes production process parameters and technology iteration data. The production process parameters are obtained through the factory production management system. The production management system records the reaction temperature, pressure, and raw material ratio in real time. The unit of reaction temperature is degrees Celsius (°C), the unit of pressure is kilopascals (kPa), and the raw material ratio is the mass percentage of each raw material, expressed as a percentage (%). The technology iteration data is recorded through the enterprise R & D log, including the application time of the new process and the percentage of energy efficiency improvement. The percentage of energy efficiency improvement is calculated by dividing the difference in unit product energy consumption between the old and new processes by the energy consumption of the old process. The calculation formula is that the percentage of energy efficiency improvement is equal to (old process energy consumption - new process energy consumption) divided by old process energy consumption multiplied by 100%.
[0024] Collect monitoring indicators for the equipment efficiency dimension. The equipment efficiency dimension includes equipment operation efficiency and maintenance status parameters. The equipment operation efficiency is calculated by monitoring the ratio of energy consumption to output of the equipment in real time. The calculation formula is that the equipment operation efficiency is equal to the total mass of the output product (kg) divided by the total energy consumption of the equipment (kWh), and the unit is kilograms per kilowatt-hour (kg / kWh); the maintenance status parameters are obtained through the equipment maintenance record. The maintenance record includes the date of the last maintenance and the frequency of fault codes. The frequency of fault codes is counted as the number of times a specific fault code appears during every 1000 hours of operation (times / 1000h).
[0025] Collect monitoring indicators for the environmental parameter dimension. The environmental parameter dimension includes temperature, humidity, and air pressure monitoring values. The temperature and humidity are collected in real time through Internet of Things sensors. The sensor model selected is a temperature and humidity integrated sensor that meets industrial environmental standards. The temperature measurement accuracy is ±0.5 degrees Celsius (°C), the humidity measurement accuracy is ±3%, and the data collection frequency is once per minute; the air pressure is obtained through the data interface of the weather station. The data interface of the weather station updates the air pressure data hourly, and the unit is kilopascals (kPa).
[0026] Align the monitoring indicators of all the above dimensions according to the timestamp to construct a multi-dimensional dataset. The specific alignment method is as follows: Based on the minute-level timestamp, linearly interpolate and fill the hour-level or day-level data (such as the air pressure of the weather station) to the minute-level to ensure that all dimension data has the same time granularity. The dataset is stored as a two-dimensional table. The rows represent the time series samples, and the columns represent the monitoring indicators of each dimension. Missing values are processed using the forward filling method, that is, filling the current missing value with the valid value of the previous moment.
[0027] Perform standardized preprocessing on the multi-dimensional dataset using the Z-score standardization method. The Z-score standardization is calculated separately for each monitoring indicator. The calculation formula is that the standardized value is equal to the original value minus the mean of the indicator and then divided by the standard deviation of the indicator. After processing, the mean of each indicator of the data is 0, and the standard deviation is 1. The mean and standard deviation are calculated based on the training set data and the same parameters are used in the test set data to avoid data leakage.
[0028] Based on the standardized multi-dimensional dataset, the chi-square test or mutual information method is used to calculate the correlation between each dimensional feature and the carbon emission factor. The implementation steps of the chi-square test include: discretizing the continuous monitoring indicators by the equal-width binning method, and the number of bin intervals is dynamically adjusted according to the data distribution. The specific rules are as follows: if the data distribution is approximately normal, the number of bins is 10; if the data is skewed, the number of bins is 5; count the actual frequency of the carbon emission factor in each bin interval, and calculate the theoretical expected frequency, where the expected frequency is equal to the number of samples in the bin multiplied by the average distribution probability of the overall carbon emission factor; calculate the chi-square statistic, and the calculation formula is the sum of the squares of the differences between the actual frequency and the expected frequency in each bin divided by the expected frequency; obtain the P-value by looking up the chi-square distribution table based on the chi-square statistic and the degrees of freedom, and retain the features with a P-value less than 0.05; the 0.05 significance threshold is determined according to the general standard of statistical significance level.
[0029] The implementation steps of the mutual information method include: using the kernel density estimation method to fit the joint probability distribution of the continuous monitoring indicators and the carbon emission factor, selecting the Gaussian kernel as the kernel function, and automatically calculating the bandwidth parameter by the Silverman rule. The calculation formula is that the bandwidth is equal to 1.06 times the sample standard deviation multiplied by the negative one-fifth power of the sample size; calculate the marginal probability distribution of the monitoring indicators and the carbon emission factor, and the marginal probability is obtained by integrating the joint probability distribution; calculate the mutual information value based on the entropy formula, and the expected value of the logarithm ratio of the joint probability and the marginal probability in the entropy formula is the mutual information value; sort the mutual information values of all features in descending order, and retain the top 30% of the features as significantly correlated features; the 30% quantile threshold is determined through experimental verification, and the experimental results show that the F1 score of the model reaches more than 0.85 under this threshold.
[0030] Take the union of the screening results of the chi-square test and the mutual information method, and retain the features determined to be significantly correlated by at least one method to form a key feature subset. The key feature subset includes fuel type, calorific value, regional energy structure, seasonal energy consumption fluctuation, equipment operation efficiency, and temperature and humidity indicators. The key feature subset is stored in the form of a two-dimensional table, with rows representing time series samples and columns representing the screened feature indicators, which is used for the subsequent construction step of the correlation matrix.
[0031] S2. Establish the correlation matrix between the key feature subset and the carbon emission factor, extract the cross-dimensional main factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional main factor, and generate hierarchical weight parameters. The specific implementation steps are as follows: Based on the key feature subset, the Spearman rank correlation coefficient of each dimensional feature and the carbon emission factor is calculated. The Spearman rank correlation coefficient is calculated by the following steps: sort the dimensional features and carbon emission factors in the key feature subset according to the numerical size and assign a rank, where the rank is the ranking position of the data point in the sequence to which it belongs. For example, the data point with the smallest numerical value has a rank of 1, the second smallest has a rank of 2, and so on; calculate the sum of the squares of the rank differences of each pair of data points, where the sum of the squares of the rank differences of the corresponding data points is accumulated and summed; generate the correlation coefficient value according to the formula, which is 1 minus 6 times the sum of the squares of the rank differences divided by the sample size multiplied by the square of the sample size minus 1. The correlation coefficient value ranges from -1 to 1, and the larger the absolute value, the stronger the linear correlation. The preset correlation threshold is set to 0.3, which is determined by cross-validation experiments. The specific experimental method is: try 0.2, 0.3, and 0.4 as thresholds on the training set, calculate the features selected under each threshold to build a model, and evaluate the prediction accuracy on the test set. The results show that when the threshold is 0.3, the root mean square error of the model is 15% lower than that of other thresholds, so 0.3 is selected as the optimal threshold. Features with an absolute value of correlation coefficient greater than 0.3 are selected to generate an initial correlation matrix. The rows of the initial correlation matrix represent the features in the key feature subset, the columns represent the carbon emission factors, the matrix elements are the corresponding Spearman rank correlation coefficient values, and the matrix dimension is M×N, where M is the number of features in the key feature subset and N is the number of carbon emission factors.
[0032] Perform eigenvalue decomposition on the initial correlation matrix, calculate the eigenvalues and corresponding eigenvectors of the matrix, and the eigenvalue decomposition is achieved by decomposing the matrix into the product of the eigenvector matrix and the diagonal eigenvalue matrix. Arrange the eigenvalues in descending order and calculate the cumulative contribution rate. The cumulative contribution rate is the percentage of the sum of the first N eigenvalues to the total sum of all eigenvalues. For example, the sum of the first three eigenvalues is 80% of the total, and the sum of the first four eigenvalues is 90% of the total, then the cumulative contribution rates are 80% and 90% respectively. The preset contribution threshold is set to 85%, that is, when the cumulative contribution rate reaches 85%, the principal component extraction is stopped. This threshold is set according to the industry practice of principal component analysis to ensure that the main information is retained while reducing the data dimension. The principal components corresponding to the first N eigenvectors that meet the cumulative contribution rate threshold are used as cross-dimensional principal factors. The cross-dimensional principal factors characterize the global coupling relationship between the characteristics of each dimension in the key feature subset and the carbon emission factor. For example, the first principal component can be interpreted as a comprehensive influencing factor of fuel properties and regional energy structure.
[0033] The kernel density estimation method is used to fit the joint probability distribution of the cross-dimensional principal factors and the carbon emission factors. The kernel function of the kernel density estimation selects the Gaussian kernel. The Gaussian kernel function is a symmetric bell-shaped curve, and its width is controlled by the bandwidth parameter. The bandwidth parameter is calculated by the Silverman's rule. The calculation formula of the Silverman's rule is that the bandwidth is equal to 1.06 multiplied by the sample standard deviation multiplied by the negative one-fifth power of the sample size. Based on the joint probability distribution and the marginal probability distribution, the non-linear dependence strength between the cross-dimensional principal factors and the carbon emission factors is calculated through the mutual information entropy formula. The calculation steps of the mutual information entropy include: calculating the logarithm of the ratio of the product of the joint probability distribution and the marginal probability distribution, where the ratio is the joint probability distribution divided by the product of the marginal probability distributions; obtaining the expected value of all possible values, that is, performing a weighted sum on the logarithm of the ratio under the joint probability distribution, and the weight is the corresponding joint probability value. The range of the non-linear dependence strength is from 0 to positive infinity. The larger the value, the stronger the explanatory ability of the cross-dimensional principal factor for the carbon emission factor. For example, the non-linear dependence strength of a certain principal factor is 1.2, indicating that its explanatory ability for the carbon emission factor is significantly higher than that of other principal factors with a dependence strength of 0.8.
[0034] Based on the non-linear dependence strength of each cross-dimensional principal factor, the entropy weight method is used to calculate the weight allocation ratio. The implementation steps of the entropy weight method include: normalizing the non-linear dependence strength into a probability distribution. The normalization method is to divide the dependence strength value of each principal factor by the sum of the dependence strength values of all principal factors, so that the sum of the normalized values is 1; calculating the information entropy value of each principal factor, and the information entropy value is the cumulative sum of the product of the negative logarithm of the normalized value and itself; calculating the difference coefficient according to the information entropy value, and the difference coefficient is equal to 1 minus the information entropy value. The larger the difference coefficient, the higher the weight of the principal factor should be; the final weight is the ratio of the difference coefficient to the sum of the difference coefficients of all principal factors. The generated hierarchical weight parameters are stored in vector form. The vector elements represent the importance weights of each cross-dimensional principal factor in the carbon emission factor modeling, and the sum of the weights is 1. For example, the weights of three principal factors are 0.5, 0.3, and 0.2 respectively, indicating that the first principal factor has the highest contribution degree to the carbon emission factor.
[0035] Among them, the Spearman correlation coefficient threshold of 0.3 is determined by cross-validation. Experiments show that the root mean square error of the model is reduced by 15% under this threshold; the cumulative contribution rate threshold of 85% is set according to the industry convention of principal component analysis to ensure that the main information is retained after dimensionality reduction; the bandwidth parameter of the kernel density estimation is automatically calculated by the Silverman's rule to avoid the deviation of manual intervention.
[0036] In this step S2, a correlation matrix is constructed through the Spearman correlation coefficient to capture the non-linear correlation between features and carbon emission factors, breaking through the limitations of traditional linear correlation analysis; principal component analysis is used to extract cross-dimensional principal factors to solve the problem of multi-dimensional data redundancy and retain key coupling information; kernel density estimation and mutual information entropy are used to quantify the intensity of non-linear dependence and accurately identify the influence weights of cross-scale dynamic features; the entropy weight method combines information entropy to dynamically allocate weights, avoiding subjective preset bias. Compared with the data fusion methods based on mean aggregation or fixed weights in the prior art, this step significantly improves the model's analytical ability for local abnormal signals and cross-dimensional interaction effects through non-linear statistics and dynamic weight allocation, solves the problem of distorted feature contribution degrees caused by linear assumptions in traditional methods, and enhances the sensitivity and scenario generalization of carbon emission factor modeling.
[0037] S3. Construct a feature importance network based on gradient masking using hierarchical weight parameters, optimize the feature distribution of the multi-dimensional dataset through feature channel scaling, and generate an initial prediction model. The specific implementation steps are as follows: Take the hierarchical weight parameters as the initial feature importance weights, and dynamically update the weight gradients of each feature channel through gradient backpropagation to generate a dynamic feature importance mask. The hierarchical weight parameters are derived from the calculation results of the non-linear dependence intensity of cross-dimensional principal factors and are stored in vector form. The vector elements represent the initial importance weights of each feature channel. Gradient backpropagation is implemented based on a neural network architecture. The input of the neural network is the multi-dimensional dataset, and the output is the estimated value of the predicted carbon emission factor. The loss function uses the mean square error. During the backpropagation process, calculate the gradient of the loss function with respect to the weights of each feature channel. The gradient value reflects the contribution degree of the feature channel to the prediction error, and dynamically update the initial importance weights to generate a dynamic feature importance mask. The dynamic feature importance mask is a weight matrix with the same channel dimension as the multi-dimensional dataset. The matrix element values range from 0 to 1, and the larger the value, the higher the importance of the corresponding feature channel. For example, if the initial weight of a certain feature channel is 0.5 and the weight is increased to 0.9 after gradient update, it indicates that the contribution degree of this channel to the prediction result has been significantly enhanced.
[0038] Perform a per-channel multiplication operation on the dynamic feature importance mask and the multi-dimensional dataset to suppress low-importance feature noise and enhance key feature signals, generating an optimized multi-dimensional feature distribution. The per-channel multiplication operation is performed separately for each feature channel of the multi-dimensional dataset, multiplying the original data value of the feature channel by the weight value of the corresponding channel in the dynamic feature importance mask to achieve feature channel scaling. For example, if the original data value of a feature channel is 100 and the mask weight is 0.8, the scaled data value is 80; if the mask weight is 0.2, the scaled data value is 20. Through this operation, the data values of low-weight feature channels are suppressed, and the data values of high-weight feature channels are retained or enhanced, forming an optimized multi-dimensional feature distribution. The optimized multi-dimensional feature distribution is stored in the form of a two-dimensional table, with rows representing time series samples and columns representing the scaled feature channel data values.
[0039] Select the gradient boosting tree model architecture based on the optimized multi-dimensional feature distribution, and use cross-validation and Bayesian optimization to adjust the model hyperparameters to generate an initial prediction model. The tree depth of the gradient boosting tree model architecture is set to 3 to 8 layers, and the learning rate is set to 0.01 to 0.1. The hyperparameter range is determined based on experience and experimental verification. For example, if the tree depth exceeds 8 layers, it is likely to cause overfitting, and if it is less than 3 layers, the model will be underfitted. The cross-validation uses 5-fold cross-validation, dividing the optimized multi-dimensional feature distribution into 5 subsets. In turn, 4 subsets are used as the training set and 1 subset is used as the validation set, and the average value of the model prediction accuracy is calculated as the evaluation index. Bayesian optimization models the relationship between hyperparameters and model accuracy through Gaussian process regression, iteratively searches for the optimal hyperparameter combination, and the number of iterations is set to 50 times. When the accuracy improvement in 10 consecutive iterations is less than 1%, the optimization is terminated in advance. This termination condition is determined through experimental verification. Experiments show that this threshold can balance the computational efficiency and the optimization effect. The initial prediction model is stored in the serialized file format, including the trained gradient boosting tree structure and hyperparameter configuration, which are used for subsequent incremental learning and transfer learning steps.
[0040] Among them, the gradient boosting tree depth range (3 - 8 layers) is determined through experimental verification. Experiments show that exceeding this range is likely to cause overfitting or underfitting; the number of Bayesian optimization iterations, 50 times, is determined through convergence testing to ensure the balance between search efficiency and accuracy; the early termination condition (improvement in 10 consecutive iterations < 1%) is set based on the statistical results of historical training data to avoid ineffective calculations.
[0041] In this step S3, the adaptive scaling of feature channels is achieved through a dynamic feature importance mask, breaking through the limitations of traditional static weight allocation, accurately identifying key features and suppressing noise interference. By combining gradient boosting trees and Bayesian optimization, it dynamically adapts to the characteristics of data distribution and improves the model's ability to analyze cross-dimensional interaction effects. Compared with the fixed feature weights or single model architecture in the existing technology, this step dynamically updates the feature importance through gradient backpropagation, combines cross-validation and hyperparameter optimization, effectively solves the problem of insufficient model sensitivity caused by feature redundancy and noise, and enhances the accuracy of carbon emission factor prediction and the response ability to local abnormal events.
[0042] S4. Build dedicated sub-models for multiple preset dimensions, and use a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate multi-scenario fusion carbon emission factors. The specific implementation steps are as follows: Based on the optimized multi-dimensional feature distribution, gradient boosting tree sub-models corresponding to the fuel property dimension, regional feature dimension, and process technology dimension are trained respectively. The optimized multi-dimensional feature distribution is from the output of the feature channel scaling step and is stored in the form of a two-dimensional table. The rows represent time series samples, and the columns represent the scaled feature channel data values. For the fuel property dimension, the feature channels related to fuel type, calorific value, and carbon content are selected from the optimized multi-dimensional feature distribution. The fuel type identifies the columns whose feature names contain keywords such as "fuel type", "calorific value", and "carbon content" through string matching. For example, the columns with feature names "fuel type_coal" and "calorific value_MJ / kg" are selected; for the regional feature dimension, the feature channels related to regional energy structure and policy intensity are selected by matching the fields "regional energy" and "policy intensity" in the feature names through regular expressions. For example, "regional energy_coal power ratio" and "policy intensity_carbon tax rate" are selected; for the process technology dimension, the feature channels related to production process parameters and technology iteration data are selected according to the identifiers "reaction temperature", "raw material ratio", and "technology iteration" in the feature names. For example, "reaction temperature_°C" and "technology iteration_energy efficiency improvement percentage" are selected. The tree depth of each sub-model is set to 3 to 5 layers, and the learning rate is set to 0.05 to 0.1. The tree depth range is determined through cross-validation experiments. The experiment shows that when the tree depth exceeds 5 layers, the validation set error increases by 15%, indicating a significant increase in the overfitting risk; the learning rate range is set according to the convergence of gradient boosting tree training. When the learning rate is lower than 0.05, the model convergence speed drops by 30%, and when it is higher than 0.1, oscillations occur.
[0043] Calculate the dynamic weights based on the historical prediction errors of each gradient boosting tree sub-model, and sum the prediction results of the gradient boosting tree sub-models by weighted summation according to the dynamic weights. The historical prediction errors are calculated by the root mean square error of the sub-model on the validation set. The validation set data is divided in chronological order, and the data in the most recent 30% time window is taken as the validation set to ensure that the evaluation results reflect the latest state of the model. The calculation formula of the dynamic weight is that the weight is equal to the reciprocal of the sub-model error divided by the sum of the reciprocals of all sub-model errors. For example, if the error of the fuel property sub-model is 10, the error of the regional feature sub-model is 15, and the error of the process technology sub-model is 20, then the weight of the fuel property is (1 / 10) / (1 / 10 + 1 / 15 + 1 / 20) = 0.5, the weight of the regional feature is 0.33, and the weight of the process technology is 0.17. If the historical error of a certain sub-model is 0, its weight is forcibly set to 1, and the weights of other sub-models are 0 to avoid division-by-zero errors; if the errors of all sub-models are 0, the weights are evenly distributed. The weighted summation operation is to multiply the predicted values of each sub-model by their dynamic weights and then accumulate them to generate a preliminary fusion result. For example, if the predicted value of the fuel property sub-model is 100, the predicted value of the regional feature sub-model is 110, and the predicted value of the process technology sub-model is 90, then the fusion result is 100×0.5 + 110×0.33 + 90×0.17 = 101.6.
[0044] Perform normalization processing on the weighted summation result and output the multi-scenario fusion carbon emission factor. The normalization processing uses the maximum-minimum normalization method to linearly map the preliminary fusion result to the interval of 0 to 1. The normalization formula is that the fusion factor is equal to the preliminary fusion value minus the minimum value and then divided by the difference between the maximum value and the minimum value. The maximum value and the minimum value are determined by statistical analysis of historical data, and the statistical time window is the past 12 months. For example, if the historical minimum fusion value is 50 and the maximum fusion value is 150, then the current fusion value of 101.6 is normalized to (101.6 - 50) / (150 - 50) = 0.516. If the maximum value and the minimum value are equal, skip the normalization step and directly output the original fusion value to avoid division-by-zero errors. The normalized fusion factor is stored in a standardized format, retaining four decimal places of precision for subsequent uncertainty calibration steps.
[0045] Among them, the tree depth range (3 - 5 layers) is determined by cross-validation. Experiments show that going beyond this range leads to overfitting or underfitting; the learning rate range (0.05 - 0.1) is set based on training convergence, and experimental data supports its effectiveness; the validation set time window (30%) is determined by timeliness analysis to ensure that the model evaluation reflects the latest data distribution.
[0046] S5. Update the initial prediction model based on the incremental learning mechanism and generate historical validation data, and transfer historical knowledge to the dedicated sub-model and filter the input data noise by combining transfer learning. The specific implementation steps are as follows: The newly added monitoring data is input into the initial prediction model, and the model parameters are incrementally updated through an online sequential extreme learning machine to generate the updated initial prediction model and historical verification data. The newly added monitoring data is sourced from multi-dimensional monitoring metrics collected in real time, with the same format as the training data, including feature channel data in dimensions of fuel properties, regional characteristics, and process technologies. For example, the dimension of fuel properties includes indicators such as fuel type, calorific value, and carbon content, and the dimension of regional characteristics includes indicators such as regional energy structure and policy intensity.
[0047] The incremental learning process of the online sequential extreme learning machine includes: inputting the newly added data into the model in chronological order, calculating the model output error, updating the model weights based on error backpropagation, and setting the update step size to be between 0.001 and 0.01. The step size range is determined through experimental verification. The experimental method is to test the impact of different step sizes on the model convergence speed and stability on the historical dataset. The results show that when the step size is less than 0.001, the number of iterations required for the model to converge increases by 50%, and when the step size is greater than 0.01, the error fluctuation range of the validation set exceeds 20%. Therefore, the range of 0.001 to 0.01 is selected as the reasonable range. The historical verification data is stored as a two-dimensional table containing timestamps, input features, and model prediction results for subsequent uncertainty calibration steps. The storage format is a Parquet file with the default compression ratio.
[0048] Extract the weight distribution characteristics of the updated initial prediction model and adapt them to the weight space of the gradient boosting tree sub-model through a feature mapping function. The weight distribution characteristics are obtained by extracting the weight matrix of the fully connected layer of the initial prediction model. The rows of the matrix represent the input feature dimensions, and the columns represent the output nodes. For example, if the input feature dimension is 50 and the output node is 1, the dimension of the weight matrix is 50×1. The feature mapping function uses an attention mechanism, which realizes the calculation by computing the similarity between the weights of each tree node in the gradient boosting tree sub-model and the weights of the initial model. The similarity calculation method is cosine similarity, and the calculation formula is the dot product of two vectors divided by the product of the vector norms, with a value range of -1 to 1. The weights of the tree nodes with a similarity higher than 0.8 are replaced with the corresponding weights of the initial model, and the nodes with a similarity lower than 0.8 retain their original weights. For example, if the weight of the fuel property dimension in the initial model is 0.9, and the similarity of a certain tree node corresponding to the feature is 0.85, then the weight of this node is updated to 0.9; if the similarity is 0.6, the original weight of 0.5 remains unchanged. The similarity threshold of 0.8 is determined through a transfer effectiveness experiment. The experimental method is to compare the prediction accuracy of the sub-model under different thresholds. The results show that when the threshold is 0.8, the root mean square error of the model is reduced by 8% compared to the threshold of 0.7.
[0049] Detect outliers in the newly added monitoring data based on the Isolation Forest algorithm, and input the data after removing the outliers into the Gradient Boosting Tree sub-model. The Isolation Forest algorithm divides the data space by constructing multiple isolation trees. The number of isolation trees is set to 100, and the maximum depth of each tree is 8 layers. The parameters are set according to the default configuration of the scikit-learn library. When detecting outliers, calculate the path length of the data points to judge the degree of abnormality. The path length threshold is set to 1.5 times less than the average path length of the data points. The threshold is determined based on the statistical distribution of historical data. The statistical method is to calculate the path length distribution of the data in the past three months. For example, if the average path length of historical normal data is 10 and the standard deviation is 2, then the threshold is 10 + 2×1.5 = 13. After detecting outliers, mark and remove the abnormal data, and input the remaining data into the Gradient Boosting Tree sub-model for prediction. If the proportion of outliers in a single batch of data exceeds 30%, trigger the manual review mechanism, suspend the automatic prediction process, and notify the administrator by email.
[0050] Among them, the step size range (0.001 - 0.01) of the online sequential extreme learning machine is determined through convergence experiments. The experiments show that when the step size is less than 0.001, the convergence speed decreases by 50%, and when the step size is greater than 0.01, the model oscillates; the similarity threshold of 0.8 is determined through transfer effectiveness tests. The experiments show that when the threshold is lower than this value, the model accuracy decreases by 12%; the Isolation Forest path length threshold is set through the statistics of three months of historical data to ensure that the threshold is dynamically adapted to the data distribution.
[0051] S6. Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor based on historical verification data, and output the dynamically optimized evaluation results. The specific implementation steps are as follows: Generate the probability distribution of the multi-scenario fusion carbon emission factor based on Monte Carlo simulation, and calculate the confidence interval with the confidence level being the preset threshold. The execution steps of Monte Carlo simulation include: randomly sampling from the historical verification data to generate a simulation data set. The historical verification data comes from the output of the incremental learning step and is stored as a two-dimensional table containing timestamps, input features, and model prediction results. The data format is a Parquet file, and the compression ratio is the default value. The random sampling method uses sampling with replacement, and the sampling volume each time is 80% of the total amount of historical data. For example, if the historical data contains 1000 samples, then 800 samples are sampled each time. The sampling ratio is determined through experimental verification. The experiments show that a sampling volume of 80% can balance the calculation efficiency and result stability.
[0052] After the simulated dataset is generated, the distribution statistics of the carbon emission factors for multi-scenario fusion are calculated iteratively. The number of iterations is set to 10,000 times. In each iteration, the mean and standard deviation of the fusion factor are calculated to generate a probability density function. The probability density function is fitted by the kernel density estimation method, and the Gaussian kernel is selected as the kernel function. The bandwidth parameter is calculated by the Silverman's rule. The preset threshold of the confidence level is set to 95%. The upper and lower bounds of the confidence interval are calculated as the 2.5% quantile and 97.5% quantile of the probability density function. For example, if the simulated dataset shows that the mean of the fusion factor is 100 and the standard deviation is 10, then the 95% confidence interval is from 80 to 120.
[0053] The upper and lower bounds of the confidence interval are adjusted in combination with the preset rules to generate a dynamically optimized evaluation result. The preset rules include the statistical laws of historical data and external constraint conditions. The statistical laws of historical data are determined by analyzing the fluctuation range of the fusion factor in the past year. The specific method is to calculate the maximum value, minimum value, and standard deviation of the historical data. For example, if the historical maximum value is 130, the minimum value is 70, and the standard deviation is 15, then the adjusted upper and lower bounds of the confidence interval shall not exceed the range of the maximum and minimum values, and the fluctuation range of the standard deviation is limited within 1.5 times of the historical standard deviation. The external constraint conditions are realized by parsing the carbon emission intensity limit clauses in the policy documents. The parsing method is to use regular expressions to match the keywords in the policy text (such as "upper limit of carbon emission factor" and "allowed fluctuation range"), extract the numerical constraints and convert them into thresholds. For example, if the policy requires that the carbon emission factor of a certain industry shall not exceed 110, after matching "upper limit 110", the upper bound threshold of the confidence interval is adjusted from 120 to 110. The adjusted dynamically optimized evaluation result is stored in a standardized format, including the upper and lower bounds of the confidence interval, the adjustment basis, and the timestamp, which is used for visual display and decision support. The storage format is a JSON file, and the encoding method is UTF-8.
[0054] Among them, the 10,000 Monte Carlo simulation iterations are determined through a convergence test. The test method is to run 5,000, 10,000, and 20,000 iterations respectively, and calculate the fluctuation range of the confidence interval. The results show that the fluctuation range is stable within 2% when the number of iterations is 10,000. The 95% confidence level is set according to the general statistical standard and the reliability requirements of the industrial scenario. Experiments show that the 95% confidence level can cover more than 90% of the actual data distribution. The historical data fluctuation range (maximum value 130, minimum value 70) is obtained through descriptive statistical analysis. The analysis window is the rolling 12-month data to ensure dynamic adaptation to data changes.
[0055] In this embodiment, through cross-dimensional principal factor extraction and non-linear dependence strength quantification, the limitations of traditional linear aggregation methods are broken through, and the problem of distorted feature contribution degrees in multi-source heterogeneous data fusion is solved. A dynamic feature importance network is constructed based on hierarchical weight parameters, and local abnormal signals are enhanced through gradient masking and feature channel scaling, replacing fixed weights or manual experience settings, and improving the response sensitivity of the model to microscopic dynamic features. A dedicated sub-model is constructed for multi-dimensional scenarios and a dynamic weight allocation mechanism is designed to overcome the deficiency of the single model's adaptability to cross-scenario data distribution differences. Through the collaborative optimization of incremental learning and transfer learning, the dual generalization ability of model parameters in the time dimension and the scenario dimension is achieved. Uncertainty interval calibration combines Monte Carlo simulation and dynamic rule feedback, dynamically fusing statistical inference results with external constraint conditions, and avoiding the failure risk of static confidence intervals when policy changes or data distributions shift. The data transfer and logical closed-loop between each step form a complete technical chain. For example, the key feature subset drives the generation of cross-dimensional principal factors, the hierarchical weight parameters guide the optimization of the feature importance network, and the dynamic evaluation results reverse-constrain the model update, forming an enhanced loop of "data screening - feature fusion - model iteration - result calibration". This embodiment systematically solves the problems of sensitivity attenuation, scenario bias, and evaluation inaccuracy in multi-scale data fusion by coupling non-linear statistics, dynamic weight migration, and cross-scenario generalization mechanisms.
[0056] Embodiment 2: Figure 2 A structural schematic diagram of a multi-dimensional dynamic carbon emission factor modeling system according to the present invention is given. A multi-dimensional dynamic carbon emission factor modeling system includes: Multi-dimensional acquisition module: Collect multi-dimensional monitoring data to construct a multi-dimensional data set, and screen a key feature subset significantly related to the carbon emission factor based on statistical tests; Cross-dimensional analysis module: Establish an association matrix between the key feature subset and the carbon emission factor, extract cross-dimensional principal factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional principal factor, and generate hierarchical weight parameters; Weight optimization module: Construct a gradient-mask-based feature importance network based on hierarchical weight parameters, optimize the feature distribution of the multi-dimensional data set through feature channel scaling, and generate an initial prediction model; Dynamic fusion module: Construct dedicated sub-models for multiple preset dimensions, and use a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor; Incremental migration module: Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, and combine transfer learning to transfer historical knowledge to the dedicated sub-model and filter input data noise; Dynamic calibration module: Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor according to the historical verification data, and output a dynamically optimized evaluation result.
[0057] The multi-dimensional acquisition module provides input for the cross-dimensional analysis module by collecting multi-dimensional monitoring data and screening key features; the cross-dimensional analysis module extracts cross-dimensional principal factors and quantifies non-linear dependencies, and the generated hierarchical weight parameters are input into the weight optimization module; the weight optimization module constructs a feature importance network based on the weight parameters, and the optimized feature distribution drives the dynamic fusion module to train multi-dimensional sub-models and fuse the outputs; the multi-scenario carbon emission factors generated by the dynamic fusion module are input into the incremental migration module, which updates the model parameters through incremental learning and migrates historical knowledge, while filtering out noise data to ensure the input quality; the historical verification data output by the incremental migration module is input into the dynamic calibration module, which combines Monte Carlo simulation and dynamic rules to calibrate the uncertainty interval, and finally generates a dynamically optimized evaluation result.
[0058] Each module forms a collaborative mechanism through data flow and logical closed-loop: the cross-dimensional analysis module solves the problem of distorted feature contribution degree caused by linear aggregation of cross-scale data in traditional methods; the weight optimization module enhances the ability to identify local abnormal signals through dynamic feature scaling; the dynamic fusion module uses dynamic weight allocation to solve the model bias problem caused by the distribution differences of multi-scenario data; the incremental migration module enables the model to continuously adapt to the dynamic changes of data; the dynamic calibration module improves the reliability and compliance of the evaluation results through the fusion of statistics and rules. The overall solution systematically solves the problems of insufficient sensitivity, poor scenario generalization, and inaccurate evaluation in carbon emission factor modeling through modular design and data transfer mechanism.
[0059] All the calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and threshold selections in the calculations are set by those skilled in the art according to the actual situation.
[0060] It should be noted that the present invention can be deployed on the device itself to achieve embedded applications, or can also run on a PC or other terminals with a user interface, so as to meet various hardware environments and usage requirements.
[0061] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center wirelessly or wiredly (such as infrared, wireless, microwave, etc.); where the wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission includes infrared, microwave, etc. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0062] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0063] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0064] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0065] In addition, in each embodiment of the present application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0066] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0067] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0068] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-dimensional dynamic carbon emission factor modeling method, characterized in that Including: S1. Collect multi-dimensional monitoring data to construct a multi-dimensional data set, and screen a key feature subset significantly related to carbon emission factors based on statistical tests; S2. Establish an association matrix between the key feature subset and carbon emission factors, extract cross-dimensional main factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional main factor, and generate hierarchical weight parameters; S3. Construct a feature importance network based on gradient masking based on hierarchical weight parameters, optimize the feature distribution of the multi-dimensional data set through feature channel scaling, and generate an initial prediction model; S4. Construct dedicated sub-models for multiple preset dimensions, and adopt a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor; S5. Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, and transfer historical knowledge to the dedicated sub-model and filter input data noise by combining transfer learning; S6. Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor according to the historical verification data, and output a dynamically optimized evaluation result.
2. The multi-dimensional dynamic carbon emission factor modeling method according to claim 1, characterized in that Collect multi-dimensional monitoring data to construct a multi-dimensional data set, and screen a key feature subset significantly related to carbon emission factors based on statistical tests, including: Collect monitoring indicators of fuel property dimension, regional characteristic dimension, time dynamic dimension, process technology dimension, equipment efficiency dimension and environmental parameter dimension to construct a multi-dimensional data set; After preprocessing the multi-dimensional data set by standardization, calculate the correlation between the features of each dimension and carbon emission factors by using chi-square test or mutual information method, and retain the features with a correlation higher than the preset significance threshold to form a key feature subset.
3. The multi-dimensional dynamic carbon emission factor modeling method according to claim 2, wherein, Among them, the fuel property dimension includes fuel type, calorific value and carbon content indicators, the regional characteristic dimension includes regional energy structure and policy intensity indicators, and the time dynamic dimension includes seasonal energy consumption fluctuations and equipment operation cycle indicators.
4. A multi-dimensional dynamic carbon emission factor modeling method according to claim 1, characterized in that Establish an association matrix between the key feature subset and carbon emission factors, extract cross-dimensional main factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional main factor, and generate hierarchical weight parameters, including: Calculate the Spearman rank correlation coefficient between the features of each dimension in the key feature subset and carbon emission factors, and screen the features with the absolute value of the Spearman rank correlation coefficient greater than the preset correlation threshold to generate an initial association matrix; Perform eigenvalue decomposition on the initial association matrix, and extract the principal components with a cumulative contribution rate exceeding the preset contribution threshold as cross-dimensional main factors; Adopt the kernel density estimation method to fit the joint probability distribution of the cross-dimensional main factor and the carbon emission factor, and calculate the non-linear dependence strength between the two based on the mutual information entropy formula; Calculate the weight allocation ratio based on the entropy weight method combined with the non-linear dependence strength of each cross-dimensional main factor to generate hierarchical weight parameters.
5. A multi-dimensional dynamic carbon emission factor modeling method according to claim 1, characterized in that Construct a feature importance network based on gradient masking based on hierarchical weight parameters, optimize the feature distribution of the multi-dimensional data set through feature channel scaling, and generate an initial prediction model, including: Take the hierarchical weight parameters as the initial feature importance weights, dynamically update the weight gradients of each feature channel through gradient backpropagation, and generate a dynamic feature importance mask; Perform a channel-wise multiplication of the dynamic feature importance mask with the multi-dimensional dataset to suppress low-importance feature noise and enhance key feature signals, generating an optimized multi-dimensional feature distribution; Based on the optimized multi-dimensional feature distribution, select the gradient boosting tree model architecture, and use cross-validation and Bayesian optimization to adjust the model hyperparameters to generate an initial prediction model.
6. The multi-dimensional dynamic carbon emission factor modeling method according to claim 1, characterized in that Construct dedicated sub-models for multiple preset dimensions, and adopt a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor, including: Based on the optimized multi-dimensional feature distribution, train gradient boosting tree sub-models corresponding to the fuel property dimension, regional feature dimension, and process technology dimension respectively; Calculate the dynamic weights according to the historical prediction errors of each gradient boosting tree sub-model, and perform a weighted sum of the prediction results of the gradient boosting tree sub-models according to the dynamic weights; Normalize the weighted sum result and output the multi-scenario fusion carbon emission factor.
7. A multi-dimensional dynamic carbon emission factor modeling method according to claim 1, characterized in that Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, and transfer historical knowledge to the dedicated sub-models and filter the input data noise in combination with transfer learning, including: Input the newly added monitoring data into the initial prediction model, and incrementally update the model parameters through an online sequential extreme learning machine to generate an updated initial prediction model and historical verification data; Extract the weight distribution characteristics of the updated initial prediction model and adapt them to the weight space of the gradient boosting tree sub-model through a feature mapping function; Detect outliers in the newly added monitoring data based on the isolation forest algorithm, and input the data after removing the outliers into the gradient boosting tree sub-model.
8. A multi-dimensional dynamic carbon emission factor modeling method according to claim 1, characterized in that Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor according to the historical verification data and output the dynamic optimization evaluation result, including: Generate the probability distribution of the multi-scenario fusion carbon emission factor based on Monte Carlo simulation, and calculate the confidence interval with the confidence level being the preset threshold; Adjust the upper and lower bounds thresholds of the confidence interval in combination with the preset rules to generate the dynamic optimization evaluation result.
9. A multi-dimensional dynamic carbon emission factor modeling method according to claim 8, characterized in that The execution steps of the Monte Carlo simulation include: randomly sampling from the historical verification data to generate a simulation dataset, and iteratively calculating the distribution statistics of the multi-scenario fusion carbon emission factor; the preset rules include the statistical laws of historical data and external constraint conditions.
10. A multi-dimensional dynamic carbon emission factor modeling system for implementing a multi-dimensional dynamic carbon emission factor modeling method according to any one of claims 1-9, characterized in that, Including: Multi-dimensional acquisition module: Collect multi-dimensional monitoring data to construct a multi-dimensional dataset, and based on statistical tests, screen out the key feature subsets that are significantly correlated with the carbon emission factor; Cross-dimensional analysis module: Establish the correlation matrix between the key feature subsets and the carbon emission factor, extract the cross-dimensional principal factors through principal component analysis, quantify the non-linear dependence strength of each cross-dimensional principal factor, and generate hierarchical weight parameters; Weight optimization module: Construct a feature importance network based on gradient masks based on the hierarchical weight parameters, optimize the feature distribution of the multi-dimensional dataset through feature channel scaling, and generate an initial prediction model; Dynamic fusion module: Construct dedicated sub-models for multiple preset dimensions, and adopt a dynamic weight allocation mechanism to fuse the output results of the dedicated sub-models to generate a multi-scenario fusion carbon emission factor; Incremental transfer module: Update the initial prediction model based on the incremental learning mechanism and generate historical verification data, and transfer historical knowledge to the dedicated sub-models and filter the input data noise in combination with transfer learning; Dynamic calibration module: Calibrate the uncertainty interval of the multi-scenario fusion carbon emission factor based on historical verification data and output the dynamically optimized evaluation result.
Citation Information
Patent Citations
Carbon emission measuring and calculating model, comparison evaluation method and application thereof
CN117521912A
Integrated circuit process parameter optimization method and system based on machine learning
CN119067028A
Carbon emission integrated prediction method considering driving factor fluctuation, equipment and medium
CN119514868A
Permafrost upper limit prediction method and device based on modal decomposition, terminal and storage medium
CN119885090A
Service proxy method and system based on Dores front-end node
CN119938335A
Cited By
Grape growth environment monitoring method and system based on artificial intelligence
CN120991963A
Model training method and device, energy carbon prediction method, equipment and system, and medium
CN120995108A
Energy and carbon emission comprehensive monitoring system based on micro-service architecture
CN121235715A