Machine learning-based industrial park carbon asset evaluation method and system
Patent Information
- Application Number
- CN202611153963.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]鉴于上述的分析,本发明实施例旨在提供一种基于机器学习的工业园区碳资产评估方法及系统,用以解决现有工业园区碳资产评估效率低、数据权重及碳资产等级阈值设定主观及行业适应性不足的技术问题
[0017]与现有技术相比,本发明至少可实现如下有益效果之一:
Smart Images

Figure CN122819682A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of machine learning, artificial intelligence, industrial data analysis and mining, and in particular to a method and system for assessing carbon assets in industrial parks based on machine learning. Background Technology
[0002] Carbon asset assessment of industrial parks is a crucial foundational task supporting carbon quota management, screening for low-carbon technology upgrades, and evaluation of the effectiveness of green transformation. With the inclusion of typical industrial parks such as steel, coal chemical, electrolytic aluminum, new energy vehicles, and data centers within the scope of carbon management, the assessment targets exhibit characteristics such as diverse industry types, complex processes, and significant differences in energy structures. Scientifically, objectively, and efficiently assessing the current status of carbon assets in industrial parks and proactively judging their emission reduction potential is of great significance for guiding differentiated carbon reduction pathway planning for these parks.
[0003] Currently, existing technical solutions for assessing carbon assets or carbon emissions in industrial parks mainly fall into two categories: The first type is an evaluation scheme based on comprehensive weighting and multi-attribute decision-making. This type of scheme first constructs an indicator system that includes dimensions such as carbon emission intensity, energy consumption structure, and resource utilization efficiency. Then, it uses the analytic hierarchy process, entropy weighting method, or combined weighting method to determine the weight of the indicators. Finally, it uses grey relational analysis or fuzzy comprehensive evaluation method to calculate the comprehensive score of each park. The second type is emission accounting and benchmarking schemes based on historical data backtesting. These schemes focus on carbon emission data verification after the end of a specific compliance period. By calculating absolute or relative indicators such as emissions per unit of output and total energy consumption, the actual situation of the industrial park is compared with the preset industry benchmark value to determine whether the carbon emission level of the industrial park meets the standards.
[0004] However, existing assessment models are computationally inefficient. Each new sample requires repeated global matrix normalization and distance iteration calculations, causing computation time to increase dramatically with the sample size. This makes it impossible to conduct rapid batch assessments of a large number of industrial parks and to provide real-time feedback within seconds. Furthermore, carbon asset rating thresholds rely on expert experience or fixed quantiles, which are highly subjective and lack objective data-driven extraction methods, making dynamic calibration difficult as industrial park samples accumulate.
[0005] Therefore, there is an urgent need in this field for a standardized and rapid assessment method that can significantly improve assessment efficiency, enable rapid prediction of new samples, and have the ability to objectively classify carbon asset thresholds. Summary of the Invention
[0006] Based on the above analysis, the embodiments of the present invention aim to provide a machine learning-based method and system for assessing carbon assets in industrial parks, in order to solve the technical problems of low efficiency, subjective setting of data weights and carbon asset level thresholds, and insufficient industry adaptability in existing industrial park carbon asset assessments.
[0007] The objective of this invention is mainly achieved through the following technical solutions: This invention provides a machine learning-based method for assessing carbon assets in industrial parks, comprising the following steps: Carbon asset data of various types of industrial parks were acquired and preprocessed to obtain standardized data for each type of industrial park. Determine the objective weights of each dimension of the standardized data; based on the standardized data and the corresponding objective weights of each dimension, calculate the comprehensive score of the carbon asset status of each industrial park; Using carbon asset data from various types of industrial parks as sample data and the corresponding comprehensive carbon asset status score as sample label, we construct training sample sets for each type; and use the training sample sets for each type to train corresponding machine learning proxy models. The carbon asset data of the industrial park to be evaluated is input into the corresponding type of machine learning proxy model to obtain the corresponding comprehensive score of the current status of carbon assets; the corresponding comprehensive carbon asset level is obtained based on the comprehensive score of the current status of carbon assets.
[0008] Furthermore, a projection pursuit method combined with a genetic algorithm is used to objectively assign weights to the data in each dimension of the standardized data, determining the objective weights of the data in each dimension of the standardized data, including: Projecting the data of each dimension in the standardized data onto a preset unit projection direction, we obtain the one-dimensional projection value of each dimension in the standardized data of each industrial park in the projection direction. A projection objective function is constructed based on the one-dimensional projection values. The objective function includes intra-class spacing and intra-class density. The intra-class spacing reflects the overall dispersion of the one-dimensional projection values of each dimension of data, and the intra-class density reflects the local clustering of the one-dimensional projection values of each dimension of data. Using the projection objective function as the fitness function, a genetic algorithm is used to iteratively optimize the projection direction by performing selection, crossover, and mutation operations to maximize the projection objective function. The components of the optimal projection direction are normalized and used as objective weights for each dimension of the standardized data.
[0009] Furthermore, based on the standardized data and the corresponding objective weights of each dimension, the comprehensive score of the carbon asset status of each industrial park is calculated using the approximation of the ideal solution ranking method, including: A weighted decision matrix is constructed by combining the standardized data with the objective weights of the corresponding dimensions. Based on the weighted decision matrix, the positive ideal solution and the negative ideal solution of each carbon asset data are determined; wherein, the positive ideal solution is composed of the maximum value of each dimension of each carbon asset data in all industrial parks, and the negative ideal solution is composed of the minimum value of each dimension of each carbon asset data in all industrial parks. Calculate the first Euclidean distance from the standardized data of each industrial park to the positive ideal solution, and the second Euclidean distance to the negative ideal solution; Based on the first Euclidean distance and the second Euclidean distance, calculate the comprehensive score of the current status of carbon assets for each industrial park.
[0010] Furthermore, using carbon asset data from various types of industrial parks as sample data and the corresponding comprehensive carbon asset status scores as sample labels, training sample sets for each type are constructed, including: The carbon asset data of each type of industrial park is paired with the corresponding comprehensive score of the current status of carbon assets to form an initial training sample set for the corresponding type. The Monte Carlo simulation method is used to expand the initial training sample set. Based on the mean and standard deviation of the carbon asset data of each type of industrial park, simulated carbon asset data of each type of industrial park covering the parameter space are generated. The objective weights of each dimension of the simulated carbon asset data for each type of industrial park are determined, and the corresponding comprehensive score of the simulated carbon asset status is calculated. Based on the simulated carbon asset data for each type and the corresponding comprehensive score of the simulated carbon asset status, a simulation training sample set for each type is obtained. The initial training sample sets of each type and the corresponding simulated training sample sets are merged to obtain the training sample sets of the corresponding types.
[0011] Furthermore, using various types of training sample sets, random forest models and backpropagation neural network models are trained respectively to obtain corresponding types of machine learning surrogate models, including: The training sample sets of each type are divided into training subsets and test subsets of each type according to a preset ratio; Using training subsets of various types, we train random forest models and BP neural network models respectively to obtain initial random forest models and initial BP neural network models of the corresponding types; Using test subsets of various types, the performance of the corresponding initial random forest model and initial BP neural network model is evaluated, and the mean squared error, mean absolute error and coefficient of determination are calculated. Compare the determination coefficients of the initial random forest model and the initial BP neural network model of the corresponding type, and combine them with the mean squared error and the mean absolute error; select the initial random forest model or the initial BP neural network model with the higher determination coefficient and the smaller mean squared error and mean absolute error as the corresponding type of machine learning surrogate model; When the comparison results of the coefficient of determination are inconsistent with the comparison results of the mean square error and the mean absolute error, the coefficient of determination shall prevail.
[0012] Furthermore, based on the comprehensive score of the current status of carbon assets, the corresponding comprehensive carbon asset rating is obtained, including: Find the optimal cut-off point for the overall score of the current status of carbon assets corresponding to various carbon asset data within the same type of industrial park. Based on the comprehensive score of the current status of carbon assets corresponding to the optimal segmentation point, this type of industrial park is divided into low-level and high-level groups to obtain the second-level threshold. ; The optimal split point determination method is recursively applied to independently divide the low-level group and the high-level group, and the first-level threshold is obtained for each group. Second-level threshold ; ; Based on the first, second, and third grading thresholds, the overall score range of the current status of carbon assets is divided into four grading ranges; Based on the overall score of the carbon asset status of the industrial park to be evaluated falling into the corresponding grade range, the corresponding overall carbon asset grade is obtained.
[0013] Furthermore, using the Bootstrap resampling method, the optimal split point for the overall carbon asset status score corresponding to various carbon asset data within the same type of industrial park is obtained, including: From the comprehensive carbon asset status scores corresponding to all current carbon asset data of the same type of industrial park, samples with replacement are drawn with the same number of samples as the current sample to form a resampling sample set. The carbon asset status scores in the resampled sample set are sorted in ascending order, the difference between adjacent scores is calculated, and the position corresponding to the largest difference is taken as the current resampled segmentation position. Repeat the above steps to reach the preset number of times to obtain the corresponding segmentation position and obtain the comprehensive score of the carbon asset status corresponding to the segmentation position. The arithmetic mean of the overall carbon asset status scores corresponding to all split points is used as the optimal split point.
[0014] Furthermore, the types of industrial parks include, but are not limited to, steel, coal chemical, electrolytic aluminum, new energy vehicle, and data center industrial parks.
[0015] Furthermore, the carbon asset data includes general data and specific data for each type of industrial park; The general data includes carbon quota coverage rate, carbon quota surplus rate per unit output value, CCER development coefficient, CCER gap hedging rate, carbon quota holding value, carbon quota value volatility, CCER holding value, CCER investment return index, and carbon quota trading turnover rate. The distinctive data of steel industrial parks include the contribution rate of source-storage synergy for green electricity consumption and the revenue coefficient of waste heat and waste energy recovery; Key data for coal chemical industrial parks include the carbon capture contribution rate of CCUS projects, the intensity of wind-solar-storage synergistic emission reduction, and the CCUS benefit coefficient. Key data for electrolytic aluminum industrial parks include green electricity emission reduction intensity and peak-shaving carbon benefit coefficient; Key data for new energy vehicle industrial parks include the cascade utilization coefficient of power batteries and the carbon emission reduction rate of green electricity manufacturing. The key data for data center industrial parks includes PUE (Power Usage Effectiveness) carbon premium and load-aggregated carbon emission reductions.
[0016] The present invention also discloses a rapid assessment system for carbon assets in industrial parks based on machine learning. The system includes a data acquisition and preprocessing module M1, a comprehensive score calculation module M2, a proxy model training and optimization module M3, and an assessment and rating determination module M4. The data acquisition and preprocessing module M1 is used to acquire carbon asset data of multiple types of industrial parks and preprocess them to obtain standardized data corresponding to each type of industrial park. The comprehensive score calculation module M2 is used to determine the objective weights of each dimension of the standardized data; and to calculate the comprehensive score of the carbon asset status of each industrial park based on the standardized data and the corresponding objective weights of each dimension. The proxy model training and optimization module M3 is used to construct a corresponding type of training sample set by using carbon asset data of various types of industrial parks as sample data and the corresponding comprehensive score of carbon asset status as sample label; and to train the corresponding type of machine learning proxy model using the training sample sets of each type. The assessment and rating module M4 is used to input the carbon asset data of the industrial park to be assessed into the corresponding type of machine learning agent model to obtain the corresponding comprehensive score of the current status of carbon assets; and to obtain the corresponding comprehensive carbon asset rating based on the comprehensive score of the current status of carbon assets.
[0017] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: 1. This invention significantly improves the efficiency of carbon asset assessment in industrial parks, achieving real-time prediction within seconds. Compared to existing assessment models, which require repeated global matrix normalization and distance iteration calculations for each new industrial park carbon asset data sample, resulting in a sharp increase in computation time with the sample size and failing to achieve rapid batch assessment and real-time feedback, this invention constructs a machine learning proxy model (random forest model / BP neural network model selection) to complete the complex multi-attribute decision calculation process offline. During online assessment, new carbon asset data only needs to be preprocessed and input into the machine learning proxy model to output the comprehensive carbon asset score of the industrial park within seconds, greatly improving assessment efficiency and meeting the needs of large-scale batch carbon asset assessment and real-time interaction scenarios. 2. This invention avoids interference from subjective weighting of carbon asset data and collinearity of data indicators (highly overlapping or mutually extrapolating relationships), thus improving the objectivity of the scoring. Compared to traditional methods such as the analytic hierarchy process (AHP), which rely on expert experience to determine weights and are highly subjective, and where there are often complex nonlinear relationships between indicators such as carbon emission intensity and energy consumption structure, fixed weights are insufficient to reflect the inherent patterns in the data, this invention uses objective weighting to determine weights, avoiding human intervention. Simultaneously, the machine learning proxy model automatically mines the mapping relationship between carbon asset data and the comprehensive score of the current carbon asset status through a data-driven approach, without needing to pre-determine linear or functional relationships between indicators, effectively avoiding the problem of multicollinearity and making the carbon asset assessment results more objective and data-supported. 3. This invention uses a data-driven, dynamically defined carbon asset rating threshold to achieve adaptive calibration of carbon asset ratings. Existing carbon asset rating thresholds (such as A / B / C levels) typically rely on expert experience or fixed quantiles (such as trisections) for rough determination, which is highly subjective and cannot be dynamically updated with the accumulation of samples. When the overall level of the industrial park improves, the old thresholds may lose their discriminative power. This invention uses a data-driven method to adaptively classify the comprehensive carbon asset rating based on the comprehensive score of the industrial park's carbon asset status predicted by a machine learning proxy model. The rating threshold can be dynamically calibrated as the training data sample expands, ensuring that the carbon asset rating classification always matches the actual carbon asset level distribution of the park, thus improving the scientific rigor and adaptability of the carbon asset rating assessment standard. 4. This invention achieves a closed loop of "assessment-learning-prediction," possessing continuous evolution capabilities. In contrast, traditional solutions rely on one-time calculation models that lack learning capabilities and fail to provide feedback from new data for system upgrades. This invention constructs a complete closed loop from carbon asset data preprocessing, comprehensive carbon asset status score calculation, model training, to carbon asset level determination. As the accumulated industrial park data samples increase, the machine learning proxy model is periodically incrementally trained and the carbon asset level threshold is updated, enabling the assessment system to possess self-evolution and self-optimization capabilities, achieving continuous adaptation to the differentiated characteristics of different types of industrial parks.
[0018] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0019] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0020] Figure 1 This is a flowchart of a machine learning-based carbon asset assessment method for industrial parks, as described in an embodiment of the present invention. Figure 2 This is a schematic diagram of the random forest model generation in an embodiment of the present invention; Figure 3 This is a topology diagram of the BP neural network model in an embodiment of the present invention; Figure 4 This is a schematic diagram of the functional modules of a machine learning-based carbon asset assessment system for industrial parks, as described in an embodiment of the present invention. Detailed Implementation
[0021] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0022] This invention provides a machine learning-based method and system for assessing carbon assets in industrial parks, which combines current status assessment with potential prediction and integrates objective weighting with rapid prediction. It provides comprehensive and reliable technical support for low-carbon transformation path planning and carbon asset management optimization in typical high-energy-consuming industrial parks, and offers an efficient and credible technical solution for promoting the green and high-quality development of industrial parks.
[0023] Example 1: A specific embodiment of the present invention discloses a method for assessing carbon assets in industrial parks based on machine learning, such as... Figure 1 As shown, it includes the following steps: Step S1: Obtain carbon asset data for various types of industrial parks and preprocess them to obtain standardized data for each type of industrial park. Step S2: Determine the objective weights of each dimension of the standardized data; based on the standardized data and the corresponding objective weights of each dimension, calculate the comprehensive score of the carbon asset status of each industrial park. Step S3: Using carbon asset data of various types of industrial parks as sample data and the corresponding comprehensive carbon asset status score as sample label, construct training sample sets for each type; use the training sample sets for each type to train corresponding machine learning proxy models. Step S4: Input the carbon asset data of the industrial park to be evaluated into the corresponding type of machine learning agent model to obtain the corresponding comprehensive score of the current status of carbon assets; and obtain the corresponding comprehensive carbon asset level based on the comprehensive score of the current status of carbon assets.
[0024] Step S1 includes steps S11-S12.
[0025] Step S11: Obtain carbon asset data for various types of industrial parks.
[0026] Industrial park types include, but are not limited to, steel, coal chemical, electrolytic aluminum, new energy vehicle, and data center industrial parks.
[0027] To achieve standardized and horizontally comparable assessment of carbon assets across different types of industrial parks, general data and industry-specific data are considered as an organic whole. Following the principles of systematicity, comparability, and operability, and referencing the general frameworks in carbon emission accounting and carbon asset management, this system combines the technological characteristics and emission reduction technologies of typical parks such as steel, coal chemical, electrolytic aluminum, new energy vehicles, and data centers. The carbon asset data is divided into a complete hierarchical system comprising nine general data points across three criterion layers: carbon asset quantity, carbon asset value, and carbon asset management; and differentiated characteristic data for five typical industrial park types. This system enables unified measurement of carbon assets across different types of industrial parks, while accurately characterizing the differentiated contributions of key emission reductions in various industries through characteristic data, providing a standardized and comprehensive data foundation for subsequent carbon asset assessment modeling.
[0028] The carbon asset data includes general data and specific data for each type of industrial park; The general data includes carbon quota coverage rate, carbon quota surplus rate per unit output value, CCER development coefficient, CCER gap hedging rate, carbon quota holding value, carbon quota value volatility, CCER holding value, CCER investment return index, and carbon quota trading turnover rate. The distinctive data of steel industrial parks include the contribution rate of source-storage synergy for green electricity consumption and the revenue coefficient of waste heat and waste energy recovery; Key data for coal chemical industrial parks include the carbon capture contribution rate of CCUS projects, the intensity of wind-solar-storage synergistic emission reduction, and the CCUS benefit coefficient. Key data for electrolytic aluminum industrial parks include green electricity emission reduction intensity and peak-shaving carbon benefit coefficient; Key data for new energy vehicle industrial parks include the cascade utilization coefficient of power batteries and the carbon emission reduction rate of green electricity manufacturing. The key data for data center industrial parks includes PUE (Power Usage Effectiveness) carbon premium and load-aggregated carbon emission reductions.
[0029] To achieve standardized and horizontally comparable assessment of carbon assets in different types of industrial parks, this invention obtains a hierarchical carbon asset dataset, which includes two levels: general data and specialized data.
[0030] a. Obtain general data for all types of industrial parks.
[0031] The general data is applicable to all types of industrial parks, including three criteria layers: carbon asset quantity, carbon asset value, and carbon asset management, totaling nine core data points.
[0032] (1) Carbon asset quantity layer data, including carbon quota coverage rate, carbon quota surplus rate per unit output value, CCER development coefficient and CCER gap hedging rate.
[0033] ①S1: Carbon quota coverage rate, used to measure the extent to which the carbon quotas obtained by an industrial park cover the park's actual carbon emissions, reflecting the rationality of quota allocation. See below: Formula (1) in, This indicates the quota coverage factor within the compliance period; This indicates the amount of carbon allowances obtained by the industrial park during the accounting period, expressed in tons of carbon dioxide equivalent. ; This indicates the actual carbon emissions of the park during the accounting period, expressed in tons of carbon dioxide equivalent. .
[0034] ②S2: Carbon quota surplus rate per unit output value, used to measure the net quota surplus corresponding to every 10,000 yuan of total industrial output value generated in the industrial park. See below: Formula (2) in, The unit output value quota surplus rate is expressed in tons of CO2 per 10,000 yuan. The carbon quota surplus of an industrial park is the difference between the carbon quota obtained by the industrial park and the carbon quota actually used, expressed in tons of CO2. The amount of free carbon credits obtained by the industrial park; This represents the actual carbon emissions of the industrial park. The total annual output value of the industrial park is expressed in ten thousand yuan. It refers to the value of products produced in the year, including both sold and unsold products.
[0035] ③S3: CCER Development Factor, used to quantify the ability of an industrial park to offset its total carbon emissions by developing Certified Emission Reduction (CCER) projects. See below: Formula (3) in, This represents the CCER development coefficient, with a value range of [0, 1]. This represents the emission reductions of CCER projects developed and successfully registered by the park during the accounting period, expressed in tons of carbon dioxide equivalent. ; This represents the actual carbon emissions of the park within the same accounting period, expressed in tons of carbon dioxide equivalent. .
[0036] ④S4: CCER Gap Hedging Ratio, used to measure the extent to which the total amount of CCERs held by the park can make up for its net quota gap. See below: Formula (4) in, For CCER gap hedging ratio, This is due to insufficient quota.
[0037] (2) Carbon asset value layer data, including carbon quota holding value, carbon quota value volatility, CCER holding value and CCER investment return index.
[0038] ①J1: Carbon Allowance Position Value, used to measure the total value of the park's net surplus allowances calculated based on the current average carbon market transaction price. See below: Formula (5) in, The value of the quota holdings is in ten thousand yuan. The price represents the current market price of carbon allowances at the time of assessment, expressed in yuan per ton of carbon dioxide equivalent.
[0039] ②J2: Carbon allowance value volatility, used to measure the change in the value of carbon allowance assets due to market price fluctuations. See below: Formula (6) in, This indicates the volatility of the quota's market value; The value represents the current market price of carbon allowances at the assessment point in yuan per ton of carbon dioxide equivalent. This indicates the selected carbon quota benchmark price, expressed in yuan per ton of carbon dioxide equivalent, and is usually the average price for the same period of the previous compliance cycle.
[0040] ③J3: CCER Holding Value, used to measure the total monetary value of the CCERs held in the park, calculated based on the current average transaction price in the voluntary emission reduction market. See below: Formula (7) in, This represents the CCER holding value. This represents the average transaction price of CCER listings during the current period. In actual valuation, The average transaction price of the day is used.
[0041] ④J4: CCER Investment Return Index, used to assess the return level of specific carbon reduction projects within the park at the carbon asset level. See below: Formula (8) in, This represents the return on investment coefficient for carbon emission reduction projects; This refers to the total revenue actually obtained from carbon asset trading due to the implementation of the emission reduction project, expressed in RMB 10,000. It mainly includes the revenue from the sale of CCERs generated by the project in the market and the revenue from the sale of quota surplus caused by the emission reduction project in the market. This refers to the total investment cost invested in the emission reduction project, expressed in ten thousand yuan, including capital expenditures and project-specific operation and maintenance costs.
[0042] (3) Carbon asset management data, including carbon quota trading turnover rate.
[0043] L1: Carbon allowance trading turnover rate, used to measure the frequency with which the industrial park actively participates in carbon market trading. See below: Formula (9) in, Carbon quota trading turnover rate; This refers to the total annual trading volume of carbon allowances in the industrial park. The average annual carbon allowance holdings of the industrial park; The amount of carbon allowances purchased annually by the industrial park; The number of carbon credits sold by the industrial park in a given year; This represents the total carbon allowance holdings of the industrial park at the beginning of the year. This represents the total quota holdings of the industrial park at the end of the year.
[0044] a2. Obtain characteristic data for various types of industrial parks.
[0045] For five typical industrial parks—steel, coal chemical, electrolytic aluminum, new energy vehicles, and data centers—specific data was collected to represent the carbon asset contribution of their key emission reduction technologies.
[0046] (1) Obtain characteristic data of steel industrial parks.
[0047] ①G1: Contribution rate of green electricity consumption through source-storage synergy, used to measure the carbon emission reduction benefits of energy storage systems in steel industrial parks in improving the consumption rate of fluctuating renewable energy. See below: Formula (10) in, Contribution rate to the integrated utilization of energy sources and storage for green electricity; The amount of renewable energy consumed by the energy storage system during the accounting period specifically refers to the amount of photovoltaic or wind power that should have been discarded when the renewable energy power generation output exceeds the load of the steel industrial park. The data needs to be accurately measured through the park's energy management system (EMS). Carbon emission substitution factor (in tons) for corresponding energy types This refers to the carbon emission factor that can be avoided by recycling 1 unit of this energy source from purchased energy sources. The carbon emission factor is updated annually.
[0048] ②G2: Waste heat and energy recovery benefit coefficient, used to measure the carbon asset contribution of recovering secondary energy sources such as blast furnace gas, converter gas, and sintering waste heat. See below: Formula (11) in, The carbon gain coefficient for waste heat and waste energy recovery; The total amount of waste heat and energy recovered in the steel industrial park during the accounting period, such as the amount of blast furnace gas recovered, converter gas recovered, and sintering waste heat power generation, is obtained from the energy balance sheet or online monitoring system of the steel industrial park. The carbon emission substitution factor for the corresponding energy type, per ton This refers to the carbon emission coefficient of purchased energy that can be avoided by recycling 1 unit of this energy. This represents the total revenue of the industrial park during the same period.
[0049] (2) Obtain characteristic data of coal chemical industrial parks.
[0050] ①M1: Carbon Capture, Utilization, and Storage (CCUS) project contribution rate, used to measure the proportion of carbon dioxide actually captured by CCUS projects in a coal chemical industrial park to the total carbon emissions from the park's processes. See below: Formula (12) in, Contribution to carbon capture in the CCUS project; For the actual capture of CCUS projects during the accounting period Quantity, unit is tons This data must be obtained based on a continuous emissions monitoring system or material balance, including amounts used for geological storage or chemical utilization.
[0051] ②M2: Wind-Solar-Storage Synergistic Emission Reduction Intensity, used to measure the carbon emission reduction contribution achieved by replacing purchased thermal power through the coordinated operation of wind power, photovoltaic, and energy storage systems. See below: Formula (13) in, To reduce emissions intensity through the synergistic effect of wind, solar, and energy storage; This refers to the portion of the annual cumulative discharge from the energy storage system in the coal chemical industrial park that originates from wind and solar power generation, expressed in megawatt-hours (MWH).
[0052] ③M3: CCUS Benefit Factor, used to measure the efficiency with which CCUS projects convert captured carbon dioxide into carbon assets. See below: Formula (14) in, CCUS earnings coefficient; For the revenue of CCUS during the accounting period; It is to utilize The revenue from crude oil extracted by oil displacement; for Revenue from chemical products synthesized from raw materials (such as dimethyl carbonate, polyurethane, etc.); The proceeds from the sale of CCERs or other carbon credits in the carbon market; Financial support or cost reductions obtained for CCUS projects. The total cost of the CCUS project throughout its entire lifecycle; Depreciation and amortization costs for fixed asset investments (equipment and infrastructure for collection, compression, transportation, injection, etc.); Operation and maintenance costs include energy consumption (electricity, steam), adsorbent / solvent loss, labor, repairs, monitoring, etc.
[0053] (3) Special data of the electrolytic aluminum industrial park.
[0054] ①D1: Green Electricity Emission Reduction Intensity, used to measure the contribution of green electricity consumed per ton of electrolytic aluminum produced to emission reduction in the electrolytic aluminum industrial park. See below: Formula (15) in, Indicates the intensity of green electricity emission reduction; The total amount of green electricity consumed by the park during the accounting period is expressed in megawatt-hours (MWh). Data can be obtained from the park's electricity purchase account, green certificate transaction records, and distributed photovoltaic power generation statistics. The carbon emission factor corresponding to coal-fired power generation.
[0055] ②D2: Peak-shaving carbon benefit coefficient, used to measure the carbon emission reduction benefits obtained by electrolyzers in participating in grid demand response by utilizing the interruptible characteristics of loads, thereby replacing coal-fired power generation. See below: Formula (16) in, Carbon asset coefficient for peak shaving of interruptible loads; The actual electricity transferred by the industrial park for demand response or ancillary services during the accounting period is expressed in megawatt-hours (MWh). The data is sourced from the settlement statement of the power grid demand response platform or the virtual power plant dispatch record. The total annual output value of the electrolytic aluminum industrial park is expressed in ten thousand yuan. This represents the carbon emissions corresponding to the marginal coal-fired power plants that are being replaced.
[0056] (4) Special data of new energy vehicle industrial parks.
[0057] ①X1: The power battery cascade utilization coefficient, used to measure the carbon emission reduction generated by the cascade utilization of retired power batteries in new energy vehicle industrial parks, and its contribution to the conversion into tradable carbon assets. See below: Formula (17) in, It is the carbon asset coefficient for the secondary use of power batteries; The total capacity of retired power batteries used in the industrial parks of new energy enterprises during the accounting period is expressed in megawatt-hours. The data comes from the park's battery recycling ledger or the operation record of the secondary utilization project. The carbon emission factor of new battery production is measured in tons of carbon dioxide equivalent per megawatt-hour.
[0058] ②X2: Green Electricity Manufacturing Carbon Reduction Rate, used to measure the proportion of carbon emission reductions achieved through large-scale consumption of green electricity throughout the entire process of vehicle manufacturing and battery production, relative to the total emissions of the industrial park. See below: Formula (18) in, This refers to the green electricity consumed by the park through methods such as direct green electricity connection, factory photovoltaics, and green certificate trading.
[0059] (5) The unique data of the data center industrial park.
[0060] ①Z1: PUE (Power Usage Effectiveness) carbon premium, used to measure the proportion of additional carbon emissions from infrastructure energy consumption such as cooling and power distribution to the total emissions of the park. See below: Formula (19) in, It is the PUE (Power Usage Effectiveness) carbon premium coefficient, in percentage form; PUE is the average energy efficiency of the data center industrial park during the accounting period, which is equal to the total power consumption divided by the power consumption of IT equipment. The data comes from the EMS (Energy Management System) of the data center park. Total power consumption of IT equipment (MWh).
[0061] ②Z2: Load Aggregated Carbon Emission Reduction, used to measure the direct carbon emission reduction achieved by data centers acting as virtual power plants, aggregating internal loads to participate in peak shaving and valley filling on the power grid. See below: Formula (20) in, The carbon emission reduction coefficient is the load-polymerization carbon emission reduction factor. To determine how data centers actually reduce their electricity consumption from the grid through load shifting or energy storage discharge, the calculation needs to be based on a comparison between the baseline load and the actual load during the response period. The average carbon emission factor of the power grid; This represents the actual total electricity consumption during the data center's accounting period.
[0062] Step S12: Preprocess the carbon asset data of the acquired industrial parks of various types to obtain the corresponding standardized data of each type of industrial park.
[0063] The raw data values of various indicators of carbon asset data from multiple types of industrial parks are subjected to min-max normalization to eliminate the influence between different dimensions and obtain a standardized evaluation matrix.
[0064] The general and specific data obtained in step S11 will be standardized in terms of dimensions and scale, laying a calculable data foundation for subsequent objective weighting and comprehensive scoring of the current status of carbon assets.
[0065] The specific data for each type of industrial park is different. For example, steel industrial parks focus on waste heat and energy recovery, coal chemical industrial parks focus on CCUS, and data center industrial parks focus on PUE. They must be modeled separately for each type and cannot be treated the same.
[0066] Each type of industrial park implements a complete process from data preprocessing to carbon asset level threshold extraction.
[0067] This invention includes, but is not limited to, assessing carbon assets in five types of industrial parks: steel industrial parks, coal chemical industrial parks, electrolytic aluminum industrial parks, new energy vehicle industrial parks, and data center industrial parks. For one type of industrial park, assuming there are... A data sample of industrial parks to be evaluated, each sample containing n carbon asset assessment data (from step S11), the data indicator matrix is represented as follows. Carbon asset data is obtained through step S11.
[0068] This refers to the number of carbon asset data samples to be assessed for a certain type of industrial park (such as a steel industrial park), i.e., how many steel industrial parks are participating in the assessment, with one carbon asset data point per industrial park. This refers to the total number of data indicators used in this type of industrial park, including both general data and data unique to this type of park. For example, a steel industrial park. There are 9+2=11 in total, 9+3=12 in coal chemical industry, and 9+2=11 in data center. For the first The park is in the first The raw data value of a carbon asset, for example, the calculated value of the "waste heat and waste energy recovery revenue coefficient" for a certain steel industrial park is 0.35.
[0069] The same data indicator, such as the general data "carbon quota coverage rate", can be compared in the original data values of steel industrial parks and coal chemical industrial parks. However, the special data "waste heat and waste energy recovery revenue coefficient" is intended for comparison between steel industrial parks and cannot be compared across industries. This is the reason for the classification assessment.
[0070] The min-max normalization method was used for preprocessing. Before preprocessing the carbon asset data, the "good" direction for each data indicator was first clarified: is it better to be higher or larger, or better to be lower or smaller? As shown in Table 1.
[0071] Table 1: Positive and Negative Data
[0072] The positive data is preprocessed as follows: Formula (21) The first The original data value of a positive data point is mapped to the [0,1] interval by subtracting the minimum value of that data in the corresponding type of industrial park (characteristic data) or the minimum value of that data in all industrial parks (general data), and then dividing by the range (maximum value minus minimum value).
[0073] For positive data: The larger the value of this data indicator, the closer it is to 1 after standardization. The smaller the value of this data indicator, the closer it is to 0 after standardization.
[0074] Example: Suppose the "waste heat and waste energy recovery revenue coefficients" of the five steel industrial parks are 0.1, 0.25, 0.4, 0.55 and 0.7 respectively, then 0.1→0, 0.7→1, and 0.4→0.5.
[0075] The negative data is preprocessed as follows: Formula (22) The first The original data values of a positive data point are mapped to the [0,1] interval by taking the maximum value, subtracting that value, and then dividing by the range.
[0076] For negative data: The smaller the original value of the data (i.e., the better the performance of the park), the closer it is to 1 after standardization; The larger the original value of this data indicator (i.e. the worse the performance of the park), the closer it is to 0 after standardization.
[0077] Example: Suppose the PUE values of the 5 data centers are 1.1, 1.25, 1.4, 1.6 and 1.9 respectively, then 1.1→1 (best), 1.9→0 (worst), 1.4→0.56.
[0078] All carbon asset data are standardized to positive dimensionless values. For a given type of industrial park, the preprocessed standardized data matrix is denoted as... .
[0079] For formulas (21)-(22), if it is a standardization of general data, For the maximum and minimum values of this original data across all industrial park types; if it is standardization of characteristic data, These are the maximum and minimum values of the original data for the corresponding industrial park type.
[0080] Step S1 involves acquiring a hierarchical carbon asset dataset by obtaining general data from industrial parks and characteristic data from different types of industrial parks, and then preprocessing it to obtain corresponding standardized data. This provides a unified, comparable, and clearly structured data foundation for subsequent assessment and modeling.
[0081] Step S2 includes steps S21-S22.
[0082] The projection pursuit method is used to project high-dimensional index data into a one-dimensional space. With the goal of maximizing the projection objective function, the optimal projection direction is solved by the efficient global search capability of the genetic algorithm. The optimal projection direction vector is normalized and used as the objective weight of each dimension of data in the standardized data, thus avoiding the subjective bias of traditional weighting methods such as the analytic hierarchy process. Then, a weighted decision matrix is constructed to determine the positive and negative ideal solutions of each dimension of data. The Euclidean distance from the standardized data to the positive and negative ideal solutions is calculated. Finally, the comprehensive score of the carbon asset status of each industrial park is obtained. The score value is between 0 and 1, and the closer it is to 1, the better the carbon asset status.
[0083] The comprehensive score of the current status of carbon assets provides high-quality label data for subsequent machine learning proxy models, and also lays an objective ranking foundation for the extraction of carbon asset level thresholds.
[0084] Step S21: Determine the objective weights of the standardized data.
[0085] The projection pursuit method combined with a genetic algorithm is used to objectively assign weights to the data in each dimension of the standardized data, determining the objective weights of the data in each dimension of the standardized data, including: Projecting the data of each dimension in the standardized data onto a preset unit projection direction, we obtain the one-dimensional projection value of each dimension in the standardized data of each industrial park in the projection direction. A projection objective function is constructed based on the one-dimensional projection values. The objective function includes intra-class spacing and intra-class density. The intra-class spacing reflects the overall dispersion of the one-dimensional projection values of each dimension of data, and the intra-class density reflects the local clustering of the one-dimensional projection values of each dimension of data. Using the projection objective function as the fitness function, a genetic algorithm is used to iteratively optimize the projection direction by performing selection, crossover, and mutation operations to maximize the projection objective function. The components of the optimal projection direction are normalized and used as objective weights for each dimension of the standardized data.
[0086] To objectively determine the weights of each dimension of the standardized data and accurately assess the current status of carbon assets in industrial parks, a projection pursuit method combined with a genetic algorithm is used to objectively assign weights to each dimension of the standardized data based on the aforementioned general and characteristic data. Then, the TOPSIS (Technique for Order Preference by Similarity to Ideal Solution) method is applied to calculate the benchmark score for carbon asset status assessment. The objective weights of each dimension of the standardized data are obtained by normalizing the optimal projection direction of the maximizing projection index function, effectively avoiding subjective weighting bias and fully exploring the nonlinear information in high-dimensional data.
[0087] Existing subjective weighting methods, such as the analytic hierarchy process, rely on expert experience, while objective weighting methods, such as the entropy weighting method, only focus on the degree of data dispersion and ignore the interactive coupling relationship between indicators. Neither of these methods can fully and objectively reflect the true contribution of data from each dimension in the comprehensive evaluation.
[0088] The core issue this step aims to address is how to objectively extract the true contribution (i.e., weight) of each dimension of data to the comprehensive assessment of the park's carbon assets from a standardized data matrix using a data-driven approach.
[0089] The "projection pursuit method" combined with the "genetic algorithm" is used to objectively assign weights to the data in each dimension of the standardized data.
[0090] The basic idea of projection tracking is to visualize complex problems. High-dimensional data is difficult to observe and analyze directly. It is necessary to find an optimal perspective to project it into a one-dimensional space so that the inherent laws of the data (such as the clustering distribution of data samples, extreme value characteristics, etc.) can be most clearly revealed from this perspective.
[0091] Specifically, in the context of this invention: High-dimensional data: Each park's data sample includes multi-dimensional data (9 general data points + several unique data points); Projection direction: an n-dimensional unit vector that determines the angle from which the data is observed; One-dimensional projection value: The projection score of the standardized data of each park in the projection direction, which is the weighted sum of the standardized data values of each dimension of the industrial park; Optimal projection direction: The direction that makes the projected data sample points as dispersed as possible (large intra-class spacing) and the local clustering reasonable (moderate intra-class density).
[0092] The magnitude of each component in the optimal projection direction naturally reflects the relative importance of each data point in the evaluation of carbon assets in the five types of industrial parks, i.e., the objective weights sought.
[0093] First, project the n-dimensional data onto the unit direction. On the vector, where For the first The projection direction components corresponding to the dimensional data satisfy the following conditions: (Unit vector constraint), to obtain the first Standardized data of each industrial park in various dimensions under the current projection direction One-dimensional projection value As shown below: Formula (23) in, This refers to the standardized data after preprocessing in step S12.
[0094] In essence, it refers to the data from each dimension of standardized data. For a weighted sum of weights, when the projection direction is determined, The larger the value, the better the overall performance of the industrial park from the current perspective.
[0095] By traversing different projection directions, we can obtain the projection distribution of the same set of data under different viewpoints, and find the best viewpoint that makes the data features most prominent.
[0096] The objective function of the projection pursuit model is constructed as follows: Formula (24) in, It is the intra-class distance, reflecting the overall dispersion of standardized data samples. The larger the value, the more dispersed the sample points are after projection, and the easier it is to separate samples of different categories. Intra-class density reflects the local clustering characteristics of standardized data samples.
[0097] The physical meaning of formula (24) is to find an optimal projection direction so that the projection values of the standardized data samples of all industrial parks in a certain type in this direction are as dispersed as possible as possible in the whole (to facilitate the differentiation of good and bad) and maintain a reasonable clustering structure locally (to avoid outliers dominating).
[0098] Pursue overall dispersion (global distinguishability). The pursuit of local order (rationality of local clustering) is achieved by using the product of the two as the objective even function, which takes into account both global and local characteristics, so that the projection direction can distinguish between good and bad, and will not distort the overall evaluation structure due to individual outliers.
[0099] The calculation is as follows: Formula (25) in, This is the mean of the projected values of all standardized data samples for a given type of industrial park.
[0100] The larger the value, the more dispersed the distribution of standardized data sample points of each industrial park after projection, which means that from this perspective, parks at different levels can be better distinguished. If a certain projection direction makes all industrial parks of a certain type have roughly the same score ( If the value is very small, then this perspective is ineffective in distinguishing between superior and inferior products.
[0101] Intraclass density This refers to the distance between nearby points. The calculation is as follows: Formula (26) Formula (27) in, and All are standardized data sample indexes. , , Standardized data samples With standardized data samples The absolute distance in the projection direction. ; For high-density windows, the following is typically used: Its physical meaning is the threshold distance for determining whether two projected points are adjacent; if the projected distance between two data samples is less than 1, the threshold is set at 1. If so, it is considered a local cluster; For a unit step function, only when ,Right now Only when this condition is met will the term contribute a non-zero value.
[0102] It measures the combined density and uniformity of the projected samples within a local area, by accumulating all samples that satisfy the condition of a distance less than [a certain value]. The remaining distance between data sample pairs This reflects whether the local focusing of the projection point is reasonable.
[0103] Finally, a genetic algorithm was used to optimize the projection pursuit model to obtain the objective weights of each dimension of the standardized data.
[0104] Genetic algorithms are global optimization algorithms that simulate the natural evolutionary process. Their advantages include: ① Global search capability: Through population parallel search, it is not easy to get trapped in local optima; ② No gradient information required: Optimization can be achieved as long as the objective function value can be calculated, which is suitable for black-box optimization problems such as projection pursuit; ③ Strong robustness: No requirements are placed on the continuity and differentiability of the objective function.
[0105] Genetic algorithms simulate the "selection-crossover-mutation" mechanism in biological evolution to efficiently find the optimal projection direction in the search space.
[0106] The optimized model for the projection pursuit model is as follows: Formula (28) The projection direction vector; For the projection direction component; The main parameters of the genetic algorithm are set as follows, for example: population size 200, crossover probability 0.8, mutation probability 0.05, and number of generations 500.
[0107] by The fitness function is iteratively optimized through selection, crossover, and mutation operations until the fitness converges. The final optimal projection direction vector is obtained. , The optimal projection direction component is denoted by n, which represents the total number of data indicators, the sum of the number of general data and the number of special data.
[0108] Optimal projection direction vector The magnitude of each component objectively reflects the relative importance of the corresponding indicator in distinguishing the quality of carbon assets in the park.
[0109] Larger: indicates the first The greater the contribution of dimensional data to distinguishing the quality of carbon assets in a park, the greater its weight should be. Smaller: indicates the first The smaller the contribution of dimensional data indicators to distinguishing the quality of carbon assets in a park, the smaller their weight should be.
[0110] Will After normalization, it serves as the objective weight of each dimension of the standardized data. : Formula (29) This yields the objective weight vectors of each dimension of the standardized data. This refers to the combined weighting result based on the projection pursuit-genetic algorithm.
[0111] This step projects the high-dimensional index data into a one-dimensional space using projection pursuit. Maximizing the optimal projection direction is used as the selection criterion, and the global search capability of a genetic algorithm is utilized to solve for the optimal projection direction. This direction is then normalized and used as the objective weight for each indicator, achieving a completely data-driven, objective weighting process without any expert intervention. This weighting mechanism effectively avoids the subjective bias of traditional methods and fully exploits the nonlinear structural information in high-dimensional data, providing a scientific and objective weighting basis for the subsequent evaluation of the TOPSIS (Topology for Approximating Ideal Solutions) method.
[0112] Step S22: Based on standardized data and the objective weights of each dimension of the corresponding standardized data, calculate the comprehensive score of the current status of carbon assets of each industrial park.
[0113] Based on the standardized data and the corresponding objective weights of each dimension, the comprehensive score of the carbon asset status of each industrial park is calculated using the approximation of the ideal solution ranking method, including: A weighted decision matrix is constructed by combining the standardized data with the objective weights of the corresponding dimensions. Based on the weighted decision matrix, the positive ideal solution and the negative ideal solution of each carbon asset data are determined; wherein, the positive ideal solution is composed of the maximum value of each dimension of each carbon asset data in all industrial parks, and the negative ideal solution is composed of the minimum value of each carbon asset data in all industrial parks. Calculate the first Euclidean distance from each dimension of the standardized data of each industrial park to the positive ideal solution, and the second Euclidean distance to the negative ideal solution; Based on the first Euclidean distance and the second Euclidean distance, calculate the comprehensive score of the current status of carbon assets for each industrial park.
[0114] The Top-Ideal Solution Ranking Method (TOPSIS) was used to calculate the overall score of the current status of carbon assets in each park.
[0115] The basic logic of the TOPSIS method: Among multiple alternative solutions, we first construct the most ideal solution (the optimal values of each dimension in all standardized data) and the least ideal solution (the worst values of each dimension in all standardized data). The most ideal solution and the least ideal solution are two virtual solutions. Then we calculate the distance between each real solution and these two virtual solutions.
[0116] The optimal solution is the one that is closest to the most ideal solution and furthest from the least ideal solution.
[0117] The m industrial parks to be evaluated are the alternative options, and each standardized data point corresponds to n dimensions of data, which are the evaluation dimensions.
[0118] The technical problem to be solved in this step is: how to transform the standardized data matrix obtained in the previous steps. The results are integrated with the objective weight vector to obtain a quantitative score between 0 and 1 that is sortable and comparable, namely the comprehensive score of the current status of carbon assets, which serves as the final assessment result of the current status of carbon assets in each park.
[0119] (1) Construct the weighted decision matrix. Multiply the standardized data matrix by the weight vector, as shown below: Formula (30) in, For the first The park is in the first The weighted standardized value of each data indicator .
[0120] Matrix format: All parks Arranged as Weighted decision matrix .
[0121] By multiplying by the objective weights of each dimension of data This allows data indicators that are more important for carbon asset assessment to have a greater influence in subsequent distance calculations. This is a manifestation of objective weighting. The greater the weight of a data indicator, the greater its fluctuation will have an impact on the final score.
[0122] (2) Determine the positive and negative ideal solutions. The positive ideal solution takes the maximum value of each dimension of the carbon asset data, and the negative ideal solution takes the minimum value of each dimension of the carbon asset data. Formula (31) in, For the first The positive ideal value (optimal solution) of dimensional data. For the first The negative ideal value (worst solution) of the dimensional data index.
[0123] equal to the The maximum value in the column, i.e., the best performing indicator among all work parks; equal to the The minimum value in the column, that is, the worst performing park for this indicator among all parks.
[0124] Formula (32) It is a virtual, perfect industrial park, with every dimension of data reaching the highest level among similar industrial parks; It is a virtual worst industrial park, with each dimension of data at the lowest level among similar industrial parks.
[0125] (3) Calculate the Euclidean distance. Calculate the first... The Euclidean distances from each dimension of the standardized data to the positive and negative ideal solutions are shown below: Formula (33) in, For the first The smaller the value of the standardized data of an industrial park to the first Euclidean distance of the positive ideal solution, the closer the industrial park is to the perfect level. For the first The larger the value of the standardized data of an industrial park to the second Euclidean distance from the negative ideal solution, the closer the industrial park is to the worst-case scenario.
[0126] (4) Calculate the overall score of the current status of carbon assets in the industrial park. This is the comprehensive score of the industrial park's carbon asset status, with a value between 0 and 1, as shown below: Formula (34) in, For the first The overall TOPSIS score of each park, i.e., the score of the first park. The comprehensive score of the carbon asset status of each industrial park. This score serves as the final quantitative result of the carbon asset status assessment of the park, and also as the output label for subsequent machine learning proxy model training.
[0127] The closer the value is to 1, the closer the industrial park is to the positive ideal solution and the farther it is from the negative ideal solution, indicating that the overall carbon asset level is optimal. The closer a value is to 0, the further the industrial park is from the positive ideal solution and the closer it is to the negative ideal solution, indicating the worst overall carbon asset level.
[0128] Overall score of current carbon asset status It is a dimensionless relative index, limited to the range of 0-1, which facilitates horizontal comparison; Even if new or removed industrial parks are assessed, each industrial park... To maintain comparability, if there are significant changes in carbon asset data, the rankings and scores of each industrial park need to be recalculated. This is where subsequent steps involve introducing machine learning proxy models to learn this mapping relationship, thereby avoiding duplicate calculations.
[0129] Step S2 objectively determines the weights of each indicator through projection pursuit-genetic algorithm, and calculates the comprehensive score of the carbon asset status of each industrial park by combining the TOPSIS method, so as to provide high-quality supervision labels and ranking basis for subsequent machine learning agent models.
[0130] Step S3 includes steps S31-S32.
[0131] After step S2, the comprehensive score of the carbon asset status of each park can be accurately calculated. However, the TOPSIS method has technical flaws: ① High computational cost: Each time an industrial park to be evaluated is added or deleted, the entire carbon asset dataset needs to be re-executed with normalization, projection tracking and weighting, positive and negative ideal determination and distance iteration calculation. ② Poor scalability: When the number of industrial parks expands from dozens to hundreds, the computation time increases exponentially, which cannot meet the needs of rapid batch evaluation; ③ Inability to respond in real time: For newly established companies in the industrial park, it is impossible to provide assessment feedback within seconds.
[0132] To address the issues of high computational cost and inability to quickly evaluate new samples in the TOPSIS model under large-scale sample conditions, this invention constructs a machine learning proxy model.
[0133] Through machine learning, the system learns the mapping relationship from input carbon asset data to output comprehensive scores of the current status of carbon assets. After training, the evaluation of new carbon asset data samples no longer needs to execute TOPSIS, but can directly obtain predicted scores in seconds through machine learning proxy models.
[0134] Step S31: Using the carbon asset data as sample data and the corresponding comprehensive score of the current status of carbon assets as sample labels, construct a training sample set.
[0135] Using carbon asset data from various types of industrial parks as sample data, and the corresponding comprehensive carbon asset status score as sample label, a training sample set for each type is constructed, including: The carbon asset data of each type of industrial park is paired with the corresponding comprehensive score of the current status of carbon assets to form an initial training sample set for the corresponding type. The Monte Carlo simulation method is used to expand the initial training sample set. Based on the mean and standard deviation of the carbon asset data of each type of industrial park, simulated carbon asset data of each type of industrial park covering the parameter space are generated. The objective weights of each dimension of the simulated carbon asset data for each type of industrial park are determined, and the corresponding comprehensive score of the simulated carbon asset status is calculated. Based on the simulated carbon asset data for each type and the corresponding comprehensive score of the simulated carbon asset status, a simulation training sample set for each type is obtained. The initial training sample sets of each type and the corresponding simulated training sample sets are merged to obtain the training sample sets of the corresponding types.
[0136] Using the n-dimensional data determined in step S1 as the input carbon asset data sample X (raw data values of general data + characteristic data), and the TOPSIS score calculated in step S2. Y is used as the output for training labels.
[0137] Each training sample is Among them, data samples Sample Labels .
[0138] Initially, the data samples in the park are usually limited, for example, there are only 20-30 parks of a certain type. For machine learning models, especially neural networks, it is easy to overfit when the sample size is insufficient. That is, they perform very well on the training set, but their predictive ability is very poor on new data samples.
[0139] Due to the limited initial sample size, Monte Carlo simulation was used to expand the sample. For the initial carbon asset data of each type of industrial park, the mean was calculated based on the original sample of m parks. and standard deviation Assume each type of park has r initial data points, as shown below: Formula (35) in, Represented as the first The sample data in the th... The original data values of the original input variables are non-standardized values, such as carbon quotas, actual carbon emissions, total output value, CCER emission reductions, and energy storage consumption. For the first The sample mean of each original input variable; For the first The sample standard deviation of each original input variable is estimated using an unbiased estimate (denominator is m-1).
[0140] Sampling strategy: The raw carbon asset data, such as carbon allowances, actual carbon emissions, and total output, are assumed to follow a normal distribution. Perform normal distribution sampling.
[0141] Latin hypercube sampling: Technical issue: Simple random sampling may result in "clustering" in the parameter space, causing some regions to be densely packed with samples while others are sparse, affecting the model's coverage of the entire parameter space.
[0142] Solution: Use Latin Hypercube Sampling (LHS) to generate 1000 feature vectors.
[0143] Latin Hypercube Sampling (LHS) core principle: The distribution of each parameter is divided into N small intervals, where N is the number of samples. In each subinterval of each parameter, one and only one data sample is drawn; The sampled values of each parameter are randomly combined to form N data sample points.
[0144] Technical effect: Latin hypercube sampling (LHS) ensures that sample points uniformly cover the entire parameter space, avoiding the "clustering" problem of simple random sampling, and enabling the trained model to have good predictive ability in any region of the parameter space.
[0145] Tag generation: After preprocessing each set of simulated carbon asset data, the process in step S2 is called to calculate the comprehensive score of the current status of carbon assets. Construct a simulated dataset with sample labels.
[0146] Latin hypercube sampling is used to generate 1000 sets of feature vectors to ensure uniform sampling coverage of the parameter space. After preprocessing each set of simulated carbon asset data, step S2 is called to calculate the comprehensive score of the current status of carbon assets, thus constructing a labeled simulated training sample set.
[0147] The initial training sample sets of each type of industrial park and the simulated training sample sets of each type are merged to obtain the training sample sets of the corresponding type of industrial park.
[0148] The training sample sets of each type are divided into 80% training set and 20% test set by stratified sampling.
[0149] The training set is used to learn and train the model parameters; the test set is used for the final evaluation of the model performance.
[0150] Step S32: Using the training sample sets of various types of industrial parks, train the corresponding machine learning agent models.
[0151] For example, in this invention, five machine learning agent models will be trained for five different types of industrial parks. The different types of industrial parks will be trained independently using different machine models.
[0152] Using training sample sets of various types, we trained corresponding random forest models and BP neural network models respectively. We then selected the better-performing model between the random forest model and the BP neural network model as the machine learning proxy model for the corresponding type of industrial park.
[0153] Using various types of training sample sets, random forest models and backpropagation neural network models are trained respectively to obtain corresponding types of machine learning surrogate models, including: The training sample sets of each type are divided into training subsets and test subsets of each type according to a preset ratio; Using training subsets of various types, we train random forest models and BP neural network models respectively to obtain initial random forest models and initial BP neural network models of the corresponding types; Using test subsets of various types, the performance of the corresponding initial random forest model and initial BP neural network model is evaluated, and the mean squared error, mean absolute error and coefficient of determination are calculated. Compare the determination coefficients of the initial random forest model and the initial BP neural network model of the corresponding type, and combine them with the mean squared error and the mean absolute error; select the initial random forest model or the initial BP neural network model with the higher determination coefficient and the smaller mean squared error and mean absolute error as the corresponding type of machine learning surrogate model; When the comparison results of the coefficient of determination are inconsistent with the comparison results of the mean square error and the mean absolute error, the coefficient of determination shall prevail.
[0154] For example, the preset ratio is 8:2.
[0155] Training and selecting machine learning agent models: Two algorithm models, Random Forest (RF) and Backpropagation Neural Network (BPNN), were used respectively. The models were trained on the training sample sets of each type of industrial park, and the best one was selected as the evaluation standard tool after comparing their performance.
[0156] For each type of industrial park, a corresponding machine learning agent model is obtained; in this invention, five machine learning agent models of corresponding types are trained for five types of industrial parks.
[0157] The raw carbon asset data is used as input to the machine learning proxy model, and the TOPSIS evaluation value obtained in step S2 and the comprehensive score of the current status of carbon assets are used. As the output of the model (random forest model or BP neural network model), it shows the rapid transmission from actual production data of industrial parks to carbon asset assessment results.
[0158] (1) Train a random forest model for the corresponding type of industrial park 1) Random Forest Model Construction Random forest is an algorithm based on the Bagging ensemble learning framework. It improves the generalization ability and stability of the random forest model by constructing multiple independent decision trees and combining their predictions. The generation process of the random forest model is as follows: Figure 2 As shown.
[0159] The steps for building a random forest are as follows: The first step is Bootstrap sample generation. From the corresponding type of training subset, N samples are randomly selected with replacement to form a new training subset. This process is repeated. Second-rate( (The number of decision trees in the forest) The training subsets are independent of each other. The samples that are not selected are called "out-of-bag (OOB) data" and can be used for internal error evaluation.
[0160] For example, the original training subset contains 5 data samples. ; The first Bootstrap sampling, drawing 5 samples with replacement each time, might result in:
[0161] The second sampling may result in:
[0162] The third draw might result in: .
[0163] The size of each subset remains 5, but: Some samples, such as A, C, and E, were drawn repeatedly multiple times; Some samples, such as D, do not appear even once in a certain subset.
[0164] The second step is feature random selection. Traditional decision trees examine all features to find the optimal split point when splitting a node. Random forests, however, randomly select only one feature from all M samples (the original training subset + the training subset generated from the bootstrap samples) at each node of each decision tree during splitting. The feature subset is then used to find the optimal feature from the subset for splitting.
[0165] The third step is decision tree construction and ensemble. For each Bootstrap sample, a complete CART (Classification and Regression Tree) decision tree is grown without pruning. This process is repeated to generate... These decision trees together form a random forest.
[0166] For a new input sample, each tree will give a prediction value, and the final output of the random forest model is the arithmetic mean of the prediction values of all decision trees.
[0167] For the present invention, the parameter values are exemplarily as follows: Number of decision trees A sufficiently large number of trees can avoid overfitting and increase the model's prediction accuracy, but this will correspondingly increase computation time. For example, in this invention... When the number of trees is set to 500, the mean squared error of the random forest model tends to stabilize after the number of trees exceeds 500.
[0168] Number of candidate features for nodes Set as Where M is the total number of samples. For regression tasks, It can achieve good performance. This parameter controls the intensity of randomness: The smaller the value, the weaker the correlation between trees, but the accuracy of a single tree may decrease. The larger the tree, the higher the accuracy of a single tree, but the stronger the correlation. The optimal balance point between time deviation and variance.
[0169] Minimum number of leaf node samples Set to 5. This parameter controls the growth depth of the decision tree. The larger the tree, the shallower the tree, which increases model bias but decreases variance. The smaller the value, the deeper the tree, which may lead to overfitting. For medium-sized datasets, a value of 5 to 10 can achieve a good balance between bias and variance.
[0170] 2) Random Forest Model Training and Output On the training subset, the optimal hyperparameters determined by the above experiments are used ( , , Construct a random forest model. Complete Bootstrap sampling, decision tree growth, and ensemble.
[0171] After training, the random forest model outputs a predicted score for the test set or a new sample: Formula (36) in, It is carbon asset data from a test set or a new sample; It is the predicted value of the k-th decision tree. It is the final predicted score obtained through the random forest model, which is the comprehensive score of the current status of carbon assets.
[0172] The average of all votes cast by all trees preserves the predictive power of the number of individuals while reducing the random fluctuation of a single number through averaging.
[0173] (2) Train the BP neural network model for the corresponding type of industrial park 1) Construction of BP neural network model Backpropagation (BP) neural networks are multilayer feedforward neural networks trained using an error backpropagation algorithm. The signal propagates forward from the input layer through the hidden layers, generating predicted values. Then, the error between the predicted and true values is calculated, and the propagation proceeds backward along the same path, adjusting the connection weights and biases between neurons layer by layer.
[0174] The network structure includes an input layer, hidden layers, and an output layer, with fully connected neurons between adjacent layers. The number of nodes in the data layer equals the number of data features, and each node corresponds to one dimension of the carbon asset data. The hidden layer has 10 nodes and is a single hidden layer. The output layer has 1 node and outputs a comprehensive score of the current status of carbon assets.
[0175] The BP neural network model calculates the output through forward propagation and adjusts the weights layer by layer through back propagation to minimize the prediction error.
[0176] The topology diagram of the BP neural network model is as follows: Figure 3 As shown.
[0177] The first step is forward propagation. Input data is processed layer by layer from the input layer through the hidden layers, and finally passed to the output layer to generate the predicted value. The output of each layer serves as the input of the next layer, and nonlinearity is introduced into the neurons of each layer through activation functions.
[0178] The second step is error calculation. The predicted value of the comprehensive score of the carbon asset status of the output layer is compared with the actual value, and the loss function is calculated.
[0179] The third step is backpropagation. The error is backpropagated from the output layer to the hidden and input layers. Based on the gradient descent principle, the connection weights and biases are adjusted layer by layer to make the predicted value gradually approach the true value. This process is repeated until the loss function converges or the preset number of training rounds is reached.
[0180] For this invention, the parameter settings for using a BP neural network for prediction are as follows: Key parameters of the BP neural network model include the number of hidden layer nodes, learning rate, and momentum factor, which were determined experimentally.
[0181] Number of input layer nodes: equal to the number of input features.
[0182] Number of hidden layer nodes: A single hidden layer structure is used. The number of hidden layer nodes was determined through cross-validation experiments.
[0183] The specific method is as follows: On the training subset, perform 5-fold cross-validation on the number of candidate nodes (e.g., 5~14), calculate the mean squared error of each validation set, and select the number of nodes with the smallest error as the final value. Experiments have determined that the optimal number of hidden layer nodes is 10.
[0184] Output layer node count: The output is a comprehensive score of the current status of carbon assets, which is a continuous value, so the output layer node count is set to 1.
[0185] Hidden layer activation function: The hyperbolic tangent sigmoid function is used, with an output range of [-1, 1] and a zero-center property, which helps accelerate convergence. See below: Formula (37) Output layer activation function: A linear function is used to accommodate the need for continuous value output.
[0186] Formula (38) Training algorithm: Levenberg-Marquardt algorithm is used.
[0187] Learning rate: Initial value is set to 0.01. An adaptive adjustment strategy is used during training: if the loss function decreases in two consecutive iterations, the learning rate is multiplied by 1.05; if the loss function increases, the learning rate is multiplied by 0.7 and the weights are updated backward.
[0188] Momentum factor: Set to 0.9 to smooth the weight update trajectory and prevent getting trapped in local optima.
[0189] Early stopping mechanism: The training set is divided into an 80% training subset and a 20% validation subset. After each training round, the validation set error is calculated. If the validation set error does not decrease for 6 consecutive iterations, training is terminated and the system reverts to the weight state with the smallest validation set error to prevent overfitting.
[0190] Maximum number of training epochs: set to 1000. In actual training, the early stopping mechanism is usually triggered between 200 and 400 epochs.
[0191] Target error: set as Training will be terminated early when the training error falls below this value.
[0192] 2) BP neural network model training On the corresponding training subset, a BP neural network model is constructed using the structure and training parameters of the BP neural network model determined above.
[0193] The BP neural network model iteratively optimizes the weights through forward and backward propagation.
[0194] After training is complete, for the test subset or new samples The BP neural network model outputs a prediction score, as shown below:
[0195] Formula (39) in, This represents the final prediction score of the BP neural network model. This represents the number of hidden layer nodes. For the first Input carbon asset data values in each dimension; Let be the weights from the i-th node in the input layer to the j-th node in the hidden layer. Let be the bias of the j-th node in the hidden layer. tansig is the activation function for the hidden layer. Let J be the weight from the j-th node in the hidden layer to the output layer. For output layer bias, The activation function for the output layer is linear.
[0196] enter →Weighted summation → Tansig nonlinear transformation → Weighted summation again → Linear output score.
[0197] (3) Performance comparison and selection of random forest model and BP neural network model On a test subset of the corresponding type of industrial park (accounting for 20% of the training sample set, not used for training and validation), the following performance metrics were calculated for the random forest model and the BP neural network model, respectively: ①Mean Squared Error (MSE) Defined as the average of the squares of the differences between the predicted and actual values, it is calculated as follows: Formula (40) in, To test the number of samples in the subset, Let be the comprehensive score of the current status of carbon assets for the i-th carbon asset data sample. The model predicts a score.
[0198] The MSE ranges from [0, +∞). The smaller the MSE, the higher the prediction accuracy of the model. MSE assigns higher weights (square amplification) to larger errors, making it suitable for detecting situations with large prediction bias.
[0199] ②Mean Absolute Error (MAE) Defined as the average of the absolute values of the differences between the predicted and actual values, it is calculated as follows: Formula (41) The MAE value ranges from [0, +∞). The smaller the MAE, the smaller the average bias of the model's predictions. Compared to MSE, MAE is less sensitive to outliers and better reflects the model's average prediction error on most samples.
[0200] ③ Coefficient of Determination ) Defined as the proportion of variance explained by the model to the total variance (the variance of the true values), it is calculated as follows: Formula (42) in, This represents the average score of the overall status of real carbon assets in the test set.
[0201] The range of values for is ( [∞,1], generally the closer to 1, the better the model fit.
[0202] =1 indicates that the model makes a perfect prediction; =0 indicates that the model's prediction performance is the same as directly using the mean; <0 indicates that the model's prediction performance is worse than using the mean.
[0203] Select test subset A model with higher MSE and lower MAE is used as a proxy model for the final carbon asset status assessment. In case of conflict, the model with higher MSE and lower MAE shall prevail. As the standard.
[0204] This yields machine learning proxy models for five types of industrial parks.
[0205] Step S3 constructs a machine learning agent model for each type of industrial park, replacing the computationally expensive TOPSIS iteration process in Step S2 with second-level prediction, thereby achieving a rapid assessment of the comprehensive carbon asset score of the new sample.
[0206] Step S4, specifically.
[0207] The overall carbon asset rating is derived based on the comprehensive score of the current carbon asset status, including: Find the optimal cut-off point for the overall score of the current status of carbon assets corresponding to various carbon asset data within the same type of industrial park. Based on the comprehensive score of the current status of carbon assets corresponding to the optimal segmentation point, this type of industrial park is divided into low-level and high-level groups to obtain the second-level threshold. ; The optimal split point determination method is recursively applied to independently divide the low-level group and the high-level group, and the first-level threshold is obtained for each group. Second-level threshold ; ; Based on the first, second, and third grading thresholds, the overall score range of the current status of carbon assets is divided into four grading ranges; Based on the overall score of the carbon asset status of the industrial park to be evaluated falling into the corresponding grade range, the corresponding overall carbon asset grade is obtained.
[0208] Extracting grading thresholds and constructing carbon asset grading standards: This step is based on the carbon asset status assessment score (range 0~1) output by the machine learning proxy model in step S3, but this score itself lacks intuitive decision-making significance. By using an objective threshold extraction method, the park is divided into different carbon asset contribution levels, realizing the visualization output of the carbon asset contribution assessment results.
[0209] Extracting level thresholds using the maximum interval method based on Bootstrap: The overall carbon asset status scores of all carbon asset data samples within the same type of industrial park are sorted in ascending order, as shown below: Formula (43) Define the interval (first difference) between adjacent order statistics as: Formula (44) The maximum margin method posits that finding the maximum margin... Get the index of the maximum value This position represents the optimal dividing point between the low and high groups, and the corresponding comprehensive score of the current carbon asset status is the dividing threshold, as shown below: Formula (45) The corresponding TOPSIS current performance threshold is: Formula (46) Based on this threshold, all parks are divided into two subsets, with the lower subset being... The high group is .
[0210] Using the Bootstrap resampling method, the optimal split point for the overall carbon asset status score corresponding to all carbon asset data within the same type of industrial park was obtained, including: From the comprehensive carbon asset status scores corresponding to all current carbon asset data of the same type of industrial park, samples with replacement are drawn with the same number of samples as the current sample to form a resampling sample set. The carbon asset status scores in the resampled sample set are sorted in ascending order, the difference between adjacent scores is calculated, and the position corresponding to the largest difference is taken as the current resampled segmentation position. Repeat the above steps to reach the preset number of times to obtain the corresponding segmentation position, obtain the comprehensive score of the current status of carbon assets corresponding to the segmentation position, and calculate the arithmetic mean of the comprehensive scores of the current status of carbon assets corresponding to all segmentation positions as the optimal segmentation point.
[0211] To improve robustness, the Bootstrap resampling method is adopted: m samples are drawn with replacement from the original sample, and the process is repeated B=500 times. A threshold is calculated each time, and the final threshold is the arithmetic mean of the 500 results.
[0212] Construct a multi-level evaluation standard matrix: If it is necessary to divide into four levels (A, B, C, D), the maximum margin method can be repeatedly applied for recursive binary search.
[0213] Specifically, First binary search: The first binary search divides all samples into a "high" group and a "low" group to obtain the intermediate threshold. ; Second binary search: Perform maximum margin method independently on the high-value group again to obtain the highest threshold.
[0214] Third binary search: Perform a binary search independently on the low-value group to obtain the lowest threshold. ; Ultimately, three progressive thresholds were obtained. This corresponds to four level ranges.
[0215] The final result is four levels. The rules for classifying the levels are shown in Table 1: Table 1: Four-Quadrant Classification Standards for Carbon Asset Assessment
[0216] in, These are the three level thresholds recursively extracted using the Bootstrap maximum interval method.
[0217] Finally, based on real-time carbon asset data, a comprehensive carbon asset status score between 0 and 1 is quickly output through a machine learning proxy model, and the corresponding level label (A / B / C / D) is automatically matched according to the above threshold as the assessment conclusion.
[0218] This invention further proposes an objective threshold extraction method based on the Bootstrap maximum interval method, replacing the traditional approach that relies on expert experience or fixed quantile division. It constructs a four-quadrant evaluation standard matrix of A (excellent), B (good), C (average), and D (improved) through a recursive binary strategy. This method achieves dynamic calibration and adaptive updating of the threshold, continuously optimizing as park samples accumulate, significantly enhancing the objectivity and adaptability of the evaluation, and providing a clear and operable grading tool for park carbon quota management and green transformation effectiveness evaluation.
[0219] The carbon asset data of the industrial park to be evaluated is input into the machine learning proxy model of the corresponding type of industrial park to obtain the corresponding comprehensive score of the current status of carbon assets; the corresponding comprehensive carbon asset level is obtained based on the comprehensive score of the current status of carbon assets.
[0220] Step S4 objectively extracts the carbon asset level threshold using the Bootstrap maximum interval method, converting continuous scores into four levels: A / B / C / D, providing a standardized output for intuitive positioning and horizontal comparison of the park's carbon asset status.
[0221] Example 2: A specific embodiment of the present invention discloses a machine learning-based carbon asset assessment system for industrial parks, thereby implementing the machine learning-based carbon asset assessment method for industrial parks described in Embodiment 1. The specific implementation methods of each module are as described in the corresponding sections of Embodiment 1.
[0222] like Figure 4 As shown, the system includes a data acquisition and preprocessing module M1, a comprehensive score calculation module M2, a proxy model training and optimization module M3, and an evaluation and rating determination module M4. The data acquisition and preprocessing module M1 is used to acquire carbon asset data of multiple types of industrial parks and preprocess them to obtain standardized data corresponding to each type of industrial park. The comprehensive score calculation module M2 is used to determine the objective weights of each dimension of the standardized data; and to calculate the comprehensive score of the carbon asset status of each industrial park based on the standardized data and the corresponding objective weights of each dimension. The proxy model training and optimization module M3 is used to construct a corresponding type of training sample set by using carbon asset data of various types of industrial parks as sample data and the corresponding comprehensive score of carbon asset status as sample label; and to train the corresponding type of machine learning proxy model using the training sample sets of each type. The assessment and rating module M4 is used to input the carbon asset data of the industrial park to be assessed into the corresponding type of machine learning agent model to obtain the corresponding comprehensive score of the current status of carbon assets; and to obtain the corresponding comprehensive carbon asset rating based on the comprehensive score of the current status of carbon assets.
[0223] Since the system in this embodiment and the method in Embodiment 1 are related and can be referenced from each other, this description is redundant and will not be repeated here. Because this system embodiment shares the same principle as the above method embodiment, it also possesses the corresponding technical effects of the above method embodiment.
[0224] In summary, the machine learning-based carbon asset assessment method and system for industrial parks according to embodiments of the present invention have the following beneficial effects: 1. This invention significantly improves the efficiency of carbon asset assessment in industrial parks, achieving real-time prediction within seconds. Compared to existing assessment models, which require repeated global matrix normalization and distance iteration calculations for each new industrial park carbon asset data sample, resulting in a sharp increase in computation time with the sample size and failing to achieve rapid batch assessment and real-time feedback, this invention constructs a machine learning proxy model (random forest model / BP neural network model selection) to complete the complex multi-attribute decision calculation process offline. During online assessment, new carbon asset data only needs to be preprocessed and input into the machine learning proxy model to output the comprehensive carbon asset score of the industrial park within seconds, greatly improving assessment efficiency and meeting the needs of large-scale batch carbon asset assessment and real-time interaction scenarios. 2. This invention avoids interference from subjective weighting of carbon asset data and collinearity of data indicators (highly overlapping or mutually extrapolating relationships), thus improving the objectivity of the scoring. Compared to traditional methods such as the analytic hierarchy process (AHP), which rely on expert experience to determine weights and are highly subjective, and because indicators such as carbon emission intensity and energy consumption structure often have complex nonlinear relationships, fixed weights cannot reflect the inherent patterns of the data. This invention uses objective weighting to determine weights, avoiding human intervention. Simultaneously, the machine learning proxy model automatically mines the mapping relationship between carbon asset data and the comprehensive score of the current carbon asset status through a data-driven approach, without needing to pre-determine linear or functional relationships between indicators, effectively avoiding the problem of multicollinearity and making the carbon asset assessment results more objective and data-supported. 3. This invention uses a data-driven, dynamically defined carbon asset rating threshold to achieve adaptive calibration of carbon asset ratings. Existing carbon asset rating thresholds (such as A / B / C levels) typically rely on expert experience or fixed quantiles (such as trisections) for rough determination, which is highly subjective and cannot be dynamically updated with the accumulation of samples. Furthermore, when the overall level of the industrial park improves, the old thresholds may lose their discriminative power. This invention uses a data-driven method to adaptively classify the comprehensive carbon asset rating based on the comprehensive score of the industrial park's carbon asset status predicted by a machine learning proxy model. The rating threshold can be dynamically calibrated as the training data sample expands, ensuring that the carbon asset rating classification always matches the actual carbon asset level distribution of the park, thus improving the scientific rigor and adaptability of the carbon asset rating assessment standard. 4. This invention achieves a closed loop of "assessment-learning-prediction," possessing continuous evolution capabilities. In contrast, traditional solutions rely on one-time calculation models that lack learning capabilities and fail to provide feedback from new data for system upgrades. This invention constructs a complete closed loop from carbon asset data preprocessing, comprehensive carbon asset status score calculation, model training, to carbon asset level determination. As the accumulated industrial park data samples increase, the machine learning proxy model is periodically incrementally trained and the carbon asset level threshold is updated, enabling the assessment system to possess self-evolution and self-optimization capabilities, achieving continuous adaptation to the differentiated characteristics of different types of industrial parks.
[0225] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0226] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A machine learning-based method for assessing carbon assets in industrial parks, characterized in that, include: Carbon asset data of various types of industrial parks were acquired and preprocessed to obtain standardized data for each type of industrial park. Determine the objective weights of each dimension of the standardized data; based on the standardized data and the corresponding objective weights of each dimension, calculate the comprehensive score of the carbon asset status of each industrial park; Using carbon asset data from various types of industrial parks as sample data and the corresponding comprehensive carbon asset status score as sample label, we construct training sample sets for each type; and use the training sample sets for each type to train corresponding machine learning proxy models. The carbon asset data of the industrial park to be evaluated is input into the corresponding type of machine learning proxy model to obtain the corresponding comprehensive score of the current status of carbon assets; the corresponding comprehensive carbon asset level is obtained based on the comprehensive score of the current status of carbon assets.
2. The machine learning-based carbon asset assessment method for industrial parks according to claim 1, characterized in that, The projection pursuit method combined with a genetic algorithm is used to objectively assign weights to the data in each dimension of the standardized data, determining the objective weights of the data in each dimension of the standardized data, including: Projecting the data of each dimension in the standardized data onto a preset unit projection direction, we obtain the one-dimensional projection value of each dimension in the standardized data of each industrial park in the projection direction. A projection objective function is constructed based on the one-dimensional projection values. The objective function includes intra-class spacing and intra-class density. The intra-class spacing reflects the overall dispersion of the one-dimensional projection values of each dimension of data, and the intra-class density reflects the local clustering of the one-dimensional projection values of each dimension of data. Using the projection objective function as the fitness function, a genetic algorithm is used to iteratively optimize the projection direction by performing selection, crossover, and mutation operations to maximize the projection objective function. The components of the optimal projection direction are normalized and used as objective weights for each dimension of the standardized data.
3. The industrial park carbon asset assessment method based on machine learning according to claim 1, characterized in that, Based on the standardized data and the corresponding objective weights of each dimension, the comprehensive score of the carbon asset status of each industrial park is calculated using the approximation of the ideal solution ranking method, including: A weighted decision matrix is constructed by combining the standardized data with the objective weights of the corresponding dimensions. Based on the weighted decision matrix, the positive ideal solution and the negative ideal solution of each carbon asset data are determined; wherein, the positive ideal solution is composed of the maximum value of each dimension of each carbon asset data in all industrial parks, and the negative ideal solution is composed of the minimum value of each dimension of each carbon asset data in all industrial parks. Calculate the first Euclidean distance from the standardized data of each industrial park to the positive ideal solution, and the second Euclidean distance to the negative ideal solution; Based on the first Euclidean distance and the second Euclidean distance, calculate the comprehensive score of the current status of carbon assets for each industrial park.
4. The machine learning-based carbon asset assessment method for industrial parks according to claim 3, characterized in that, Using carbon asset data from various types of industrial parks as sample data, and the corresponding comprehensive carbon asset status score as sample label, a training sample set for each type is constructed, including: The carbon asset data of each type of industrial park is paired with the corresponding comprehensive score of the current status of carbon assets to form an initial training sample set for the corresponding type. The Monte Carlo simulation method is used to expand the initial training sample set. Based on the mean and standard deviation of the carbon asset data of each type of industrial park, simulated carbon asset data of each type of industrial park covering the parameter space are generated. The objective weights of each dimension of the simulated carbon asset data for each type of industrial park are determined, and the corresponding comprehensive score of the simulated carbon asset status is calculated. Based on the simulated carbon asset data for each type and the corresponding comprehensive score of the simulated carbon asset status, a simulation training sample set for each type is obtained. The initial training sample sets of each type and the corresponding simulated training sample sets are merged to obtain the training sample sets of the corresponding types.
5. The machine learning-based carbon asset assessment method for industrial parks according to claim 1, characterized in that, Using various types of training sample sets, random forest models and backpropagation neural network models are trained respectively to obtain corresponding types of machine learning surrogate models, including: The training sample sets of each type are divided into training subsets and test subsets of each type according to a preset ratio; Using training subsets of various types, we train random forest models and BP neural network models respectively to obtain initial random forest models and initial BP neural network models of the corresponding types; Using test subsets of various types, the performance of the corresponding initial random forest model and initial BP neural network model is evaluated, and the mean squared error, mean absolute error and coefficient of determination are calculated. Compare the determination coefficients of the initial random forest model and the initial BP neural network model of the corresponding type, and combine them with the mean squared error and the mean absolute error; select the initial random forest model or the initial BP neural network model with the higher determination coefficient and the smaller mean squared error and mean absolute error as the corresponding type of machine learning surrogate model; When the comparison results of the coefficient of determination are inconsistent with the comparison results of the mean square error and the mean absolute error, the coefficient of determination shall prevail.
6. The machine learning-based carbon asset assessment method for industrial parks according to claim 1, characterized in that, The overall carbon asset rating is derived based on the comprehensive score of the current carbon asset status, including: Find the optimal cut-off point for the overall score of the current status of carbon assets corresponding to various carbon asset data within the same type of industrial park. Based on the comprehensive score of the current status of carbon assets corresponding to the optimal segmentation point, this type of industrial park is divided into low-level and high-level groups to obtain the second-level threshold. ; The optimal split point determination method is recursively applied to independently divide the low-level group and the high-level group, and the first-level threshold is obtained for each group. Second-level threshold ; ; Based on the first, second, and third grading thresholds, the overall score range of the current status of carbon assets is divided into four grading ranges; Based on the overall score of the carbon asset status of the industrial park to be evaluated falling into the corresponding grade range, the corresponding overall carbon asset grade is obtained.
7. The machine learning-based carbon asset assessment method for industrial parks according to claim 6, characterized in that, Using the Bootstrap resampling method, the optimal split point for the overall carbon asset status score corresponding to various carbon asset data within the same type of industrial park was obtained, including: From the comprehensive carbon asset status scores corresponding to all current carbon asset data of the same type of industrial park, samples with replacement are drawn with the same number of samples as the current sample to form a resampling sample set. The carbon asset status scores in the resampled sample set are sorted in ascending order, the difference between adjacent scores is calculated, and the position corresponding to the largest difference is taken as the current resampled segmentation position. Repeat the above steps to reach the preset number of times to obtain the corresponding segmentation position and obtain the comprehensive score of the carbon asset status corresponding to the segmentation position. The arithmetic mean of the overall carbon asset status scores corresponding to all split points is used as the optimal split point.
8. The machine learning-based carbon asset assessment method for industrial parks according to any one of claims 1-7, characterized in that, Industrial park types include, but are not limited to, steel, coal chemical, electrolytic aluminum, new energy vehicle, and data center industrial parks.
9. The machine learning-based carbon asset assessment method for industrial parks according to claim 8, characterized in that, The carbon asset data includes general data and specific data for each type of industrial park; The general data includes carbon quota coverage rate, carbon quota surplus rate per unit output value, CCER development coefficient, CCER gap hedging rate, carbon quota holding value, carbon quota value volatility, CCER holding value, CCER investment return index, and carbon quota trading turnover rate. The distinctive data of steel industrial parks include the contribution rate of source-storage synergy for green electricity consumption and the revenue coefficient of waste heat and waste energy recovery; Key data for coal chemical industrial parks include the carbon capture contribution rate of CCUS projects, the intensity of wind-solar-storage synergistic emission reduction, and the CCUS benefit coefficient. Key data for electrolytic aluminum industrial parks include green electricity emission reduction intensity and peak-shaving carbon benefit coefficient; Key data for new energy vehicle industrial parks include the cascade utilization coefficient of power batteries and the carbon emission reduction rate of green electricity manufacturing. The key data for data center industrial parks includes PUE (Power Usage Effectiveness) carbon premium and load-aggregated carbon emission reductions.
10. A rapid carbon asset assessment system for industrial parks based on machine learning, characterized in that, The system includes a data acquisition and preprocessing module M1, a comprehensive score calculation module M2, a proxy model training and optimization module M3, and an evaluation and grade determination module M4. The data acquisition and preprocessing module M1 is used to acquire carbon asset data of multiple types of industrial parks and preprocess them to obtain standardized data corresponding to each type of industrial park. The comprehensive score calculation module M2 is used to determine the objective weights of each dimension of the standardized data; and to calculate the comprehensive score of the carbon asset status of each industrial park based on the standardized data and the corresponding objective weights of each dimension. The proxy model training and optimization module M3 is used to construct a corresponding type of training sample set by using carbon asset data of various types of industrial parks as sample data and the corresponding comprehensive score of carbon asset status as sample label; and to train the corresponding type of machine learning proxy model using the training sample sets of each type. The assessment and rating module M4 is used to input the carbon asset data of the industrial park to be assessed into the corresponding type of machine learning agent model to obtain the corresponding comprehensive score of the current status of carbon assets. The overall carbon asset rating is obtained based on the comprehensive score of the current status of carbon assets.