Carbon neutralization path dynamic planning method and system based on carbon asset matching

By constructing a four-dimensional asset matrix and a time-varying value evaluation model, combining geographical matching and timing matching, applying the improved Bellman equation and reinforcement learning model, dynamically adjusting the carbon neutrality path, the dynamic and synergistic problems of carbon neutrality path planning in the existing technology are solved, and efficient carbon asset management and emission reduction effects are achieved.

CN120494403APending Publication Date: 2025-08-15HONG KONG CHINA (SHENZHEN) CARBON ASSET OPERATION CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510621084.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing carbon neutrality path planning lacks dynamic nature and fails to effectively coordinate carbon asset trading and emission reduction measures, resulting in increased costs and poor emission reduction effects, and insufficient multi-target optimization, making it difficult to implement.

Method used

By constructing a four-dimensional asset matrix and a time-varying value evaluation model, combining supply-side geographical matching and demand-side timing matching, the improved Bellman equation and reinforcement learning model are applied, and the carbon neutrality paths are dynamically adjusted, and decision-making is optimized in real time.

Benefits of technology

The efficiency and utilization efficiency of carbon asset management have been achieved, costs have been reduced, the ability and adaptability of carbon neutrality paths have been ensured, and comprehensive cost control and emission reduction goals have been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494403A_ABST
    Figure CN120494403A_ABST
Patent Text Reader

Abstract

The invention provides a carbon neutralization path dynamic planning method and system based on carbon asset matching, and the method comprises the following steps: obtaining multi-source carbon asset data in an enterprise, a park or a region, and carrying out the preprocessing of the multi-source carbon asset data; the carbon asset value is dynamically evaluated by constructing the time-varying value evaluation model and comprehensively considering the change of factors such as policies, markets and technologies along with time, and meanwhile, various carbon assets are managed in a classified manner by constructing the four-dimensional asset matrix, so that the carbon asset management efficiency and benefit are remarkably improved; through supply side geographical matching and demand side time sequence matching, carbon asset supply and demand are accurately connected from space and time dimensions, cost and loss are reduced, an improved Bellman equation is applied, decision weight is dynamically adjusted under triple constraint conditions, carbon asset transaction and emission reduction measures are scientifically coordinated, and the utilization efficiency of carbon assets is improved; by constructing a reinforcement learning model, path planning is continuously optimized, and the scheme is ensured to be able to land and adapt to changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of carbon neutrality technology, and in particular to a carbon neutrality path dynamic planning method based on carbon asset matching and a system thereof. Background Art

[0002] Existing technologies have obvious shortcomings in carbon neutrality path planning: On the one hand, traditional carbon accounting models are mostly static and fail to fully consider the dynamic changes in carbon assets over time due to factors such as policy changes, market fluctuations, and technological advances. For example, when a company formulated an emission reduction plan based on a static carbon accounting model, it failed to foresee the significant fluctuations in the value of carbon quotas caused by policy adjustments, resulting in significant losses in the carbon trading market. On the other hand, most current carbon asset transactions and emission reduction measures are independent of each other and lack an effective coordination mechanism. This results in companies not fully considering the implementation of their own emission reduction measures when conducting carbon asset transactions, resulting in increased transaction costs and suboptimal emission reduction results. Furthermore, existing carbon neutrality pathway planning often focuses on a single objective, such as maximizing emissions reductions, while neglecting the coordinated optimization of multiple objectives, including cost, policy compliance, and technical feasibility. This makes the resulting pathway difficult to implement in practice. To this end, a carbon neutrality path dynamic planning method and system based on carbon asset matching are proposed. Summary of the Invention

[0003] In view of this, the embodiments of the present invention hope to provide a carbon neutrality path dynamic planning method and system based on carbon asset matching to solve or alleviate the technical problems existing in the existing technology and at least provide a beneficial option.

[0004] To solve the above technical problems, this application adopts a technical solution: a carbon neutrality path dynamic planning method based on carbon asset matching, comprising the following steps: Step 1: Obtain multi-source carbon asset data within an enterprise, park, or region, and pre-process the multi-source carbon asset data; Step 2: Based on the multi-source carbon asset data obtained, a four-dimensional asset matrix is constructed to classify and manage carbon assets, and a time-varying value assessment model is constructed to dynamically assess the value of carbon assets; Step 3: Based on the acquired multi-source carbon asset data, perform supply-side geographic matching and demand-side temporal matching, and generate matching results; Step 4: Combine the carbon asset value and matching results, apply the improved Bellman equation, and calculate and output the carbon neutrality path and implementation effect data under the triple constraints; Step 5: Build a reinforcement learning model based on multi-source carbon asset data and implementation effect data; Step 6: Establish a real-time monitoring system to monitor multi-source carbon asset data in real time, and dynamically adjust carbon asset values, matching results, and carbon neutrality paths based on the real-time monitoring data; Step 7. Establish evaluation indicators, regularly evaluate the implementation effect data of the carbon neutrality path, and optimize the parameter adjustment strategy of the reinforcement learning model based on the effect evaluation results.

[0005] As a further preferred embodiment of the present technical solution, in step 2, the four-dimensional asset matrix is constructed by using carbon emission reduction, carbon sink, carbon quota and green certificate as four dimensions; The time-varying value assessment model is: V(t) = f(policy intensity, market volatility, technology decay rate); Among them, V(t) represents the value of carbon assets at different time points t; f represents the mapping relationship between carbon asset value and policy intensity, market volatility, and technological decay rate; policy intensity is quantified based on the support or restriction of policies and regulations on carbon emission reduction; market volatility is measured based on the amplitude and frequency of changes in carbon asset market prices; and the technological decay rate is calculated based on the degree of reduction in the effectiveness of the relevant technology over time since its application.

[0006] As a further preferred embodiment of the present technical solution, in step 4, the improved Bellman equation is: minΣ[αC 减排 +(1−α)C 交易 ]; Among them, the adaptive weight factor α is dynamically adjusted according to the maturity, cost-effectiveness of emission reduction technology, the stability of carbon trading market prices and cost factors. 减排 represents the emission reduction cost, C 交易 represents the cost of carbon trading; The improved Bellman equation dynamically adjusts the emission reduction cost C according to the actual situation 减排 and transaction costs C 交易 The weight relationship between them; The triple constraints include policy compliance constraints, capital liquidity constraints and technical feasibility constraints.

[0007] As a further preferred embodiment of the present technical solution, in step 5, the method for constructing the reinforcement learning model includes the following steps: Step 501: Organize the multi-source carbon asset data and implementation effect data and divide them into training set, validation set and test set; Step 502: Build a reinforcement learning model based on the deep Q network architecture, and construct a neural network including an input layer, multiple hidden layers, and an output layer; Step 503: Use the training set to train the built reinforcement learning model, and use the adaptive learning rate adjustment algorithm to automatically adjust the learning rate according to the change of the loss function during the training process; Step 504: Evaluate and optimize the trained reinforcement learning model using verification; Step 505: Use the test set to test the optimized reinforcement learning model.

[0008] As a further preferred embodiment of the present technical solution, in step 3, the supply-side geographical matching is based on the emission source location and the carbon sink project location, by calculating the transportation distance between the two and using the formula D=1 / (1+transportation distance) to obtain the spatial matching degree D; Among them, D represents the spatial matching degree, which ranges from 0 to 1. The larger the value, the higher the degree of matching between the supply-side carbon assets and the demand-side in geographical space; "transportation distance" refers to the actual distance from the carbon sink project location to the emission source location; The demand side timing matching calculates the time difference Δt according to the time nodes set by short-term goals and long-term plans, and uses the formula T=e (-λ∣Δt∣) Calculate the time matching degree T; Among them, T represents the time matching degree, and its value range is between 0 and 1. The larger the value, the higher the temporal matching degree between the time demand on the demand side and the actual carbon asset supply. λ is a constant greater than 0, called the attenuation factor, which determines the degree of influence of the time difference on the time matching degree. Δt represents the time difference calculated based on the time nodes set for short-term goals and long-term plans. e is a natural constant.

[0009] As a further preferred embodiment of the present technical solution, in step one, the multi-source carbon asset data includes carbon emission reduction, carbon sink, carbon quota, green certificate, emission source location, surrounding carbon sink project location, policy and regulatory information, market price fluctuation data and technological development dynamics; the pre-processing of the multi-source carbon asset data includes data cleaning and data normalization.

[0010] As a further preferred embodiment of the present technical solution, in step seven, the evaluation indicators include carbon emission reduction target completion rate, cost control status and policy compliance.

[0011] To solve the above technical problems, another technical solution adopted in this application is: a carbon neutrality path dynamic planning system based on carbon asset matching, the system includes: a data acquisition and processing module, a carbon asset management and evaluation module, a supply and demand matching module, a path calculation module, a reinforcement learning model construction module, a real-time monitoring and adjustment module, and an effect evaluation and optimization module; The data acquisition and processing module is configured to obtain multi-source carbon asset data within an enterprise, park or region, and pre-process the multi-source carbon asset data; The carbon asset management and assessment module is configured to construct a four-dimensional asset matrix based on the acquired multi-source carbon asset data, classify and manage carbon assets, and construct a time-varying value assessment model to dynamically assess the value of carbon assets; The supply and demand matching module is configured to perform supply-side geographic matching and demand-side temporal matching based on the acquired multi-source carbon asset data, and generate a matching result; The path calculation module is configured to combine the carbon asset value and the matching results, apply the improved Bellman equation, and calculate and output the carbon neutrality path and implementation effect data under the triple constraint conditions; The reinforcement learning model construction module is configured to construct a reinforcement learning model based on multi-source carbon asset data and implementation effect data; The real-time monitoring and adjustment module is configured to establish a real-time monitoring system, monitor multi-source carbon asset data in real time, and dynamically adjust carbon asset values, matching results, and carbon neutrality paths based on the real-time monitoring data; The effect evaluation and optimization module is configured to formulate evaluation indicators, regularly evaluate the implementation effect data of the carbon neutrality path, and optimize the parameter adjustment strategy of the reinforcement learning model based on the effect evaluation results.

[0012] As a further preferred embodiment of the present technical solution, the system further includes a display and interaction module, which is used to present the calculation results and analysis reports in a visual manner and receive instructions and parameter information input by the user.

[0013] As a further preferred embodiment of the present technical solution, the data acquisition and processing module adopts a distributed data acquisition architecture to simultaneously acquire multi-source carbon asset data from multiple data sources in parallel.

[0014] The embodiment of the present invention adopts the above technical solution, which has the following advantages: 1. This invention constructs a time-varying value assessment model, comprehensively considering the changes in policy, market, and technology factors over time to dynamically assess the value of carbon assets. At the same time, by constructing a four-dimensional asset matrix, it systematically classifies and efficiently manages various types of carbon assets, helping enterprises optimize asset allocation, avoid value misjudgments and resource mismatches, and significantly improve the efficiency and benefits of carbon asset management. 2. This invention precisely matches carbon asset supply and demand from spatial and temporal dimensions through supply-side geographic matching and demand-side temporal matching, reducing costs and losses. It also uses an improved Bellman equation to dynamically adjust decision weights under triple constraints, scientifically coordinating carbon asset trading and emission reduction measures, improving carbon asset utilization efficiency, and achieving effective control of comprehensive costs. 3. This invention balances cost, compliance, and technical goals through triple constraints, while building a reinforcement learning model and combining it with real-time monitoring and dynamic adjustment mechanisms. It continuously optimizes path planning based on multi-source data and evaluation results to ensure that the plan can be implemented and adapt to changes, and efficiently promote the achievement of carbon neutrality goals.

[0015] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 This is a flow chart of a carbon neutrality path dynamic planning method based on carbon asset matching according to the present invention; Figure 2 A schematic diagram of the process of constructing a reinforcement learning model method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a carbon neutrality path dynamic planning system based on carbon asset matching in the present invention. DETAILED DESCRIPTION

[0018] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0019] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0020] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.

[0021] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0022] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.

[0023] Figure 1 This is a flow chart of a carbon neutrality path dynamic planning method based on carbon asset matching according to an embodiment of the present invention. It should be noted that if there are substantially the same results, the method of this application is not based on Figure 1 The process sequence shown is limited. Figure 1-Figure 2 As shown: A carbon neutrality path dynamic planning method based on carbon asset matching includes the following steps: Step 1: Obtain multi-source carbon asset data within an enterprise, park, or region, and pre-process the multi-source carbon asset data; Specifically, first, clarify the types of multi-source carbon asset data required, including core carbon asset data such as carbon emission reduction, carbon sink, carbon quota, and green certificate, as well as auxiliary data such as emission source location, surrounding carbon sink project location, policy and regulatory information, market price fluctuation data, and technology development dynamics; obtain carbon emission reduction data and other data from the company's internal energy management system and production record database; obtain carbon quota, policy and regulatory information through the websites of relevant government departments and industry databases; use the carbon trading market platform to obtain market price fluctuation data; and use scientific research institute reports and technical information platforms to collect technology development dynamics data; Then, data collection tools and techniques are used to acquire data from identified data sources. For structured data, such as data from an enterprise's energy management system, data is extracted through database query statements. For unstructured data, such as policy and regulatory documents, text mining techniques are used to extract information. Geographic data on emission source locations and carbon sink project locations can also be collected through sensor networks. Finally, the collected data is cleaned, duplicate data is identified and removed to avoid data redundancy affecting the analysis results; and the collected data is normalized to make the data more suitable for subsequent model analysis and calculations.

[0024] Step 2: Based on the multi-source carbon asset data obtained, a four-dimensional asset matrix is constructed to classify and manage carbon assets, and a time-varying value assessment model is constructed to dynamically assess the value of carbon assets; The specific implementation methods for constructing a four-dimensional asset matrix and classifying and managing carbon assets are as follows: First, based on the multi-source carbon asset data obtained, carbon emission reduction, carbon sink, carbon quota and green certificate are clearly defined as the four dimensions of the four-dimensional asset matrix. These four dimensions basically cover the main carbon asset types within an enterprise, park or region, and can fully reflect the status of carbon assets. Then, the multi-source carbon asset data is classified and organized according to four dimensions; for example, the carbon emission reduction data achieved by enterprises through energy-saving transformation and the use of clean energy are classified into the "carbon emission reduction" dimension; the carbon sink data corresponding to the carbon sink projects such as forests and wetlands owned by enterprises are classified into the "carbon sink" dimension; the carbon quota data allocated by the government or obtained by enterprises through transactions are classified into the "carbon quota" dimension; and the data related to various green certificates obtained by enterprises are classified into the "green certificate" dimension. Next, a matrix structure is constructed based on these four dimensions. The rows of the matrix can represent different time periods (such as monthly, quarterly, and annually), and the columns correspond to carbon emission reductions, carbon sinks, carbon quotas, and green certificates. Each matrix cell is filled with the specific carbon asset value of that dimension within the corresponding time period, thus forming a complete four-dimensional asset matrix. Finally, a four-dimensional asset matrix is used to classify and manage carbon assets. Through the matrix, we can clearly view the quantity, change trends and other information of various carbon assets at different time points; for example, we can analyze the growth of carbon emission reductions in different time periods to judge the effectiveness of corporate emission reduction measures; compare the allocation and use of carbon quotas, and rationally plan carbon quota trading strategies; evaluate the acquisition of green certificates, and provide a basis for companies to establish a good green image; The specific implementation methods for constructing a time-varying value assessment model and dynamically assessing the value of carbon assets are as follows: First, the main factors affecting the value of carbon assets are determined to be policy intensity, market volatility, and technology decay rate. Policy intensity can be quantified based on the support or restriction of carbon emission reduction by policies and regulations, such as the amount of carbon emission reduction subsidies granted by the government and the standards for levying carbon emission taxes. Market volatility is measured by the magnitude and frequency of changes in the market price of carbon assets, which can be obtained by analyzing historical price data of the carbon trading market. Technology decay rate is calculated based on the degree to which the effectiveness of the relevant technology decreases over time since its application. For example, the proportion of emission reduction efficiency of a certain emission reduction technology that gradually decreases with the increase of its service life. Then, a time-varying value assessment model V(t) = f(policy intensity, market volatility, technology decay rate) is constructed. Machine learning algorithms (such as neural networks, decision trees, etc.) or mathematical modeling methods (such as regression analysis) can be used to determine the specific form of the function f. By collecting historical data, using policy intensity, market volatility, and technology decay rate as input variables and carbon asset value as output variable, the model is trained and optimized to obtain a model that can accurately reflect the changes in carbon asset value over time. Finally, the constructed time-varying value assessment model is used in combination with real-time data on policy intensity, market volatility, and technological decay rate to dynamically assess the value of carbon assets at different time points t. For example, when policies change, such as when new carbon emission reduction incentive policies are introduced, the model will recalculate the value of carbon assets based on changes in policy intensity. When prices in the carbon trading market fluctuate significantly, this can also be promptly reflected in the assessment of carbon asset value. Through dynamic assessment, enterprises, parks, or regions can adjust their carbon asset management strategies in a timely manner to maximize the value of carbon assets.

[0025] Step 3: Based on the acquired multi-source carbon asset data, perform supply-side geographic matching and demand-side temporal matching, and generate matching results; The specific implementation methods of supply-side geographic matching are as follows: First, the emission source location and carbon sink project location information are extracted from multi-source carbon asset data. The emission source location is accurate to the geographical location of the company's specific production facilities, and the carbon sink project location is clear with its geographical coordinates, such as the longitude and latitude of forest carbon sink projects, or the regional range coordinates of wetland carbon sinks. Then, using Geographic Information System (GIS) technology or professional map calculation tools, calculate the transportation distance between each emission source and each carbon sink project. Distance calculation is accurately measured based on actual traffic routes, topography and other factors, such as road curvature and the presence of obstacles, to avoid errors caused by simple straight-line distance calculations; Next, the spatial matching degree D is calculated according to the formula D = 1 / (1 + transportation distance). When the transportation distance is short, the D value is close to 1, indicating that the supply-side carbon assets and the demand-side carbon assets are highly matched in geographical space. When the transportation distance is long, the D value approaches 0, indicating a low matching degree. For example, if the transportation distance between an emission source and a carbon sink project is 10 kilometers, D = 1 / (1+10) ≈ 0.091. Finally, the spatial matching degree between each emission source and different carbon sink projects is organized into a list or matrix form to clearly present the matching relationship between each emission source and carbon sink project, facilitating subsequent comprehensive analysis; The specific implementation of demand-side timing matching is as follows: First, determine the time points involved in carbon asset supply and demand based on the short-term goals and long-term plans of the enterprise, park, or region. For example, if the short-term goal is to reduce carbon emissions by a certain percentage in the next quarter, and the long-term plan is to achieve carbon neutrality in the next five years, the specific time points for carbon asset supply and demand can be determined accordingly. Then, for each point in time, the time difference Δt between the carbon asset demand and the potential carbon asset supply time is calculated. For example, if a company plans to purchase carbon sinks to offset carbon emissions in the first quarter of 2025, and a carbon sink project is expected to be mature and available for use in the fourth quarter of 2024, Δt = the first quarter of 2025 - the fourth quarter of 2024 = 1 quarter. Next, determine the attenuation factor λ. The attenuation factor λ is greater than 0 and determines the degree of influence of the time difference on the time matching degree. This can be determined based on historical data, industry experience, or through simulation analysis. λ varies in different industries and scenarios. For example, the energy industry has large fluctuations, so the λ value is high; stable industries have low values. Then, using the formula T=e (-λ∣Δt∣) Calculate the time matching degree T, which ranges from 0 to 1. The larger the value, the higher the temporal matching degree between the demand side time demand and the actual carbon asset supply. Assuming λ=0.2, Δt=1, then T=e (-0.2×1) ≈0.819; Finally, the temporal matching degree between carbon asset demand and supply at each time point is summarized to form a list of matching results in the time dimension, which is combined with the supply-side geographical matching results to provide comprehensive matching data for subsequent analysis; The generation of matching results is as follows: First, by integrating the spatial matching degree D of the supply-side geographic matching and the temporal matching degree T of the demand-side temporal matching, a new data structure, such as a two-dimensional array, can be constructed to integrate the matching degree of each emission source with each carbon sink project at different time points, thus clearly presenting the comprehensive matching of carbon assets in the spatial and temporal dimensions. Then, the comprehensive matching results are analyzed to screen out carbon asset supply and demand combinations with high spatial and temporal matching. A matching threshold can be set, such as combinations with spatial matching D>0.5 and temporal matching T>0.6 as priority considerations, providing favorable data support for subsequent carbon neutrality path planning.

[0026] Step 4: Combine the carbon asset value and matching results, apply the improved Bellman equation, and calculate and output the carbon neutrality path and implementation effect data under the triple constraints; Specifically, first, summarize the carbon asset value data obtained in step 2, including the value of carbon emission reductions, carbon sinks, carbon quotas, and green certificates at different time points; as well as the matching results in step 3, namely the spatial matching degree D of supply-side geographic matching and the temporal matching degree T of demand-side temporal matching. These data are then sorted according to emission sources, carbon sink projects, and time dimensions to ensure data consistency and relevance, providing a complete data foundation for subsequent calculations. Then, comprehensively considering the maturity, cost-effectiveness, and price stability and cost factors of emission reduction technologies, for emission reduction technologies with high maturity and good cost-effectiveness, and when prices in the carbon trading market fluctuate greatly and costs are high, the α value should be appropriately increased, indicating a preference for achieving carbon neutrality through emission reduction. Conversely, if emission reduction technologies are costly and immature, and the carbon trading market is stable and costs are low, the α value should be lowered, indicating a greater reliance on carbon trading. For example, if the emission reduction technology adopted by the enterprise has been widely used, the unit emission reduction cost is low, and prices in the carbon trading market fluctuate sharply, α can be set to 0.7. Next, based on the actual situation of the enterprise, calculate the emission reduction costs, such as the one-time investment and subsequent operating costs incurred by investing in new emission reduction equipment and improving production processes, and amortize them according to a certain calculation period (such as annually); the carbon trading cost is determined based on the market price and the number of transactions, taking into account additional costs such as transaction fees; for example, an enterprise plans to install new desulfurization equipment with a total investment of 1 million yuan, a service life of 10 years, an annual operating cost of 100,000 yuan, and an annual emission reduction of 50,000 tons. The calculated emission reduction cost per ton is C 减排 =30 yuan; if the price of carbon quota per ton in the carbon trading market is 50 yuan and the transaction fee rate is 2%, then the carbon transaction cost per ton is C 交易 is 51 yuan; Then, α, C 减排 and C 交易 Substitute into the improved Bellman equation minΣ[αC 减排+(1−α)C 交易 ], the costs of different emission source and carbon sink project combinations, as well as at different time points, are accumulated to obtain the total cost of each combination scheme; for example, for an enterprise with three emission sources and two carbon sink projects, the cost of each emission source and carbon sink project combination in each year is calculated separately during the next five-year planning period, and then accumulated to obtain the total cost of different combinations; Finally, among the options that meet the triple constraints of policy compliance constraint check, capital liquidity constraint assessment and technical feasibility constraint review, the option with the lowest total cost is selected as the optimal carbon neutrality path; the implementation effect data under this path is output, including total carbon emission reduction, total carbon trading volume, and total implementation cost; the matching status of each emission source and carbon sink project, such as which emission sources and which carbon sink project have achieved effective matching in space and time; and the contribution of this path to the company's carbon emission targets, such as the expected proportion of corporate carbon emissions reduction during the planning period, etc., to provide a clear reference basis for corporate decision-making and carbon neutrality path implementation.

[0027] Step 5: Build a reinforcement learning model based on multi-source carbon asset data and implementation effect data; Specifically, first, collect the multi-source carbon asset data within the enterprise, park or region obtained in step one, covering core data such as carbon emission reduction, carbon sink, carbon quota, green certificate, as well as auxiliary data such as emission source location, location of surrounding carbon sink projects, policy and regulatory information, market price fluctuation data, and technological development dynamics; at the same time, organize the implementation effect data of the carbon neutrality path output in step four, including total carbon emission reduction, total carbon trading volume, total implementation cost, the matching status of each emission source and carbon sink project, and the contribution of the path to the enterprise's carbon emission targets; Then, the multi-source carbon asset data and the carbon neutrality path implementation effect data are combined to form a state space; for example, the state can be represented as a vector S=[C reduction , C trading , C cost , D, T, P, M, …], where C reduction is the carbon emission reduction, C trading is the carbon trading volume, C cost is the total implementation cost, D is the spatial matching degree, T is the temporal matching degree, P is the policy and regulation related indicator, M is the market price fluctuation indicator, etc.; at the same time, define the actions that the model can take, which correspond to different carbon neutrality path adjustment strategies; for example, actions can include increasing or decreasing the investment in carbon emission reduction measures, adjusting the number and timing of carbon trading, selecting different carbon sink projects for matching, etc.; and design a reward function R (S, A) to evaluate the pros and cons of each action; the design of the reward function should take into account multiple factors, such as cost reduction, achievement of carbon emission reduction targets, policy compliance, etc.; for example, the reward function can be set as: R(S,A)=α1×ΔC reduction -α2×ΔC cost +α3×I policy ; Where, ΔC reduction is the increase in carbon emission reduction after taking action, ΔC cost is the change in cost, I policy is the policy compliance indicator (compliance is 1, non-compliance is -1), α1, α2, and α3 are the corresponding weight coefficients; Next, select a reinforcement learning algorithm, including policy-based algorithms and value-based algorithms. Policy-based algorithms, such as the Deep Deterministic Policy Gradient (DDPG) algorithm, are suitable for problems in continuous action spaces. The DDPG algorithm combines deep neural networks and policy gradient methods to directly output actions by learning a deterministic policy. Value-based algorithms, such as the Deep Q-Network (DQN) algorithm, are suitable for problems in discrete action spaces. The DQN algorithm estimates the value of each state-action pair by learning a Q function, thereby selecting the optimal action. Subsequently, the reinforcement learning model is trained. The specific training process is as follows: Randomly initialize the parameters of the reinforcement learning model, such as the weights and biases of the neural network; The model is in the initial state S0 and selects an action A0 according to the current strategy. After executing the action, the environment will feedback the next state S1 and the corresponding reward R0. The experience (S0, A0, R0, S1) obtained from each interaction is stored in the experience replay buffer. A batch of experience data is randomly sampled from the buffer periodically for training to break the correlation between the data and improve the stability of training. Based on the selected reinforcement learning algorithm, the model parameters are updated using the sampled empirical data; for example, in the DQN algorithm, the parameters of the Q network are updated by minimizing the error between the predicted Q value and the target Q value; Finally, the reinforcement learning model is evaluated and optimized as follows: Evaluate the trained reinforcement learning model using an independent test dataset. Evaluation metrics may include cumulative rewards, carbon reduction target achievement rate, cost reduction rate, etc. If the model evaluation results are not ideal, you can adjust the parameters of the reinforcement learning algorithm (such as learning rate, discount factor, etc.), optimize the design of the reward function, or increase the amount of training data to improve the performance of the model.

[0028] Step 6: Establish a real-time monitoring system to monitor multi-source carbon asset data in real time, and dynamically adjust carbon asset values, matching results, and carbon neutrality paths based on the real-time monitoring data; Specifically, first, deploy diverse data collection terminals within enterprises, parks, or regions for various carbon asset data sources. For carbon emission reduction and energy consumption data, sensors can be installed on production equipment and energy metering instruments to collect data in real time. For carbon sequestration projects, use satellite remote sensing, ground monitoring stations, and other equipment to obtain carbon sequestration change data. Connect to carbon trading platforms and government policy release websites through network interfaces to obtain real-time carbon trading prices, policy and regulatory updates, and other information. Build a stable and efficient data transmission network to ensure that collected multi-source carbon asset data can be transmitted to the monitoring center in a timely and accurate manner. Use a combination of wired networks (such as fiber optics) and wireless networks (such as 5G). For collection terminals that are close to the monitoring center and have large data volumes, prioritize wired networks to ensure stable data transmission. For widely distributed and remote terminals, use 5G networks for rapid data transmission, and configure data encryption technology to ensure data transmission security. Establish a monitoring center equipped with high-performance servers and data storage devices to receive, store, and process real-time collected data. Develop specialized data monitoring software with functions such as real-time data display, early warning settings, and data analysis to facilitate managers' intuitive understanding of dynamic changes in carbon asset data. Then, multi-source carbon asset data is continuously acquired according to the set collection frequency; for example, carbon trading market price data is collected once per second or minute to ensure that price fluctuations can be captured in a timely manner; for carbon emission reduction data in the production process, it is collected and updated every hour or every day according to the production cycle and equipment operation status; the newly collected data is stored in the database of the monitoring center, overwriting the old data to ensure the real-time nature of the data; at the same time, the real-time data is analyzed using monitoring software to set reasonable data thresholds and variation ranges; when carbon asset data exceeds the threshold or shows abnormal fluctuations, the system automatically issues an early warning message; for example, if the carbon trading price rises or falls sharply in a short period of time, exceeding the set fluctuation range, the system immediately sends a text message or email notification to the management personnel to remind them of the existing risks; Then, based on the real-time monitoring of policy and regulatory changes, market price fluctuations, and technological development trends, the time-varying value assessment model in step 2 is used to recalculate the value of carbon assets at different time points. For example, when the government introduces a new carbon emission reduction subsidy policy and the policy intensity changes, the model re-evaluates the value of carbon assets such as carbon emission reduction and carbon sinks, and adjusts the carbon asset management strategy in a timely manner; and based on the real-time information on the location of emission sources and carbon sink projects, the supply-side geographical matching and demand-side temporal matching are re-performed. If a carbon sink project reduces its carbon sink volume or changes its location due to a natural disaster, its relationship with the emission source is recalculated. Spatial matching degree; according to the adjustment of the enterprise's production plan, the change of the time node of the carbon neutrality target, etc., the time matching degree is recalculated and the matching results are updated to provide more accurate data support for carbon neutrality path planning; combined with the adjusted carbon asset value and the matching results, the improved Bellman equation is applied again to recalculate the carbon neutrality path implementation effect data under the triple constraint conditions. For example, due to the change in carbon asset value, the emission reduction cost and carbon trading cost change, and the total carbon emission reduction, total carbon trading volume, total cost and other indicators are recalculated to promptly discover the problems in the original carbon neutrality path and provide a basis for path optimization; Finally, the adjusted carbon asset value, matching results and path implementation effect data are promptly fed back to other relevant modules in the system, such as the carbon asset management and evaluation module, the path calculation module, the reinforcement learning model construction module, etc. Each module re-analyzes and calculates based on the new data, so that the entire system can quickly adapt to changes in actual conditions; at the same time, the process and results of each data adjustment are recorded in detail in the monitoring system, including the reasons for the adjustment, the specific data of the adjustment, the effect evaluation after the adjustment and other information. These records not only help to trace the history of data changes, but also provide a reference basis for subsequent system optimization and decision analysis.

[0029] Step 7: Develop evaluation indicators, regularly assess the effectiveness of the carbon neutrality path implementation data, and optimize the parameter adjustment strategy of the reinforcement learning model based on the effectiveness evaluation results; Specifically, first, determine a reasonable evaluation cycle based on actual conditions and needs, such as monthly, quarterly, or annual evaluations. Short-term evaluations (e.g., monthly) can promptly identify problems during the implementation of the pathway and enable adjustments. Long-term evaluations (e.g., annual) can help assess the overall effectiveness of the pathway from a macro perspective. At the same time, collect data on the implementation of the carbon neutrality pathway within the evaluation cycle, including data related to the various evaluation indicators mentioned above, and clean and organize the data to ensure its accuracy and consistency. Then, using the established evaluation indicators, we analyze and calculate the collected data. By comparing the indicator values in different evaluation cycles, we can observe the changing trends in the implementation of the carbon neutrality path. For example, we can compare the carbon emission reduction rates in adjacent years to determine whether the emission reduction effect is continuously improving. Next, we study the relationship between the evaluation results and the various parameters in the reinforcement learning model. For example, if we find that the carbon emission reduction rate does not meet the expected target, we analyze whether it is due to an unreasonable setting of the reward function in the model, causing the model to prefer a strategy with lower cost but poor emission reduction effect when making decisions. Subsequently, based on the analysis results, the parameter adjustment strategy of the reinforcement learning model was optimized. The specific measures are as follows: Reward function optimization: If the evaluation results show that certain key indicators (such as carbon emission reduction) are improving slowly, the weight related to this indicator in the reward function can be adjusted to increase the reward intensity for this indicator and guide the model to focus more on optimizing this indicator. For example, the reward weight for carbon emission reduction can be increased from 0.3 to 0.5. Learning rate adjustment: If the model converges too slowly or oscillates during training, adjust the learning rate. If the evaluation results show no significant improvement in model performance, lower the learning rate appropriately to make the model more cautious when updating parameters. If the model stagnates for a long time, increase the learning rate appropriately to speed up convergence. Exploration rate adjustment: For models using the ε-greedy strategy, adjust the exploration rate ε based on the evaluation results. In the early stages of the model, a higher exploration rate can be set to allow the model to fully explore different strategies. As the evaluation results show that the model performance gradually stabilizes, the exploration rate can be reduced to make the model more inclined to choose the optimal strategy that has been learned. Finally, after adjusting the parameter adjustment strategy, the reinforcement learning model is retrained using the updated data. During the training process, the model's performance indicators are continuously monitored, and the model is verified using the validation set to ensure that the model can achieve better results under the new parameter adjustment strategy. Through continuous iterative optimization, the reinforcement learning model can more accurately adapt to the dynamic changes of the carbon neutrality path and improve the efficiency and effectiveness of carbon neutrality path planning.

[0030] In one implementation, specifically, in step 2, the four-dimensional asset matrix is constructed from carbon emission reduction, carbon sink, carbon quota and green certificate as four dimensions; among them, carbon emission reduction is the reduction of carbon dioxide emissions achieved by enterprises, parks or regions by means of improving production technology, optimizing energy structure, improving energy utilization efficiency, etc. For example, enterprises introduce advanced energy-saving equipment to reduce energy consumption per unit product, thereby reducing carbon dioxide emissions. This part of the reduced emissions is carbon emission reduction; carbon sink is achieved by afforestation, forest management, vegetation restoration, etc., using plant photosynthesis to absorb carbon dioxide in the atmosphere and fix it in vegetation and soil, thereby reducing the concentration of carbon dioxide in the atmosphere. A carbon quota is a carbon emission quota allocated to an enterprise for a certain period of time by the government or relevant management departments based on factors such as the enterprise's industry category, production scale, and historical emission levels. If the enterprise's actual emissions are lower than the quota, it can sell the remaining quota on the carbon trading market; otherwise, it needs to purchase additional quotas from the market; a green certificate is a certification of green electricity produced by renewable energy power generation enterprises, proving that the electricity is clean energy generated by renewable energy. Green certificates can be traded on the market, which helps promote the development and consumption of renewable energy.

[0031] The time-varying value assessment model is: V(t) = f(policy intensity, market volatility, technology decay rate); Among them, V(t) represents the value of carbon assets at different time points t; for example, at the current time point t1, due to the strong policy support for carbon emission reduction, the market demand for carbon assets is strong, and the relevant technology is stable and effective, and the calculated carbon asset value V(t1) of a certain enterprise is 1 million yuan; after a period of time to t2, the policy changes, market fluctuations lead to a drop in carbon asset prices, and technology also shows a certain attenuation, at this time V(t2) becomes 800,000 yuan.

[0032] f represents the mapping relationship between carbon asset value and policy intensity, market volatility, and technological decay rate. The mapping relationship can be determined through a variety of mathematical methods, such as linear regression, nonlinear regression, and machine learning algorithms (such as neural networks and decision trees). By training the model with a large amount of historical data, the function form that most accurately describes the relationship between these variables is found. For example, after analyzing historical data and training the model, the functional relationship is determined to be V(t) = 20×policy intensity + 15×market volatility - 10×technological decay rate + 50.

[0033] Policy intensity is quantified based on the degree of support or restriction on carbon emission reduction by policies and regulations; Specifically, carbon emission reduction policies introduced by the government, such as carbon taxes, subsidies, and carbon emission quota allocation policies, have a significant impact on the value of carbon assets. These can be quantified in terms of the degree of policy incentives or restrictions, coverage, and enforcement strength. For example, in carbon tax policies, quantitative indicators are set based on the tax amount per ton of carbon emissions; in subsidy policies, the proportion of subsidies to the company's emission reduction costs is measured. For example, if a region provides high subsidies to companies using clean energy, and the subsidy amount accounts for 30% of the company's clean energy equipment investment, this ratio can be used as a numerical value to measure the policy strength. Supportive policies, such as subsidies and tax exemptions, will reduce corporate emission reduction costs and increase the value of carbon assets. Taking subsidy policies as an example, after companies receive subsidies, the actual value of carbon emission reductions increases, and the corresponding carbon asset value increases; restrictive policies, such as raising carbon taxes and tightening carbon emission quotas, will increase corporate carbon emission costs, prompting companies to pay attention to carbon assets and increase their value. If carbon taxes are raised, companies will increase their emission reduction efforts to reduce costs, which will increase the value of carbon emission reductions and carbon sinks.

[0034] Market volatility is measured by the magnitude and frequency of changes in carbon asset market prices; Specifically, market volatility is measured by the amplitude and frequency of changes in carbon asset market prices. The amplitude of change can be calculated by taking the difference between the highest and lowest carbon asset prices over a period of time (e.g., a week or a month) and dividing it by the initial price. The frequency can be calculated by counting the number of times the price fluctuates by more than a certain threshold (e.g., 5%) over a certain period of time. Assuming that the highest price of a carbon allowance is 50 yuan / ton, the lowest price is 40 yuan / ton, and the initial price is 45 yuan / ton within a month, the amplitude of change is (50-40) ÷ 45 ≈ 0.22. If the number of times the price fluctuates by more than 5% within that month is three, this is the frequency indicator. When market prices fluctuate violently, the uncertainty of carbon asset value increases; in the price rising stage, the asset value of companies holding carbon assets increases, stimulating companies to increase their carbon asset reserves; in the price falling stage, the value of companies' carbon assets shrinks, prompting companies to adjust their strategies, such as selling some carbon assets; when prices in the carbon trading market rise rapidly, the value of carbon quotas held by companies increases significantly, and companies will pay more attention to the appreciation potential of carbon assets.

[0035] The technology decay rate is calculated based on the degree to which the effectiveness of the relevant technology decreases over time since its application; Specifically, it is calculated based on the degree to which the effectiveness of the relevant technology decreases over time since its application. For emission reduction technologies, the emission reduction efficiency at the initial stage of technology application and the current stage can be compared. For example, in the initial stage of application of a certain emission reduction technology, the unit investment can achieve an emission reduction of 100 tons / year. After 5 years of use, the unit investment emission reduction drops to 80 tons / year. The technology attenuation rate is (100-80) ÷ 100 ÷ 5 = 0.04. For carbon sequestration technology, the change in carbon sequestration over time can be observed. Technological decay will cause the value of carbon assets to decline; as the effectiveness of technology decreases, the cost of achieving the same amount of carbon emissions reduction or carbon sink increases, and the value of carbon assets decreases accordingly; if the carbon capture technology used by a company decays due to technological decay, the cost of capturing each ton of carbon dioxide increases, the asset value corresponding to its carbon emissions reduction will decrease.

[0036] The specific application of the time-varying value assessment model is as follows: Collect information on policies and regulations, carbon asset market prices, and technology application data to quantify policy intensity, measure market volatility, and calculate technology decay rates. Substitute this data into a time-varying valuation model, V(t) = f(policy intensity, market volatility, technology decay rate), and calculate the value of carbon assets at different time points t using a specific algorithm (e.g., a functional relationship determined by linear regression or a neural network algorithm). Assuming that, through analysis, the policy intensity at a certain point in time is 0.8 (out of a maximum score of 1), the market volatility is 0.15, and the technology decay rate is 0.03, substituting these data into the model yields a carbon sink asset value of 45 yuan per ton at that time. Enterprises optimize their carbon asset management strategies based on the model evaluation results. When the model shows a clear upward trend in the value of carbon assets, enterprises can increase their carbon asset investments. If the value declines, enterprises can adjust emission reduction technologies and sell some carbon assets. For example, if the model predicts that the value of a certain type of carbon asset will continue to rise in the next six months, enterprises can reserve such carbon assets in advance to obtain value-added benefits.

[0037] In one implementation, specifically, in step 4, the improved Bellman equation is: minΣ[αC 减排 +(1−α)C 交易 ]; Specifically, the Bellman equation is usually used to solve the optimal strategy in dynamic programming. In the carbon neutral path planning scenario, the improved Bellman equation minΣ[αC 减排 +(1−α)C 交易 The goal is to minimize the total cost, which is the weighted sum of the emission reduction cost and the carbon trading cost. By continuously accumulating these costs at different decision moments (such as different time periods), a carbon neutrality path with the minimum total cost can be found. Among them, the adaptive weight factor α is dynamically adjusted according to the maturity, cost-effectiveness of emission reduction technology, the stability of carbon trading market prices and cost factors. 减排 represents the emission reduction cost, C 交易 represents the cost of carbon trading; The improved Bellman equation dynamically adjusts the emission reduction cost C according to actual conditions 减排 and transaction costs C 交易 The weight relationship between them; Specifically, the adaptive weight factor α is a value between 0 and 1, which represents the relative importance of emission reduction costs and carbon trading costs in the total cost; the larger the value of α, the greater the weight of emission reduction costs in the total cost, that is, the more inclined to achieve carbon neutrality through emission reduction measures; conversely, the smaller α is, the more dependent on carbon trading to achieve the goal; α will be dynamically adjusted according to the maturity of emission reduction technology, cost-effectiveness, and the stability of carbon trading market prices and cost factors; for example, if the emission reduction technology is already very mature, the unit emission reduction cost is low, and the carbon trading market price fluctuates greatly and the cost is high, then the value of α can be appropriately increased, which means that it is more inclined to achieve carbon neutrality through its own emission reduction measures; conversely, if the emission reduction technology is still immature and the cost is high, but the carbon trading market is stable and the cost is low, then the value of α can be reduced, and more reliance can be placed on carbon trading; Emission reduction costs refer to the costs invested by an enterprise, park, or region to reduce carbon emissions, including but not limited to the cost of purchasing energy-saving and emission reduction equipment, the cost of improving production processes, and the cost of training employees to use new technologies. Emission reduction costs need to be calculated by taking into account various factors. For example, an enterprise plans to install a new set of exhaust gas treatment equipment to reduce carbon emissions. The equipment costs 1 million yuan to purchase, has a service life of 10 years, and an annual maintenance cost of 100,000 yuan. At the same time, it costs 50,000 yuan to train employees to adapt to the new equipment. Then, over this 10-year usage cycle, the average annual emission reduction cost needs to be reasonably allocated to calculate these costs. Carbon trading costs refer to the costs incurred by enterprises in buying and selling carbon emission rights in the carbon trading market. When an enterprise's own carbon emissions exceed the allocated carbon quota, it needs to purchase additional quotas from the market. Conversely, if the enterprise's carbon emissions are lower than the quota, it can sell the remaining quota. Carbon trading costs are mainly determined by the price and transaction volume of the carbon trading market, and additional costs such as handling fees incurred during the transaction must also be considered. For example, the price of each ton of carbon quota in the carbon trading market is 50 yuan, and the enterprise needs to purchase 100 tons of quotas. The transaction fee rate is 2%, then the carbon trading cost is 50×100×(1+2\%)=5100 yuan.

[0038] The triple constraints include policy compliance constraints, capital liquidity constraints, and technical feasibility constraints; Among them, policy compliance constraints mean that the carbon neutrality path planning of an enterprise or region must comply with relevant national and local policies and regulations. These policies include total carbon emission control targets, carbon emission reduction ratio requirements, emission standards for specific industries, etc. Policy compliance constraints can ensure that planning schemes are carried out within a legal and compliant framework, avoiding penalties for policy violations, and also helping to promote the realization of carbon neutrality goals for the entire society. For example, if a region stipulates that enterprises must reduce their total carbon emissions by 20% within the next five years, then the carbon neutrality path planning of enterprises must meet this policy requirement. Liquidity constraints refer to taking into account the financial status and capital budget of the enterprise to ensure that the carbon neutrality path planning will not lead to a tight capital chain for the enterprise. When implementing emission reduction measures or conducting carbon trading, a certain amount of capital needs to be invested. If the capital investment is too large and too concentrated, it will affect the normal operation of the enterprise. Liquidity constraints can ensure that the enterprise can maintain a good financial health in the process of achieving the carbon neutrality goal, avoiding project interruptions or financial difficulties caused by funding problems. For example, when formulating plans, enterprises need to take into account the annual capital expenditure to ensure that there will be no capital shortage due to the purchase of large amounts of emission reduction equipment or large-scale carbon trading. Technical feasibility constraints mean that the emission reduction technologies and carbon sink technologies used in the planning scheme must be practical, including the maturity, reliability, and availability of supporting infrastructure of the technology. Technical feasibility constraints can avoid the use of technologies that have not yet been verified or do not have the conditions for implementation, thereby ensuring that the planning scheme can be implemented smoothly. For example, if an enterprise plans to adopt a new type of carbon capture technology, but the technology is still in the laboratory stage and lacks experience and supporting equipment for large-scale applications, then it does not meet the technical feasibility constraints and the plan needs to be adjusted.

[0039] In one implementation, specifically, in step five, the method for constructing a reinforcement learning model includes the following steps: Step 501: Organize the multi-source carbon asset data and implementation effect data and divide them into training set, validation set and test set; Specifically, first, the collected multi-source carbon asset data, such as carbon emission reductions, carbon sinks, carbon quotas, green certificates, emission source locations, policy and regulatory information, as well as carbon neutrality implementation effect data, such as carbon emission reductions, carbon trading volume, total costs, and the matching status of various emission sources and carbon sink projects, will be summarized and integrated to ensure data accuracy and consistency. The data will also be cleaned to remove duplicate, erroneous, or missing values. Then, the sorted data is divided into training set, validation set and test set according to a certain ratio; the common division ratio is 70% training set, 15% validation set and 15% test set; among them, the training set is used to train the reinforcement learning model to allow the model to learn the patterns in the data; the validation set is used to evaluate the performance of the model during the training process to adjust the model's hyperparameters and avoid overfitting; the test set is used to finally evaluate the generalization ability of the trained model to ensure that the model can perform well on unseen data.

[0040] Step 502: Build a reinforcement learning model based on the deep Q network architecture, and construct a neural network including an input layer, multiple hidden layers, and an output layer; Specifically, the input layer receives preprocessed data that constitutes the state of the reinforcement learning model. For example, the number of carbon assets, market prices, policy indicators, and key parameters of the current carbon neutrality path are encoded as neurons in the input layer. The number of neurons in the input layer depends on the dimension of the state vector to ensure that the state information can be fully input into the network. Multiple hidden layers are set up to extract complex features from the data. The number of hidden layer neurons and the number of layers can be adjusted according to the complexity of the data and the performance of the model. Typically, the hidden layer uses the ReLU (Rectified Linear Unit) activation function, i.e., f(x) = max(0, x), which can effectively alleviate the vanishing gradient problem and accelerate model convergence. The number of neurons in the output layer corresponds to the number of actions the model can take. Each neuron outputs a value, which represents the Q value (action value) of taking the corresponding action in the current state. For example, if the model's actions include adjusting carbon emission reduction strategies, selecting different carbon sink projects, conducting carbon trading, etc., the number of neurons in the output layer is equal to the total number of these actions. The output layer generally does not use an activation function and directly outputs the Q value.

[0041] Step 503: Use the training set to train the built reinforcement learning model, and use the adaptive learning rate adjustment algorithm to automatically adjust the learning rate according to the change of the loss function during the training process; Specifically, during training, the model sequentially acquires states from the training set and selects actions based on the current policy. The policy is usually based on the ε-greedy algorithm, where the model randomly selects actions with a probability of ε and selects the action with the largest Q value with a probability of 1-ε. After executing the action, the model receives a reward and a new state from the environment. These experiences (state, action, reward, new state) are stored in the experience replay buffer. Adaptive learning rate adjustment algorithms, such as the Adam (Adaptive Moment Estimation) algorithm, are used. During training, the algorithm automatically adjusts the learning rate according to changes in the loss function. When the loss function decreases rapidly, the learning rate is appropriately increased to speed up training. When the loss function decreases slowly or starts to increase, the learning rate is reduced to prevent the model from missing the optimal solution. By continuously randomly sampling experience data from the experience replay buffer, the error (such as the mean square error) between the predicted Q value and the target Q value is calculated, and the backpropagation algorithm is used to update the weights of the neural network, so that the model gradually learns the optimal strategy.

[0042] Step 504: Evaluate and optimize the trained reinforcement learning model using verification; Specifically, the validation set is used to evaluate the model in training. During the evaluation process, the model selects actions based on the states in the validation set according to the trained strategy and calculates evaluation indicators such as cumulative rewards, carbon reduction target achievement rate, and cost control effect. The performance of the model is judged based on these indicators. For example, higher cumulative rewards, higher carbon reduction target achievement rate, and better cost control effect indicate better model performance. Optimize the model based on the evaluation results. If the model's performance on the validation set does not meet expectations, adjust the model's hyperparameters, such as the number of hidden layers and neurons, the ε value in the ε-greedy algorithm, and the size of the experience replay buffer. You can also try adjusting the design of the reward function to enable the model to learn a strategy that better meets actual needs. Repeated adjustments and evaluations are performed until the model achieves good performance on the validation set.

[0043] Step 505: Test the optimized reinforcement learning model using the test set; Specifically, the optimized model is finally tested using the test set. This process is similar to the validation set evaluation, but the test set has not been used during training and validation, and can more realistically reflect the model's generalization ability. Various evaluation indicators of the model on the test set are calculated, such as cumulative rewards, carbon emission reduction target achievement rate, and cost control effect, and compared with the expected targets. Analyze the test results. If the model performs well on the test set, it means that the model can be effectively applied to actual carbon neutrality path planning. If the model performs poorly, further analysis is needed to determine whether it is due to data quality issues, unreasonable model structure, or improper training methods. After improving the problems, retrain and test the model until the model meets the requirements.

[0044] In one implementation, specifically, in step three, the supply-side geographic matching is based on the emission source location and the carbon sink project location, by calculating the transportation distance between the two and using the formula D=1 / (1+transportation distance) to obtain the spatial matching degree D; Among them, D represents the spatial matching degree, which ranges from 0 to 1. The larger the value, the higher the degree of matching between the supply-side carbon assets and the demand-side in geographical space; "transportation distance" refers to the actual distance from the carbon sink project location to the emission source location; Assume there is an emission source and a carbon sink project, and the transportation distance between the two is 50 kilometers as calculated by GIS. Substituting this into the formula yields: D=1 / 1+50≈0.0196; This result indicates that the carbon sequestration project has a low degree of geographical matching with the emission source.

[0045] Demand-side timing matching calculates the time difference Δt based on the time nodes set by short-term goals and long-term plans, and uses the formula T=e (-λ∣Δt∣) Calculate the time matching degree T; Where T represents the time matching degree, and its value range is between 0 and 1. The larger the value, the higher the temporal matching degree between the demand side's time demand and the actual carbon asset supply. λ is a constant greater than 0, called the attenuation factor, which determines the degree of influence of the time difference on the time matching degree. Δt represents the time difference calculated based on the time nodes set by the short-term goal and the long-term plan. e is a natural constant. Assuming the attenuation factor λ = 0.1, the short-term target time node on the demand side is December 31, 2025, and the actual supply time of carbon assets is June 30, 2026. The time difference Δt = 180 days (for the convenience of calculation, assuming a year of 360 days); substituting it into the formula, we can get: T = e (-0.1×180) ≈2.06×10 -8 ; This result shows that the temporal matching between the demand side’s temporal demands and the supply of carbon assets is extremely low.

[0046] In one implementation, specifically, in step one, the multi-source carbon asset data includes carbon emission reductions, carbon sinks, carbon quotas, green certificates, emission source locations, surrounding carbon sink project locations, policy and regulatory information, market price fluctuation data, and technological development trends; the multi-source carbon asset data is preprocessed, including data cleaning and data normalization; wherein, the purpose of data cleaning is to remove duplicate data, process missing values, and correct erroneous data; the purpose of data normalization is to unify different types of data to the same scale, eliminate the impact of dimension, and improve the accuracy and stability of the model.

[0047] In one implementation, specifically, in step seven, the evaluation indicators include carbon emission reduction target completion rate, cost control status and policy compliance; among them, the carbon emission reduction target completion rate is obtained by calculating the ratio of the actual carbon emission reduction amount to the set emission reduction target amount; the cost control status is obtained by comparing the actual carbon asset investment cost with the budgeted cost; and the policy compliance is obtained by checking whether the path implementation process strictly complies with the relevant policy and regulatory requirements.

[0048] In summary, the embodiment of the present invention provides a dynamic planning method for a carbon neutrality path based on carbon asset matching. By acquiring and preprocessing multi-source carbon asset data, a four-dimensional asset matrix is constructed to classify and manage carbon assets and its value is dynamically evaluated using a time-varying value assessment model; supply-side geographic matching and demand-side temporal matching are performed to generate matching results; combining the carbon asset value and matching results, the improved Bellman equation is applied to calculate the carbon neutrality path implementation effect data under triple constraints; a reinforcement learning model is constructed based on multi-source data and implementation effects; a real-time monitoring system is established to dynamically adjust relevant data; and an evaluation index is formulated to optimize the reinforcement learning model parameter adjustment strategy, which effectively solves the problems of static evaluation, insufficient coordination, and multi-objective imbalance in carbon neutrality path planning in the existing technology, significantly improves the efficiency and benefits of carbon asset management, and efficiently promotes the achievement of carbon neutrality goals.

[0049] Figure 3 This is a functional module diagram of a carbon neutrality path dynamic planning system based on carbon asset matching in an embodiment of the present application. Figure 3 As shown, a carbon neutrality path dynamic planning system based on carbon asset matching, the system includes: data acquisition and processing module, carbon asset management and evaluation module, supply and demand matching module, path calculation module, reinforcement learning model construction module, real-time monitoring and adjustment module and effect evaluation and optimization module; A data acquisition and processing module is configured to obtain multi-source carbon asset data within an enterprise, park or region, and pre-process the multi-source carbon asset data; The carbon asset management and assessment module is configured to construct a four-dimensional asset matrix based on the acquired multi-source carbon asset data, classify and manage carbon assets, and build a time-varying value assessment model to dynamically assess the value of carbon assets; A supply and demand matching module is configured to perform supply-side geographic matching and demand-side temporal matching based on the acquired multi-source carbon asset data, and generate matching results; The path calculation module is configured to combine the carbon asset value and matching results, apply the improved Bellman equation, and calculate and output the carbon neutrality path and implementation effect data under the triple constraint conditions; A reinforcement learning model building module configured to build a reinforcement learning model based on multi-source carbon asset data and implementation effect data; A real-time monitoring and adjustment module is configured to establish a real-time monitoring system to monitor multi-source carbon asset data in real time and dynamically adjust carbon asset values, matching results, and carbon neutrality paths based on the real-time monitoring data; The effect evaluation and optimization module is configured to formulate evaluation indicators, regularly evaluate the effect data of the carbon neutrality path implementation, and optimize the parameter adjustment strategy of the reinforcement learning model based on the effect evaluation results.

[0050] In one embodiment, specifically, the system further includes a display and interaction module, which is used to present the calculation results and analysis reports in a visual manner and receive instructions and parameter information input by the user.

[0051] In one implementation, specifically, the data acquisition and processing module adopts a distributed data acquisition architecture to simultaneously acquire multi-source carbon asset data from multiple data sources in parallel.

[0052] For other details about the technical solutions for implementing each module in the carbon neutrality path dynamic planning system based on carbon asset matching in the above embodiment, please refer to the description of the carbon neutrality path dynamic planning method based on carbon asset matching in the above embodiment, which will not be repeated here.

[0053] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For system-related embodiments, since they are generally similar to method-related embodiments, their description is relatively simple. For relevant details, refer to the description of the method-related embodiments.

[0054] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0055] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0056] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.

[0057] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0058] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.

[0059] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0060] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A carbon neutrality path dynamic planning method based on carbon asset matching, characterized in that: The following steps are involved: Obtain multi-source carbon asset data within an enterprise, park or region, and pre-process the multi-source carbon asset data; Based on the multi-source carbon asset data obtained, a four-dimensional asset matrix is constructed to classify and manage carbon assets, and a time-varying value assessment model is constructed to dynamically assess the value of carbon assets; Based on the acquired multi-source carbon asset data, perform supply-side geographic matching and demand-side temporal matching and generate matching results; Combining carbon asset value and matching results, the improved Bellman equation is applied to calculate and output the carbon neutrality path and implementation effect data under triple constraints; Build a reinforcement learning model based on multi-source carbon asset data and implementation effect data; Establish a real-time monitoring system to monitor multi-source carbon asset data in real time, and dynamically adjust carbon asset values, matching results, and carbon neutrality paths based on real-time monitoring data; Establish evaluation indicators, regularly evaluate the implementation effect data of the carbon neutrality path, and optimize the parameter adjustment strategy of the reinforcement learning model based on the effect evaluation results.

2. A carbon neutrality path dynamic planning method based on carbon asset matching according to claim 1, characterized in that: The four-dimensional asset matrix is constructed from carbon emission reduction, carbon sink, carbon quota and green certificate as four dimensions; The time-varying value assessment model is: V(t) = f(policy intensity, market volatility, technology decay rate); Among them, V(t) represents the value of carbon assets at different time points t; f represents the mapping relationship between carbon asset value and policy intensity, market volatility, and technological decay rate; policy intensity is quantified based on the support or restriction of policies and regulations on carbon emission reduction; market volatility is measured based on the amplitude and frequency of changes in carbon asset market prices; and the technological decay rate is calculated based on the degree of reduction in the effectiveness of the relevant technology over time since its application.

3. A carbon neutrality path dynamic planning method based on carbon asset matching according to claim 1, characterized in that: The improved Bellman equation is: minΣ[αC 减排 +(1−α)C 交易 ]; Among them, the adaptive weight factor α is dynamically adjusted according to the maturity, cost-effectiveness of emission reduction technology, the stability of carbon trading market prices and cost factors. 减排 represents the emission reduction cost, C 交易 represents the cost of carbon trading; The improved Bellman equation dynamically adjusts the emission reduction cost C according to the actual situation 减排 and transaction costs C 交易 The weight relationship between them; The triple constraints include policy compliance constraints, capital liquidity constraints and technical feasibility constraints.

4. A carbon neutrality path dynamic planning method based on carbon asset matching according to claim 1, characterized in that: The method for constructing a reinforcement learning model comprises the following steps: Step 501: Organize the multi-source carbon asset data and implementation effect data and divide them into training set, validation set and test set; Step 502: Build a reinforcement learning model based on the deep Q network architecture, and construct a neural network including an input layer, multiple hidden layers, and an output layer; Step 503: Use the training set to train the built reinforcement learning model, and use the adaptive learning rate adjustment algorithm to automatically adjust the learning rate according to the change of the loss function during the training process; Step 504: Evaluate and optimize the trained reinforcement learning model using verification; Step 505: Use the test set to test the optimized reinforcement learning model.

5. The carbon neutrality path dynamic planning method based on carbon asset matching according to claim 1 is characterized by: The supply-side geographical matching is based on the location of the emission source and the location of the carbon sink project. The transportation distance between the two is calculated and the spatial matching degree D is obtained using the formula D=1 / (1+transportation distance). Where D represents the spatial matching degree, ranging from 0 to 1. A larger value indicates a higher degree of geographical matching between the carbon assets on the supply side and the demand side. "Transportation distance" refers to the actual distance from the carbon sink project location to the emission source location. The demand-side timing matching calculates the time difference Δt according to the time nodes set by short-term goals and long-term plans, and uses the formula T=e (-λ∣Δt∣) Calculate the time matching degree T; Among them, T represents the time matching degree, and its value range is between 0 and 1. The larger the value, the higher the temporal matching degree between the time demand on the demand side and the actual carbon asset supply. λ is a constant greater than 0, called the attenuation factor, which determines the degree of influence of the time difference on the time matching degree. Δt represents the time difference calculated based on the time nodes set for short-term goals and long-term plans. e is a natural constant.

6. A carbon neutrality path dynamic planning method based on carbon asset matching according to claim 1, characterized in that: The multi-source carbon asset data includes carbon emission reductions, carbon sinks, carbon quotas, green certificates, emission source locations, locations of surrounding carbon sink projects, policy and regulatory information, market price fluctuation data, and technological development trends; The pre-processing of multi-source carbon asset data includes data cleaning and data normalization.

7. A carbon neutrality path dynamic planning method based on carbon asset matching according to claim 1, characterized in that: The evaluation indicators include carbon emission reduction target completion rate, cost control and policy compliance.

8. A carbon neutrality path dynamic planning system based on carbon asset matching, applied to the carbon neutrality path dynamic planning method based on carbon asset matching according to any one of claims 1 to 7, characterized in that: The system includes: data acquisition and processing module, carbon asset management and evaluation module, supply and demand matching module, path calculation module, reinforcement learning model construction module, real-time monitoring and adjustment module and effect evaluation and optimization module; The data acquisition and processing module is configured to obtain multi-source carbon asset data within an enterprise, park or region, and pre-process the multi-source carbon asset data; The carbon asset management and assessment module is configured to construct a four-dimensional asset matrix based on the acquired multi-source carbon asset data, classify and manage carbon assets, and construct a time-varying value assessment model to dynamically assess the value of carbon assets; The supply and demand matching module is configured to perform supply-side geographic matching and demand-side temporal matching based on the acquired multi-source carbon asset data, and generate a matching result; The path calculation module is configured to combine the carbon asset value and the matching results, apply the improved Bellman equation, and calculate and output the carbon neutrality path and implementation effect data under the triple constraint conditions; The reinforcement learning model construction module is configured to construct a reinforcement learning model based on multi-source carbon asset data and implementation effect data; The real-time monitoring and adjustment module is configured to establish a real-time monitoring system, monitor multi-source carbon asset data in real time, and dynamically adjust carbon asset values, matching results, and carbon neutrality paths based on the real-time monitoring data; The effect evaluation and optimization module is configured to formulate evaluation indicators, regularly evaluate the implementation effect data of the carbon neutrality path, and optimize the parameter adjustment strategy of the reinforcement learning model based on the effect evaluation results.

9. A carbon neutrality path dynamic planning system based on carbon asset matching according to claim 8, characterized in that: The system further comprises a display and interaction module, which is used to present the calculation results and analysis reports in a visual manner and receive instructions and parameter information input by the user.

10. A carbon neutrality path dynamic planning system based on carbon asset matching according to claim 8, characterized in that: The data acquisition and processing module adopts a distributed data acquisition architecture to obtain multi-source carbon asset data from multiple data sources in parallel.

Citation Information

Cited By

  • Green carbon emission statistics and carbon asset management method and system

    CN121189609A

  • Industrial park pollution reduction and carbon reduction comprehensive decision path analysis method

    CN121257895A

  • Enterprise carbon reduction path optimization method based on CCER and energy-saving technical improvement cooperation

    CN121436336A

  • Multi-dimensional carbon neutralization path simulation and decision-making system driven by enterprise energy consumption data

    CN121961619A

  • Enterprise energy consumption data-driven multi-dimensional carbon neutrality pathway simulation and decision-making system

    CN121961619B