Carbon asset transaction strategy generation method and system based on multi-source information fusion and reinforcement learning

CN122840802APending Publication Date: 2026-09-29SHENZHEN ZHONGTIAN BIM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611171836.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-04
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

(1)多源碳要素数据未能有效整合与利用,易形成数据孤岛,导致决策依据不全面,且未结合企业自身风险偏好、资金约束、履约紧迫度等个性化特征构建决策模型,策略通用性强、针对性差,无法适配不同企业的实际运营与风险承受能力;

Benefits of technology

本发明通过多源碳要素数据的融合预处理与企业碳资产画像构建,打破传统碳交易决策的数据孤岛局限,将企业风险偏好、资金约束、履约紧迫度等个性化特征融入碳资产决策体系,再基于预设碳交易规则搭建定制化的碳资产交易强化学习环境,精准定义适配企业实际需求的状态空间、动作空间与奖励函数,依托强化学习模型的自适应迭代训练实现交易决策逻辑的自主优化,无需依赖人工经验制定策略,最终可基于实时碳资产状态快速输出合规、低成本、低风险的最优碳资产交易策略,大幅提升企业碳资产履约管控与交易决策的效率、针对性及合理性,有效降低企业碳交易成本与履约风险,适配不同企业的碳资产运营管理需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840802A_ABST
    Figure CN122840802A_ABST
Patent Text Reader

Abstract

This invention relates to the field of carbon trading technology, specifically providing a method and system for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning. The method includes: acquiring multi-source carbon element data and performing fusion preprocessing to obtain a carbon element fusion dataset; constructing an enterprise carbon asset profile containing carbon asset decision-making characteristics based on the dataset; building a carbon asset trading reinforcement learning environment according to preset carbon trading rules, defining the state space, action space, and reward function based on the profile; iteratively training a carbon asset trading reinforcement learning model in the carbon asset trading reinforcement learning environment until convergence; and finally, inputting real-time carbon asset state data and outputting the optimal carbon asset trading strategy through the trained model. This invention can eliminate data silos, achieve adaptive model optimization, and output compliant, low-cost, and low-risk carbon asset trading strategies, effectively improving the efficiency and rationality of enterprise carbon asset compliance management and trading decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon trading technology, specifically to a method and system for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning. Background Technology

[0002] With the gradual improvement of the carbon trading market and the continuous advancement of the "dual carbon" goals, corporate carbon asset trading and compliance management have become routine management needs. Corporate carbon asset decisions involve multiple sources and dimensions of information, including internal quota holdings, CCER inventory, financial status, compliance period, as well as external market carbon prices, trading rules, and policy requirements. The above data is scattered and lacks effective integration, making it difficult to form a unified decision-making basis.

[0003] Current carbon asset trading strategies rely heavily on human experience and judgment, or are based on simple analysis of data from a single dimension, which has significant shortcomings. (1) The multi-source carbon element data has not been effectively integrated and utilized, which easily leads to data silos, resulting in incomplete decision-making basis. Furthermore, the decision-making model has not been constructed in combination with the enterprise's own risk preferences, capital constraints, and urgency of fulfilling obligations. The strategy is highly generalized but poorly targeted, and cannot be adapted to the actual operation and risk tolerance of different enterprises. (2) Traditional intelligent decision-making methods lack adaptive learning mechanisms and fail to build customized reinforcement learning environments based on enterprise characteristics, making it difficult for the model to autonomously optimize transaction decision logic; (3) The company is unable to quickly output the optimal trading strategy that is compliant, low-cost and low-risk based on the real-time carbon asset status. In the process of carbon asset trading, the company is likely to face problems such as high costs and high compliance risks.

[0004] To address the shortcomings of the existing technologies, this application provides a method and system for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning, thereby solving the aforementioned problems. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides a method and system for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning, in order to solve the problems in existing technologies.

[0006] One embodiment of the present invention provides a method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning, comprising the following steps: Acquire multi-source carbon element data, fuse and preprocess the multi-source carbon element data to obtain a carbon element fusion dataset; Based on the carbon element fusion dataset, construct a corporate carbon asset profile that includes carbon asset decision-making characteristics; A carbon asset trading reinforcement learning environment is constructed based on preset carbon trading rules, and the state space, action space, and reward function of the reinforcement learning environment are defined based on the enterprise carbon asset profile. Based on the defined state space, action space, and reward function, the pre-built carbon trading reinforcement learning model is iteratively trained in the carbon asset trading reinforcement learning environment to obtain the trained carbon trading reinforcement learning model. Acquire real-time carbon asset status data and output the optimal carbon asset trading strategy through a trained carbon trading reinforcement learning model.

[0007] This application also relates to a carbon asset trading strategy generation system based on multi-source information fusion and reinforcement learning, including: The data processing module is used to acquire multi-source carbon element data, fuse and preprocess the multi-source carbon element data to obtain a carbon element fusion dataset. The first construction module is used to build a corporate carbon asset profile containing carbon asset decision-making characteristics based on the carbon element fusion dataset. The second construction module is used to build a carbon asset trading reinforcement learning environment according to the preset carbon trading rules, and to define the state space, action space and reward function of the reinforcement learning environment according to the enterprise carbon asset profile. The model training module is used to iteratively train a pre-built carbon trading reinforcement learning model in a carbon asset trading reinforcement learning environment based on a defined state space, action space, and reward function, so as to obtain a trained carbon trading reinforcement learning model. The strategy generation module is used to acquire real-time carbon asset status data and output the optimal carbon asset trading strategy through a trained carbon trading reinforcement learning model.

[0008] This application also relates to a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning.

[0009] This application also relates to a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning.

[0010] The carbon asset trading strategy generation method and system based on multi-source information fusion and reinforcement learning provided in the above embodiments have the following beneficial effects: This invention breaks through the data silos of traditional carbon trading decisions by fusing and preprocessing multi-source carbon element data and constructing corporate carbon asset profiles. It integrates personalized characteristics such as corporate risk preferences, financial constraints, and compliance urgency into the carbon asset decision-making system. Based on preset carbon trading rules, it builds a customized carbon asset trading reinforcement learning environment, accurately defining the state space, action space, and reward function to suit the actual needs of enterprises. Relying on the adaptive iterative training of the reinforcement learning model, it achieves autonomous optimization of the trading decision logic without relying on human experience to formulate strategies. Ultimately, it can quickly output the optimal carbon asset trading strategy that is compliant, low-cost, and low-risk based on real-time carbon asset status, significantly improving the efficiency, pertinence, and rationality of corporate carbon asset compliance management and trading decisions, effectively reducing corporate carbon trading costs and compliance risks, and adapting to the carbon asset operation and management needs of different enterprises. Attached Figure Description

[0011] Figure 1 A flowchart illustrating a carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning, provided as an embodiment of the present invention; Figure 2 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions in the embodiments of the present invention will now be clearly and completely described in conjunction with the accompanying drawings.

[0013] Reference Figure 1 One embodiment of the present invention provides a method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning, comprising the following steps: S10. Obtain multi-source carbon element data, fuse and preprocess the multi-source carbon element data to obtain a carbon element fusion dataset. S20. Based on the carbon element fusion dataset, construct a corporate carbon asset profile that includes carbon asset decision-making characteristics; S30. Construct a carbon asset trading reinforcement learning environment based on preset carbon trading rules, and define the state space, action space and reward function of the reinforcement learning environment based on the enterprise carbon asset profile. S40. Based on the defined state space, action space and reward function, the pre-built carbon trading reinforcement learning model is iteratively trained in the carbon asset trading reinforcement learning environment to obtain the trained carbon trading reinforcement learning model. S50: Acquire real-time carbon asset status data and output the optimal carbon asset trading strategy through a trained carbon trading reinforcement learning model.

[0014] In this embodiment, as described in steps S10-S50 above, the core is a fully intelligent decision-making mechanism that integrates multi-source carbon element data fusion preprocessing → personalized enterprise carbon asset profile construction → customized reinforcement learning environment setup → adaptive iterative model training → real-time optimal strategy output. This mechanism addresses the core technical pain points of traditional carbon asset trading strategies, such as isolated multi-source data, lack of personalized strategy adaptation, lack of adaptive learning capabilities, reliance on human experience, and low decision-making efficiency. It deeply integrates the personalized characteristics of internal enterprises with external market dynamics, leveraging the autonomous optimization characteristics of reinforcement learning to automatically generate carbon asset trading strategies that prioritize compliance, optimize costs, and controllable risks. This significantly improves the targeting, rationality, and efficiency of enterprise carbon trading decisions, effectively reduces compliance risks and transaction costs, and adapts to the carbon asset operation and management needs of different enterprises, as detailed below: Step S10 is the multi-source carbon element data collection and fusion preprocessing stage. Its core is to break down data silos and integrate scattered data, addressing the problems of ineffective integration of multi-source carbon data and incomplete decision-making basis in traditional decision-making. This provides a high-quality, unified data foundation for subsequent profile construction. Specifically, multi-source carbon element data encompasses internal corporate carbon data (such as carbon quota holdings, CCER inventory, carbon trading budget, compliance progress, quota carryover balance, historical trading records, and cash flow capacity) and external carbon market-related data (such as real-time carbon prices, price fluctuation trends, market regulatory rule adjustments, industry quota allocation standards, compliance cycle rules, and quota allocation mechanism updates). Data fusion and preprocessing include cleaning (removing outliers and filling in missing values), standardization (unifying data formats and statistical standards), correlation matching (establishing a mapping relationship between internal data and external carbon market-related data), and feature extraction, ultimately forming a fused carbon element data set. The dataset contains three main categories of core data: ① Carbon trading risk-related data (carbon price fluctuation range, historical trading default risk records, risk coefficients for changes in market regulatory rules, etc.); ② Carbon asset funding-related data (carbon trading budget, capital turnover flexibility parameters, transaction cost control thresholds, etc.); ③ Quota compliance-related data (current quota holdings, effective CCER stock, quota gap / surplus data, remaining compliance period, historical compliance completion status, etc.). This dataset is comprehensive and logically consistent, and is used to extract carbon trading risk data, carbon asset funding data, and quota compliance data in subsequent steps, breaking the limitations of "fragmented internal data and disconnected external carbon market data" in traditional decision-making.

[0015] Step S20 involves constructing a personalized corporate carbon asset profile. Its core objective is to uncover the personalized characteristics of corporate carbon asset decision-making, addressing the problem of traditional strategies being "highly generalizable but lacking specificity," and providing a fundamental basis for subsequent customization of the reinforcement learning environment. Specifically, based on a carbon element fusion dataset, it extracts and quantifies the unique characteristics of corporate carbon assets from three core decision-making dimensions: risk appetite, financial constraints, and compliance requirements. Instead of relying on generalized decision-making models, it uses data quantification to transform the company's own risk tolerance, cash flow flexibility, and compliance urgency into calculable decision features. This constructs a corporate carbon asset profile that accurately reflects the company's current carbon asset operation status and decision-making needs, making subsequent strategy generation more aligned with the company's actual operating conditions.

[0016] Step S30 involves constructing a customized reinforcement learning environment and defining its core elements. The core of this step is to build an intelligent decision-making scenario tailored to the company's needs, addressing the problem of traditional intelligent methods lacking customized environments and models that are difficult to adapt to actual business situations. This provides a simulation space for reinforcement learning model training that closely matches real-world carbon trading scenarios. Specifically, first, based on preset rules such as carbon market trading rules, quota compliance cycle rules, and market trading rule adjustment mechanisms, a reinforcement learning environment capable of dynamically simulating the entire carbon trading process is constructed. Then, using the company's carbon asset profile as the core, the three core elements of the environment are precisely defined: state space (covering the company's internal carbon asset status and external market dynamics), action space (limiting the scope of trading behaviors that comply with the company's risk and financial constraints), and reward function (matching the company's compliance and cost optimization goals). This ensures that the reinforcement learning environment is highly consistent with the company's actual needs and real market rules, providing accurate scenario simulation support for model training.

[0017] Step S40 is the adaptive iterative training phase of the reinforcement learning model. Its core is to enable the model to autonomously optimize its trading decision-making logic, addressing the problem of traditional strategies "relying on human experience and lacking autonomous optimization capabilities," thus achieving an intelligent upgrade of the decision-making logic. Specifically, based on the defined state space, action space, and reward function, a pre-built reinforcement learning model (such as DQN, PPO, or other models adapted to sequential decision-making) is placed in a carbon asset trading reinforcement learning environment. Through iterative cycles of "model executing trading actions → environment providing reward signals → model updating parameters based on rewards," the model continuously learns the decision-making logic of "compliance, low cost, and low risk" in simulated trading scenarios, gradually optimizing its strategy generation capabilities until the model converges to a stable state. This yields a trained carbon trading reinforcement learning model with autonomous decision-making capabilities, enabling the generation of carbon asset trading strategies to be independent of human experience.

[0018] Step S50 is the output stage of the real-time optimal carbon asset trading strategy. Its core is to rapidly respond to the real-time carbon asset status, solving the problems of "low efficiency and inability to adapt to dynamic markets" in traditional decision-making, and providing enterprises with timely and accurate trading guidance. Specifically, it collects real-time data on the enterprise's current carbon asset status (such as the latest allowance holdings, CCER inventory, cash balance, and remaining compliance period) and real-time market data (such as current carbon prices and dynamic adjustments to market trading rules). This real-time data is input into a pre-trained reinforcement learning model. Based on its self-learning decision-making logic, the model outputs the current optimal trading strategy (including the timing and scale of allowance purchases / sales, CCER offsetting ratios, and unused allowance carryover schemes), ensuring that enterprises can take timely, compliant, and cost-effective carbon asset trading actions in a dynamically changing carbon market, effectively balancing compliance risks and trading costs.

[0019] In one embodiment, step S20 specifically includes the following steps: S21. Extract carbon trading risk data, carbon asset funding data, and quota compliance data from the carbon element fusion dataset; S22. Based on the carbon price volatility sensitivity and historical trading risk threshold, the carbon trading risk data is quantitatively analyzed to obtain a quantitative value of carbon asset risk preference; S23. Based on the carbon trading budget ratio and capital turnover flexibility, the carbon asset capital data is matched and analyzed to obtain the quantitative value of carbon asset capital constraints. S24. Based on the quota shortfall rate and the weight of the remaining compliance period, perform time-weighted calculation on the quota compliance data to obtain a quantitative value of carbon asset compliance urgency. S25. The carbon asset risk preference quantification value, carbon asset capital constraint quantification value, and carbon asset compliance urgency quantification value are fused to obtain carbon asset decision characteristics, and a corporate carbon asset profile is constructed based on the carbon asset decision characteristics.

[0020] In this embodiment, as described in steps S21-S25 above, the core is a process of quantifying enterprise carbon asset characteristics by classifying and extracting core carbon element data, quantifying and calculating three major decision dimensions, and constructing a multi-dimensional feature fusion profile. This addresses the technical pain points of traditional carbon trading decisions, which fail to incorporate the unique characteristics of enterprises, have highly generalized strategies, and lack specificity. It transforms the enterprise's vague risk tolerance, financial situation, and compliance needs into calculable and modelable quantitative characteristics, forming a unique enterprise carbon asset profile. This provides a personalized, quantifiable, and implementable core basis for the precise definition of the state space, action space, and reward function of the reinforcement learning environment in subsequent step S30, ensuring that the generated strategies align with the enterprise's actual operations and risk tolerance. Specifically: Step S21 is the core data classification and extraction step for carbon elements. The core of this step is to extract three key types of data required for decision-making from the unified fusion data, addressing the problem of mixed multi-source data that cannot be directly used for feature quantification. Specifically, based on the carbon element fusion dataset generated in step S10, it is precisely split and extracted according to the decision-making dimension: carbon trading risk data (for risk preference assessment), carbon asset funding data (for funding constraint assessment), and quota compliance data (for compliance urgency assessment). These three types of data correspond one-to-one with the content of the fusion dataset. After extraction, they are categorized and organized, providing a clear and single-dimensional data foundation for subsequent itemized quantification calculations.

[0021] Step S22 is the quantitative analysis of carbon asset risk appetite. Its core is to transform a company's risk tolerance into a standardized value, addressing the problem of ambiguous risk characteristics that cannot be used for model calculations. Specifically, the quantitative analysis involves: using carbon price volatility sensitivity (a company's tolerance for carbon price fluctuations) as the volatility risk coefficient and historical trading risk thresholds (the maximum acceptable trading risk for a company in the past) as the risk ceiling benchmark, normalizing the carbon trading risk data into a score. This involves calculating the ratio of the actual carbon price fluctuation to the volatility sensitivity, then comparing and scoring it against the historical trading risk thresholds. Through numerical mapping, the analysis results are transformed into standardized values ​​within the [0,1] range, ultimately yielding a quantitative value for carbon asset risk appetite. The magnitude of this value directly represents the level of a company's risk tolerance, providing a quantitative basis for subsequent action space constraints.

[0022] Step S23 is the carbon asset funding constraint matching analysis stage. Its core is to perform a compatibility calculation between the company's financial strength and trading needs to obtain quantitative indicators of funding that can be directly used for model constraints. Specifically, the matching analysis refers to: using the carbon trading budget ratio (the proportion of carbon trading funds to the company's total working capital) as the funding scale benchmark and the capital turnover elasticity (the company's ability to flexibly allocate funds for carbon trading) as the adjustable space coefficient, performing hierarchical matching and weighted calculations on the carbon asset funding data: determining the level of funds the company can invest based on the budget ratio, and determining the floating ratio based on the capital turnover elasticity. After matching and calculating these two, they are mapped to standardized values ​​within the range of [0,1] to obtain the quantitative value of carbon asset funding constraints. The smaller the value, the stronger the funding constraint and the lower the tradable amount, providing funding constraints for subsequently limiting the trading scope.

[0023] Step S24 is the time-weighted calculation of carbon asset compliance urgency. Its core is the introduction of a time dimension, highlighting the increasing urgency as the compliance deadline approaches, thus dynamically quantifying compliance pressure. Specifically, the time-weighted calculation involves: first, calculating the quota gap rate based on the quota shortfall and the total compliance quota; then, setting dynamic time weights based on the remaining compliance period (the closer to the compliance deadline, the larger the weight coefficient); multiplying and summing the quota gap rate with the corresponding time weights to complete the time-weighted calculation, yielding a quantified value of carbon asset compliance urgency. A higher value indicates greater urgency, providing a core basis for setting the subsequent reward function coefficients.

[0024] Step S25 involves the fusion of carbon asset decision-making features and the construction of a corporate profile. The core of this step is to integrate three types of quantitative features into a unified decision-making feature, forming a unique carbon asset profile to address the issue that a single feature cannot fully reflect the enterprise's decision-making attributes. Specifically, the quantitative values ​​of carbon asset risk preference, funding constraints, and compliance urgency are vector-fused and normalized to form a complete and logically consistent carbon asset decision-making feature. Then, using this decision-making feature as the core, a corporate carbon asset profile is constructed that comprehensively represents the enterprise's carbon asset operating status, risk tolerance, funding constraints, and compliance needs. This profile will be directly used in the definition of the three core elements of the reinforcement learning environment in step S30, enabling the customization of the decision-making model for the enterprise.

[0025] In one embodiment, step S30 specifically includes the following steps: S31. Construct a carbon asset trading reinforcement learning environment including a dynamic simulation module, a rule update module, and a compliance constraint module based on preset carbon trading rules. The preset carbon trading rules include carbon market trading rules, quota compliance cycle rules, and trading rule adjustment mechanisms. S32. Based on the decision-making characteristics of carbon assets and combined with external carbon market correlation data, define the state space of the reinforcement learning environment; S33. Based on the quantitative values ​​of carbon asset risk preference and carbon asset capital constraint, limit the scope of quota trading, CCER offset ratio and quota carry-over scope to define the action space of the reinforcement learning environment. S34. Taking corporate compliance goals and cost optimization needs as objectives, and based on the quantifiable value of carbon asset compliance urgency, set compliance reward coefficients, cost penalty coefficients, and risk control weights, and define the reward function for the reinforcement learning environment.

[0026] In this embodiment, as described in steps S31-S34 above, the core is a customized reinforcement learning environment process: building a customized reinforcement learning environment → defining the enterprise's personalized state space → limiting the action space with risk and capital constraints → designing a multi-objective-oriented reward function. This addresses the technical pain points of traditional intelligent decision-making methods, such as the lack of customized environments adapted to the actual situation of enterprises, the generalization of state / action / reward functions, and the difficulty in aligning models with enterprise needs and market rules. It deeply binds the personalized characteristics of the enterprise's carbon asset profile with real carbon market rules, constructing a "company-specific + market-adapted" reinforcement learning environment. This provides a precise, dynamic, and realistic simulation scenario for model training in step S40, ensuring that the decision logic learned by the model not only conforms to the enterprise's risk tolerance and financial situation but also meets the carbon market trading rules and compliance requirements. Specifically, as follows: Step S31 involves building a customized carbon asset trading reinforcement learning environment. Its core function is to simulate the entire process and rule constraints of real carbon trading, addressing the problems of static traditional reinforcement learning environments and their disconnect from market rules. This provides a high-fidelity simulation scenario for model training. Specifically, the preset carbon trading rules strictly adhere to the actual operational logic of the carbon market, including carbon market trading rules (such as quota trading methods, trading hours, and minimum trading units), quota compliance cycle rules (such as compliance deadlines and quota calculation standards), and trading rule adjustment mechanisms (such as updates to market regulatory rules and changes to quota allocation mechanisms). The constructed reinforcement learning environment comprises three core modules: ① Dynamic Simulation Module: Simulates real-world trading scenarios such as carbon price fluctuations, transaction matching, and quota transfers; in the dynamic simulation module, carbon price fluctuations are implemented using a geometric Brownian motion stochastic process, with the specific formula as follows:

[0027] in, For differential operators, for Carbon price at all times For carbon price drift rate, For time infinitesimal element, The carbon price volatility is derived from the standard deviation of historical trading prices in the national carbon market over the past three years, and is based on publicly available and calculable data. For standard Wiener process (Brownian motion), The differential increment of the standard Wiener process; The transaction matching adopts the open matching rules of price priority and time priority, and the quota transfer is calculated in real time according to the quota holding change logic within the performance period; ② Rule update module: Real-time adaptation to trading rule adjustment mechanism to ensure that environmental rules are synchronized with market dynamics; ③ Performance constraint module: Rigidly implement quota performance cycle rules and limit the boundaries of performance-related actions; The three elements work together to form a reinforcement learning environment that is "dynamically updatable, rule-adaptable, and performance-constrained," thereby overcoming the limitations of traditional static environments.

[0028] This reinforcement learning environment can be precisely integrated with subsequent model training. It can convert carbon price fluctuation data, transaction matching results, and quota transfer status output by the dynamic simulation module into state feature vectors that the model can recognize in real time. The rule update module can update the constraints of model training in sync, and the compliance constraint module can directly limit the compliance boundaries of the model's output actions. This enables the reinforcement learning environment and model training to form a closed-loop linkage, providing a highly realistic, strongly constrained, and dynamically adaptable training foundation for subsequent model training, and ensuring that the model training scenario is highly consistent with the real carbon trading scenario.

[0029] Step S32 defines the enterprise's personalized state space. Its core purpose is to enable the reinforcement learning environment to accurately perceive the enterprise's internal state and external market dynamics, addressing the problem of traditional state spaces being too generic and unable to reflect the enterprise's personalized characteristics. Specifically, it uses the carbon asset decision-making characteristics obtained in step S25 (quantified values ​​of carbon asset risk appetite, capital constraints, and compliance urgency) as the core, combined with external carbon market-related data (such as real-time carbon prices, price fluctuation trends, industry quota supply and demand, and market rule adjustments) to construct a multi-dimensional state vector. The state space specifically covers two dimensions: "internal enterprise state" (risk appetite level, available funds, current quota gap / surplus, and remaining compliance time) and "external market state" (current carbon price, price fluctuation range, and market rule effectiveness). This ensures that the environment can comprehensively and in real-time capture key information affecting carbon trading decisions, providing the model with accurate state input.

[0030] Step S33 is the risk-capital dual-constraint action space limitation step. Its core is to limit the boundaries of the transaction actions that the model can execute, so as to avoid the model outputting strategies that exceed the company's risk tolerance and capital capacity, thus solving the problems of unconstrained action space and poor strategy feasibility in traditional models. Specifically, the system employs a dual constraint mechanism: a quantitative value for carbon asset risk appetite (risk tolerance) and a quantitative value for carbon asset funding constraints (available funding). This mechanism limits the scope of quota trading by: ① determining the maximum volume of quota purchases / sales based on the funding constraint value (e.g., a funding constraint value of 0.3 corresponds to a maximum purchase volume of no more than 5,000 tons), and limiting the trading frequency based on the risk appetite value (e.g., a conservative enterprise with a risk appetite value of 0.2 can trade no more than twice a day); ② limiting the CCER offset ratio by: setting the maximum CCER offset ratio based on risk appetite and funding constraints, combined with the quota compliance cycle rules (e.g., for enterprises with tight funding and conservative risk tolerance, the offset ratio is capped at 30%); ③ limiting the quota carryover scope by: setting the carryover period (e.g., a maximum carryover of one compliance cycle) and the carryover volume (e.g., the carryover volume cannot exceed 20% of the current holdings) based on the constraints. This integrated approach ensures that all executable actions within the enterprise's capabilities and risk tolerance are within the enterprise's capacity and risk tolerance.

[0031] Step S34 is the design stage of the compliance-cost-risk multi-objective-oriented reward function. The core is to guide the model to learn the decision logic of "compliance first, cost best, and risk controllable", which solves the problem of the traditional reward function being too singular and unable to balance the needs of multiple objectives. Specifically, guided by the core objectives of corporate compliance and cost optimization, the system dynamically sets parameters based on the quantifiable value of carbon asset compliance urgency: ① Compliance reward coefficient: positively correlated with the quantifiable value of compliance urgency (e.g., the coefficient is set to 1.2 when urgency is 0.8, and 0.6 when urgency is 0.3), encouraging the model to prioritize compliance goals during critical compliance periods; ② Cost penalty coefficient: negatively correlated with the quantifiable value of compliance urgency (e.g., the coefficient is set to 0.4 when urgency is 0.8, avoiding excessive costs due to rush purchases; the coefficient is set to 1.0 when urgency is 0.3, strictly controlling transaction costs during normal periods); ③ Risk control weight: set according to carbon trading risk management requirements (e.g., fixed at 0.3), used to balance the impact of compliance rewards and cost penalties, preventing the model from ignoring risks in pursuit of compliance or violating rules to control costs; Substituting these three factors into a preset formula (e.g., reward value = compliance achievement × compliance reward coefficient - transaction cost × cost penalty coefficient × risk control weight), the resulting reward function can accurately guide the model to make optimal decisions balancing compliance, cost, and risk under different compliance stages and different risk and financial conditions.

[0032] In one embodiment, step S32 specifically includes the following steps: S321. Extract the quantitative values ​​of carbon asset risk preference, carbon asset funding constraint, and carbon asset compliance urgency from the carbon asset decision-making characteristics. S322. Obtain external carbon market correlation data, including carbon price fluctuation data, quota gap data, and funding constraint data; S323. The carbon asset decision-making characteristics are correlated and fused with external carbon market related data to obtain a carbon trading status information set, and the state space of the reinforcement learning environment is defined based on the carbon trading status information set.

[0033] In this embodiment, as described in steps S321-S323 above, the core is the refined construction process of the state space through the extraction of personalized carbon asset decision-making features of enterprises → collection of external carbon market related data → fusion of internal and external multi-dimensional data. This addresses the technical pain points of traditional reinforcement learning state spaces, such as the disconnect between internal features and external market data, one-sided state representation, and inability to fully reflect the carbon trading decision-making environment. It organically combines three types of personalized quantitative features of the enterprise—risk, capital, and compliance—with real-time dynamic data from the external carbon market to form a complete and coordinated set of carbon trading state information. This defines a reinforcement learning state space that better reflects the actual situation of enterprises and the realities of the market, providing comprehensive, accurate, and highly recognizable state inputs for subsequent model training, further improving the model's decision-making perception capabilities and the accuracy of strategy generation. Specifically, as follows: Step S321 is the extraction of personalized carbon asset decision-making features for enterprises. Its core is to extract three key quantitative features that characterize the enterprise's own attributes, providing a personalized internal benchmark for the state space. Specifically, from the carbon asset decision-making features generated in step S25, quantitative values ​​for carbon asset risk appetite, carbon asset funding constraints, and carbon asset compliance urgency are extracted. These three types of values ​​have been standardized and normalized, and can intuitively and quantitatively reflect the enterprise's carbon asset risk tolerance, funding constraints, and compliance urgency. As a core component reflecting the enterprise's own situation in the state space, this ensures that the state space possesses unique personalized attributes specific to the enterprise.

[0034] Step S322 involves acquiring external carbon market-related data. The core of this step is collecting key external market data that influences carbon trading decisions, providing an external market environment benchmark for the state space. Specifically, this includes acquiring external carbon market-related data related to carbon asset trading, such as carbon price fluctuation data (e.g., real-time carbon prices, price fluctuations, and trends), quota gap data (e.g., overall industry quota supply and demand gap, regional quota surplus), and funding constraint data (e.g., average market transaction costs, industry capital turnover level). This data objectively reflects the current external carbon market environment, complementing the company's internal characteristics and avoiding a lack of environmental awareness due to the state space relying solely on internal data.

[0035] Step S323 involves the integration of internal and external data and the definition of the state space. Its core is to unify and integrate the enterprise's internal personalized characteristics with external market data to construct a comprehensive, collaborative, and modelable state space. Specifically, the extracted three types of carbon asset decision-making features are aligned dimensionally, correlated with features, and fused with the acquired external carbon market data. This eliminates discrepancies in data definitions and logical fragmentation, forming a carbon trading state information set that includes both the enterprise's internal personalized state and the external market environment state. Based on this state information set, each fused feature vector is defined as a state in the reinforcement learning environment. This constructs a complete state space covering both the enterprise's own attributes and market dynamics, enabling the reinforcement learning model to simultaneously perceive internal decision-making conditions and the external market environment, significantly improving the completeness of state representation and the rationality of decision-making.

[0036] In one embodiment, step S33 specifically includes the following steps: S331. Extract the quantitative values ​​of carbon asset risk preference and carbon asset funding constraints from the carbon asset decision-making characteristics as the first constraint condition; S332. Based on the first constraint condition, constrain the magnitude range of quota buying and selling to obtain the quota trading range; S333. Based on the first constraint and the quota compliance cycle rules, constrain the upper limit of the proportion of CCER that can be used to offset carbon emissions, and obtain the CCER offset ratio. S334. Based on the first constraint and the quota fulfillment cycle rules, constrain the carry-over period and carry-over amount of unused quotas in the current period to obtain the quota carry-over range; S335. Integrate the quota transaction range, CCER offset ratio, and quota carry-over range to define the action space of the reinforcement learning environment.

[0037] In this embodiment, as described in steps S331-S335 above, the core is a personalized action space construction process that extracts risk and capital constraints → limits the scope of quota trading → constrains CCER offset ratios → controls the scope of quota carryover → integrates multi-dimensional actions. This process addresses the technical pain points of traditional reinforcement learning action spaces, which fail to consider enterprise risk tolerance and financial status, suffer from generalized action ranges, and are prone to generating trading actions beyond the enterprise's capabilities or non-compliant. By using the enterprise's own risk appetite and capital constraints as the core limiting conditions, and combining quota compliance cycle rules, precise boundary limits are defined for various carbon trading-related actions. This constructs a reinforcement learning action space that is enterprise-specific, risk-controllable, capital-suitable, and compliant, ensuring that the model can only execute actions that conform to the enterprise's actual capabilities and market rules during training and strategy generation. This further enhances the feasibility and security of the trading strategy, as detailed below: Step S331 is the first constraint extraction stage, the core of which is to determine the key basis for limiting the action space and solve the problem of the lack of enterprise-specific constraint benchmarks for the action space. Specifically, the standardized quantitative values ​​of carbon asset risk appetite and carbon asset funding constraints are extracted from the carbon asset decision-making characteristics, and these two are used together as the first constraint. This constraint directly reflects the enterprise's risk tolerance and the scale of its available funds, and is the fundamental basis for limiting the boundaries of all subsequent transaction actions, ensuring that the action space is consistent with the enterprise's own attributes from the source.

[0038] Step S332 is the quota trading range constraint step, the core of which is to limit the reasonable scale of quota buying and selling to avoid situations where the trading volume exceeds the company's capital capacity or risk tolerance. Specifically, the maximum and minimum scales of quota buying and selling are determined based on the quantitative value of capital constraint in the first constraint condition, and the trading amplitude and frequency are adjusted in combination with the quantitative value of risk preference. This results in a quota trading range that suits the company's capital and risk level, ensuring the rationality and controllability of trading actions in terms of trading volume.

[0039] Step S333 is the CCER offset ratio constraint step. Its core is to limit the upper limit of CCER offset usage under the premise of compliance, balancing compliance needs and risk control. Specifically, based on the first constraint condition and combined with the quota compliance cycle rules, an upper limit is imposed on the proportion of CCERs that can be used to offset carbon emissions, resulting in the CCER offset ratio. This ensures that enterprises can meet compliance requirements through CCERs while avoiding compliance and cost risks due to excessively high offset ratios, ensuring that CCER usage aligns with the enterprise's actual situation and market rules.

[0040] Step S334 is the constraint step for the scope of quota carryover. The core of this step is to regulate the carryover behavior of unused quotas, ensuring that quota management is compliant and adaptable to the company's operational pace. Specifically, based on the first constraint condition and the quota performance cycle rules, the carryover period (such as the number of performance cycles allowed for carryover) and the carryover volume (such as the maximum carryover quota percentage) of unused quotas in the current period are subject to dual constraints to obtain the scope of quota carryover. This avoids performance risks caused by disorderly quota carryover and ensures that quota carryover actions are adapted to the company's financial and risk situation within the rule framework.

[0041] Step S335 is the multi-dimensional action integration and action space definition stage. Its core is to integrate the range of actions under various constraints into a unified, executable action space. Specifically, it standardizes and integrates the limited quota trading range, CCER offset ratio, and quota carry-over range to form a complete set covering the three core carbon trading actions: quota trading, CCER offset, and quota carry-over. This defines the action space of the reinforcement learning environment. All actions within this space meet the requirements of controllable risk, adequate funding, and compliance, providing clear, reasonable, and feasible action options for model training.

[0042] In one embodiment, step S34 specifically includes the following steps: S341. Set the urgency quantification value of carbon asset compliance as the second constraint, and set corporate compliance targets and cost optimization requirements; S342. Based on the second constraint and the established corporate compliance targets, set a compliance reward coefficient that is positively correlated with the quantified value of carbon asset compliance urgency. S343. Based on the second constraint and the set cost optimization requirements, set a cost penalty coefficient that is negatively correlated with the quantified value of carbon asset compliance urgency. S344. Based on the second constraint and carbon trading risk management requirements, set risk control weights to balance the impact of compliance reward coefficients and cost penalty coefficients. S345. Substitute the compliance reward coefficient, cost penalty coefficient, and risk control weight into a preset formula to obtain the reward function of the reinforcement learning environment, wherein the reward function is output-oriented to maximize compliance rewards and minimize cost penalties.

[0043] In this embodiment, as described in steps S341-S345 above, the core is a customized reward function design process that involves setting compliance urgency constraints → dynamically configuring dual-objective coefficients → balancing risk weights → constructing a multi-objective reward function. This addresses the technical pain points of traditional reinforcement learning reward functions, such as their single objective, inability to dynamically adapt to compliance cycles, and difficulty in balancing compliance and transaction costs. Using carbon asset compliance urgency as the core dynamic constraint, and focusing on the dual objectives of corporate compliance and cost optimization, differentiated reward, penalty, and balancing coefficients are set to construct a multi-objective-oriented reward function that prioritizes compliance, optimizes costs, and controls risks. This provides precise optimization guidance for the reinforcement learning model, allowing the model to autonomously learn carbon trading decision-making logic that meets the actual needs of enterprises during iterative training. Specifically, as follows: Step S341 is the second constraint and optimization objective setting stage. Its core is to determine the core constraint benchmark and optimization direction of the reward function, addressing the issue of the reward function lacking a personalized objective orientation for enterprises. Specifically, the quantifiable value of carbon asset compliance urgency obtained in step S24 is set as the second constraint of the reward function. This value directly determines the dynamic magnitude of the reward and penalty coefficients. Simultaneously, the core objectives of enterprise carbon trading are clarified: compliance objectives (ensuring quota fulfillment and avoiding violations) and cost optimization needs (controlling transaction costs and reducing compliance expenses). This serves as the fundamental guideline for the design of the reward function, ensuring a high degree of alignment between the reward function and the enterprise's core needs.

[0044] Step S342 involves setting the compliance reward coefficient. Its core function is to use positive incentives to guide the model to prioritize compliance and address the issue of insufficient compliance priority during critical compliance periods. Specifically, the compliance reward coefficient is positively correlated with the quantifiable value of carbon asset compliance urgency: the higher the urgency, the closer the company is to the compliance deadline and the greater the compliance pressure, resulting in a higher compliance reward coefficient; conversely, the lower the urgency, the lower the compliance reward coefficient. This incentivizes the model to prioritize actions that achieve compliance goals during critical compliance phases, ensuring the company's compliance baseline.

[0045] Step S343 involves setting the cost penalty coefficient, the core of which is to control transaction costs through a negative penalty constraint model to address the problem of unreasonable cost management at different compliance stages. Specifically, the cost penalty coefficient is negatively correlated with the quantified value of carbon asset compliance urgency: when compliance urgency is high, the cost penalty coefficient is appropriately reduced to avoid failure to comply on time due to excessive cost control; when compliance urgency is low, the cost penalty coefficient is increased to strictly constrain the model to reduce transaction costs and avoid unnecessary expenditures, thus achieving a dynamic balance between compliance urgency and cost management.

[0046] Step S344 involves setting risk control weights, the core of which is balancing the impact of compliance incentives and cost penalties to address the issues of extreme model decisions and uncontrolled risks. Specifically, based on carbon trading risk management requirements and the second constraint, risk control weights are set to adjust the overall impact of compliance reward coefficients and cost penalty coefficients. This prevents the model from ignoring trading risks in pursuit of compliance rewards, and also prevents the model from failing to meet compliance requirements in order to reduce cost penalties, thus ensuring a reasonable balance between compliance, cost, and risk in model decisions.

[0047] Step S345 is the multi-objective reward function construction stage. Its core is to integrate various coefficients to form a standardized reward function, providing clear optimization criteria for model training. Specifically, the compliance reward coefficient, cost penalty coefficient, and risk control weight are substituted into the following preset calculation formula, and the final reinforcement learning reward function is obtained through weighted calculation:

[0048] in, The final reward value obtained after the model performs a single transaction action; The compliance achievement level (value range [0,1], where 1 represents full compliance and 0 represents complete non-compliance). The compliance reward coefficient set for step S342 (with a value range of [0.5, 2.0], and the coefficient increases with the higher the urgency of compliance). The relative transaction cost (value range [0,1], where 1 represents the transaction cost reaching the preset upper limit, and 0 represents the transaction cost being 0); The cost penalty coefficient set for step S343 (within the range of [0.3, 1.2], and the coefficient decreases as the urgency of fulfillment increases); The risk control weight set for step S344 (fixed value such as 0.3, or dynamically adjusted according to the company's risk appetite).

[0049] This reward function is designed to maximize compliance rewards and minimize cost penalties. Each time the model executes a transaction, it receives a reward based on this formula: higher compliance achievement results in a larger reward; higher transaction costs result in a smaller reward. Risk control weights further buffer the impact of cost penalties, preventing extreme decisions. This guides the model to continuously iterate and optimize towards the optimal decision-making direction of compliance, low cost, and low risk.

[0050] In one embodiment, step S40 specifically includes the following steps: S41. Obtain the completed carbon asset trading reinforcement learning environment, as well as the defined state space, action space, and reward function; S42. Initialize the pre-built carbon trading reinforcement learning model and train the model with corporate compliance, cost optimization and risk control as training objectives. S43. Based on the state space and action space, execute carbon trading strategy interaction in the carbon asset trading reinforcement learning environment, and calculate the corresponding reward feedback according to the reward function; S44. Iteratively update the model parameters of the carbon trading reinforcement learning model based on the reward feedback; S45. Repeat the strategy interaction and model parameter update steps until the carbon trading reinforcement learning model meets the preset convergence condition, and obtain the trained carbon trading reinforcement learning model.

[0051] In this embodiment, as described in steps S41-S45 above, the core is an adaptive model iterative training process that involves acquiring the reinforcement learning environment and core elements → model initialization and target setting → strategy interaction and reward feedback calculation → iterative updating of model parameters → convergence judgment and training completion. This process addresses the technical pain points of traditional carbon trading decision-making methods, which lack autonomous learning optimization mechanisms, rely on human experience, and have decision logic that is difficult to adapt to dynamic markets. By relying on a pre-customized reinforcement learning environment and precisely defined states, actions, and reward functions, the carbon trading reinforcement learning model can autonomously interact, experiment, and optimize in a simulated scenario. Ultimately, this results in a trained model with "compliance awareness, optimal cost, and controllable risk," providing stable and intelligent core model support for subsequent real-time optimal carbon asset trading strategy output. Specifically, as follows: Step S41 involves acquiring the reinforcement learning environment and core elements. Its core purpose is to build upon the previously customized results, establishing a complete foundation for model training and addressing the issue of insufficient suitable scenarios and rules for model training. Specifically, it involves acquiring the carbon asset trading reinforcement learning environment constructed in step S31, as well as the state space, action space, and reward function defined in steps S32-S34. These environments and elements are already tailored to the specific characteristics of enterprises and the real rules of the carbon market, requiring no further generalization adjustments. They serve directly as the complete foundation for model training, ensuring a high degree of consistency between the training scenario and the actual needs of enterprises.

[0052] Step S42 is the reinforcement learning model initialization and training objective setting stage. Its core is to complete the initial model configuration and clarify the optimization direction, addressing the problems of undirected model training and unreasonable initial states. Specifically, the pre-built carbon trading reinforcement learning model refers to an initial neural network model built based on a public reinforcement learning framework. Pre-building specifically means: using the PPO (Proximity Point Optimization) network as the basic architecture, completing the basic model construction process of network layer definition, activation function setting, and initial weight random initialization; the pre-training data uses publicly available historical trading data from the national carbon market, anonymized corporate carbon compliance data, and publicly available quota transfer data. The pre-training process involves: converting the publicly available historical data into standardized state vectors, performing unsupervised pre-training on the initial model, learning the basic feature distribution of carbon trading, and obtaining a pre-built model with stable initial weights. The pre-built carbon trading reinforcement learning model (PPO) undergoes routine initialization operations such as weight initialization and learning rate setting. The specific network architecture of the model is as follows: input layer (state feature vector dimension × 1) → 2 fully connected hidden layers (128 neurons per layer, ReLU activation function) → output layer (action space dimension × 1, Softmax activation function). The model input format is the standardized state feature vector defined in step S32, and the output format is the probability distribution vector of each trading action in the action space. The model input is a standardized state vector normalized to the [0,1] interval, and the model output is a probability distribution vector with the same length as the action space. By selecting the action index corresponding to the maximum probability, and based on the unique correspondence between the preset action index and the specific trading action, the specific executable trading action is directly mapped to achieve accurate mapping from state to action. At the same time, corporate compliance, cost optimization, and risk control are set as the core training objectives of the model, so that the model always optimizes decisions around the three objectives throughout the entire iteration process, ensuring the rationality and applicability of the model output strategy from the training source.

[0053] Step S43 is the strategy interaction and reward feedback calculation stage. Its core is to allow the model to execute decision-making actions in a simulated environment and obtain optimization guidance, addressing the problem that the model cannot autonomously perceive the merits of its decisions. Specifically, based on the enterprise and market state information in the current state space, the model selects and executes corresponding carbon trading actions within a limited action space. The state-action mapping is implemented as follows: the state feature vector is input into the model network, the action probability distribution is output, the action index with the highest probability is selected, and mapped to the corresponding quota trading, CCER offsetting, and quota carry-over specific trading actions in the action space, completing the strategy interaction with the reinforcement learning environment. Then, based on the reward function constructed in step S34, the compliance, cost level, and risk level of this trading action are comprehensively calculated to obtain the corresponding reward feedback value. This reward value directly represents the merits of the current decision-making action, providing a quantitative basis for model parameter updates. The strategy interaction process is a closed-loop interaction process between the reinforcement learning agent and the carbon trading simulation environment. The state space is composed of enterprise and market state information, and the action space is composed of quota trading, CCER offsetting, and quota carry-over. The state-action mapping is realized through forward inference of the model. The reward feedback value comprehensively represents the quality of the decision action and provides a quantitative basis for subsequent parameter iteration and update.

[0054] Step S44 is the model parameter iterative update stage, the core of which is to optimize the model's decision-making logic based on reward feedback, thereby achieving adaptive learning and upgrading of the model. Specifically, based on the reward feedback value obtained in step S43, the parameter update mechanism corresponding to reinforcement learning (PPO policy gradient update) is used to adjust the network weights, decision thresholds, and other parameters of the carbon trading reinforcement learning model. This makes the model more inclined to choose trading actions that can obtain higher rewards in the next interaction, gradually correcting and optimizing its own carbon trading decision-making logic, and improving the accuracy of strategy generation. The parameter iterative update process is implemented based on the proximal policy optimization (PPO) algorithm. By calculating the policy gradient to update the model's network weights and decision thresholds, the model's output decision actions are gradually optimized towards higher rewards, achieving adaptive learning and upgrading of the model's decision logic.

[0055] Step S45 is the iterative and convergence determination stage. Its core is to continuously train the model to reach a stable optimal state, resulting in a final, usable model. Specifically, the strategy interaction in step S43 and the parameter update process in step S44 are repeated until the carbon trading reinforcement learning model meets the preset convergence conditions (e.g., the model reward value tends to stabilize, the loss function fluctuation is less than a preset threshold, and the strategy shows no significant change in consecutive iterations). At this point, the model has fully learned the optimal decision logic adapted to the enterprise characteristics and market rules. Training stops, and the trained carbon trading reinforcement learning model is obtained, which can be directly used to output the optimal trading strategy under subsequent real-time carbon asset conditions. The iterative process constitutes the complete training loop of the reinforcement learning model. Through multiple strategy interactions and parameter updates, the model gradually learns the characteristics of the carbon trading environment and decision rules until the model reward value tends to stabilize, the loss function fluctuation is less than a preset threshold, and the strategy shows no significant change in consecutive iterations, thus meeting the convergence conditions and obtaining a stable, optimal final model.

[0056] Furthermore, based on the trained carbon trading reinforcement learning model obtained from steps S41 to S45 above, the real-time carbon asset status data can be processed according to the following reasoning logic to output the optimal carbon asset trading strategy. The specific reasoning execution process is as follows: (1) According to the data collection requirements of step S50, acquire the real-time carbon asset status data of the enterprise and the real-time data of the external carbon market, and standardize and organize them according to the data format of the state space to form a real-time state feature vector that the model can directly identify. (2) Input the real-time state feature vector into the carbon trading reinforcement learning model that has been trained. Based on the decision logic learned through iterative training, the model traverses and evaluates all feasible carbon trading actions within the compliance action space defined in step S33. (3) The model combines the optimization guidance of the reward function to predict and calculate the expected reward value of each action, and automatically selects the optimal transaction action that can simultaneously meet the three major objectives of performance compliance, cost optimization and risk control. (4) Specific implementation of model output mapping: The optimal action index output by the model is mapped one-to-one to the specific trading action preset in the action space. The selected optimal trading action is parsed and transformed into the optimal carbon asset trading strategy that can be directly executed. Specifically, it includes clear decision-making content such as the timing and magnitude of quota buying / selling, CCER offset ratio, and unused quota carry-over scheme, and completes the output of the optimal carbon asset trading strategy. The reasoning process is the forward reasoning process of the trained model. Through real-time state feature vector input, compliance action space traversal, expected reward value prediction, optimal action screening, and action index mapping, the direct output of real-time carbon asset state to optimal carbon asset trading strategy is realized, ensuring the compliance, optimality and executability of the strategy output.

[0057] In one embodiment, a carbon asset trading strategy generation system based on multi-source information fusion and reinforcement learning is provided. This system corresponds to the carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning described in the previous embodiment. The system includes: The data processing module is used to acquire multi-source carbon element data, fuse and preprocess the multi-source carbon element data to obtain a carbon element fusion dataset. The first construction module is used to build a corporate carbon asset profile containing carbon asset decision-making characteristics based on the carbon element fusion dataset. The second construction module is used to build a carbon asset trading reinforcement learning environment according to the preset carbon trading rules, and to define the state space, action space and reward function of the reinforcement learning environment according to the enterprise carbon asset profile. The model training module is used to iteratively train a pre-built carbon trading reinforcement learning model in a carbon asset trading reinforcement learning environment based on a defined state space, action space, and reward function, so as to obtain a trained carbon trading reinforcement learning model. The strategy generation module is used to acquire real-time carbon asset status data and output the optimal carbon asset trading strategy through a trained carbon trading reinforcement learning model.

[0058] Specific limitations regarding the carbon asset trading strategy generation system based on multi-source information fusion and reinforcement learning can be found in the limitations of the carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning described above, and will not be repeated here. Each module in the aforementioned carbon asset trading strategy generation system based on multi-source information fusion and reinforcement learning can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0059] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 2As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database is used for data storage, data processing, and data analysis. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning.

[0060] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning.

[0061] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements a method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning.

[0062] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0063] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0064] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating carbon asset trading strategies based on multi-source information fusion and reinforcement learning, characterized in that, Includes the following steps: Acquire multi-source carbon element data, fuse and preprocess the multi-source carbon element data to obtain a carbon element fusion dataset; Based on the carbon element fusion dataset, construct a corporate carbon asset profile that includes carbon asset decision-making characteristics; A carbon asset trading reinforcement learning environment is constructed based on preset carbon trading rules, and the state space, action space, and reward function of the reinforcement learning environment are defined based on the enterprise carbon asset profile. Based on the defined state space, action space, and reward function, the pre-built carbon trading reinforcement learning model is iteratively trained in the carbon asset trading reinforcement learning environment to obtain the trained carbon trading reinforcement learning model. Acquire real-time carbon asset status data and output the optimal carbon asset trading strategy through a trained carbon trading reinforcement learning model.

2. The carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in claim 1, characterized in that, The step of constructing a corporate carbon asset profile containing carbon asset decision-making characteristics based on a carbon element fusion dataset specifically includes the following steps: Extract carbon trading risk data, carbon asset funding data, and quota compliance data from the carbon element fusion dataset; Based on the carbon price volatility sensitivity and historical trading risk threshold, the carbon trading risk data is quantitatively analyzed to obtain a quantitative value of carbon asset risk preference. Based on the proportion of carbon trading budget and the flexibility of capital turnover, the carbon asset capital data is matched and analyzed to obtain the quantitative value of carbon asset capital constraints. Based on the quota shortfall rate and the weight of the remaining compliance period, the quota compliance data is time-weighted to obtain a quantified value of carbon asset compliance urgency. The carbon asset risk preference quantification value, carbon asset funding constraint quantification value, and carbon asset compliance urgency quantification value are fused to obtain carbon asset decision characteristics, and a corporate carbon asset profile is constructed based on the carbon asset decision characteristics.

3. The carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in claim 2, characterized in that, The steps of constructing a carbon asset trading reinforcement learning environment based on preset carbon trading rules and defining the state space, action space, and reward function of the reinforcement learning environment based on the enterprise's carbon asset profile specifically include the following steps: A carbon asset trading reinforcement learning environment is constructed based on preset carbon trading rules, including a dynamic simulation module, a rule update module, and a compliance constraint module. The preset carbon trading rules include carbon market trading rules, quota compliance cycle rules, and trading rule adjustment mechanisms. Based on the decision-making characteristics of carbon assets and combined with external carbon market data, the state space of the reinforcement learning environment is defined. Based on the quantitative values ​​of carbon asset risk appetite and carbon asset capital constraints, the scope of quota trading, CCER offset ratio, and quota carry-over scope are limited to define the action space of the reinforcement learning environment. With the goals of corporate compliance and cost optimization as the objectives, and based on the quantifiable value of carbon asset compliance urgency, a compliance reward coefficient, a cost penalty coefficient, and a risk control weight are set, and a reward function for the reinforcement learning environment is defined.

4. The carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in claim 3, characterized in that, The step of defining the state space of the reinforcement learning environment based on carbon asset decision-making characteristics and combined with external carbon market correlation data specifically includes the following steps: Extract the quantitative values ​​of carbon asset risk preference, carbon asset funding constraint, and carbon asset compliance urgency from the carbon asset decision-making characteristics; Obtain external carbon market correlation data, including carbon price fluctuation data, quota gap data, and funding constraint data; The carbon asset decision-making characteristics are correlated and fused with external carbon market data to obtain a carbon trading status information set, and the state space of the reinforcement learning environment is defined based on the carbon trading status information set.

5. The carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in claim 3, characterized in that, The step of defining the action space of the reinforcement learning environment by limiting the scope of quota trading, the CCER offset ratio, and the quota carry-over scope based on the quantitative values ​​of carbon asset risk preference and carbon asset capital constraints specifically includes the following steps: The quantitative values ​​of carbon asset risk preference and carbon asset funding constraints in carbon asset decision-making characteristics are extracted as the first constraint condition. Based on the first constraint, the range of the quota purchase and sale quantities is constrained to obtain the quota trading range; Based on the first constraint and the quota compliance cycle rules, the upper limit of the proportion of CCERs that can be used to offset carbon emissions is constrained, and the CCER offset ratio is obtained. Based on the first constraint and the quota fulfillment cycle rules, the carry-over period and carry-over amount of unused quotas in the current period are constrained to obtain the quota carry-over range; The quota transaction range, CCER offset ratio, and quota carry-over range are integrated to define the action space of the reinforcement learning environment.

6. The carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in claim 3, characterized in that, The steps described above, which aim to balance corporate compliance goals and cost optimization needs, and define the reward function for the reinforcement learning environment by setting compliance reward coefficients, cost penalty coefficients, and risk control weights based on the quantifiable value of carbon asset compliance urgency, specifically include the following steps: Set the urgency quantification of carbon asset compliance as the second constraint, and set corporate compliance targets and cost optimization needs; Based on the second constraint and the established corporate compliance targets, a compliance reward coefficient that is positively correlated with the metric value of carbon asset compliance urgency is set. Based on the second constraint and the set cost optimization requirements, a cost penalty coefficient that is negatively correlated with the quantified value of carbon asset compliance urgency is set. Based on the second constraint and carbon trading risk management requirements, risk control weights are set to balance the impact of compliance reward coefficients and cost penalty coefficients. Substituting the compliance reward coefficient, cost penalty coefficient, and risk control weight into a preset formula yields the reward function of the reinforcement learning environment, wherein the reward function is output-oriented towards maximizing compliance rewards and minimizing cost penalties.

7. The carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in claim 1, characterized in that, The step of iteratively training a pre-built carbon trading reinforcement learning model in a carbon asset trading reinforcement learning environment based on a defined state space, action space, and reward function to obtain a trained carbon trading reinforcement learning model specifically includes the following steps: Obtain the completed carbon asset trading reinforcement learning environment, as well as the defined state space, action space, and reward function; Initialize the pre-built carbon trading reinforcement learning model and train the model with corporate compliance, cost optimization and risk control as training objectives; Based on the state space and action space, carbon trading strategy interactions are executed in the carbon asset trading reinforcement learning environment, and corresponding reward feedback is calculated according to the reward function. The model parameters of the carbon trading reinforcement learning model are iteratively updated based on the reward feedback. Repeat the strategy interaction and model parameter update steps until the carbon trading reinforcement learning model meets the preset convergence condition, and obtain the trained carbon trading reinforcement learning model.

8. A carbon asset trading strategy generation system based on multi-source information fusion and reinforcement learning, used to implement the steps of the carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in any one of claims 1-7, characterized in that, include: The data processing module is used to acquire multi-source carbon element data, fuse and preprocess the multi-source carbon element data to obtain a carbon element fusion dataset. The first construction module is used to build a corporate carbon asset profile containing carbon asset decision-making characteristics based on the carbon element fusion dataset. The second construction module is used to build a carbon asset trading reinforcement learning environment according to the preset carbon trading rules, and to define the state space, action space and reward function of the reinforcement learning environment according to the enterprise carbon asset profile. The model training module is used to iteratively train a pre-built carbon trading reinforcement learning model in a carbon asset trading reinforcement learning environment based on a defined state space, action space, and reward function, so as to obtain a trained carbon trading reinforcement learning model. The strategy generation module is used to acquire real-time carbon asset status data and output the optimal carbon asset trading strategy through a trained carbon trading reinforcement learning model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the carbon asset trading strategy generation method based on multi-source information fusion and reinforcement learning as described in any one of claims 1-7.