A Data-Driven Multi-Agent Energy and Carbon Collaborative Decision-Making Method for Low-Carbon Industrial Parks
By constructing a multi-agent dynamic Stackelberg game model and combining it with an improved multi-agent deep reinforcement learning and multi-objective particle swarm optimization algorithm, the collaborative optimization problem of multi-agent decision-making in low-carbon industrial parks was solved. This enabled accurate carbon emission tracing and real-time optimization of economic costs, thereby improving the low-carbon efficiency of park operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2026-04-22
- Publication Date
- 2026-05-26
AI Technical Summary
Existing optimization methods for low-carbon industrial parks struggle to coordinate independent decision-making among multiple stakeholders with system-wide collaborative goals when faced with scenarios involving high proportions of renewable energy access and multiple stakeholders. Their carbon emission accounting is crude and they cannot achieve accurate optimization decisions. Furthermore, traditional optimization algorithms suffer from high computational complexity when dealing with mixed decision-making problems involving high dimensions, nonlinearity, and strong constraints, making it difficult to obtain satisfactory solutions.
A multi-agent dynamic Stackelberg game model is constructed, which introduces time-varying grid-side carbon intensity and green electricity carbon offsetting benefits. Combined with an improved multi-agent deep reinforcement learning and multi-objective particle swarm optimization algorithm, a hybrid intelligent solution framework is formed to achieve accurate source tracing and optimization of carbon responsibility.
It has achieved precise and coordinated optimization of the economic and carbon costs of various entities within the low-carbon park, improved the accuracy and efficiency of decision-making, enabled dynamic responses to fluctuations in energy structure and carbon emissions, and enhanced the economic efficiency and low-carbon nature of park operation.
Smart Images

Figure CN122089347A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-carbon park scheduling and control technology, and in particular to a data-driven multi-entity energy and carbon collaborative decision-making method for low-carbon parks. Background Technology
[0002] With the high penetration of distributed renewable energy, the widespread integration of diverse loads (such as electric vehicles and flexible loads), and the continuous improvement of carbon market mechanisms, the energy system of industrial parks exhibits typical characteristics such as the intertwined interests of multiple stakeholders, the close coupling of energy flow and carbon flow, and strong uncertainty on both the source and load sides. This brings unprecedented challenges to its operation management and decision optimization. Although significant progress has been made in the research on the optimized operation of low-carbon industrial parks, existing theoretical methods still have obvious shortcomings in addressing the following three major challenges in the future scenario of high-proportion renewable energy integration and multi-stakeholder participation: The contradiction between independent decision-making by multiple stakeholders and the goal of system collaboration is difficult to reconcile, and there is a lack of effective methods for modeling and solving game equilibrium. Most existing studies employ centralized optimization or single-agent decision-making models, assuming a global control center capable of unconditionally allocating all resources. This severely contradicts the reality that energy suppliers, users, and other entities within the industrial park have independent property rights and diverse interests. While some studies have introduced game theory, they often simplify the decision-making complexity of the stakeholders and fail to fully characterize their rich strategy space. More importantly, simultaneously considering the dual-objective conflict of economic efficiency and low-carbon sustainability in game models dramatically increases the dimensionality and complexity of decision-making. Formalizing such multi-objective dynamic games and solving for their equilibrium strategies presents a significant challenge. Existing methods struggle to handle this hybrid problem of nested game optimization, leading to strategies that often deviate from the true equilibrium and fail to provide effective guidance for the actual operation of industrial parks.
[0003] Carbon emission accounting methods are static and crude, failing to support accurate carbon cost management and real-time optimization decisions. Most current studies measure carbon emissions using national or regional annual average carbon emission factors. This static method, where each kilowatt-hour represents an emission value, has a fatal flaw: it completely ignores the significant fluctuations in the power grid's energy structure over time and fails to reflect the real-time carbon offsetting benefits of absorbing local distributed green electricity. Consequently, it cannot accurately reflect the true carbon responsibility of various entities within the industrial park at different times, leading to distorted carbon cost calculations. Any optimization decisions made based on this distorted carbon signal will have significantly reduced emission reduction effects and may even be misleading. Therefore, developing a dynamic carbon footprint factor embedding model that can accurately trace sources and update in real time is the primary prerequisite for achieving true energy-carbon synergistic optimization, and this is precisely the weak link in current research.
[0004] The high-dimensional, nonlinear, and strongly constrained hybrid decision-making problem constructed by the park energy optimization method poses an insurmountable computational obstacle for traditional optimization algorithms. The multi-agent game problem integrating dynamic carbon footprints is essentially a high-dimensional (decision variables grow exponentially with the number of users), nonlinear (objective functions and constraints contain numerous nonlinear terms), dual-objective (economic and carbon cost), and strongly constrained (constraints on the operation of various physical equipment and energy balance) mixed integer optimization problem, belonging to the NP-hard category. Traditional mathematical programming methods (such as Mixed Integer Linear Programming (MILP)) either sacrifice accuracy due to model simplification or fail to find a satisfactory solution within an effective timeframe due to exponential computational complexity. Although intelligent algorithms have been applied, traditional single-agent reinforcement learning struggles to handle unsteady multi-agent environments, while common metaheuristic algorithms suffer from inherent defects in solving large-scale, multi-objective, constrained problems, including slow convergence, susceptibility to local optima, and low-quality solutions.
[0005] Therefore, there is an urgent need to develop a novel dynamic game optimization method that can efficiently and stably solve the intelligent decision-making problem of energy and carbon coordination among multiple stakeholders in industrial parks, and approximate the optimal or equilibrium solution. Summary of the Invention
[0006] The purpose of this invention is to solve the multi-agent game problem of integrated dynamic carbon footprint and to provide a multi-agent energy and carbon collaborative decision optimization method for low-carbon industrial parks.
[0007] The objective of this invention can be achieved through the following technical solutions: A multi-stakeholder energy and carbon collaborative decision-making method for low-carbon industrial parks, comprising the following steps: A multi-agent dynamic Stackelberg game model is constructed, and the time-varying grid-side carbon intensity and green electricity carbon offsetting benefits are introduced into the multi-agent game decision-making process to establish a multi-objective optimization model integrating dynamic carbon footprint factors: Energy suppliers, as leaders, formulate and publish pricing strategies; users, as followers, aim to minimize the sum of economic energy costs and carbon costs calculated based on dynamic carbon factors, and solve for the optimal energy use strategy; the power grid, as the external environment maker, provides dynamic grid carbon intensity factors and renewable energy forecasts as external boundary conditions. A hybrid intelligent solution framework combining an improved multi-agent deep reinforcement learning algorithm and an improved multi-objective particle swarm optimization algorithm is used to solve a multi-agent dynamic Stackelberg game model. At the game strategy learning level, the improved multi-agent deep reinforcement learning algorithm is used to approximate the Stackelberg equilibrium. At the individual optimization solution level, the improved multi-objective particle swarm optimization algorithm is embedded in the follower agent to solve the follower's optimal response under the current price strategy. In the improved multi-agent deep reinforcement learning algorithm, all leaders and followers are modeled as agents, each with its own Actor network and Critic network. An attention layer is introduced into the Critic network. In the formula: and These are coding networks; This is the global state; For intelligent agents i The action; For intelligent agents j For intelligent agents i Attention weights.
[0008] As a preferred technical solution, the multi-agent dynamic Stackelberg game model is specifically constructed as follows: The first phase involves leader decision-making: energy suppliers develop and publish pricing strategies based on their forecasts of user response, renewable energy output, and grid carbon intensity. The second stage involves follower decision-making: After observing the pricing strategy, each user aims to minimize the sum of the economic cost of energy use and the carbon cost calculated based on the dynamic carbon factor, and solves for their optimal energy use strategy. The leader's strategy space is the energy price it sets, including the price at which the leader sells electricity to followers, the price at which the leader buys electricity from followers, and the price of natural gas during all periods of the scheduling cycle; the leader's utility function aims to maximize the difference between its own revenue and costs, which include the leader's operating costs and carbon emission costs. The follower's strategy space is a scheduling plan for adjustable loads, including the follower's adjustable load power in the corresponding time period, the follower's energy storage charging / discharging power in the corresponding time period, and the follower's energy storage state of charge in the corresponding time period; the follower's utility function aims to minimize total cost, including the user's economic cost and carbon emission cost.
[0009] As a preferred technical solution, the multi-agent dynamic Stackelberg game model extends the economic goal optimization problem of the game agents into a dual-objective optimization problem of economy and environment by introducing a dynamic carbon footprint factor. The leader's objective function is to minimize the difference between the sum of operating costs and carbon costs and the revenue: In the formula, Operating costs for leaders; The carbon emission costs for leaders; For the benefit of the leader; Among them, the carbon emission costs of leaders It is expressed as follows: In the formula, This refers to the carbon price coefficient. for t The amount of electricity purchased from the grid during a given time period; Let be the carbon intensity coefficient of the power grid during time period t; for t Natural gas purchase capacity during specific time periods; The carbon intensity coefficient of natural gas; for t Total power generation of distributed renewable energy within the park during the specified time period; for t Real-time carbon offset coefficient of green electricity for a given period; The objective function of the followers is to minimize the sum of operating costs and carbon costs; In the formula, The economic cost to followers; The carbon emission costs for followers; Among them, the carbon emission costs of followers It is expressed as follows: In the formula, This refers to the carbon price coefficient. For users i exist t The amount of electricity purchased from the leader during a given period; for t Carbon intensity coefficient of the power grid during a given time period; For users i exist t The amount of gas purchased from the leader during the specified period; This is the carbon intensity coefficient of natural gas.
[0010] As a preferred technical solution, the green electricity real-time carbon offset coefficient It represents the amount of fossil fuel carbon emissions that can be replaced by a unit of green electricity generation over its entire life cycle, and is set based on real-time environmental benefits or a fixed value based on life cycle assessment.
[0011] As a preferred technical solution, the carbon intensity coefficient of the power grid It reflects the indirect carbon emission responsibility borne by purchasing a unit of electricity from the public grid. The value depends on the real-time energy structure of the grid and is obtained through official data released by the regional grid or AI-based predictive models.
[0012] As a preferred technical solution, the improved multi-agent deep reinforcement learning algorithm adopts a centralized training-decentralized execution framework for dynamic game playing, with the specific settings as follows: The Actor network takes local observations as input and outputs deterministic actions, relying only on local information during execution; the Critic network takes the global state and the actions of all agents as input and outputs... Q The value is used to evaluate the quality of the Actor network's output action, while the Critic network only uses global information during training; All homogeneous follower agents share the parameters of the Actor and Critic networks; during training, all experience tuples are stored in a common experience replay buffer for all agents to sample and learn. The environment encompasses the entire park's energy system, including physical equipment and its operational constraints. The state space contains public information and environmental parameters that all agents rely on to make decisions, specifically including historical electricity price and gas price sequences, historical grid carbon intensity sequences, historical total load sequences, photovoltaic and wind power output prediction sequences, the current time period, and date type; The action space is where each agent outputs its decision action based on its current state. The leader agent's action is the set price signal; the follower agent's action is the change in adjustable load power in the next time period and the planned charging and discharging power of energy storage. Reward function: The reward for the leader agent is its utility function value for the day, and the reward for the follower agent is the negative value of the total cost for the day.
[0013] As a preferred technical solution, the improved multi-objective particle swarm optimization algorithm is used as an embedded optimizer, employing Pareto sorting and non-dominated solution selection mechanisms: Pareto sorting is used to maintain an external archive set and save non-dominated solutions found during the iteration process; the crowding degree of each solution is calculated, and solutions with high crowding degree are retained first; from the converged Pareto front, TOPSIS is used to select the optimal compromise solution that is closest to the ideal solution as the final action of the agent.
[0014] As a preferred technical solution, the improved multi-objective particle swarm optimization algorithm employs adaptive inertia weights: In the formula: This represents the current iteration number; This represents the maximum number of iterations. , These are the initial and final values of the inertia weight; This is the diversity compensation term calculated based on the population distribution variance.
[0015] As a preferred technical solution, the solution process of the hybrid intelligent solution framework that integrates the improved multi-agent deep reinforcement learning algorithm and the improved multi-objective particle swarm optimization algorithm is as follows: Initialize the Actor network, Critic network, and experience replay buffer for all agents; The following process is executed repeatedly within each training round: Initialize the environment to obtain the initial state; Traverse all intelligent agents: Leader agent: Selects a release pricing strategy based on the current strategy; Follower agent: Receives the price strategy, calls the improved multi-objective particle swarm optimization solver, and solves the internal optimization model with the goal of minimizing the sum of energy economic cost and carbon cost calculated based on dynamic carbon factor to obtain the optimal response action; All agents perform joint actions, the environment transitions to a new state, and rewards are calculated; Store the experience tuple into the experience replay buffer; Experience data is randomly sampled from the experience replay buffer; Training is performed by traversing all agents: The Critic network is trained in a concentrated manner to update the Critic network parameters while minimizing the TD error loss; The Actor network parameters are updated using a gradient ascent strategy. Output the trained policy network to obtain an approximate Stackelberg equilibrium policy.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1) This invention constructs a park energy-carbon synergy optimization model based on dynamic carbon footprint embedding and multi-agent Stackelberg game. It creatively constructs a three-layer Stackelberg game framework involving energy suppliers, multiple users, and the power grid by deeply embedding time-varying dynamic carbon footprint factors into the utility functions of multi-agent decision-making within the park. The real-time carbon offset coefficient of green electricity and the dynamic carbon intensity coefficient on the grid side are accurately quantified and introduced as core parameters into the multi-objective optimization model for leaders and followers, enabling them to simultaneously weigh economic and carbon costs in their decisions. This method overcomes the limitations of traditional models using static carbon factors, achieving accurate and real-time traceability of energy carbon emission responsibility. By defining the strategy space of each agent and a utility function containing carbon costs, and aiming to solve the Stackelberg equilibrium, it formally describes the competitive interaction behavior of various agents within the park under carbon constraints within a unified framework. This provides a complete mathematical model foundation for analyzing their energy-carbon synergy strategies and solves the key problem that existing methods cannot coordinate the interests of multiple agents with dynamic carbon objectives.
[0017] 2) This invention proposes a hybrid intelligent solution algorithm that integrates multi-agent attention reinforcement learning (MADRL) and improved multi-objective particle swarm optimization (MPSO). A novel hybrid intelligent algorithm (MADRL-IPSO) is designed to efficiently solve the aforementioned complex game equilibrium. An improved multi-agent deep deterministic policy gradient algorithm is used as the backbone framework to simulate continuous policy learning among multiple agents. Its innovation lies in introducing an attention mechanism into the Critic network, enabling each agent to adaptively focus on the most relevant information from other agents, significantly improving the accuracy of policy evaluation in a multi-agent environment. Simultaneously, addressing the challenge of complex constraints in DRL agent action decisions, the invention incorporates the improved MPSO algorithm as an embedded optimizer. This IPSO possesses adaptive inertial weights, Pareto ranking, and non-dominated solution selection mechanisms, enabling efficient solution of high-dimensional, constrained multi-objective optimization problems faced by follower users. This hierarchical fusion architecture, where MADRL handles the game and IPSO handles the optimization, effectively solves the problems of traditional methods struggling to solve high-dimensional nonlinear game models and easily getting trapped in local optima, ensuring both computational efficiency and accuracy of the equilibrium solution.
[0018] 3) This invention establishes a digital twin-driven two-stage closed-loop optimization control mechanism for the energy and carbon systems of industrial parks, encompassing day-ahead and real-time phases. It integrates models and algorithms to form a closed-loop decision support system driven by digital twin technology. A virtual mapping of the park's physical system is constructed using the digital twin, integrating real-time source, load, and carbon intensity prediction data, as well as historical operational data, to provide a dynamically updated training and simulation environment for the MADRL-IPSO hybrid algorithm. Based on this, a two-stage rolling optimization mechanism is designed: in the day-ahead phase, a pre-scheduling scheme for optimal price and energy consumption plan is solved with a 24-hour cycle; in the real-time phase, the plan is rolled over based on ultra-short-term forecast data. This mechanism, through a closed loop of real-time perception, virtual simulation, intelligent decision-making, physical execution, and feedback optimization, enables the system to continuously adapt to fluctuations in renewable energy output and grid carbon intensity, achieving dynamic adjustment and self-optimization of decision-making schemes. This effectively improves the economic efficiency and low-carbon operation of the park's energy system, overcoming the inherent limitations of static optimization schemes in handling uncertainties. Attached Figure Description
[0019] Figure 1 This is a flowchart of a multi-entity energy and carbon collaborative decision-making optimization method for low-carbon industrial parks according to the present invention.
[0020] Figure 2 This is a decision sequence diagram for the present invention.
[0021] Figure 3 This is a diagram of the centralized training-distributed execution (CTDE) framework of the present invention.
[0022] Figure 4This is a flowchart of the MADRL-IPSO hybrid solution algorithm of the present invention.
[0023] Figure 5 This is the leader benefit optimization process in a specific embodiment of the present invention.
[0024] Figure 6 This illustrates the trend of synergistic optimization of revenue and carbon cost in a specific embodiment of the present invention.
[0025] Figure 7 This is the optimized energy price signal in a specific embodiment of the present invention.
[0026] Figure 8 This is a comparison of user demand response effects in specific embodiments of the present invention.
[0027] Figure 9 This is a specific embodiment of the system carbon emission optimization process.
[0028] Figure 10 The above describes the multi-objective optimization performance indicators in a specific embodiment of the present invention.
[0029] Figure 11 This is a specific embodiment of the invention that optimizes the benefits and carbon costs in a coordinated manner. Detailed Implementation
[0030] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0031] Example 1 This invention proposes a data-driven intelligent decision-making and dynamic game optimization method for multi-agent energy carbon coordination in low-carbon / zero-carbon industrial parks. By constructing a Stackelberg game model integrating dynamic carbon footprint factors, the time-varying grid carbon intensity and green electricity carbon offsetting benefits are introduced into the agent decision-making process, achieving accurate source tracing and cost internalization of carbon responsibility. To address the model solving challenge, a hybrid intelligent solution strategy integrating multi-agent attention reinforcement learning and an improved multi-objective particle swarm optimization algorithm is designed, effectively improving the computational efficiency and accuracy of game equilibrium. Simulation results show that the proposed method can significantly reduce system carbon emissions while ensuring economic benefits, promoting source-load coordination and efficient renewable energy consumption. The overall research methodology and structure are as follows: Figure 1 As shown.
[0032] 1.1 Multi-Agent Stackelberg Game Framework: The multi-agent dynamic Stackelberg game modeling in Part 1 forms the starting point and theoretical foundation of this research. Traditional industrial park energy optimization often assumes a centralized control center, neglecting the independent interests and strategic interactions of multiple decision-making agents (such as energy suppliers and diverse users). To address this deficiency, the Stackelberg game model from game theory is introduced, abstracting the industrial park energy system into a three-tiered decision-making structure consisting of leaders (energy suppliers), followers (diverse users), and external environment makers (the power grid). The core of this model lies in formally characterizing the master-follower relationships and strategic dependencies among the agents: the leader first sets energy price signals, followers optimize their own energy consumption plans based on these signals, and the power grid provides time-varying electricity prices and carbon intensity signals as external boundary conditions. By clearly defining the strategy spaces of each agent (such as prices and load scheduling plans) and a dual-objective utility function centered on economic and carbon costs, this model provides a rigorous mathematical description for analyzing the competitive and cooperative behavior of multiple agents under carbon constraints, laying a game-theoretic foundation for subsequent optimization modeling. The key innovation in this part is that the operation of the park has shifted from the traditional single-entity optimization or completely centralized control paradigm to a more realistic multi-entity non-cooperative game paradigm, which can more realistically reflect the tension between individual rationality and the overall goals of the system.
[0033] 1.2 Multi-objective Optimization Model Integrating Dynamic Carbon Factors: The multi-objective optimization model integrating dynamic carbon footprint factors in Part Two is a core link bridging game theory and optimization practice, and represents a significant breakthrough in carbon management methodology. Addressing the carbon accounting distortion problem caused by the prevalent use of static average carbon emission factors in existing research, this part innovatively proposes the concept and modeling method of dynamic carbon footprint factors. Specifically, the model introduces two key time-varying parameters: the real-time carbon offset coefficient of green electricity and the grid-side dynamic carbon intensity coefficient. The former quantifies the real-time carbon emission reduction benefits brought by a unit of local renewable energy generation, while the latter accurately reflects the time-varying indirect carbon emission responsibility undertaken from purchasing electricity from the grid.
[0034] Building upon this foundation, these two dynamic factors are deeply embedded into the Stackelberg game framework described in Part One, constructing multi-objective optimization models at both the leader and follower levels. The leader's objective function simultaneously considers its operational revenue, costs, and the carbon costs generated by its operations; the follower's objective is to minimize the sum of its energy economic costs and the carbon costs calculated based on dynamic carbon factors. Constraints encompass system power balance, physical limitations of various equipment (such as gas turbines, energy storage, and adjustable loads), and user energy demand. This model represents a fundamental shift from static carbon accounting to dynamic carbon traceability, enabling carbon costs to be internalized in real-time and accurately into each agent's decision-making process, just like economic costs, providing a computable mathematical model for achieving true energy-carbon flow synergistic optimization.
[0035] 1.3 Solution Strategy of Fusion Intelligent Algorithms: The solution method of fusion intelligent algorithms in Part Three is a computational engine designed to address the high complexity (high dimensionality, nonlinearity, strong constraints, NP-hardness) of the models constructed in the first two parts. It is the key to propelling this framework from theory to application. Facing the bottleneck of traditional optimization algorithms in solving such problems, an innovative hybrid intelligent solution strategy (MADRL-IPSO) is proposed. The core idea of this strategy is hierarchical fusion and collaborative division of labor.
[0036] At the game strategy learning level, an improved multi-agent deep reinforcement learning (MADRL) algorithm is adopted, specifically the MADDPG algorithm based on the centralized training-distributed execution (CTDE) framework. The improvement lies in introducing an attention mechanism into the Critic network, enabling each agent (representing a decision-making entity) to more accurately assess its own policy value under the influence of other agents, thereby learning to approximate the Stackelberg equilibrium more efficiently.
[0037] At the individual optimization level, to address the issue that the action outputs of DRL agents cannot directly satisfy complex constraints, an improved multi-objective particle swarm optimization algorithm (IPSO) is embedded in each follower agent. This IPSO features adaptive inertial weights, Pareto ranking, and a TOPSIS decision mechanism, specifically designed for efficiently solving high-dimensional, constrained multi-objective optimization problems (i.e., their optimal response) for users given prices. This architecture, where MADRL handles game-theoretic interactions and IPSO handles internal optimization, cleverly decomposes the complex hybrid problem. It leverages both the advantages of DRL in policy space exploration and the strengths of metaheuristic algorithms in complex constraint optimization, thereby significantly improving computational efficiency while maintaining solution accuracy.
[0038] In summary, the overall framework presents a clear progressive logic and close module connections: Stackelberg game modeling defines the basic structure and rules of the problem (Who-How); the multi-objective optimization model integrating dynamic carbon factors endows the problem with precise quantitative connotations and objectives (What); and the solution method integrating intelligent algorithms provides an effective tool for overcoming computational challenges (How to solve). These three parts are interconnected, forming a complete research chain from problem characterization and model construction to algorithm solution. Finally, the effectiveness of the framework is verified through case studies, demonstrating its superior performance in coordinating economic and environmental goals, guiding user demand response, and reducing system carbon emissions. This provides systematic theoretical and methodological support and feasible technical implementation paths for the intelligent operation and management of low-carbon / zero-carbon parks.
[0039] 2. Multi-agent dynamic Stackelberg game modeling.
[0040] This section aims to formalize the core stakeholders and their interactions within low-carbon industrial parks into a dynamic Stackelberg game model. This model defines a three-tiered decision-making framework consisting of a leader (energy supplier), multiple followers (multiple users), and an external constraint maker (power grid). By clarifying the decision-making order, strategy space, and objective function of each staker to maximize their own utility, a game-theoretic foundation is laid for subsequent multi-objective optimization of integrated carbon costs.
[0041] 2.1 Definition and Roles of Game Players: The operation of the park's energy system involves multiple stakeholders with different objectives and decision-making powers. The main participants are defined as follows: Leader: Energy suppliers (such as integrated energy system operators in industrial parks). As pioneers and rule-makers in the game, their core strategy is to maximize their own profits and reduce carbon costs while meeting user needs by setting internal energy price signals and dispatching distributed energy resources.
[0042] Followers: A diverse set of users, including industrial users, commercial buildings, residential areas, etc. As followers in a game, after observing the price signals released by the leader, users non-cooperatively formulate their optimal energy use plans with the goal of minimizing their total energy costs (economic costs + carbon costs).
[0043] External participant and constraint setter: the power grid. It is modeled as a provider of the external environment and rule boundaries. It provides the industrial park with the main grid purchase and sale price and grid-connected power constraints. Its key role is to provide a time-varying grid carbon intensity factor that reflects the overall power generation structure of the entire grid; this factor is a crucial dynamic external input to the model.
[0044] 2.2 Game Order and Strategy Space: This game is a sequential game with complete information, a single leader, and multiple followers. Its decision sequence is as follows: Figure 2 As shown.
[0045] Game order: Phase 1 (Leadership Decision): At the beginning of each scheduling cycle (or day-ahead phase), energy suppliers formulate and publish their pricing strategies based on their forecasts of user response, renewable energy output, and grid carbon intensity.
[0046] Phase Two (Follower Decisions): Each User i After observing the price strategy Π, independently solve its cost minimization problem to determine its optimal energy use strategy. X i .
[0047] Equilibrium Achievement: The leader can anticipate the users' decision-making patterns; therefore, the pricing strategy it formulates in the first stage is actually the optimal decision under the predicted optimal user response function. When both sides reach a consensus on their strategies, the system reaches Stackelberg equilibrium.
[0048] Strategy Space: The leader's strategy space (SL) mainly includes the energy prices it sets.
[0049] (1) In the formula: It is the set of all time periods within a scheduling cycle; for t The price at which the time-of-use leader sells electricity to users (RMB / kWh); for t The price at which the time-of-use leader purchases electricity from users (RMB / kWh); for t Natural gas price per hour (yuan / m³) 3 ).
[0050] No. i Individual user strategy space This mainly includes the scheduling plan for its adjustable load.
[0051] (2) In the formula: For users i exist t Adjustable load power (kW) for different time periods; For users i Energy storage t The charging / discharging power (kW) during the time period, with discharge being positive; For users i Energy storaget State of charge during a given time period.
[0052] Utility function (payment function): The utility function of each entity is the negative of its cost or benefit, and all take into account both economic and low-carbon objectives.
[0053] Followers (users) i utility function : Its goal is to minimize total cost. .
[0054] (3) In the formula: For users i The economic costs mainly include the cost of purchasing electricity / gas from the leader, equipment operation and maintenance costs, etc. For users i The carbon emission cost is calculated based on its electricity / gas consumption and the corresponding dynamic carbon footprint factor. The carbon price coefficient (yuan / kgCO2) monetizes the cost of carbon emissions.
[0055] The utility function of the leader (energy supplier) Its goal is to maximize its own utility, that is, the difference between benefits and costs.
[0056] (4) In the formula: For revenue, it mainly refers to income from selling electricity / gas to users. Operating costs include the cost of purchasing electricity from the grid, the cost of purchasing natural gas, the operation and maintenance costs of distributed renewable energy generation, and the depreciation costs of other equipment. Carbon emission costs refer to the costs associated with the carbon emissions generated by the leader's own operations (such as carbon emissions from gas turbine power generation and purchasing electricity from the grid).
[0057] The objective of this Stackelberg game is to find its equilibrium solution (Stackelberg Equilibrium, SE). In equilibrium, the following conditions are satisfied: Strategy of all followers (users) Leadership strategy The optimal response is: (5) Anticipating the user's optimal response function Afterwards, the leader's strategy This is the optimal strategy it can adopt, namely: (6) At the equilibrium point No single entity can gain better self-utility by unilaterally deviating from its current strategy.
[0058] 3. A multi-objective optimization model integrating dynamic carbon footprint factors.
[0059] This section is the core of the entire game theory framework, concretizing the utility functions of the game actors defined in Part 1 into a computable mathematical optimization model. The key innovation lies in the introduction of a dynamic carbon footprint factor, extending the traditional single-economic-objective optimization problem into a dual-objective economic-environment optimization problem. The modeling method for the carbon footprint factor will be explained in detail below, and complete mathematical descriptions of the optimization problems at both the leader and follower levels will be provided, including the objective functions and all key constraints.
[0060] 3.1 Dynamic Carbon Footprint Factor Modeling: Accurate carbon accounting and source tracing are the cornerstones for achieving coordinated energy and carbon optimization. This model abandons the traditional constant carbon emission factor and introduces the following time-varying parameters to more accurately characterize the green attributes of energy.
[0061] Real-time carbon offset coefficient of green electricity This coefficient represents the amount of fossil fuel carbon emissions that can be replaced by a unit of green electricity generation (such as photovoltaic or wind power) over its entire life cycle. A positive value can be considered a negative carbon emission. This coefficient can be set based on real-time environmental benefits or a fixed value based on life cycle assessment (LCA).
[0062] (7) In the formula: for t Total carbon offset (gCO2) of green electricity generated in the park during the specified time period; for t Total power generation (kW) of distributed renewable energy within the park during the specified time period; for t Real-time carbon offset coefficient (gCO2 / kWh) of green electricity during a given period.
[0063] Dynamic carbon intensity coefficient on the grid side This coefficient reflects the indirect carbon emission responsibility borne for each unit of electricity purchased from the public grid. Its value depends on the real-time energy structure of the grid (the proportion of thermal power, hydropower, nuclear power, etc.), and is a parameter that changes dramatically over time. It can be obtained through official data released by the regional grid or AI-based predictive models.
[0064] (8) In the formula: for t Carbon intensity coefficient of the power grid during a given time period (gCO2 / kWh).
[0065] Load-side carbon emission calculation principle: Adhering to the principle of shared responsibility. The carbon emission responsibility for each kilowatt-hour of electricity consumed by a user or leader is allocated based on the source of the electricity: The carbon intensity of electricity consumed from the grid is: .
[0066] The carbon intensity of consuming locally distributed green electricity is (i.e., generating carbon offsetting benefits).
[0067] The carbon intensity of natural gas consumption is constant. .
[0068] 3.2 Leader (Energy Supplier) Optimization Model: As the energy hub of the park, the leader's decision-making needs to simultaneously meet energy balance and equipment operation constraints, and seek the optimal trade-off between economic benefits and carbon costs.
[0069] 3.2.1 Objective Function: Leaders seek to minimize total costs (the sum of operating costs and carbon costs minus revenues), which is equivalent to maximizing their utility function. .
[0070] (9) 1) Operating costs ( ): (10) In the formula: For time step, , They are respectively t The power (kW) purchased from and sold to the grid during a given period, and typically ; for t Natural gas purchase capacity (kWth) during the time period; , , These are the corresponding energy prices; for t Operation and maintenance costs (in yuan) of all equipment (photovoltaics, wind turbines, gas turbines, energy storage, etc.) during the period.
[0071] 2) Carbon emission costs ( ): (11) In the formula: Carbon valence coefficient (yuan / gCO2); The carbon intensity coefficient of natural gas (gCO2 / kWhth).
[0072] 3) Profits ( ): (12) In the formula: Total number of users; For users i exist t Electricity (kW) purchased from the leader during the specified time period; For users i exist t The amount of gas power (kWth) purchased from the leader during a given period.
[0073] 3.2.2 Constraints Electric power balance constraints: (13) In the formula: , , These represent the power output (kW) of photovoltaic, wind turbine, and gas turbine respectively during time period t; The charging / discharging power (kW) of the energy storage belonging to the leader; The leader's own load (kW).
[0074] Equipment operating constraints (taking gas turbines and energy storage as examples): Gas turbine (MT) gradeing constraints: (14) In the formula: , These are the uphill and downhill ramp rate limits (kW / h) for the gas turbine.
[0075] Energy Storage System (ESS) Operational Constraints: (15) (16) In the formula: The state of charge of the leader's energy storage during time period t; , For charging and discharging efficiency; Rated energy storage capacity (kWh); , These represent the lower and upper limits allowed for the state of charge.
[0076] Power exchange constraints with the power grid: (17) In the formula: , The contractual upper limit (kW) for the power exchanged with the grid.
[0077] 3.3 Followers (users) i Optimization model: Each user acts as an independent decision-making entity, and after receiving price signals from the leader, adjusts their energy consumption behavior to minimize total cost.
[0078] 3.3.1 Objective Function: User i The goal is to minimize total cost.
[0079] (18) 1) Economic cost ( ): (19) In the formula: For users i The operation and maintenance costs of the equipment.
[0080] 2) Carbon emission costs ( ): (20) Note: Users typically do not directly own renewable energy sources, so their carbon emissions mainly come from purchased electricity and gas. If users have rooftop solar power, a corresponding negative carbon term should be added to the model.
[0081] 3.3.2 Constraints: Energy demand constraint: The basic load must be met.
[0082] (twenty one) In the formula: , users respectively i exist t Fixed load and adjustable load power (kW) for different time periods.
[0083] Adjustable load constraints (taking temperature-controlled load as an example): (twenty two) (twenty three) In the formula: The set of allowable operating segments for adjustable loads; This represents the total energy (kWh) required for the load to complete within one cycle. The indoor temperature should be maintained within a comfortable range.
[0084] User-side energy storage operation constraints: similar to those of leader energy storage constraints.
[0085] Electric vehicle (EV) charging constraints (if any): (twenty four) (25) In the formula: Plan off-grid time for electric vehicles; This represents the desired state of charge when the device is off-grid.
[0086] 4. Solution method integrating intelligent algorithms.
[0087] To address the computational complexity challenge of the previously constructed model, the dynamic Stackelberg game model is a high-dimensional, nonlinear, and complex system with multiple constraints, making it difficult for traditional optimization methods to directly solve for its equilibrium. Therefore, a hybrid intelligent solution framework integrating improved multi-agent deep reinforcement learning (MADRL) and improved particle swarm optimization (IPSO) is proposed. MADRL is used to simulate continuous policy interactions and learning among multiple agents to approximate the game equilibrium; while IPSO serves as an embedded optimizer to efficiently solve the high-dimensional, constrained decision problem for each agent under a given state. This hybrid strategy combines the advantages of global policy search and local exact optimization.
[0088] 4.1 Improved Multi-Agent Deep Reinforcement Learning (MADRL) for Dynamic Game Theory. Deep reinforcement learning (DRL) possesses a powerful ability to learn through trial and error in complex and unknown environments, making it highly suitable for simulating multi-agent decision-making processes. This invention employs, as follows: Figure 3 The Centralized Training-Distributed Execution (CTDE) framework is shown.
[0089] 1) Agent and Environment Setup: Agent: Each decision-making entity in the game (1 leader + N user followers) is modeled as an independent agent.
[0090] Environment: The entire park's energy system, including the power grid, renewable energy, loads, energy storage, and other physical equipment and their operational constraints.
[0091] 2) State Space (S): States It includes public information and environmental parameters on which all agents rely to make decisions during time period t.
[0092] (26) In the formula: Historical electricity and gas price series; Historical power grid carbon intensity sequence; This is a historical total load sequence; For the output prediction sequence of photovoltaic and wind power; The current time period; For date type (weekday / weekend).
[0093] 3) Action Space (A): Each agent outputs its decision action based on its current state.
[0094] Leader agent actions Its actions are based on established price signals.
[0095] (27) In the formula: These are the electricity sales price, electricity purchase price, and gas price set by the leaders for the t+1 period.
[0096] Follower agent i-action Its action is the adjustment of the electricity consumption plan.
[0097] (28) In the formula: This represents the change in adjustable load power during the t+1 time period; The planned charge and discharge power for energy storage during period t+1.
[0098] 4) Reward Function (R): The reward function is directly linked to the utility function of each agent, guiding its learning direction.
[0099] Leader rewards : (29) That is, its utility function value on that day, which encourages it to formulate a price strategy that maximizes net income in the long run.
[0100] Follower i reward : (30) This means that the negative value of its total daily cost encourages it to minimize costs by adjusting its energy consumption plan.
[0101] 5) Algorithm and Network Structure: An improvement is made based on the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm. MADDPG is an extension of DDPG in multi-agent scenarios, and its core idea is CTDE.
[0102] Network structure: Each agent has its own Actor network ( ) and Critic network ( ).
[0103] Actor Network ( ): Input local observations Output deterministic actions It relies only on local information during execution.
[0104] Critic Network ( ): Input global state and the actions of all intelligent agents It outputs Q-values to evaluate the quality of the Actor's actions, using global information only during training.
[0105] Improvement 1: Employ the Critic network with an attention mechanism.
[0106] Traditional MADDPG's Critic network simply concatenates the actions of other agents, making it difficult to effectively handle situations with a large number of agents and complex relationships. Introducing an attention layer into the Critic network allows it to automatically focus on the most relevant information from other agents in the current state, thus more accurately evaluating the Q-value.
[0107] (31) In the formula: and For encoding networks; The attention weight of agent j to agent i is calculated by the query-key mechanism.
[0108] Improvement point 2: Parameter sharing and experience playback All homogeneous follower agents share the parameters of the Actor and Critic networks. This significantly reduces model complexity, speeds up training, and ensures policy consistency. All empirical tuples It is stored in a public experience replay buffer for all agents to sample and learn from.
[0109] 4.2 An improved Particle Swarm Optimization (IPSO) algorithm is used for internal optimization. At each time step of the DRL, when the follower agent needs to make a decision based on the current electricity price, the decision itself is a complex constrained optimization problem. Directly using the neural network to output actions is difficult to guarantee that all constraints are satisfied. Therefore, IPSO is used as the agent's internal solver, which is called by the Actor network to solve for the optimal response in the current state.
[0110] The improvements of IPSO are as follows: 1) Adaptive inertia weight: Inertia weight Balancing global exploration with local exploitation capabilities. Employing a linear decrease strategy, and dynamically increasing diversity when population diversity is too low. To escape local optima.
[0111] (32) In the formula: This represents the current iteration number; This represents the maximum number of iterations. , These are the initial and final values of the inertia weight; This is the diversity compensation term calculated based on the population distribution variance.
[0112] 2) Multi-objective processing and decision-making: The user optimization problem is essentially a dual-objective problem of economic cost and carbon cost. IPSO maintains an external archive through Pareto Ranking to store the non-dominated solutions (Pareto fronts) found during the iteration process.
[0113] Crowding Distance Calculation: To ensure the distribution of the frontier, the crowding degree of each solution is calculated, and solutions with high crowding degree are retained first.
[0114] Final Decision: From the converged Pareto front, the TOPSIS (Technique for Order Preference by Similarity to Ideal Solution) method is used to select the optimal compromise solution that is closest to the ideal solution, which is then used as the agent's final action. .
[0115] 3) Constraint handling mechanism: A constraint handling mechanism based on the feasible solution priority criterion is adopted. When comparing two particles: if both are feasible, the objective function values are compared; if one is feasible and the other is not, the feasible solution is selected; if neither is feasible, the solution with the smaller constraint violation is selected.
[0116] 4.3 Hybrid Solution Process: (e.g.) Figure 4 As shown, the entire solution process of the algorithm is a nested loop structure. Initialization: Initialize the Actor network for all agents. Critic Network and experience replay buffer .
[0117] The following process is executed repeatedly within each training round: Initialize the environment and obtain the initial state. .
[0118] Traverse all intelligent agents: Leader agent: Based on the current policy Select Action (Published price); Follower agent: Receives the price, invokes the IPSO solver, and minimizes... To obtain the optimal response action, we solve its internal optimization model. .
[0119] All agents perform a joint action. The environment shifts to a new state And calculate the reward .
[0120] experience tuples Store in the experience replay buffer Inside.
[0121] Randomly from the experience replay buffer A small batch of empirical data was sampled.
[0122] Training is performed by traversing all agents: Focus on training the Critic network to minimize the TD error: Update the Critic network parameters to minimize the loss: ; Update the Actor network: Update target network parameters: .
[0123] Output: The trained policy network This approximates the Stackelberg equilibrium strategy.
[0124] Example 2 This embodiment provides a specific implementation example, simulating a medium-sized low-carbon industrial park, comprising one energy supplier (leader) and five typical user groups (followers): heavy industry, light industry, commercial, office, and residential users. The park is equipped with a 200kW photovoltaic power generation system and interacts with the main grid. The optimization period is 24 hours, with a time resolution of 1 hour. The key parameters for system operation are set as shown in the table below: Table 1 System Basic Parameter Configuration The carbon intensity coefficient of the power grid adopts a time-varying model to simulate the carbon emission characteristics of the actual power grid: carbon intensity during peak carbon emission periods (18:00-21:00) is 0.8-1.0 kgCO2 / kWh; carbon intensity during low carbon periods (night and noon) is 0.1-0.4 kgCO2 / kWh; and the maximum offset coefficient of green electricity carbon offset is 0.7 kgCO2 / kWh (noon).
[0125] The convergence behavior of the system was observed through 100 rounds of game optimization iteration. Figure 5 The study demonstrates the trend of leader payoffs with optimization rounds, showing that payoffs increase rapidly in the first 40 rounds before gradually converging to a stable state, indicating that the proposed MADRL-IPSO hybrid algorithm has good convergence performance.
[0126] Figure 6 This demonstrates the simultaneous optimization process of benefits and carbon costs, both of which show an improving trend, proving that the proposed framework can effectively coordinate economic and environmental goals.
[0127] Optimized energy price signals such as Figure 7 As shown, the system exhibits significant carbon sensitivity. Compared to traditional time-of-use pricing, the price curve generated by the optimized model is highly coupled with the grid carbon intensity curve: during the evening period (18:00-21:00) when carbon intensity is high, the price is further increased, reflecting not only higher carbon costs but also suppressing high-carbon electricity demand during this period through price leverage; while during the midday period (12:00-14:00) when photovoltaic output is abundant and grid carbon intensity is low, the price is relatively lower than the traditional time-of-use pricing. This linkage between carbon price and electricity price accurately transmits dynamic carbon costs to users, providing key economic incentives for guiding low-carbon energy consumption behavior. This demonstrates that the constructed game theory framework can endogenously generate intelligent price signals that promote the low-carbon operation of the system.
[0128] like Figure 8As shown, guided by the optimized carbon-sensitive price signal, the total user load curve underwent a fundamental reshaping, exhibiting three typical characteristics: peak-valley shifting and matching, with peak load significantly shifting from the traditional high-carbon evening period (18:00-21:00) to the low-carbon midday period (12:00-14:00), which effectively matches the peak output of local photovoltaic power generation, greatly improving the local consumption level of clean energy. Peak shaving and valley filling, by incentivizing users to adjust adjustable loads (such as electric vehicles and flexible loads), reduced peak load by approximately 15% while increasing valley load, decreasing the peak-valley load difference rate from 45% to 32%, significantly enhancing system stability. Carbon emission optimization, effectively suppressing electricity consumption during periods of high grid carbon intensity while stimulating electricity demand during low-carbon periods, optimizing the carbon flow distribution of the system from the source. This result verifies that the proposed model can effectively coordinate the decentralized decisions of diverse users, guiding their self-interested behavior in a direction consistent with the system's low-carbon and stable operation goals.
[0129] The optimization process for total carbon emissions of the system is as follows: Figure 9 As shown, after game-theoretic optimization, total carbon emissions decreased from approximately 1200 kg CO2 initially to approximately 950 kg CO2, a reduction of 20.8%. This significant achievement stems from two synergistic effects: first, supply-side optimization, where dynamic carbon costs incentivize energy suppliers to prioritize the dispatch of green electricity such as photovoltaic power and optimize the operating strategies of equipment like gas turbines; second, load-side response, where users proactively shift peak loads and fill valleys based on price signals incorporating carbon costs. This demonstrates that by embedding dynamic carbon footprint factors into a multi-agent decision-making model, this invention can simultaneously stimulate the emission reduction potential on both the source and load sides.
[0130] To comprehensively evaluate the economic and environmental benefits, Figure 10 The normalized performance metrics for the two objectives are presented. It can be seen that both metrics show a steady upward trend and reach a relatively balanced state in the later stages of optimization, demonstrating the effectiveness of the proposed method in multi-objective optimization.
[0131] Table 2 presents the quantitative results, comprehensively demonstrating the overall performance improvement of this invention. In terms of economics, the leader's total revenue increased by 20.0%, proving that this game theory framework can protect the economic interests of the core operating entity while promoting decarbonization. Regarding decarbonization, the system's total carbon emissions decreased by 20.8%, and the green electricity consumption rate increased by 14 percentage points, showing significant effectiveness. In terms of system efficiency, the reduction in peak-valley load difference and the decrease in the proportion of carbon costs together indicate that both the economic and environmental costs of system operation have been optimized. This demonstrates that the framework successfully achieves an effective trade-off and synergistic improvement between economics and decarbonization.
[0132] Table 2 Comparison of key indicators before and after optimization Carbon price is a key policy lever for balancing economic and environmental goals. Sensitivity analysis shows that at low carbon prices (0.05 yuan / kgCO2), carbon costs account for a relatively small proportion of total costs, and model decisions tend to prioritize economic efficiency, resulting in limited emission reduction effects. At high carbon prices (0.20 yuan / kgCO2), carbon costs become the dominant factor, driving the system to adopt more aggressive emission reduction strategies, further reducing total carbon emissions, but leaders' gains are sacrificed due to increased operating costs. The equilibrium carbon price (0.10 yuan / kgCO2) achieves the best balance between economic efficiency and low carbon emissions. This provides a quantitative reference for carbon market policymaking, indicating the existence of an optimal carbon price range that maximizes total social welfare.
[0133] When the photovoltaic capacity increased to 300kW, the optimization results showed a stronger low-carbon orientation: the electricity price signal during the midday period was further reduced to incentivize the consumption of surplus green electricity, and total carbon emissions dropped to approximately 800 kgCO2. This demonstrates that the framework has good scalability and can adapt to the trend of the park's renewable energy ratio continuing to increase in the future.
[0134] Compared with static carbon accounting methods, the proposed dynamic carbon footprint method has advantages over traditional static carbon factor methods in the following aspects: Scientific decision-making. Dynamic factors accurately reflect the true carbon responsibility of electricity consumption behavior at different times, avoiding the misleading decisions that may arise from static methods, such as saving electricity during low-carbon periods and consuming electricity during high-carbon periods. The optimization results are shown in Table 2 and the comparative analysis. Under the same conditions, the emission reduction effect of dynamic carbon accounting is improved by approximately 8 percentage points, and the economic cost is reduced by approximately 12% when achieving the same emission reduction target. This highlights the necessity of dynamic carbon traceability for achieving precise carbon management and optimal cost.
[0135] Compared to centralized optimization, which assumes the existence of an omnipotent central scheduler, the distributed game framework has more advantages: model realism, respecting the reality of independent property rights and autonomous decision-making among multiple entities within the park, resulting in strategies with more practical guidance; computational scalability, the MADRL-IPSO hybrid algorithm effectively avoids the dimensionality curse problem that arises in centralized optimization as the number of users increases through distributed solution; system robustness, the decentralized decision-making structure makes it more tolerant to prediction errors or failures of individual entities.
[0136] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A data-driven, multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks, characterized in that, include: A multi-agent dynamic Stackelberg game model is constructed, and the time-varying grid-side carbon intensity and green electricity carbon offsetting benefits are introduced into the multi-agent game decision-making process to establish a multi-objective optimization model integrating dynamic carbon footprint factors: Energy suppliers, as leaders, formulate and publish pricing strategies; users, as followers, aim to minimize the sum of economic energy costs and carbon costs calculated based on dynamic carbon factors, and solve for the optimal energy use strategy; the power grid, as the external environment maker, provides dynamic grid carbon intensity factors and renewable energy forecasts as external boundary conditions. A hybrid intelligent solution framework combining an improved multi-agent deep reinforcement learning algorithm and an improved multi-objective particle swarm optimization algorithm is used to solve a multi-agent dynamic Stackelberg game model. At the game strategy learning level, the improved multi-agent deep reinforcement learning algorithm is used to approximate the Stackelberg equilibrium. At the individual optimization solution level, the improved multi-objective particle swarm optimization algorithm is embedded in the follower agent to solve the follower's optimal response under the current price strategy. In the improved multi-agent deep reinforcement learning algorithm, all leaders and followers are modeled as agents, each with its own Actor network and Critic network. An attention layer is introduced into the Critic network. In the formula: and These are coding networks; This is the global state; For intelligent agents i The action; For intelligent agents j For intelligent agents i Attention weights.
2. The data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 1, characterized in that, The multi-agent dynamic Stackelberg game model is specifically constructed as follows: The first phase involves leader decision-making: energy suppliers develop and publish pricing strategies based on their forecasts of user response, renewable energy output, and grid carbon intensity. The second stage involves follower decision-making: After observing the pricing strategy, each user aims to minimize the sum of the economic cost of energy use and the carbon cost calculated based on the dynamic carbon factor, and solves for their optimal energy use strategy. The leader's strategy space is the energy price it sets, including the price at which the leader sells electricity to followers, the price at which the leader buys electricity from followers, and the price of natural gas during all periods of the scheduling cycle; the leader's utility function aims to maximize the difference between its own revenue and costs, which include the leader's operating costs and carbon emission costs. The follower's strategy space is a scheduling plan for adjustable loads, including the follower's adjustable load power in the corresponding time period, the follower's energy storage charging / discharging power in the corresponding time period, and the follower's energy storage state of charge in the corresponding time period; the follower's utility function aims to minimize total cost, including the user's economic cost and carbon emission cost.
3. The data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 2, characterized in that, The multi-agent dynamic Stackelberg game model extends the economic objective optimization problem of the game agents into a dual economic-environmental objective optimization problem by introducing a dynamic carbon footprint factor. The leader's objective function is to minimize the difference between the sum of operating costs and carbon costs and the revenue: In the formula, For the leader's operating costs; The carbon emission costs for leaders; For the benefit of the leader; Among them, the carbon emission costs of leaders It is expressed as follows: In the formula, This refers to the carbon price coefficient. For time step; for t The amount of electricity purchased from the grid during a given time period; Let be the carbon intensity coefficient of the power grid during time period t; for t Natural gas purchase capacity during specific time periods; The carbon intensity coefficient of natural gas; for t Total power generation of distributed renewable energy within the park during the specified time period; for t Real-time carbon offset coefficient of green electricity for a given period; The objective function of the followers is to minimize the sum of operating costs and carbon costs; In the formula, The economic cost to followers; The carbon emission costs for followers; where, the carbon emission costs for followers It is expressed as follows: In the formula, This refers to the carbon price coefficient. For time step; For users i exist t The amount of electricity purchased from the leader during a given period; for t Carbon intensity coefficient of the power grid during a given time period; For users i exist t The amount of gas purchased from the leader during the specified period; This is the carbon intensity coefficient of natural gas.
4. The data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 3, characterized in that, The green electricity real-time carbon offset coefficient It represents the amount of fossil fuel carbon emissions that can be replaced by a unit of green electricity generation over its entire life cycle, and is set based on real-time environmental benefits or a fixed value based on life cycle assessment.
5. A data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 3, characterized in that, The carbon intensity coefficient of the power grid It reflects the indirect carbon emission responsibility borne by purchasing a unit of electricity from the public grid. The value depends on the real-time energy structure of the grid and is obtained through official data released by the regional grid or AI-based predictive models.
6. The data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 1, characterized in that, The improved multi-agent deep reinforcement learning algorithm employs a centralized training-decentralized execution framework for dynamic game playing, with the specific settings as follows: The Actor network takes local observations as input and outputs deterministic actions, relying only on local information during execution; the Critic network takes the global state and the actions of all agents as input and outputs... Q The value is used to evaluate the quality of the Actor network's output action, while the Critic network only uses global information during training; All homogeneous follower agents share the parameters of the Actor and Critic networks; during training, all experience tuples are stored in a common experience replay buffer for all agents to sample and learn. The environment encompasses the entire park's energy system, including physical equipment and its operational constraints. The state space contains public information and environmental parameters that all agents rely on to make decisions, specifically including historical electricity price and gas price sequences, historical grid carbon intensity sequences, historical total load sequences, photovoltaic and wind power output prediction sequences, the current time period, and date type; The action space is where each agent outputs its decision action based on its current state. The leader agent's action is the set price signal; the follower agent's action is the change in adjustable load power in the next time period and the planned charging and discharging power of energy storage. Reward function: The reward for the leader agent is its utility function value for the day, and the reward for the follower agent is the negative value of the total cost for the day.
7. The data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 1, characterized in that, The improved multi-objective particle swarm optimization algorithm is used as an embedded optimizer, employing Pareto sorting and non-dominated solution screening mechanisms: Pareto sorting is used to maintain an external archive set and to save non-dominated solutions found during the iteration process. Calculate the crowding degree of each solution and prioritize retaining solutions with high crowding degree; from the converged Pareto front, use TOPSIS to select the optimal compromise solution that is closest to the ideal solution as the agent's final action.
8. The data-driven multi-entity energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 1, characterized in that, The improved multi-objective particle swarm optimization algorithm employs adaptive inertia weights: In the formula: This represents the current iteration number; This represents the maximum number of iterations. , These are the initial and final values of the inertia weight; This is the diversity compensation term calculated based on the population distribution variance.
9. A data-driven multi-agent energy and carbon collaborative decision-making method for low-carbon industrial parks according to claim 1, characterized in that, The solution process of the hybrid intelligent solution framework that integrates the improved multi-agent deep reinforcement learning algorithm and the improved multi-objective particle swarm optimization algorithm is as follows: Initialize the Actor network, Critic network, and experience replay buffer for all agents; The following process is executed repeatedly within each training round: Initialize the environment to obtain the initial state; Traverse all intelligent agents: Leader agent: Selects a release pricing strategy based on the current strategy; Follower agent: Receives the price strategy, calls the improved multi-objective particle swarm optimization solver, and solves the internal optimization model with the goal of minimizing the sum of energy economic cost and carbon cost calculated based on dynamic carbon factor to obtain the optimal response action; All agents perform joint actions, the environment transitions to a new state, and rewards are calculated; Store the experience tuple into the experience replay buffer; Experience data is randomly sampled from the experience replay buffer; Training is performed by traversing all agents: The Critic network is trained in a concentrated manner to update the Critic network parameters while minimizing the TD error loss; The Actor network parameters are updated using a gradient ascent strategy. Output the trained policy network to obtain an approximate Stackelberg equilibrium policy.