Two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning
By constructing a two-layer dynamic game framework and an asynchronous multi-agent reinforcement learning algorithm, the problems of multi-party conflicts of interest and dynamic constraints in power demand response are resolved, the efficient consumption of distributed energy and load curve optimization are achieved, and the stability of the power grid and the efficiency of multi-agent collaboration are improved.
Patent Information
- Application Number
- CN202511067292.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Existing electricity demand response methods fail to effectively coordinate the conflicts of interest among multiple parties, making it difficult to achieve efficient consumption of distributed energy and load curve optimization. In addition, the dimension of the strategy space explodes in large-scale user groups, and traditional methods have shortcomings in dynamic constraint processing.
A two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning is adopted to construct a two-layer dynamic game framework, and an asynchronous multi-agent reinforcement learning algorithm is designed. Through the Stackelberg-Nash equilibrium and partially observable Markov game model, strategic interaction and optimization between utility companies and consumers are realized.
It improves the efficiency of multi-agent collaboration, realizes the efficient consumption of distributed energy and load curve optimization, solves the problems of lack of interest coordination mechanism and insufficient dynamic constraint processing in traditional methods, and ensures the convergence speed of the algorithm and the stability of power grid operation.
Smart Images

Figure CN120563276B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of smart grid optimization scheduling and distributed energy management, and specifically to a two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning. Background Art
[0002] With the continuous development of smart grid and distributed energy technologies, the operation mode and supply-demand relationship of the power system have undergone profound changes. The intermittent and volatile nature of distributed energy resources connected to the grid has posed challenges to the stable operation of the grid and has also changed the supply and demand structure of the electricity market. Electricity demand response, as an effective means, can guide users to rationally adjust their electricity consumption behavior, improve the flexibility and reliability of the grid, and optimize the allocation of power resources. In this context, how to achieve efficient interaction between the grid and users and the efficient consumption of distributed energy resources through scientific methods has become a major research topic in the power sector. In-depth research and effective solutions to these issues will help promote the development of a cleaner, more efficient, and more intelligent power system to meet the needs of sustainable social and economic development.
[0003] In the early days, traditional demand response (DR) approaches primarily focused on static regulation based on price signals or incentive mechanisms. For example, some approaches employed centralized optimization algorithms, such as linear programming and model predictive control, attempting to coordinate grid operation with user electricity demand to a certain extent. However, these approaches often decoupled the utility's objectives from the user's electricity costs, failing to fully consider the game-playing behavior of users as rational decision-makers. Furthermore, while traditional multi-agent approaches, such as Q-learning and DDPG, attempt to address the coordination issues among multiple agents, they encounter the dimensionality explosion of the strategy space when dealing with large user groups. Furthermore, previous studies have mostly modeled demand response as a static game, ignoring dynamic constraints such as the decay of the energy storage system's state of charge and the time-varying nature of electricity prices. Furthermore, information asymmetry, such as errors in user-side distributed energy resource output forecasts and incomplete transparency of utility company strategies, complicates the traditional assumption of a perfect information game.
[0004] Therefore, a new method is urgently needed to resolve the conflicts of interests and dynamic coordination problems among multiple parties, to achieve efficient consumption of distributed energy, optimization of load curves, and coordination of interests among multiple parties, to improve the efficiency of multi-agent collaboration, and at the same time ensure the convergence speed of the algorithm and the stability of power grid operation. Summary of the Invention
[0005] In order to achieve efficient consumption of distributed energy resources (DER), load curve optimization and coordination of interests of multiple parties, this application provides a two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning.
[0006] In a first aspect, the present application provides a two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning, comprising:
[0007] S1. Construct a two-layer non-cooperative game framework, including: construct a two-layer dynamic game model between utility companies and consumers; wherein, the upper-layer utility company strategy includes: maximizing the grid revenue by adjusting the charging and discharging power of the energy storage system, setting time-of-use electricity prices and DR incentive prices; the lower-layer consumer strategy includes: considering the external factors of weather factors and the internal factors of consumer consumption level, and adaptively adjusting the electricity load curve, the proportion of distributed energy use and the DR response based on internal and external factors to minimize the electricity cost in the non-cooperative Nash game; the equilibrium goal includes: convergence to the Stackelberg-Nash equilibrium after the interaction of the two parties' strategies; S2. Complete the partially observable Markov game modeling, transform the two-layer non-cooperative game into a partially observable Markov game, define the agent, state space, observation space, action space, reward function and value function; S3. Design an asynchronous multi-agent reinforcement learning algorithm as the game equilibrium solution engine. The solution process includes: adopting Using an asynchronously updated multi-agent proximal policy optimization algorithm, in the Stackelberg stage, the upper-level utility company prioritizes updating the policy network, and consumers update the shared policy encoder parameters based on the upper-level utility company's policy; in the Nash stage, consumers use independent policy decoders to perform Nash game optimization under a fixed utility company's policy, and the policy network adopts a shared encoder and independent decoder structure to ensure decentralized execution; and monitor the convergence speed and stability indicators of the Stackelberg-Nash equilibrium in real time, and dynamically adjust the policy update frequency and step size of the utility company and the consumer; S4, implement centralized training and priority experience replay to achieve efficiency optimization of the asynchronous multi-agent reinforcement learning algorithm, including: adopting a centralized training-distributed execution architecture, the utility company and the consumer share the policy network encoder parameters; through the priority experience replay mechanism, the historical trajectories are weighted sampled according to the action advantage value, and the high-reward policy trajectories are trained first to accelerate convergence to the equilibrium strategy.
[0008] By adopting the above scheme, considering the strategic preferences of consumers in different scenarios, a two-layer dynamic game model is constructed to enable utility companies and consumers to maximize their profits and minimize their costs respectively, and form a closed-loop strategy coupling mechanism; the game is modeled as a partially observable Markov game environment to introduce differentiated reward functions, thereby comprehensively considering the benefits of all parties; an asynchronous multi-agent proximal strategy optimization algorithm including the Stackelberg stage and the Nash stage is used to dynamically update the strategy, so that both parties can iteratively update the strategy according to different rules to achieve reasonable strategy interaction; strategy training is carried out through a centralized training decentralized execution architecture and a priority experience replay mechanism, which can improve generalization, prioritize training of high-return strategy trajectories, and accelerate convergence to the equilibrium strategy, thereby achieving efficient consumption of distributed energy in the power grid and suppressing load fluctuations.
[0009] Preferably, the construction of the two-layer non-cooperative game framework in step S1 includes the following sub-steps: S101, defining a game hierarchy; including: establishing a non-cooperative game architecture with two layers of decision-making entities, i.e., a two-layer non-cooperative game architecture, wherein the upper layer is the utility company as the leader, which uses the Stackelberg game optimization strategy to maximize the grid revenue; the lower layer is the group of electricity consumers as followers, which considers the external factors of weather factors and the internal factors of consumers' own consumption levels, and minimizes the individual electricity cost through the non-cooperative Nash game; the specific mathematical model is expressed as follows:
[0010] , ,and ;in, and They correspond to the utility companies and The strategy vector of each consumer, i.e. the decision variable containing the time series; for The corresponding optimal response strategy is To correspond The optimal response strategy; Strategic feasible domain for utility companies; Utility Company Strategies for Consumers The feasible domain under Take 1 or 2, when Taking 1 corresponds to the consumer with good consumption level considering weather factors using the utility company strategy The feasible domain under Taking 2 corresponds to the consumer with average consumption level considering weather factors using the utility company strategy The feasible domain under Characterization Consumers in other consumer strategies and utility company strategies The utility of the following; Characterization Consumers with good consumption level considering weather factors have other consumer strategies and utility company strategies The utility of the following; Characterization Consumers with average spending power considering weather factors have other consumer strategies and utility company strategies The utility of the following; The utility of the utility company under the consumer's optimal response strategy; Represent the consumer set, take When the weather factor is taken into account, it represents consumers with a good consumption level. When , it represents consumers with average consumption level considering weather factors; Indicates that except When the Stackelberg-Nash equilibrium is reached, the double constraint is satisfied: the utility company's strategy Maximize your own utility under the optimal response of all consumers, the consumer strategy The Nash equilibrium is formed as follows:
[0011] , ;
[0012] S102. Upper-level utility company strategy modeling. In the demand response framework, the utility company's optimization model formula includes:
[0013] , ; Among them, the decision variables It contains eight types of control parameters on the time series; Indicates the amount of electricity purchased from the main network. To sell electricity to users, is the abandoned distributed energy power, and Respectively represent the charging power and discharging power of the energy storage system, Represents the amount of charging using distributed energy, The amount of electricity purchased for the grid, is the expected demand response power; For electricity sales revenue, The cost of purchasing electricity, For distributed energy consumption income, flexibility costs for demand response; It refers to parameter maximization;
[0014] S103, lower-level consumer strategy modeling, the optimization model of the i-th consumer is:
[0015] , ;in, , ; Indicates the The amount of electricity purchased by a consumer from distributed energy resources, Indicates the The amount of electricity a consumer buys from a utility company, The value range is determined by a pre-built distributed energy use or grid power purchase decision model that takes into account weather and consumption levels. The pre-built distributed energy use or grid power purchase decision model that takes into account weather and consumption levels adopts a neural network algorithm, and the input is the current Considering the consumer consumption level type and weather factors under weather factors, the output is The value range of is generated by training based on historical weather factors, the consumption level type of each consumer under weather factors, and the ratio of the amount of electricity each consumer actually purchased from distributed energy resources to the amount of electricity each consumer purchased from the utility company. Indicates the Each consumer can reduce load, Indicates the consumers can transfer load, and and The corresponding value range is determined by a pre-built decision model for dynamically adjusting the electricity load curve considering weather and consumption levels; the pre-built decision model for dynamically adjusting the electricity load curve considering weather and consumption levels adopts a neural network, and the input is the current Considering the consumer consumption level type and weather factors under weather factors, the output is and The corresponding value range is generated by training based on historical weather factors, the consumption level type of each consumer under weather factors, the actual load reduction range of each consumer, and the load transfer range of each consumer; For the What a consumer spends on electricity from a utility company; For the electricity price Absorbing the benefits of distributed energy generation; For the The utility function of each consumer participating in load dispatch and transfer;
[0016] S104. Establish a two-tier game interaction mechanism; the utility company, as the leader, dominates the game through the charging and discharging power and regulation of the energy storage system, the allocation of demand response capacity, the setting of the time-of-use electricity price mechanism, and the incentive pricing of distributed energy consumption; the consumer group, as the follower, considers the external factors of weather factors and the internal factors of consumers' own consumption level, and participates in the two-tier interaction by adaptively utilizing load curve scheduling, power consumption period shifting, and the proportion of distributed energy and grid power purchase based on internal and external factors; at the same time, non-cooperative competition is formed among consumers at the same level to compete for limited distributed energy consumption quotas and bid for demand response incentive compensation; cross-tier strategy coupling includes: the utility company affects the consumer's feasible domain through decision variables, and consumers react to the utility company's demand response flexibility cost item through demand response incentive compensation, forming a closed-loop game link.
[0017] By adopting the above scheme, based on the Stackelberg-Nash dynamic game theory, a non-cooperative game architecture with two-layer decision-making entities is established. The upper layer is designed as the grid operator (UC) as the leader, and the Stackelberg game is used to optimize the benefits. The lower layer is the electricity consumer group as the follower. Considering the external factors such as weather factors and the internal factors such as consumers' own consumption levels, consumers are divided into detailed categories. The individual electricity cost is minimized through the non-cooperative Nash game. The closed-loop strategy coupling mechanism makes the strategies of both parties influence and interact with each other, promoting the convergence of the strategies of both parties to the equilibrium state after interaction, thereby realizing the efficient consumption of distributed energy and the suppression of load fluctuations.
[0018] Preferably, the step S2 of completing the partially observable Markov game modeling includes the following sub-steps: S201, defining a POMG tuple; including: based on the POMG theory, regarding the utility company and each consumer as POMG agents, and the POMG tuple is expressed as:
[0019] ;
[0020] Where, is the state space; is the set of observations of the utility company; is the consumer's observation set; It is the utility company’s action set; is the consumer’s action set; is the utility company’s reward function, is the consumer’s reward function, is the collection of transfer nuclei;
[0021] S202. Constructing a value function system; including: given strategy , the state value functions of the utility company as the leader and the consumer as the follower are defined as: , ; where the random strategy of the leading utility company is is a set of probability distributions of actions given an observation; the random joint strategy of the follower is defined as ,in, It is the centralized element of utility company observation; It is the concentrated element of consumer observation; is the action set element of UC; the utility company is the reward function of the agent , the reward function of the consumer as an intelligent agent ,in, The weight ratio is set according to the type of consumer consumption level under weather factors. The weight ratios of consumers with different consumption level types under weather factors are different.
[0022] S203. Define the advantage function. The advantage function formula for the utility company as a leader and the consumer as a follower is: , Where, is an element of the state space set; It is the centralized element of utility company observation; It is the selected element of the utility company observation set; It is the concentrated element of consumer observation; It is the action-focused element of UC; It is the action-focused element that unites consumers;
[0023] S204. Establishing an equilibrium solution mechanism; Strategies for utility companies , the consumer's Nash equilibrium is a joint strategy , for each utility company's strategy , the consumer's best response strategy is a given utility company strategy The Nash equilibrium of consumers is: ; Assuming that consumers always adopt The utility company's strategy to maximize the value function is ,Right now: ;Utility company policies via asynchronous policy updates and consumer strategies Interactive iteration ultimately satisfies: , .
[0024] By adopting the above scheme, the "strategy confrontation logic" of game theory is transformed into the "state-action-reward" optimization problem of reinforcement learning. Based on this transformation, the complex subject interaction is deconstructed into a computable intelligent agent decision-making process. At the same time, the dynamic and information asymmetry modeling capabilities of POMG are utilized to make up for the shortcomings of traditional game theory in real-time and uncertainty processing. At the same time, differentiated value functions are set, and then the POMG value function system is used to quantify the advantages and disadvantages of strategies, providing optimization targets for solving the subsequent game equilibrium.
[0025] Preferably, the design of the asynchronous multi-agent reinforcement learning algorithm in the Stackelberg stage in step S3 includes the following sub-steps: S301-1, defining the multi-agent network architecture; adopting an improved multi-agent network structure, including: the actor network input is the sum of all consumer observations , the evaluation network output is the median value of the consumer's joint action ;
[0026] S301-2. Design an asynchronous gradient update mechanism; implement the Stackelberg game through dual-time-scale asynchronous updates; update the parameters of the utility company's policy network and the consumer's policy network as follows: , Where, 、 Both represent the learning rate for parameter update, satisfying To ensure leadership priority; represents the total derivative of the utility company's policy loss function, Represents the partial derivative of the consumer strategy loss function; where the corresponding game stage type is matched according to the current Stackelberg stage Value range and The value range is different, and the current Stackelberg stage corresponds to the type of game stage in which it is located. Value range and Value range; the current Stackelberg stage corresponds to the game stage type including the early game, mid-game and game near equilibrium period, which is determined based on the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium;
[0027] S301-3. Design the policy network loss function. The total derivative of the utility company's policy loss function represents the utility company's optimization strategy based on the best response of consumers. Based on the Stackelberg multi-agent algorithm, the policy network loss function is defined as follows: , Where, Indicates the expected value, Indicates the clipping interval; clip is a truncation function. When the importance sampling exceeds the specified upper or lower limit, the corresponding upper or lower limit is returned. and represents importance sampling; the loss function of the utility company and consumer judgment network is defined as: , ;in, and denote the outputs of the utility company evaluation network and the consumer evaluation network, respectively, and denote the outputs of the old UC evaluation network and the old consumer evaluation network respectively; the learning dynamics of the evaluation network is defined as: , ;in, and Represents the learning rate for evaluating network parameter updates.
[0028] By adopting the above scheme, an asynchronous multi-agent reinforcement learning algorithm is designed to achieve the mapping from theoretical equilibrium to actual strategy. In the Stackelberg stage, a multi-agent network including an actor network and a judge network and an asynchronous gradient update mechanism are set up to enable the utility company to prioritize updating the policy network and output parameters to form strategic guidance for consumers. Based on the latest strategy of the utility company, consumers use a shared encoder to extract common features, providing a basis for subsequent Nash games. By real-time monitoring of the type of game stage, the learning rates of the utility company and consumer strategy networks in the Stackelberg stage are dynamically adjusted, which improves the pertinence and effectiveness of strategy updates and accelerates the entire two-layer non-cooperative demand response system to reach a stable Stackelberg-Nash equilibrium.
[0029] Preferably, the design of the asynchronous multi-agent reinforcement learning algorithm in the Nash stage in step S3 includes the following sub-steps: S302, considering that consumers are independent and self-interested individuals, individuals tend to maximize their own utility ; The loss function of the policy network and the judgment network is expressed as: ,
[0030] ; The parameter update method is as follows: ;in, represents the old policy network decoder parameters, Represents the old judgment network parameters; the consumer policy network decoder and the judgment network are trained based on the multi-agent reinforcement learning algorithm; in the execution phase, the consumer's policy network It is constructed as a connection between an encoder and a decoder; wherein, the type of game stage is matched according to the current Nash stage 、 The value range is different, and the type of game stage in which the current Nash stage corresponds to is different. 、 Value range; the current Nash stage corresponds to the game stage type including the early game, mid-game and game close to equilibrium period, which is specifically determined based on the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium.
[0031] By adopting the above scheme, consumers generate individual strategies through independent decoders and design corresponding loss functions to enable consumers to achieve self-interest equilibrium under the constraints given by the utility company, avoiding the explosion of the strategy space. By dynamically adjusting the learning rate of the consumer strategy network in the Nash stage by real-time monitoring of the type of game stage, the targetedness and effectiveness of strategy updates are improved, and the entire two-layer non-cooperative demand response system is accelerated to reach a stable Stackelberg-Nash equilibrium.
[0032] Preferably, the implementation of centralized training and priority experience replay in step S4 includes the following sub-steps: S401, neural network structure design, including: the utility company strategy network has 5 layers, including: 3 parallel multi-layer perceptrons, 1 embedding layer, 2 gated recurrent unit layers and 2 fully connected layers, the input of the utility company strategy network is , the output is The consumer policy network consists of an encoder and a decoder. The encoder of the consumer policy network contains two parallel MLPs, one embedding layer, one GRU layer, and two fully connected layers. The encoder of the consumer policy network inputs the sum of the observation states. , output the sum of consumer actions The decoder consists of 3 parallel MLPs, 1 embedding layer, 1 GRU layer, and 2 fully connected layers. The decoder input is the observation state and , output consumer actions The judgment network of the utility company and the consumer has 5 layers, including: 1 security mask layer, 2 GRU layers and 2 fully connected layers; the input of the utility company judgment network is the moment The state of the Stackelberg game environment , whose parameters are updated based on the loss function; in the Stackelberg game environment, the encoders of the utility company's policy network and the consumer's policy network are updated based on the set parameter update formula; in the Nash game environment, consumers update the decoder parameters based on the multi-agent reinforcement learning algorithm by sharing the encoder parameters;
[0033] S402, constraint processing, including: using safe mask actions to filter invalid actions; setting the following to ensure that the policy networks of different agents have the same dimensional input: , Where, Replacing the advantage function that violates the constraint with ; The safety mask layer interrupts the update process of invalid actions during training. During training, the safety mask action prevents the consumer action from exceeding the feasible domain; Considering that consumers may have different action spaces, in order to ensure the scalability of the policy network when parameters are shared, the action mask is applied so that the action space not included in the consumer is not updated during training; In the loss function, Compute actions to block consumers from policy updates that do not include, An indicator representing the consumer's action space;
[0034] S403, data processing specification: the original load power of the consumer satisfies the following relationship: ;in, represents the raw load data from the load dataset, Indicates the compensation amount for the electricity purchased by the consumer; the consumer demand response power satisfies the following relationship:
[0035] ;in, For the Maximum demand response power of each consumer; load data , Maximum demand response power , distributed energy generation and real-time market electricity purchase prices is used as the environment state; considering that market electricity charges are generally settled monthly, an asynchronous multi-agent reinforcement learning algorithm is set up to use monthly load data for training;
[0036] S404, Prioritized Experience Replay; PER is estimated based on the advantage-corrected state value, and the probability of sampling the utility company's strategy trajectory and the consumer's strategy trajectory is expressed as:
[0037] , ;in, and Represent the utility company's strategic trajectory and consumer strategy trajectory The playback buffer; and Indicates the priority, 、 Is the strategy trajectory and The priority of the association, Indicates the action advantage and Ranking of trajectories when sorting;
[0038] S405. Convergence condition determination: Use the asynchronous multi-agent reinforcement learning algorithm to calculate the KL divergence of the original strategy and the updated strategy to determine whether equilibrium has been reached. When all updated strategies meet the following formula for N consecutive rounds compared with the old strategies in the strategy pool, the convergence condition is met: , ; Where N is a positive integer.
[0039] By adopting the above scheme, using asynchronous game strategy updates and safety mask constraints, and replacing the advantage function that violates the constraints, the gradient propagation of actions beyond the feasible domain can be interrupted during the training phase, preventing consumer actions from exceeding the feasible domain, ensuring the effectiveness of training and the feasibility of the strategy; utilizing the priority experience replay mechanism, high-reward strategy trajectories are trained first to accelerate convergence to the equilibrium strategy; the priority is calculated based on the ranking position of the strategy trajectory to more reasonably determine the importance of different trajectories; setting up independent strategy trajectory storage areas for power companies and consumers can clearly distinguish the strategy trajectories of the two, facilitate targeted training and optimization, and accelerate convergence to the equilibrium strategy.
[0040] Preferably, the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium and the dynamic adjustment of the strategy update frequency and step size of the utility company and the consumer include: for the utility company, monitoring and obtaining the first strategy-related parameter index after each strategy update, including: the rate of change of the grid revenue, the adjustment range of the charging and discharging power of the energy storage system, the time-of-use electricity price and the impact of the DR incentive price on the electricity sales; for the consumer, monitoring and obtaining the second strategy-related parameter index after each strategy update, including: the fluctuation range of the electricity cost, the adjustment range of the electricity load curve, the ratio of distributed energy to grid power purchase and DR The rate of change of the response quantity; by calculating the variance or standard deviation of any parameter indicator of the first strategy indicator and the second strategy indicator in a preset number of consecutive game rounds, it is judged whether the variance or standard deviation of any parameter indicator is greater than the preset variance or preset standard deviation of the corresponding parameter indicator. If it is greater, it is determined that the equilibrium point fluctuates greatly and the current degree of stability is low; otherwise, it is determined that the equilibrium point fluctuates slightly and the current degree of stability is good; by calculating the difference between any parameter indicator of the first strategy indicator and the second strategy indicator in adjacent game rounds, it is judged whether the difference of any parameter indicator is greater than the preset difference of the corresponding parameter indicator; if it is greater, it is judged that the convergence speed is fast, otherwise it is judged that the convergence speed is average; the current game stage is determined according to the judged convergence speed and stability level, and different game stages are equipped with matching convergence speeds and stability levels.
[0041] By adopting the above scheme, the relevant indicators of Stackelberg-Nash equilibrium are tracked in real time, and these indicators are used to evaluate the stability and convergence speed of the equilibrium point to determine the current game stage and assist in subsequent strategy updates and adjustments.
[0042] Preferably, the neural network structure design in step S401 also includes: in the process of implementing centralized training and priority experience replay to achieve efficiency optimization of the asynchronous multi-agent reinforcement learning algorithm, introducing a transfer learning mechanism to determine the initial parameters of the utility company and consumer strategy network, including: collecting historical demand response data under different regions, different seasons and different electricity consumption peak and valley mode combination scenarios, and when training for new scenarios, using the utility company and consumer strategy network encoder parameters obtained from training in historical scenarios whose similarity to the new scenario is greater than a preset similarity as initialization parameters.
[0043] By adopting the above solution, taking into account the differences in regions, seasons, and peak and trough electricity consumption patterns, demand response training and acquisition are often carried out separately for different scenarios. Then, the initial parameters of the strategy network are set in combination with the transfer learning mechanism under different scenario conditions to assist in more accurate completion of demand response acquisition.
[0044] In a second aspect, the present application provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method as described above.
[0045] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a program stored and executable on the memory, wherein the program implements the steps of the above method when executed by the processor.
[0046] In summary, this application has the following beneficial effects: It constructs a two-layer dynamic game model, considers consumer heterogeneity, and specifically designs a model structure for utility companies and consumers, enabling them to reach equilibrium through strategic interaction. This addresses the lack of an interest coordination mechanism in traditional methods, maximizes grid operating revenue, minimizes consumer electricity costs, and enhances the synergy of interests among multiple parties. It models a partially observable Markov game and designs differentiated reward functions for both consumers and utility companies, taking into account dynamic constraints such as the decay of the energy storage system's state of charge and time-varying electricity prices. This overcomes the shortcomings of traditional methods in modeling temporal dynamics, optimizes grid economics, distributed energy consumption efficiency, and user comfort. It employs an asynchronous multi-agent proximal policy optimization algorithm, combined with a dynamic adaptive policy update mechanism based on real-time monitoring of Stackelberg-Nash equilibrium indicators, a centralized training decentralized execution architecture, and a prioritized experience replay mechanism. This addresses the dimensionality explosion problem of the policy space in traditional multi-agent methods, improves the efficiency of multi-agent collaboration, accelerates convergence to an equilibrium strategy, and ensures algorithm convergence speed. Together, these achieve efficient consumption of distributed energy resources (DER), load curve optimization, and synergy among multiple parties. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow chart of the method described in a specific embodiment; Figure 2 is a flow chart of step S1 in the method described in a specific embodiment; Figure 3 is a flow chart of step S2 in the method described in a specific embodiment; Figure 4 is a flow chart of step S3 in the method described in a specific embodiment; Figure 5 is a flow chart of step S4 in the method described in a specific embodiment; Figure 6 The framework and update process of the asynchronous multi-agent reinforcement learning algorithm in the method described in the specific embodiment; Figure 7 This is the specific process of the asynchronous multi-agent reinforcement learning algorithm in the method described in the specific embodiment. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0049] like Figure 1 As shown, an embodiment of the present application discloses a two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning, and the specific steps include: S1, constructing a two-layer non-cooperative game framework.
[0050] Considering that traditional single-level optimization only focuses on the unilateral goals of the utility company (UC) or the user, and ignores the strategic confrontation between the two parties (for example, users may resist electricity price strategies that are unfavorable to themselves), a two-level non-cooperative game framework is set up to characterize the optimal strategy combination under the conflict of interests between the two parties through non-cooperative games, ensuring that the solution is executable in actual interactions; specifically, the game structure defines the conflicting subjects and the strategy interaction, such as Figure 2 As shown in FIG, the specific steps are as follows: construct a two-layer dynamic game model between utility companies and consumers; wherein, the upper-layer utility company strategy includes: maximizing the grid revenue by adjusting the charging and discharging power of the energy storage system, setting time-of-use electricity prices and DR incentive prices; the lower-layer consumer strategy includes: considering the external factor restrictions of weather factors and the internal factor restrictions of consumers' own consumption levels, by adaptively adjusting the electricity load curve, the proportion of distributed energy use and the DR response amount based on internal and external factor restrictions, to minimize the electricity cost in the non-cooperative Nash game; the equilibrium goal includes: convergence to the Stackelberg-Nash equilibrium after the interaction of the two parties' strategies; S101, define the game hierarchy.
[0051] Based on the Stackelberg-Nash dynamic game theory, a non-cooperative game architecture with two layers of decision-making entities is established. The upper layer of this architecture is the leading utility company (UC), which optimizes its operating strategy (e.g., electricity pricing, energy storage scheduling, DER incentives, etc.) through Stackelberg game theory to maximize grid revenue. The lower layer is the follower group of electricity consumers, who, taking into account external constraints such as weather factors and internal constraints such as their own consumption levels, use non-cooperative Nash game theory to make autonomous decisions (e.g., load shifting, DER usage ratio) to minimize their individual electricity costs. The specific mathematical model is expressed as follows:
[0052] ; ; ;in, and They correspond to the utility companies and The strategy vector of each consumer, i.e. the decision variable containing the time series; for The corresponding optimal response strategy is To correspond The optimal response strategy; Strategic feasible domain for utility companies; Utility Company Strategies for Consumers The feasible domain under Take 1 or 2, when Taking 1 corresponds to the consumer with good consumption level considering weather factors using the utility company strategy The feasible domain under Taking 2 corresponds to the consumer with average consumption level considering weather factors using the utility company strategy The feasible domain under Characterization Consumers in other consumer strategies and utility company strategies The utility of the following; Characterization Consumers with good consumption level considering weather factors have other consumer strategies and utility company strategies The utility of the following; Characterization Consumers with average spending power considering weather factors have other consumer strategies and utility company strategies The utility of the following; The utility of the utility company under the consumer's optimal response strategy; Represent the consumer set, take When the weather factor is taken into account, it represents consumers with a good consumption level. When , it represents consumers with average consumption level considering weather factors; Indicates that except All consumers managed by UC actively participate in demand response and remain completely rational in load scheduling to maximize benefits. Consumers are not equipped with distributed energy resources (DER) or energy storage systems (ESS), and DER and ESS data are fully open to consumers and UC.
[0053] It is clear that UC and consumers act alternately in time sequence, and both parties predict the expected utility of the future T time span based on the current state: and , the DR framework is modeled as a dynamic programming problem under incomplete information, requiring both parties to update their strategies through rolling horizon optimization (RHC).
[0054] When the Stackelberg-Nash equilibrium is reached, the dual constraint is satisfied: the utility company's strategy Maximize your own utility under the optimal response of all consumers, the consumer strategy The Nash equilibrium is formed as follows: , ; S102, upper-layer UC strategy modeling.
[0055] In the demand response framework, the optimization model of the utility company (UC) can be formally expressed as follows: , ; Among them, the decision variables It contains eight types of control parameters on the time series; Indicates the amount of electricity purchased from the main network. To sell electricity to users, is the abandoned distributed energy power, and Respectively represent the charging power and discharging power of the energy storage system, Represents the amount of charging using distributed energy, The amount of electricity purchased for the grid, is the expected demand response power; For electricity sales revenue, The cost of purchasing electricity, For distributed energy consumption income, flexibility costs for demand response; It refers to parameter maximization; the objective function consists of four parts: electricity sales revenue The formula is: Among them, time-of-use tiered electricity prices Peak and valley periods And the user's cumulative electricity consumption A piecewise linear function of Purchase electricity from UC for users; electricity purchase cost The formula is: Among them, the real-time electricity purchase price Superimpose random disturbances that follow a truncated normal distribution (mean 0, variance 0.03, and margins ±3%) ; Distributed energy consumption benefits The formula is: ;in, is the DER consumption incentive coefficient, is the DER power generation, For discarded power, Penalty price for abandonment; demand response flexibility costs The formula is: ; Among them, through the marginal benefit coefficient and basic compensation unit price Adjusting User DR Power and target value degree of matching.
[0056] Under the constraints of the consumer's optimal response strategy, the definition of the feasible domain XU of the public utility company (UC) strategy is shown in formula (6), and its specific mathematical representation is as follows:
[0057] Among them, formula (11) defines the upper bound of the abandoned power, formulas (12) and (13) define the charging power and output power range of the energy storage system (ESS), formula (14) defines the expected power boundary of demand response, formula (17) represents the energy conservation constraint of the distribution network, and formula (18) prevents the simultaneous charging and discharging of the ESS through non-coincidence constraints. The state of charge dynamics of the ESS is described by formula (19):
[0058] ;in, The real-time charge status of ESS, is the maximum energy storage capacity, 、 They represent the charge and discharge efficiency coefficients, is the natural attenuation factor of the state of charge; the piecewise function of the ESS charging power is the charging power function Designed to: ;in, is the SOC threshold, is the charging power attenuation ratio after reaching the threshold, is the rated charging power of ESS; the model realizes adaptive charging and discharging control of energy storage system through the state of charge (SOC) feedback mechanism.
[0059] S103, Lower-level consumer strategy modeling. The optimization model of a consumer is shown in formula (21), and the consumer strategy is shown in formula (22). Constraint set must be met ;
[0060] ; ;in, , ; Considering consumers’ UC strategy Optimize its own cost as a premise; Formula (21) contains three parts: First, the expenditure of purchasing electricity from UC can be expressed as : ;in, Indicates the The amount of electricity purchased by a consumer from UC; secondly, the consumer pays the electricity price The benefits of absorbing distributed energy generation can be expressed as : Among them, consumers give priority to DER electricity (because the price is lower than UC electricity), reflecting the non-cooperative game relationship between consumers; is the amount of electricity purchased by consumers from DER. The DER electricity price is: , and are the slope and intercept parameters of the DER price; then, the utility function of consumers participating in load dispatch and transfer is : ;in, For the electricity comfort price, ; To reduce the load, It is a transferable load; consumers participate in DR through demand response incentives.
[0061] Based on the Stackelberg game hypothesis, the feasible domain of consumer strategy and UC strategy Related; Consumer Strategy Feasible Domain Defined as:
[0062]
[0063] Among them, Equation (26) represents DER power purchase and balance; Equation (28) represents consumers' active participation in demand response; Equation (29) is the adjustable load constraint; Equation (30) is the transferable load constraint; Equation (31) indicates that the DR power of all consumers does not exceed the UC expected value.
[0064] In addition, in the above formula 、 , it is necessary to keep the satisfaction ratio of different consumers within a reasonable range to ensure the rationality of subsequent strategic interactions. The value range is determined by a pre-built distributed energy use or grid electricity purchase decision model that considers weather and consumption levels. The pre-built distributed energy use or grid electricity purchase decision model that considers weather and consumption levels uses a neural network algorithm. The input is the consumer consumption level type and weather factors under the current i-th consumer considering weather factors, and the output is The value range of is generated by training based on historical weather factors, the consumption level type of each consumer under weather factors, and the ratio of the electricity actually purchased by each consumer from distributed energy resources to the electricity purchased by each consumer from the utility company.
[0065] Accordingly, and The specific value must also be within a reasonable range to ensure the rationality of subsequent strategy interactions. and The corresponding value range is determined according to the pre-built decision model for dynamic adjustment of the electricity load curve considering weather and consumption level; the pre-built decision model for dynamic adjustment of the electricity load curve considering weather and consumption level adopts a neural network, the input of which is the consumer consumption level type and weather factors under the current i-th consumer considering weather factors, and the output is and The corresponding value range is generated through training based on historical weather factors, the consumption level type of each consumer under weather factors, the actual load reduction range of each consumer, and the load transfer range of each consumer.
[0066] S104. Establish a two-layer game interaction mechanism. Considering that the UC strategy directly limits the consumer's feasible decision space (such as time-of-use electricity prices affecting the choice of electricity consumption period), and the consumer's electricity consumption behavior (such as load peak-valley difference) reacts to the UC's operating costs, a closed-loop causal chain is formed, corresponding to the strategic interaction of the two-layer game; specifically, the demand response framework is implemented through the following dual game structure: leader layer (Stackelberg game). As the leader, the utility company (UC) dominates the game through the following four-dimensional strategy space: energy storage system (ESS) charging and discharging power regulation; demand response capacity allocation ; Time-of-use electricity price mechanism (TUTT) setting; Distributed energy consumption incentive pricing .
[0067] Follower layer (Nash game). The consumer group acts as a follower, taking into account the external constraints of weather factors and the internal constraints of the consumer's own consumption level, and participates in the two-layer interaction through the following decision variables based on the adaptability of internal and external constraints: load curve scheduling ; Shifting electricity usage periods DER / grid power purchase ratio At the same time, non-cooperative competition is formed among consumers at the same level: competing for limited DER consumption quotas ; Bid for DR incentive compensation.
[0068] Cross-layer strategy coupling; UC through Influencing the consumer's feasible domain , and consumers through Counteraction to UC Cost items form a closed-loop game chain.
[0069] S2. Modeled as a partially observable Markov game.
[0070] First, considering that the UC cannot fully know the consumer's DER output, actual electricity usage preferences, etc., and the consumer cannot predict the UC's real-time strategy adjustments (such as sudden ESS charging and discharging), POMG is used to model this information asymmetry through the local observation space; secondly, the temporal state transition of POMG is used to decompose the game into continuous stages. In each stage, the strategy interaction between the UC and the consumer forms a dynamic equilibrium sequence, which is close to the real-time dispatch demand of the power grid. In addition, the current state in POMG (such as ESS charge state, real-time electricity price) only depends on the previous state and action, rather than the entire historical trajectory, which can simplify the complexity of dynamic modeling; finally, POMG is combined with reinforcement learning to directly fit the continuous strategy space through neural networks to improve decision-making accuracy; in summary, it can be judged that the two-layer non-cooperative game structure is logically compatible with the POMG characteristics, and the structured conversion of game elements to POMG tuples is selected to complete the closed loop from game interaction to POMG state transition, such as Figure 3 As shown, the specific steps include: S201, defining a POMG tuple.
[0071] Based on POMG, the utility company (UC) and each consumer are considered as POMG agents. A stage-wise version of a general and asynchronous mobile Markov game is defined by the tuple:
[0072] ;
[0073] In the formula, the state space yes The set of UC observations yes The set of consumers' observations yes A collection of UC action sets yes A collection of actions of joint consumers yes The collection of The action set of a consumer yes A collection of is the number of steps in each round; is the utility company's reward function, is the consumer’s reward function, is the set of transfer kernels; the reward function of the UC agent is defined as: , by the current state Calculated; the reward function of the consumer agent is defined as ;in, The weight ratio is set according to the type of consumer consumption level under weather factors. Consumers with different types of consumer consumption levels under weather factors have different matching weight ratios. Specifically, consumers with good consumption levels under weather factors are more pursuing environmentally friendly consumption, and are more inclined to energy generation. They will not be too concerned about the expenditure on purchasing electricity or adjusting load scheduling. Therefore, Compared to 、 On the contrary, considering the weather factors, the consumption level of consumers is average. Compared to 、 The proportion is even smaller; is the set of transfer cores.
[0074] S202. Construct a value function system.
[0075] In addition to the POMG modeling corresponding tuples constructed based on the corresponding mapping of the elements of the two-level non-cooperative game framework mentioned above: decision-making subject-agent, strategy space-action space (UC actions and consumer actions), game state-state space (global state and local state), information asymmetry-observation space (UC can observe the global state, consumers only observe the local state), benefit target-reward function (UC reward and consumer reward), it is also necessary to construct a value function based on the purpose of equilibrium goal. That is, by solving the Nash equilibrium value function of POMG, it is equivalent to the Stackelberg-Nash equilibrium (SNE) of the two-level game. Specifically, it includes:
[0076] Leader's random strategy is a set of probability distributions of actions under given observations; meanwhile, the random joint strategy of the follower is defined as ; Given a strategy , the state value functions of the leader (UC) and follower (consumer) are defined as follows: , .
[0077] S203. Define advantage function.
[0078] The advantage functions of the leader (UC) and follower (consumer) are defined as follows: , ; In this framework, the advantage function intuitively expresses whether an action performs better or worse than average in a given state.
[0079] S204. Establish a balance solution mechanism.
[0080] Strategies for UCs , the consumer's Nash equilibrium is a joint strategy , for each utility company's strategy , the consumer's optimal response strategy is defined as "parameter maximization" ( ), where the consumer's best response strategy is is a given utility company strategy The Nash equilibrium of consumers is: .
[0081] The Stackelberg-Nash equilibrium (SNE) of UC is the "best response to the best response"; under the assumption that consumers always adopt The utility company's strategy to maximize the value function is ,Right now:
[0082] .
[0083] Through asynchronous policy updates, utility company policies and consumer strategies Interactive iteration ultimately satisfies: , .
[0084] S3. Design an asynchronous multi-agent reinforcement learning algorithm (SN-MAPPO).
[0085] To further accurately solve the game equilibrium, the asynchronous update of SN-MAPPO is designed to simulate the real decision sequence, adapt to time-varying constraints, avoid synchronization conflicts, and obtain a more accurate solution to the game equilibrium. In addition, to further improve the pertinence and effectiveness of strategy updates and accelerate the entire two-layer non-cooperative demand response system to achieve a stable Stackelberg-Nash equilibrium, the convergence speed and stability indicators of the Stackelberg-Nash equilibrium are monitored in real time, and the strategy update frequency and step size of the utility company and the consumer are dynamically adjusted; Figure 4 As shown, the specific steps include: S301-1, defining the multi-agent network architecture.
[0086] Adopting an improved multi-agent network structure: the actor network input is the sum of all consumer observations , the evaluation network output is the median value of the consumer's joint action .
[0087] This algorithm is trained based on a "centralized training, decentralized execution" paradigm. During centralized training, the UC and consumers share global observation information. This global observation updates the policy network, outputting parameters such as time-of-use electricity prices and ESS power, thereby providing "policy guidance" for consumers. During decentralized execution, consumers rely solely on local observations to generate policies, achieving a self-interested equilibrium within the constraints imposed by the UC and avoiding "strategy space explosion." Accordingly, a shared encoder is used to extract common features (such as power and price trends) between the UC and consumers, which are shared by all agents. Independent decoders are used, allowing consumers to independently optimize their policies (such as load scheduling) to achieve self-interested decision-making in the Nash game.
[0088] S301-2. Design an asynchronous gradient update mechanism.
[0089] The asynchronous gradient update mechanism is a core feature of the SN-MAPPO algorithm. To avoid the "strategy interference" problem of synchronous updates and accelerate convergence to the Stackelberg-Nash equilibrium (SNE), the Stackelberg game is implemented through dual-time-scale asynchronous updates. In the Stackelberg phase, the UC prioritizes updating the policy network, and the consumer updates the shared encoder parameters based on the latest UC policy. In the Nash phase, the consumer performs Nash game optimization (asynchronous to the UC update) using an independent decoder under a fixed UC policy. Specifically, the UC policy network parameter updates and the consumer policy network parameter updates are as follows: , Where, 、 Both represent the learning rate for parameter update, satisfying To ensure leadership priority; represents the total derivative of the utility company's policy loss function, represents the partial derivative of the consumer strategy loss function; is defined as follows: .
[0090] Among them, the corresponding game stage type is matched according to the current Stackelberg stage. Value range and The value range is different, and the current Stackelberg stage corresponds to the type of game stage in which it is located. Value range and Value range; Among them, the current Stackelberg stage corresponds to the game stage type including the early game, mid-game and near-equilibrium period, which is determined based on the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium; Specifically, in the early game, as the leader in the Stackelberg game, the utility company gives priority to adjusting its strategy and can choose to speed up the update and iteration speed of the utility company's strategy network parameters, so that it can try more different energy storage system charging and discharging strategies, time-of-use electricity price setting schemes and DR incentive price combinations in a shorter time; the lower-level consumer agent synchronously updates the shared strategy encoder parameters based on the updated strategy of the upper-level utility company. As the game progresses, when the monitoring module detects that the Stackelberg-Nash equilibrium point gradually stabilizes and the convergence speed reaches a certain threshold, the dynamic adaptive strategy update mechanism will choose to reduce the frequency of strategy updates. The lower-level consumers also reduce the frequency of strategy updates and further optimize electricity costs by fine-tuning the power consumption strategy under the relatively stable strategy of the utility company. Correspondingly, according to the early game, mid-game and near-equilibrium period, Value range and The value range gradually becomes smaller.
[0091] Specifically, the following are: for utility companies, monitoring and obtaining the first strategy-related parameter indicators after each strategy update, including: the rate of change of grid revenue, the adjustment range of energy storage system charging and discharging power, the impact of time-of-use electricity prices and DR incentive prices on electricity sales; for consumers, monitoring and obtaining the second strategy-related parameter indicators after each strategy update, including: the fluctuation range of electricity costs, the adjustment range of electricity load curves, the ratio of distributed energy to grid power purchases, and DR The rate of change of the response quantity; by calculating the variance or standard deviation of any parameter indicator of the first strategy indicator and the second strategy indicator in a preset number of consecutive game rounds, it is judged whether the variance or standard deviation of any parameter indicator is greater than the preset variance or preset standard deviation of the corresponding parameter indicator. If it is greater, it is determined that the equilibrium point fluctuates greatly and the current degree of stability is low. Otherwise, it is determined that the equilibrium point fluctuates slightly and the current degree of stability is good. By calculating the difference of any parameter indicator of the first strategy indicator and the second strategy indicator in adjacent game rounds, it is judged whether the difference of any parameter indicator is greater than the preset difference of the corresponding parameter indicator. If it is greater, it is judged that the convergence speed is fast, otherwise it is judged that the convergence speed is average. The current game stage is determined according to the judged convergence speed and stability. Different game stages are equipped with matching convergence speeds and stability levels. For example, good convergence speed and good stability match the game close to equilibrium period, average convergence speed and average stability match the early stage of the game, and the rest correspond to matching the mid-game period.
[0092] S301-3. Design the strategy network loss function.
[0093] The total derivative of the UC strategy loss function indicates that the UC optimizes the strategy based on the best response of the consumer. Based on the Stackelberg multi-agent algorithm, the UC strategy loss function design includes an importance sampling truncation term. The truncation interval is set to avoid the extreme values of importance sampling from having too much impact on the strategy update and stabilize the strategy optimization process. In addition, an entropy regularization term is set to encourage strategy exploration, which allows the power company strategy network to try more different strategies, thereby discovering a better strategy combination, enhancing the diversity and adaptability of the strategy, and helping to improve the performance of the overall system and the long-term benefits of the power company. Specifically, the loss function of the corresponding strategy network is defined as follows:
[0094] ;in, Indicates the expected value, Indicates the clipping interval; clip is a truncation function. When the importance sampling exceeds the specified upper or lower limit, the corresponding upper or lower limit is returned. and Represents importance sampling; specifically: .
[0095] To encourage exploration, an entropy regularization term is applied to the loss function, as follows:
[0096] ;in, and Represents the KL divergence between the new policy and the old policy.
[0097] Correspondingly, the loss function of UC and consumer judgment network is defined as: ,
[0098] ;in, and denote the outputs of the utility company evaluation network and the consumer evaluation network, respectively, and Represent the outputs of the old UC evaluation network and the old consumer evaluation network respectively.
[0099] The learning dynamics of the critic network is defined as: ;in, and Represents the learning rate for judging network parameter updates; MAPPO usually adds an entropy regularization term to the loss function according to the following formula.
[0100] S302. Design the loss functions of the strategy network and the evaluation network.
[0101] Since consumers are independent and self-interested individuals, they tend to maximize their own utility. , based on the Nash stage: the loss function of the policy network and the judgment network is expressed as follows:
[0102] ,
[0103] .
[0104] The parameter update method is as follows: , ;in, represents the old policy network decoder parameters, Represents the old judgment network parameters; the consumer policy network decoder and the judgment network are trained based on MAPPO. In the execution phase, the consumer's policy network It is constructed as a connection between an encoder and a decoder.
[0105] Among them, the type of game stage corresponding to the current Nash stage is matched 、 The value range is different, and the type of game stage in which the current Nash stage corresponds to is different. 、 The value range; the current Nash stage corresponds to the game stage type including the early game, mid-game and game close to equilibrium period, which is specifically determined based on the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium; accordingly, according to the early game, mid-game and game close to equilibrium period, The value range gradually becomes smaller.
[0106] Specifically, such as Figure 6 As shown in the figure, the framework and update process of SN-MAPPO are specifically divided into Stackelberg game and Nash game environments, and the specific design of strategy network and evaluation network parameter update and loss function is carried out.
[0107] S4. Implement centralized training and prioritized experience replay.
[0108] In order to further improve the efficiency and effectiveness of the algorithm, the centralized training decentralized execution (CTDE) architecture and the PER mechanism are selected. The CTDE architecture improves generalization by sharing policy network encoder parameters, while the PER mechanism performs weighted sampling of historical trajectories based on action advantage values, giving priority to training high-reward policy trajectories to accelerate convergence. Figure 5 As shown, the specific steps include: S401, designing a neural network structure.
[0109] Strategy Encoder parameters are shared parameters, shared by all consumer agents; while the decoder parameters is a characteristic parameter, which is updated by consumers in the Nash game environment. In the Stackelberg game environment, the UC strategy network and Consumer Strategy Network The encoder is updated based on formulas (40) and (41); in the Nash game environment, consumers share the encoder parameters , the decoder parameters are updated based on the MAPPO algorithm. Since the UC's actions affect the Nash game results, the UC strategy and the consumer strategy are updated asynchronously, and the final strategy converges to the stochastic Nash equilibrium (SNE).
[0110] In order to achieve the above parameter update, the specific design of the neural network structure includes:
[0111] First, the UC policy network consists of 5 layers: 3 parallel multi-layer perceptrons (MLPs), 1 embedding layer, 2 gated recurrent unit (GRU) layers, and 2 fully connected layers. The input and output of the UC policy network are and The consumer policy network consists of an encoder and a decoder. The encoder of the consumer policy network contains two parallel MLPs, one embedding layer, one GRU layer, and two fully connected layers. The encoder of the consumer policy network inputs the sum of the observation states. , output the sum of consumer actions The decoder consists of 3 parallel MLPs, 1 embedding layer, 1 GRU layer, and 2 fully connected layers. The decoder input is the observation state and , output consumer actions The UC and consumer critic network consists of 5 layers: 1 security mask layer, 2 GRU layers and 2 fully connected layers. The input of the UC critic network is the moment The state of the Stackelberg game environment , whose parameters are updated based on the loss function; and, considering that consumers in the POMG framework may have suboptimal strategies due to incomplete observation information (e.g., they cannot perceive the DER output of other users), an attention mechanism is introduced into the consumer strategy network to assist users in focusing on key observation information (e.g., DER output in neighboring areas and real-time electricity prices) and reducing redundant information interference.
[0112] Secondly, in SN-MAPPO, the encoder of the consumer policy network is shared by all consumers. The consumer policy network encoder and the UC policy network are updated based on formulas (43) and (44), and the consumer policy network decoder is updated based on formula (52). The consumer policy network encoder aims to obtain the Stackelberg equilibrium between the UC and the consumer, while the policy network decoder is used to obtain the Nash equilibrium between consumers. The MLP is used to capture the range differences in power and electricity prices. The output vectors of the MLP are concatenated and input into the embedding layer to realize the vectorization of the observation information in the UC and consumer policy network. SN-MAPPO uses a batch normalization layer before each activation layer of the policy network to improve convergence.
[0113] Furthermore, given the real-world differences between cities, the demand responses between utility companies and consumers in different regions vary significantly. Consequently, the demand response between utility companies and consumers is typically determined based on a single city. Therefore, during the implementation of centralized training and prioritized experience replay to optimize the efficiency of the asynchronous multi-agent reinforcement learning algorithm, a transfer learning mechanism is introduced to determine the initial parameters of the utility and consumer policy networks. This involves collecting historical demand response data for scenarios based on different regions, seasons, and peak / valley patterns (several preset peak / valley patterns statistically divided according to the numerical ranges corresponding to peak, peak, flat, and valley periods). When training new scenarios (untrained scenarios with regions, seasons, and peak / valley patterns), the encoder parameters of the utility and consumer policy networks trained in historical scenarios (trained scenarios with regions, seasons, and peak / valley patterns) with a similarity greater than a preset similarity to the new scenario are used as initialization parameters.
[0114] S402: Constraint processing.
[0115] Through the centralized training decentralized execution architecture and the priority experience replay mechanism for policy training, the efficiency of multi-agent collaboration can be improved and convergence can be accelerated. On this basis, the consumer policy network setting includes a safety mask layer to filter invalid actions. By replacing the advantage function that violates the constraints, the gradient propagation of actions outside the feasible domain can be interrupted during the training phase, preventing consumer actions from exceeding the feasible domain, ensuring the effectiveness of training and the feasibility of the strategy. To filter invalid actions, safety mask actions are used. To ensure that the policy networks of different agents have inputs of the same dimensionality:
[0116] ; Among them, the SM operator Replacing the advantage function that violates the constraint with ;Safety Mask LayerThe safety mask layer can interrupt the update process of invalid actions during training. During training, the safety mask action prevents the consumer action from going outside the feasible region.
[0117] Since consumers may have different action spaces, to ensure the scalability of the policy network when parameters are shared, an action mask is applied so that the action spaces not included in the consumer are not updated during training. Compute actions to block consumers from policy updates that do not include, An indicator representing the consumer's action space.
[0118] S403. Design data processing specifications.
[0119] In order to further ensure the accuracy of the training data, the original load power of the consumer is designed to meet the following relationship:
[0120] ;in, represents the raw load data from the load dataset, Indicates the amount of compensation for the electricity purchased by the consumer; the consumer demand response power satisfies the following relationship:
[0121] ;in, For the Maximum demand response power of each consumer; load data , Maximum demand response power , distributed energy generation and real-time market electricity purchase prices is used as the environment state; considering that market electricity charges are generally settled on a monthly basis, an asynchronous multi-agent reinforcement learning algorithm is set up to use monthly load data for training.
[0122] S404. Apply Priority Experience Replay (PER).
[0123] By weighting historical trajectories based on action dominance, we prioritize training high-reward strategies (e.g., those with efficient DER absorption and reduced load peak-to-valley differences), reducing ineffective training and improving convergence efficiency. Specifically, PER is estimated based on the dominance-corrected state value, and the probability of sampling a specific strategy trajectory is expressed as: , ;in, and represent the utility company's strategic trajectory and consumer strategy trajectory The playback buffer; and Indicates the priority, 、 Is the strategy trajectory and The priority of the association, Indicates the action advantage and The ranking of the tracks when sorting.
[0124] S405: Determine the convergence condition.
[0125] The SN-MAPPO algorithm maintains a policy pool to store the original policies before updates. The algorithm calculates the KL divergence between the original and updated policies to determine whether equilibrium has been reached. Convergence is achieved when all updated policies meet the following equation for 50 consecutive rounds compared to the old policies in the policy pool:
[0126] ; Using KL divergence threshold detection as the convergence condition judgment, when the KL divergence of the updated strategy and the old strategy for 50 consecutive rounds is lower than the set threshold, the training is terminated. This can accurately determine whether the algorithm has reached a balanced state, ensure the stability and convergence of the strategy, avoid unnecessary iterations, and save computing resources and time; Applying the above solution, if Figure 7 As shown, based on the two-layer non-cooperative game structure with the utility company and the consumer group as the decision-making subjects, POMG and SN-MAPPO are linked to solve SNE, which specifically includes: initializing the Stackelberg game environment, initializing the UC and consumer strategy network, judging whether SNE is achieved, and outputting the UC strategy and consumer strategy; if the judgment result is that SNE is not achieved, asynchronous interaction and strategy update are performed accordingly, including: in the Stackelberg game stage, action collection, execution and storage, calculation of advantages and losses, updating of the judgment network and strategy network, updating of the Nash game environment, parameter synchronization, action collection, execution and storage, calculation of advantages and losses, updating of the judgment network and strategy network for each consumer agent, and then repeating the steps to achieve SNE.
[0127] In summary, this application addresses the limitations of traditional electricity demand response (DR) methods in the smart grid environment and proposes a two-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning. By constructing a Stackelberg-Nash two-layer non-cooperative game framework and combining it with the asynchronous multi-agent proximal strategy optimization algorithm (SN-MAPPO), efficient consumption of distributed energy resources (DER), load curve optimization and multi-party interest coordination are achieved.
[0128] The embodiments of the present application also disclose a computer-readable storage medium. Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and execute the above-mentioned two-tier non-cooperative demand response method based on asynchronous multi-agent reinforcement learning. The computer-readable storage medium includes, for example, various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk. The embodiments of the present application also disclose a computer device. Specifically, the computer device includes a memory and a processor, and the memory stores a computer program that can be loaded by the processor and execute the above-mentioned two-tier non-cooperative demand response method based on asynchronous multi-agent reinforcement learning.
Claims
1. A two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning, characterized in that: include: S1. Construct a two-tier non-cooperative game framework, including: constructing a two-tier dynamic game model between utility companies and consumers; wherein the upper-tier utility company strategy includes: maximizing grid revenue by adjusting the charging and discharging power of the energy storage system, setting time-of-use electricity prices, and DR incentive prices; the lower-tier consumer strategy includes: minimizing electricity costs in a non-cooperative Nash game by adaptively adjusting the electricity load curve, distributed energy utilization ratio, and DR response based on internal and external constraints, taking into account external constraints such as weather factors and internal constraints such as consumer consumption levels; the equilibrium goal includes: convergence to a Stackelberg-Nash equilibrium after the interaction of the two strategies; S2. Complete the partially observable Markov game modeling, transform the two-level non-cooperative game into a partially observable Markov game, and define the agent, state space, observation space, action space, reward function, and value function; S3. Design an asynchronous multi-agent reinforcement learning algorithm as the game equilibrium solution engine. The solution process includes: using an asynchronously updated multi-agent proximal policy optimization algorithm. In the Stackelberg phase, the upper-level utility company prioritizes updating the policy network, and consumers update the shared policy encoder parameters based on the upper-level utility company's policy. In the Nash phase, consumers use independent policy decoders to optimize the Nash game under a fixed utility company policy. The policy network uses a shared encoder and independent decoder structure to ensure decentralized execution. The convergence speed and stability indicators of the Stackelberg-Nash equilibrium are monitored in real time, and the policy update frequency and step size of the utility company and consumers are dynamically adjusted. S4. Implement centralized training and prioritized experience replay to optimize the efficiency of asynchronous multi-agent reinforcement learning algorithms, including: adopting a centralized training-distributed execution architecture, where utilities and consumers share policy network encoder parameters; using a prioritized experience replay mechanism, weighted sampling of historical trajectories based on action advantage values, prioritizing training of high-reward policy trajectories, and accelerating convergence to an equilibrium strategy.
2. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 1 is characterized in that: The construction of the two-layer non-cooperative game framework in step S1 includes the following sub-steps: S101. Define the game hierarchy; this includes: establishing a non-cooperative game architecture with two layers of decision-makers, namely a two-layer non-cooperative game architecture. In this architecture, the upper layer is the leader, the utility company, which uses Stackelberg game optimization strategy to maximize grid revenue; the lower layer is the follower group of electricity consumers, which considers external factors such as weather factors and internal factors such as consumers' consumption levels, and minimizes individual electricity costs through non-cooperative Nash game. The specific mathematical model is expressed as follows: ; in, and They correspond to the utility companies and The strategy vector of each consumer, i.e. the decision variable containing the time series; for The corresponding optimal response strategy is To correspond The optimal response strategy; Strategic feasible domain for utility companies; Utility Company Strategies for Consumers The feasible domain under Take 1 or 2, when Taking 1 corresponds to the consumer with good consumption level considering weather factors using the utility company strategy The feasible domain under Taking 2 corresponds to the consumer with average consumption level considering weather factors using the utility company strategy The feasible domain under Characterization Consumers in other consumer strategies and utility company strategies The utility of the following; Characterization Consumers with good consumption level considering weather factors have other consumer strategies and utility company strategies The utility of the following; Characterization Consumers with average spending power considering weather factors have other consumer strategies and utility company strategies The utility of the following; The utility of the utility company under the consumer's optimal response strategy; Represent the consumer set, take When the weather factor is taken into account, it represents consumers with a good consumption level. When , it represents consumers with average consumption level considering weather factors; Indicates that except Individuals other than consumers; When the Stackelberg-Nash equilibrium is reached, the dual constraint is satisfied: the utility company's strategy Maximize your own utility under the optimal response of all consumers, the consumer strategy The Nash equilibrium is formed as follows: ; S102. Upper-level utility company strategy modeling. In the demand response framework, the utility company's optimization model formula includes: ; Among them, the decision variables It contains eight types of control parameters on the time series; Indicates the amount of electricity purchased from the main network. To sell electricity to users, is the abandoned distributed energy power, and Respectively represent the charging power and discharging power of the energy storage system, Represents the amount of charging using distributed energy, The amount of electricity purchased for the grid, is the expected demand response power; For electricity sales revenue, The cost of purchasing electricity, For distributed energy consumption income, flexibility costs for demand response; It refers to parameter maximization; S103, Lower-level Consumer Strategy Modeling, The optimization model for a consumer is: ; in, , ; Indicates the The amount of electricity purchased by a consumer from distributed energy resources, Indicates the The amount of electricity a consumer buys from a utility company, The value range is determined by a pre-built distributed energy use or grid power purchase decision model that takes into account weather and consumption levels. The pre-built distributed energy use or grid power purchase decision model that takes into account weather and consumption levels adopts a neural network algorithm, and the input is the current Considering the consumer consumption level type and weather factors under weather factors, the output is The value range of is generated by training based on historical weather factors, the consumption level type of each consumer under weather factors, and the ratio of the amount of electricity each consumer actually purchased from distributed energy resources to the amount of electricity each consumer purchased from the utility company. Indicates the Each consumer can reduce load, Indicates the consumers can transfer load, and and The corresponding value range is determined by a pre-built decision model for dynamically adjusting the electricity load curve considering weather and consumption levels; the pre-built decision model for dynamically adjusting the electricity load curve considering weather and consumption levels adopts a neural network, and the input is the current Considering the consumer consumption level type and weather factors under weather factors, the output is and The corresponding value range is generated by training based on historical weather factors, the consumption level type of each consumer under weather factors, the actual load reduction range of each consumer, and the load transfer range of each consumer; For the What a consumer spends on electricity from a utility company; For the electricity price Absorbing the benefits of distributed energy generation; For the The utility function of each consumer participating in load dispatch and transfer; S104. Establish a two-tier game interaction mechanism; As leaders, utility companies dominate the game through energy storage system charging and discharging power and regulation, demand response capacity allocation, time-of-use electricity price mechanism setting, and distributed energy consumption incentive pricing: As followers, the consumer group takes into account external constraints such as weather factors and internal constraints such as consumer consumption levels. They participate in two-tier interactions by adaptively utilizing load curve scheduling, power consumption time shifting, and the proportion of distributed energy and grid power purchases based on these constraints. At the same time, non-cooperative competition is formed among consumers at the same level, competing for limited distributed energy consumption quotas and bidding for demand response incentive compensation. Cross-layer strategy coupling includes: utilities influence consumers' feasible domains through decision variables, and consumers react to utilities' demand response flexibility cost items through demand response incentive compensation, forming a closed-loop game link.
3. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 2 is characterized in that: The step S2 of completing the partially observable Markov game modeling includes the following sub-steps: S201. Define a POMG tuple, including: Based on the POMG theory, consider the utility company and each consumer as POMG agents. The POMG tuple is expressed as: ; Where, is the state space; is the set of observations of the utility company; is the consumer's observation set; It is the utility company’s action set; is the action set of the i-th consumer; is the utility company’s reward function, is the consumer’s reward function, is the collection of transfer nuclei; S202. Constructing a value function system; including: given strategy , the state value functions of the utility company as the leader and the consumer as the follower are defined as: ; where the random strategy of the leading utility company is is a set of probability distributions of actions given an observation; the random joint strategy of the follower is defined as ,in, It is the centralized element of utility company observation; It is the concentrated element of consumer observation; is the action set element of UC; the utility company is the reward function of the agent , the reward function of the consumer as an intelligent agent ,in, The weight ratio is set according to the type of consumer consumption level under weather factors. The weight ratios of consumers with different consumption level types under weather factors are different. S203. Define the advantage function. The advantage function formula for the utility company as a leader and the consumer as a follower is: Where, is an element of the state space set; It is the utility company’s observation concentration element; It is the concentrated element of consumer observation; It is the action-focused element of UC; It is the action-focused element that unites consumers; S204. Establishing an equilibrium solution mechanism; Strategies for utility companies , the consumer's Nash equilibrium is a joint strategy , for each utility company's strategy , the consumer's best response strategy is a given utility company strategy The Nash equilibrium of consumers is: ; Assuming that consumers always adopt The utility company's strategy to maximize the value function is ,Right now: ;Utility company policies via asynchronous policy updates and consumer strategies Interactive iteration ultimately satisfies: 。 4. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 3 is characterized in that: The design of the asynchronous multi-agent reinforcement learning algorithm in the Stackelberg phase described in step S3 includes the following sub-steps: S301-1. Define the multi-agent network architecture; Adopting an improved multi-agent network structure, including: the actor network input is the sum of all consumer observations , the evaluation network output is the median value of the consumer's joint action ; S301-2. Design an asynchronous gradient update mechanism; implement the Stackelberg game through dual-time-scale asynchronous updates; update the parameters of the utility company's policy network and the consumer's policy network as follows: Where, 、 Both represent the learning rate for parameter update, satisfying To ensure leadership priority; represents the total derivative of the utility company's policy loss function, Represents the partial derivative of the consumer strategy loss function; where the corresponding game stage type is matched according to the current Stackelberg stage Value range and The value range is different, and the current Stackelberg stage corresponds to the type of game stage in which it is located. Value range and Value range; the current Stackelberg stage corresponds to the game stage type including the early game, mid-game and game near equilibrium period, which is determined based on the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium; S301-3. Design the policy network loss function. The total derivative of the utility company's policy loss function represents the utility company's optimization strategy based on the best response of consumers. Based on the Stackelberg multi-agent algorithm, the policy network loss function is defined as follows: Where, Indicates the expected value, Indicates the clipping interval; clip is a truncation function. When the importance sampling exceeds the specified upper or lower limit, the corresponding upper or lower limit is returned. and represents importance sampling; the loss function of the utility company and consumer judgment network is defined as: ;in, and denote the outputs of the utility company evaluation network and the consumer evaluation network, respectively, and denote the outputs of the old UC evaluation network and the old consumer evaluation network respectively; the learning dynamics of the evaluation network is defined as: ;in, and Represents the learning rate for evaluating network parameter updates.
5. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 3 is characterized in that: The design of the asynchronous multi-agent reinforcement learning algorithm in step S3 includes the following sub-steps in the Nash phase: S302. Considering that consumers are independent and self-interested individuals, individuals tend to maximize their own utility. ; The loss function of the policy network and the judgment network is expressed as: , ; The parameter update method is as follows: ;in, represents the old policy network decoder parameters, Represents the old judgment network parameters; the consumer policy network decoder and the judgment network are trained based on the multi-agent reinforcement learning algorithm; in the execution phase, the consumer's policy network It is constructed as a connection between an encoder and a decoder; wherein, the type of game stage is matched according to the current Nash stage 、 The value range is different, and the type of game stage in which the current Nash stage corresponds to is different. 、 Value range; the current Nash stage corresponds to the game stage type including the early game, mid-game and game close to equilibrium period, which is specifically determined based on the real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium.
6. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 4 is characterized in that: The implementation of centralized training and priority experience replay in step S4 includes the following sub-steps: S401. Neural network structure design, including: The utility company strategy network has 5 layers, including: 3 parallel multi-layer perceptrons, 1 embedding layer, 2 gated recurrent unit layers and 2 fully connected layers. The input of the utility company strategy network is , the output is ; The consumer policy network consists of an encoder and a decoder; the encoder of the consumer policy network contains 2 parallel MLPs, 1 embedding layer, 1 GRU layer and 2 fully connected layers, and the encoder of the consumer policy network inputs the sum of the observation states , output the sum of consumer actions The decoder consists of 3 parallel MLPs, 1 embedding layer, 1 GRU layer, and 2 fully connected layers. The decoder input is the observation state and , output consumer actions sum; The judgment network of utility companies and consumers has 5 layers, including: 1 security mask layer, 2 GRU layers and 2 fully connected layers; the input of the utility company judgment network is the time The state of the Stackelberg game environment , whose parameters are updated based on the loss function; In the Stackelberg game environment, the encoders of the utility company's policy network and the consumer's policy network are updated based on the set parameter update formula; in the Nash game environment, consumers update the decoder parameters based on the multi-agent reinforcement learning algorithm by sharing the encoder parameters; S402, constraint processing, including: using safe mask actions to filter invalid actions; setting the following to ensure that the policy networks of different agents have the same dimensional input: Where, Replacing the advantage function that violates the constraint with ; The safety mask layer interrupts the update process of invalid actions during training. During training, the safety mask action prevents the consumer action from exceeding the feasible domain; Considering that consumers may have different action spaces, in order to ensure the scalability of the policy network when parameters are shared, the action mask is applied so that the action space not included in the consumer is not updated during training; In the loss function, Compute actions to block consumers from policy updates that do not include, An indicator representing the consumer's action space; S403, data processing specification: the original load power of the consumer satisfies the following relationship: ;in, represents the raw load data from the load dataset, Indicates the compensation amount for the electricity purchased by the consumer; the consumer demand response power satisfies the following relationship: ;in, For the Maximum demand response power of each consumer; load data , Maximum demand response power , distributed energy generation and real-time market electricity purchase prices is used as the environment state; considering that market electricity charges are generally settled monthly, an asynchronous multi-agent reinforcement learning algorithm is set up to use monthly load data for training; S404, Prioritized Experience Replay; PER is estimated based on the advantage-corrected state value, and the probability of sampling a specific strategy trajectory is expressed as: ;in, and Represent the utility company's strategic trajectory and consumer strategy trajectory The playback buffer; and Indicates the priority, 、 Is the strategy trajectory and The priority of the association, Indicates the action advantage and Ranking of trajectories when sorting; S405. Convergence condition determination: Use the asynchronous multi-agent reinforcement learning algorithm to calculate the KL divergence of the original strategy and the updated strategy to determine whether equilibrium has been reached. When all updated strategies meet the following formula for N consecutive rounds compared with the old strategies in the strategy pool, the convergence condition is met: ; Where N is a positive integer.
7. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 1 is characterized in that: The real-time monitoring of the convergence speed and stability index of the Stackelberg-Nash equilibrium and the dynamic adjustment of the strategy update frequency and step size of the utility company and the consumer include: For utilities, the first strategy-related parameters are monitored and obtained after each strategy update, including: the rate of change in grid revenue, the adjustment range of energy storage system charging and discharging power, and the impact of time-of-use electricity prices and DR incentive prices on electricity sales. For consumers, the second strategy-related parameters are monitored and obtained after each strategy update, including: the fluctuation range of electricity costs, the adjustment range of electricity load curves, the ratio of distributed energy resources to grid power purchases, and the rate of change of DR response volume. By calculating the variance or standard deviation of any parameter indicator of the first strategy indicator and the second strategy indicator in a preset number of consecutive game rounds, it is determined whether the variance or standard deviation of any parameter indicator is greater than the preset variance or preset standard deviation of the corresponding parameter indicator. If it is greater, it is determined that the equilibrium point fluctuates greatly and the current stability level is low. Otherwise, it is determined that the equilibrium point fluctuates slightly and the current stability level is good. By calculating the difference between any parameter indicator of the first strategy indicator and the second strategy indicator in adjacent game rounds, it is determined whether the difference between any parameter indicator is greater than the preset difference of the corresponding parameter indicator; if so, it is determined that the convergence speed is fast, otherwise it is determined that the convergence speed is average; The current game stage is determined based on the judged convergence speed and stability level, and different game stages are equipped with corresponding convergence speeds and stability levels.
8. The two-layer non-cooperative demand response optimization method based on asynchronous multi-agent reinforcement learning according to claim 6 is characterized in that: The neural network structure design in step S401 further includes: In the process of implementing centralized training and prioritized experience replay to optimize the efficiency of the asynchronous multi-agent reinforcement learning algorithm, a transfer learning mechanism is introduced to determine the initial parameters of the utility company and consumer strategy network, including: Historical demand response data are collected based on different regions, different seasons, and different combinations of electricity peak and valley patterns. When training new scenarios, the utility company and consumer strategy network encoder parameters obtained from historical scenarios whose similarity to the new scenario is greater than the preset similarity are used as initialization parameters.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 8.
10. A computer device, characterized in that: The computer device includes a memory, a processor, and a program stored and executable on the memory, and when the program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Voltage reactive power optimization method based on double-layer reinforcement learning power grid-user cooperation
CN115313407A
Power distribution network-microgrid group master-slave game optimization scheduling method based on multi-agent reinforcement learning algorithm
CN118611067A