A power grid voltage regulation method based on multi-agent policy evolution

By using a multi-agent policy evolution method, the state and actions of building agents are defined, and a multi-agent Markov decision process is established. This solves the problems of model dependence and long computation time in traditional power grid voltage control algorithms, and realizes rapid power grid voltage regulation.

CN117239761BActive Publication Date: 2026-05-29ZHEJIANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2023-02-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional grid voltage control algorithms rely on complete grid models, which cannot effectively cope with the uncertainties of renewable energy penetration and human behavior. They also have long computation times and cannot solve the curse of dimensionality and the exploration-utilization balance problem in the decision-making process.

Method used

A multi-agent policy evolution method is adopted to define the state and actions of building agents, establish a multi-agent partial observation Markov decision process, train building agents to regulate power grid voltage, and update the policy using gradient descent.

Benefits of technology

It enables rapid voltage regulation in complex power grid environments and can provide real-time policy actions for each agent, making it suitable for industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117239761B_ABST
    Figure CN117239761B_ABST
Patent Text Reader

Abstract

The application discloses a kind of power grid voltage regulation method based on multi-agent strategy evolution, comprising: defining the state of each building agent, the action that can be taken and the change of own energy storage state after action;Establish multi-agent partial observation Markov decision process with building agent as PQ type load;Building agent is mapped to multi-agent strategy evolution process for power grid environment change and decision process, to carry out power grid voltage regulation;In multi-agent strategy evolution process, including the training of building agent: building agent takes action, obtains the state value of power grid through power flow constraint, and gives the reward of each building agent according to state value, forms a experience trajectory;According to the preset experience pool threshold, experience trajectory is divided, and strategy is updated.The application can give the strategy action of each agent in real time for complex source network load storage scene, with the characteristics of fast operation speed, and can be applied to industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-agent reinforcement learning in the field of power electronic control technology, and specifically relates to a grid voltage regulation method based on multi-agent policy evolution. Background Technology

[0002] With the surge in renewable energy sources such as solar, hydro, and wind power, the widespread adoption of rooftop photovoltaics, and the extensive use of HVAC systems, the large-scale use of power electronic inverters has impacted power system voltage stability. Currently, traditional control algorithms, such as optimal power flow, suffer from two main problems: 1. Most planning algorithms rely on comprehensive domain knowledge and accurate grid models. When faced with the large-scale penetration of renewable energy and unpredictable human user behavior, these methods may fail. 2. Faced with models containing complex optimization objectives and constraints, traditional control algorithms require time to solve for every moment of grid operation, resulting in excessive computational overhead.

[0003] With the rapid development of multi-agent reinforcement learning algorithms, new methods have been provided for voltage control under partial observation conditions in power grids. Currently, power grid environments still only have single-agent models or models that only support training a few agents, and many algorithms are based on single-agent policy gradients, which cannot solve the exploration-exploitation balance problem caused by the curse of dimensionality and the non-stationarity of the decision-making process. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a power grid voltage regulation method and apparatus based on multi-agent strategy evolution.

[0005] To achieve the above-mentioned technical objectives, the technical solution of the present invention is as follows: A first aspect of the present invention provides a grid voltage regulation method based on multi-agent strategy evolution, comprising:

[0006] Define the state of each building agent, the actions it can take, and the changes in its own energy storage state after taking an action;

[0007] A multi-agent partially observed Markov decision process with building agents as PQ-type loads is established, and the building agents' decision-making process in response to changes in the power grid environment is mapped into a multi-agent policy evolution process, thereby regulating the power grid voltage.

[0008] The multi-agent policy evolution process includes training the building agents: the building agents take actions, obtain the state value of the power grid through power flow constraints, and give each building agent a reward based on the state value, forming an experience trajectory; the experience trajectory is divided according to the preset experience pool threshold, and the similarity between the current policy and the target policy is calculated based on KL divergence, and the policy is updated by gradient descent.

[0009] A second aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-described grid voltage regulation method based on multi-agent strategy evolution.

[0010] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-described grid voltage regulation method based on multi-agent strategy evolution.

[0011] Compared with existing technologies, the advantages of this invention are as follows: This invention provides a grid voltage regulation method based on multi-agent policy evolution. By training building agents, a multi-agent partially observed Markov decision process is established with the building agents as PQ-type loads. The building agents' decisions in response to grid environment changes are mapped to a multi-agent policy evolution process for grid voltage regulation. This invention can provide real-time policy actions for each agent in complex source-grid-load-storage scenarios, has a high computation speed, and can be widely applied in industrial scenarios. Attached Figure Description

[0012] Figure 1 This is a training flowchart of the present invention;

[0013] Figure 2 This is a diagram illustrating the strategy update;

[0014] Figure 3 This is a schematic diagram of an electronic device provided by the present invention. Detailed Implementation

[0015] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0016] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0017] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0018] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.

[0019] See Figure 1 This invention proposes a grid voltage regulation method based on multi-agent strategy evolution, the method specifically including the following steps:

[0020] Step S1: Define the state of each building agent, the actions it can take, and the changes in its own energy storage state after taking an action.

[0021] Configure the status of each building system agent, including the current timestamp, indoor temperature, constant load power consumption, and air conditioning system power consumption.

[0022] Specifically, as an implementation example, the climate zones of each building system intelligent agent are set, including Hohhot, Yan'an, Wenzhou, Guangzhou, Xishuangbanna, and Wuhan. The temperature changes of these climate zones in 2017-2018 and the corresponding changes in temperature demand of commercial markets are selected. The data is sampled every 15 minutes, and the missing data is supplemented by linear interpolation. A total of 8760 data points are collected, including the timestamp, indoor temperature, constant load power consumption, and air conditioning system power consumption.

[0023] Each building system intelligent agent includes a photovoltaic module, a battery module, a thermal energy storage module, and an air conditioning system. The actions taken by the intelligent agent are set, and the changes in the energy storage state after the actions are taken are obtained.

[0024] Photovoltaic modules are used to convert solar energy into electrical energy. The active power generation of the photovoltaic module is given by simulation based on real data. In this example, the annual power generation is given using PVLib based on the local irradiance of the climate zone, and the reactive power generation q is calculated. k PV It is determined by the following formula:

[0025]

[0026] Among them, a k The maximum reactive power ratio of the inverter output, s kmax For inverter capacity, p k PV To simulate a given active power generation using PVLib simulation.

[0027] Batteries are used to store electrical energy for use in air conditioning systems to balance the grid voltage. Their maximum stored energy decreases slowly with charging and discharging.

[0028] The formula for calculating the capacity of a storage battery is as follows:

[0029]

[0030] Among them, C new For the new capacity, c loss Here, |E| represents the attenuation coefficient, C is the old capacity, C0 is the original capacity, and |E| represents the attenuation coefficient. in|out | represents the absolute value of charging and discharging.

[0031] The energy storage module includes a cold water tank and a hot water tank, which are used for temperature regulation. The state of charge formula for the energy storage module is as follows:

[0032] SOC t+1 =SOC t +max{min{a t ·C,Q tmax -Q b},-Q dem}

[0033] Among them, SOC t+1 For the next charging state, SOC t a represents the current charging state. t Q represents the percentage of thermal energy storage output, where C is the maximum capacity and Q is the maximum capacity. tmax Maximum output of the air conditioning system, Q dem The heat energy required by the building system.

[0034] The air conditioning system, used to regulate temperature, activates when the energy storage module cannot meet the building's temperature requirements, supplementing the insufficient temperature. The air conditioning system's operating efficiency is first defined based on the indoor-outdoor temperature difference.

[0035]

[0036]

[0037] Among them, COP c It refers to the air conditioner's electrothermal conversion cooling efficiency, COP. h It refers to the heating efficiency of an air conditioner through electrothermal conversion, η. tech T is the technical efficiency coefficient (usually taken as 0.2 or 0.3). ctg For the target cooling temperature, T h tg For the target heating temperature, T outdoor This refers to the outdoor temperature.

[0038] Then define the air conditioner output formula:

[0039] Q t+1 =C·(SOC) t+1 -SOC t )+Q dem

[0040] Among them, Q t+1 The SOC (State of Charge) is the air conditioning output value for the next moment. t+1 For the next charging state of the thermal energy storage module, SOC t Q represents the current charging state of the thermal energy storage module. dem Let Q represent the building system's temperature requirement (i.e., the target temperature). This formula reflects that when the energy storage module is insufficient to meet the building system's temperature requirement, the shortfall is provided by the air conditioning system. In this case, Q... dem The temperature requirements of buildings must be met.

[0041] Finally, the relationship between the actual heat energy converted by the air conditioner and the air conditioner output is defined as E. t :

[0042]

[0043] In the formula, COP t This refers to the actual heat conversion efficiency of the air conditioner.

[0044] Thus, by completing the construction of the building system electrothermal coupling model, four actions of the intelligent agent can be defined: the percentage of reactive power output by the photovoltaic module, the percentage of charge and discharge of the battery module, and the percentage of thermal energy stored by the energy storage module (divided into two actions: power output from the hot water storage tank and power output from the cold water storage tank).

[0045] Step S2: Establish a multi-agent partially observed Markov decision process with building agents as PQ-type loads, and map the building agents' decision-making process in response to changes in the power grid environment into a multi-agent strategy evolution process, thereby regulating the power grid voltage.

[0046] Specifically, as an implementation example, the buildings are first divided into different types based on their daily indoor temperature requirements. The difference between indoor and outdoor temperatures determines the varying air conditioning output efficiency of the building's intelligent agents, thus defining different agent types. Secondly, the selected power grid topology is determined. In this example, the IEEE-33 power grid topology is used as the basis. This grid has a rated voltage of 12.66kV and 32 nodes excluding voltage sources. The maximum total load of all nodes is 5084.26kW and 2547.32kVar, respectively. Simultaneously, the connections 17-18, 21-25, and 32-22 are disconnected to prevent grid loops and ensure the stability of power flow calculations. The maximum load that each node can drive is determined based on the power grid topology. In this example, each power grid node is configured to have 6 intelligent agents connected to it, randomly assigned 10 types of intelligent agents, and the target demand for the intelligent agents is scaled proportionally to the number of agents and the maximum load of the nodes.

[0047] Calculate the reachability matrix based on the adjacency matrix of the power grid topology:

[0048] tmpA = A 0 +A 1 +LA n-1 +A n

[0049]

[0050] Here, n is the number of neighboring nodes that each node can observe. This defines a partial observation range for each node. The agent under this node can only obtain the voltage and active and reactive power values ​​of the nodes that can be observed by this node, and cannot know the global information.

[0051] Based on the thermal energy storage, electrical energy storage, photovoltaic active power generation, and target demand of each intelligent agent, and combined with the air conditioning output formula, the active power demand of each intelligent agent can be calculated. The reactive power generation of the intelligent agent can be calculated based on the output of its photovoltaic grid-connected inverter. Therefore, the intelligent agent that has completed its actions can be defined as a PQ-type load, according to grid power flow constraints:

[0052]

[0053]

[0054] Where, p k pv p represents the active power generation from photovoltaic power. k L q represents the sum of the building's air conditioning system and constant load. L k v is the total reactive power of the load. k For node voltage, g kj For mutual conductance, b kjFor mutual impedance, θ jk The phase difference is used to determine the voltage, active power, and reactive power values ​​of the local node and neighboring nodes. Based on the deep learning strategy, each agent provides actions for photovoltaic, thermal energy storage, and electrical energy storage. The air conditioning system provides the required power value based on the actions and demand. After all agents have completed their actions, the power flow calculation formula updates the state value of the entire power grid. At the same time, a reward is given to each agent based on the state, thus completing an empirical trajectory.

[0055] The multi-agent policy evolution process includes training the building agents: the building agents take actions, obtain the state value of the power grid through power flow constraints, and give each building agent a reward based on the state value, forming an experience trajectory; the experience trajectory is divided according to the preset experience pool threshold, and the similarity between the current policy and the target policy is calculated based on KL divergence, and the policy is updated by gradient descent.

[0056] Specifically, the following steps are included:

[0057] Step S201: Determine the climate zone data to be used for training the building agent, initialize network parameters and experience pool, determine the number of training rounds, and start training.

[0058] Specifically, as an implementation example, the climate zone data used to train the model is selected. This data is a function of outdoor temperature changing over time, which, together with the agent's indoor temperature, determines the efficiency of the air conditioning system. Based on the reinforcement learning framework, each agent has a corresponding evaluation network and a policy network. Thus, the policy gradient can be expressed as:

[0059]

[0060] Wherein, J(θ) i The estimated value is the agent's response to policy π. i Expected reward, π i Let Q be the policy network for the i-th agent, where Q is the probability of the agent adopting a given policy given its observations. i π To provide a fitted estimate of the true reward for the i-th agent, given the current global state and the actions of each agent, output the expected reward of the agent under the current state and current action. First, initialize π. i and Q i π Parameters of two neural networks. Among them, the evaluation network Q... i π A multilayer state perceptron (MLP) model is selected, which has three linear layers with weight dimensions of 1182×128, 128×64, and 64×1, respectively. The ReLU activation function is used for the intermediate layers. The policy network π... iAn RNN network is used to record previous experience. This RNN network contains an input linear layer, a GRU neuron, and an output linear layer. The dimension of the input linear layer is 1182×64, the dimension of the GRU neuron is 64×64, and the dimension of the output linear layer is 64×384. The activation function from the input linear layer to the GRU neuron is the ReLU function, and the activation function from the input linear layer to the output linear layer is the sigmoid function, which maps the action space to the range [-1,1].

[0061] Then initialize the experience pool, divide the experience pool into good and bad categories, and set the experience pool threshold to -1000.

[0062] Finally, the climate zone data and agent data are read in, the number of training rounds is determined to be 384, and training begins.

[0063] In step S202, the agent uses the model to determine the action under the current observation, the reward calculated from the power grid environment, and the next state to form an empirical trajectory.

[0064] Specifically, each agent uses the voltage, active power, and reactive power values ​​of its own node and neighboring nodes as observation values. i Based on the deep learning strategy, the actions of photovoltaic, thermal energy storage, and electrical energy storage are given. The air conditioning system gives the required power value based on the actions and demand. All agents follow the strategy π. i Give the corresponding action a i After each agent completes its action, the power flow calculation formula updates the state value of the entire power grid at the next moment and the individual observation values ​​of each agent. i At the same time, a reward s is given to each agent based on its state. i , forming an empirical trajectory (o i ,a i ,r i ,o i ′).

[0065] Step S203: Based on the reward of the current experience trajectory, determine whether the experience trajectory is divided into a good trajectory pool or a bad trajectory pool.

[0066] Specifically, calculate the average reward from the most recent experiences. Then compare the new member experience reward R i and The size, if new experience rewards R i Greater than average reward If the experience trajectory is not good, it will be added to the good experience pool; otherwise, it will be added to the bad experience pool.

[0067] Step S204: Calculate the similarity between the current policy and the good and bad policies based on KL divergence, and update the policy using gradient descent.

[0068] Specifically, a policy π is estimated using either maximum likelihood estimation or a neural network fitting method on the existing good and bad experience pools. P (a i |o i ) and π N (a i |o i ), calculate the KL divergence respectively:

[0069]

[0070] Here, π represents the current agent's policy. This method updates the policy by minimizing the similarity between the current policy and a good policy while maximizing the similarity between the current policy and a bad policy.

[0071] like Figure 3 As shown, this application provides an electronic device including a memory 101 for storing one or more programs and a processor 102. When the one or more programs are executed by the processor 102, they implement the method as described in any of the first aspects above.

[0072] The system also includes a communication interface 103. The memory 101, processor 102, and communication interface 103 are electrically connected directly or indirectly to each other to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.

[0073] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0074] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor 102, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0075] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can also be implemented in other ways. The method and system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0076] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0077] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by processor 102, the computer program implements the methods described in any of the first aspects above. If the functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory 101 (ROM), random access memory 101 (RAM), magnetic disks, or optical disks.

[0078] In summary, this invention provides a grid voltage regulation method based on multi-agent policy evolution. By training building agents, a multi-agent partially observed Markov decision process is established, with the building agents as PQ-type loads. The building agents' decisions in response to grid environment changes are mapped to a multi-agent policy evolution process for grid voltage regulation. This invention can provide real-time policy actions for each agent in complex source-grid-load-storage scenarios, features high computation speed, and can be widely applied in industrial scenarios.

[0079] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only.

[0080] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A grid voltage regulation method based on multi-agent strategy evolution, characterized in that, include: Define the state of each building agent, the actions it can take, and the changes in its own energy storage state after taking an action; A multi-agent partially observed Markov decision process is established with building agents as PQ-type loads. The decision-making process of building agents in response to changes in the power grid environment is mapped to a multi-agent policy evolution process for power grid voltage regulation; this includes: First, the buildings are divided into different types based on their different indoor temperature requirements throughout the day. Second, the selected power grid topology is determined. Based on the power grid structure, the maximum load that each node can drive is determined. Then, the number of agents selected under each power grid node is specified, and different types of agents are randomly assigned. The target demand of the agents is scaled proportionally to the number of agents and the maximum load of the node. Based on the thermal energy storage, electrical energy storage, photovoltaic active power generation and target demand of each intelligent agent, the active power demand of each intelligent agent is calculated in combination with the air conditioning output, and the reactive power generation of the intelligent agent is calculated based on the output of the photovoltaic grid-connected inverter. Thus, the intelligent agent that has completed the action is defined as a PQ type load. Each agent observes the voltage, active power, and reactive power values ​​of its own node and neighboring nodes, and gives actions for photovoltaic, thermal energy storage, and electrical energy storage according to the strategy. The air conditioning system gives the required power value according to the actions and target requirements, thereby regulating the grid voltage. The multi-agent policy evolution process includes training the building agents: the building agents take actions, obtain the state value of the power grid through power flow constraints, and give each building agent a reward based on the state value, forming an experience trajectory; the experience trajectory is divided according to the preset experience pool threshold, and the similarity between the current policy and the target policy is calculated based on KL divergence, and the policy is updated by gradient descent.

2. The power grid voltage regulation method based on multi-agent strategy evolution according to claim 1, characterized in that, The state of each building agent is defined as follows: the timestamp of each building agent, indoor temperature, constant load power consumption, and air conditioning system power consumption; Each building system intelligent agent includes a photovoltaic module, a battery module, a thermal energy storage module, and an air conditioning system; the actions taken by each building intelligent agent include: the percentage of reactive power output by the photovoltaic module, the percentage of charge and discharge of the battery module, and the percentage of thermal energy stored by the energy storage module; among them, the percentage of thermal energy stored by the energy storage module is divided into two actions: the output of the thermal storage tank and the output of the cold storage tank.

3. The power grid voltage regulation method based on multi-agent strategy evolution according to claim 2, characterized in that, Photovoltaic modules are used to convert solar energy into electrical energy. The reactive power output of a photovoltaic module is determined by the following formula: ; in, This represents the maximum reactive power ratio of the inverter output. For inverter capacity, Active power generation; Storage batteries are used to store electrical energy for use in air conditioning systems. The formula for calculating the capacity of a storage battery is as follows: ; in, For the new capacity, The attenuation coefficient is... For old capacity, For the original capacity, The absolute values ​​of charging and discharging; The energy storage module includes a cold water storage tank and a hot water storage tank, and the state of charge formula is as follows: ; in, To determine the charging state for the next moment. The current charging status. Percentage of thermal energy storage output. For maximum capacity, Maximum output of the air conditioning system The heat energy required by the building system; An air conditioning system is used to achieve the target temperature when the energy storage module cannot meet the building's temperature requirements. The air conditioning system first defines its operating efficiency based on the temperature difference between indoors and outdoors: ; Among them, among them, It refers to the air conditioner's electrothermal conversion cooling efficiency. It refers to the heating efficiency of the air conditioner's electrothermal conversion. For efficiency coefficient, For the target cooling temperature, For the target heating temperature, Outdoor temperature; Then define the air conditioner output formula: ; in, This is the air conditioning output value for the next moment. This indicates the charging state of the thermal energy storage module at the next moment. This indicates the current charging status of the thermal energy storage module. The target temperature for the building; Finally, the relationship between the actual heat energy converted by the air conditioner and the air conditioner output is defined: 。 4. The power grid voltage regulation method based on multi-agent strategy evolution according to claim 1, characterized in that, The formula for calculating power flow constraints in a power grid is as follows: ; in, This refers to the active power generation from photovoltaic power. This is the sum of the building's air conditioning system and constant load. This represents the total reactive power of the load. For node voltage, , These are mutual conductance and mutual impedance, respectively. The phase difference is used to determine the voltage, active power, and reactive power values ​​of the local node and its neighboring nodes.

5. The power grid voltage regulation method based on multi-agent strategy evolution according to claim 1, characterized in that, The training of the building agent includes: each building agent includes a corresponding evaluation network and policy network, initializing network parameters, initializing the experience pool, setting the experience pool threshold, dividing the experience pool into a good experience pool and a bad experience pool, using climate zone data as training data, and training the building agent. Each agent uses the voltage, active power, and reactive power values ​​of its own node and neighboring nodes as observations. Based on the strategy, the photovoltaic, thermal energy storage, and electrical energy storage systems provide their actions, and the air conditioning system provides the required power values ​​based on the actions and demands; all intelligent agents follow the strategy. Give the corresponding action After each agent completes its action, the power flow calculation formula updates the state value of the entire power grid at the next moment and the individual observation values ​​of each agent. At the same time, a reward is given to each agent based on its state. This forms an empirical trajectory. ; Calculate the average reward Then compare the experience rewards for new members. and Size, if new experience rewards are added Greater than average reward Then add the experience trajectory to the good experience pool, and vice versa. The similarity between the current policy and the good and bad policies is calculated based on KL divergence, and the policy is updated using gradient descent.

6. The power grid voltage regulation method based on multi-agent strategy evolution according to claim 5, characterized in that, Calculating the similarity between the current policy and good and bad policies based on KL divergence specifically includes: Estimate a policy using maximum likelihood estimation or neural network fitting for both the good and bad experience pools, respectively. and Calculate the KL divergence separately, using the following formula: ; in, This represents the current agent's strategy.

7. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the power grid voltage regulation method based on multi-agent strategy evolution as described in any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the grid voltage regulation method based on multi-agent policy evolution as described in any one of claims 1-6.