Reactive voltage regulation method and device based on large model and multi-agent reinforcement learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明提供了一种基于大模型和多智能体强化学习的无功调压方法及装置,用以解决如何对高比例新能源接入的配电网进行无功调压的技术问题
本发明的基于大模型和多智能体强化学习的无功调压方法中,混合奖励函数采用全局奖励+局部奖励的协同结构,有效破解了多智能体强化学习中的信用分配难题;其中,局部奖励为各智能体提供了明确的安全边界与贡献度量,全局奖励则引导各智能体关注系统整体的经济运行目标;这种差异化设计使智能体在策略更新时能够精准平衡局部电压安全与全局网损优化,避免了单一奖励结构下安全性与经济性目标的相互制约。奖励函数的动态更新机制将LLM对奖励函数的调整嵌入至强化学习的训练闭环内,从根本上改变了传统闭环外调整需经历多个完整训练周期的模式,减少了试错迭代带来的时间成本,显著提升了智能体学习效率。本发明方法实现了从人工设计奖励函数向大模型辅助生成与优化奖励函数的跨越,降低了奖励函数设计门槛,避免了繁琐的反复调参与验证过程,节约了高昂的时间成本与人工成本。
Smart Images

Figure CN122338834B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electrical automation technology, and in particular to a reactive power voltage regulation method and device based on large model and multi-agent reinforcement learning. Background Technology
[0002] As the penetration rate of distributed renewable energy sources, represented by photovoltaics, continues to increase in distribution networks, the power system operation mode is shifting from centralized control and unidirectional power flow to a new operating mode of distributed coordinated control and bidirectional power flow. Due to the intermittent and random nature of distributed photovoltaic power output, its large-scale integration can easily lead to increased fluctuations in the injected power at distribution network nodes; at the same time, the bidirectional power flow characteristics introduced by its grid connection can easily lead to voltage rise at local nodes, exacerbating the risk of voltage exceeding limits.
[0003] Traditional methods for voltage control in distribution networks often employ reactive power regulation strategies based on mathematical models, achieving coordinated control of reactive power resources by solving deterministic models. However, these methods are highly dependent on the accuracy of system models and parameters. With the increasing penetration of distributed power sources and rapid changes in operating conditions, they suffer from limitations in computational efficiency and real-time response, making it difficult to meet the real-time voltage control requirements under complex operating environments. In recent years, deep reinforcement learning, with its advantages of model-free operation and autonomous interactive learning, has been gradually introduced into the field of reactive power regulation in distribution networks. However, existing single-agent modeling methods are prone to bottlenecks such as state-action space expansion and increased computational complexity when the number of controlled objects increases. To overcome this limitation, distributed reactive power cooperative control based on MADRL (Multi-Agent DRL) is considered a better solution. The performance ceiling of multi-agent systems is largely constrained by the design quality of the reward function. However, the design of reward functions in existing MADRL systems typically relies on expert experience and manual trial and error, resulting in long development cycles, high costs, and difficulties in credit allocation among multiple agents. In reward function optimization, existing hyperparameter optimization methods are mostly limited to numerical optimization within a pre-defined reward function framework. In reactive power regulation scenarios in distribution networks, fixed-structure reward functions often struggle to simultaneously satisfy voltage safety and system economic objectives. Large Language Models (LLMs), with their powerful semantic understanding, logical reasoning, and contextual learning capabilities, have made rapid progress in the field of artificial intelligence. LLMs can not only optimize the parameters of reward functions but also dynamically reconstruct their mathematical structure. Although research using LLMs to design reward functions has made initial progress in several fields, related research in the field of reactive power regulation in distribution networks remains relatively scarce. Furthermore, existing research generally employs a single reward structure, and adjustments to the reward function typically occur outside the training loop, significantly limiting learning efficiency and final performance.
[0004] Therefore, a new technical solution is urgently needed to address the technical problem of reactive power regulation in distribution networks with a high proportion of renewable energy access. Summary of the Invention
[0005] This invention provides a reactive power voltage regulation method and device based on large model and multi-agent reinforcement learning, which solves the technical problem of how to perform reactive power voltage regulation on distribution networks with a high proportion of renewable energy access.
[0006] To achieve the above objectives, this invention provides a reactive power regulation method based on large model and multi-agent reinforcement learning, comprising: The photovoltaic controller is used as an intelligent agent; a reward function is generated for each intelligent agent based on the large language model; the reward function includes a network loss penalty value calculation term and a voltage limit violation penalty average value calculation term for nodes under the target intelligent agent; A reactive power voltage regulation model is constructed based on the distribution network operation constraints with the goal of minimizing network loss. A Markov decision process is constructed based on the reactive power voltage regulation model, the agent, and the reward function. Iterative training is performed based on the MAAC algorithm and the Markov decision process. After each preset number of iterations, a preset criterion is determined. If the criterion is satisfied, the reward functions are updated based on the large language model. After the iteration is completed, each intelligent agent controls the photovoltaic inverters within its respective area.
[0007] Preferably, the reward functions generated for each agent based on the large language model include: The large language model generates network loss penalty calculation terms and voltage limit violation average penalty calculation terms based on the agent and the pre-acquired distribution network topology, and generates a reward function based on the network loss penalty calculation terms and voltage limit violation average penalty calculation terms. The network loss penalty calculation term includes a network loss coefficient; the voltage limit violation average penalty calculation term includes a voltage penalty function, and the voltage penalty function adopts a symmetrical structure for upper and lower limits violation.
[0008] Preferably, the reward function includes: For the reward function of agent g , is represented as: ; ; ; in, For local reward terms, it represents the average voltage over-limit penalty of all nodes within the control area of agent g; This is a global reward item, representing the network loss penalty value; Represents a set of intelligent agents; This represents the total number of nodes within the control area of agent g; It is a voltage penalty function; For nodes The measured voltage amplitude; This is the voltage reference value; This is the network loss coefficient; This is due to network loss.
[0009] Preferably, after each preset number of iterations, a preset criterion is determined. If the criterion is satisfied, the reward functions are updated based on the large language model, including: For each agent, after each preset number of iterations, the mean and variance of the network voltage exceedance rate and network loss for each iteration are calculated. If the mean is higher than the preset mean standard, the variance is lower than the preset variance standard, and the iteration has not yet ended, the network loss coefficient and voltage penalty function in each reward function are updated based on the large language model, and the iterative training is restarted.
[0010] Preferably, updating the network loss coefficient and voltage penalty function in each reward function based on the large language model includes: Obtain first information; the first information includes network loss and voltage over-limit rate, first amplitude and second amplitude in the control area of each intelligent agent; the first amplitude includes the maximum voltage amplitude that exceeds the preset voltage amplitude upper limit; the second amplitude includes the minimum voltage amplitude that exceeds the preset voltage amplitude lower limit; the large language model updates the network loss coefficient and voltage penalty function according to the first information and preset update rules.
[0011] Preferred preset update rules include: Prioritize adjusting the voltage penalty function to ensure the elimination of voltage over-limit situations in each region before adjusting the network loss coefficient. In the voltage penalty function, design a gradient distribution of voltage over-limit penalty based on the voltage over-limit rate, the first amplitude, and the second amplitude. The voltage penalty function can be a linear penalty or a non-linear penalty. The adjustment of the network loss coefficient and the voltage penalty function must ensure the elimination of voltage over-limit situations and the reduction of network losses.
[0012] Preferably, the reactive power voltage regulation model constructed based on distribution network operation constraints with the objective of minimizing network losses includes: The objective function includes: ; in, For network loss; For the set of routes; For the line The resistance; and The lines are respectively The active and reactive power; The parent node of the node. For nodes The voltage amplitude; Distribution network operation constraints include: Power balance constraints: Power balance constraints in the distribution network are described using the Distflow model; Node voltage constraint: Node voltage deviation is controlled within ±5% of the rated voltage; Photovoltaic inverter operating constraints: The active power and apparent power of the photovoltaic inverter are maintained within the preset range.
[0013] Preferably, the construction of a Markov decision process based on a reactive power regulation model, an agent, and a reward function includes: Using the reactive power regulation model as the interactive environment, the reactive power regulation problem in the distribution network is described as a partially observable Markov decision process involving multiple agents, including a seven-tuple. ;in, A collection of intelligent agents, o For the joint observation space set of intelligent agents, S For a set of state spaces, For the action space set, Let be the state transition probability. R For joint awards, Discount factor; State space set ;in, and These represent the active and reactive load sets of all nodes in the network, respectively. V Represents the set of node voltage amplitudes across the entire network; and These represent the sets of active and reactive power of the photovoltaic inverters in the network, respectively. Joint observation space set of intelligent agents o In the observation space of agent g ;in, , and These represent the active load, reactive load, and voltage amplitude of all nodes within the area controlled by agent g, respectively. , Represent the active and reactive power of the photovoltaic inverter controlled by agent g, respectively; the joint observation space set of the agent. and ; Action space set In the middle, region z The set of actions of all photovoltaic inverters in the system is represented as , Indicates the region z The collection of all photovoltaic inverters in the system. Indicates photovoltaic inverter Intelligent agent actions; photovoltaic inverter Actual output reactive power With agent actions The mapping relationship is represented as ,in For photovoltaic inverters Rated apparent power, For photovoltaic inverters The active power; when The inverter absorbs reactive power from the system, and conversely, the inverter injects reactive power into the system. For state transition probability Represented as This indicates that the distribution network is in control action. Under the influence of the action, the system changes from state Evolved to The dynamic process is simulated using the Distflow power flow calculation method.
[0014] Preferably, iterative training based on the MAAC algorithm and Markov decision process includes: Based on the MAAC algorithm, an independent evaluation network and policy network are built for each agent, and a target evaluation network and a target policy network corresponding to the evaluation network and policy network are built respectively. An attention mechanism is introduced into the evaluation network of each agent. Based on the built evaluation network, policy network, target evaluation network and target policy network, centralized learning and iterative training are carried out in each agent using the joint observation space set and action space set of all agents.
[0015] The present invention also provides a reactive power voltage regulation device based on large model and multi-agent reinforcement learning, which is used in the method of the present invention. The device includes a first module, a second module, a third module, a fourth module, a fifth module and a sixth module. The first module is used to treat the photovoltaic controller as an intelligent agent; based on the large language model, a reward function is generated for each intelligent agent; the reward function includes a network loss penalty value calculation term and a voltage over-limit penalty average value calculation term for the nodes under the target intelligent agent; The second module is used to construct a reactive power voltage regulation model based on the operating constraints of the distribution network with the goal of minimizing network losses. The third module is used to construct a Markov decision process based on a reactive power regulation model, an agent, and a reward function. The fourth module is used for iterative training based on the MAAC algorithm and Markov decision process; The fifth module is used to determine the preset criteria after each preset round of iteration. If the criteria are met, the reward functions are updated based on the large language model. The sixth module is used to control the photovoltaic inverters in their respective areas based on each intelligent agent after the iteration is completed.
[0016] The present invention has the following beneficial effects: In the reactive power voltage regulation method based on large model and multi-agent reinforcement learning of this invention, the hybrid reward function adopts a collaborative structure of global reward + local reward, effectively solving the credit allocation problem in multi-agent reinforcement learning. Local rewards provide clear safety boundaries and contribution metrics for each agent, while global rewards guide each agent to focus on the overall economic operation goal of the system. This differentiated design enables agents to accurately balance local voltage safety and global network loss optimization during policy updates, avoiding the mutual constraints between safety and economic goals under a single reward structure. The dynamic update mechanism of the reward function embeds the LLM's adjustment of the reward function into the training loop of reinforcement learning, fundamentally changing the traditional mode where adjustments outside the loop require multiple complete training cycles, reducing the time cost of trial and error iteration, and significantly improving the learning efficiency of the agents. This invention achieves a leap from manually designing reward functions to large model-assisted generation and optimization of reward functions, lowering the threshold for reward function design, avoiding tedious repeated tuning and verification processes, and saving high time and labor costs.
[0017] The reactive power voltage regulation device based on large model and multi-agent reinforcement learning of the present invention, when used in the method of the present invention, has the same beneficial effects as the method of the present invention.
[0018] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0019] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart illustrating a preferred embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of the voltage over-limit rate convergence curve of the method of the present invention according to a preferred embodiment of the present invention.
[0021] Figure 3 This is a schematic diagram of the network loss convergence curve of the method of the present invention according to a preferred embodiment of the present invention.
[0022] Figure 4 This is a schematic diagram illustrating the trend of the system's highest voltage amplitude in a preferred embodiment of the present invention.
[0023] Figure 5 This is a schematic diagram illustrating the trend of the minimum voltage amplitude variation in a preferred embodiment of the present invention. Detailed Implementation
[0024] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.
[0025] See Figure 1 In a preferred embodiment of the present invention, a reactive power regulation method based on large model and multi-agent reinforcement learning is provided, comprising: F1. Treat the photovoltaic controller as an agent; generate reward functions for each agent based on the large language model; the reward function includes a network loss penalty calculation term and a voltage over-limit penalty average calculation term for the nodes under the target agent.
[0026] In a preferred embodiment of the present invention, generating reward functions for each agent based on a large language model includes: The large language model generates network loss penalty calculation terms and voltage limit violation average penalty calculation terms based on the agent and the pre-acquired distribution network topology, and generates a reward function based on the network loss penalty calculation terms and voltage limit violation average penalty calculation terms. The network loss penalty calculation term includes a network loss coefficient; the voltage limit violation average penalty calculation term includes a voltage penalty function, and the voltage penalty function adopts a symmetrical structure for upper and lower limits violation.
[0027] In a preferred embodiment of the present invention, the reward function includes: For the reward function of agent g , is represented as: ; ; ; in, For local reward terms, it represents the average voltage over-limit penalty of all nodes within the control area of agent g; This is a global reward item, representing the network loss penalty value; Represents a set of intelligent agents; This represents the total number of nodes within the control area of agent g; It is a voltage penalty function; For nodes The measured voltage amplitude; This is the voltage reference value; This is the network loss coefficient; This is due to network loss.
[0028] In a preferred embodiment of the present invention, the design of prompt words when generating the reward function through a large language model includes: 1) Expert Role Setting: You are an expert in power system operation and reinforcement learning. Please design a hybrid reward function for the multi-agent deep reinforcement learning task in reactive power regulation of power systems. 2) Reinforcement learning environment characteristics: Multiple agents are located in a 141-node power distribution network system, which is divided into nine regions. The nodes of each region are allocated and configured with photovoltaic inverters according to the actual scenario. Each region has a reinforcement learning agent to control the photovoltaic inverters in the region. There are a total of 9 agents in the environment. 3) Task Objectives and Requirements: This reactive power regulation task aims to adjust the reactive power of the photovoltaic inverter to minimize voltage exceedances and reduce distribution network losses. Please design a hybrid reward function for each agent based on this objective. 4) Reward function generation constraints: The hybrid reward for each agent consists of a global reward term and a local reward term. The global reward term is the network loss penalty, and the local reward term is the node voltage limit violation penalty within the controlled area of each agent. The voltage amplitude of each node should be within the range of [0.95, 1.05]. Voltage penalty will only be applied if the node voltage amplitude exceeds the limit. After meeting this safety requirement, the network loss will be reduced to the maximum extent. 5) Basic template for reward functions: a standardized Python code paradigm that covers function name, function parameter definition, and function return value definition.
[0029] F2. Construct a reactive power voltage regulation model based on distribution network operation constraints, with the objective of minimizing network losses. F2 specifically includes: The objective function includes: ; in, For network loss; For the set of routes; For the line The resistance; and The lines are respectively The active and reactive power; The parent node of the node. For nodes The voltage amplitude; Distribution network operation constraints include: Power balance constraints: The power balance constraints in the distribution network are described using the Distflow model, specifically including: ; ; ; in, and They are nodes The active and reactive load power; and They are nodes Active and reactive power of photovoltaic equipment; This is a set of nodes, where node 0 is the substation busbar, which is usually modeled as a balancing node. For nodes The set of child nodes. This describes the voltage magnitude relationship between two adjacent nodes. and These are the resistance and reactance of line i, respectively.
[0030] Node voltage constraints: According to the distribution network voltage safety regulations, node voltage deviations are controlled within ±5% of the rated voltage, specifically including: ; in, For nodes The measured voltage amplitude.
[0031] Photovoltaic inverter operating constraints: The active power and apparent power of the photovoltaic inverter are maintained within a preset range, specifically including: ; ; in, For nodes The rated apparent power of a photovoltaic inverter; For nodes The photovoltaic inverter represents the maximum available active power under current sunlight conditions. This is achieved by adjusting the reactive power output of the distributed photovoltaic inverter. It can support node voltage without affecting active power generation, thus achieving active voltage control.
[0032] F3. Constructing a Markov decision process based on a reactive power regulation model, an agent, and a reward function. F3 specifically includes: Using the reactive power regulation model as the interactive environment, the reactive power regulation problem in the distribution network is described as a partially observable Markov decision process involving multiple agents, including a seven-tuple. ;in, A collection of intelligent agents, o For the joint observation space set of intelligent agents, S For a set of state spaces, For the action space set, Let be the state transition probability. R For joint awards, The discount factor is used to measure the importance of future rewards relative to current rewards.
[0033] The photovoltaic controllers in each region are modeled as intelligent agents to coordinate the control of the photovoltaic inverters within that region. The set of intelligent agents is defined as follows: ,in The number of intelligent agents.
[0034] In partially observable Markov decision-making processes, the agent cannot acquire state-space information, but can only perceive local observational space information within its own region. The state-space set characterizes the operational features of the distribution network. ;in, and These represent the active and reactive load sets of all nodes in the network, respectively. V Represents the set of node voltage amplitudes across the entire network; and These represent the sets of active and reactive power of the photovoltaic inverters in the network, respectively. Joint observation space set of intelligent agents o In the observation space of agent g ;in, , and These represent the active load, reactive load, and voltage amplitude of all nodes within the area controlled by agent g, respectively. , Represent the active and reactive power of the photovoltaic inverter controlled by agent g, respectively; the joint observation space set of the agent. and ; Since the reactive power output of a photovoltaic inverter is continuously adjustable, the reactive power voltage regulation task can be modeled as a continuous motion control task. Action space set In the middle, region z The set of actions of all photovoltaic inverters in the system is represented as , Indicates the region z The collection of all photovoltaic inverters in the system. Indicates photovoltaic inverter Intelligent agent actions; photovoltaic inverter Actual output reactive power With agent actions The mapping relationship is represented as ,in For photovoltaic inverters Rated apparent power, For photovoltaic inverters The active power; when The inverter absorbs reactive power from the system, and conversely, the inverter injects reactive power into the system. For state transition probability Represented as This indicates that the distribution network is in control action. Under the influence of the action, the system changes from state Evolved to The dynamic process of the distribution network is driven by power flow equations, network topology, photovoltaic inverter output disturbances, and load changes. State updates exhibit characteristics of continuity, coupling, and uncertainty. This embodiment uses the Distflow power flow calculation method to simulate the state transition process of the distribution network.
[0035] The reward function for the agents has been described above. In a preferred embodiment of this invention, a hybrid reward function is constructed with the assistance of a large language model. This function not only includes the global objective of reducing network losses but also introduces a voltage control objective specific to the agent's territory. This provides differentiated rewards for each agent, balancing overall cooperation and individual contributions, thereby guiding them to effectively consider both network losses and voltage safety during policy generation. For partially observable Markov decision processes, the joint reward includes... .
[0036] F4 is based on the MAAC algorithm and Markov decision process for iterative training. F4 specifically includes: Based on the MAAC (Multi-Agent Attention Actor-Critic) algorithm, an independent evaluation network and policy network are built for each agent, along with a target evaluation network and a target policy network corresponding to the evaluation and policy networks, respectively. A "centralized training-distributed execution" paradigm is adopted: during the training phase, the evaluation network of each agent incorporates an attention mechanism. Based on the built evaluation network, policy network, target evaluation network, and target policy network, centralized learning and iterative training are performed in each agent using the joint observation space and action space set of all agents. During the execution phase, the agents achieve distributed autonomous decision-making based solely on local observations.
[0037] For an agent g, the evaluation network and the policy network are respectively denoted as: and The objective evaluation network and the objective policy network are respectively represented as: and ; Each evaluation network receives joint observations from all agents during the training phase. and joint actions ( (Number of agents), output the action to be taken by the target agent in the current state. Long-term cumulative return expectation In this process, the evaluation network first encodes the observation state and action information of each agent into feature vectors. Then, it uses an attention mechanism to calculate the association weights between the target agent's feature vector and the feature vectors of other agents, and aggregates them into a weighted feature vector containing global interaction information. This allows the target agent to selectively extract key information from other agents during parameter updates. Evaluation network parameters By minimizing the current estimate With target value loss function Update.
[0038] ; in, This indicates an evaluation of network parameters. The loss function; The parameters representing the agent's evaluation network; This represents an experience replay buffer, used to store quadruple experience samples generated by the interaction between the agent and the power distribution network environment. ,in This serves as the joint observation space for all agents in the next moment; Indicates the experience replay buffer Random sampling experience samples and the estimated value With target value Calculate the expected value by squared deviations between the two values.
[0039] In calculating the target value At that time, MAAC incorporates the action entropy of the target policy distribution into the long-term return, including: ; in, The target policy network generates the next action. Then, seek long-term future returns. Expectations; This is a discount factor used to balance immediate rewards with future long-term benefits; This represents the target evaluation network estimate; Temperature parameter used to balance instant rewards Weights relative to action entropy; Action entropy is used to encourage agents to actively explore the environment and prevent control strategies from converging prematurely to local optima. Indicates the local observation at the next moment. Under these circumstances, the target policy network of agent g generates actions. The probability of; This indicates that the agent g will determine its next action based on its target policy network and local observations. The generated deterministic actions; This represents the network parameters of the target policy. It is used in calculating the target value. At that time, the next action of all agents is generated by their respective target policy networks.
[0040] During the training phase, the agent policy network is based on local observations. Generate Actions It receives global value information from the evaluation network. Subsequently, the policy network iteratively updates its parameters according to the maximum entropy policy gradient theorem. This drives the agent's actions to converge towards ensuring voltage safety and minimizing network losses, including: ; ; ; in, This represents the expected cumulative reward of the agent's policy; This represents the expectation of the policy gradient with respect to the entropy regularization term, where Indicates joint observation From the experience replay buffer Random sampling Indicates action It is determined by the current strategy Generated; The gradient of the function is expressed with respect to the network parameters of the agent g's policy. Represents the multi-agent advantage function; Indicating joint observation o and linked actions Below, the agent g evaluates the network output evaluation value; Represents the baseline value for multiple agents. Given that the actions of other agents are fixed, this represents the expected cumulative reward of agent g based on the current strategy. This means that when calculating the above baseline value, the motion space is excluded. The actions of agent g. During the execution phase, the policy networks of each agent do not require information exchange; they only rely on the observed state of their respective regions. As input, reactive power regulation actions are generated. The policy network output layer outputs normalized actions through the tanh activation function, and further maps the actions to actual reactive power output that meets power constraints.
[0041] To further enhance the stability of the training process, the MAAC algorithm introduces a soft update mechanism for the target network to iteratively update its parameters. During training, this soft update mechanism uses a moving average approach to perform small-scale iterations of the target network parameters. After each parameter update step of the main network, the tracking update of the target network parameters includes: ; ; in, The soft update coefficient is set to 0.01.
[0042] F5. After each preset number of iterations, a preset criterion is determined. If the criterion is met, the reward functions are updated based on the large language model.
[0043] F5 specifically includes: For each agent, after each preset number of iterations, the mean and variance of the network voltage exceedance rate and network loss for each iteration are calculated. If the mean is higher than the preset mean standard, the variance is lower than the preset variance standard, and the iteration has not yet ended, the network loss coefficient and voltage penalty function in each reward function are updated based on the large language model, and the iterative training is restarted.
[0044] In a preferred embodiment of the present invention, the preset number of rounds is set to 20, and the overall network voltage over-limit rate is... and network loss include: ; ; ; Where T represents the number of time steps in a training round, and in this paper, T=240; Indicates the total number of nodes in the network; express Time Node The voltage amplitude; pu and pu represents the preset lower limit of voltage amplitude and the preset upper limit of voltage amplitude, respectively; express Network loss at all times.
[0045] In a preferred embodiment of the present invention, updating the network loss coefficient and voltage penalty function in each reward function based on the large language model includes: First information is obtained; the first information includes network loss and voltage over-limit rate, first amplitude and second amplitude in the control area of each agent; the first amplitude includes the maximum voltage amplitude that exceeds the preset voltage amplitude upper limit; the second amplitude includes the minimum voltage amplitude that exceeds the preset voltage amplitude lower limit; the large language model updates the network loss coefficient and voltage penalty function according to the first information and preset update rules to drive the agent correction strategy and further realize the coordinated optimization of voltage over-limit and network loss.
[0046] In a preferred embodiment of the present invention, the preset update rules include: Prioritize adjusting the voltage penalty function to ensure the elimination of voltage over-limit situations in each region before adjusting the network loss coefficient. In the voltage penalty function, design a gradient distribution of voltage over-limit penalty based on the voltage over-limit rate, the first amplitude, and the second amplitude. The voltage penalty function can be a linear penalty or a non-linear penalty. The adjustment of the network loss coefficient and the voltage penalty function must ensure the elimination of voltage over-limit situations and the reduction of network losses.
[0047] In a preferred embodiment of the present invention, when updating the network loss coefficient and voltage penalty function in each reward function through a large language model, the prompt word design includes: 1) Current reward function: The Python code for the reward function used in the current training phase.
[0048] 2) Agent Training Progress: The total number of training rounds for the agent is 800. We are currently in the kth round, and the agent will be used for a 10-round test. This involves recording the number of training rounds completed and the number of upcoming test rounds for the agent.
[0049] 3) Test results of global and local performance indicators of the agent: The global performance indicator of the agent is network loss, and the local performance indicators of each region include average voltage limit exceedance rate, maximum voltage limit exceedance rate, maximum voltage amplitude exceeding the upper limit, and minimum voltage amplitude exceeding the lower limit.
[0050] 4) Thought Chain Suggestions: Design different levels of voltage limit violation penalties based on the severity of voltage limit violations; voltage limit violation penalties can be linear or non-linear; prioritize adjusting voltage penalties to ensure the elimination of voltage limit violations in each region before increasing the network loss coefficient w, thereby reducing network loss; the voltage penalty and network loss penalty in the hybrid reward function should be balanced to achieve the effect of eliminating voltage limit violations and reducing network loss. This provides a thought process for optimizing the reward function of a large model, guiding the large model to decompose the reward function design task into several sub-tasks, improving the logical rigor and reliability of the results in the large model's inference process.
[0051] F6. After the iteration is completed, each agent controls the photovoltaic inverter in its respective area.
[0052] In the reactive power voltage regulation method based on large model and multi-agent reinforcement learning of this invention, the hybrid reward function adopts a collaborative structure of global reward + local reward, effectively solving the credit allocation problem in multi-agent reinforcement learning. Local rewards provide clear safety boundaries and contribution metrics for each agent, while global rewards guide each agent to focus on the overall economic operation goal of the system. This differentiated design enables agents to accurately balance local voltage safety and global network loss optimization during policy updates, avoiding the mutual constraints between safety and economic goals under a single reward structure. The dynamic update mechanism of the reward function embeds the LLM's adjustment of the reward function into the training loop of reinforcement learning, fundamentally changing the traditional mode where adjustments outside the loop require multiple complete training cycles, reducing the time cost of trial and error iteration, and significantly improving the learning efficiency of the agents. This invention achieves a leap from manually designing reward functions to large model-assisted generation and optimization of reward functions, lowering the threshold for reward function design, avoiding tedious repeated tuning and verification processes, and saving high time and labor costs.
[0053] In a preferred embodiment of the present invention, a reactive power regulation device based on large model and multi-agent reinforcement learning is also provided for use in the method of the present invention. The device includes a first module, a second module, a third module, a fourth module, a fifth module, and a sixth module. The first module is used to treat the photovoltaic controller as an intelligent agent; based on the large language model, a reward function is generated for each intelligent agent; the reward function includes a network loss penalty value calculation term and a voltage over-limit penalty average value calculation term for the nodes under the target intelligent agent; The second module is used to construct a reactive power voltage regulation model based on the operating constraints of the distribution network with the goal of minimizing network losses. The third module is used to construct a Markov decision process based on a reactive power regulation model, an agent, and a reward function. The fourth module is used for iterative training based on the MAAC algorithm and Markov decision process; The fifth module is used to determine the preset criteria after each preset round of iteration. If the criteria are met, the reward functions are updated based on the large language model. The sixth module is used to control the photovoltaic inverters in their respective areas based on each intelligent agent after the iteration is completed.
[0054] The reactive power voltage regulation device based on large model and multi-agent reinforcement learning of the present invention, when used in the method of the present invention, has the same beneficial effects as the method of the present invention.
[0055] Verification section: To verify the effectiveness of the method of this invention, the training process of the intelligent agent was implemented in a Python 3.9.23 virtual environment, and accelerated using the PyTorch deep learning framework. The hyperparameter settings of the MAAC algorithm used are shown in Table 1. Table 1. Hyperparameters of the MAAC Algorithm ; Figures 2 to 3 The training convergence curve of the agent in the method of this invention is shown, where the blue markers indicate the moments when the large language model adjusts the reward function. In the initial 300 rounds of training, both the voltage limit violation rate and network loss exhibit significant fluctuations, reflecting the inadequacy of the initially designed reward function in balancing safety constraints and economic objectives. As training progresses, the LLM dynamically reconstructs the mathematical form and weight coefficients of the hybrid reward function based on test feedback, thereby guiding the agent's optimization direction. Under this continuous dynamic guidance, the agent's strategy gradually converges to the optimal solution: the voltage limit violation rate steadily decreases and converges to near zero after 600 rounds; simultaneously, the network loss remains stably low while ensuring voltage safety constraints, ultimately achieving a synergistic optimization of safety and economy.
[0056] To comprehensively evaluate the model's control performance, 5760 samples were randomly selected from the test set over 12 days to test the agent's performance. The number of voltage limit violations, the maximum voltage limit violation, the minimum voltage limit violation, and the mean and variance of network loss are shown in Table 2. Table 2 Mean and Variance Table ; The results show that, compared with the method before optimization, the method of the present invention has achieved significant improvements in the number of voltage limit violations, the maximum value of voltage limit violations, the minimum value of voltage limit violations, and the mean and variance of network loss in the test set, and has eliminated the number of voltage limit violations, proving the effectiveness and superiority of the method of the present invention in suppressing voltage limit violations.
[0057] To visually demonstrate the control performance of the method of this invention on a typical day, one day was randomly selected from the test set, with a total of 480 sets of sample data to test the agent. The maximum, minimum, mean, and variance of the network loss of the agent on a typical day were statistically analyzed, as shown in Table 3: Table 3. Maximum, minimum, mean, and variance of network loss on a typical day. ; Figures 4 to 5 The study demonstrates the temporal variation trend of voltage extreme values within a typical day. Observations show that, guided by the method of this invention, the agent's actions not only strictly adhere to voltage safety constraints but also search for a more economical operating point within the feasible region, fully verifying the superiority of the method in suppressing voltage exceedances and reducing network losses.
[0058] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A reactive power voltage regulation method based on large model and multi-agent reinforcement learning, characterized in that, include: The photovoltaic controller is used as an intelligent agent; Reward functions are generated for each agent based on a large language model; the reward function includes a network loss penalty calculation term and a voltage limit violation penalty average calculation term for nodes under the target agent; A reactive power voltage regulation model is constructed based on the distribution network operation constraints with the goal of minimizing network losses; a Markov decision process is constructed based on the reactive power voltage regulation model, the agent, and the reward function; Iterative training is performed based on the MAAC algorithm and the Markov decision process. After each preset number of iterations, a preset criterion is determined. If the criterion is met, the reward functions are updated based on the large language model, including: For each agent, after each preset number of iterations, the mean and variance of the network voltage exceedance rate and network loss for each iteration are calculated. If the mean is higher than the preset mean standard, the variance is lower than the preset variance standard, and the iteration has not yet ended, the network loss coefficient and voltage penalty function in each reward function are updated based on the large language model, and the iterative training is restarted. After the iteration is completed, each agent controls the photovoltaic inverter in its respective area.
2. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 1, characterized in that, Based on a large language model, reward functions are generated for each agent, including: The large language model generates the network loss penalty value calculation item and the voltage over-limit penalty average value calculation item based on the agent and the pre-acquired distribution network topology, and generates a reward function based on the network loss penalty value calculation item and the voltage over-limit penalty average value calculation item; the network loss penalty value calculation item includes a network loss coefficient; the voltage over-limit penalty average value calculation item includes a voltage penalty function, and the voltage penalty function adopts a symmetrical structure for upper and lower limits.
3. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 2, characterized in that, The reward function includes: For the reward function of agent g , represented as: ; ; ; in, For local reward terms, it represents the average voltage over-limit penalty of all nodes within the control area of agent g; This is a global reward item, representing the network loss penalty value; Represents a set of intelligent agents; This represents the total number of nodes within the control area of agent g; It is a voltage penalty function; For nodes The measured voltage amplitude; This is the voltage reference value; This is the network loss coefficient; This is due to network loss.
4. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 3, characterized in that, Updating the network loss coefficient and voltage penalty function in each reward function based on the large language model includes: Obtain first information; the first information includes network loss and voltage over-limit rate, first amplitude and second amplitude in the control area of each intelligent agent; the first amplitude includes the maximum voltage amplitude that exceeds the preset voltage amplitude upper limit; the second amplitude includes the minimum voltage amplitude that exceeds the preset voltage amplitude lower limit; the large language model updates the network loss coefficient and the voltage penalty function according to the first information and the preset update rules.
5. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 4, characterized in that, The preset update rules include: Prioritize adjusting the voltage penalty function to ensure the elimination of voltage over-limit situations in each region before adjusting the network loss coefficient. The voltage penalty function is designed with a gradient distribution based on the voltage over-limit rate, the first amplitude, and the second amplitude. The voltage penalty function can be a linear or nonlinear penalty. The adjustment of the network loss coefficient and the voltage penalty function must ensure the elimination of voltage over-limit situations and the reduction of network losses.
6. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 5, characterized in that, A reactive power voltage regulation model based on distribution network operation constraints and aiming to minimize network losses is constructed, including: The objective function includes: ; in, For network loss; For the set of routes; For the line The resistance; and The lines are respectively The active and reactive power; The parent node of the node. For nodes The voltage amplitude; The power distribution network operation constraints include: Power balance constraints: Power balance constraints in the distribution network are described using the Distflow model; Node voltage constraint: Node voltage deviation is controlled within ±5% of the rated voltage; Photovoltaic inverter operating constraints: The active power and apparent power of the photovoltaic inverter are maintained within the preset range.
7. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 6, characterized in that, The Markov decision process constructed based on the aforementioned reactive power regulation model, agent, and reward function includes: Using the aforementioned reactive power regulation model as the interactive environment, the reactive power regulation problem in the distribution network is described as a partially observable Markov decision process involving multiple agents, including a seven-tuple. ;in, A collection of intelligent agents, o For the joint observation space set of intelligent agents, S For a set of state spaces, For the action space set, Let be the state transition probability. R For joint awards, Discount factor; State space set ;in, and These represent the active and reactive load sets of all nodes in the network, respectively. V Represents the set of node voltage amplitudes across the entire network; and These represent the sets of active and reactive power of the photovoltaic inverters in the network, respectively. Joint observation space set of intelligent agents o In the observation space of agent g ;in, , and These represent the active load, reactive load, and voltage amplitude of all nodes within the area controlled by agent g, respectively. , Represent the active and reactive power of the photovoltaic inverter controlled by agent g, respectively; the joint observation space set of the agent. and ; Action space set In the middle, region z The set of actions of all photovoltaic inverters in the system is represented as , Indicates the area z The collection of all photovoltaic inverters in the system. Indicates photovoltaic inverter Intelligent agent actions; photovoltaic inverter Actual output reactive power With agent actions The mapping relationship is represented as ,in For photovoltaic inverters Rated apparent power, For photovoltaic inverters The active power; when The inverter absorbs reactive power from the system, and conversely, the inverter injects reactive power into the system. For state transition probability Represented as This indicates that the distribution network is in control action. Under the influence of the action, the system changes from state Evolved to The dynamic process is simulated using the Distflow power flow calculation method.
8. The reactive power regulation method based on large model and multi-agent reinforcement learning according to claim 7, characterized in that, Iterative training based on the MAAC algorithm and the Markov decision process includes: Based on the MAAC algorithm, an independent evaluation network and policy network are built for each agent, and a target evaluation network and a target policy network corresponding to the evaluation network and policy network are built respectively. An attention mechanism is introduced into the evaluation network of each agent. Based on the built evaluation network, policy network, target evaluation network and target policy network, centralized learning and iterative training are carried out in each agent using the joint observation space set and action space set of all agents.
9. A reactive power voltage regulation device based on large model and multi-agent reinforcement learning, used in the method described in any one of claims 1 to 8, characterized in that, The device includes a first module, a second module, a third module, a fourth module, a fifth module, and a sixth module; The first module is used to treat the photovoltaic controller as an intelligent agent; and to generate reward functions for each intelligent agent based on the large language model; the reward function includes a network loss penalty value calculation term and a voltage over-limit penalty average value calculation term for the nodes under the target intelligent agent; The second module is used to construct a reactive power voltage regulation model based on the operating constraints of the distribution network with the goal of minimizing network losses; The third module is used to construct a Markov decision process based on the reactive power regulation model, the agent, and the reward function. The fourth module is used for iterative training based on the MAAC algorithm and the Markov decision process. The fifth module is used to determine the preset criteria after each preset round of iteration. If the criteria are met, the reward functions are updated based on the large language model. The sixth module is used to control the photovoltaic inverters in their respective areas based on each intelligent agent after the iteration is completed.