Voltage control method and device of power distribution network, computer device and storage medium
Patent Information
- Application Number
- CN202610724195.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-05-25
AI Technical Summary
传统主动电压控制方法通常依赖精确潮流模型和集中优化求解,在节点规模增大、工况快速变化时,存在计算开销高、实时性不足、通信负担重等问题
[0009]在本申请实施例中,提供一种配电网的电压控制方法、装置、计算机设备以及存储介质,根据各个逆变器智能体的状态空间数据以及反映配电网的拓扑结构的物理先验矩阵集合进行动作推理以及动作执行,结合动作执行后计算的各个逆变器智能体的奖励值进行逆变器智能体更新,提高了电压控制的准确性以及效率。
Smart Images

Figure CN122267936B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial intelligent control technology, and in particular to a voltage control method, device, computer equipment, and storage medium for a power distribution network. Background Technology
[0002] With the high proportion of distributed photovoltaic inverters connected to the distribution network, node voltage is easily affected by fluctuations in photovoltaic output and load, leading to situations where it exceeds the upper or lower limit. Traditional active voltage control methods typically rely on accurate power flow models and centralized optimization solutions. However, when the node scale increases and operating conditions change rapidly, these methods suffer from high computational overhead, insufficient real-time performance, and heavy communication burden. Summary of the Invention
[0003] Based on this, the purpose of the present invention is to provide a voltage control method, device, computer equipment, and storage medium for a power distribution network. The method performs action reasoning and action execution based on the state space data of each inverter agent and the physical prior matrix set reflecting the topology of the power distribution network. The method updates the inverter agents by combining the reward values of each inverter agent calculated after the action execution, thereby improving the accuracy and efficiency of voltage control.
[0004] In a first aspect, embodiments of this application provide a voltage control method for a power distribution network, comprising the following steps:
[0005] Based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents in the distribution network, and the local attention bias matrix of each inverter agent in a preset physical prior matrix set, the target node embedding representation and action space data of each inverter agent at each target time are obtained. The distribution network is configured with an agent update network, which includes a main value network and a graph encoder. The local attention bias matrix is used to characterize the physical correlation strength between nodes within the local region of each inverter agent. Voltage control and reward calculation are performed based on the action space data of each inverter agent at each target time to obtain the global reward value of the distribution network at each target time. The global state space data, global reward value, local observation data of each inverter agent, target node embedding representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time are combined to construct training samples for each target time step. The target node embedding representation of each inverter agent at each target time and the action space data of the same inverter agent in the training samples are spliced and mapped to obtain the agent global input feature set at each target time. The agent global input feature set includes the agent global input features of each inverter agent. The global input feature set of the agent at each target time and the global physical bias matrix of the agent in the physical prior matrix set are input into the graph encoder in the main value network for encoding processing to obtain the global encoding result at each target time; the global encoding results at each target time are aggregated to obtain the global value estimate at each target time. Based on the global value estimate at each target moment, each inverter agent is updated to obtain several target inverter agents. Based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the set of physical prior matrices, voltage control of the distribution network is performed.
[0006] Secondly, embodiments of this application provide a voltage control device for a power distribution network, comprising: The action reasoning module is used to obtain the target node embedding representation and action space data of each inverter agent at each target time based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents of the distribution network, and the local attention bias matrix of each inverter agent in a preset physical prior matrix set. The distribution network is configured with an agent update network, which includes a main value network; the main value network includes a graph encoder; and the local attention bias matrix is used to characterize the physical association strength between nodes within the local region of each inverter agent. The reward calculation module is used to perform voltage control and reward calculation based on the action space data of each inverter agent at each target time, so as to obtain the global reward value of the distribution network at each target time. The data processing module is used to combine the global state space data, global reward value, local observation data of each inverter agent, target node embedded representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time to construct training samples for each target time step. The global value estimation module is used to concatenate and map the target node embedding representation of each inverter agent at each target time and the action space data of the same inverter agent in the training samples to obtain the global input feature set of the agent at each target time. The global input feature set of the agent includes the global input features of each inverter agent. The global input feature set of the agent at each target time and the global physical bias matrix of the agent in the physical prior matrix set are input into the graph encoder in the main value network. The graph encoder is processed through multiple attention layers, residual connection layers, feedforward networks and normalization layers to obtain the global encoding result at each target time. The global encoding results at each target time are aggregated to obtain the global value estimate at each target time. The agent update module is used to update each inverter agent based on the global value estimate at each target time, and obtain several target inverter agents. The voltage control module is used to perform voltage control of the distribution network based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the physical prior matrix set.
[0007] Thirdly, embodiments of this application provide a computer device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, it implements the steps of the voltage control method for the power distribution network as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a storage medium storing a computer program that, when executed by a processor, implements the steps of the voltage control method for the power distribution network as described in the first aspect.
[0009] In this application embodiment, a voltage control method, device, computer equipment, and storage medium for a power distribution network are provided. Action reasoning and action execution are performed based on the state space data of each inverter agent and the physical prior matrix set reflecting the topology of the power distribution network. The inverter agents are updated by combining the reward values of each inverter agent calculated after the action execution, thereby improving the accuracy and efficiency of voltage control.
[0010] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description
[0011] Figure 1 This is a schematic flowchart of a voltage control method for a power distribution network provided in one embodiment of this application; Figure 2A schematic flowchart of step S6 in a voltage control method for a power distribution network provided in another embodiment of this application; Figure 3 This is a schematic flowchart of step S1 in a voltage control method for a power distribution network provided in one embodiment of this application; Figure 4 This is a flowchart illustrating step S2 of a voltage control method for a power distribution network provided in one embodiment of this application. Figure 5 This is a flowchart illustrating step S21 of a voltage control method for a power distribution network provided in one embodiment of this application. Figure 6 This is a schematic flowchart of step S22 in a voltage control method for a power distribution network provided in one embodiment of this application. Figure 7 This is a flowchart illustrating step S4 of a voltage control method for a power distribution network provided in one embodiment of this application. Figure 8 A schematic flowchart of step S4 in a voltage control method for a power distribution network provided in another embodiment of this application; Figure 9 A schematic diagram of the structure of a voltage control device for a power distribution network provided in one embodiment of this application; Figure 10 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation
[0012] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0013] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0014] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0015] Please see Figure 1 , Figure 1 The following is a flowchart illustrating a voltage control method for a power distribution network according to an embodiment of this application. The method includes the following steps: S1: Based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents of the distribution network, and the local attention bias matrix of each inverter agent in the preset physical prior matrix set, the target node embedding representation and action space data of each inverter agent at each target time are obtained.
[0016] The execution entity of the voltage control method for the power distribution network in this application is the control equipment for the voltage control method of the power distribution network (hereinafter referred to as the control equipment). In an optional embodiment, the control equipment may be a computer device, a server, or a server cluster composed of multiple computer devices.
[0017] In this embodiment, the control device obtains the target node embedding representation and action space data of each inverter agent at each target time based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents of the distribution network, and the local attention bias matrix of each inverter agent in the preset physical prior matrix set.
[0018] The distribution network is a network structure consisting of buses and lines. Some buses are connected to distributed renewable energy devices and equipped with inverters with reactive power regulation capabilities. Each inverter agent corresponds one-to-one with each inverter in the distribution network and is configured to independently sense the operating status of the inverter, the voltage fluctuation of each node in the local area corresponding to the inverter with load, and the changes in the output of distributed renewable energy devices, as local observation data. Each inverter agent generates an inverter control action based on the operating status of its local area.
[0019] Specifically, the global state space data includes the load active power vector, load reactive power vector, distributed renewable energy active power vector, inverter reactive power vector, voltage amplitude vector, and voltage phase angle vector of several nodes.
[0020] The local observation data includes node feature vectors of each node within the local area of the inverter. The node feature vectors include node load active power, node load reactive power, voltage amplitude, voltage phase angle, distributed renewable energy device installation marker variables, distributed renewable energy active power output, and inverter reactive power output.
[0021] The distribution network is configured with an agent update network, which includes a main value network; the main value network includes a graph encoder; the local attention bias matrix is used to characterize the physical correlation strength between nodes in the local area of each inverter agent. In an optional embodiment, in order to explicitly introduce the structural features and electrical coupling features of the real power grid in the subsequent local observation encoding stage and centralized value assessment stage, step S8 is also included: constructing the physical prior matrix set.
[0022] Please see Figure 2 , Figure 2 A flowchart illustrating step S8 of a voltage control method for a power distribution network provided in another embodiment of this application includes steps S81 to S86, as detailed below: S81: Obtain the physical prior information of the distribution network.
[0023] In this embodiment, the control device obtains the physical prior information of the distribution network, wherein the physical prior information includes the distribution network topology, the bus location of each inverter, the local area division result of the distribution network, and the power flow calculation result; the power flow calculation result includes the reactive power-voltage sensitivity of each node in the local area of each inverter to the corresponding inverter.
[0024] S82: Based on the bus positions of each inverter, construct a first adjacency matrix and a second adjacency matrix; normalize the first adjacency matrix to obtain the topological adjacency matrix of the agent; normalize the second adjacency matrix to obtain the sensitivity adjacency matrix of the agent.
[0025] In this embodiment, the control device constructs a first adjacency matrix and a second adjacency matrix based on the location of each inverter on the bus. The first adjacency matrix includes the topological proximity of each inverter's bus to the other inverters' buses in the distribution network. The second adjacency matrix includes the intensity of the impact of each inverter's reactive power change on the voltage of the other inverters' buses.
[0026] The control device normalizes the first adjacency matrix to obtain the topological adjacency matrix of the agent; and normalizes the second adjacency matrix to obtain the sensitivity adjacency matrix of the agent.
[0027] S83: Construct the local attention bias matrix based on the local topology bias matrix and local sensitivity bias matrix of each inverter agent to obtain the local attention bias matrix of each inverter agent.
[0028] In this embodiment, the control device constructs a local attention bias matrix based on the local topology bias matrix and local sensitivity bias matrix of each inverter agent, thereby obtaining the local attention bias matrix of each inverter agent, as described below:
[0029] In the formula, For the first i Local attention bias matrix of each inverter agent The first topological bias weight coefficient, For the first i Local topological bias matrix of each inverter agent This is the first sensitivity bias weighting coefficient. For the first i Local sensitivity bias matrix of each inverter agent.
[0030] S84: Based on the power distribution network topology, construct the local topology matrix of each inverter, the local topology matrix including the topological proximity between nodes in the local area; normalize the local topology matrix of each inverter to obtain the local topology bias matrix of each inverter agent.
[0031] In this embodiment, the control device constructs a local topology matrix for each inverter based on the power distribution network topology, wherein the local topology matrix includes the topological proximity between nodes within a local area.
[0032] The control device normalizes the local topology matrix of each inverter to obtain the local topology bias matrix of each inverter agent.
[0033] S85: Based on the power flow calculation results, construct the local sensitivity matrix of each inverter. The local sensitivity matrix includes the sensitivity response of each node in the local area to the corresponding inverter. Normalize the local sensitivity matrix of each inverter to obtain the local sensitivity bias matrix of each inverter agent.
[0034] In this embodiment, the control device constructs a local sensitivity matrix for each inverter based on the power flow calculation results. The local sensitivity matrix includes the sensitivity response of each node in the local area to the corresponding inverter.
[0035] S86: Construct the global physical bias matrix of the agent based on the agent topological adjacency matrix and the agent sensitivity adjacency matrix to obtain the global physical bias matrix of the agent.
[0036] In this embodiment, the control device constructs the agent's global physical bias matrix based on the agent's topological adjacency matrix and agent sensitivity adjacency matrix, as described below:
[0037] In the formula, This is the global physical bias matrix of the agent. The second topological bias weighting coefficient. This is the topological adjacency matrix of the intelligent agents. This is the second sensitivity bias weighting coefficient. This is the adjacency matrix for the agent's sensitivity.
[0038] Please see Figure 3 , Figure 3 The flowchart of step S1 in the voltage control method for a power distribution network provided in one embodiment of this application includes steps S11 to S14, as follows: S11: Perform linear mapping and region identifier embedding on the node feature vectors of each node in the local observation data of each inverter agent at each target time to obtain the node feature representation of each inverter agent at each target time.
[0039] In this embodiment, the control device performs linear mapping and region identifier embedding on the node feature vectors of each node in the local observation data of each inverter agent at each target time to distinguish different local regions and obtain the node feature representation of each inverter agent at each target time.
[0040] Specifically, the control device performs linear mapping and region identifier embedding based on the node feature vectors of each node in the local observation data of each inverter agent at each target time and a preset embedding vector calculation algorithm to obtain the embedding vectors of each node of each inverter agent at each target time. The embedding vector calculation algorithm is as follows:
[0041] In the formula, For the first t The target time of the first i The first inverter agent m The embedding vector of each node. For mapping weight parameters, For the first t The target time of the firsti The first inverter agent m The node feature vector of each node. For mapping bias parameters, For the first i A local area identifier for a local area of an inverter smart agent.
[0042] The control device concatenates the embedding vectors of each node of the same inverter agent at the same target time to obtain the node feature representation of each inverter agent at each target time.
[0043] S12: Input the node feature representations and local attention bias matrices of each inverter agent at each target time into a preset local encoder, and process them sequentially through the multi-layer attention layer, residual connection layer, feedforward network and normalization layer of the local encoder to obtain the local encoding results of each inverter agent at each target time.
[0044] In this embodiment, the control device inputs the node feature representations of each inverter agent at each target time and the local attention bias matrix into a preset local encoder. The local encoder processes the data through multiple attention layers, residual connection layers, feedforward networks, and normalization layers to obtain the local encoding results of each inverter agent at each target time.
[0045] Specifically, the control device inputs the node feature representations of each inverter agent at each target time and the local attention bias matrix into the local encoder. In any attention layer, local physical bias attention calculation is performed according to a preset local physical bias attention calculation algorithm to obtain the local physical bias attention representations of each inverter agent at each target time output by each attention layer. The local attention bias matrix is introduced as an additive bias into the local attention calculation to guide the attention mechanism to prioritize node pairs with stronger electrical coupling. This allows the local encoder to consider both data-driven characteristics and grid physical laws when modeling node interaction relationships. Nodes that are closer to the inverter topologically or more sensitive to its reactive power regulation receive higher weights in attention allocation. The local physical bias attention calculation algorithm is as follows:
[0046] In the formula, For the first l The first layer of attention layer output i Local physical bias attention representation of each inverter agent. For normalized activation functions, , , The first lThe first layer of attention layer i The query matrix, key matrix, and value matrix of each inverter agent are obtained by linear mapping based on the node feature representations of each inverter agent. The dimension of the key vector. For the first l The first layer of attention layer i A mask matrix for each inverter agent, wherein the mask matrix is constructed based on local observation data of the inverter agents.
[0047] The control device inputs the local physical bias attention representation of each inverter agent at each target time and the node feature representation of each inverter agent at each target time, which are output from each attention layer, into the residual connection layer for residual connection to obtain the residual connection representation of each inverter agent at each target time.
[0048] The control device inputs the residual connection representation of each inverter agent at each target time into the feedforward network to perform hidden layer calculation, thereby obtaining the hidden layer representation of each inverter agent at each target time.
[0049] The control device inputs the hidden layer representation of each inverter agent at each target time into the normalization layer for normalization processing to obtain the local encoding result of each inverter agent at each target time.
[0050] S13: Obtain the target node position index of each inverter agent; based on the target node position index of each inverter agent, extract the target node embedding from the corresponding local encoding result to obtain the target node embedding representation of each inverter agent at each target time.
[0051] In this embodiment, the control device obtains the target node location index of each inverter agent, wherein the target node location index indicates the location of the node where the inverter is placed.
[0052] The control device extracts the target node embedding from the corresponding local encoding results based on the target node position index of each inverter agent, and obtains the target node embedding representation of each inverter agent at each target time. The target node embedding representation represents the local region state observed by the current inverter agent from its inverter perspective.
[0053] S14: The target node embedding representation of each inverter agent at each target time is input into a preset shared strategy head for timing modeling and normalized action generation, so as to obtain the normalized action of each inverter agent at each target time as the action space data, and thus obtain the action space data of each inverter agent at each target time.
[0054] In this embodiment, the control device embeds the target node representation of each inverter agent at each target time into a preset shared strategy header for timing modeling and normalized action generation, thereby obtaining the normalized actions of each inverter agent at each target time as the action space data, and thus obtaining the action space data of each inverter agent at each target time.
[0055] Specifically, the shared strategy head includes a timing modeling unit and an action output layer. The control device inputs the target node embedding representation of each inverter agent at each target time into the timing modeling unit for timing modeling, thereby obtaining the timing hidden state representation of each inverter agent at each target time.
[0056] The control device inputs the time-series hidden state representation of each inverter agent at each target time into the action output layer for normalized action generation, thereby obtaining the normalized actions of each inverter agent at each target time, as described below:
[0057] In the formula, For the first t The target time of the first i Normalized actions of each inverter agent This is the upper bound of the motion amplitude. For action weight parameters, For the first t The target time of the first i The temporal hidden state representation of each inverter agent. For action bias parameters, It is the hyperbolic tangent function.
[0058] S2: Perform voltage control and reward calculation based on the action space data of each inverter agent at each target time to obtain the global reward value of the distribution network at each target time.
[0059] In this embodiment, the control device performs voltage control and reward calculation based on the action space data of each inverter agent at each target time to obtain the global reward value of the distribution network at each target time.
[0060] Please see Figure 4 , Figure 4 The flowchart of step S2 in the voltage control method for a power distribution network provided in one embodiment of this application includes steps S21 to S22, as follows: S21: Execute actions based on the inverter action space data set at each target time to obtain the voltage control execution index of the distribution network at each target time.
[0061] In this embodiment, the control device performs actions based on the inverter action space data set at each target time to obtain the voltage control execution index of the distribution network at each target time. The voltage control execution index includes the voltage amplitude of several voltage monitoring nodes and the reactive power output vector of each inverter.
[0062] Please see Figure 5 , Figure 5 The flowchart of step S21 in the voltage control method for a power distribution network provided in one embodiment of this application includes steps S211 to S212, as follows: S211: Calculate the maximum adjustable reactive power range for each inverter agent at each target time by using the distributed renewable energy active power vector and the rated apparent power limit for each inverter agent at each target time.
[0063] In this embodiment, the control device calculates the maximum adjustable reactive power range based on the distributed renewable energy active power vector of each inverter agent at each target time and the rated apparent power limit, thereby obtaining the maximum adjustable reactive power range of each inverter agent at each target time. The maximum adjustable reactive power range is:
[0064] In the formula, For the first t The target time of the first i The maximum adjustable reactive power range of an inverter agent. For the first i The rated apparent power limit of an inverter smart agent. For the first t The target time of the first i Distributed new energy active power vector of an inverter intelligent agent.
[0065] S212: Based on the action space data of each inverter intelligent entity at each target time and the maximum adjustable reactive power range, perform instruction mapping to obtain the reactive power control instructions of each inverter intelligent entity at each target time; perform actions based on the reactive power control instructions of each inverter intelligent entity at each target time to obtain the voltage control execution index of the distribution network at each target time.
[0066] In this embodiment, the control device performs command mapping based on the action space data of each inverter agent at each target time and the maximum adjustable reactive power range to obtain the reactive power control commands for each inverter agent at each target time, as described below:
[0067] In the formula, For the first t The target time of the first i Reactive power control commands for each inverter intelligent agent For the first t The target time of the first i Action space data of an inverter intelligent agent.
[0068] In an optional embodiment, to improve execution safety, the control device further performs amplitude limiting processing on the reactive power control commands of each inverter intelligent entity at each target time, obtaining the amplitude-limited reactive power control commands of each inverter intelligent entity at each target time, as described below:
[0069] In the formula, The first after the amplitude limiting process t The target time of the first i Reactive power control commands for each inverter intelligent agent This is a truncation function used to ensure that the control command does not exceed the current physical capacity limit of the inverter. In another optional embodiment, the control device can also perform speed limiting or smoothing on the difference between the current command and the reactive power control command executed in the previous cycle to suppress control jitter.
[0070] The control equipment executes actions according to the reactive power control instructions of each inverter intelligent agent at each target time, and obtains the voltage control execution index of the distribution network at each target time. The voltage control execution index includes the voltage amplitude of several nodes that need to be monitored and the reactive power output vector of each inverter.
[0071] S22: Calculate the reward based on the voltage control execution index of the distribution network at each target time to obtain the global reward value of the distribution network at each target time.
[0072] In this embodiment, the control device performs reward calculations based on the voltage control execution indicators of the distribution network at each target time to obtain the global reward value of the distribution network at each target time.
[0073] Please see Figure 6 , Figure 6 The flowchart of step S22 in the voltage control method for a power distribution network provided in one embodiment of this application includes steps S221 to S222, as follows: S221: Calculate the voltage limit penalty based on the voltage amplitude of each voltage detection node at each target time to obtain the voltage limit penalty value of each voltage detection node at each target time; calculate the reactive power regulation loss based on the reactive power output vector of each inverter at each target time and the preset weighting coefficient to obtain the reactive power regulation loss value of each inverter at each target time.
[0074] In this embodiment, the control device calculates the voltage limit penalty based on the voltage amplitude of each voltage detection node at each target time, and obtains the voltage limit penalty value of each voltage detection node at each target time.
[0075] The control device calculates the reactive power regulation loss based on the reactive power output vector of each inverter at each target time and the preset weighting coefficient, and obtains the reactive power regulation loss value of each inverter at each target time. The weighting coefficient is used to balance the voltage control effect and the reactive power regulation cost.
[0076] S222: Subtract the voltage amplitude of each voltage detection node at the same target time from the reactive power regulation loss value of each inverter to obtain the local reward value of each voltage detection node and each inverter at each target time; and calculate the weighted average of the local reward values of each voltage detection node and each inverter at each target time to obtain the global reward value of the distribution network at each target time.
[0077] In this embodiment, the control device subtracts the voltage amplitude of each voltage detection node at the same target time from the reactive power regulation loss value of each inverter to obtain the local reward value of each voltage detection node and each inverter at each target time.
[0078] The control equipment performs a weighted average of the local reward values of each voltage detection node and each inverter at each target time to obtain the global reward value of the distribution network at each target time. Under the conditions of load fluctuation and changes in distributed renewable energy output, it coordinates the reactive power output of each inverter to keep the voltage of all network nodes within a safe range as much as possible, while minimizing the cost of reactive power regulation. The global reward value is:
[0079] In the formula, For the first t The global reward value at each target time. V For a set of voltage sensing nodes, For the first t The target time of the first j The voltage amplitude of each voltage detection node, For the first t The target time of the first jVoltage over-limit penalty value for each voltage detection node The weighting coefficient is used to balance the voltage control effect with the reactive power regulation cost. For the first t The reactive power regulation loss value of the inverter at a target time.
[0080] S3: Combine the global state space data, global reward value, local observation data of each inverter agent, target node embedded representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time to construct training samples for each target time step.
[0081] In this embodiment, the control device obtains the global state space data of the agent voltage control network at the next moment and the local observation set of the inverter agent.
[0082] The control device combines the global state space data, global reward value, local observation data of each inverter agent, target node embedded representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time to construct training samples for each target time, and stores them in a preset shared experience playback pool.
[0083] S4: The target node embedding representations of each inverter agent at each target time and the action space data of the same inverter agent in the training samples are concatenated and mapped to obtain the global input feature set of the agent at each target time.
[0084] In this embodiment, the control device splices and maps the target node embedding representation of each inverter agent at each target time and the action space data of the same inverter agent in the training samples to obtain the agent global input feature set at each target time. The agent global input feature set includes the agent global input features of each inverter agent.
[0085] S5: Input the agent's global input feature set at each target time and the agent's global physical bias matrix in the physical prior matrix set into the graph encoder in the main value network for encoding processing to obtain the global encoding result at each target time; aggregate the global encoding results at each target time to obtain the global value estimate at each target time.
[0086] In this embodiment, the control device inputs the agent's global input feature set at each target time and the agent's global physical bias matrix in the physical prior matrix set into the graph encoder in the main value network. The graph encoder then performs encoding processing through multiple attention layers, residual connection layers, feedforward networks, and normalization layers to obtain the global encoding result at each target time.
[0087] Specifically, the control device inputs the global input feature set of the agent at each target time and the global physical bias matrix of the agent into the graph encoder in the main value network. In any attention layer, global physical bias attention is calculated according to a preset global physical bias attention calculation algorithm to obtain the global physical bias attention representation of the agent at each target time output by each attention layer. The interaction of inverter agents with valid physical associations is retained, and higher interaction priority is given to inverter agent pairs that are closer in topology or have stronger electrical coupling. The global physical bias attention calculation algorithm is as follows:
[0088] In the formula, For the first t The target time of the first l The global physical bias attention representation of the agent's global other attention layer outputs the attention layer. , , The first t The target time of the first l The attention layer contains the agent's global query matrix, key matrix, and value matrix, which are obtained by linear mapping based on the agent's global input feature set at the corresponding target time. This is the global physical bias matrix of the agent. The graph mask matrix is constructed based on the edge set of the agent interaction graph, which is constructed based on the agent topological adjacency matrix and the agent sensitivity adjacency matrix. The agent interaction graph is used to indicate the coupling relationship of inverter agents in a real power distribution network.
[0089] By constructing a physically guided agent interaction graph in the main value network and performing attention interactions only on strongly coupled or necessary connected inverter agents, the redundant modeling error caused by fully connected value assessment can be reduced, and the accuracy of characterizing the collaborative relationship of multiple inverter agents can be improved.
[0090] The control device inputs the global physical bias attention representation of the agent at each target time step, output from each attention layer, and the agent's global input feature set at each target time step, into the residual connection layer for residual connection to obtain the residual connection representation at each target time step. The control device then inputs the agent's global residual connection representation at each target time step into the feedforward network for hidden layer computation to obtain the hidden layer representation at each target time step. Finally, the control device inputs the agent's global hidden layer representation at each target time step into the normalization layer for normalization processing to obtain the global encoding result at each target time step.
[0091] The control device aggregates the global encoding results at each target time to obtain a global value estimate for each target time. This estimate is used to evaluate the impact of the joint actions of multiple inverter agents on the active voltage control effect of the entire distribution network, providing global value feedback for the shared policy network. The global value estimate is as follows:
[0092] In the formula, For parameters Global value estimation of the output of a centralized value network. For the first t The global encoding result at each target time point. It is an aggregate function. This is the output mapping function.
[0093] S6: Based on the global value estimate at each target time, update each inverter agent to obtain several target inverter agents.
[0094] In this embodiment, the control device updates each inverter agent based on the global value estimate at each target time, obtaining several target inverter agents. By introducing a set of physical prior matrices, the agent voltage control network can prioritize key nodes that are more closely related to the current inverter's electrical relationship, reducing the interference of invalid interactions and irrelevant features on policy learning, thereby improving the physical consistency of local representations.
[0095] The agent update network also includes a shared policy network; the model parameters of the action policy networks of all inverter agents call the model parameters of the shared policy network. See also... Figure 7 , Figure 7 The flowchart of step S6 in the voltage control method for a power distribution network provided in one embodiment of this application includes steps S61 to S64, as follows: S61: Based on the voltage amplitude of each node in the local area of each inverter agent in the training samples at each target time, labels are constructed to obtain the label data of the training samples at each target time.
[0096] In this embodiment, the control device constructs labels based on the voltage amplitude of each node in the local area of each inverter agent in the training samples at each target time, and obtains the label data of the training samples at each target time. This data is used to introduce auxiliary supervision signals that are directly related to the voltage safety status and short-time voltage evolution of the distribution network, so as to enhance the physical consistency and training stability of the local encoder output representation.
[0097] Specifically, the tag data includes true tags for the local voltage exceedance ratio of each inverter agent and true tags for the local average voltage deviation. The true tags for the local voltage exceedance ratio are used to characterize the voltage safety risk in a local area at the current moment. The true tags for the local voltage exceedance ratio are:
[0098] In the formula, For the first t The target time of the first i Real-time labeling of the local voltage over-limit ratio of each inverter smart agent. For the first i Number of nodes in a local region For the first i A local area, , These are the lower voltage safety limit and the upper voltage safety limit, respectively.
[0099] The local average voltage deviation true label characterizes the short-time evolution trend of the local area voltage under the current action. The local average voltage deviation true label is:
[0100] In the formula, For the first t The target time of the first i The local average voltage deviation of each inverter agent is accurately labeled. This is the reference voltage.
[0101] S62: Obtain the target node embedding representation of each inverter agent at each target time; input the target node embedding representation of each inverter agent at each target time into the preset local voltage over-limit ratio prediction head and average voltage deviation prediction head for prediction, and obtain the local voltage over-limit ratio prediction label and local average voltage deviation prediction label of each inverter agent at each target time.
[0102] In this embodiment, the control device obtains the target node embedding representation of each inverter agent at each target time. For specific implementation, please refer to step S13, which will not be repeated here.
[0103] The control device inputs the target node embedding representation of each inverter intelligent agent at each target time into the preset local voltage over-limit ratio prediction head and average voltage deviation prediction head for prediction, and obtains the local voltage over-limit ratio prediction label and local average voltage deviation prediction label of each inverter intelligent agent at each target time.
[0104] S63: Calculate the auxiliary loss based on the actual label of the local voltage over-limit ratio, the predicted label of the local voltage over-limit ratio, the actual label of the local average voltage deviation, and the predicted label of the local average voltage deviation for each inverter agent at each target time.
[0105] In this embodiment, the control device performs loss calculations based on the actual local voltage exceedance ratio label, the predicted local voltage exceedance ratio label, the actual local average voltage deviation label, and the predicted local average voltage deviation label of each inverter agent at each target time, to obtain the sub-auxiliary loss of each inverter agent at each target time. The sub-auxiliary loss is as follows:
[0106] In the formula, For the first t The target time of the first i Sub-auxiliary loss of each inverter agent The first loss weighting coefficient, For the first t The target time of the first i The mean squared error loss value calculated from the actual label of the local voltage over-limit ratio and the predicted label of the local voltage over-limit ratio of each inverter agent. For the first t The target time of the first i Real-time labeling of the local voltage over-limit ratio of each inverter smart agent. For the first t The target time of the first i Local voltage over-limit prediction label for each inverter agent. This is the second loss weighting coefficient. For the first t The target time of the first i The mean square error loss value calculated from the actual local average voltage deviation label and the predicted local average voltage deviation label of each inverter agent. For the first t The target time of the first i The local average voltage deviation of each inverter agent is accurately labeled. For the first t The target time of the firsti Local average voltage deviation prediction label for each inverter agent.
[0107] The control device obtains the sub-auxiliary losses of each inverter agent at each target time and averages or sums them to obtain the auxiliary loss. By introducing two auxiliary tasks, "local voltage over-limit ratio prediction" and "local average voltage deviation prediction", the shared encoder's ability to represent current risks and short-term dynamic trends can be enhanced, thereby improving training stability and policy convergence speed.
[0108] S64: Calculate the main loss based on the global value estimate at each target time to obtain the main control loss; update the shared policy network based on the auxiliary loss and the main control loss to obtain the updated first shared policy network; update the action policy network of each inverter agent based on the updated first shared policy network to obtain several target inverter agents.
[0109] In this embodiment, the control device performs main loss calculation based on the global value estimate at each target time to obtain the main control loss. Specifically, the control device performs average summation based on the global value estimate at each target time to obtain the average summation result as the main control loss.
[0110] The control device updates the shared policy network based on the auxiliary loss and the main control loss to obtain the updated first shared policy network. Specifically, the control device accumulates the auxiliary loss and the main control loss to obtain a first total loss, and performs backpropagation based on the obtained first total loss to update the shared policy network to obtain the updated first shared policy network.
[0111] The control device updates the action policy network of each inverter agent according to the updated first shared policy network, and obtains several target inverter agents.
[0112] The agent voltage control network also includes a target policy network and a target value network. Please refer to [link / reference]. Figure 8 , Figure 8 The flowchart of step S6 in the voltage control method for a power distribution network provided in another embodiment of this application includes steps S65 to S67, as follows: S65: Based on the local observation data, local attention bias matrix, and target policy network of each inverter agent at each target time in the training samples at each target time, perform action reasoning for the next time step to obtain the predicted action space data for each target time step.
[0113] In this embodiment, the control device performs action reasoning for the next time step based on the local observation data, local attention bias matrix, and target policy network of each inverter agent at each target time step in the training samples at each target time step, thereby obtaining the predicted action space data for each target time step. The predicted action space data includes the action space data of each inverter agent at the next time step.
[0114] S66: Based on the predicted action space data at each target time, the global reward value at each target time in the training samples, and the target value network, value estimation is performed to obtain the target value estimate at each target time; based on the target value estimate at each target time and the global value estimate, value network loss is calculated to obtain the value network loss.
[0115] In this embodiment, the control device performs value estimation based on the predicted action space data at each target time, the global reward value at each target time in the training samples, and the target value network to obtain the target value estimate at each target time. The target value estimate is as follows:
[0116] In the formula, For the first t Target value estimation at each target time. For the first t The global reward value at each target time. For preset coefficients, For the first t The termination marker at each target time. For parameters The global value estimate output by the target value network. For the first t Local observation data of the inverter agent at the target time +1, For the first t Predicted action space data for each target time point.
[0117] The control device calculates the value network loss based on the target value estimates at each target time and the global value estimate, and obtains the value network loss, wherein the value network loss is:
[0118] In the formula, For the loss of value network, This is the loss function for the value network.
[0119] S67: Based on the value network loss and auxiliary loss, update the target policy network and target value network to obtain the updated target policy network and updated target value network; based on the updated target policy network and updated target value network, update the first shared policy network and main value network respectively to obtain the updated second shared policy network and updated main value network; based on the updated second shared policy network, update the action policy network of each inverter agent to obtain several target inverter agents.
[0120] In this embodiment, the control device updates the target policy network and the target value network based on the value network loss and the auxiliary loss, obtaining the updated target policy network and the updated target value network. Specifically, the control device accumulates the auxiliary loss and the value network loss to obtain a second total loss, and performs backpropagation based on the obtained second total loss to update the target policy network and the target value network, obtaining the updated target policy network and the updated target value network.
[0121] The control device updates the first shared policy network and the main value network according to the updated target policy network and the updated target value network, respectively, to obtain the updated second shared policy network and the updated main value network.
[0122] The control device updates the action policy network of each inverter agent according to the updated second shared policy network, and obtains several target inverter agents.
[0123] S7: Based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the set of physical prior matrices, perform voltage control of the distribution network.
[0124] In this embodiment, the control device performs voltage control of the distribution network based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the physical prior matrix set. Action reasoning and execution are performed based on the state space data of each inverter agent and the physical prior matrix set reflecting the topology of the distribution network. The inverter agents are then updated based on the reward values calculated after the action execution, thus improving the accuracy and efficiency of voltage control.
[0125] Please refer to Figure 9 , Figure 9This is a schematic diagram of the structure of a voltage control device for a distribution network according to an embodiment of this application. This device can be implemented in whole or in part through software, hardware, or a combination of both. The voltage control device 9 for the distribution network includes: The action reasoning module 91 is used to obtain the action space data of each inverter agent at each target time based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents of the distribution network, and the physical prior matrix set. The reward calculation module 92 is used to perform voltage control and reward calculation based on the action space data of each inverter agent at each target time, so as to obtain the global reward value of the distribution network at each target time. The data processing module 93 is used to combine the global state space data, global reward value, local observation data of each inverter agent, target node embedded representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time to construct training samples for each target time step. The global value estimation module 94 is used to splice and map the target node embedding representation of each inverter agent at each target time and the action space data of the same inverter agent in the training samples to obtain the agent global input feature set at each target time, wherein the agent global input feature set includes the agent global input features of each inverter agent. The global input feature set of the agent at each target time and the global physical bias matrix of the agent in the physical prior matrix set are input into the graph encoder in the main value network. The graph encoder is processed through multiple attention layers, residual connection layers, feedforward networks and normalization layers to obtain the global encoding result at each target time. The global encoding results at each target time are aggregated to obtain the global value estimate at each target time. The agent update module 95 is used to update each inverter agent based on the global value estimate at each target time, and obtain several target inverter agents. The voltage control module 96 is used to perform voltage control of the distribution network based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the physical prior matrix set.
[0126] In this embodiment, the action reasoning module obtains the action space data of each inverter agent at each target time based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents in the distribution network, and the physical prior matrix set. The reward calculation module performs voltage control and reward calculation based on the action space data of each inverter agent at each target time to obtain the global reward value of the distribution network at each target time. The data processing module combines the global state space data, global reward value, local observation data of each inverter agent, target node embedding representation, action space data, global state space data of the distribution network at the next target time, and local observation data of each inverter agent at the next target time to construct training samples for each target time. The global value estimation module combines the target node embedding representation of each inverter agent at each target time with the target value of the same inverter agent in the training samples. The action space data of the agents are spliced and mapped to obtain the global input feature set of the agents at each target time. The global input feature set of the agents includes the global input features of each inverter agent. The global input feature set of the agents at each target time and the global physical bias matrix of the agents in the physical prior matrix set are input into the graph encoder in the main value network. The graph encoder is processed through multiple attention layers, residual connection layers, feedforward networks and normalization layers to obtain the global encoding result of each target time. The global encoding results of each target time are aggregated to obtain the global value estimate of each target time. Through the agent update module, each inverter agent is updated according to the global value estimate of each target time to obtain several target inverter agents. Through the voltage control module, the voltage control of the distribution network is performed based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent and the physical prior matrix set. Action reasoning and execution are performed based on the state space data of each inverter agent and the set of physical prior matrices reflecting the topology of the distribution network. The inverter agents are then updated by combining the reward values of each inverter agent calculated after the action execution, which improves the accuracy and efficiency of voltage control.
[0127] Please refer to Figure 10 , Figure 10 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. The computer device 10 includes: a processor 101, a memory 102, and a computer program 103 stored in the memory 102 and executable on the processor 101. The computer device can store multiple instructions, which are adapted to be loaded and executed by the processor 101. Figures 1 to 8For the method steps and specific execution process, please refer to [link / reference]. Figures 1 to 8 Specific details will not be elaborated here.
[0128] The processor 101 may include one or more processing cores. The processor 101 connects to various parts within the server using various interfaces and lines. It executes instructions, programs, code sets, or instruction sets stored in the memory 102, and retrieves data from the memory 102 to perform various functions and process data for the voltage control device 9 of the power distribution network. Optionally, the processor 101 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 101 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for the touch screen; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 101.
[0129] The memory 102 may include random access memory (RAM) or read-only memory. Optionally, the memory 102 may include a non-transitory computer-readable storage medium. The memory 102 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 102 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch instructions), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. Optionally, the memory 102 may also be at least one storage device located remotely from the aforementioned processor 101.
[0130] This application embodiment also provides a storage medium that can store multiple instructions, which are adapted to be loaded and executed by a processor as described above. Figures 1 to 8For the method steps and specific execution process, please refer to [link / reference]. Figures 1 to 8 Specific details will not be elaborated here.
[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0132] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the algorithm. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0134] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0136] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0137] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms.
[0138] This invention is not limited to the above-described embodiments. If any modifications or variations to this invention do not depart from the spirit and scope of this invention, and if such modifications and variations fall within the scope of the claims and equivalent technologies of this invention, then this invention also intends to include such modifications and variations.
Claims
1. A voltage control method for a power distribution network, characterized in that, Includes the following steps: Based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents in the distribution network, and the local attention bias matrix of each inverter agent in a preset physical prior matrix set, the target node embedding representation and action space data of each inverter agent at each target time are obtained. The distribution network is configured with an agent update network, which includes a main value network and a graph encoder. The local attention bias matrix is used to characterize the physical association strength between nodes within the local region of each inverter agent. Voltage control and reward calculation are performed based on the action space data of each inverter agent at each target time to obtain the global reward value of the distribution network at each target time. The global state space data, global reward value, local observation data of each inverter agent, target node embedding representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time are combined to construct training samples for each target time step. The target node embedding representation and action space data of each inverter agent at each target time are spliced and mapped to obtain the agent global input feature set at each target time, wherein the agent global input feature set includes the agent global input features of each inverter agent. The global input feature set of the agent at each target time and the global physical bias matrix of the agent in the physical prior matrix set are input into the graph encoder in the main value network for encoding processing to obtain the global encoding result at each target time; the global encoding results at each target time are aggregated to obtain the global value estimate at each target time. Based on the global value estimate at each target moment, each inverter agent is updated to obtain several target inverter agents. Based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the set of physical prior matrices, voltage control of the distribution network is performed.
2. The voltage control method for a power distribution network according to claim 1, characterized in that, It also includes the step of: constructing the set of physical prior matrices; The construction of the physical prior matrix set includes the following steps: Obtain the physical prior information of the distribution network, wherein the physical prior information includes the distribution network topology, the bus location of each inverter, the local area division results of the distribution network, and the power flow calculation results; the power flow calculation results include the reactive power-voltage sensitivity of each node in the local area of each inverter to the corresponding inverter. Based on the bus locations of each inverter, a first adjacency matrix and a second adjacency matrix are constructed. The first adjacency matrix is normalized to obtain the topological adjacency matrix of the intelligent agent. The second adjacency matrix is normalized to obtain the sensitivity adjacency matrix of the intelligent agent. The first adjacency matrix includes the topological proximity of each inverter's bus to other inverters' buses in the distribution network. The second adjacency matrix includes the intensity of the impact of reactive power changes of each inverter on the voltage of other inverters' buses. The local attention bias matrix of each inverter agent is constructed by using the local topology bias matrix and local sensitivity bias matrix of each inverter agent. Based on the power distribution network topology, a local topology matrix is constructed for each inverter, which includes the topological proximity between nodes within a local area; the local topology matrix of each inverter is normalized to obtain the local topology bias matrix of each inverter agent. Based on the power flow calculation results, a local sensitivity matrix for each inverter is constructed. The local sensitivity matrix includes the sensitivity response of each node in the local region to the corresponding inverter. The local sensitivity matrix of each inverter is normalized to obtain the local sensitivity bias matrix of each inverter agent. The global physical bias matrix of the agent is constructed based on the agent's topological adjacency matrix and the agent's sensitivity adjacency matrix, thus obtaining the agent's global physical bias matrix.
3. The voltage control method for a power distribution network according to claim 2, characterized in that, The method of obtaining the target node embedding representation and action space data of each inverter agent at each target time based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents of the distribution network, and the local attention bias matrix of each inverter agent in a preset physical prior matrix set, includes the following steps: Linear mapping and region identifier embedding are performed on the node feature vectors of each node in the local observation data of each inverter agent at each target time to obtain the node feature representation of each inverter agent at each target time, wherein the node feature representation includes the node feature representation of each node. The node feature representations and local attention bias matrices of each inverter agent at each target time are input into a preset local encoder, and then processed sequentially through multiple attention layers, residual connection layers, feedforward networks and normalization layers of the local encoder to obtain the local encoding results of each inverter agent at each target time. Obtain the target node position index of each inverter agent; based on the target node position index of each inverter agent, extract the target node embedding from the corresponding local encoding result to obtain the target node embedding representation of each inverter agent at each target time, wherein the target node embedding representation characterizes the local region state observed by the current inverter agent from its inverter perspective. The target node of each inverter agent at each target time is embedded into a preset shared strategy header for time-series modeling and normalized action generation, thereby obtaining the normalized actions of each inverter agent at each target time as the action space data.
4. The voltage control method for a power distribution network according to claim 3, characterized in that, The step of performing voltage control and reward calculation based on the action space data of each inverter agent at each target time to obtain the global reward value of the distribution network at each target time includes the following steps: The action is executed based on the inverter action space data set at each target time, and the voltage control execution index of the distribution network at each target time is obtained. The reward is calculated based on the voltage control execution index of the distribution network at each target time, and the global reward value of the distribution network at each target time is obtained.
5. The voltage control method for a power distribution network according to claim 4, characterized in that, The step of executing actions based on the inverter action space data set at each target time to obtain the voltage control execution index of the distribution network at each target time includes the following steps: The maximum adjustable reactive power range is calculated based on the distributed new energy active power vector of each inverter agent at each target time and the rated apparent power limit, thus obtaining the maximum adjustable reactive power range of each inverter agent at each target time. Based on the action space data of each inverter intelligent agent at each target time and the maximum adjustable reactive power range, the command mapping is performed to obtain the reactive power control command of each inverter intelligent agent at each target time; based on the reactive power control command of each inverter intelligent agent at each target time, the action is executed to obtain the voltage control execution index of the distribution network at each target time.
6. The voltage control method for a power distribution network according to claim 5, characterized in that: The voltage control execution indicators include the voltage amplitude of several voltage monitoring nodes and the reactive power output vector of each inverter. The step of calculating rewards based on the voltage control execution indicators of the distribution network at each target time to obtain the global reward value of the distribution network at each target time includes the following steps: Voltage over-limit penalty is calculated based on the voltage amplitude of each voltage detection node at each target time to obtain the voltage over-limit penalty value of each voltage detection node at each target time; reactive power regulation loss is calculated based on the reactive power output vector of each inverter at each target time and the preset weight coefficient to obtain the reactive power regulation loss value of each inverter at each target time, wherein the weight coefficient is used to balance the voltage control effect and the reactive power regulation cost. The voltage amplitude of each voltage detection node at the same target time is subtracted from the reactive power regulation loss value of each inverter to obtain the local reward value of each voltage detection node and each inverter at each target time. The global reward value of the distribution network at each target time is obtained by weighted averaging the local reward values of each voltage detection node and each inverter at each target time.
7. The voltage control method for a power distribution network according to claim 6, characterized in that: The agent update network also includes a shared policy network; the model parameters of the action policy networks of all inverter agents call the model parameters of the shared policy network. The step of updating each inverter agent based on the global value estimate at each target time to obtain several target inverter agents also includes the following steps: Labels are constructed based on the voltage amplitude of each node in the local area of each inverter agent in the training samples at each target time, and the label data of the training samples at each target time is obtained. The label data includes the true label of the local voltage over-limit ratio and the true label of the local average voltage deviation of each inverter agent. Obtain the target node embedding representation of each inverter agent at each target time; input the target node embedding representation of each inverter agent at each target time into the preset local voltage over-limit ratio prediction head and average voltage deviation prediction head for prediction, and obtain the local voltage over-limit ratio prediction label and local average voltage deviation prediction label of each inverter agent at each target time. The auxiliary loss is obtained by calculating the loss based on the true label of the local voltage over-limit ratio, the predicted label of the local voltage over-limit ratio, the true label of the local average voltage deviation, and the predicted label of the local average voltage deviation of each inverter agent at each target time. The main loss is calculated based on the global value estimate at each target time to obtain the main control loss; the shared policy network is updated based on the auxiliary loss and the main control loss to obtain the updated first shared policy network; the action policy network of each inverter agent is updated based on the updated first shared policy network to obtain several target inverter agents.
8. The voltage control method for a power distribution network according to claim 7, characterized in that: The agent update network also includes a target policy network and a target value network; The step of updating each inverter agent based on the training samples and the set of physical prior matrices at each target time to obtain several target inverter agents includes the following steps: Based on the local observation data, local attention bias matrix, and target policy network of each inverter agent at each target time in the training samples at each target time, the action reasoning of the next time step is performed to obtain the predicted action space data of each target time step, wherein the predicted action space data includes the action space data of each inverter agent at the next time step. Value estimation is performed based on the predicted action space data at each target time, the global reward value at each target time in the training samples, and the target value network to obtain the target value estimate at each target time; value network loss is calculated based on the target value estimate at each target time and the global value estimate to obtain the value network loss. Based on the value network loss and auxiliary loss, the target policy network and target value network are updated to obtain the updated target policy network and updated target value network; based on the updated target policy network and updated target value network, the first shared policy network and main value network are updated respectively to obtain the updated second shared policy network and updated main value network; based on the updated second shared policy network, the action policy network of each inverter agent is updated to obtain several target inverter agents.
9. A voltage control device for a power distribution network, characterized in that, include: The action reasoning module is used to obtain the target node embedding representation and action space data of each inverter agent at each target time based on the global state space data of the distribution network at each target time within a preset time period, the local observation data of several inverter agents of the distribution network, and the local attention bias matrix of each inverter agent in a preset physical prior matrix set. The distribution network is configured with an agent update network, which includes a main value network; the main value network includes a graph encoder; and the local attention bias matrix is used to characterize the physical association strength between nodes within the local region of each inverter agent. The reward calculation module is used to perform voltage control and reward calculation based on the action space data of each inverter agent at each target time, so as to obtain the global reward value of the distribution network at each target time. The data processing module is used to combine the global state space data, global reward value, local observation data of each inverter agent, target node embedded representation, action space data, global state space data of the distribution network at the next time step of the target time, and local observation data of each inverter agent at the next time step of the target time to construct training samples for each target time step. The global value estimation module is used to concatenate and map the target node embedding representation of each inverter agent at each target time and the action space data of the same inverter agent in the training samples to obtain the global input feature set of the agent at each target time. The global input feature set of the agent includes the global input features of each inverter agent. The global input feature set of the agent at each target time and the global physical bias matrix of the agent in the physical prior matrix set are input into the graph encoder in the main value network. The graph encoder is processed through multiple attention layers, residual connection layers, feedforward networks and normalization layers to obtain the global encoding result at each target time. The global encoding results at each target time are aggregated to obtain the global value estimate at each target time. The agent update module is used to update each inverter agent based on the global value estimate at each target time, and obtain several target inverter agents. The voltage control module is used to perform voltage control of the distribution network based on the global state space data of the distribution network at the current time after the preset time period, the local observation data of each target inverter agent, and the physical prior matrix set.
10. A computer device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the voltage control method for a power distribution network as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Distributed real-time control method for voltage of power distribution network
CN115333152A
Method and apparatus for controlling voltage of distributed photovoltaic power distribution network
WO2018214810A1