A power distribution network intelligent agent cooperative simulation method, device, equipment and storage medium
By combining physical information neural networks and Monte Carlo tree search algorithms, a collaborative simulation method for distribution network intelligent agents is constructed, which solves the problems of high computational complexity and low reliability in distribution network simulation and achieves efficient and accurate distributed resource scheduling decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-31
AI Technical Summary
Existing power distribution network simulation technologies suffer from high computational complexity, simulation results that do not conform to physical laws, and low reliability when dealing with the high uncertainty of distributed energy resources and multi-entity interactions.
A collaborative simulation method for power distribution network agents is constructed by combining a physical information neural network with a multi-agent interaction model and a Monte Carlo tree search algorithm. The trained physical information neural network predicts the global state, the multi-agent interaction model generates a candidate scheduling decision set, and the Monte Carlo tree search algorithm optimizes the decision to ensure that the simulation results conform to physical laws.
It achieves efficient and reliable power distribution network simulation, and can process the scheduling decisions of distributed resources in real time in large-scale scenarios, improving simulation efficiency and the accuracy of results, and adapting to complex dynamic scenarios.
Smart Images

Figure CN122495556A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of intelligent simulation technology of power systems, and in particular to a method, device, equipment and storage medium for collaborative simulation of intelligent agents in distribution networks. Background Technology
[0002] Distribution network simulation technology is a core supporting tool for power system planning, operation, and security analysis. With the increasing penetration of distributed energy sources (such as photovoltaic power generation and energy storage systems) in distribution networks, the operating characteristics of distribution networks are becoming increasingly complex, exhibiting strong uncertainty, high-dimensional dynamics, and multi-agent interactions. This places higher demands on simulation technology. Existing simulation technologies include traditional numerical simulation methods, pure data-driven simulation methods, and traditional multi-agent simulation methods.
[0003] However, traditional numerical simulation methods are computationally complex and slow to solve problems under large-scale power grids; purely data-driven simulation methods lack physical constraints, which can easily lead to low reliability of simulation results; while traditional multi-agent simulation methods can model the interaction relationships between distributed resources, their decision-making process often ignores the physical constraints of the power grid, resulting in infeasible or unsafe scheduling strategies. Therefore, how to construct a simulation and deduction technique with high reliability, high efficiency, and conformity to physical laws is an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a method, apparatus, equipment, and storage medium for collaborative simulation of distribution network intelligent agents, in order to solve the problems of low technical efficiency, low reliability, and non-compliance with physical laws in the simulation and deduction of distribution networks in the prior art.
[0005] According to one aspect of the present invention, a method for collaborative simulation of intelligent agents in a power distribution network is provided, the method comprising: The current state vector of the distribution network is input into a trained physical information neural network to obtain the global predicted state of the distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the distribution network. Based on the global predicted state, at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step is output through a multi-agent interaction model; wherein, each agent in the multi-agent interaction model corresponds to a distributed resource in the distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next time step and a state set corresponding to the joint scheduling action. The Monte Carlo tree search algorithm is used, combined with the physical information neural network and the multi-agent interaction model, to optimize the candidate scheduling decision set and obtain the target scheduling decision set. The scheduling decisions in the target scheduling decision set are used as the target scheduling decisions for each agent.
[0006] According to another aspect of the present invention, a power distribution network intelligent agent collaborative simulation device is provided, the device comprising: The input module is used to input the current state vector of the distribution network into the trained physical information neural network to obtain the global predicted state of the distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the distribution network. The candidate module is used to output at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step based on the global predicted state through a multi-agent interaction model; wherein each agent in the multi-agent interaction model corresponds to a distributed resource in the distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next time step and a state set corresponding to the joint scheduling action. The optimization module is used to optimize the candidate scheduling decision set by employing the Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, to obtain the target scheduling decision set. The determination module is used to take the scheduling decisions in the target scheduling decision set as the target scheduling decisions for each agent.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the power distribution network intelligent agent collaborative simulation method according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the power distribution network intelligent agent collaborative simulation method according to any embodiment of the present invention.
[0009] This invention discloses a method, apparatus, device, and storage medium for collaborative simulation of power distribution networks using intelligent agents. The method includes: inputting the state vector of the power distribution network at the current moment into a trained physical information neural network to obtain the global predicted state of the power distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the power distribution network; based on the global predicted state, outputting at least one candidate scheduling decision set for all distributed resources in the power distribution network at the next moment through a multi-agent interaction model; wherein each agent in the multi-agent interaction model corresponds to a distributed resource in the power distribution network, and each candidate scheduling decision set includes joint scheduling actions that the multiple agents can execute at the next moment and the state set corresponding to the joint scheduling actions; using a Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, to optimize the candidate scheduling decision set to obtain a target scheduling decision set; and using the scheduling decisions in the target scheduling decision set as the target scheduling decisions for each agent. This method, by combining a physical information neural network and a Monte Carlo tree search algorithm, can accurately achieve real-time simulation of the power distribution network, solving the problems of low technical efficiency, low reliability, and non-compliance with physical laws in existing power distribution network simulations.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a collaborative simulation method for power distribution network intelligent agents provided in Embodiment 1 of the present invention; Figure 2 A system architecture diagram of a cooperative simulation method for power distribution network intelligent agents provided in an embodiment of the present invention; Figure 3 A schematic diagram comparing the voltage prediction accuracy of the method of this embodiment and the baseline method is provided for an embodiment of the present invention. Figure 4 A schematic diagram comparing the computational efficiency of the method of this embodiment of the invention with a baseline method, provided for an embodiment of the invention; Figure 5 A graph showing the relationship between computation time and the number of agents provided in an embodiment of the present invention; Figure 6 A schematic diagram comparing the voltage over-limit response time of the method of this embodiment and the baseline method provided for this embodiment of the invention; Figure 7 A schematic diagram illustrating the robustness comparison between the method of this embodiment and a baseline method in a noisy environment, provided for an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a power distribution network intelligent agent collaborative simulation device provided in Embodiment 2 of the present invention; Figure 9 This is a schematic diagram of the electronic device used in the power distribution network intelligent agent collaborative simulation method according to an embodiment of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention. It should be understood that the various steps described in the method embodiments of the present invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, any variations of the terms "comprising" and "having," etc., are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0018] Currently, there are three main technical solutions in the field of power distribution network simulation, but all of them have obvious technical defects: (1) Traditional numerical simulation methods Traditional numerical simulation methods, such as the forward backward substitution method, calculate the state of the distribution network by iteratively solving nonlinear power flow equations. The computational complexity of this type of method is typically O(n^2). 2 At the [level] level, as the number of nodes and simulation scenarios increase, the computation time grows dramatically, making it difficult to meet the real-time simulation requirements of highly dynamic and multi-scenario applications. Furthermore, traditional methods lack the ability to adaptively handle the random fluctuations of distributed energy resources, making it difficult to effectively cope with dynamic scenarios such as sudden changes in photovoltaic output.
[0019] (2) Pure data-driven simulation method Purely data-driven methods, such as Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNNs), offer improved computational speed. However, due to the lack of explicit constraints on the physical laws of power grids (such as Kirchhoff's laws and nodal power balance equations), they are prone to simulation distortion and overfitting in scenarios with scarce or out-of-distribution data. Furthermore, the outputs of purely data-driven methods may violate fundamental physical laws, severely limiting the reliability and practicality of simulation results.
[0020] (3) Traditional multi-agent simulation method While existing multi-agent simulation methods can model the interactions between distributed resources, their decision-making processes often ignore the physical constraints of the power grid, making the scheduling strategies generated by the agents infeasible or unsafe in real power grids. Furthermore, traditional multi-agent methods are inadequate in handling uncertainty, lacking effective quantification and robust mechanisms to cope with random disturbances.
[0021] In recent years, Physically Informed Neural Networks (PINNs) have addressed the physical consistency problem to some extent by embedding the control equations into the loss function. However, existing PINN methods still fall short in high-dimensional uncertain decision-making and multi-agent cooperative optimization, particularly in handling distributed resource scheduling problems with stochastic and game-theoretic characteristics. While Monte Carlo Tree Search (MCTS) excels at sequential optimization in complex decision spaces, its traditional form relies on accurate state transition models, making it difficult to apply directly in physically constrained environments like power systems.
[0022] Therefore, there is an urgent need for a power distribution network simulation and deduction method that can deeply integrate physical mechanism constraints and data-driven intelligence, and at the same time have the ability of multi-agent collaborative decision-making, in order to solve the above-mentioned technical problems.
[0023] Example 1 Figure 1 This is a flowchart illustrating a distribution network intelligent agent collaborative simulation method provided in Embodiment 1 of the present invention. This method is applicable to simulating the actual state of a distribution network. The method can be executed by a distribution network intelligent agent collaborative simulation device, which can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes, but is not limited to, devices such as computers.
[0024] like Figure 1 As shown in Embodiment 1 of the present invention, a collaborative simulation method for power distribution network intelligent agents includes the following steps: S110. Input the state vector of the distribution network at the current moment into the trained physical information neural network to obtain the global predicted state of the distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the distribution network.
[0025] The state vector can be a set of physical quantities representing the operating state of the distribution network. For example, the state vector may include parameters such as node voltage magnitude, active power, reactive power, branch power, and switch status. The physical information neural network can be a deep learning model that integrates prior physical knowledge with neural networks. The physical information neural network embeds physical laws such as the power flow equations, electrical constraints, and dynamic equations of the distribution network. The global predicted state can be the global state vector of the distribution network predicted by the physical information neural network. The distribution network branch power flow (DistFlow) equations can be a set of algebraic equations describing the physical relationship between the voltage at both ends of a branch and the power transmitted through the branch in the distribution network.
[0026] In this embodiment, the state vector of the distribution network at the current moment can be input into a trained physical information neural network to obtain the global predicted state of the distribution network at the next moment.
[0027] For example, the distribution network can be defined in time. t The state vector below is ,in For time t The node voltage amplitude (unit: pu). Indicates time t The active power below, Indicates time t The reactive power (units: MW and Mvar) is given. The distribution network behavior described by the DistFlow equations can be formalized into partial differential equations: ; in, For nonlinear functions of the dynamic behavior of the distribution network, Voltage amplitude, Active power Reactive power This is the parameter vector of the physical information neural network, i.e., the weights of the neural network.
[0028] In one embodiment, training the physical information neural network includes: training an initial physical information neural network based on historical operating data of the distribution network to obtain a predicted state; determining the current loss value of the initial physical information neural network based on the predicted state and a fusion loss function; when the loss value does not meet the training termination condition, adjusting the network parameters of the initial physical information neural network based on the loss value to obtain an updated physical information neural network and continuing training; when the loss value of the updated physical information neural network meets the training termination condition, using the updated physical information neural network as the trained physical information neural network.
[0029] Historical operating data can refer to the operating data of the distribution network over a historical period, and the predicted state can refer to the predicted state of the distribution network. The initial physical information neural network can refer to an untrained physical information neural network. The fusion loss function can be a loss function that includes both data-driven and physical constraint loss terms. The training termination condition can refer to the conditions for ending training, and can be set according to actual conditions. For example, the training termination condition can be set to the loss value being less than a preset loss value, or the loss value changing less than a preset value during a preset number of iterations. Network parameters can be the parameters of the physical information neural network.
[0030] In this embodiment, the initial physical information neural network can be trained using historical operating data of the distribution network to obtain the predicted state. Based on the predicted state and the fusion loss function, the current loss value of the initial physical information neural network can be determined. If the loss value does not meet the training termination condition, the network parameters of the initial physical information neural network are adjusted based on the loss value to obtain the updated physical information neural network and continue to train and calculate the loss value. If the loss value of the updated physical information neural network meets the training termination condition, the updated physical information neural network can be used as the trained physical information neural network.
[0031] In one embodiment, the fusion loss function for: ; ; ; in, For data-driven loss terms, The number of training samples. For the distribution network in the first The actual state of each physical constraint configuration point For the distribution network in the first Predicted state of each physical constraint configuration point; For physical constraint loss terms, The weights of the physical constraint loss term, The number of physical constraint configuration points in the distribution network. For time, For nonlinear functions of the dynamic behavior of the distribution network, For the distribution network in the first Predicted voltage amplitude at each physically constrained configuration point For the distribution network in the first Predicted active power for each physically constrained configuration point For the distribution network in the first Predicted reactive power at each physically constrained configuration point These are the network parameters of the physical information neural network; The update equation for the network parameters is: ; in, The learning rate of the physical information neural network. The iteration number of the physical information neural network. This is to fuse the gradient of the loss function with respect to the network parameters.
[0032] In this embodiment, the physical information neural network can use state variables and timet As input, output the global predicted state at the next time step. Network parameters can be minimized by the fusion loss function. Learning is conducted, in which data-driven loss terms are used. It can be defined as the mean squared error between the predicted value and the actual data, and the physical constraint loss term. It can be defined as the sum of the mean squares of the residuals of the DistFlow equation.
[0033] In the physical constraint loss term, for the physical constraint configuration point in the distribution network physical functions This can be further elaborated as follows: ; in, Configure points for physical constraints The equivalent capacitance reflects voltage inertia; Configure points for physical constraints voltage amplitude, Configure points for physical constraints active power, Configure points for physical constraints reactive power, Configure points for physical constraints The set of neighbors; In time t Injected active power at the physical constraint configuration point. In time t Active power of the load at the physical constraint configuration point; In time t Time from physical constraint configuration point Flow to adjacent physical constraint configuration points active power, Satisfying the line power flow equation , Configure points for physical constraints voltage amplitude, Configure points for physical constraints and The electrical conductance of the lines between them Configure points for physical constraints and The susceptance of the line between them Configure points for physical constraints and The voltage phase angle difference between them; for the reactive power part It has active power A similar structure. The equation is essentially a dynamic expression of Kirchhoff's current law and power balance equation in the time domain. By embedding a loss function in PINNs, it can be ensured that the output of the neural network always satisfies the physical laws of the power grid.
[0034] When training a physical information neural network (PEN), historical operating data can be used for pre-training to learn the dynamic behavior of the distribution network under standard operating conditions. Then, transfer learning can be used to fine-tune the network parameters to adapt to specific scenarios. The network parameters can be updated using gradient descent. ; in, The learning rate of the physical information neural network. The iteration number of the physical information neural network. This is to fuse the gradient of the loss function with respect to the network parameters.
[0035] S120. Based on the global predicted state, output at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step through a multi-agent interaction model; wherein, each agent in the multi-agent interaction model corresponds to a distributed resource in the distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next time step and a state set corresponding to the joint scheduling action.
[0036] The multi-agent interaction model refers to a model describing the interaction between multiple agents in a power distribution network. This model includes multiple agents, each corresponding to a distributed resource in the network. Distributed resources can refer to resources such as photovoltaic systems, energy storage systems, and charging piles. The candidate scheduling decision set can contain joint scheduling actions that multiple agents can execute in the next time step, as well as the corresponding set of states. The joint scheduling action can refer to a combination of actions, and the state set can refer to a combination of states.
[0037] In this embodiment, the global predicted state can be input into the multi-agent interaction model to output at least one candidate scheduling decision set for all agents (i.e., all distributed resources in the distribution network) at the next time step.
[0038] In one embodiment, the step of outputting at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step based on the global predicted state through a multi-agent interaction model includes: obtaining the current state of each agent; inputting the current state of each agent and the global predicted state into the multi-agent interaction model to obtain multiple candidate action combinations; each candidate action combination contains the predicted actions of all agents at the next time step; inputting each candidate action combination and the current state into a stochastic differential equation to obtain a state combination corresponding to each candidate action combination; the state combination contains the predicted states of all agents at the next time step; and constructing at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step based on each candidate action combination and the corresponding state combination.
[0039] Here, "current state" refers to the agent's current real-time situation. For example, if the agent is a photovoltaic system, the state could include current power generation, remaining battery charge, and equipment temperature. "Action" refers to the operation the agent needs to perform, such as how much electricity the photovoltaic system generates, how much energy storage to charge / discharge, or how much load to adjust. "Candidate action combination" refers to a combination of all agent candidate actions, with each combination containing the predicted actions for all agents in the next time step. "State combination" refers to a combination of the states corresponding to all agent candidate actions, with each state combination containing the predicted states for all agents in the next time step.
[0040] In this embodiment, the current state of each agent can be obtained, and the current state of each agent and the global predicted state can be input into the multi-agent interaction model to obtain multiple candidate action combinations for each agent. Each candidate action combination contains the predicted actions of all agents in the next time step. Each candidate action combination and the current state can be input into a stochastic differential equation to obtain the state combination corresponding to each candidate action combination. Based on each candidate action combination and the state combination corresponding to the candidate action combination, at least one candidate scheduling decision set for all distributed resources in the distribution network in the next time step can be constructed. Each candidate scheduling decision set contains a candidate action combination and the state combination corresponding to the candidate action combination.
[0041] For example, each distributed resource in the distribution network is modeled as an autonomous agent, and the set of agents is denoted as . ,in K The number of agents. Each agent... In time t The state is The action is Decision-making strategies of intelligent agents Local state and global predicted state Mapping to Action : ; in, The neural network parameterizes the output to satisfy physical constraints such as power limits (e.g., power balance constraints of the power grid, equipment not exceeding limits, etc.).
[0042] To handle environmental uncertainties, stochastic differential equations are introduced to describe the agent's state. Evolution: ; in, It is a drift term, representing deterministic evolution; It is a diffusion term, representing the intensity of random noise; It is the differential of the Wiener process, representing Gaussian white noise. This equation is approximated by Bayesian variational inference of PINNs to quantify the uncertainty in the derivation and ensure that the simulation remains robust in noisy environments.
[0043] S130. Using the Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, the candidate scheduling decision set is optimized to obtain the target scheduling decision set.
[0044] Among them, the Monte Carlo tree search algorithm can refer to a heuristic decision optimization algorithm based on sampling and iterative search. By simulating the execution of different decision sequences multiple times in the action space, iteratively updating the value evaluation of each state-action pair, and finally selecting the decision with the best long-term benefit from the candidate decisions.
[0045] In this embodiment, the Monte Carlo tree search algorithm can be used, combined with a physical information neural network and a multi-agent interaction model, to optimize the candidate scheduling decision set. The optimal candidate scheduling decision set is found from all candidate scheduling decision sets and used as the target scheduling decision set.
[0046] In one embodiment, the step of using the Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model to optimize the candidate scheduling decision set to obtain the target scheduling decision set includes: for each candidate scheduling decision set, using the Monte Carlo tree search algorithm to search the action space, and combining the physical information neural network and the multi-agent interaction model to determine the subsequent scheduling decision set corresponding to the candidate scheduling decision set; updating the action value of the candidate scheduling decision set based on the state-action pair value function, wherein the action value is used to guide the direction of searching the action space; evaluating the long-term value of the candidate scheduling decision set in the long run based on the candidate scheduling decision set and the corresponding subsequent scheduling decision set, combined with an objective function; wherein the objective function has the value of minimizing long-term simulation error as the optimization objective; and taking the candidate scheduling decision set with the best long-term value among all candidate scheduling decision sets as the target scheduling decision set.
[0047] In this context, the action space refers to the set of all legal scheduling actions that can be executed by all agents during the distribution network scheduling process. The subsequent scheduling decision set can be the set of scheduling decisions corresponding to the agents at times after the candidate scheduling decision set; there can be multiple subsequent scheduling decision sets. The state-action pair value function can be a function used to evaluate the value of actions; action value can refer to the action's score. The optimization objective of the objective function can be to minimize the value of long-term simulation error. Long-term value can refer to the value generated after the candidate scheduling decision set is executed in the distribution network over a long period. Long-term simulation error can refer to the sum of deviations between the distribution network operating state obtained from long-term extrapolation after executing the candidate scheduling decision set and the actual operating state.
[0048] In this embodiment, for each candidate scheduling decision set, a Monte Carlo tree search algorithm is used to search the action space. Combined with a physical information neural network and a multi-agent interaction model, the subsequent scheduling decision set corresponding to the candidate scheduling decision set is determined. For example, after obtaining the candidate scheduling decision set, it can be executed, and the state variables of the distribution network after execution are input into the physical information neural network to obtain the global predicted state. This global predicted state is then input into the multi-agent interaction model to obtain the actions and states of each agent after executing the candidate scheduling decision set, and the subsequent scheduling decision set is constructed. The process of constructing the subsequent scheduling decision set is similar to that of constructing the candidate scheduling decision set. During each construction of the subsequent scheduling decision set, the action value of the candidate scheduling decision set can be updated based on the state-action pair value function. This action value can be used to guide the direction of subsequent search of the action space. Subsequently, for each candidate scheduling decision set, after searching for a preset number of subsequent scheduling decision sets, the long-term value of the candidate scheduling decision set in the long run can be evaluated based on the candidate scheduling decision set and its corresponding subsequent scheduling decision set, combined with the objective function. Finally, the candidate scheduling decision set with the best long-term value among all candidate scheduling decision sets is taken as the target scheduling decision set.
[0049] In one embodiment, the objective function is: ; in, Indicates at time The expected value of the cumulative discount return, For expectation operator, To extrapolate the time domain, To deduce the first in the time domain Time step index As a discount factor, ; The value function for the state-action pair is: ; in, It is the value of the state-action pair. It's the learning rate. It's an instant reward. , It is to perform an action The next state after that, This is a possible action for the next state.
[0050] In this embodiment, in the inference optimization layer, a rolling time-domain optimization algorithm coupled with Monte Carlo tree search and a physical information neural network can be used. At each time step... t The objective function aims to minimize the long-term simulation error. ; in, Indicates at time The expected value of the cumulative discount return, Represents the expectation operator. It refers to the time domain of the deduction, i.e., the total number of time steps. To deduce the first in the time domain Time step index It is a discount factor, which is a scalar. It is used to weigh immediate and future errors.
[0051] Monte Carlo tree search algorithm selects the optimal agent policy by searching the action space, and the state-action pair value function. The update rules are as follows: ; in, It is the value of the state-action pair. It is the learning rate, which is a scalar. It is an immediate reward, defined as a negative simulation error, i.e. , It is to perform an action The next state after that, This represents the possible actions for the next state. This process is tightly integrated with PINNs: PINNs provide predicted states. The calculated loss is used as input to MCTS to ensure that the simulation is performed under physical constraints.
[0052] S140. The scheduling decisions in the target scheduling decision set are used as the target scheduling decisions for each agent.
[0053] In this embodiment, the target scheduling decision set includes the scheduling decisions of each agent, and the scheduling decisions in the target scheduling decision set can be used as the target scheduling decisions of each agent.
[0054] This invention provides a method for collaborative simulation of power distribution networks using intelligent agents, comprising: inputting the state vector of the power distribution network at the current moment into a trained physical information neural network to obtain the global predicted state of the power distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the power distribution network; based on the global predicted state, outputting at least one candidate scheduling decision set for all distributed resources in the power distribution network at the next moment through a multi-agent interaction model; wherein each agent in the multi-agent interaction model corresponds to a distributed resource in the power distribution network, and each candidate scheduling decision set includes joint scheduling actions that the multiple agents can execute at the next moment and the state set corresponding to the joint scheduling actions; using a Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, to optimize the candidate scheduling decision set to obtain a target scheduling decision set; and using the scheduling decisions in the target scheduling decision set as the target scheduling decisions for each agent. This method, by combining a physical information neural network and a Monte Carlo tree search algorithm, can accurately achieve real-time simulation of the power distribution network, solving the problems of low technical efficiency, low reliability, and non-compliance with physical laws in existing power distribution network simulations.
[0055] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.
[0056] In one embodiment, the step of inputting the state vector of the distribution network at the current moment into a trained physical information neural network to obtain the global predicted state of the distribution network at the next moment includes: inputting the state vector of the distribution network at the current moment into a graph neural network to obtain the embedding vector of the distribution network; the topology graph of the graph neural network is constructed based on the topology of the distribution network; and inputting the embedding vector into the trained physical information neural network to obtain the global predicted state of the distribution network at the next moment.
[0057] In this embodiment, a graph neural network topology graph can be constructed based on the distribution network topology. The current state vector of the distribution network is input into the graph neural network to obtain the embedding vector of the distribution network. This embedding vector is then input into a trained physical information neural network to obtain the global predicted state of the distribution network at the next time step. By introducing a graph neural network, neighbor information can be aggregated through a message passing mechanism, reducing the computational load.
[0058] For example, the distribution network topology can be modeled as a topology graph. ,in, For a set of nodes, Let be the set of edges. Graph neural networks can be used to aggregate neighbor information through message passing. The inter-layer update of a graph neural network is represented as: ; in, It is a topology node in the distribution network In the Layer embedding vectors, It is a topology node in the distribution network In the Layer embedding vectors, It is a topology node The neighborhood group, and It is a learnable function (such as a multilayer perceptron). It is the edge The eigenvectors (such as line impedance) can be used to construct the graph. Combining this graph structure with PINNs can significantly reduce computational complexity through distributed computing, enabling the method to scale linearly to scenarios with hundreds of nodes.
[0059] Furthermore, this embodiment can also introduce Lyapunov functions. To verify the stability of the deduction process, among which, Let be the state error vector. By analyzing its time derivative, we prove that the error converges asymptotically, ensuring the reliability of the method in dynamic environments.
[0060] This invention provides a method and system for collaborative simulation and optimization of distribution network agents based on physical information neural networks and Monte Carlo tree search. By embedding the DistFlow physical equations of the distribution network as constraints into the neural network training process, the physical rationality of the simulation output can be guaranteed. By modeling distributed resources as autonomous agents and using the MCTS algorithm to achieve multi-agent collaborative decision optimization, the uncertainty problem in the operation of the distribution network can be effectively handled. By introducing graph neural networks to accelerate state reasoning, real-time simulation of large-scale distribution networks can be realized.
[0061] Based on the technical solutions of the above embodiments, this invention provides several specific implementation methods.
[0062] As one specific implementation method of this embodiment. Figure 2 A system architecture diagram of a cooperative simulation method for power distribution network intelligent agents provided in an embodiment of the present invention is shown below. Figure 2 As shown, the overall architecture of the system adopts a three-layer integrated design, which consists of a physical modeling layer, an intelligent agent interaction layer, and an inference and optimization layer from bottom to top.
[0063] The physical modeling layer, located at the bottom of the architecture, is responsible for constructing Physical Information Neural Networks (PINNs) embedded with DistFlow equation constraints. It encodes the physical laws of the distribution network into the loss function of the neural network, outputting state predictions that satisfy the network's operational rules. The agent interaction layer, located in the middle layer, models distributed resources (such as photovoltaic systems and energy storage systems) as autonomous agents. Based on the global state predictions provided by PINNs, it performs local autonomous decision-making through a neural network-parameterized decision strategy. The inference and optimization layer, located at the top layer, uses the Monte Carlo tree search algorithm to achieve collaborative rolling time-domain optimization among multiple agents under physical constraints. These three layers are tightly coupled through data flow: the physical modeling layer outputs state predictions, the agent interaction layer generates behavioral decisions, and the inference and optimization layer achieves global collaboration through rolling optimization.
[0064] As a specific implementation method of this embodiment, the IEEE 33-node distribution network simulation is used as an example: (1) Experimental environment configuration: Experimental verification was carried out based on the IEEE 33-node standard distribution network model. The topology of this model includes 33 nodes and 32 lines, with a base voltage of 12.66kV and a total load of 3.715MW+2.3Mvar. The line parameters adopt the standard impedance matrix, and the node load is distributed according to the typical ratio of residential and industrial / commercial.
[0065] The distributed resource configuration is as follows: Nodes 6, 18, and 22 are configured with photovoltaic systems, with a total installed capacity of 1.2MW; Nodes 8 and 25 are configured with energy storage systems, each with a capacity of 0.5MWh. The simulation platform is built based on Python 3.8 and TensorFlow 2.4, with hardware configuration including an Intel Xeon CPU (2.4GHz, 14 cores), 128GB of RAM, and running Ubuntu 20.04. Experimental data comes from the actual operation records of a power grid from January 2020 to December 2021, including load, photovoltaic output, and energy storage dispatch data, with a sampling interval of 15 minutes. The training set and test set are randomly divided in a 7:3 ratio.
[0066] (2) Accuracy verification: The method of the present invention is compared with three baseline methods: traditional numerical method (forward and backward substitution), pure data-driven LSTM simulation method, and intelligent agent simulation method without physical constraints. Figure 3 A schematic diagram comparing the voltage prediction accuracy of the method of this embodiment of the invention with that of a baseline method is provided for an embodiment of the invention, as shown in the figure. Figure 3As shown, experimental results demonstrate that the root mean square error (RMSE) and mean absolute error (MAE) of the voltage prediction method of this invention are 0.021 pu and 0.015 pu, respectively, significantly outperforming the benchmark methods. Table 1 compares the voltage prediction accuracy; it can be seen that the RMSE of the traditional numerical method is 0.038 pu, and the accuracy of the pure LSTM method and the unconstrained agent method is also significantly lower than that of the method of this invention. The accuracy advantage stems from the embedding of the physical loss term in PINNs, which makes the network output follow the DistFlow equation, effectively reducing data overfitting.
[0067] Table 1 Comparison of Voltage Prediction Accuracy ; (3) Efficiency analysis: Figure 4 A schematic diagram comparing the computational efficiency of the method of this embodiment of the invention with a baseline method is provided for an embodiment of the invention. Figure 4 As shown, the computation time of each method varies with the number of agents during a one-hour simulation. Table 2 is a comparison of computational efficiency. In a scenario with 100 agents, the method of this invention takes an average of 5.3 seconds (standard deviation 0.4 seconds), while traditional numerical methods, LSTM methods, and unconstrained agent methods take 12.5 seconds (standard deviation 1.2 seconds), 8.1 seconds (standard deviation 0.7 seconds), and 9.7 seconds (standard deviation 0.9 seconds), respectively. The efficiency improvement is attributed to the parallel computation of PINNs and the topology optimization of graph neural networks, resulting in a computational complexity approximately linear O(n). Traditional methods, due to iterative solving of nonlinear equations, have a complexity close to O(n). 2 With 50 agents, the method of this invention takes 3.1 seconds, while the traditional method takes 7.8 seconds.
[0068] Table 2 Comparison of computational efficiency ; Figure 5 A graph showing the relationship between computation time and the number of agents is provided for embodiments of the present invention, such as... Figure 5 As shown, the computation time of each method varies with different numbers of agents (10-100).
[0069] (4) Dynamic scene examples: Figure 6 A schematic diagram comparing the voltage over-limit response time of the method of this embodiment of the invention and the baseline method is provided for an embodiment of the invention. Figure 6As shown, a voltage over-limit event under high photovoltaic penetration is simulated. During the period from 14:00 to 16:00 at node 21, the method of this invention successfully predicted the voltage over-limit (peak value 1.048 pu) and stabilized the voltage at 1.02 pu (safe limit 1.05 pu) by coordinating the energy storage output through the intelligent agent. In contrast, the traditional method and LSTM method experienced a brief voltage over-limit to 1.06 pu, lasting approximately 12 minutes, due to response delay. PINNs provide accurate state prediction, and the MCTS algorithm continuously optimizes the agent's actions (energy storage charging and discharging strategy), avoiding control failures caused by computational lag in traditional methods.
[0070] (5) Robustness assessment: Figure 7 A schematic diagram illustrating the robustness comparison between the method of this embodiment and a baseline method in a noisy environment, as provided in this embodiment of the invention, is shown below. Figure 7 As shown, Gaussian noise with a standard deviation of 5% of the measured value was injected to simulate data error. Table 3 is a robustness comparison table; the method of this invention has the lowest voltage fluctuation variance, at 0.0012 pu. 2 The variances of the traditional numerical method, the LSTM method, and the unconstrained agent method are 0.0028 pu. 2 0.0035 pu 2 and 0.0041pu 2 In noisy environments, the voltage prediction error of node 25 increased by only 8% using the method of this invention, while the LSTM method and the unconstrained agent method increased by 25% and 32%, respectively.
[0071] Table 3 Robustness Comparison Table ; Experiments show that the voltage prediction root mean square error of the method of the present invention is reduced to 0.021 pu (45% improvement in accuracy), the calculation time is shortened to 5.3 seconds (57% improvement in efficiency), the voltage over-limit duration is reduced from 12 minutes to near zero, and it is also significantly better than existing methods in terms of uncertainty robustness.
[0072] Compared with the prior art, the present invention has the following significant advantages: (1) The simulation accuracy has been greatly improved. By embedding the DistFlow physical equations as soft constraints into the neural network training process, the physical rationality of the simulation output is fundamentally guaranteed. Experiments show that the root mean square error (RMSE) of voltage prediction is reduced from 0.038 pu in the traditional method to 0.021 pu, an improvement of approximately 45% in accuracy. Ablation experiments confirm that physical constraint embedding can improve simulation accuracy by more than 30%, while reducing the physical violation rate from 12% to 0.5%.
[0073] (2) The computational efficiency is significantly improved. Thanks to the parallel computing capabilities of physical information neural networks and the topology-aware optimization of graph neural networks, the computational complexity has been reduced from O(n) of traditional methods. 2 The computation time was reduced to approximately linear O(n). In a scenario with 100 agents, the simulation computation time was reduced from 12.5 seconds to 5.3 seconds, an efficiency improvement of approximately 57%.
[0074] (3) Real-time dynamic response capability By employing a closed-loop collaborative mechanism of physical information neural networks and Monte Carlo tree search, real-time prediction and proactive control of dynamic events such as voltage exceedances are achieved. In photovoltaic fluctuation scenarios, the duration of voltage exceedances is reduced from 12 minutes to near zero, effectively avoiding control failures caused by computational lag in traditional methods.
[0075] (4) Robustness to strong uncertainty By using stochastic differential equation modeling and Bayesian variational inference, the uncertainties in the simulation were effectively quantified. Under 5% Gaussian noise interference, the voltage fluctuation variance was only 0.0012 pu. 2 The percentages are 42.9%, 34.3%, and 29.3% for traditional numerical methods, LSTM methods, and unconstrained agent methods, respectively, which are significantly better than the comparison methods.
[0076] (5) Good scalability The message passing mechanism of graph neural networks is naturally adapted to the topology of power distribution networks, enabling the method to be linearly extended to scenarios with hundreds of nodes, providing a feasible technical path for intelligent simulation and digital twin construction of large-scale power distribution networks.
[0077] Example 2 Figure 8 This is a schematic diagram of the structure of a power distribution network intelligent agent collaborative simulation device provided in Embodiment 2 of the present invention. The device is applicable to the simulation of the actual state of the power distribution network. The device can be implemented by software and / or hardware and is generally integrated on an electronic device.
[0078] like Figure 8 As shown, the device includes: The input module 210 is used to input the state vector of the distribution network at the current moment into the trained physical information neural network to obtain the global predicted state of the distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the distribution network. The candidate module 220 is used to output at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step based on the global predicted state through a multi-agent interaction model; wherein, each agent in the multi-agent interaction model corresponds to a distributed resource in the distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next time step and a state set corresponding to the joint scheduling action. The optimization module 230 is used to optimize the candidate scheduling decision set by employing the Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, to obtain the target scheduling decision set. The determination module 240 is used to take the scheduling decisions in the target scheduling decision set as the target scheduling decisions for each intelligent agent.
[0079] This embodiment provides a power distribution network intelligent agent collaborative simulation device, comprising: an input module, used to input the state vector of the power distribution network at the current moment into a trained physical information neural network to obtain the global predicted state of the power distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the power distribution network; a candidate module, used to output at least one candidate scheduling decision set for all distributed resources in the power distribution network at the next moment based on the global predicted state through a multi-agent interaction model; wherein, each agent in the multi-agent interaction model corresponds to a distributed resource in the power distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next moment and a state set corresponding to the joint scheduling action; an optimization module, used to optimize the candidate scheduling decision set by using a Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, to obtain a target scheduling decision set; and a determination module, used to take the scheduling decisions in the target scheduling decision set as the target scheduling decisions for each agent. This device, by combining physical information neural networks and Monte Carlo tree search algorithms, can accurately achieve real-time simulation of power distribution networks, solving the problems of low technical efficiency, low reliability, and non-compliance with physical laws in existing technologies for simulating power distribution networks.
[0080] Furthermore, the training of the physical information neural network includes: The initial physical information neural network is trained based on historical operating data of the distribution network to obtain the predicted state; Based on the predicted state and the fusion loss function, the current loss value of the initial physical information neural network is determined; When the loss value does not meet the training termination condition, the network parameters of the initial physical information neural network are adjusted based on the loss value to obtain an updated physical information neural network and continue training. When the loss value of the updated physical information neural network meets the training termination condition, the updated physical information neural network is taken as the trained physical information neural network.
[0081] Furthermore, the fusion loss function for: ; ; ; in, For data-driven loss terms, The number of training samples. For the distribution network in the first The actual state of each physical constraint configuration point For the distribution network in the first Predicted state of each physical constraint configuration point; For physical constraint loss terms, The weights of the physical constraint loss term, The number of physical constraint configuration points in the distribution network. For time, For nonlinear functions of the dynamic behavior of the distribution network, For the distribution network in the first Predicted voltage amplitude at each physically constrained configuration point For the distribution network in the first Predicted active power for each physically constrained configuration point For the distribution network in the first Predicted reactive power at each physically constrained configuration point These are the network parameters of the physical information neural network; The update equation for the network parameters is: ; in, The learning rate of the physical information neural network. The iteration number of the physical information neural network. This is to fuse the gradient of the loss function with respect to the network parameters.
[0082] Furthermore, candidate module 220 includes: The current state of each agent is obtained, and the current state of each agent and the global prediction state are input into the multi-agent interaction model to obtain multiple candidate action combinations; each candidate action combination contains the predicted actions of all agents in the next time step. Each candidate action combination and the current state are input into a stochastic differential equation to obtain a state combination corresponding to each candidate action combination; the state combination contains the predicted state of all agents in the next time step; Based on each of the candidate action combinations and the corresponding state combinations, at least one candidate scheduling decision set for all distributed resources in the distribution network at the next moment is constructed.
[0083] Furthermore, the optimization module 230 includes: For each candidate scheduling decision set, the Monte Carlo tree search algorithm is used to search the action space, and the physical information neural network and the multi-agent interaction model are combined to determine the subsequent scheduling decision set corresponding to the candidate scheduling decision set. The action value of the candidate scheduling decision set is updated based on the state-action pair value function, and the action value is used to guide the direction of searching the action space. Based on the candidate scheduling decision set and the corresponding subsequent scheduling decision set, the long-term value of the candidate scheduling decision set in the long run is evaluated by combining the objective function; the objective function takes minimizing the value of long-term simulation error as the optimization objective. The set of candidate scheduling decisions that has the best long-term value among all candidate scheduling decisions is taken as the target scheduling decision set.
[0084] Furthermore, the objective function is: ; in, Indicates at time The expected value of the cumulative discount return, For expectation operator, To extrapolate the time domain, To deduce the first in the time domain Time step index As a discount factor, ; The value function for the state-action pair is: ; in, It is the value of the state-action pair. It's the learning rate. It's an instant reward. , It is to perform an action The next state after that, This is a possible action for the next state.
[0085] Furthermore, the input module 210 includes: The current state vector of the distribution network is input into a graph neural network to obtain the embedding vector of the distribution network; the topology graph of the graph neural network is constructed based on the topology of the distribution network. The embedded vector is input into the trained physical information neural network to obtain the global predicted state of the power distribution network at the next moment.
[0086] The above-mentioned distribution network intelligent agent collaborative simulation device can execute the distribution network intelligent agent collaborative simulation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0087] Example 3 Figure 9 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0088] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0089] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0090] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the distribution network intelligent agent collaborative simulation method.
[0091] In some embodiments, the distribution network intelligent agent cooperative simulation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the distribution network intelligent agent cooperative simulation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the distribution network intelligent agent cooperative simulation method by any other suitable means (e.g., by means of firmware).
[0092] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0093] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0094] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0096] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0097] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0098] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0099] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A collaborative simulation method for intelligent agents in a power distribution network, characterized in that, The method includes: The current state vector of the distribution network is input into a trained physical information neural network to obtain the global predicted state of the distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the distribution network. Based on the global predicted state, at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step is output through a multi-agent interaction model; wherein, each agent in the multi-agent interaction model corresponds to a distributed resource in the distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next time step and a state set corresponding to the joint scheduling action. The Monte Carlo tree search algorithm is used, combined with the physical information neural network and the multi-agent interaction model, to optimize the candidate scheduling decision set and obtain the target scheduling decision set. The scheduling decisions in the target scheduling decision set are used as the target scheduling decisions for each agent.
2. The method according to claim 1, characterized in that, The training of the physical information neural network includes: The initial physical information neural network is trained based on historical operating data of the distribution network to obtain the predicted state; Based on the predicted state and the fusion loss function, determine the current loss value of the initial physical information neural network; When the loss value does not meet the training termination condition, the network parameters of the initial physical information neural network are adjusted based on the loss value to obtain an updated physical information neural network and continue training. When the loss value of the updated physical information neural network meets the training termination condition, the updated physical information neural network is taken as the trained physical information neural network.
3. The method according to claim 2, characterized in that, The fusion loss function for: ; ; ; in, For data-driven loss terms, The number of training samples. For the distribution network in the first The actual state of each physical constraint configuration point For the distribution network in the first Predicted state of each physical constraint configuration point; For physical constraint loss terms, The weights of the physical constraint loss term, The number of physical constraint configuration points in the distribution network. For time, For nonlinear functions of the dynamic behavior of the distribution network, For the distribution network in the first Predicted voltage amplitude at each physically constrained configuration point For the distribution network in the first Predicted active power for each physically constrained configuration point For the distribution network in the first Predicted reactive power at each physically constrained configuration point These are the network parameters of a physical information neural network; The update equation for the network parameters is: ; in, The learning rate of the physical information neural network. The iteration number of the physical information neural network. This is to fuse the gradient of the loss function with respect to the network parameters.
4. The method according to claim 1, characterized in that, Based on the global predicted state, the multi-agent interaction model outputs at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step, including: The current state of each agent is obtained, and the current state of each agent and the global prediction state are input into the multi-agent interaction model to obtain multiple candidate action combinations; each candidate action combination contains the predicted actions of all agents in the next time step. Each candidate action combination and the current state are input into a stochastic differential equation to obtain a state combination corresponding to each candidate action combination; the state combination contains the predicted state of all agents in the next time step; Based on each of the candidate action combinations and the corresponding state combinations, at least one candidate scheduling decision set for all distributed resources in the distribution network at the next moment is constructed.
5. The method according to claim 1, characterized in that, The Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, is used to optimize the candidate scheduling decision set to obtain the target scheduling decision set, including: For each candidate scheduling decision set, the Monte Carlo tree search algorithm is used to search the action space, and the physical information neural network and the multi-agent interaction model are combined to determine the subsequent scheduling decision set corresponding to the candidate scheduling decision set. The action value of the candidate scheduling decision set is updated based on the state-action pair value function, and the action value is used to guide the direction of searching the action space. Based on the candidate scheduling decision set and the corresponding subsequent scheduling decision set, the long-term value of the candidate scheduling decision set in the long run is evaluated by combining the objective function; the objective function takes minimizing the value of long-term simulation error as the optimization objective. The set of candidate scheduling decisions that has the best long-term value among all candidate scheduling decisions is taken as the target scheduling decision set.
6. The method according to claim 5, characterized in that, The objective function is: ; in, Indicates at time The expected value of the cumulative discount return, For expectation operator, To extrapolate the time domain, To deduce the first in the time domain Time step index As a discount factor, ; The value function for the state-action pair is: ; in, It is the value of the state-action pair. It's the learning rate. It's an instant reward. , It is to perform an action The next state after that, This is a possible action for the next state.
7. The method according to claim 1, characterized in that, The step of inputting the current state vector of the distribution network into a trained physical information neural network to obtain the global predicted state of the distribution network at the next moment includes: The current state vector of the distribution network is input into a graph neural network to obtain the embedding vector of the distribution network; the topology graph of the graph neural network is constructed based on the topology of the distribution network. The embedded vector is input into the trained physical information neural network to obtain the global predicted state of the power distribution network at the next moment.
8. A collaborative simulation device for intelligent agents in a power distribution network, characterized in that, The device includes: The input module is used to input the current state vector of the distribution network into the trained physical information neural network to obtain the global predicted state of the distribution network at the next moment; the physical information neural network is constructed based on the branch power flow equations of the distribution network. The candidate module is used to output at least one candidate scheduling decision set for all distributed resources in the distribution network at the next time step based on the global predicted state through a multi-agent interaction model; wherein, each agent in the multi-agent interaction model corresponds to a distributed resource in the distribution network, and each candidate scheduling decision set includes a joint scheduling action that the multiple agents can execute at the next time step and a state set corresponding to the joint scheduling action. The optimization module is used to optimize the candidate scheduling decision set by employing the Monte Carlo tree search algorithm, combined with the physical information neural network and the multi-agent interaction model, to obtain the target scheduling decision set. The determination module is used to take the scheduling decisions in the target scheduling decision set as the target scheduling decisions for each agent.
9. An electronic device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the power distribution network intelligent agent collaborative simulation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the power distribution network intelligent agent collaborative simulation method according to any one of claims 1-7.