Power distribution network voltage control method and device and electronic equipment
By constructing a voltage regulation model for the target node and training the strategy network using real-time observation data and historical data, precise and real-time voltage control under low perception conditions was achieved, solving the problems of accuracy and applicability of voltage control in distribution networks and improving the reliability of control.
Patent Information
- Application Number
- CN202411119934.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-08-15
AI Technical Summary
The accuracy, real-time performance, and applicability of existing power distribution network voltage control are relatively low, especially in situations with low perception, where it is difficult to achieve precise and rapid voltage regulation.
By constructing an initial voltage regulation model for the policy network, feature extraction network, and value network of the target node, training it using real-time observation data of the target node, obtaining the target output power, and making independent decisions at each target node, decentralized voltage control is achieved.
Precise, real-time voltage control of target nodes was achieved with low sensitivity, improving the accuracy and applicability of voltage control in the distribution network and reducing reliance on node communication.
Smart Images

Figure CN119029897B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device and electronic equipment for controlling the voltage of a power distribution network. Background Technology
[0002] With the integration of a large amount of renewable energy into the distribution network, the randomness and volatility of its output have led to problems such as reverse power flow overload and bidirectional voltage limit exceedances, seriously affecting the safe and stable operation of the distribution network. Therefore, voltage control in the distribution network is crucial. Traditional voltage regulation equipment in the distribution network includes group-switched capacitors and on-load tap-changing transformers. However, because they only support discrete regulation and have long switching times, they cannot meet the requirements of real-time voltage control. In contrast, distributed photovoltaic systems, static var compensators (SVCs), and other devices connected through inverters can achieve continuous regulation of reactive power output and have the advantages of fast response speed and low regulation cost. Therefore, they can coordinate and regulate various fast reactive power regulation resources to fully tap their reactive power regulation potential.
[0003] Existing voltage regulation technologies for power distribution networks can be broadly categorized into two types: mathematical optimization methods and data-driven methods. Mathematical optimization methods generally require precise physical models and system network parameters. For example, robust optimization methods are used to model the uncertainties of distributed power sources and loads, transforming the non-convex voltage control problem model into a solvable form through second-order cone relaxation and linearization techniques, and then iteratively finding the optimal solution. Another example is the use of day-ahead discrete device control and intraday rolling optimization correction, which effectively improves system voltage stability. Data-driven methods can perform grid voltage regulation through deep reinforcement learning. This approach learns the control strategies of agents from historical interaction experience, and can be trained offline beforehand. During execution, it only requires feedforward computation of the neural network, enabling rapid decision-making and significantly improving control speed. For instance, mathematical models can be combined with data-driven methods: day-ahead voltage optimization is based on mixed-integer second-order cone programming, while intraday real-time regulation is based on multi-agent reinforcement learning, reducing the dependence of control on models and communication costs.
[0004] Mathematical optimization methods are highly dependent on the power flow model and data of the distribution network. However, in actual distribution networks, there are often missing parameters, making it difficult to obtain all parameters in real time and to establish an accurate physical model. Therefore, mathematical optimization methods are unlikely to achieve ideal results under low-perception conditions and have low applicability. Among the existing deep reinforcement learning algorithms, neural networks are relatively simple. When applied to solve large-scale distribution network systems, they are difficult to achieve real-time and accurate control of the distribution network.
[0005] There is currently no effective solution to the problems of low accuracy, real-time performance, and applicability of voltage control in power distribution networks. Summary of the Invention
[0006] The purpose of this application is to provide a method, device, and electronic device for controlling distribution network voltage, so as to solve the problems of low accuracy, real-time performance, and applicability of existing distribution network voltage control.
[0007] To solve the above-mentioned technical problems, the first aspect of this specification provides a method for controlling the voltage of a power distribution network, including:
[0008] Acquire real-time observation data of target nodes in the target distribution network;
[0009] The real-time observation features are input into the target voltage regulation model corresponding to the target node, and the target output power corresponding to the target node is output. The target voltage regulation model of the target node is trained as follows: an initial voltage regulation model of the target node, including a policy network, a feature extraction network, and a value network, is constructed; first historical observation data of N nodes in the target distribution network and second historical observation data of the target node are obtained, where N is a positive integer; the policy network of the target node is adjusted based on the first and second historical observation data; the observation features of the target node are extracted using the feature extraction network of the target node; the adjusted policy network is optimized using the policy network and the observation features, and the optimized policy network is used as the target voltage regulation model.
[0010] The target node is controlled based on the target output power.
[0011] In some embodiments of this specification, the target distribution network includes N nodes, and the target node is one of the nodes selected from the N nodes. Each target node corresponds to a target voltage regulation model, and the target voltage regulation model corresponding to each target node independently decides the target output power based on the real-time observation data of the target node.
[0012] In some embodiments of this specification, before constructing the initial voltage regulation model of the target node, which includes a policy network, a feature extraction network, and a value network, the following steps are included:
[0013] Construct a distribution network system model for the target distribution network, which includes a distributed photovoltaic model, a static var compensator model, a load model, and a power flow model.
[0014] Based on the power distribution network system model, determine the node status information, observation status information, action information, and action report information of the target power distribution network;
[0015] Accordingly, the initial voltage regulation model for the target node, including a policy network, a feature extraction network, and a value network, is constructed as follows:
[0016] Based on node state information, observation state information, action information, and action reward information, the target node is constructed, including a policy network, a feature extraction network, and a value network.
[0017] In some embodiments of this specification, the feature extraction network of the target node is used to extract the observed features of the target node, including:
[0018] The first historical observation data is input into the adjusted policy network, and the historical policy actions of the target node are output.
[0019] Obtain the third historical observation data after the target node executes the historical strategy action;
[0020] The importance of the target node in the target distribution network is determined by calculating the correlation between multiple target nodes in the target distribution network;
[0021] The feature extraction network is used to extract features from the historical policy actions and the third historical observation data based on the determined importance and the topological mask of the region where the target node is located, to obtain the observation features.
[0022] In some embodiments of this specification, the feature extraction network includes a linear layer, a first self-attention encoder, a second self-attention encoder, and a selection layer;
[0023] Accordingly, after obtaining the third historical observation data after the target node executes the historical policy action, it includes:
[0024] The linear layer is used to perform a linear transformation on the observation sequence composed of the third historical observation data and the historical strategy actions;
[0025] The first self-attention encoder is used to perform self-attention encoding on the linearly transformed observation sequence based on the topological mask of the target region where the target node is located, and the first observation action vector is output.
[0026] The observation action vector is self-attention encoded using the second self-attention encoder, and the second observation action vector is output.
[0027] The attention weight of the target node is calculated using the selection layer, and the second observation action vector is feature extracted based on the attention weight to obtain the observation features; wherein the attention weight is determined by the correlation between the observation action vectors of other nodes in the target distribution network and the second observation action vector.
[0028] In some embodiments of this specification, the feature extraction network is represented by the following formula:
[0029]
[0030] Among them, g i (o i ,a i ) represents the encoding function for the observation data and policy actions of the i-th target node, ζ i f represents the sum of attention weights of the i-th target node to other target nodes in the target distribution network. i Represents a fully connected layer, V represents a linear transformation, ReLU represents the neural network activation function, and ε represents the full connection layer. ij This represents the attention weight of the i-th target node relative to the j-th target node.
[0031] In some embodiments of this specification, the adjusted policy network is optimized using the policy network and the observation features, including:
[0032] The observed features are input into the value function of the value network, and combined with the reward function of the target node, the value of the historical policy actions output by the policy network is determined.
[0033] Based on the determined value and the optimization function of the target node, the adjusted policy network is optimized.
[0034] The second aspect of this specification provides a control device for distribution network voltage, comprising:
[0035] The data acquisition module is used to acquire real-time observation data of target nodes in the target distribution network;
[0036] The model calculation module is used to input real-time observation features into the target voltage regulation model corresponding to the target node and output the target output power corresponding to the target node. The target voltage regulation model of the target node is trained as follows: an initial voltage regulation model of the target node, including a policy network, a feature extraction network, and a value network, is constructed; first historical observation data of N nodes in the target distribution network and second historical observation data of the target node are obtained, where N is a positive integer; the policy network of the target node is adjusted based on the first and second historical observation data; the observation features of the target node are extracted using the feature extraction network of the target node; the adjusted policy network is optimized using the policy network and the observation features, and the optimized policy network is used as the target voltage regulation model.
[0037] A voltage control module is used to control the target node based on the target output power.
[0038] A third aspect of this specification provides an electronic device, comprising: a memory and a processor, wherein the processor and the memory are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to implement the steps of the above method.
[0039] A fourth aspect of this specification provides a computer storage medium storing computer program instructions that, when executed by a processor, implement the steps of the above-described method.
[0040] The distribution network voltage control method provided in this specification involves acquiring real-time observation data of target nodes in the target distribution network; inputting the real-time observation features into the target voltage regulation model corresponding to the target node; and outputting the target output power corresponding to the target node. The target voltage regulation model of the target node is trained as follows: constructing an initial voltage regulation model of the target node including a policy network, a feature extraction network, and a value network; acquiring first historical observation data of N nodes in the target distribution network and second historical observation data of the target node, where N is a positive integer; adjusting the policy network of the target node based on the first and second historical observation data; extracting observation features of the target node using the feature extraction network of the target node; optimizing the adjusted policy network using the policy network and observation features, and using the optimized policy network as the target voltage regulation model; and controlling the target node based on the target output power. In the embodiments of this specification, the voltage control method for the distribution network described above allows for training of the target voltage regulation model for each target node using historical observation data from multiple nodes in the target distribution network. Each target node can train its own target voltage regulation model. In practical application decision-making, the control of the target output power of each target node can be independent of information from other nodes and communication between nodes, achieving decentralized voltage control of the target distribution network. That is, through centralized training and decentralized decision execution of multiple target nodes, precise and real-time control of the target output power of the target node can be achieved even when the target distribution network has low awareness and insufficient observation data. Furthermore, when training the target voltage regulation model, the observation features of the target nodes are extracted using a feature extraction network. When optimizing the policy network using a value network, the extracted observation features can be combined with the feature extraction network, improving the applicability of the method in large-scale distribution networks and increasing the accuracy and reliability of distribution network voltage control. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0042] Figure 1 The diagram shown is a schematic of a power distribution network voltage control method provided in an embodiment of this specification.
[0043] Figure 2 The diagram shown is a schematic of a training method for the target voltage regulation model provided in an embodiment of this specification.
[0044] Figure 3 The diagram shown is a schematic representation of the feature extraction process provided in an embodiment of this specification.
[0045] Figure 4 The diagram shown is a schematic representation of the voltage regulation decision-making process provided in an embodiment of this specification.
[0046] Figure 5 The figure shown is a schematic diagram of the initial voltage regulation model provided in the embodiments of this specification;
[0047] Figure 6 The diagram shown is a schematic of the distribution network voltage provided in the embodiments of this specification;
[0048] Figure 7 The figure shown is a schematic diagram of the node voltage variation curve provided in the embodiment of this specification;
[0049] Figure 8 The diagram shown is a schematic of a power distribution network voltage control device provided in an embodiment of this specification.
[0050] Figure 9 The diagram shown is a schematic of an electronic device provided in an embodiment of this specification. Detailed Implementation
[0051] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0052] As mentioned above, the accuracy, real-time performance, and applicability of current distribution network voltage control are relatively low. To address these issues, this specification provides a method for controlling distribution network voltage. This method involves acquiring real-time observation data of target nodes in the target distribution network; inputting the real-time observation features into the target voltage regulation model corresponding to the target node; and outputting the target output power corresponding to the target node. The target voltage regulation model of the target node is trained as follows: constructing an initial voltage regulation model for the target node, including a strategy network, a feature extraction network, and a value network; acquiring first historical observation data of N nodes in the target distribution network and second historical observation data of the target node, where N is a positive integer; adjusting the strategy network of the target node based on the first and second historical observation data; extracting observation features of the target node using the feature extraction network of the target node; optimizing the adjusted strategy network using the strategy network and observation features, and using the optimized strategy network as the target voltage regulation model; and controlling the target node based on the target output power.
[0053] In the embodiments of this specification, the voltage control method for the distribution network described above allows for training of the target voltage regulation model for each target node using historical observation data from multiple nodes in the target distribution network. Each target node can train its own target voltage regulation model. In practical application decision-making, the control of the target output power of each target node can be independent of information from other nodes and communication between nodes, achieving decentralized voltage control of the target distribution network. That is, through centralized training and decentralized decision execution of multiple target nodes, precise and real-time control of the target output power of the target node can be achieved even when the target distribution network has low awareness and insufficient observation data. Furthermore, when training the target voltage regulation model, the observation features of the target nodes are extracted using a feature extraction network. When optimizing the policy network using a value network, the extracted observation features can be combined with the feature extraction network, improving the applicability of the method in large-scale distribution networks and increasing the accuracy and reliability of distribution network voltage control.
[0054] It is understood that the methods described in the embodiments of this specification can be applied to electronic devices, which can refer to electronic devices with data computing, processing, and storage capabilities. These electronic devices can be terminals such as PCs (Personal Computers), tablets, smartphones, wearable devices, and intelligent robots; they can also be servers. A server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0055] The method for controlling the voltage of the power distribution network provided in the embodiments of this application will now be described in conjunction with the accompanying drawings.
[0056] Figure 1 The diagram illustrates a method for controlling distribution network voltage according to an embodiment of this application. While this specification provides method operation steps or device structures as shown in the following embodiments or figures, the method or device may include more or fewer operation steps or module units through conventional or non-inventive means. In steps or structures where there is no logically necessary causal relationship, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or figures of this specification. When the method or module structure is applied in actual devices, servers, or terminal products, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or figures (e.g., in a parallel processor or multi-threaded processing environment, or even in a distributed processing or server cluster implementation environment). Figure 1 As shown, the method may include:
[0057] S101: Obtain real-time observation data of the target node in the target distribution network.
[0058] It is understandable that the target node can be a node selected from all nodes in the target distribution network based on the target distribution network topology, which requires voltage regulation. Furthermore, voltage regulation of the target node can be achieved solely through the observation data of the target node. Real-time observation data can be the status information of all devices connected to the target node, such as the active and reactive power output of the distributed photovoltaic devices connected to the target node, the reactive power output of the static var compensator (SVC), the active and reactive power demand of the target node load, and the node voltage, specifically related to the devices connected to the target node.
[0059] S102: Input the real-time observed features into the target voltage regulation model corresponding to the target node, and output the target output power corresponding to the target node.
[0060] The target voltage regulation model for the target node can be achieved by constructing an initial voltage regulation model for the target node, including a policy network, a feature extraction network, and a value network. Based on historical observation data from the target distribution network and the feature extraction and value networks, the policy network is adjusted and optimized. The trained target voltage regulation model can consist only of the adjusted and optimized policy network. This model is then configured in the adjustable transformer area connected to the target node, with each adjustable transformer area serving as the corresponding agent for the target node. Each adjustable transformer area for the target node can correspond to a separate agent used to make decisions regarding the target output power of the target node. The specific training process of the target voltage regulation model for the target node will be described below with reference to the accompanying figures, and will not be elaborated upon here.
[0061] In some embodiments of this specification, the target distribution network may include N nodes, and the target node is one of the nodes selected from the N nodes. Each target node may correspond to a target voltage regulation model, and the target voltage regulation model corresponding to each target node can independently decide the target output power based on the real-time observation data of the target node.
[0062] In some embodiments of this specification, the target output power output by the target voltage regulation model can be the reactive power output value of the reactive power regulating device connected to the target node, and the reactive power output value can meet the reactive power output constraint of the corresponding reactive power regulating device.
[0063] S103: Control the target node based on the target output power.
[0064] In some embodiments of this specification, multiple target nodes can simultaneously generate corresponding target output power based on the target voltage regulation model corresponding to each target node, thereby enabling voltage control of the target distribution network through the collaborative action of the agents corresponding to multiple target nodes. In other embodiments, the decision-making and actions of the agents of each target node can be performed separately. For example, the agent of target node A can acquire the real-time observation data of the corresponding node at time T1 and generate a strategy action using the target voltage regulation model corresponding to that node. The agent of target node A can then control the equipment based on the generated strategy action. Similarly, the agent of target node B can acquire the real-time observation data of the corresponding node at time T2 and generate a strategy action using the target voltage regulation model corresponding to that node. The agent of target node B can then control the equipment based on the generated strategy action. The strategy action needs to be executed for the agent corresponding to target node A at time T2. That is, agents executing actions at different times need to acquire real-time observation data only after the agent's action at the previous time has been completed, to avoid interference between the decision-making actions of different agents and affecting the accuracy of the decision-making and control of other agents.
[0065] The following is combined Figure 2 The training process of the target voltage regulation model for the target node in the embodiments of this specification is described. Specifically, as follows... Figure 2 As shown, the target voltage regulation model of the target node can be trained in the following way:
[0066] S201: Construct an initial voltage regulation model for the target node, including a policy network, a feature extraction network, and a value network.
[0067] It can be understood that the policy network can be used for action decisions of target nodes, the feature extraction network can be used to extract the observation features of target nodes, and the value network can be used to evaluate the value of the actions decided by the policy network, and the policy network can be optimized and adjusted based on the value evaluated by the value network. Specifically, the feature extraction network can be used to extract the policy actions output by the policy network and the observation features of the observation data obtained after executing the policy actions from the perspective of the agent corresponding to the target node. Then, when the value network performs value evaluation, it can achieve value evaluation based on the observation features output by the feature extraction network. The feature extraction network can reduce the data processing volume of the value network and can achieve more refined control of the target node based on the importance of the target node in the target distribution network.
[0068] In some embodiments of this specification, the strategy network can construct an experience pool based on expert knowledge, domain knowledge base, voltage regulation rules, and voltage regulation rules in special scenarios corresponding to the equipment connected to the target node in the target distribution network. The strategy network can generate decision actions based on the constructed experience pool.
[0069] In some embodiments of this specification, before constructing the initial voltage regulation model of the target node, which includes a policy network, a feature extraction network, and a value network, the following may be included:
[0070] Construct a distribution network system model for the target distribution network, which includes a distributed photovoltaic model, a static var compensator model, a load model, and a power flow model.
[0071] Based on the power distribution network system model, determine the node status information, observation status information, action information, and action report information of the target power distribution network;
[0072] Accordingly, the initial voltage regulation model for constructing the target node, including the policy network, feature extraction network, and value network, may include:
[0073] Based on node state information, observation state information, action information, and action reward information, the target node is constructed, including a policy network, a feature extraction network, and a value network.
[0074] In some embodiments of this specification, the distributed photovoltaic (PV) model can correspond to the distributed PV devices connected to the target node. Distributed PV devices typically operate in maximum power point tracking (MPPT) mode to ensure their active power output is at its maximum. Under this constraint, the active power output mainly depends on sunlight, and can be expressed as:
[0075]
[0076] in, P can represent the active power output of distributed photovoltaic devices in maximum power point tracking mode. pv,e This can represent the rated value of distributed photovoltaic equipment; I, I e These represent the actual light intensity and the rated light intensity of the distributed photovoltaic equipment, respectively.
[0077] Furthermore, if the actual light intensity can follow a β distribution, i.e., I~B(α,β), then the adjustable active power range of the distributed photovoltaic system can be expressed as:
[0078]
[0079] Among them, P pv It can represent the active power output of distributed photovoltaic equipment.
[0080] The reactive power regulation capability of distributed photovoltaic (PV) equipment is limited by its capacity, active power, and maximum allowable power factor angle. Therefore, its reactive power adjustable range can be expressed as:
[0081]
[0082] Among them, Q pv It can represent the reactive power output of distributed photovoltaic devices; The maximum reactive power output of distributed photovoltaic (PV) equipment can be expressed by the following formula:
[0083]
[0084] in, S can represent the maximum permissible power factor angle of distributed photovoltaic (PV) equipment. N It can represent the installed capacity of distributed photovoltaic equipment.
[0085] It is understandable that the above formulas (1) to (4) can be represented as distributed photovoltaic models for distributed photovoltaic equipment.
[0086] In some embodiments of this specification, the Static Var Compensator (SVC) corresponding to the Static Var Compensator model can be a reactive power compensation device that can smoothly adjust the equivalent reactance of the distribution network it is connected to, thereby changing the reactive power compensation value at the grid connection point (i.e., the node in the target distribution network). Connecting it in parallel with the system can achieve the function of rapidly adjusting the reactive power of the system. The reactive power adjustable range of the SVC can be expressed as:
[0087]
[0088] in, This can represent the minimum reactive power output of the SVC. Q can represent the maximum reactive power output of the SVC. svc The reactive power output of the SVC can be represented. The static var compensator model can be represented by formula (5).
[0089] In some embodiments of this specification, due to the uncertainty of the load at the target node, the active power and reactive power generally follow a normal distribution, and can be expressed as follows:
[0090]
[0091] Among them, P L Q L σ can represent the active power and reactive power output by the load, respectively. P σ Q The standard deviations of active power and reactive power can be represented by μ, respectively. P μ Q These can represent the expected active power and reactive power, respectively. Specifically, the expected active and reactive power of the load can be assumed to satisfy an average distribution. Formulas (6) and (7) above can be represented as a load model.
[0092] In some embodiments of this specification, the power flow model can be represented by the following formula:
[0093]
[0094] Among them, P i Q i These can be represented as the active power and reactive power injected into node i in the target distribution network, respectively, θ ij G can represent the phase angle difference between the two ends of the limit ij between node i and node j in the target distribution network. ij B ij V can represent the conductance and susceptance of line ij, respectively. i V j These can represent the voltage amplitudes of node i and node j in the target distribution network, respectively.
[0095] Furthermore, based on the above formulas (1) to (8), the observable equipment information of each node in the target distribution network can be obtained, including: the active and reactive power output of the distributed photovoltaic equipment connected to the target node, the reactive power output of the static var compensator (SVC), the active and reactive power demand of the target node load, and the node voltage. Furthermore, based on the determined observable quantities, the observation data input to the strategy network can be determined, and the output quantity of the strategy network corresponding to each node can be defined as the reactive power output value of the equipment connected to that node, and this value needs to satisfy the reactive power output constraints of each device. The specific construction process of the initial target voltage regulation model will be introduced below in conjunction with the embodiments, and will not be elaborated here.
[0096] S202: Obtain the first historical observation data of N nodes in the target distribution network and the second historical observation data of the target node, where N is a positive integer.
[0097] It can be understood that the N nodes in the target distribution network can be all nodes in the target distribution network. During the training process, historical observation data of all nodes in the target distribution network are used for training, which can result in a target voltage regulation model with higher data processing accuracy. That is, during the training process, multiple target node agents undergo intensive multi-agent chemical training.
[0098] S203: Adjust the policy network of the target node based on the first historical observation data and the second historical observation data.
[0099] It is understandable that during the adjustment of the policy network, the second historical observation data can be input into the policy network, the policy network can make policy actions based on the second historical observation data, the agent can execute the policy actions, update the device state based on the state transition matrix of the device connected to the target node, and adjust the policy network based on the updated device state and the corresponding update policy of the policy network.
[0100] S204: Use the feature extraction network of the target node to extract the observation features of the target node.
[0101] It can be understood that the observation features extracted by the feature extraction network can be derived from the policy actions taken by the policy network based on the importance of the target nodes, as well as the new observation data after the actions. The importance of the target nodes can be obtained by correlation analysis based on the network topology of the target distribution network and the policy actions of multiple target nodes, as well as the new observation data after the actions.
[0102] In some embodiments of this specification, extracting the observed features of the target node using the feature extraction network of the target node may include:
[0103] The first historical observation data is input into the adjusted policy network, and the historical policy actions of the target node are output.
[0104] Obtain the third historical observation data after the target node executes the historical strategy action;
[0105] The importance of the target node in the target distribution network is determined by calculating the correlation between multiple target nodes in the target distribution network;
[0106] The feature extraction network is used to extract features from the historical policy actions and the third historical observation data based on the determined importance and the topological mask of the region where the target node is located, to obtain the observation features.
[0107] It is understandable that the third historical observation data can be the data obtained by the agent corresponding to the target node after executing historical policy actions and updating its state based on the state transition matrix. This data, obtained by the agent based on the new environmental state, can be used to evaluate the value of the second historical observation data-historical policy action pair. The importance of the target node in the target distribution network can be determined through a self-attention encoding mechanism.
[0108] like Figure 3 This illustrates the process of observation feature extraction using the feature extraction network in the embodiments of this specification. For details, please refer to... Figure 3 As shown, in some embodiments of this specification, the feature extraction network may include a linear layer 301, a first self-attention encoder 302, a second self-attention encoder 303, and a selection layer 304;
[0109] Accordingly, after obtaining the third historical observation data after the target node executes the historical strategy action, the process may include: using the linear layer to perform a linear transformation on the observation sequence composed of the third historical observation data and the historical strategy action; using the first self-attention encoder to perform self-attention encoding on the linearly transformed observation sequence based on the topology mask of the target area where the target node is located, and outputting a first observation action vector; using the second self-attention encoder to perform self-attention encoding on the observation action vector, and outputting a second observation action vector; using the selection layer to calculate the attention weight of the target node, and performing feature extraction on the second observation action vector based on the attention weight, to obtain the observation features; wherein the attention weight is determined by the correlation between the observation action vectors of other nodes in the target distribution network and the second observation action vector.
[0110] It can be understood that the first self-attention encoder can be a masked self-attention encoder, and the mask it carries can be a topology mask generated based on the network topology of the target distribution network. Specifically, the topology mask for each agent is not the same and can be obtained from the topology of the region where the agent is located. The topology mask is not fixed and can be adjusted according to the topological links between nodes. The second attention encoder can be a maskless self-attention encoder.
[0111] For details, please refer to Figure 3 As shown, the specific process by which a feature extraction network extracts observed features can include: defining the observation sequence Θ. i The region Λ where the intelligent agent i is located in the distribution area. i The state vectors of all nodes within the topology are used, and the observation sequence can include historical policy actions and third-party historical observation data. Each state variable is linearly transformed and then input into an L-layer self-attention encoder with a mask. Topological mask. For region Λ i The corresponding adjacency matrix, and Let be a symmetric matrix consisting of 0 and -∞. If nodes are connected, the corresponding matrix element is 0; otherwise, it is -∞. Next, after passing through a self-attention encoder without a mask and a selection layer, and defining the output as the vector corresponding to the node i where the agent is located, the feature vector λ can be obtained. i (Observational features). By introducing a selection layer, the feature extraction network can extract observational features from the perspective of the agent.
[0112] In some embodiments of this specification, the selection layer in the feature extraction network can extract key features using a self-attention encoder. The self-attention encoder can employ a transformer structure, embedding and encoding the input sequence. The input to the self-attention encoder layer yields Q (query), K (key), and V (value), which can be represented as vectors of fixed dimensions. For each agent, its own observation-action vector, after self-attention encoding, serves as the key, while the observation-action vectors of other agents serve as the query. The correlation between agents is calculated using the dot product method, and the calculated correlation can be used as attention weights. By assigning different weights to other agents, the value network can extract the observation features of the agent most relevant to its own reward.
[0113] S205: Optimize the adjusted strategy network using the strategy network and the observation features, and use the optimized strategy network as the target voltage regulation model.
[0114] In some embodiments of this specification, a reward function and gradient update strategy corresponding to the policy network can be constructed. Then, when optimizing the policy network based on historical policy networks and observed features, the network update can be driven by the reward function as the core, and the policy gradient can be updated in the direction of maximizing the reward function until the reward function no longer increases and converges. This indicates that the policy network optimization is complete, and the optimized policy network is used as the target voltage regulation model corresponding to the target node for voltage regulation of the target node.
[0115] In some embodiments of this specification, optimizing the adjusted policy network using the policy network and the observation features may include:
[0116] The observed features are input into the value function of the value network, and combined with the reward function of the target node, the value of the historical policy actions output by the policy network is determined.
[0117] Based on the determined value and the optimization function of the target node, the adjusted policy network is optimized.
[0118] The following section, with reference to the accompanying diagram, describes the training process of the target voltage regulation model using deep reinforcement learning (DEL) as an example. This training process may include agent model construction and multi-agent solution.
[0119] In some embodiments of this specification, the construction of the intelligent agent model may include a decision process analysis stage and a model definition stage for the voltage regulation model of the target distribution network.
[0120] Specifically, in the decision-making process analysis phase, considering the low-perception characteristics of the target distribution network, each reactive power device in the target distribution network can only observe a portion of local state variables. Therefore, the decision-making of the target node can be represented as a Partially Observable Markov Decision Process (POMDP). That is, the voltage control problem is transformed into a standard distributed partially observable Markov decision process, constructing a distributed voltage regulation framework with the objectives of minimizing voltage deviation and network losses. The interaction process between the agent corresponding to each target node and the environment can be described as follows: Figure 4 As shown. Reference Figure 4 As shown, this process can be specifically described as a decision-making process in which an agent makes policy actions based on observation data obtained from the environment. After all the actions of the agents, the external environment updates its state based on the state transition probability matrix. The agent obtains a reward based on the new environmental state and starts a new round of decision-making process. The purpose of DEL is to maximize the cumulative reward.
[0121] All information about the POMDP problem corresponding to the target node can be integrated into a single tuple. Where Ω can represent a set of agents; O can represent the environmental state space, containing observation information from all nodes in the external environment used for decision-making; O can represent the joint observation space, o i可以 For data observable by agent i; It can represent the joint action space, a i It can be the action value for agent i to make decisions; It can represent the environmental state transition probability function that satisfies the Markov property; γ can represent the discount factor, which is used to characterize the effect of the current action on the agent's reward in the next step, and γ∈[0,1]; It can represent the global reward, r i The reward that can be allocated to agent i.
[0122] Specifically, during the model definition phase, relevant elements in the initial voltage regulation model can be defined, which may include the following:
[0123] 1) Intelligent agent: Each adjustable transformer substation connected to the distribution network can be regarded as an intelligent agent, and a transformer substation connected to node i corresponds to a single intelligent agent.
[0124] 2) Environmental State Space: The physical environment for voltage regulation is the distribution network; therefore, S can represent the set of state information for all nodes in the distribution network, where...
[0125]
[0126] S can include N s i s i P can represent the state vector of node i. pv,i Q pv,i Q can represent the active and reactive power outputs of the distributed photovoltaic system connected to node i, respectively. svg,i P can represent the reactive power output of the SVC connected to node i. L,i Q L,i U can represent the active and reactive power demand of node i, respectively. i θ can represent the voltage at node i. i It can represent the phase angle of node i, and N is the number of nodes in the distribution network.
[0127] 3) Observation Space: Since the strategy aims to achieve distributed voltage control with low awareness, each agent only needs to observe the device information of its own grid-connected node and make decisions based on the monitored information. Therefore, the observation space of the agent can be represented by the following formula:
[0128] Oi ={P pv,i Q pv,i Q svg,i ,P L,i Q L,i U i ,θ i} Formula (10)
[0129] 4) Action Space: To achieve voltage control, the action value a decided by agent i is... i It can be the reactive power output value of the reactive power regulating resource, and it must meet the reactive power output constraints of each component.
[0130] 5) Reward Function: The updating of the agent policy network is guided by maximizing the reward. Therefore, a reward function can be designed to guide the agent to reduce voltage deviation and network loss. The reward function can be expressed by the following formula:
[0131]
[0132] Where, k U k loss The coordination factors N can represent voltage and network loss return, respectively. U This can represent a node in the distribution network where voltage exceeds the limit, r i P can represent the reward value of node i. loss The system loss of a distribution network can be represented by the following formula:
[0133]
[0134] It can be understood that formula (9) is the global state space, which represents the set of state information of all nodes in the distribution network. The subscript i here belongs to the range of all nodes in the distribution network. Formula (10) is a partial observation space for a single agent. Each agent can only observe the electrical information of its own grid-connected node. The subscript i here represents the agent's own grid-connected point.
[0135] In some embodiments of this specification, multi-agent problem solving may include reinforcement learning and feature extraction. It is understood that after model definition, the defined model can be solved, that is, the model parameters can be optimized and adjusted.
[0136] Specifically, for reinforcement learning solutions, a multi-agent reinforcement learning algorithm can be used. This means each agent is assigned a specific reinforcement learning algorithm, and multiple agents learn from each other's algorithms collectively, sharing information during the learning process. The reinforcement learning algorithm combines value assessment and policy generation; its specific structure can be as follows: Figure 5 As shown. Reference Figure 5As shown, it can include a value network and a policy network. The two Q-networks in the value assessment network are related to the environment (i.e.,...) Figure 5 The Q-value network evaluates the benefits generated by the Environment and actions, selecting the smaller Q-value, while the policy network is responsible for generating the optimal action. Understandably, a feature extraction network can be placed within the value network to extract features from the input data before the value network performs value evaluation.
[0137] In the aforementioned chemical strengthening process, while maximizing the benefit, a concept is introduced that can be expressed by the following formula:
[0138]
[0139] Where T can represent the optimization time length, t can represent time, E can represent the expected value, and s can represent the expected value. t It can represent the state space at time t, a t It can represent the action space at time t, and has ρ can represent the weighting coefficient, and H can represent the entropy of the policy action π. The entropy value of policy action π can be calculated using the following formula:
[0140] H(π(·|s t ))=-∑a t π(a t |s t )logπ(a t |s t ) Formula (14)
[0141] Reinforcement learning algorithms can use a value function Q(s,a) to evaluate the value of policy action π, and update the gradient of the policy network using a gradient update formula, where the value function and the gradient update formula can be expressed as follows:
[0142]
[0143] Where τ can be a coefficient, and θ can be a parameter of the value network Q.
[0144] Specifically, feature extraction can be achieved using methods such as... Figure 3The feature extraction network shown is implemented as described above. It is understood that current MASAC algorithms mostly employ fully connected neural networks (FNNs), which simply concatenate the state vectors observed by the agent and directly input them into the policy and value networks for computation. While this method has good control performance in small systems, it leads to excessively long vector dimensions when applied to large and complex power distribution networks, increasing the processing burden on the agent. Furthermore, considering the varying importance of different nodes in a power distribution network, simple concatenation cannot specifically address the situation of each node for more refined adjustments. In this embodiment, a feature extraction network considering network topology is constructed based on a self-attention encoder to understand the topological structure of the controlled area, identify the location of key nodes, and obtain more concise high-dimensional features. The specific structure of the feature extraction network and the feature extraction process can be referred to the preceding description and will not be repeated here.
[0145] In some embodiments of this specification, it is possible to Figure 3 The observation sequence input into the feature extraction network is defined as follows:
[0146]
[0147] Furthermore, if we use Ξ to represent the feature extraction network, then the feature extraction network can be represented by the following formula:
[0148]
[0149] Among them, g i (o i ,a i ζ can represent the encoding function for the observation data and policy actions of the i-th target node. i f can represent the sum of attention weights of the i-th target node to other target nodes in the target distribution network. i ReLU can represent a fully connected layer, V represents a linear transformation, and ReLU can represent the activation function of a neural network. ε ij This can represent the attention weight of the i-th target node relative to the j-th target node.
[0150] The beneficial effects of the power distribution network voltage regulation method in the embodiments of this specification will be further introduced below in combination with specific application scenarios.
[0151] like Figure 6 The diagram illustrates the topology of the target distribution network in the application scenario of Embodiment 1 of this specification. This target distribution network may include 33 nodes, and nodes 3, 15, 17, and 22 are selected as target nodes. Since voltage exceedances are more likely at system end nodes, node 17 is used as an example for verification. The voltage variation of this node over 24 hours can be shown as follows... Figure 7 As shown. Reference Figure 7 As shown, compared to no control, local droop control (QV-Droop), and particle swarm (PSO) control, after adopting the MASAC voltage control method proposed in the embodiments of this specification, through centralized training and decentralized execution, it can ensure that the voltage of node 17 does not exceed the limit within 24 hours, reduce the maximum voltage deviation of node 17, make the overall voltage curve more gradual, and reduce fluctuations.
[0152] The distribution network voltage control method provided in this specification establishes a voltage control model under low-perception conditions. It treats each adjustable transformer substation as an agent and transforms the voltage control problem into a standard distributed partially observable Markov decision process. A distributed voltage regulation framework is constructed with the objectives of minimizing voltage deviation and network losses. Based on this, an improved MASAC algorithm based on a feature extraction network is designed for offline centralized training. The trained agents are then used for online distributed voltage regulation of the transformer substations, ultimately achieving rapid voltage stabilization. This method can be applied to distribution network voltage regulation scenarios of a certain scale under low-perception conditions. Furthermore, a self-attention encoder is introduced to construct the feature extraction network, assigning attention weights to adjustable transformer substations at different nodes. This allows the value network to focus more on substations with stronger voltage regulation capabilities and enables on-site distributed reactive power and voltage regulation of substations without relying on communication, even when measurement data is insufficient. The dynamic network topology of the distribution network is also considered. Based on graph theory, the radial topology of the distribution network is represented using a topology mask and input into the feature extraction network, enabling the agents to identify the locations of key nodes during training.
[0153] Based on the above-described method for controlling distribution network voltage, one or more embodiments of this specification also provide a device for controlling distribution network voltage. The device may include an apparatus (including a distributed system), software (application), module, plug-in, server, client, etc., using the method described in the embodiments of this specification, combined with necessary hardware implementation. Based on the same innovative concept, the devices in one or more embodiments provided in this specification are as described in the following embodiments. Since the implementation schemes and methods for solving the problem are similar, the implementation of specific devices in the embodiments of this specification can refer to the implementation of the aforementioned method, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated. Figure 8 The diagram shown is a schematic representation of a power distribution network voltage control device provided in an embodiment of this application. Figure 8 As shown, the voltage control device 800 of the power distribution network may include:
[0154] The data acquisition module 801 is used to acquire real-time observation data of the target nodes in the target distribution network.
[0155] The model calculation module 802 is used to input real-time observation features into the target voltage regulation model corresponding to the target node and output the target output power corresponding to the target node. The target voltage regulation model of the target node is trained in the following manner: constructing an initial voltage regulation model of the target node including a policy network, a feature extraction network, and a value network; acquiring first historical observation data of N nodes in the target distribution network and second historical observation data of the target node, where N is a positive integer; adjusting the policy network of the target node based on the first and second historical observation data; extracting observation features of the target node using the feature extraction network of the target node; optimizing the adjusted policy network using the policy network and the observation features, and using the optimized policy network as the target voltage regulation model.
[0156] Voltage control module 803 is used to control the target node based on the target output power.
[0157] The descriptions and functions of the above units can be understood by referring to the section on control methods for distribution network voltage, and will not be repeated here.
[0158] This specification also provides an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 901 and a memory 902, wherein the processor 901 and the memory 902 may be connected via a bus or other means. Figure 9 Taking the example of a connection between China and Israel via a bus.
[0159] Processor 901 can be a Central Processing Unit (CPU). Processor 901 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.
[0160] Memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the power distribution network voltage control method in this embodiment of the invention (e.g., Figure 8 The data acquisition module 801, model calculation module 802, and voltage control module 803 are shown. The processor 901 executes various functional applications and data processing by running non-transient software programs, instructions, and modules stored in the memory 902, thereby realizing the power distribution network voltage control method in the above method embodiment.
[0161] The memory 902 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor 901, etc. Furthermore, the memory 902 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 902 may optionally include memory remotely located relative to the processor 901, and these remote memories may be connected to the processor 901 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0162] The one or more modules are stored in the memory 902, and when executed by the processor 901, they perform the following: Figure 1 The method for controlling the voltage of the power distribution network in the illustrated embodiment.
[0163] The specific details of the aforementioned electronic device can be understood by referring to the relevant descriptions and effects in the above method embodiments, and will not be repeated here.
[0164] This specification also provides a computer storage medium storing computer program instructions, which, when executed, implement the steps of the above-described power distribution network voltage control method.
[0165] This specification also provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the above-described power distribution network voltage control method.
[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium can also include combinations of the above types of memory.
[0167] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, please refer to each other. The focus of each embodiment is to describe the differences from other embodiments.
[0168] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions.
[0169] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0170] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute certain parts of the methods of various embodiments of this application.
[0171] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.
[0172] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0173] Although this application has been described through embodiments, those skilled in the art will know that this application has many modifications and variations without departing from the spirit of this application, and it is intended that the appended claims cover such modifications and variations without departing from the spirit of this application.
Claims
1. A method for controlling voltage in a power distribution network, characterized in that, include: Acquire real-time observation data of target nodes in the target distribution network; The real-time observation features are input into the target voltage regulation model corresponding to the target node, and the target output power corresponding to the target node is output. The target voltage regulation model of the target node is trained as follows: an initial voltage regulation model of the target node, including a policy network, a feature extraction network, and a value network, is constructed; first historical observation data of N nodes in the target distribution network and second historical observation data of the target node are obtained, where N is a positive integer; the policy network of the target node is adjusted based on the first and second historical observation data; the observation features of the target node are extracted using the feature extraction network of the target node; the adjusted policy network is optimized using the policy network and the observation features, and the optimized policy network is used as the target voltage regulation model. The target node is controlled based on the target output power; The feature extraction network includes a linear layer, a first self-attention encoder, a second self-attention encoder, and a selection layer; it utilizes the target node's feature extraction network to extract the observed features of the target node, including: Input the first historical observation data into the adjusted policy network, and output the historical policy actions of the target node; Obtain the third historical observation data after the target node executes the historical strategy action; A linear transformation is performed on the observation sequence composed of the third historical observation data and historical strategy actions using a linear layer; The first self-attention encoder is used to perform self-attention encoding on the linearly transformed observation sequence based on the topological mask of the target region where the target node is located, and the first observation action vector is output. The first observation action vector is self-attentionally encoded using a second self-attention encoder, and the second observation action vector is output. The attention weights of the target node are calculated using a selection layer, and the second observation action vector is extracted based on the attention weights to obtain the observation features. The attention weights are determined by the correlation between the observation action vectors of other nodes in the target distribution network and the second observation action vector.
2. The method according to claim 1, characterized in that, The target distribution network includes N nodes. The target node is one of the nodes selected from the N nodes. Each target node corresponds to a target voltage regulation model, and the target voltage regulation model corresponding to each target node independently decides the target output power based on the real-time observation data of the target node.
3. The method according to claim 1, characterized in that, Before constructing the initial voltage regulation model for the target node, which includes a policy network, a feature extraction network, and a value network, the following steps are included: Construct a distribution network system model for the target distribution network, which includes a distributed photovoltaic model, a static var compensator model, a load model, and a power flow model. Based on the power distribution network system model, determine the node status information, observation status information, action information, and action report information of the target power distribution network; Accordingly, the initial voltage regulation model for the target node, including a policy network, a feature extraction network, and a value network, is constructed as follows: Based on node state information, observation state information, action information, and action reward information, the target node is constructed, including a policy network, a feature extraction network, and a value network.
4. The method according to claim 1, characterized in that, Using the feature extraction network of the target node, the observed features of the target node are extracted, including: The first historical observation data is input into the adjusted policy network, and the historical policy actions of the target node are output. Obtain the third historical observation data after the target node executes the historical strategy action; The importance of the target node in the target distribution network is determined by calculating the correlation between multiple target nodes in the target distribution network; The feature extraction network is used to extract features from the historical policy actions and the third historical observation data based on the determined importance and the topological mask of the region where the target node is located, to obtain the observation features.
5. The method according to any one of claims 1-4, characterized in that, The feature extraction network is represented by the following formula: ; Among them, g i (o i ,a i ) represents the encoding function for the observation data and policy actions of the i-th target node. Let represent the sum of attention weights of the i-th target node to other target nodes in the target distribution network. f i This represents a fully connected layer, where V represents a linear transformation, and ReLU represents the neural network activation function. This represents the attention weight of the i-th target node relative to the j-th target node. Represents a topological mask. Let Ω represent the feature extraction network, and let Ω represent the set of agents.
6. The method according to claim 1, characterized in that, Optimizing the adjusted policy network using the policy network and the observation features includes: The observed features are input into the value function of the value network, and combined with the reward function of the target node, the value of the historical policy actions output by the policy network is determined. Based on the determined value and the optimization function of the target node, the adjusted policy network is optimized.
7. A control device for distribution network voltage, characterized in that, include: The data acquisition module is used to acquire real-time observation data of target nodes in the target distribution network; The model calculation module is used to input real-time observation features into the target voltage regulation model corresponding to the target node and output the target output power corresponding to the target node. The target voltage regulation model of the target node is trained as follows: an initial voltage regulation model of the target node, including a policy network, a feature extraction network, and a value network, is constructed; first historical observation data of N nodes in the target distribution network and second historical observation data of the target node are obtained, where N is a positive integer; the policy network of the target node is adjusted based on the first and second historical observation data; the observation features of the target node are extracted using the feature extraction network of the target node; the adjusted policy network is optimized using the policy network and the observation features, and the optimized policy network is used as the target voltage regulation model. A voltage control module is used to control the target node based on the target output power; The feature extraction network includes a linear layer, a first self-attention encoder, a second self-attention encoder, and a selection layer; it utilizes the target node's feature extraction network to extract the observed features of the target node, including: Input the first historical observation data into the adjusted policy network, and output the historical policy actions of the target node; Obtain the third historical observation data after the target node executes the historical strategy action; A linear transformation is performed on the observation sequence composed of the third historical observation data and historical strategy actions using a linear layer; The first self-attention encoder is used to perform self-attention encoding on the linearly transformed observation sequence based on the topological mask of the target region where the target node is located, and the first observation action vector is output. The first observation action vector is self-attentionally encoded using a second self-attention encoder, and the second observation action vector is output. The attention weights of the target node are calculated using a selection layer, and the second observation action vector is extracted based on the attention weights to obtain the observation features. The attention weights are determined by the correlation between the observation action vectors of other nodes in the target distribution network and the second observation action vector.
8. An electronic device, characterized in that, include: A memory and a processor, the processor and the memory being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to implement the steps of the method according to any one of claims 1 to 6.
9. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent power distribution network voltage safety control method, device, equipment and medium thereof
CN116826762A