Model establishment method for power distribution network node voltage control mode selection, dynamic selection method for power distribution network node voltage control mode, control system and power system

By adopting the attention mechanism and deep reinforcement learning model in the new distribution network to calculate the node voltage feature weights and frequency curves, the complexity and accuracy problems of selecting the node voltage control mode of the new distribution network are solved, and the control efficiency and safety of the power system are improved.

CN120675093AActive Publication Date: 2025-09-19JINCHENG POWER SUPPLY COMPANY OF STATE GRID SHANXI ELECTRIC POWER
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510785958.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The node voltage control of the new distribution network faces the problems of changeable flow direction, complex node voltage control requirements, low accuracy of analysis results of existing methods, and poor control effect.

Method used

Using the attention mechanism and deep reinforcement learning model, we collect node voltage features, calculate weights, construct state space and reward functions, train the deep reinforcement learning model, output node occurrence frequency, and draw a frequency curve to select the voltage control mode.

Benefits of technology

It achieves scientific and reasonable selection of voltage control mode for distribution network nodes, improves the effectiveness and accuracy of control, and ensures the safe and stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675093A_ABST
    Figure CN120675093A_ABST
Patent Text Reader

Abstract

The invention provides a model establishment method for power distribution network node voltage control mode selection, a dynamic selection method for the power distribution network node voltage control mode, a control system and a power system. The establishment method comprises the following steps: acquiring voltage characteristics of each node of a power distribution network at a certain moment; on the basis of the voltage characteristics, an attention mechanism is adopted to carry out weight calculation on each node, and a weight calculation result is used as an attribute degree value; inputting the attribute degree value into a first deep reinforcement learning model for training to obtain an occurrence frequency of each node when the model reaches an optimal target state; according to the occurrence frequency of each node when the model reaches the optimal target state at multiple moments, creating an occurrence frequency curve of each node at continuous moments; matching a corresponding node voltage control mode according to the occurrence frequency curve corresponding to each node; the method comprehensively considers the mutual influence degree between the voltage characteristics of each node, realizes the comprehensive judgment of the node voltage control mode of the power distribution network, and is suitable for the field of power systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power systems, and specifically to a model establishment method for selecting a voltage control mode for a distribution network node, a dynamic selection method for a voltage control mode for a distribution network node, a control system, and a power system. Background Art

[0002] In the power system, controlling the node voltage of the distribution network is a core measure to ensure the safe, reliable and economical operation of the power system. For today's new distribution networks, voltage control faces greater challenges than traditional distribution network voltage control, mainly reflected in the following aspects: First, the traditional distribution network topology is a radial structure with a single power flow direction. Its voltage control mainly relies on centralized control of the voltage at the upper node. However, due to the integration of a large number of renewable energy sources, the new distribution network combines centralized and distributed systems (such as rooftop photovoltaics) with variable power flow directions. Therefore, relying solely on centralized control of the upper node cannot achieve voltage stability. Second, the degree of automation of traditional distribution network nodes is low, and functions such as telemetry and telesignaling are imperfect. This cannot provide effective data information for analysis when formulating voltage control strategies. Therefore, the voltage control of traditional distribution networks generally does not consider the participation of nodes in medium and low voltage distribution networks (10kV and below). However, the voltage control of new distribution networks needs to cover all voltage levels: high voltage (110kV, 35kV), medium voltage (10kV), and low voltage (380V / 220V). It has higher requirements for voltage control optimization strategies and requires the coordinated cooperation of networks at all voltage levels.

[0003] Because new distribution networks connect to a wider range of loads, including distributed photovoltaics, microgrids, smart parks, smart substations, air conditioners, electric vehicles, and other controllable resources, they offer a broader range of voltage regulation options compared to traditional distribution networks. Therefore, node voltage control in new distribution networks is no longer limited to the traditional single-phase control mode, and local control has become an important option. Furthermore, due to the integration of a large number of renewable energy stations, rooftop photovoltaics, electric vehicles, air conditioners, and other controllable load resources, new distribution networks possess significant voltage control potential, enabling voltage regulation at both medium and low voltage distribution network nodes, playing an important regulatory role in the research and development of voltage control strategies. Therefore, the main research direction for voltage control in new distribution networks lies in the coordinated control between high, medium, and low voltage distribution networks.

[0004] Currently, the basic methods of coordinated control include centralized control, local control, and hierarchical control. Centralized control primarily achieves voltage balance in the regional distribution network by adjusting parameters such as reactive power, main transformer gear position, and active power at important nodes in the network. Local control primarily achieves voltage stability at nodes experiencing voltage violations through power control and main transformer gear position adjustments. Hierarchical control, encompassing both centralized and local control, primarily achieves voltage balance in the regional distribution network by implementing hierarchical adjustments at nodes of different voltage levels and in different regions. Regardless of the control mode employed, determining the node type to be controlled and the corresponding control mode is crucial.

[0005] At present, there are three main methods for selecting the voltage control mode of new distribution network nodes: the first is to determine the voltage-sensitive nodes of the regional distribution network based on voltage sensitivity, and achieve voltage balance of the regional power grid by controlling the nodes with the largest voltage sensitivity; the second is to control special nodes in the distribution network where voltage exceeds the limit frequently or the regulation capacity is large (usually nodes such as active power transmission, reactive power imbalance, serious voltage exceeding the limit, and new energy stations are selected for control) to achieve on-site balance of the nodes; the third is to use genetic optimization algorithms, particle swarm optimization algorithms and other means to optimize the distribution network, perform optimal control by setting the objective function, and finally generate a node control plan that meets the optimization goal.

[0006] The three selection methods mentioned above all have significant drawbacks in practical applications. First, although the voltage sensitivity selection mechanism is the most common, voltage sensitivity calculation relies on power flow calculation. Power flow calculation for new distribution networks involves modeling operations for new energy station node models, load models, and other models, which is more difficult than traditional power flow calculation. At the same time, voltage sensitivity cannot quickly adapt to changes in grid topology and operating parameters, resulting in more difficult and inaccurate analysis results, and cannot provide a screening mechanism for selecting voltage control modes for distribution network nodes. Second, although selecting specific nodes for voltage control based on the cause of voltage over-limit is a common method in actual power grid projects, this method mainly focuses on the symptoms of voltage over-limit, such as active power transmission and reactive power imbalance. Therefore, voltage regulation at these nodes is a crude regulation method. Although node selection is easy and the control mode is changeable, it ignores the overall internal mechanism of the power grid, cannot achieve precise control of the distribution network, and the control effect is poor. Third, many existing optimization algorithms are usually only satisfied with the grid operation objective function. Although they cover multiple objectives such as voltage stability, active power reduction, and control cost, they ignore the selection of node control mode, making it difficult to achieve ideal results in application. Summary of the Invention

[0007] In order to solve one of the above technical defects, the present application provides a model establishment method for selecting a distribution network node voltage control mode, a dynamic selection method for a distribution network node voltage control mode, a control system and a power system.

[0008] According to a first aspect of the present application, a model establishment method for selecting a voltage control mode for a distribution network node is provided, comprising:

[0009] Collect voltage characteristics of each node in the distribution network at time t1, wherein the voltage characteristics include active power, reactive power, voltage RMS and voltage phase angle;

[0010] Based on the voltage characteristics of each node at time t1, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value to measure the relationship between the voltages of each node in the distribution network at time t1;

[0011] Constructing a state space, action space, and reward function of a deep reinforcement learning model based on the attribute values;

[0012] Training each state space variable in the state space one by one so that the deep reinforcement learning model reaches the optimal target state during training, and outputting the frequency of occurrence of each node when it reaches the optimal target state;

[0013] The successfully trained deep reinforcement learning model is used as the first deep reinforcement learning model for the final selection of the voltage control mode of the distribution network node.

[0014] Preferably, the method of calculating the weight of each node based on the voltage characteristics of each node at time t1 by using an attention mechanism, and using the weight calculation result as the attribute value for measuring the relationship between the voltages of each node in the distribution network at time t1, specifically includes:

[0015] According to the active power, reactive power, voltage RMS and voltage phase angle of each node in the distribution network at time t1, a voltage characteristic matrix is ​​established to represent the voltage change characteristics of each node in the distribution network at time t1;

[0016] According to each voltage characteristic matrix, the node characteristic matrix of the distribution network at time t1 is constructed;

[0017] An attention mechanism is used to calculate the similarity between the features of each node in the node feature matrix, and an attention mechanism matrix is ​​obtained according to the similarity calculation result;

[0018] After normalizing the node feature matrix, an attention weight matrix is ​​calculated based on the attention mechanism matrix and the normalized node feature matrix;

[0019] The attention weight matrix is ​​used as the attribute value for measuring the relationship between the voltages of the nodes in the distribution network at time t1.

[0020] Preferably, constructing the state space, action space and reward function of the deep reinforcement learning model according to the attribute value specifically includes:

[0021] Constructing a state space of a deep reinforcement learning model based on the attention weight matrix, wherein variables in the state space correspond one-to-one to weight elements in the attention weight matrix;

[0022] Constructing an action space of a deep reinforcement learning model based on the mutual transfer actions between state variables in the state space, wherein each transfer action is executed by an ε-greedy strategy;

[0023] According to the attention weight matrix, a reward function of the deep reinforcement learning model is constructed: when training a certain state space variable, if the weight element between the current state space variable and other state space variables is 0, the reward function value is 0; if the weight element between the current state space variable and other state space variables is not 0, the reward function value is the sum of the weight element between the two and the next state space variable.

[0024] Preferably, the deep reinforcement learning model is a DQN network model, and the training of each state space variable in the state space one by one so that the deep reinforcement learning model reaches the optimal target state during training and outputs the frequency of occurrence of each node when it reaches the optimal target state specifically includes:

[0025] Select a state space variable from the state space as the initial state space variable, input it into the training Q network of the first deep reinforcement learning model for training, and generate the corresponding initial training Q value;

[0026] Execute the action space according to the ε-greedy strategy, calculate the corresponding reward function value according to the reward function, and form the experience data set of each action information to build the experience pool;

[0027] Select at least one set of experience data sets from the experience pool, input them into the training Q network and the target Q network of the first deep reinforcement learning model respectively, and calculate the corresponding training Q value of the training Q network and the target value of the target Q network respectively;

[0028] Determine whether the training Q value reaches the target value. If it does, determine whether the model has reached the number of iterations. If it does not, reselect the initial state space variables for a new round of training.

[0029] If the model reaches the number of iterations, it determines whether the model has calculated the target state for each node; if the model does not reach the number of iterations, it executes the action space according to the ε-greedy strategy and reselects the next state space variable for training until the number of iterations is reached;

[0030] If the model calculates the target state for each node, the frequency of occurrence of each node when the model reaches the optimal target state is counted and output; if there are nodes that have not been calculated for the target state, the initial state space variables are reselected for a new round of training.

[0031] Preferably, the voltage characteristic matrix is ​​expressed as:

[0032]

[0033] Among them, x i.t1 is the voltage characteristic matrix of the i-th node at time t1, p i is the active power of the i-th node, q i is the reactive power of the ith node, U i is the effective value of the voltage at the i-th node, is the voltage phase angle of the i-th node;

[0034] The node feature matrix is ​​expressed as:

[0035]

[0036] Among them, X t1 is the node characteristic matrix of the distribution network at time t1, n1 is the number of nodes in the distribution network;

[0037] The attention mechanism matrix is ​​expressed as:

[0038]

[0039] Among them, e t1 is the attention mechanism matrix at time t1, e t1 Each element e in i.j Used to represent the node feature matrix X t1 Medium x i.t1 and x j.t1 similarity;

[0040] The element e i.j Expressed as:

[0041]

[0042] Where ReLU is the activation function, a is the trainable parameter vector, W is the trainable parameter matrix used to perform linear transformation on the input features, || represents the concatenation of vectors, and k represents the number of node features contained in the node feature matrix, k = 1, ..., n1;

[0043] The normalized node feature matrix is ​​expressed as The elements in Expressed as:

[0044]

[0045] Among them, E[X t1 ] and Var[X t1 ] are the node feature matrices X t1 The mean and variance of , ε1 is the variance control bias, γ1 and β are the learnable parameters of the normalized calculation layer of the dynamic convolutional neural network, and ⊙ represents the Hadamard product;

[0046] The attention weight matrix is ​​expressed as:

[0047]

[0048] Among them, W B.t1 is the attention weight matrix at time t1, σ(·) is the activation function, is the similarity matrix, for The transposed matrix of

[0049] The similarity matrix Expressed as:

[0050]

[0051] Preferably, the state space is represented as s t1 , the action space is represented as a t1 ;

[0052] The reward function is expressed as:

[0053]

[0054] Among them, r i is the reward function corresponding to the training of the i-th node. When the weight element w between the i-th node and the j-th node ij When it is 0, the reward function value is 0. When the weight element w between the i-th node and the j-th node ij When it is not 0, the reward function value is the weight element w ij and the next state space variable s j of and;

[0055] The ε-greedy strategy is expressed as:

[0056]

[0057] Among them, π(a t1 |s t1 ) is the ε-greedy strategy function; ε2 is the random selection probability, which is ε2=max(0.01, 1-n2 / N), where n2 is the current training number of the model and N is the total training number; argmaxQ(a t1 ,s t1 ) is the maximum Q value corresponding to the optimal action; when the action is the optimal choice, its selection probability is 1-ε2, otherwise the selection probability is ε2.

[0058] Preferably, the empirical data set is represented as:

[0059] (s i , a t1 , r i , s j )

[0060] Among them, s i is the current state space variable being trained, a t1 is the currently executed action, r i is the corresponding reward function value, s j is the next state space variable selected according to the state transition strategy;

[0061] When d sets of experience data sets are selected from the experience pool and input into the training Q network and target Q network of the first deep reinforcement learning model respectively, d <d m , where d m is the data capacity of the experience pool;

[0062] During the training process, the target Q network is calculated using the following formula:

[0063]

[0064] Among them, m is the current iteration number, T is the total number of iterations set in advance, and y m is the target value of the target Q network for the mth time. When m reaches T, the immediate reward r obtained by training the network Q m is the reward value; when m does not reach T, the long-term cumulative reward value is the reward value; γ2 is the discount coefficient;

[0065] During the training process, the following formula is used to update the loss function of the training Q network:

[0066]

[0067] Among them, m is the current iteration number, T is the total number of iterations set in advance, and y m is the target value of the mth target Q network, θ m is the m-th loss function.

[0068] According to a second aspect of the present application, a method for dynamically selecting a voltage control mode for a distribution network node is provided, comprising:

[0069] Collect voltage characteristics of each node in the distribution network at time t2, wherein the voltage characteristics include active power, reactive power, voltage RMS value and voltage phase angle;

[0070] Based on the voltage characteristics of each node at time t2, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value to measure the relationship between the voltages of each node in the distribution network at time t2;

[0071] Inputting the attribute degree value into a first deep reinforcement learning model for training to obtain the occurrence frequency of each node when the model reaches the optimal target state, wherein the first deep reinforcement learning model is the first deep reinforcement learning model established by the aforementioned model establishment method;

[0072] According to the frequency of each node appearing when the model reaches the optimal target state at multiple moments, create a frequency curve of each node appearing at consecutive moments;

[0073] According to the characteristics of the occurrence frequency curve corresponding to each node, the corresponding node voltage control mode is matched.

[0074] According to the third aspect of the present application, a distribution network node voltage control system is provided, comprising a module for implementing the model establishment method for selecting the distribution network node voltage control mode as described above; or comprising a module for implementing the dynamic selection method of the distribution network node voltage control mode as described above.

[0075] According to a fourth aspect of the present application, a power system is provided, comprising the distribution network node voltage control system as described above.

[0076] The dynamic selection method of the voltage control mode of the distribution network node provided in this application directly uses the active power, reactive power, voltage RMS and voltage phase angle of each node in the distribution network as node features, dynamically calculates the weights of the node features of the distribution network through a dynamic attention mechanism, and uses the calculated dynamic weight results as the attribute value of the node mode selection. Then, the reward function of the deep reinforcement learning network is designed based on the attribute value, and the deep reinforcement learning network with this unique reward function is used to train each node. The frequency of occurrence of each node when the model reaches the optimal target state is calculated, and then the frequency curve of occurrence of each node at consecutive moments is plotted. Finally, the corresponding node voltage control mode is matched according to the characteristics of the frequency curve of occurrence of each node;

[0077] This method comprehensively considers the active power, reactive power, voltage RMS and voltage phase angle of each node as node voltage characteristics. Compared with the traditional single consideration method, it improves the comprehensiveness and accuracy of the node voltage control mode analysis; adopts the dynamic attention mechanism to calculate the weight of each node voltage feature, and constructs a new training model based on the weight calculation results. It not only effectively, accurately and comprehensively expresses the voltage change characteristics of each node, but also comprehensively considers the mutual relationship and mutual influence degree between the voltage characteristics of each node, and realizes the comprehensive judgment of the distribution network node voltage control mode; adopts deep reinforcement learning to carry out characteristic node voltage changes and analysis decisions at dynamic moments, fully developing the node voltage characteristics of each node. The algorithm has exerted its strong adaptability and generalization ability, and has effectively solved the complex decision-making problem of continuous changes in distribution network voltage; for the first time, "node occurrence frequency" is proposed as a characterization variable that is better used to quantify the characteristics of node voltage changes. It can reflect the level of change of voltage at different nodes in different time periods and different working conditions, and can also reflect the node voltage control potential and response ability, providing a reliable basis for the selection of node voltage control mode; the node occurrence frequency curve at continuous moments is used as the basis for selecting the node voltage control mode, which fully considers the dynamic performance of the node characteristic variables, achieves the purpose of multi-type and multi-time control mode selection, and realizes the flexibility and adaptability of distribution network node voltage control.

[0078] In summary, this application can provide a scientific and reasonable auxiliary decision-making basis for the selection of distribution network node voltage control mode, can improve the effectiveness and accuracy of node voltage control mode selection, and thus can improve the efficiency and effect of distribution network voltage control, and ensure the safe and stable operation of the power system.

[0079] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the contents indicated in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0081] Figure 1 A flowchart of a method for dynamically selecting a voltage control mode for a distribution network node provided by one embodiment of the present application;

[0082] Figure 2 A flowchart of a model establishment method for selecting a voltage control mode for a distribution network node provided by one embodiment of the present application;

[0083] Figure 3 A graph showing the frequency of occurrence of distribution network nodes provided in one embodiment of the present application. DETAILED DESCRIPTION

[0084] In order to make the technical solutions and advantages of the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless they conflict.

[0085] In view of some problems existing in the existing technology:

[0086] In a first aspect, embodiments of the present application provide a method for dynamically selecting a voltage control mode for a distribution network node. This method can be performed by a device for dynamically selecting a voltage control mode for a distribution network node, or by a component configured within the device for dynamically selecting a voltage control mode for a distribution network node, such as a chip or chip system. The method can also be implemented by a logic module or software having some or all of the functions of a device for dynamically selecting a voltage control mode for a distribution network node. This application is not limited to this.

[0087] For example, Figure 1 As shown, the dynamic selection method of the distribution network node voltage control mode includes:

[0088] S101, collecting voltage characteristics of each node in the distribution network at time t2, wherein the voltage characteristics include active power, reactive power, voltage RMS, and voltage phase angle;

[0089] S102, based on the voltage characteristics of each node at time t2, using the attention mechanism to calculate the weight of each node, and using the weight calculation result as the attribute value for measuring the relationship between the voltages of each node in the distribution network at time t2;

[0090] S103. Input the attribute degree values ​​into a first deep reinforcement learning model for training to obtain the frequency of occurrence of each node when the model reaches an optimal target state. The first deep reinforcement learning model is not a conventional deep reinforcement learning model, but a deep reinforcement learning model established using a model establishment method specifically provided in an embodiment of the present application. The specific establishment process is described in detail below.

[0091] S104, creating an appearance frequency curve of each node at consecutive moments based on the appearance frequency of each node when the model reaches the optimal target state at multiple moments;

[0092] S105 , matching a corresponding node voltage control mode according to the characteristics of the occurrence frequency curve corresponding to each node.

[0093] In certain possible implementations of the first aspect, based on the voltage characteristics of each node at time t2, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value for measuring the relationship between the voltages of each node in the distribution network at time t2, specifically including:

[0094] According to the active power, reactive power, voltage RMS and voltage phase angle of each node in the distribution network at time t2, a characteristic matrix is ​​established to represent the voltage change characteristics of each node in the distribution network at time t2;

[0095] According to each characteristic matrix, the node characteristic matrix of the distribution network at time t2 is constructed;

[0096] An attention mechanism is used to calculate the similarity between the features of each node in the node feature matrix, and an attention mechanism matrix is ​​obtained according to the similarity calculation result;

[0097] After normalizing the node feature matrix, an attention weight matrix is ​​calculated based on the attention mechanism matrix and the normalized node feature matrix;

[0098] The attention weight matrix is ​​used as the attribute value for measuring the relationship between the voltages of the nodes in the distribution network at time t2.

[0099] As an extremely important link in the embodiment of this application, the establishment of the first deep reinforcement learning model plays a key and decisive role in the rationality, accuracy and effectiveness of the entire node voltage control mode selection result. Therefore:

[0100] In a second aspect, embodiments of the present application provide a model building method for selecting a voltage control mode for a distribution network node. This method can be performed by a model building device for selecting a voltage control mode for a distribution network node, or by components configured within the model building device for selecting a voltage control mode for a distribution network node, such as a chip or chip system. Furthermore, this method can be implemented by a logic module or software having some or all of the functions of a model building device for selecting a voltage control mode for a distribution network node. This application is not limited to this.

[0101] For example, Figure 2 As shown, the model establishment method for selecting the voltage control mode of the distribution network node includes:

[0102] S201, collecting voltage characteristics of each node in the distribution network at time t1, wherein the voltage characteristics include active power, reactive power, voltage RMS, and voltage phase angle;

[0103] S202: Based on the voltage characteristics of each node at time t1, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value for measuring the relationship between the voltages of each node in the distribution network at time t1;

[0104] S203: Constructing a state space, action space, and reward function of a deep reinforcement learning model based on the attribute values;

[0105] S204: Training each state space variable in the state space one by one, so that the deep reinforcement learning model reaches the optimal target state during training, and outputting the occurrence frequency of each node when it reaches the optimal target state;

[0106] S205. Use the successfully trained deep reinforcement learning model as the first deep reinforcement learning model ultimately used for selecting the voltage control mode of the distribution network node.

[0107] In certain possible implementations of the second aspect, based on the voltage characteristics of each node at time t1, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value for measuring the relationship between the voltages of each node in the distribution network at time t1, specifically including:

[0108] According to the active power, reactive power, voltage RMS and voltage phase angle of each node in the distribution network at time t1, a voltage characteristic matrix is ​​established to represent the voltage change characteristics of each node in the distribution network at time t1;

[0109] According to each voltage characteristic matrix, the node characteristic matrix of the distribution network at time t1 is constructed;

[0110] An attention mechanism is used to calculate the similarity between the node features in the node feature matrix (i.e., the importance of one node feature relative to another node feature), and an attention mechanism matrix is ​​obtained according to the similarity calculation results;

[0111] After normalizing the node feature matrix (normalizing each element in the node feature matrix to generate feature data with a mean of 0 and a variance of 1), the attention weight matrix is ​​calculated based on the attention mechanism matrix and the normalized node feature matrix;

[0112] The attention weight matrix is ​​used as the attribute value for measuring the relationship between the voltages of the nodes in the distribution network at time t1.

[0113] In certain possible implementations of the second aspect, constructing the state space, action space, and reward function of the deep reinforcement learning model based on the attribute values ​​specifically includes:

[0114] Constructing a state space of a deep reinforcement learning model based on the attention weight matrix, wherein variables in the state space correspond one-to-one to weight elements in the attention weight matrix;

[0115] Constructing an action space of a deep reinforcement learning model based on the mutual transfer actions between state variables in the state space, wherein each transfer action is executed by an ε-greedy strategy;

[0116] According to the attention weight matrix, a reward function of the deep reinforcement learning model is constructed: when training a certain state space variable, if the weight element between the current state space variable and other state space variables is 0, the reward function value is 0; if the weight element between the current state space variable and other state space variables is not 0, the reward function value is the sum of the weight element between the two and the next state space variable.

[0117] In certain possible implementations of the second aspect, the deep reinforcement learning model is a DQN network model, and the training of each state space variable in the state space one by one, so that the deep reinforcement learning model reaches an optimal target state during training, and outputting the frequency of occurrence of each node when the optimal target state is reached, specifically includes:

[0118] S2041. Randomly select a state space variable from the state space as the initial state space variable, input it into the training Q network of the first deep reinforcement learning model for training, and generate a corresponding initial training Q value (i.e., value);

[0119] S2042. Execute the action space according to the ε-greedy strategy, calculate the corresponding reward function value according to the reward function, and form the action information of each action into an experience data set to build an experience pool;

[0120] S2043. Randomly select at least one set of experience data sets from the experience pool, input them into the training Q network and the target Q network of the first deep reinforcement learning model, and calculate the training Q value of the corresponding training Q network and the target value of the target Q network respectively;

[0121] S2044: Determine whether the training Q value reaches the target value. If so, execute step S2045 to determine whether the model has reached the number of iterations. If not, reselect the initial state space variables for a new round of training.

[0122] If the model reaches the number of iterations, step S2046 is executed to determine whether the model has calculated the target state for each node; if the model has not reached the number of iterations, the action space is executed according to the ε-greedy strategy, and the next state space variable is reselected for training until the number of iterations is reached;

[0123] If the model has performed target state calculation for each node, execute step S2047 to count and output the frequency of occurrence of each node when the model reaches the optimal target state; if there are nodes that have not performed target state calculation, reselect the initial state space variables for a new round of training to solve the common local optimal problem.

[0124] In some possible implementations of the second aspect, the voltage characteristic matrix is ​​expressed as:

[0125]

[0126] Among them, x i.t1 is the voltage characteristic matrix of the i-th node at time t1, p i is the active power of the i-th node, q i is the reactive power of the ith node, U i is the effective value of the voltage at the i-th node, is the voltage phase angle of the i-th node;

[0127] The node feature matrix is ​​expressed as:

[0128]

[0129] Among them, X t1 is the node characteristic matrix of the distribution network at time t1, n1 is the number of nodes in the distribution network;

[0130] The attention mechanism matrix is ​​expressed as:

[0131]

[0132] Among them, e t1 is the attention mechanism matrix at time t1, e t1 Each element e in i.j Used to represent the node feature matrix X t1 Medium x i.t1 and x j.t1 The similarity of (can also be described as x i.t1 Relative to x j.t1 importance of

[0133] The element e i.j Expressed as:

[0134]

[0135] Where ReLU is the activation function, a is the trainable parameter vector, W is the trainable parameter matrix used to perform linear transformation on the input features, || represents the concatenation of vectors, and k represents the number of node features contained in the node feature matrix, k = 1, ..., n1;

[0136] The normalized node feature matrix is ​​expressed as The elements in Expressed as:

[0137]

[0138] Among them, E[X t1 ] and Var[X t1 ] are the node feature matrices X t1 , ε1 is the variance control bias, γ1 and β are the learnable parameters of the normalization calculation layer of the dynamic convolutional neural network, and ⊙ represents the Hadamard product (i.e., element-by-element multiplication operation);

[0139] The variance control bias ε1 generally takes an extremely small value. In practical applications, by controlling the variance of the data, the sensitivity of the model to outliers can be reduced, thereby improving the stability of the model and the prediction accuracy. The variance is the normalized distribution amplitude of the data. If there is a very small deviation in the distribution amplitude due to individual outliers (extreme values), it is necessary to set the control bias for correction. Usually, this value is 0. It is set when there is a deviation. The specific setting in engineering calculations needs to consider the deviation after the outlier calculation; γ1 and β are the learnable network parameters of the dynamic convolutional neural network, which are mainly used for the linear extraction of features of normalized data samples. The network parameters are randomly generated during network initialization and are updated according to the characteristics of each element during training to achieve feature extraction.

[0140] The attention weight matrix is ​​expressed as:

[0141]

[0142] Among them, W B.t1 is the attention weight matrix at time t1, σ(·) is the activation function, and the embodiment of the present application uses the sigmoid function as the activation function. is the similarity matrix, for The transposed matrix of

[0143] The sigmoid function is a relatively classic activation function. Its main function is to nonlinearly map the input eigenvalue to the (0, 1) or (-1, 1) region to achieve feature compression. The sigmoid function has good continuity and smoothness. The activation function used in the embodiment of the present application has good applicability for the feature calculation of power grid nodes at continuous moments. At the same time, the good smoothness is conducive to the optimal calculation of model training.

[0144] The similarity matrix Expressed as:

[0145]

[0146] In some possible implementations of the second aspect, the state space is represented by s t1 , state space s t1 The state variables in the one-to-one correspondence are the attention weight matrix W B.t1 The weight elements in the action space are represented as a t1 In this embodiment of the present application, the state transition is used as the action space, a t1 The physical meaning of is represented as a continuous action from the target state of node i to the target state of the next node j under the action of the ε-greedy strategy;

[0147] The reward function is expressed as:

[0148]

[0149] Among them, r i is the reward function corresponding to the training of the i-th node. When the weight element w between the i-th node and the j-th node ij When it is 0, it means that the voltage changes between the two nodes have no mutual influence, and the reward function value is 0. When the weight element w between the i-th node and the j-th node is ij When it is not 0, it means that the voltage changes between the two nodes have a mutual influence relationship, and the reward function value is the weight element w ij and the next state space variable s jWhen designing the reward function, in order to guide the reinforcement learning model state target to calculate and train toward the optimal target, the next state value is added to the reward function, that is, s j , which means that after the i-th node is selected as the target node, the model will select the state value of the next node according to the ε-greedy strategy, that is, the node weight.

[0150] According to the attention weight matrix W B.t1 The design of the reward function is a relatively important innovation of the embodiment of the present application. Traditional graph networks often use the adjacency matrix as the input feature quantity of network information when applying artificial intelligence models. However, in reality, in voltage control, there may be correlation between nodes that are not directly connected during voltage changes. When the adjacency matrix is ​​used, the relationship between nodes that are not directly connected physically is directly represented as "0", which directly ignores the correlation between electrical quantities. Therefore, the embodiment of the present application directly uses the attention weight matrix to design the reward function, which can effectively reflect the degree of mutual influence of node voltage changes and provide a comprehensive, scientific and reliable basis for the selection of the final control mode.

[0151] The ε-greedy strategy is expressed as:

[0152]

[0153] Among them, π(a t1 |s t1 ) is the ε-greedy strategy function; ε2 is the random selection probability, which is ε2=max(0.01, 1-n2 / N), where n2 is the current training number of the model and N is the total training number; argmaxQ(a t1 ,s t1 ) is the maximum Q value corresponding to the optimal action; this strategy indicates that when the action is the optimal choice, its selection probability is 1-ε2, and otherwise the selection probability is ε2. In the early stages of model training, the embodiment of the present application sets ε2 to a numerical value close to 1-n2 / N to represent the probability, enabling the model to fully calculate the optimal target state of the node. As the number of training times increases, the optimal target state is selected with the highest probability, accelerating the calculation speed.

[0154] In some possible implementations of the second aspect, the experience dataset is represented as:

[0155] (s i , a t1 , r i , s j )

[0156] Among them, s i is the current state space variable being trained, at1 is the currently executed action, r i is the corresponding reward function value, s j is the next state space variable selected according to the state transition strategy (i.e., ε-greedy strategy);

[0157] When d sets of experience data sets are selected from the experience pool and input into the training Q network and target Q network of the first deep reinforcement learning model respectively, d <d m , where d m is the data capacity of the experience pool;

[0158] During the training process, the target Q network is calculated using the following formula:

[0159]

[0160] Among them, m is the current iteration number, T is the total number of iterations set in advance, and y m is the target value of the target Q network for the mth time. When m reaches T, the immediate reward r obtained by training the network Q m is the reward value; when m does not reach T, the long-term cumulative reward value (immediate reward r m and the optimal Q value γargmaxQ(a m ,s m ) is the reward value; γ2 is the discount coefficient. The introduction of γ2 is to limit the reward value to avoid infinite reward value;

[0161] During the training process, the following formula is used to update the loss function of the training Q network:

[0162]

[0163] Among them, m is the current iteration number, T is the total number of iterations set in advance, and y m is the target value of the mth target Q network, θ m is the mth loss function; the embodiment of the present application is based on the back propagation algorithm of the neural network, and recalculates the θ value after the loss function is updated and calculated, so that the Q value calculated by the training Q network is closer to the calculated target value of the target Q network.

[0164] It should be noted that the t2 moment mentioned in the mode selection method in the first aspect and the t1 moment mentioned in the model establishment method in the second aspect both refer to "a certain moment" in the operation process of the distribution network node. In order to prevent description confusion and reference confusion, the embodiment of the present application uses t1 and t2 to label a certain moment involved in the two different processes respectively to clearly distinguish the mode selection process and the model establishment process. In actual applications, t1 and t2 can be the same moment or different moments, and no limitation is made here.

[0165] In practical applications, this system can collect grid node data information based on the grid dispatch SCADA system, perform attention weight matrix calculation and first deep reinforcement learning model training every period (for example, 15 minutes), obtain the node occurrence frequency corresponding to each node at each time, and draw the node occurrence frequency curve. Among them, the frequency of each node at time t2 obtained by the first deep reinforcement learning model can be expressed as: L t2 (1) L t2 (2),……,L t2 (n); The frequency of occurrence of node i at consecutive moments can be expressed as: L1(i), L2(i), ..., L t (i) L t+1 (i), ..., L h (i), a total of h moments.

[0166] When the node frequency curve is drawn, the node voltage control mode is selected according to the curve characteristic trend matching. In actual application, the node voltage control mode can correspond to 6 curve types, such as Figure 3 As shown:

[0167] Curve ①: The node occurrence frequency curve continues to fluctuate within a high occurrence frequency range, indicating that this type of node always plays a significant role in the overall voltage control of the distribution network and has great regulation potential, making it an optimal node for centralized control mode.

[0168] Curve ②: The node frequency curve continues to fluctuate within a relatively low range, indicating that this type of node plays a very small role in the overall voltage control of the distribution network and has a weak regulation capability, so it is usually not controlled.

[0169] Curve ③: The frequency of node appearance gradually increases with the grid operation time, indicating that this type of node can play a significant role in the global voltage control of the distribution network under most grid operating conditions and has strong regulation capabilities. It can be selected as a centralized control mode according to the operating conditions.

[0170] Curve ④: The frequency of node appearance gradually decreases with the change of grid operation time, indicating that this type of node plays a small role in the global voltage control of the distribution network under most grid operation conditions and has a weak regulation capability. It can be selected as a local control mode according to the operation conditions;

[0171] Curve ⑤: The node's frequency of occurrence decreases gradually with the grid's operating time, but then gradually increases at a certain moment. This indicates that this type of node can play a significant role in the overall voltage control of the distribution network under certain grid operating conditions and has a strong regulation capability. It is generally controlled in a local control mode, and centralized control can be selected when necessary when operating conditions are met.

[0172] Curve ⑥: The frequency of node appearance gradually increases with the grid operation time, and then gradually decreases at a certain moment. This shows that this type of node can play a significant role in the global voltage control of the distribution network under certain grid operating conditions and has a strong regulation capability. It is generally controlled as a local control mode, and when necessary, centralized control can be selected when the operating conditions are met.

[0173] Based on the above scheme, the method provided in the embodiments of the present application is:

[0174] 1. The active power, reactive power, voltage RMS, and voltage phase angle of each node are comprehensively considered as node voltage characteristics. Compared with the traditional single consideration method, this improves the comprehensiveness and accuracy of node voltage control mode analysis;

[0175] 2. A dynamic attention mechanism is used to weight the voltage characteristics of each node and construct a new training model based on the weight calculation results. This not only effectively, accurately, and comprehensively expresses the voltage variation characteristics of each node, but also comprehensively considers the mutual relationship and mutual influence between the voltage characteristics of each node, realizing the comprehensive determination of the voltage control mode of the distribution network node;

[0176] 3. Deep reinforcement learning is used to analyze and make decisions on characteristic node voltage changes at dynamic moments, giving full play to the algorithm's strong adaptability and generalization capabilities, and effectively solving the complex decision-making problem of continuously changing distribution network voltage;

[0177] 4. The "node occurrence frequency" is proposed for the first time as a variable that is well suited to quantifying the characteristics of node voltage changes. It can reflect the voltage variation level of different nodes in different time periods and under different operating conditions. It can also reflect the node voltage control potential and response ability, providing a reliable basis for the selection of node voltage control mode.

[0178] 5. The node occurrence frequency curve at continuous moments is used as the basis for selecting the node voltage control mode, which fully considers the dynamic performance of the node characteristic variables, achieves the purpose of multi-type and multi-period control mode selection, and realizes the flexibility and adaptability of distribution network node voltage control.

[0179] In a third aspect, an embodiment of the present application provides a distribution network node voltage control system. Exemplarily, the distribution network node voltage control system includes a module for implementing the model establishment method for selecting the distribution network node voltage control mode as described above; or includes a module for implementing the dynamic selection method of the distribution network node voltage control mode as described above.

[0180] In a fourth aspect, an embodiment of the present application provides a power system. Exemplarily, the power system includes the distribution network node voltage control system as described above.

[0181] In the fifth aspect, a device is provided in an embodiment of the present application, which can be any device that can implement the model establishment method for selecting the distribution network node voltage control mode as described above, or can implement the dynamic selection method of the distribution network node voltage control mode as described above. The device can be various terminal devices, such as: desktop computers, laptops, tablet computers, handheld devices, etc., which can be specifically implemented through software and / or hardware.

[0182] Exemplarily, the device includes:

[0183] Memory;

[0184] processor; and

[0185] computer programs;

[0186] In which, the computer program is stored in the memory and is configured to be executed by the processor to implement the model establishment method for selecting the distribution network node voltage control mode as described above, or is configured to be executed by the processor to implement the dynamic selection method of the distribution network node voltage control mode as described above.

[0187] In a sixth aspect, a computer-readable storage medium is provided in an embodiment of the present application. The computer-readable storage medium may be: ROM, RAM, a magnetic disk or an optical disk, etc.

[0188] Exemplarily, a computer program is stored on the computer-readable storage medium; the computer program is executed by the processor to implement the model establishment method for selecting the distribution network node voltage control mode as described above, or to implement the dynamic selection method of the distribution network node voltage control mode as described above.

[0189] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, C language, VHDL language, Verilog language, object-oriented programming language Java, and directly interpreted scripting language JavaScript, etc.

[0190] The present application is described with reference to the flowcharts of the methods, systems, and computer program products according to the embodiments of the present application. It should be understood that each process in the flowchart, as well as the combination of processes in the flowchart, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart. Figure 1 A device that specifies functions in a process or multiple processes.

[0191] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A function specified in a process or multiple processes.

[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 The steps of a specified function in a process or multiple processes.

[0193] In the description of this application, it should be understood that the terms "first," "second," and the like are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Thus, features defined with terms such as "first," "second," and the like may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0194] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0195] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A model establishment method for selecting a voltage control mode for a distribution network node, characterized by: include: Collect voltage characteristics of each node in the distribution network at time t1, wherein the voltage characteristics include active power, reactive power, voltage RMS and voltage phase angle; Based on the voltage characteristics of each node at time t1, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value to measure the relationship between the voltages of each node in the distribution network at time t1; Constructing a state space, action space, and reward function of a deep reinforcement learning model based on the attribute values; Training each state space variable in the state space one by one so that the deep reinforcement learning model reaches the optimal target state during training, and outputting the frequency of occurrence of each node when it reaches the optimal target state; The successfully trained deep reinforcement learning model is used as the first deep reinforcement learning model for the final selection of the voltage control mode of the distribution network node.

2. The model establishment method for selecting a voltage control mode for a distribution network node according to claim 1, characterized in that: Based on the voltage characteristics of each node at time t1, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value for measuring the relationship between the voltages of each node in the distribution network at time t1, specifically including: According to the active power, reactive power, voltage RMS and voltage phase angle of each node in the distribution network at time t1, a voltage characteristic matrix is ​​established to represent the voltage change characteristics of each node in the distribution network at time t1; According to each voltage characteristic matrix, the node characteristic matrix of the distribution network at time t1 is constructed; An attention mechanism is used to calculate the similarity between the features of each node in the node feature matrix, and an attention mechanism matrix is ​​obtained according to the similarity calculation result; After normalizing the node feature matrix, an attention weight matrix is ​​calculated based on the attention mechanism matrix and the normalized node feature matrix; The attention weight matrix is ​​used as the attribute value for measuring the relationship between the voltages of the nodes in the distribution network at time t1.

3. The model building method for selecting a voltage control mode for a distribution network node according to claim 2, characterized in that: The state space, action space, and reward function of the deep reinforcement learning model are constructed based on the attribute values, specifically including: Constructing a state space of a deep reinforcement learning model based on the attention weight matrix, wherein variables in the state space correspond one-to-one to weight elements in the attention weight matrix; Constructing an action space of a deep reinforcement learning model based on the mutual transfer actions between state variables in the state space, wherein each transfer action is executed by an ε-greedy strategy; According to the attention weight matrix, a reward function of the deep reinforcement learning model is constructed: when training a certain state space variable, if the weight element between the current state space variable and other state space variables is 0, the reward function value is 0; if the weight element between the current state space variable and other state space variables is not 0, the reward function value is the sum of the weight element between the two and the next state space variable.

4. The model building method for selecting a voltage control mode for a distribution network node according to claim 3, characterized in that: The deep reinforcement learning model is a DQN network model. The training of each state space variable in the state space one by one so that the deep reinforcement learning model reaches the optimal target state during training and outputs the frequency of occurrence of each node when it reaches the optimal target state specifically includes: Select a state space variable from the state space as the initial state space variable, input it into the training Q network of the first deep reinforcement learning model for training, and generate the corresponding initial training Q value; Execute the action space according to the ε-greedy strategy, calculate the corresponding reward function value according to the reward function, and form the experience data set of each action information to build the experience pool; Select at least one set of experience data sets from the experience pool, input them into the training Q network and the target Q network of the first deep reinforcement learning model respectively, and calculate the corresponding training Q value of the training Q network and the target value of the target Q network respectively; Determine whether the training Q value reaches the target value. If it does, determine whether the model has reached the number of iterations. If it does not, reselect the initial state space variables for a new round of training. If the model reaches the number of iterations, it determines whether the model has calculated the target state for each node; if the model does not reach the number of iterations, it executes the action space according to the ε-greedy strategy and reselects the next state space variable for training until the number of iterations is reached; If the model calculates the target state for each node, the frequency of occurrence of each node when the model reaches the optimal target state is counted and output; if there are nodes that have not been calculated for the target state, the initial state space variables are reselected for a new round of training.

5. The model building method for selecting a voltage control mode for a distribution network node according to claim 4, characterized in that: The voltage characteristic matrix is ​​expressed as: Among them, x i.t1 is the voltage characteristic matrix of the i-th node at time t1, p i is the active power of the i-th node, q i is the reactive power of the ith node, U i is the effective value of the voltage at the i-th node, is the voltage phase angle of the i-th node; The node feature matrix is ​​expressed as: Among them, X t1 is the node characteristic matrix of the distribution network at time t1, n1 is the number of nodes in the distribution network; The attention mechanism matrix is ​​expressed as: Among them, e t1 is the attention mechanism matrix at time t1, e t1 Each element e in i.j Used to represent the node feature matrix X t1 Medium x i.t1 and x j.t1 similarity; The element e i.j Expressed as: Where ReLU is the activation function, a is the trainable parameter vector, W is the trainable parameter matrix used to perform linear transformation on the input features, || represents the concatenation of vectors, and k represents the number of node features contained in the node feature matrix, k = 1, ..., n1; The normalized node feature matrix is ​​expressed as The elements in Expressed as: Among them, E[X t1 ] and Var[X t1 ] are the node feature matrices X t1 The mean and variance of , ε1 is the variance control bias, γ1 and β are the learnable parameters of the normalized calculation layer of the dynamic convolutional neural network, and ⊙ represents the Hadamard product; The attention weight matrix is ​​expressed as: Among them, W B.t1 is the attention weight matrix at time t1, σ(·) is the activation function, is the similarity matrix, for The transposed matrix of The similarity matrix Expressed as:

6. The model building method for selecting a voltage control mode for a distribution network node according to claim 5, characterized in that: The state space is represented as s t1 , the action space is represented as a t1 ; The reward function is expressed as: Among them, r i is the reward function corresponding to the training of the i-th node. When the weight element w between the i-th node and the j-th node ij When it is 0, the reward function value is 0. When the weight element w between the i-th node and the j-th node ij When it is not 0, the reward function value is the weight element w ij and the next state space variable s j of and; The ε-greedy strategy is expressed as: Among them, π(a t1 |s t1 ) is the ε-greedy strategy function; ε2 is the random selection probability, which is ε2=max(0.01, 1-n2 / N), where n2 is the current training number of the model and N is the total training number; argmaxQ(a t1 ,s t1 ) is the maximum Q value corresponding to the optimal action; when the action is the optimal choice, its selection probability is 1-ε2, otherwise the selection probability is ε2.

7. The model building method for selecting a voltage control mode for a distribution network node according to claim 6, characterized in that: The empirical data set is represented as: (s i ,a t1 ,r i ,s j ) Among them, s i is the current state space variable being trained, a t1 is the currently executed action, r i is the corresponding reward function value, s j is the next state space variable selected according to the state transition strategy; When d sets of experience data sets are selected from the experience pool and input into the training Q network and target Q network of the first deep reinforcement learning model respectively, d <d m , where d m is the data capacity of the experience pool; During the training process, the target Q network is calculated using the following formula: Among them, m is the current iteration number, T is the total number of iterations set in advance, and y m is the target value of the target Q network for the mth time. When m reaches T, the immediate reward r obtained by training the network Q m is the reward value; when m does not reach T, the long-term cumulative reward value is the reward value; γ2 is the discount coefficient; During the training process, the following formula is used to update the loss function of the training Q network: Among them, m is the current iteration number, T is the total number of iterations set in advance, and y m is the target value of the mth target Q network, θ m is the m-th loss function.

8. A method for dynamically selecting a voltage control mode for a distribution network node, characterized in that: include: Collect voltage characteristics of each node in the distribution network at time t2, wherein the voltage characteristics include active power, reactive power, voltage RMS value and voltage phase angle; Based on the voltage characteristics of each node at time t2, the attention mechanism is used to calculate the weight of each node, and the weight calculation result is used as the attribute value to measure the relationship between the voltages of each node in the distribution network at time t2; Inputting the attribute degree value into a first deep reinforcement learning model for training to obtain the frequency of occurrence of each node when the model reaches the optimal target state, wherein the first deep reinforcement learning model is a first deep reinforcement learning model established by the model establishment method according to any one of claims 1 to 7; According to the frequency of each node appearing when the model reaches the optimal target state at multiple moments, create the frequency curve of each node appearing at consecutive moments; According to the characteristics of the occurrence frequency curve corresponding to each node, the corresponding node voltage control mode is matched.

9. Distribution network node voltage control system, characterized by: It includes a module for implementing the model establishment method for selecting the distribution network node voltage control mode described in any one of claims 1 to 7; or it includes a module for implementing the dynamic selection method of the distribution network node voltage control mode described in claim 8.

10. Power system, characterized by: It includes the distribution network node voltage control system as claimed in claim 9.

Citation Information

Patent Citations

  • Power distribution network voltage reactive power control method and system based on safety reinforcement learning algorithm

    CN116760047A

  • Power distribution network voltage reactive power control method based on DDQN algorithm

    CN116780562A

  • Method and device for optimizing active power and reactive power of power distribution network and computer program product

    CN118748411A

  • Low-voltage distribution network control method and device, electronic equipment and storage medium

    CN119009980A

  • Power distribution network region voltage control method based on multi-agent deep reinforcement learning

    CN119070315A