Power distribution network voltage control method based on sample scene classification reinforcement learning

By using a reinforcement learning method based on sample scenario classification, the problem of reward sparsity and distribution deviation in the voltage control of power distribution networks by the RL algorithm is solved, and more efficient voltage control and lower line loss are achieved.

CN120879629APending Publication Date: 2025-10-31ZHUHAI UNIV OF SCI & TECH RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511231391.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-31
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional RL algorithms face problems of reward sparsity and experience distribution bias in distribution network voltage control, which makes it impossible for agents to be trained effectively and to cope with voltage fluctuations caused by the instability of green electricity output.

Method used

We employ a reinforcement learning approach based on sample scene classification, dividing experiences into smooth and non-smooth scene sets. We use the TD3 algorithm and scene classification experience replay technology, combined with curiosity-driven rewards and compensators, to optimize the training process.

Benefits of technology

It improved training efficiency by 37.5%, reduced line loss costs and voltage failure rate, and achieved more efficient voltage control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120879629A_ABST
    Figure CN120879629A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of power distribution, and discloses a power distribution network voltage control method based on sample scene classification reinforcement learning. Two memory buffer areas are introduced to recognize different experiences, a scene classification experience playback method is provided, training of the RL agent on typical voltage violation related experiences can be enhanced by designing experience sampling weights, and it is guaranteed that the agent can be effectively converged to a better strategy. According to the invention, an RL scheme based on a compensator is introduced to enhance the robustness of training, and the problem of reward sparseness is solved. According to the power distribution network voltage control technology based on sample scene classification reinforcement learning, the voltage safety under random photovoltaic output can be effectively ensured, and the line loss cost is reduced. The example research verifies that compared with the traditional reinforcement learning algorithm, the method provided by the invention improves the sample efficiency from 4000 times to 2500 times, and the sample efficiency is improved by 37.5%. Meanwhile, the line loss cost and the maximum voltage default rate are respectively reduced by 7.5% and 50%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power distribution, specifically relating to a power distribution network voltage control method based on sample scene classification reinforcement learning. Background Technology

[0002] With increasing emphasis on environmental protection, the deployment of renewable green electricity such as wind power, photovoltaic power, and tidal power is growing. The generation process of green electricity is affected by various factors, resulting in highly unstable power output and frequent, significant voltage fluctuations. For example, overvoltage often occurs when photovoltaic power generation peaks at midday while load demand is relatively low; at night, voltage sag is most common due to low photovoltaic power generation. In practice, significant overvoltage or undervoltage can lead to unsafe operating conditions in the distribution network, such as violations of equipment thermal limits or changes in fault current levels. Furthermore, voltage violations can increase current flow in lines, reduce the power factor, and thus increase line losses due to unnecessary heat loss. Therefore, voltage control in the distribution network is crucial for maintaining voltage stability and reducing network losses. Traditional equipment (such as on-load tap-changing transformers) cannot provide frequent regulation and generally can only coordinate the control of reactive and active power by green electricity and regional power consumption systems.

[0003] Traditionally, a widely used approach for optimizing active and reactive power in distribution networks is optimal power flow using convex relaxation techniques, such as second-order cone relaxation. However, physical model-based optimization methods heavily rely on accurate distribution network models, which may be impractical in real-world scenarios. Furthermore, the computational cost of optimization increases rapidly with network size, leading to a significant computational burden. Power Reduction (RL), as a model-free approach, has shown great potential in solving challenging control problems in power systems, including voltage control.

[0004] While Reinforcement Algorithms (RL) can handle the most complex continuous control tasks, two foreseeable problems exist in practical applications. The first is the reward sparsity problem. Typically, RL agents converge to the optimal policy based on a large amount of reward feedback. In the real world, the relatively few over / under voltage scenarios throughout the day mean that RL agents cannot obtain sufficient feedback from exploration. Secondly, when predictive information (such as photovoltaic output or load forecasting) is considered as input to describe uncertainty, a double layer of stochasticity is introduced to the RL agent. That is, the RL agent needs to face more random state transitions, which may reduce its generalization ability across different scenarios. The second problem is the distribution bias of collected experience due to the irregular occurrence of voltage violations. The agent may collect far less experience for over / under voltage scenarios than for normal scenarios, leading to a biased experience distribution. However, traditional experience replay techniques in most RL algorithms fail to work with this distribution bias problem, resulting in trained agents that are not proficient in handling voltage violations. This is because the sampling probability of experience-related violation scenarios based on the bias distribution is relatively low. Summary of the Invention

[0005] The purpose of this invention is to overcome at least one deficiency of the prior art and provide a power distribution network voltage control method based on sample scene classification reinforcement learning.

[0006] To address the problem of skewed experience distribution caused by irregular voltage violations, this invention introduces two memory buffers to identify different experiences (i.e., all experiences). Dividing the scenarios into smooth and non-smooth sets, this invention proposes a scenario classification experience replay method. By designing experience sampling weights, the RL agent can strengthen its training on typical voltage violation-related experiences, ensuring that the agent can effectively converge to a better policy. Simultaneously, this invention introduces a compensator-based RL scheme to enhance training robustness and address the reward sparsity problem.

[0007] The technical solution adopted in this invention is:

[0008] A voltage control method for power distribution networks based on sample scene classification reinforcement learning includes the following steps: Treating the distribution network as an environment, define the state space of Markov decisions. Action space and reward function ,in: For each time step Environmental information of power distribution network The state space is composed of the bus voltage amplitude and real-time active / reactive power injection, and is defined as follows: , Where, vector Let K be the operating power of all DCS systems connected to the distribution network, and K be the set of DCS devices; vector Including reactive power injection from all green inverters, N is the set of inverter devices; voltage vector B contains the voltage magnitudes of all buses; B is the set of buses; vector. For the historical output of green electricity, For the predicted active power output of green electricity; The variables in the voltage control problem are the active power of the regional power consumption system and the reactive power of the green inverter. The reactive power of the green inverter changes the reactive power setpoint. Direct control is used to regulate the active power of regional power consumption systems. The action space is defined as: The size of the motion space is ; The reward function aims to reduce line losses and mitigate the discomfort caused by voltage violations and regional power consumption systems in the distribution network, at the time step. Define rewards It consists of the following three parts: , Among the rewards The first component is the cost of controlling the total line loss of the target; rewards The second component represents the cumulative voltage violations of the distribution network, in which It is the lower voltage limit requirement. It is the upper limit requirement for voltage; reward The third component is the inadequacy of the regional power consumption system, parameters , and These are the positive weight parameters for the three reward items under consideration, obtained through parameters. To achieve the control objective, a large negative reward is directly proportional to the magnitude of the line loss; using parameters... and Constraint violations are penalized based on voltage violations and inadequacies in regional power consumption systems, respectively.

[0009] Adjust the active power according to the differences in power consumption systems in different regions. For energy storage devices (such as thermal storage, cold storage, and electrical storage), the power can be directly adjusted, as long as the active power can be effectively adjusted. That's fine. The suitability of regional power consumption systems will vary depending on the specific type of the power consumption system. For example, for energy storage systems, the power can be directly calculated, so there is no need to introduce other indirect power control parameters into the state space; the charging and discharging power can be used directly.

[0010] In some examples of distribution network voltage control methods, critique networks are used. Evaluate the expected total reward of the agent's state-action pair. The above consists of two neural networks (l=1,2) constitutes the structure. Used to assess the state Next, take action The value of, among which (l=1,2) is a network The corresponding parameters for (l=1,2) - The function satisfies the following recurrence relation: , in Used to indicate in the strategy Next state action pair The value of It is a discount factor that takes into account rewards for future steps. Representing state-action pairs The next state after that; The Critical Network Used to train an optimal actor network Determine the maximum value corresponding to each state. -function value action .

[0011] In some examples of distribution network voltage control methods, the optimal actor network is trained. At the same time, the TD3 algorithm, based on scene classification and experience playback, strengthens the accumulation of experience regarding voltage violations, and integrates all the experience... Divided into smooth scene sets and non-smooth scene sets, among which: The smoothing scenario set is defined as: all bus voltages in the distribution network, including the current state. and the next state If all strictly adhere to the upper and lower voltage limits, then it belongs to the smooth scene set, represented as: The definition of a non-smooth scenario is: the voltage of any bus in the distribution network, including the current state. and the next state If a violation of the upper or lower voltage limit exists, it belongs to the set of non-smooth scenarios, represented as: , All conversions or The corresponding set size is used and This means that during the training process, each parameter update process selects a batch of samples.

[0012] In some examples of distribution network voltage control methods, TD error is used for each scenario set. Calculate the sampling priority for each sample Thus, its sampling probability is derived. for: in A positive value ensures that the sampling probability is not zero. express of Power of 1 It is a factor that adjusts the degree of influence of priority on the empirical sampling probability distribution. At that time, all experiences are sampled with the same priority.

[0013] In some examples of distribution network voltage control methods, the set size is used to assign significant sampling weights to each scenario. Defined as: , in (0,1] is the hyperparameter of the importance sampling weights, which will be used to sub-batch Kazuko Pi The sampling transformation forms the entire batch of samples for training.

[0014] In some examples of distribution network voltage control methods, curiosity-driven rewards are introduced to evaluate the current state. and the predicted next state The similarity between states reflects the probability of future state transitions and predicts step changes in states.

[0015] In some examples of distribution network voltage control methods, the Pearson correlation coefficient is used. The similarity is calculated using the range [-1,1], and the formula is as follows: Cov and Var Let represent the covariance and variance function, respectively.

[0016] In some examples of power distribution network voltage control methods, curiosity-driven rewards are used. To intervene in the training of the intelligent agent: , , in To introduce a weighting factor for the compensator reward; It's the final reward.

[0017] In some examples of power distribution network voltage control methods, the regional power consumption system is a cooling system, and the control command for the cooling system is based on the mass flow rate of the regional cooling system's water supply. Adjust its active power The action space is defined as: , , award The third component is compensation for unsuitable temperature deviations in the terminal buildings of the cooling system; express ,parameter , and These are the positive weight parameters for the three reward items under consideration, obtained through parameters. To achieve the control objective, a large negative reward is directly proportional to the magnitude of the line loss; using parameters... and Constraint violations are penalized based on voltage inconsistency and temperature inconsistency, respectively.

[0018] In some examples of power distribution network voltage control methods, the green electricity is a photovoltaic power generation system.

[0019] In some examples of distribution network voltage control methods, the distribution network is based on the IEEE-33 bus system.

[0020] The above features can be combined arbitrarily as long as they do not conflict.

[0021] The beneficial effects of this invention are: The voltage control technology for distribution networks based on sample scenario classification reinforcement learning proposed in this invention can effectively ensure voltage safety under random photovoltaic output and reduce line loss costs. Case studies verify that, compared with traditional reinforcement learning algorithms, the proposed method improves the sample efficiency from 4000 times to 2500 times, an improvement of 37.5%. Simultaneously, line loss costs and the maximum voltage default rate are reduced by 7.5% and 50%, respectively. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the reinforcement learning algorithm framework based on sample scene classification designed in this invention.

[0023] Figure 2 This is a schematic diagram of the compensator designed in this invention.

[0024] Figure 3 This is a schematic diagram of the sample classification logic designed in this invention.

[0025] Figure 4 The chart shows a comparison of constraint violations and line loss costs for four reinforcement learning algorithms controlling voltage, with (a) daily line loss cost and (b) daily maximum voltage violation.

[0026] Figure 5 This is a comparison chart of the training results of four reinforcement learning algorithms. Detailed Implementation

[0027] The following description, in conjunction with the embodiments and accompanying drawings, provides further details. Example 1

[0028] In this embodiment, the distribution network is a distribution network based on the IEEE-33 bus system, the green electricity is a photovoltaic power generation system, and the regional power consumption system is a regional cooling system.

[0029] like Figure 1 As shown, the power distribution system operator is treated as an intelligent agent controller, sending online regulation signals to each photovoltaic inverter and district cooling system. Then, by designing the corresponding state space, action space, and reward function for the agent's tasks, the voltage control problem is reformulated as a mathematical model of a Markov decision process. For the voltage control problem, the power distribution network is considered as the environment, and the state space of the Markov decision process is defined. Action space and reward function The definition is as follows: (1) State space For each time step Environmental information of power distribution network It consists of the bus voltage amplitude and real-time active / reactive power injection. In this example, the state space is defined as: Where, vector Let K be the operating power of all DCS systems connected to the distribution network, and K be the set of DCS devices; vector Including reactive power injection from all photovoltaic inverters, where N is the set of inverter devices; voltage vector. B contains the voltage magnitudes of all buses, and B is the set of buses. Here, a vector is introduced. and This describes the historical and predicted photovoltaic active power output. Here, to capture the long-term characteristics of photovoltaic power generation, a length of [length value] is defined. The power and length of the historical time series are The predicted time series powers are as follows: , , in and These represent the actual active power injection from photovoltaic (PV) power plants and the predicted active power output from PV power plants in the distribution network, respectively. Both symbols are vectors containing all data from the PV power plant, represented as: , , Among them, the vector of the first element Indicates the first The actual active power of each photovoltaic power station This represents the system's predicted power at the corresponding time.

[0030] (2) Action space The variables in the voltage control problem are the active power output of the district cooling system and the reactive power of the photovoltaic inverter. This example designs the operation for different types of controllable devices. The intelligent inverter changes the reactive power setpoint. Direct control. For district cooling systems, this example defines the control command as the district cooling system water supply mass flow rate. Adjust its active power Therefore, the action space is defined as: , The size of the motion space is .

[0031] (3) Reward function This example, based on the objectives and constraints of the voltage control problem, uses a reward function designed to reduce line losses and mitigate voltage violations and building temperature incompatibilities in the distribution network. Therefore, at time step... This example defines a reward. It consists of the following three parts: , , , Among the rewards The first component is the cost of total line losses, which is the control objective. (Reward) The second component represents the cumulative voltage violations of the distribution network, in which It is the lower voltage limit requirement. This is the voltage upper limit requirement. (Reward) The third component is compensation for unsuitable temperature deviations in terminal buildings. express .parameter , and These are the positive weight parameters for the three reward items under consideration.

[0032] This example uses parameters. To achieve the control objective. In this case, a large negative reward is proportional to the magnitude of the line loss. Then, parameters are used. and Constraint violations are penalized separately for voltage violations and temperature incompatibilities. These three weighting parameters are designed to maintain a balance between the objective and the constraints. In practice, their values ​​can be flexibly adjusted to meet different control requirements. For example, increasing... To encourage a demand-driven response to voltage stability, or to increase To avoid significant temperature discomfort (sacrificing voltage regulation capability).

[0033] After constructing the Markov decision process, the core algorithm needs to be determined. To avoid performance issues caused by overfitting and high model variance in traditional reinforcement learning algorithms, this example employs the Dual-Delay Deep Deterministic Policy Gradient (TD3) algorithm to circumvent these drawbacks. This example deploys two neural networks. (l=1,2) is used to represent the critique of the internet. Used to assess the state Next, take action The value of, among which (l=1,2) is a network The corresponding parameters for (l=1,2). Criticism network. The aim is to estimate an important function called the action safety value function or - A function that evaluates the expected total reward of an agent's state-action pair. , - The function satisfies the following recurrence relation: , in Used to indicate in the strategy Next state action pair The value of It is a discount factor that takes into account rewards for future steps. Representing state-action pairs The next state is as follows. The TD3 algorithm has six networks, and its goal is to train an optimal actor network. The optimal action network is a mapping rule that can determine the maximum value corresponding to each state. -function value action Therefore, the primary goal during training is to obtain accurate... - Function. However. The network parameters of the function are initially randomized. Then, during training, the parameters are adjusted... - The function is iteratively updated to obtain a more accurate estimate.

[0034] Furthermore, to address the limited experience accumulation due to the low frequency of voltage violations during TD3 training, this example proposes a TD3 algorithm based on scene classification and experience playback. The key idea is to effectively capture and utilize sparse but important historical experience. A schematic diagram of the sample classification logic is shown below. Figure 3 As shown. First, the variables are given. This is used to mathematically represent experience (also known as a transition). We then categorize all experiences into two scenarios:

[0035] (1) Smooth scene

[0036] For transition If the voltage of all busbars in the distribution network, including the current state and the next state If all strictly adhere to the upper and lower voltage limits, then it belongs to the smooth scene set: .

[0037] (2) Unsmooth scenes For transition If the voltage of any bus in the distribution network, including the current state and the next state If there is a violation of the upper or lower voltage limit, it belongs to the non-smooth scenario set: .

[0038] Based on the above definition, all transformations can be divided into the two scenario sets mentioned above. and The corresponding set size is used and This is explained below. During training, each parameter update process selects a batch of samples. This example defines the sample proportion of each sub-batch to calculate the number of samples selected from each scene set. Secondly, for each scene set, the sampling priority for each experience needs to be defined. Similar to the priority experience replay method, TD error is used. Calculate the sampling priority for each sample Thus, its sampling probability is derived. for: in A positive value ensures that the sampling probability is not zero. express of Power of 1 It is a factor that adjusts the degree to which priority affects the empirical sampling probability distribution. Especially when In this case, all experiences are sampled with the same priority; in TD3, this method degenerates into traditional uniform random sampling. The two scene classifications in this example can lay the model foundation for subsequent sampling weights.

[0039] Finally, in order to improve this - To improve the estimation accuracy of the function and avoid training convergence failure, this example redefines the importance sampling weights to correct sampling bias and gradients. Considering there are two scene sets, the importance sampling weights corresponding to each transition are allocated using the set size. Defined as: in (0,1) represents the hyperparameters of the importance sampling weights. Finally, based on the above equations, two scene sets (i.e., sub-batches) are... Kazuko Pi The sampling transformation is used to form the entire batch of samples for training.

[0040] Building upon the sample scenario classification, this example introduces curiosity-driven rewards to evaluate the current state. and the predicted next state The similarity between states. Similarity can reflect the probability of future state transitions and predict step changes in states, such as a sudden increase or decrease in PV. This example uses the Pearson correlation coefficient. The similarity is calculated using the range [-1,1], and the formula is as follows: , Cov and Var Let these represent the covariance and variance functions, respectively. Then, this example defines a compensated curiosity-driven reward. To intervene in the training of the agent and solve the problem of sparse rewards: , , in To introduce a weighting factor for the compensator reward; This is the final reward. In this example, weaker similarity leads to... The smaller the value, the greater the compensation reward. This facilitates the agent's exploration. The weaker the similarity between states, the greater the probability of a state transition, thus prompting the agent to adjust the voltage in advance. Therefore, the core idea of ​​the compensator proposed in this example is to help the agent predict state transition probabilities to achieve real-time response to system voltage fluctuations. A schematic diagram of the compensator is shown below. Figure 2 As shown.

[0041] Validation was performed on a distribution network based on the IEEE-33 bus system. PV fluctuation data is based on actual data collected in Australia from March 1, 2021 to August 31, 2022, from the Pecan Street Dataport.

[0042] For the aforementioned distribution network voltage control model, this example generates the corresponding SCER-TD3 strategy. To verify the superiority of the proposed algorithm, this example compares it with three benchmarks: 1) the traditional TD3 algorithm; 2) the compensator-based TD3 algorithm (CTD3); and 3) the compensator-based Priority Experience Replay TD3 algorithm (PER-CTD3). The first benchmark, TD3, is chosen to demonstrate the effectiveness of the compensator. The other two benchmarks, CTD3 and PER-CTD3, are considered to prove the effectiveness of SCER because the traditional uniform sampling used in CTD3 and the Priority Experience Replay used in PER-CTD3 are two advanced replay techniques.

[0043] In this example, to obtain fair verification results, the reinforcement learning agents used are all configured identically, including the number of neural network layers, the number of neurons, the optimizer, and the learning rate.

[0044] In this example, the economics of the proposed algorithm are evaluated based on the average line loss cost using two months of test set data; the safety of system constraints is assessed using the daily average maximum voltage deviation. The results are presented using box plots. Figure 4 As can be seen, the average values ​​are highlighted with a red rectangle. The proposed SCER-CTD3 consistently maintains voltage safety, with its enclosure always within the range of [-0.05, 0.05] pu. In contrast, the other three strategies all achieved a minimum voltage violation of -0.2 pu, while their maximum voltage violations all exceeded 0.1 pu. Line loss costs and the maximum voltage default rate were reduced by 7.5% and 50%, respectively. Therefore, the proposed SCER-CTD3 can effectively prevent unpredictable voltage fluctuations caused by photovoltaics.

[0045] Figure 5The training process of four algorithms over 8000 episodes is demonstrated. The convergence results show that the TD3 agent struggles to converge due to sparse rewards. In contrast, the other three agents with compensators all exhibit successful convergence. In terms of convergence efficiency, CTD3 and PER-CTD3 perform similarly, stabilizing at a near-perfect reward level after 4000 episodes. However, the proposed SCER-CTD3 significantly outperforms all three benchmarks, converging to a maximum reward level close to -100, with the fastest convergence speed at 2500 episodes, representing a 37.5% improvement. Therefore, the proposed SCER-CTD3 demonstrates higher training efficiency and convergence reward.

[0046] In summary, compared with traditional reinforcement learning algorithms, the proposed sample-based scene classification reinforcement learning can improve training security and policy convergence optimality. Specifically, this method improves training sample efficiency by 37.5%, and reduces line loss cost and maximum voltage default rate by 7.5% and 50%, respectively.

[0047] The above is a further detailed description of the present invention and should not be considered as a limitation on the specific implementation of the present invention. For those skilled in the art, simple deductions or substitutions without departing from the concept of the present invention are all within the protection scope of the present invention.

Claims

1. A voltage control method for power distribution networks based on sample scene classification reinforcement learning, comprising the following steps: Treating the distribution network as the environment, we define the state space s of the Markov decision. t ∈S, action space a t ∈A and reward function r t ∈R, where: For each time step t, the environmental information s of the distribution network t ∈S consists of the bus voltage magnitude and real-time active / reactive power injection, and its state space is defined as: Where, vector Let K be the operating power of all DCS systems connected to the distribution network, and K be the set of DCS devices; vector. Including reactive power injection from all green inverters, N is the set of inverter devices; voltage vector V t =(V b,t |b∈B) contains the voltage magnitudes of all buses, where B is the set of buses; vector For the historical output of green electricity, For the predicted active power output of green electricity; The variables in the voltage control problem are the active power of the regional power consumption system and the reactive power of the green inverter. The reactive power of the green inverter changes the reactive power setpoint. Direct control is used to regulate the active power of regional power consumption systems. Action space is defined as: The size of the action space is |N|+|K|; The reward function aims to reduce line losses and mitigate voltage violations and regional power consumption issues in the distribution network. At time step t, the reward r is defined. t ∈R consists of the following three parts: Among the rewards The first component is the cost of controlling the target total line loss; reward R VV The second component of (t) represents the cumulative voltage violations of the distribution network, where V is the lower voltage limit requirement. It is the upper limit requirement of voltage; bonus R DT The third component of (t) is the inadequacy of the regional power consumption system. Parameters λ1, λ2 and λ3 are the positive weight parameters of the three reward items under consideration. The control objective is achieved through parameter λ1. Large negative rewards are proportional to the magnitude of line losses. The parameters λ2 and λ3 are used to penalize constraint violations based on voltage violations and regional power consumption system incompatibilities, respectively.

2. The power distribution network voltage control method according to claim 1, characterized in that, The expected total reward (s) of an agent's state-action pair is evaluated using a critical network Q. t ,a t The above consists of two neural networks. constitute, Used to evaluate in state s " Next, take action a " The value of , where θ l (l=1,2) is a network For the corresponding parameters, the Q-function satisfies the following recurrence relation: Q π (s " ,a " ) is used to represent state-action pairs (s) under policy π. " ,a " The value of γ, where γ∈(0,1] is a discount factor considering rewards for future steps. Representing state-action pairs (s " ,a " The next state after ) The criticism network Q is used to train an optimal actor network. Determine the maximum Q-function value for each state. action a " =π(s) " ).

3. The power distribution network voltage control method according to claim 2, characterized in that, Training the optimal actor network At the same time, the TD3 algorithm, based on scene classification and experience playback, strengthens the accumulation of experience regarding voltage violations, and integrates all the experience... Divided into smooth scene sets and non-smooth scene sets, among which: The smoothing scenario set is defined as: all bus voltages in the distribution network, including the current state s. " and the next state If all strictly adhere to the upper and lower voltage limits, then it belongs to the smooth scene set, represented as: The definition of a non-smooth scenario is: any bus voltage in a distribution network, including the current state s. " and the next state If there is a violation of the upper or lower voltage limit, it belongs to the set of non-smooth scenarios, which is represented as: All conversions or The corresponding set size is used and This means that during the training process, each parameter update process selects a batch of samples.

4. The power distribution network voltage control method according to claim 3, characterized in that, For each scenario set, the sampling priority p of each sample is calculated using the TD error δt. " Therefore, its sampling probability P is obtained. " for: p " =d " +v, Where ε is a positive value to ensure that the sampling probability is not zero. p " The power of α, where α∈[0,1] is a factor that adjusts the degree of influence of the priority on the empirical sampling probability distribution. When α=0, all empirical sampling priorities are the same.

5. The power distribution network voltage control method according to claim 3, characterized in that, Use the set size to assign important sampling weights w to each scene " Defined as: Where β∈(0,1] is the hyperparameter of the importance sampling weight, and the sub-batch Kazuko Pi The sampling transformation forms the entire batch of samples for training.

6. The power distribution network voltage control method according to claim 1, characterized in that, Introduce curiosity-driven rewards to evaluate the current state. and the predicted next state The similarity between states reflects the probability of future state transitions and predicts step changes in states.

7. The power distribution network voltage control method according to claim 6, characterized in that, Using the Pearson correlation coefficient ρ t The similarity is calculated using the range [-1, 1], and the formula is as follows: Where Cov(·) and Var(·) represent the covariance and variance functions, respectively.

8. The power distribution network voltage control method according to claim 1, characterized in that, Use curiosity-driven rewards To intervene in the training of the intelligent agent: Where λ com To introduce a weighting factor for the compensator reward; r t It's the final reward.

9. The power distribution network voltage control method according to claim 1, characterized in that, The regional power consumption system is a cooling system. The control command for the cooling system is based on the mass flow rate (m) of the regional cooling system's water supply. k,t Adjust its active power Action space is defined as: Reward R DT The third component of (t) is compensation for unsuitable temperature deviations in the terminal buildings of the cooling system; [·] + [x] + =max(0,x), where parameters λ1, λ2 and λ3 are the positive weight parameters of the three reward items under consideration. The control objective is achieved through parameter λ1, where a large negative reward is proportional to the magnitude of the line loss. Parameters λ2 and λ3 are used to punish constraint violations from the perspectives of voltage violation and temperature incompatibility, respectively.

10. The power distribution network voltage control method according to claim 1, characterized in that, The green electricity refers to a photovoltaic power generation system, and / or the distribution network is a distribution network based on the IEEE-33 bus system.