Voltage regulation method, system and device for high photovoltaic-penetration distribution transformer area, and medium
By constructing a voltage regulation framework using reinforcement learning agents and training a policy network, the voltage of a high-proportion photovoltaic power station area is regulated using flexible interconnected devices, solving the problems of voltage fluctuation and imbalance, and achieving fast and precise voltage regulation.
Patent Information
- Application Number
- PCT/CN2024/124382
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-23
- Filing Date
- 2024-10-12
- Publication Date
- 2026-02-26
AI Technical Summary
High proportions of photovoltaic (PV) grid connection to distribution substations lead to voltage over-limits and fluctuations. Existing control methods rely on substation model parameters and PV forecasts, which have insufficient applicability and response speed, making it difficult to achieve precise control.
A voltage regulation framework is constructed by defining states, actions, and rewards using reinforcement learning agents. The policy network is trained through a simulation environment, and the voltage regulation strategy is output. The energy storage and photovoltaic grid-connected power of flexible interconnected devices are used for regulation.
Under conditions of unknown transformer parameters and inaccurate photovoltaic forecasts, it achieves precise voltage control within the allowable range, reduces three-phase imbalance, has a fast response speed, and is widely applicable.
Smart Images

Figure CN2024124382_26022026_PF_FP_ABST
Abstract
Description
Voltage regulation method, system, device and medium for high-proportion photovoltaic area
[0001] The present application claims priority from the Chinese patent application No. 202411166421.1 filed on August 23, 2024, and entitled "Voltage regulation method, system, device and medium for high-proportion photovoltaic area", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of new area operation regulation technology, in particular to a voltage regulation method, system, device and medium for high-proportion photovoltaic area. BACKGROUND
[0003] As the "last mile" of power transmission, distribution areas usually have a radial structure. After large-scale household photovoltaic grid connection, the traditional low-voltage distribution network changes from a "single power source" network to a "multi-power source" network, and the distribution characteristics of power flow change fundamentally. Due to the volatility, intermittency and randomness of photovoltaic power generation, after household photovoltaic grid connection, the distribution area will face voltage out-of-limit and fluctuation, three-phase imbalance and other problems.
[0004] With the rapid development of power electronics technology, building a flexible interconnected AC / DC hybrid distribution area with AC as the main and DC as the auxiliary will be the mainstream form of future distribution area development. Compared with AC distribution network, DC distribution network has flexible topology and controllable power flow, and has the potential to suppress the uncertainty of new energy and new type of load. Therefore, it is of great significance to study how to make full use of flexible interconnection devices to solve the voltage quality problems caused by photovoltaic.
[0005] The existing voltage regulation methods for distribution areas mainly include: 1) adjusting the distribution transformer tap changer. This method can adjust the voltage of the end users in the distribution area to a certain extent, but it often leads to high or low voltage of the head-end users. 2) installing reactive power compensation devices. This method can reduce the line voltage drop to a certain extent, but it is difficult to solve the problem of voltage rise caused by photovoltaic access. 3) power-based regulation method. This method mainly reduces the photovoltaic grid-connected power and adjusts the power of distributed energy storage devices. This method relies on the configuration of energy storage and has limited regulation capacity in high-proportion photovoltaic access areas. In addition, this method needs to obtain all the topology parameters, device status and photovoltaic power prediction values of the distribution area, which is difficult to achieve under the existing distribution area collection and measurement capacity.
[0006] SUMMARY
[0007] The present application provides a voltage regulation method, system, device and medium for high-proportion photovoltaic area, which is used for accurate voltage regulation under the condition of unknown distribution area parameters and inaccurate photovoltaic prediction data.
[0008] Therefore, the first aspect of the present application provides a voltage regulation method for a high-proportion photovoltaic area, the method comprising:
[0009] S1, constructing a solution framework for voltage regulation of a high-proportion photovoltaic area by defining the state, action and reward of a reinforcement learning agent;
[0010] S2, training a policy function of the agent by interacting the agent with a simulation environment to obtain a trained policy network;
[0011] S3, inputting an actual environment state to be regulated into the trained policy network to obtain a voltage regulation strategy output by the agent.
[0012] Optionally, the definition of the state, action and reward of the reinforcement learning agent specifically comprises:
[0013] defining the state of the agent as a state of charge (SOC) of a direct-current side energy storage of a flexible interconnected device DC ;
[0014] defining the action of the agent as a charging power P ESS and a cut photovoltaic grid-connected power △P PV of the direct-current side energy storage of the flexible interconnected device;
[0015] defining the reward of the agent as:
[0016] wherein, R t (SOC DC,t |P ESS,t ) is a reward value obtained by the agent in a state SOC DC,t at a time t by executing the action P ESS,t and △P PV,t ; U a,i,t , U b,i,t , U c,i,t are voltage phasors of a node i at a time t in phases a, b and c respectively; α = 1 ∠ 120°; U max , U min are an upper limit and a lower limit of a voltage amplitude of the area respectively.
[0017] Optionally, the training of the policy function of the agent by interacting the agent with the simulation environment to obtain the trained policy network further comprises:
[0018] based on historical operation data of the area, fitting a function relationship between time, state, action and reward by using an LSTM neural network, and taking the LSTM neural network as a simulation environment of the reinforcement learning agent.
[0019] Optionally, the training of the strategy function of the agent by interacting the agent with the simulation environment comprises:
[0020] S21, initializing network parameters of a strategy network and an evaluation network of the agent;
[0021] S22, interacting the agent with the simulation environment through the current strategy network to generate a sampling trajectory;
[0022] S23, calculating a reward value according to the current generated sampling trajectory, and calculating an update gradient of the strategy network and the evaluation network according to the reward value, thereby updating the network parameters of the strategy network and the evaluation network;
[0023] S24, repeating steps S22 to S23 until a preset upper limit value of a training round is reached, thereby obtaining a trained strategy network.
[0024] The second aspect of the application provides a voltage regulation system for a high-proportion photovoltaic area, the system comprising:
[0025] A construction unit is configured to construct a solution framework for the voltage regulation problem of the high-proportion photovoltaic area by defining the state, action and reward of the reinforcement learning agent;
[0026] A training unit is configured to train the strategy function of the agent by interacting the agent with the simulation environment, thereby obtaining a trained strategy network;
[0027] An output unit is configured to input the actual environment state to be regulated into the trained strategy network, thereby obtaining the voltage regulation strategy output by the agent.
[0028] Optionally, the definition of the state, action and reward of the reinforcement learning agent comprises:
[0029] The state of the agent is defined as the state of charge (SOC) of the direct-current side energy storage of the flexible interconnected device DC ;
[0030] The action of the agent is defined as the charging power P ESS and the reduced photovoltaic grid-connected power ΔP PV of the direct-current side energy storage of the flexible interconnected device;
[0031] The reward of the agent is defined as:
[0032] In the formula, R t (SOC DC,t |P ESS,t ) is the reward of the agent in the state SOC DC,t at time t when performing the action PESS,t and ΔP PV,t obtain an environmental feedback reward value; U a,i,t , U b,i,t , U c,i,t are voltage phasors of node i in a, b, c phases at time t respectively; α = 1 ∠120°; U max , U min are the upper limit and the lower limit of the voltage amplitude of the substation respectively.
[0033] Optionally, the method further comprises: a building unit;
[0034] The building unit is configured to fit a functional relationship between time, state, action and reward based on historical operation data of the substation by using an LSTM neural network, and use the LSTM neural network as a simulation environment of the reinforcement learning agent.
[0035] Optionally, the training unit is specifically configured to:
[0036] S21, initialize network parameters of a policy network and an evaluation network of the agent;
[0037] S22, interact the agent with the simulation environment by using the current policy network to generate a sample trajectory;
[0038] S23, calculate a return value according to the current generated sample trajectory, and calculate an update gradient of the policy network and the evaluation network according to the return value, so as to update the network parameters of the policy network and the evaluation network;
[0039] S24, repeat steps S22 to S23 until a preset upper limit value of a training round is reached, to obtain a trained policy network.
[0040] The third aspect of the present application provides a voltage regulation device for a high-proportion photovoltaic substation, the device comprising a processor and a memory:
[0041] The memory is configured to store program code and transmit the program code to the processor;
[0042] The processor is configured to execute the steps of the voltage regulation method for a high-proportion photovoltaic substation according to the instructions in the program code.
[0043] The fourth aspect of the present application provides a computer readable storage medium for storing program code, the program code being used to execute the voltage regulation method for a high-proportion photovoltaic substation.
[0044] From the above technical solutions, the present application has the following advantages:
[0045] The application provides a voltage regulation method for a high-proportion photovoltaic transformer area, comprising the following steps: S1, constructing a solution framework for the voltage regulation problem of the high-proportion photovoltaic transformer area by defining the state, action and reward of a reinforcement learning agent; S2, training a strategy function of the agent by interacting the agent with a simulation environment to obtain a trained strategy network; and S3, inputting an actual environment state to be regulated into the trained strategy network to obtain a voltage regulation strategy output by the agent.
[0046] Compared with the prior art, the regulation method of the application has the following advantages:
[0047] 1) Independent of model parameters: the method provided by the application can solve the voltage regulation strategy of the transformer area under the condition that the transformer area topology parameters and photovoltaic predicted output are both uncertain. This is different from the previous method which requires complete model parameters, and has wider applicability.
[0048] 2) Faster regulation speed: the trained regulation strategy obtained by the application only needs to input the defined state quantity into the strategy network in actual application, and the decision quantity under the current state can be output, without the need to solve complex optimization problems, and the response speed is faster.
[0049] 3) Stronger voltage regulation effect: the method provided by the application can make the voltage of each node of the transformer area within the allowable range, and at the same time minimize the three-phase voltage unbalance degree. BRIEF DESCRIPTION OF DRAWINGS
[0050] Fig. 1 is a flowchart of a voltage regulation method for a high-proportion photovoltaic transformer area provided in an embodiment of the application;
[0051] Fig. 2 is a structural schematic diagram of a voltage regulation system for a high-proportion photovoltaic transformer area provided in an embodiment of the application. DETAILED DESCRIPTION
[0052] In order to enable personnel in the technical field to better understand the application scheme, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0053] Referring to Fig. 1, a voltage regulation method for a high-proportion photovoltaic transformer area provided in an embodiment of the application comprises the following steps:
[0054] Step 101, constructing a solution framework for the voltage regulation problem of the high-proportion photovoltaic transformer area by defining the state, action and reward of a reinforcement learning agent.
[0055] In one embodiment, step 101 specifically comprises:
[0056] The state of the agent is defined as the state of charge (SOC) of the DC side energy storage of the flexible interconnected device DC ;
[0057] The action of the agent is defined as the charging power (P) of the DC side energy storage of the flexible interconnected device ESS and the curtailed grid-connected power of photovoltaic (△P) PV ;
[0058] The reward of the agent is defined as:
[0059] wherein R t (SOC DC,t |P ESS,t ) is the reward value obtained by the agent in state SOC DC ,t at time t by performing action P ESS ,t and △P PV,t ; U a,i,t , U b,i,t , U c,i,t are the voltage phasors of node i in phase a, b and c at time t respectively; α = 1∠120°; U max , U min are the upper limit and lower limit of the voltage amplitude of the substation respectively.
[0060] In one embodiment, step 102 further comprises, before step 102:
[0061] Based on the historical operation data of the substation, a function relationship between time, state, action and reward is fitted by using an LSTM neural network, and the LSTM neural network is used as a simulation environment of the reinforcement learning agent.
[0062] Step 102, by interacting the agent with the simulation environment, the policy function of the agent is trained, and a trained policy network is obtained.
[0063] In one embodiment, step 102 specifically comprises:
[0064] S21, initializing the network parameters of the policy network and the evaluation network of the agent;
[0065] It should be noted that the network parameters θ and φ of the policy network π(A|S, θ) and the evaluation network V(S|φ) of the agent are initialized. Wherein, A represents an action, S represents a state, θ is a network parameter in the policy network π, the policy network π is used to fit the mapping relationship between the state S and the action A, φ is a network parameter in the evaluation network V, and the evaluation network V is used to fit the mapping relationship between the state S and the state-action value function, and the state-action value function represents the expected value of the total reward that can be obtained in the future under the current state and action.
[0066] S22, the agent is interacted with the simulation environment through the current policy network, and a sample trajectory is generated;
[0067] It should be noted that the agent is interacted with the simulation environment 24 times through the current policy network in the embodiment, and the sample trajectory as shown in formula (2) is generated: S t ,A t ,R t ,S t+1 ,A t+1 ,R t1 ,…,S t+23 ,A t+23 ,R t+23 ,S t+24 ;(2)
[0068] Wherein, t is a time period index, S t is a state in the tth time period, A t is an action in the tth time period, and R t is a reward in the tth time period.
[0069] S23, the reward value is calculated according to the sample trajectory generated at present, and the update gradient of the policy network and the evaluation network is calculated according to the reward value, so as to update the network parameters of the policy network and the evaluation network;
[0070] It should be noted that first, according to the sample trajectory generated at present, the reward value G t of the kth time period is calculated.
[0071] Wherein, γ is a discount factor, used to adjust the weight of future rewards, and is taken as 0.7. k is a time period index, and R k is a reward in the kth time period.
[0072] Then, the difference D t between the current reward value and the evaluation network output value is calculated: D t =G t -V(S t |φ);(4)
[0073] Then, the update gradient of the policy network parameter is calculated:
[0074] At the same time, the update gradient of the evaluation network parameter is calculated:
[0075] Finally, the network parameters of the policy network and the evaluation network are updated: θ = θ + ξdθ, φ = φ + ζdφ; (7)
[0076] Wherein, ξ, ζ are learning rates of the policy network and the evaluation network parameter update, respectively taking 0.1 and 0.01.
[0077] S24, repeating steps S22 to S23 until the training round reaches the preset upper limit value, obtaining the trained policy network.
[0078] Step 103, inputting the actual environment state of the voltage to be regulated into the trained policy network to obtain the voltage regulation strategy output by the agent.
[0079] The voltage regulation method for a high-proportion photovoltaic substation provided in the embodiments of the present application can solve the voltage regulation strategy of the substation under the condition that the substation topology parameters and the photovoltaic predicted output are both uncertain. This is different from the previous method which requires complete substation model parameters, and has wider applicability. Moreover, the trained regulation strategy obtained only needs to input the defined state quantity into the policy network in actual application, and can output the decision quantity under the current state, without the need to solve complex optimization problems, and has faster response speed. Further, the regulation method of the present application can make the voltage of each node of the substation within the allowable range, and at the same time minimize the three-phase voltage unbalance degree. Thus, the flexible interconnected substation voltage accurate regulation with high proportion of photovoltaic under the condition that the substation parameters are unknown and the photovoltaic prediction data are inaccurate is realized.
[0080] The above is a voltage regulation method for a high-proportion photovoltaic substation provided in the embodiments of the present application, and the following is a voltage regulation system for a high-proportion photovoltaic substation provided in the embodiments of the present application.
[0081] Please refer to FIG. 2, the voltage regulation system for a high-proportion photovoltaic substation provided in the embodiments of the present application, comprising:
[0082] The construction unit 201 is configured to define the state, action and reward of the reinforcement learning agent, thereby constructing the solving framework of the voltage regulation problem of the high-proportion photovoltaic substation.
[0083] The training unit 202 is configured to train the policy function of the agent by interacting the agent with the simulation environment, thereby obtaining the trained policy network.
[0084] The output unit 203 is configured to input the actual environment state of the voltage to be regulated into the trained policy network, thereby obtaining the voltage regulation strategy output by the agent.
[0085] In one embodiment, the construction unit 201 is specifically used to: define the state of the agent as the state of charge (SOC) of the DC-side energy storage of the flexible interconnect device. DC ;
[0086] The action of the intelligent agent is defined as the charging power P of the DC-side energy storage of the flexible interconnect device. ESS and the reduction in grid-connected photovoltaic power ΔP PV ;
[0087] The reward for the agent is defined as:
[0088] In the formula, R t (SOC DC,t |P ESS,t (SOC) represents the agent's state at time t. DC,t Next, execute action P ESS,t and △P PV,t Rewards are obtained from environmental feedback; U a,i,t U b,i,t U c,i,t Let U be the voltage phasors of node i at time t in phases a, b, and c, respectively; α = 1∠120°; U max U min These are the upper and lower limits of the voltage amplitude in the transformer area, respectively.
[0089] In one embodiment, the voltage regulation system containing a high proportion of photovoltaic power stations provided in this application further includes:
[0090] Building units;
[0091] The aforementioned construction unit is used to fit the functional relationship between time, state, action and reward using an LSTM neural network based on the historical operating data of the transformer area, and to use the LSTM neural network as a simulation environment for the reinforcement learning agent.
[0092] In one embodiment, the training unit is specifically used for:
[0093] S21. Initialize the network parameters of the agent's policy network and evaluation network;
[0094] S22. The agent interacts with the simulation environment through the current policy network to generate a sampling trajectory;
[0095] S23. Calculate the reward value based on the currently generated sampling trajectory, and calculate the update gradient of the policy network and the evaluation network based on the reward value, thereby updating the network parameters of the policy network and the evaluation network;
[0096] S24, repeating steps S22 to S23 until the training round reaches a preset upper limit value, to obtain a trained strategy network.
[0097] Further, the embodiment of the present application also provides a voltage regulation device containing a high proportion of photovoltaic blocks, the device comprising a processor and a memory:
[0098] The memory is configured to store program code and transmit the program code to the processor.
[0099] The processor is configured to execute the steps of the voltage regulation method for the high proportion of photovoltaic blocks according to the instructions in the program code.
[0100] Further, the embodiment of the present application also provides a computer readable storage medium, which is configured to store program code, and the program code is configured to execute the voltage regulation method for the high proportion of photovoltaic blocks.
[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the system and the unit described above can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0102] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and the above drawings of the present application are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0103] It should be understood that, in the application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases of only A, only B and A and B existing at the same time, wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent a, b, c, "a and b", "a and c", "b and c", or "a and b and c", wherein a, b and c can be single or multiple.
[0104] In several embodiments provided in the application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0105] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0106] In addition, each functional unit in each embodiment of the application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0107] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (English full name: Read-Only Memory, English abbreviation: ROM), a random access memory (English full name: Random Access Memory, English abbreviation: RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0108] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A voltage regulation method for a photovoltaic power station area with a high proportion of photovoltaic power, characterized in that, The method comprises the steps of: S1, constructing a solution framework for a voltage regulation problem of a high proportion photovoltaic substation by defining a state, an action and a reward of a reinforcement learning agent; S2, training a policy function of the agent by interacting the agent with a simulation environment to obtain a trained policy network; S3, inputting an actual environment state to be regulated into the trained policy network to obtain a voltage regulation strategy output by the agent.
2. The method of claim 1, wherein the voltage regulation method is applied to a high PV penetration zone. The definition of the state, the action and the reward of the reinforcement learning agent specifically comprises: The state of the agent is defined as the state of charge (SOC) of the energy storage on the DC side of the flexible interconnection device DC ; The action of the intelligent agent is defined as the charging power P of the flexible interconnected device direct current side energy storage ESS And the reduced photovoltaic grid-connected power ΔP PV ; The reward of the agent is defined as: wherein R t (SOC DC,t |P ESS,t is the reward value of the environment feedback obtained by the agent performing action P DC,t and △P ESS,t at time t; U PV,t , U a,i,t , U b,i,t , U c,i,t are the voltage phasors of node i at phase a, b, and c at time t, respectively; α = 1∠120°; U max , U min are the upper and lower limits of the voltage amplitude of the transformer substation, respectively.
3. The method of claim 1, wherein the voltage regulation method is applied to a high- PV-penetrated distribution area. The agent is further interacted with the simulation environment to train the policy function of the agent to obtain the trained policy network. The agent is further interacted with the simulation environment to train the policy function of the agent to obtain the trained policy network.
4. The method of claim 1, wherein the voltage regulation method is applied to a high PV penetration zone. The agent is further interacted with the simulation environment to train the policy function of the agent to obtain the trained policy network. S21, initializing network parameters of a policy network and an evaluation network of the agent; S22, interacting the agent with the simulation environment through the current policy network to generate a sampling trajectory; S23, calculating a reward value according to the currently generated sampling trajectory, and calculating an update gradient of the policy network and the evaluation network according to the reward value, so as to update the network parameters of the policy network and the evaluation network; S24, repeating steps S22 to S23 until a preset upper limit value of a training round is reached to obtain the trained policy network.
5. A voltage regulation system for a high photovoltaic penetration area, comprising: The method comprises the steps of: constructing a solution framework for a voltage regulation problem of a high proportion photovoltaic substation by defining a state, an action and a reward of a reinforcement learning agent; training a policy function of the agent by interacting the agent with a simulation environment to obtain a trained policy network; inputting an actual environment state to be regulated into the trained policy network to obtain a voltage regulation strategy output by the agent.
6. The high PV percentage zone containing voltage regulation system according to claim 5, wherein, The definition of the state, the action and the reward of the reinforcement learning agent specifically comprises: The state of the agent is defined as the state of charge (SOC) of the energy storage on the DC side of the flexible interconnection device DC ; The action of the intelligent agent is defined as the charging power P of the flexible interconnected device direct current side energy storage ESS and the reduced photovoltaic grid-connected power ΔP PV ; The reward of the agent is defined as: In the formula, R t (SOC DC,t |P ESS,t ) is the reward value of the environment feedback obtained by the agent performing action P DC,t and △P ESS,t at time t; U PV,t , U a,i,t , U b,i,t , U c,i,t are voltage phasors of node i at phase a, b and c at time t respectively; α = 1∠120°; U max , U min are the upper and lower limits of the voltage amplitude of the transformer area respectively.
7. The high PV percentage zone containing voltage regulation system of claim 5, wherein, The method further comprises the steps of: The method further comprises the steps of: The LSTM neural network is used to fit a functional relationship among time, state, action and reward, and is used as a simulation environment of the reinforcement learning agent. The training unit is specifically configured to:
8. The high PV percentage zone containing voltage regulation system of claim 5, wherein, S21, initializing network parameters of a policy network and an evaluation network of the agent; S22, interacting the agent with the simulation environment through the current policy network to generate a sampling trajectory; S23, calculating a reward value according to the currently generated sampling trajectory, and calculating an update gradient of the policy network and the evaluation network according to the reward value, so as to update the network parameters of the policy network and the evaluation network; S24, repeating steps S22 to S23 until a preset upper limit value of a training round is reached to obtain the trained policy network. The device comprises a processor and a memory:
9. A voltage regulating device for a high photovoltaic penetration area, characterized in that, The memory is configured to store program code and transmit the program code to the processor; The processor is configured to execute the voltage regulation method for a high proportion photovoltaic area according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store program code for executing the voltage regulation method for a high proportion photovoltaic area according to any one of claims 1-4.
Citation Information
Patent Citations
Microgrid load coordination control simulation system and modeling method
CN106777673A
Reactive voltage control method based on multi-time-scale multi-agent deep reinforcement learning
CN113363997A
Multi-agent reinforcement learning method for distributed resource collaborative scheduling
CN116542137A
Active power distribution network real-time voltage control method based on reactive power regulation of photovoltaic inverter
CN118316135A
Training action selection neural networks using a differentiable credit function
US20200175364A1
Cited By
Photovoltaic area optimization regulation and control method, system, equipment and medium
CN121965815A