A virtual power plant large-scale distributed resource intelligent regulation method
Patent Information
- Application Number
- CN202611231623.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-22
AI Technical Summary
模型驱动方法基于混合整数规划、随机鲁棒优化等数学工具,能够精细刻画潮流方程、无功补偿及资源调节等物理约束,但面对大规模分布式资源接入时,海量0-1二元变量与强耦合约束使模型产生“维数灾难”,求解效率急剧下降,难以满足在线实时调度要求;同时,分布式资源在接入时序、可调容量上的强随机性也增加了精准机理建模难度
本发明的虚拟电厂大规模分布式资源智能调控方法,利用分布式DQN离线训练得到的动作价值函数,对传统调度模型进行降维重构,剔除海量二元0-1变量与耦合约束,大幅缩减决策变量与约束规模,显著提升求解效率,有效克服维数灾难,满足大规模资源在线实时调度需求。
Smart Images

Figure CN122801408A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of power system operation optimization technology, and in particular to a method for intelligent control of large-scale distributed resources in a virtual power plant. Background Technology
[0002] With the deepening of the "dual-carbon" strategy, the installed capacity of distributed energy has surged. A large number of distributed resources, mainly electric vehicles and energy storage, have been connected to the grid in a disorderly manner, leading to increasingly prominent safety issues such as widening peak-valley load differences in the distribution network, voltage exceeding limits at nodes, and power overload in branch circuits. Virtual power plants, as an effective means of aggregating distributed resources to participate in grid operation, have control methods that directly affect the safety and economic efficiency of distribution network operation.
[0003] Currently, distributed resource regulation methods in virtual power plants are mainly divided into two categories: model-driven and data-driven. Model-driven methods, based on mathematical tools such as mixed-integer programming and stochastic robust optimization, can accurately characterize physical constraints such as power flow equations, reactive power compensation, and resource regulation. However, when faced with large-scale distributed resource access, the massive number of 0-1 binary variables and strong coupling constraints cause the model to suffer from the "curse of dimensionality," resulting in a sharp decline in solution efficiency and making it difficult to meet the requirements of online real-time scheduling. At the same time, the strong randomness of distributed resources in terms of access timing and adjustable capacity also increases the difficulty of accurate mechanism modeling. Data-driven methods, represented by deep reinforcement learning such as Deep Q-Network (DQN), can adapt to resource uncertainty without precise physical modeling. However, they are prone to the curse of dimensionality in large-scale scenarios, with low agent sampling efficiency, slow convergence, and difficulty in rigidly embedding distribution network safety constraints such as power flow and voltage into the optimization process. The decision results are prone to exceeding the system's safe operation boundary. In recent years, although model-data hybrid driven methods have been proposed, existing solutions usually reduce decision variables by means of resource aggregation, which cannot achieve fine-grained control of each distributed resource. They still fail to fundamentally solve the dimensionality curse problem of large-scale distributed resource control, thus restricting the effectiveness of practical engineering applications.
[0004] Therefore, there is an urgent need for a virtual power plant intelligent control method that can efficiently and precisely regulate large-scale distributed resources while ensuring the security constraints of the distribution network. Summary of the Invention
[0005] The purpose of this invention is to provide a method for intelligent regulation of large-scale distributed resources in a virtual power plant, so as to efficiently and precisely regulate large-scale distributed resources.
[0006] To address the aforementioned technical problems, this invention provides a method for intelligent regulation and control of large-scale distributed resources in a virtual power plant.
[0007] The present invention provides a method for large-scale distributed intelligent control of virtual power plants, comprising: A traditional scheduling model for distributed resource regulation of a virtual power plant is constructed. The traditional scheduling model includes power flow constraints and distributed resource regulation constraints. A deep reinforcement learning offline training framework based on a distributed deep Q-network is constructed, and the distributed resource adjustment process is modeled as a Markov decision process using the deep reinforcement learning offline training framework. The comprehensive action value function of the distributed resources is obtained through offline training. Based on the action value function, and combined with the second-order cone optimization technique, the traditional scheduling model is reconstructed by reducing its dimension, eliminating binary 0-1 variables and coupling constraints, and obtaining a virtual power plant mixed integer second-order cone programming model. Solving the mixed-integer second-order cone programming model yields an intelligent control scheme for distributed resources.
[0008] Furthermore, the distributed deep Q-network adopts an actor-learner architecture; The actor-learner architecture is configured with multiple actor networks. Each actor network collects experience data in parallel in an independent environment and temporarily stores it in its own local experience replay buffer. When the local experience replay buffer meets the conditions, it samples from the local buffer and calculates the initial priority of each experience data and stores it in the shared experience replay buffer. The learner network samples experiences from the shared experience replay buffer according to priority, updates the priority of experience samples based on the magnitude of temporal difference error, and uses importance sampling weights to correct biases in the network parameter update process; each actor network periodically synchronizes the parameters of the learner network.
[0009] The learner network samples experiences from the shared experience replay buffer according to priority, updates network parameters, and synchronously adjusts experience priorities; each actor network periodically synchronizes the parameters of the learner network.
[0010] Furthermore, when the learner network samples according to priority, the intensity of priority sampling is adjusted by parameters; parameters are introduced into the importance sampling weight to adjust the degree of priority experience replay bias correction; at the same time, old samples are periodically removed from the shared experience replay buffer to maintain the buffer's effectiveness.
[0011] Furthermore, the traditional scheduling model takes minimizing the operating cost of the virtual power plant as its objective function.
[0012] Furthermore, the process of modeling the distributed resource adjustment process as a Markov decision process includes: Construct a state space, which includes binary connection variables indicating whether the distributed resource is in an adjustment state, state of charge, access time, exit time, and time-of-use electricity price for the current time period; Construct an action space by dividing the continuous adjustment power range of distributed resources according to a set discretization step size to obtain a finite set of discrete actions; Design a reward function that includes a limit-crossing penalty term to ensure that the distributed resource does not exceed the limit during operation, and a state-of-charge deviation penalty term to ensure that the state of charge of the distributed resource approaches the target value when it exits the virtual power plant.
[0013] Furthermore, the formula for the reward function is: ; in, To reward scaling factor, Indicates efficiency. This represents the charge state of the k-th distributed resource at time t. and Let represent the maximum and minimum values of the state of charge of the k-th distributed resource at time t, respectively; This indicates the expected exit time for the k-th distributed resource.
[0014] Furthermore, the comprehensive action value function for obtaining distributed resources includes: When the k-th distributed resource is allocated to the adjustment area, adjustment is... At that time, its action value Equal to state The maximum motion value of all possible actions except zero power, i.e. ; When distributed resources are in the standby area, i.e. At that time, its action value Equal to adjusting power The value of action at that time, namely: ; Based on the above action values, the comprehensive action value function of the k-th distributed resource at time t is... for: ; in, This represents the state vector of the k-th distributed resource at time t; The binary decision variable representing whether the k-th distributed resource is connected to the adjustment region at time t; This indicates the adjustment power of distributed resources.
[0015] Furthermore, when reconstructing the traditional scheduling model in terms of dimensionality reduction, a second-order cone programming method is used for relaxation, and auxiliary variables are introduced. , The relaxed power flow constraints include: ; ; ; ; ; ; ; ; ; in, , and Let represent the active power, reactive power, and branch current flowing from parent node i to child node j in branch ij at time t, respectively. and Both represent the initial injection power of branch ij at time t. This represents the active power flowing from parent node j to child node c in branch jc at time t. The reactive power flowing from parent node j to child node c in branch jc at time t; Represents the set of child nodes of node j; and Let represent the resistance and reactance of branch ij, respectively; and These represent the net active load and net reactive load of child node j, respectively. and These represent the node voltages of parent node i and child node j, respectively. and These represent the lower and upper limits of the squared magnitude of the node voltage of parent node i, respectively; This represents the upper limit of the square of the current amplitude in branch ij; This indicates the magnitude of reactive power compensation at node j; and This indicates the lower and upper limits of the reactive power that the static reactive power compensation device can output; This represents the set of all nodes in a distribution network that are equipped with static var compensators. It is a universal quantifier, meaning "for any" or "for all"; This means that it is effective for all nodes j in the distribution network that are equipped with static var compensators.
[0016] Furthermore, the objective function of the virtual power plant mixed-integer second-order cone programming model is: ; Where Ω represents the set of all nodes in the distribution network, and T represents the total time set of the entire scheduling cycle; The binary decision variable represents whether the k-th distributed resource of the m-th virtual power plant is connected to the regulation area at time t.
[0017] Furthermore, the distributed resource constraints of the virtual power plant mixed-integer second-order cone programming model are as follows: ; ; in, This indicates the number of regulating positions within the virtual power plant; and These represent the maximum and minimum injected power allowed at the common coupling point, respectively; K represents the set of distributed resources within the virtual power plant. This means that the condition must be met at any time t within the scheduling period T.
[0018] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention provides a large-scale distributed intelligent control method for virtual power plants. It utilizes the action value function obtained by offline training of distributed DQN to reconstruct the traditional scheduling model by reducing the dimensionality. This eliminates a large number of binary 0-1 variables and coupling constraints, significantly reduces the scale of decision variables and constraints, significantly improves the solution efficiency, effectively overcomes the curse of dimensionality, and meets the needs of large-scale online real-time scheduling of resources.
[0019] Furthermore, this invention incorporates second-order cone optimization techniques during the reconstruction process to perform convex relaxation on non-convex power flow constraints, while retaining safety constraints such as distribution network power flow constraints and distributed resource regulation constraints, ensuring that the scheduling scheme is solved within the safe operation boundary of the power grid. Offline training utilizes real-world data to adapt to the stochasticity of distributed resource access timing and adjustable capacity. During online solving, only the current state needs to be input and the MISOCP (Mixed Integer Second Order Cone Programming) model needs to be solved, reducing the online computational burden while balancing scheduling real-time performance and power grid operational safety. Attached Figure Description
[0020] Figure 1This is a flowchart illustrating an embodiment of the virtual power plant large-scale distributed resource intelligent control method of the present invention; Figure 2 This is a schematic diagram of a virtual power plant resource scheduling scenario, representing an embodiment of the virtual power plant large-scale distributed resource intelligent control method of the present invention. Figure 3 This is a schematic diagram of the topology of a radial distribution network in one embodiment of the virtual power plant large-scale distributed resource intelligent control method of the present invention; Figure 4 This is a schematic diagram of the topology of the IEEE 33-node distribution network test system used in an embodiment of the virtual power plant large-scale distributed resource intelligent control method of the present invention. Figure 5 This is a comparison of the reward convergence curves of distributed DQN and standard DQN during the offline training phase in one embodiment of the virtual power plant large-scale distributed resource intelligent control method of the present invention. Figure 6 This is a schematic diagram of the change in the state of charge of distributed resources, represented by electric vehicles, over time within a virtual power plant, according to one embodiment of the intelligent regulation method for large-scale distributed resources in a virtual power plant of the present invention. Figure 7 This is a comparison diagram of the voltage distribution of key nodes in the distribution network at different time periods in one embodiment of the virtual power plant large-scale distributed resource intelligent control method of the present invention. Detailed Implementation
[0021] The present invention's method for large-scale distributed intelligent control of virtual power plants will now be described with reference to schematic diagrams, which illustrate preferred embodiments of the invention. It should be understood that those skilled in the art can modify the invention described herein while still achieving its advantageous effects. Therefore, the following description should be understood as being of general knowledge to those skilled in the art and is not intended to limit the invention.
[0022] The serial numbers assigned to components in this document, such as "first," "second," etc., are merely used to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this application, unless otherwise specified, include both direct and indirect connections (linkages). In the description of this invention, it should be understood that the terms "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention.
[0023] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0024] In this application, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a direct connection or an indirect connection through an intermediate medium. Furthermore, the term "electrical connection" can be a direct electrical connection or an indirect electrical connection through an intermediate medium.
[0025] The invention is described more specifically by way of example in the following paragraphs with reference to the accompanying drawings. The advantages and features of the invention will become clearer from the following description. It should be noted that the drawings are in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the invention.
[0026] The following is in conjunction with the instruction manual appendix. Figure 1 To be continued Figure 7 This paper introduces the intelligent control method for large-scale distributed resources in virtual power plants according to the present invention.
[0027] In some of these embodiments, such as Figure 1 As shown, the large-scale distributed resource intelligent control method for virtual power plants includes: Step S100: Construct a traditional scheduling model for distributed resource regulation of a virtual power plant, wherein the traditional scheduling model includes power flow constraints and distributed resource regulation constraints of the distribution network; Step S200: Construct a deep reinforcement learning offline training framework based on a distributed deep Q-network, and use the deep reinforcement learning offline training framework to model the distributed resource adjustment process as a Markov decision process, and obtain the comprehensive action value function of the distributed resources through offline training. Step S300: Based on the action value function, combined with the second-order cone optimization technique, the traditional scheduling model is reconstructed by dimensionality reduction, eliminating binary 0-1 variables and coupling constraints, to obtain a virtual power plant mixed integer second-order cone programming model; Step S400: Solve the mixed-integer second-order cone programming model to obtain an intelligent control scheme for distributed resources.
[0028] The present invention provides a large-scale distributed intelligent control method for virtual power plants. It utilizes the action value function obtained by offline training of distributed DQN to reconstruct the traditional scheduling model by reducing the dimensionality. This eliminates a large number of binary 0-1 variables and coupling constraints, significantly reduces the scale of decision variables and constraints, significantly improves the solution efficiency, effectively overcomes the curse of dimensionality, and meets the needs of large-scale online real-time scheduling of resources.
[0029] Furthermore, this invention incorporates second-order cone optimization techniques during the reconstruction process to perform convex relaxation on non-convex power flow constraints, while retaining safety constraints such as distribution network power flow constraints and distributed resource regulation constraints, ensuring that the scheduling scheme is solved within the safe operation boundary of the power grid. Offline training utilizes real-world data to adapt to the stochasticity of distributed resource access timing and adjustable capacity. During online solving, only the current state needs to be input and the MISOCP (Mixed Integer Second Order Cone Programming) model needs to be solved, reducing the online computational burden while balancing scheduling real-time performance and power grid operational safety.
[0030] In step S100, specifically, as follows Figure 2 As shown, when there is surplus regulation capacity within the virtual power plant, distributed resources such as electric vehicles are allowed to connect. After the distributed resources connect to the virtual power plant, the scheduling model determines whether the resource should be allocated to the regulation area to participate in regulation. Figure 3 The above describes a traditional scheduling model constructed based on the power flow constraints of the distribution network and with the objective function of minimizing the operating cost of the virtual power plant. The objective function of the traditional scheduling model is: ; The power flow constraints of the distribution network include: ; ; ; ; ; ; ; The distributed resource adjustment constraints include: ; ; ; ; ; ; ; ; Where Ω represents the set of all nodes in the distribution network; K represents the set of distributed resources within the virtual power plant; and T represents the total time set of the entire scheduling cycle. It is a universal quantifier, meaning "for any" or "for all"; This means that the condition must be met at any time t within the scheduling period T; This means that for any distributed resource k in set K, the following condition must be met; For another set of binary variables, and This indicates that the distributed resource is either not present in the virtual power plant or is present in the virtual power plant, respectively. Indicates the penalty factor; The binary decision variable represents whether the k-th distributed resource of the virtual power plant is connected to the regulation area at time t. and These represent the k-th distributed resource being unconnected and connected to the adjustment unit for adjustment at time t, respectively. This represents the time-of-use electricity price at time t. and These represent the initial charge state and the target charge state when accessing distributed resources, respectively. , and Let represent the active power, reactive power, and branch current flowing from parent node i to child node j in branch ij at time t, respectively. and Both represent the initial injection power of branch ij at time t. This represents the active power flowing from parent node j to child node c in branch jc at time t. The reactive power flowing from parent node j to child node c in branch jc at time t; Represents the set of child nodes of node j; This indicates the number of regulating positions within the virtual power plant; and These are the upper and lower limits of the active power of a distribution network branch. and These are the upper and lower limits of reactive power; and These represent the net active and reactive loads of child node j, respectively. and These represent the access time and exit time of the distributed resource, respectively. Indicates the access time of the k-th distributed resource. Exit time any time between All must be satisfied; and Let represent the resistance and reactance of branch ij, respectively; and These represent the net active load and net reactive load of child node j, respectively. and These represent the node voltages of parent node i and child node j, respectively. This indicates the magnitude of reactive power compensation at node j; and This indicates the lower and upper limits of the reactive power that the static reactive power compensation device can output; This represents the set of all nodes in a distribution network that are equipped with static var compensators. This means that it is effective for all nodes j in the distribution network that are equipped with static var compensators; and Let represent the maximum and minimum adjustment power of the k-th distributed resource at time t, respectively; This represents the adjustment power of the k-th distributed resource at time t; and These represent the maximum and minimum injection power allowed at the common coupling point, respectively.
[0031] In step S200, in some embodiments, the deep reinforcement learning framework based on distributed deep Q-networks (DQN) specifically adopts an actor-learner distributed deep reinforcement learning architecture. The network training process under this architecture is specifically divided into two parts: actor network training and learner network training.
[0032] Regarding actor network training, the actor-learner architecture is configured with multiple actor networks. Each actor network executes an action exploration strategy in parallel based on its own independent virtual power plant environment instance, generates experience data online in real time, and temporarily stores this experience data in its corresponding local experience replay buffer. When the local experience replay buffer meets a preset condition (e.g., when it is full), the actor network extracts samples from the local buffer, calculates the initial priority of each experience data, and then stores the priority-bearing experience data into a globally shared experience replay buffer.
[0033] Regarding learner network training, the learner network samples experiences from the shared experience replay buffer according to priority. After acquiring experience samples, the learner network calculates the corresponding temporal difference error and dynamically updates the priority of the corresponding experience sample in the shared buffer based on the magnitude of the temporal difference error, so that subsequent training is biased towards experiences with greater learning value. Simultaneously, to reduce the bias in network parameter updates caused by priority experience replay, the learner network introduces importance sampling weights for bias correction during parameter gradient updates. After the network update is complete, each actor network periodically synchronizes the latest parameters of the learner network to ensure consistency between policy evaluation and policy execution, and to maintain training stability.
[0034] Furthermore, in the process of priority sampling in the learner network described above, a first adjustment parameter (corresponding to the hyperparameter below) is introduced. The system controls the intensity of priority sampling. When this parameter approaches 0, sampling tends to be uniformly distributed; when it approaches 1, sampling is performed entirely according to priority. A second adjustment parameter (corresponding to the hyperparameter β below) is introduced into the importance sampling weight to adjust the degree of priority experience replay bias correction. Furthermore, the system also has an experience purging mechanism that periodically removes older samples from the shared experience replay buffer to maintain the timeliness and diversity of experience samples in the shared buffer, ensuring the effectiveness of model training.
[0035] Specifically, in some embodiments, the network training process for actors under the distributed DQN framework includes the following steps: Step S211: Configure multiple actor networks. Each actor network executes the exploration strategy in parallel based on its own independent virtual power plant environment instance, generates experience data in real time, and temporarily stores the experience data in its corresponding local experience replay buffer. Step S212: When the local experience playback buffer meets a preset condition (e.g., when it is full), experience data is sampled from the local buffer, and the initial priority of each piece of experience data is calculated. Then, the experience data with the priority is stored in the shared experience playback buffer. The priority calculation formula for the experience data is: ; in, Indicates the priority of the i-th empirical sample; parameter To adjust the priority sampling intensity in empirical playback, when At that time, the empirical samples were taken according to a uniform distribution, and the sampling was carried out when... In this case, sampling is performed entirely based on priority.
[0036] To reduce bias during parameter updates, this invention introduces importance sampling weights to correct the network parameter update process. The formula for calculating the importance sampling weights is as follows: ; in, The importance sampling weight corresponding to the i-th sample is represented by N; N represents the total number of samples in the shared experience replay buffer; parameters This is a hyperparameter used to adjust the degree of bias correction for priority experience playback.
[0037] Step S213: To ensure consistency between policy evaluation and policy execution and maintain the stability of the entire training process, each actor network will periodically update its own network parameters based on the current parameters of the learner network.
[0038] In some embodiments, the network training process for learners under the distributed DQN framework includes the following steps: Step S221: The learner network extracts small batches of experience samples from the shared experience replay buffer established above, according to the updated priority sampling strategy. Step S222: The learner network calculates the loss based on the sampled mini-batch experience samples using the following time-difference objective function, and updates the network parameters of the learner network:
[0039] Among them, the target value It represents multi-step rewards and integrates multi-step rewards and the double-Q bootstrapping mechanism. Its specific calculation formula is as follows:
[0040] In the formula, This represents the target Q-value calculated by combining multi-step reward and double Q-guided mechanisms, used to calculate the loss function. ; (i=1,2,...,n) represents the instantaneous reward generated at the i-th time step in the future, starting from the current time t; t represents the current action time step; n represents the horizon step size of the multi-step reward, which is approximated by the n-step real cumulative reward; This represents the discount factor, used to balance the weight between immediate rewards and long-term cumulative rewards; This represents the action value function fitted by the deep reinforcement learning network. This represents the system state at time t; This indicates the action chosen at time t; This represents the weight parameters of the current main network; This represents the weight parameters of the target network, used to stabilize the calculation of the target value and avoid oscillations during the training process.
[0041] Step S223: To ensure the effectiveness of the shared experience replay buffer and prevent outdated or invalid samples from interfering with training, the learner network periodically removes older samples from the global experience replay buffer. Simultaneously, based on the temporal difference error calculated in the previous steps, the priority of the experience samples is recalculated and adjusted. Therefore, samples with larger temporal difference errors are assigned higher priority in the shared experience replay buffer, resulting in a higher probability of being sampled in subsequent network parameter updates and a greater contribution weight to the network update. At the same time, periodically removing outdated samples prevents policy bias and ensures the timeliness of the training samples.
[0042] In step S200, in some embodiments, modeling the distributed resource regulation process as a Markov decision process includes: Construct a state space, which includes binary connection variables indicating whether the distributed resource is in an adjustment state, state of charge, access time, exit time, and time-of-use electricity price for the current time period; Construct an action space by dividing the continuous adjustment power range of distributed resources according to a set discretization step size to obtain a finite set of discrete actions; Design a reward function that includes a limit-crossing penalty term to ensure that the distributed resource does not exceed the limit during operation, and a state-of-charge deviation penalty term to ensure that the state of charge of the distributed resource approaches the target value when it exits the virtual power plant.
[0043] Specifically, preferably, the present invention models the problem of distributed resources participating in the regulation of virtual power plant regulation areas as a Markov decision process, wherein the Markov decision process consists of a quintuple. express, Among them, and Representing the state space and action space respectively. The state transition probability distribution is defined. This represents the real-valued reward function, while As a discount factor, the closer it is to 1, the more it emphasizes long-term cumulative rewards.
[0044] Specifically, the state space This includes binary connection variables indicating whether the distributed resource is in a regulated state, state of charge, access time, exit time, and the current time-of-use electricity price. This information collectively constitutes the environmental state input upon which the agent relies during its decision-making process.
[0045] The state space The formula is: ; in, A binary join variable indicating whether a distributed resource is in a regulated state; It indicates the state of charge when distributed resources are accessed, reflecting the current remaining adjustable power. Indicates the access time of distributed resources; This represents the time-of-use electricity price at time t. Indicates the exit time of the distributed resource.
[0046] Regarding the aforementioned action space Since distributed DQN is only applicable to discrete action spaces, while actual adjustment power is usually a continuous variable, the range of continuous adjustment power of distributed resources needs to be divided according to a set discretization step size to obtain a finite set of discrete actions. This finite set of discrete actions constitutes the action space. .
[0047] Specifically, let the power interval of the k-th distributed resource at time t be... And introduce discretization step size Then its discretized action set can be represented as: .
[0048] It should be noted that regarding the state transition probability distribution The state transition probability defines the probability distribution of a distributed resource transitioning from its state at time t to its state at time t+1. Since the initial state of charge of a distributed resource when connected to a virtual power plant is random, this probability distribution cannot be analyzed analytically through mechanistic modeling in a real environment. It belongs to an unknown distribution. Therefore, this invention does not... Instead of performing direct computation, it utilizes a model-free distributed DQN framework. Through experience samples generated by the interaction between the actor network and the environment, it approximates the action value function with the help of a deep neural network, thereby achieving efficient decision-making without needing to know the transition probabilities.
[0049] Regarding the reward function The goal of a distributed resource regulation system is to minimize regulation costs while meeting regulation requirements and adhering to operational constraints. To achieve this goal, the reward function for the k-th distributed resource at time t... The formula is: ; in, To reward scaling factor, Indicates efficiency. This represents the charge state of the k-th distributed resource at time t. and Let represent the maximum and minimum values of the state of charge of the k-th distributed resource at time t, respectively; This represents the expected exit time of the k-th distributed resource.
[0050] At time t, based on the probability that the agent chooses action a in state s... The action value function in this state is expressed as: ; in, Representation strategy The expected mathematical return is given below, where l is a decision step size; This represents the state vector of the k-th distributed resource at time t; This represents the adjustment power of the k-th distributed resource at time t.
[0051] Based on the above action value function, the action values of the adjustment bits and standby bits of the distributed resources obtained by offline training of the distributed DQN algorithm are as follows: When the k-th distributed resource is allocated to the adjustment area, adjustment is... At that time, its action value Equal to state The maximum motion value of all possible actions except zero power, i.e. ; When distributed resources are in the standby area, i.e. At that time, its action value Equal to adjusting power The value of action at that time, namely: ; Based on the above action values, the comprehensive action value function of the k-th distributed resource at time t is... for: ; in, This represents the state vector of the k-th distributed resource at time t; The binary decision variable representing whether the k-th distributed resource is connected to the adjustment region at time t; This represents the adjustment power of the k-th distributed resource at time t.
[0052] In step S300, since the power flow constraints are a non-convex nonlinear model, the model is difficult to solve. Therefore, when reconstructing the traditional scheduling model in a reduced dimension, the Second-Order Cone Programming (SOCP) method is used for relaxation. Auxiliary variables are introduced. The relaxed power flow constraints include: ; ; ; ; ; ; ; ; ; in, , and Let represent the active power, reactive power, and branch current flowing from parent node i to child node j in branch ij at time t, respectively. and Both represent the initial injection power of branch ij at time t. This represents the active power flowing from parent node j to child node c in branch jc at time t. The reactive power flowing from parent node j to child node c in branch jc at time t; Represents the set of child nodes of node j; and Let represent the resistance and reactance of branch ij, respectively; and These represent the net active load and net reactive load of child node j, respectively. and These represent the node voltages of parent node i and child node j, respectively. and These represent the lower and upper limits of the squared magnitude of the node voltage of parent node i, respectively; This represents the upper limit of the square of the current amplitude in branch ij; This indicates the magnitude of reactive power compensation at node j; and This indicates the lower and upper limits of the reactive power that the static reactive power compensation device can output; This represents the set of all nodes in a distribution network that are equipped with static var compensators.
[0053] Based on the comprehensive action value function obtained in step S200 The objective function of the virtual power plant mixed-integer second-order cone programming model is: ; Where Ω represents the set of all nodes in the distribution network, and T represents the total time set of the entire scheduling cycle; The binary decision variable represents whether the k-th distributed resource of the m-th virtual power plant is connected to the regulation area at time t.
[0054] Meanwhile, the distributed resource constraints of the virtual power plant mixed-integer second-order cone programming model are as follows: ; ; in, This indicates the number of regulating positions within the virtual power plant; and These represent the maximum and minimum injected power allowed at the common coupling point, respectively; K represents the set of distributed resources within the virtual power plant.
[0055] Finally, the complete virtual power plant mixed-integer second-order cone programming model includes: The objective function is: ; The constraints are: ; ; ; ; ; ; ; ; ; ; .
[0056] Regarding step S400, specifically, after the virtual power plant mixed integer second-order cone programming model is reconstructed in step S300, since the model has been transformed into a convex optimization problem through second-order cone relaxation, it can be directly solved numerically.
[0057] The solution process will directly output the optimal binary state decision variables for each time period in the model. (That is, determining which distributed resources enter the regulation zone) and simultaneously calculating the corresponding optimal regulation power. Based on the solution results, the system generates the final executable control commands for each distributed resource, thereby realizing large-scale online intelligent collaborative control of virtual power plant resources under the premise of meeting the safety constraints of the distribution network.
[0058] To verify the effectiveness of this invention, a simulation of distributed resource access and adjustment was conducted based on actual data. The distributed resource access time follows a normal distribution. The exit time follows a normal distribution. The initial state of charge upon connection follows a normal distribution within the range of [0.2, 0.6]. .
[0059] This invention employs an IEEE 33-node system, the topology of which is as follows: Figure 4 As shown, three distributed resource clusters are connected. The distributed resource, represented by electric vehicles, has a capacity of 24 kWh and a maximum regulating power of 6 kW. The desired state of charge target range is [0.9, 1]. The time-of-use pricing, training hyperparameters, and distribution network parameters in the simulation are shown in Tables 1, 2, and 3, respectively. The simulation is based on Python 3.9 and the PyTorch deep learning framework (an open-source deep learning framework). The distributed resource regulation environment is built on the Gym platform (an open-source reinforcement learning toolkit developed by OpenAI).
[0060] Table 1 Time-of-use electricity prices
[0061] Table 2 Hyperparameters of the Distributed DQN Algorithm
[0062] Table 3 Parameter Configuration in Distribution Network System
[0063] Figure 5 This paper presents a comparison of the reward curves of the distributed deep Q-network (Ape-X DQN) used in this invention with those of traditional DQN during the offline training phase. Figure 5 As can be seen, distributed DQN achieves faster cumulative reward improvement in the early stages of training, tends to converge and stabilize within approximately 30,000 training epochs, and the cumulative reward value obtained upon final convergence is significantly higher than that of traditional DQN. This verifies the effectiveness and advancement of the parallel actor-learner architecture and priority experience replay mechanism proposed in this invention in improving sample utilization efficiency and training convergence speed.
[0064] Figure 6 This paper presents a heatmap showing the temporal distribution of the State of Charge (SOC) of distributed resources, represented by electric vehicles, within a virtual power plant over a one-day simulation period. Combined with the time-of-use pricing data in Table 1, it can be seen that the SOC of each distributed resource decreases during peak hours (reflecting discharge behavior) and increases during normal and off-peak hours (reflecting charging behavior). This charging and discharging characteristic indicates that the proposed method can effectively guide distributed resources to respond to pricing signals, achieving "charging during off-peak hours and discharging during peak hours," thus verifying the positive benefits of this invention in promoting the orderly integration of distributed resources and assisting the power grid in peak shaving and valley filling.
[0065] exist Figure 7 The paper uses two typical time periods, 8:00 AM and 1:00 PM, as examples to demonstrate the voltage distribution at various nodes of the distribution network under the control of the method proposed in this invention. Figure 7 As can be seen, after applying the intelligent control method for virtual power plants proposed in this invention, the orderly adjustment of various distributed resources within the virtual power plant can maintain the voltage of each node in the distribution network within a safe and stable range [0.95, 1.0]. This strongly demonstrates that while achieving the economic goals of virtual power plants, this invention can effectively ensure the voltage safety and operational stability of the distribution network, and realize peak shaving and valley filling.
[0066] Table 4 Comparison of Charging Station Revenue
[0067] Table 5 Comparison of Solution Time for Scheduling Models
[0068] Furthermore, since the above-mentioned reward function model and action value function have completed the comprehensive integration of distributed resource adjustment actions, state of charge, adjustment costs and distribution network operation safety constraints during the definition process, this embodiment uses the total value of the simulation-obtained reward function (i.e., the cumulative action value Q value) as the core indicator for measuring the economic efficiency of virtual power plant operation, and the model solution time as the core indicator for measuring the real-time performance of online scheduling.
[0069] Tables 4 and 5 present the comparison results of the total daily reward function of distributed resources and the model solution time for different scheduling models under the condition of considering the operation constraints of the distribution network.
[0070] As shown in Table 4, in the comparison of the total daily reward function value, the corresponding value of the strategy proposed in this invention is -148.02, the traditional MISOCP model is -110.35, and the queuing theory model is -171.91. This indicates that the strategy proposed in this invention improves the total daily reward value by 13.9% compared to the traditional queuing theory model. This fully demonstrates the significant economic advantages of this invention in resource scheduling strategies.
[0071] As shown in Table 5, in the comparison of model solution time, the single scheduling solution time of the strategy proposed in this invention is 0.21s, while the traditional MISOCP model takes as much as 17.72s, and the queuing theory model takes 0.25s. The strategy proposed in this invention reduces the solution time by 98.8% compared to the traditional MISOCP model. This indicates that this invention, through distributed DQN offline dimensionality reduction and MISOCP reconstruction, greatly solves the curse of dimensionality problem of traditional models in large-scale distributed resource scenarios, and significantly improves the real-time response capability of online scheduling.
[0072] In summary, the strategy proposed in this invention can balance operational economy and real-time solution in complex scenarios involving large-scale distributed resource access to the distribution network, while effectively ensuring the safe operation of the power grid.
[0073] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for large-scale distributed intelligent control of virtual power plants, characterized in that, include: A traditional scheduling model for distributed resource regulation of a virtual power plant is constructed. The traditional scheduling model includes power flow constraints and distributed resource regulation constraints. A deep reinforcement learning offline training framework based on a distributed deep Q-network is constructed, and the distributed resource adjustment process is modeled as a Markov decision process using the deep reinforcement learning offline training framework. The comprehensive action value function of the distributed resources is obtained through offline training. Based on the action value function, and combined with the second-order cone optimization technique, the traditional scheduling model is reconstructed by reducing its dimension, eliminating binary 0-1 variables and coupling constraints, and obtaining a virtual power plant mixed integer second-order cone programming model. Solving the mixed-integer second-order cone programming model yields an intelligent control scheme for distributed resources.
2. The method for large-scale distributed intelligent control of virtual power plants according to claim 1, characterized in that, The distributed deep Q-network adopts an actor-learner architecture; The actor-learner architecture is configured with multiple actor networks. Each actor network collects experience data in parallel in an independent environment and temporarily stores it in its own local experience replay buffer. When the local experience replay buffer meets the conditions, it samples from the local buffer and calculates the initial priority of each experience data and stores it in the shared experience replay buffer. The learner network samples experiences from the shared experience replay buffer according to priority, updates the priority of experience samples based on the magnitude of temporal difference error, and uses importance sampling weights to correct biases in the network parameter update process; each actor network periodically synchronizes the parameters of the learner network. The learner network samples experiences from the shared experience replay buffer according to priority, updates network parameters, and synchronously adjusts experience priorities; each actor network periodically synchronizes the parameters of the learner network.
3. The method for large-scale distributed intelligent control of virtual power plants according to claim 2, characterized in that, When the learner network samples according to priority, a first adjustment parameter is used to adjust the intensity of priority sampling; a second adjustment parameter is introduced into the importance sampling weight to adjust the degree of priority experience replay bias correction; at the same time, old samples are periodically removed from the shared experience replay buffer to maintain the buffer's effectiveness.
4. The method for large-scale distributed intelligent regulation and control of virtual power plants according to claim 1, characterized in that, The traditional scheduling model takes minimizing the operating cost of the virtual power plant as its objective function.
5. The method for large-scale distributed intelligent control of virtual power plants according to claim 1, characterized in that, The process of modeling the distributed resource adjustment process as a Markov decision process includes: Construct a state space, which includes binary connection variables indicating whether the distributed resource is in an adjustment state, state of charge, access time, exit time, and time-of-use electricity price for the current time period; Construct an action space by dividing the continuous adjustment power range of distributed resources according to a set discretization step size to obtain a finite set of discrete actions; Design a reward function that includes a limit-crossing penalty term to ensure that the distributed resource does not exceed the limit during operation, and a state-of-charge deviation penalty term to ensure that the state of charge of the distributed resource approaches the target value when it exits the virtual power plant.
6. The method for large-scale distributed intelligent control of virtual power plants according to claim 5, characterized in that, The formula for the reward function is: ; in, To reward scaling factor, Indicates efficiency. This represents the charge state of the k-th distributed resource at time t. and Let represent the maximum and minimum values of the state of charge of the k-th distributed resource at time t, respectively; This indicates the expected exit time for the k-th distributed resource.
7. The method for large-scale distributed intelligent control of virtual power plants according to claim 1, characterized in that, The comprehensive action value function for obtaining distributed resources includes: When the k-th distributed resource is allocated to the adjustment area, adjustment is... At that time, its action value Equal to state The maximum motion value of all possible actions except zero power, i.e. ; When distributed resources are in the standby area, i.e. At that time, its action value Equal to adjusting power The value of action at that time, namely: ; Based on the above action values, the comprehensive action value function of the k-th distributed resource at time t is... for: ; in, This represents the state vector of the k-th distributed resource at time t; The binary decision variable representing whether the k-th distributed resource is connected to the adjustment region at time t; This indicates the adjustment power of distributed resources.
8. The method for large-scale distributed intelligent control of virtual power plants according to claim 7, characterized in that, When reconstructing the traditional scheduling model in terms of dimensionality reduction, a second-order cone programming method is used for relaxation, and auxiliary variables are introduced. , The relaxed power flow constraints include: ; ; ; ; ; ; ; ; ; in, , and Let represent the active power, reactive power, and branch current flowing from parent node i to child node j in branch ij at time t, respectively. and Both represent the initial injection power of branch ij at time t. This represents the active power flowing from parent node j to child node c in branch jc at time t. The reactive power flowing from parent node j to child node c in branch jc at time t; Represents the set of child nodes of node j; and Let represent the resistance and reactance of branch ij, respectively; and These represent the net active load and net reactive load of child node j, respectively. and These represent the node voltages of parent node i and child node j, respectively. and These represent the lower and upper limits of the squared magnitude of the node voltage of parent node i, respectively; This represents the upper limit of the square of the current amplitude in branch ij; This indicates the magnitude of reactive power compensation at node j; and This indicates the lower and upper limits of the reactive power that the static reactive power compensation device can output; This represents the set of all nodes in a distribution network that are equipped with static var compensators. It is a universal quantifier, meaning "for any" or "for all"; This means that it is effective for all nodes j in the distribution network that are equipped with static var compensators.
9. The method for large-scale distributed intelligent regulation of virtual power plants according to claim 8, characterized in that, The objective function of the virtual power plant mixed-integer second-order cone programming model is: ; Where Ω represents the set of all nodes in the distribution network, and T represents the total time set of the entire scheduling cycle; The binary decision variable represents whether the k-th distributed resource of the m-th virtual power plant is connected to the regulation area at time t.
10. The method for large-scale distributed intelligent regulation and control of virtual power plants according to claim 9, characterized in that, The distributed resource constraints of the virtual power plant mixed-integer second-order cone programming model are as follows: ; ; in, This indicates the number of regulating positions within the virtual power plant; and These represent the maximum and minimum injected power allowed at the common coupling point, respectively; K represents the set of distributed resources within the virtual power plant. This means that the condition must be met at any time t within the scheduling period T.