Load Frequency Control Method, Device and Medium for Low-Voltage Distribution Network in Substation Area
By adopting the edge cloud collaborative data-driven load frequency control method in the low-voltage table distribution network, combining meta reinforcement learning and multi-agent depth deterministic strategy gradient algorithm, cloud-edge collaborative control is realized, solving the problems of frequency stability and balance of interests, and improving power supply quality and control robustness.
Patent Information
- Application Number
- CN202311606231.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-11-28
AI Technical Summary
The prior art is difficult to effectively control frequency stability in low-voltage power distribution networks, especially in the case of uncertainty in distributed power output and conflicts of interest among multiple operators.
The edge cloud collaborative data-driven load frequency control method (ECCDD-LFC) is adopted to build a load frequency control model, combine cloud agents and edge agents, and use meta reinforcement learning and multi-agent deep deterministic policy gradient (ECLSMA-DMDPG) algorithm to achieve cloud-edge collaborative control.
It improves the frequency stability and power supply quality in the distribution network in the Taiwan area, balances the interests between different units, and improves the robustness and decision-making efficiency of the control system.
Smart Images

Figure CN117650542B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and particularly relates to a load frequency control (ECCDD-LFC) method, device, and storage medium for a low-voltage distribution network in a low-voltage area. Background Art
[0002] At present, progress has been made in power grids designed to accommodate a large number of distributed energy sources connected. The distribution network technology in the low-voltage area can effectively integrate the advantages of distributed power sources while meeting the load growth demand, and has been widely studied to ensure the reliability of system power supply. The low-voltage area distribution network requires effective voltage and frequency control to ensure stable operation. The frequency control method for the low-voltage area distribution network is called load frequency control (LFC). With the increase in the penetration rate of new energy, the system inertia decreases, and power imbalance will cause rapid frequency changes, thus endangering the stability of the system. Although renewable energy is highly praised for its environmental friendliness and flexibility, its output power has great uncertainty and is greatly affected by external environmental factors. Therefore, the imbalance of its output power will affect the power grid frequency. To solve the adverse effects of distributed power sources and improve the power supply quality, it is of great significance to study the load frequency control technology of the system for the safe operation of the power system.
[0003] In addition, in an isolated low-voltage area distribution network, the frequency modulation reserve capacity and generation cost of each adjustable resource change with the operating conditions. Since multiple units belong to different operators, they each pursue their own interests to maximize certain parameters. In frequency regulation, each unit needs to handle the task sharing, cost sharing, and economy issues of frequency regulation caused by power interference; to balance the needs of individual adjustable resources and the low-voltage area distribution network, the orderliness of the frequency control of the low-voltage area distribution network is crucial. For the LFC problem in the low-voltage area distribution network, researchers generally design two main solutions, centralized LFC control and distributed LFC control.
[0004] Specifically, the current technology has the following problems:
[0005] Technical Solution of the Existing Technology One
[0006] Patent Application No.: 202211325051, Patent Name: Power System Transient Voltage Stability Emergency Load Shedding Control Method, System, and Medium
[0007] A transient voltage stability emergency load shedding control method, system and medium for a power system are provided. Fault condition samples are selected; according to the fault condition samples and the previous load shedding control measure, a first state, a first completion flag and a first reward are obtained; the intelligent agent selects the current load shedding control measure in the linear decision space according to the previous load shedding control measure and simulates to obtain a second state; the strategy network parameters of the intelligent agent are updated by using the first state, the first completion flag, the first reward, the current load shedding control measure and the second state, and the current load shedding control measure is used as the new previous load shedding control measure; the second step to the fourth step are repeatedly executed until an effective load shedding control measure is obtained or a preset number of iterations is reached; the first step to the fifth step are repeatedly executed until a preset number of training rounds is reached.
[0008] Disadvantages of the prior art I
[0009] This method belongs to the offline pre-decision - online matching mode. Based on power flow equations, heuristic algorithms, and sensitivity analysis of transient process control variables, offline pre-decision emergency load shedding control is specified, and deep reinforcement learning is applied to it. However, there are currently problems in applying deep reinforcement learning to emergency load shedding control, such as difficulties in setting up the Markov decision process, an overly large decision space, and insufficient knowledge integration. Moreover, it is difficult to train the intelligent agent of this method, and the decision-making efficiency is low. For practical engineering applications, it is difficult to design a reasonable decision space and integrate domain knowledge, and it is difficult to improve the training efficiency and decision-making quality of the intelligent agent.
[0010] Technical solution of the prior art II
[0011] Patent application number: 202110806077 Patent name: Frequency regulation collaborative control method and system for large-scale distributed photovoltaic power stations
[0012] A frequency regulation collaborative control method and system for large-scale distributed photovoltaic power stations are provided. The change information of the AC system frequency signal is converted into the change information of the DC bus voltage of the grid-connected inverter, so as to deploy decentralized communication-free frequency modulation control for individual photovoltaic modules. This control method is divided into two levels. On the grid-connected inverter side, the AC system frequency deviation is introduced into the DC bus voltage control loop to adjust the energy stored in the DC bus capacitor to achieve virtual inertia control; on the photovoltaic power generation side, decentralized f-P droop control based on DC optimizers is deployed, so as to adjust the active power output of each photovoltaic power generation unit according to the change in the output voltage of the DC optimizer generated by the frequency change.
[0013] Disadvantages of the prior art II
[0014] This method is based on a centralized photovoltaic power station and adopts global maximum power point tracking technology. However, in the case of frequent local shadow shading, the power generation efficiency of the photovoltaic power station is low, which reduces the spare photovoltaic capacity for frequency regulation and limits the overall frequency regulation ability of the system. Therefore, when low-frequency or high-frequency oscillations occur in the system, the system cannot effectively regulate the frequency by itself and cannot effectively keep the frequency within a stable range. At the same time, this method does not consider economy, and uses too many photovoltaics, which is very likely to cause waste of resources. Therefore, this method is not suitable for practical engineering applications.
[0015] Technical solution of the prior art three
[0016] Patent Application No.: 202211432485, Patent Name: Shared Energy Storage Scheduling Method and System Based on Energy Frequency Regulation and Load Demand
[0017] A shared energy storage scheduling method and system based on energy frequency regulation and load demand are provided. The method includes: establishing an objective function for the collaborative scheduling of the shared energy storage system to participate in energy frequency regulation and load demand; inputting relevant parameters of the grid side, user side, and shared energy storage system into the objective function; solving the objective function according to the objective function constraint conditions and the switching cost of the load importance degree, in combination with the mixed integer linear programming algorithm, to obtain a shared energy storage configuration plan; configuring the shared energy storage system according to the shared energy storage configuration plan, and controlling the shared energy storage system to participate in the energy collaborative scheduling of the grid side and user side according to the hierarchical control strategy of the shared energy storage system.
[0018] Disadvantages of the prior art three
[0019] This method does not explore the collaborative optimization of the shared energy storage system to participate in both the grid side and the user side at the same time, resulting in the overall system not fully utilizing the existing renewable energy resources. In addition, the collaborative optimization scheduling between the grid side and the user side will cause the charge and discharge state of the shared energy storage system to switch frequently, which is likely to cause the rated capacity of the shared energy storage system to drop rapidly, reduce the economy, and furthermore, there are many unpredictable safety problems and it is difficult to be put into practical engineering applications. Summary of the invention
[0020] A load frequency control method, device, and storage medium for a low-voltage distribution network in a low-voltage area proposed by the present invention can solve at least one of the technical problems in the background technology.
[0021] To achieve the above object, the present invention adopts the following technical solutions:
[0022] A load frequency control method for a low-voltage distribution network in a low-voltage area realizes cloud-edge collaborative control by using a pre-constructed load frequency control model. Among them, the construction steps of the load frequency control model are as follows:
[0023] S1. First, propose the modeling of the LFC for the distribution network in the substation area, and model the microcomputer, distributed power supply, wind power generation, and photovoltaic power generation in the system unit of the distribution network in the substation area;
[0024] S2. Introduce edge-cloud collaborative data-driven on the basis of traditional LFC, build the ECCDD-LFC framework, use the cloud agent to collect the frequency deviation data of the distribution network in the substation area and output the total adjustment command, and use the edge agents of each unit to output the participation variables of each unit;
[0025] S3. Considering the interaction between each agent and its environment, introduce meta-reinforcement learning into the ECCDD-LFC framework, and model the base learner and the meta-learner in the MDP modeling, where the meta-learner is responsible for the initialization of the algorithm to explore the noise variance, and the base learner is responsible for the training of specific tasks;
[0026] S4. In the training of the base learner, through the MADDPG algorithm, change the gradients of the local policies of all agents to optimize the global objective, and through the ECLSMA-DMDPG algorithm, use a large number of samples to update its own network parameters.
[0027] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.
[0028] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which when executed by the processor causes the processor to execute the steps of the above method.
[0029] As can be seen from the above technical solutions, the present invention adopts the edge-cloud collaborative data-driven load frequency control method (ECCDD-LFC), fully mobilizes the node computing resources and the global optimization ability of the cloud, sets each adjustment unit as an independent decision-making edge agent, and combines the edge-cloud collaborative framework with deep meta-reinforcement learning to enable the interaction between each unit and achieve cloud-edge collaborative control. To implement this method, the present invention proposes an Emergent Computing Large-Scale Multi-Agent Deep Meta-Deterministic Policy Gradient (ECLSMA-DMDPG). This strategy adopts a centralized training and decentralized execution strategy to achieve multi-agent collaboration, adopts meta-learning technology to improve the robustness of the algorithm, and adopts emergent computing to generate high-value samples to improve performance.
[0030] Specifically, the beneficial effects of the present invention are as follows:
[0031] 1) The present invention designs an edge-cloud collaborative data-driven load frequency control (ECCDD-LFC) method. In this method, the LFC controller in the traditional distribution network control center of the substation area is replaced by an independent decision-making cloud agent, each regulation unit is set as an independent decision-making edge agent, and the edge-cloud collaborative framework is combined with deep meta-reinforcement learning to enable interaction among units and achieve edge-cloud collaborative control. This method is a complementary collaboration between cloud computing and edge computing, which can unify the advantages of centralized control and distributed control, aiming to combine the advantages of centralized LFC and distributed LFC and fully mobilize the autonomy of the edge load frequency control unit and the global optimization ability of the control center.
[0032] 2) The present invention designs an emergent computing large-scale multi-agent deep meta deterministic policy gradient (ECLSMA-DMDPG). The multi-agent deep deterministic policy gradient combines the strong perception ability of deep learning with the strong decision-making ability of reinforcement learning. It has the ability to collaborate with multiple agents and is suitable for solving distributed LFC problems. Its structure includes a parallel system, an explorer, an emergency agent, a leader, and an experience pool. This method uses a centralized training and decentralized execution strategy to achieve multi-agent collaboration, uses meta-learning technology to improve the robustness of the algorithm, and uses emergent computing to generate high-value samples to obtain a higher-performance policy.
[0033] 3) The present invention considers the optimal coordinated control among multiple units. By combining intelligent control algorithms with multiple agents, it solves the communication ability among a large number of units to a certain extent, enabling the units to make independent decisions and balancing the interests of different units at the same time. Description of the Drawings
[0034] Figure 1 is the edge-cloud collaborative data-driven load frequency control model according to the embodiment of the present invention;
[0035] Figure 2 is the edge-cloud collaborative data-driven load frequency control transfer function model according to the embodiment of the present invention;
[0036] Figure 3 is the MDP modeling of the edge-cloud collaborative data-driven load frequency control according to the embodiment of the present invention;
[0037] Figure 4 is the training process of the meta-learner according to the embodiment of the present invention;
[0038] Figure 5 is the training framework of the ECLSMA-DMDPG algorithm according to the embodiment of the present invention;
[0039] Figure 6 is the distributed training process of the algorithm according to the embodiment of the present invention. Detailed Embodiment
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0041] A load frequency control method for a low-voltage substation distribution network according to an embodiment of the present invention mainly includes the following steps:
[0042] S1. Propose the LFC modeling of the substation distribution network, and model the microcomputer, distributed power source, wind power generation, and photovoltaic power generation in the substation distribution network system unit.
[0043] S2. Introduce edge cloud collaborative data-driven (ECCDD) on the basis of traditional LFC, build an ECCDD-LFC framework, use the cloud agent to collect the frequency deviation data of the substation distribution network and output the total adjustment command, and use the edge agents of each unit to output the participation variables of each unit.
[0044] S3. Considering the interaction between each agent and its environment, introduce meta-reinforcement learning into the ECCDD-LFC framework. Model the base learner and the meta-learner in the MDP modeling, where the meta-learner is responsible for the initialization of the algorithm to explore the noise variance, and the base learner is responsible for the training of specific tasks.
[0045] S4. To improve the ability of the agent to learn new tasks and reduce the sample complexity, propose reinforcement learning training. In the training of the meta-learner, improve the meta-reinforcement learning algorithm, design the internal and external meta-parameter update rules, obtain the meta-exploration noise variance, and improve the convergence speed and environmental adaptability for different tasks. In the training of the base learner, through the MADDPG algorithm, change the gradients of the local policies of all agents to optimize the global objective. Through the ECLSMA-DMDPG algorithm, use a large number of samples to update its own network parameters.
[0046] Specifically, the main content of the load frequency control method for the low-voltage substation distribution network according to the embodiment of the present invention is as follows:
[0047] Establish an ECCDD-LFC control model as Figure 1 shown, replace the LFC controller in the traditional LFC with a cloud agent, and use the edge agent to process the edge unit and device unit of the system.
[0048] Establish an ECCDD-LFC transfer function model as Figure 2As shown in the figure, the purpose of edge computing is to place rich cloud computing resources at the "edge" of the network to achieve proximal processing of data. Edge computing sinks the capabilities of cloud computing to the edge and device sides, enabling hierarchical computing in a system where the cloud and the edge coexist. To meet the actual needs of distributed power grid connection for regulation, this patent establishes an information-energy coupling architecture and designs a collaborative "cloud-edge" optimization strategy based on the information layer. The cloud-edge collaboration LFC framework includes a cloud agent and n edge agents. The role of the cloud agent is to collect the frequency deviation data of the distribution network in the island area and output the total regulation command.
[0049] Each edge agent corresponds to each unit, and its role is to output the participation factor of each unit. The product of the total regulation command and the participation factor is the regulation command for each adjustable resource, and each adjustable resource adjusts the output of the adjustable resource according to the regulation command. Within the LFC framework of edge-cloud collaboration, secondary and tertiary regulations can be coordinated simultaneously to achieve the optimal combination of frequency deviation and generation cost.
[0050] Furthermore, the power mismatch in the distribution network of the area is used to obtain the frequency deviation by Equation (1). The cloud agent collects the frequency deviation of the distribution network in the area and generates a command to be passed to the edge agent to control the output power of diesel generators, controllable loads, and fuel cells. The power output of photovoltaic units is determined by temperature and light, and the power output of wind turbines is determined by wind speed. Due to the uncertainty of weather conditions, photovoltaic units and wind turbines are regarded as uncertain disturbances, and there are also random disturbances in load control.
[0051] The frequency stability of the distribution network in the area depends on the active power balance of the system, and the system frequency deviation and the power deviation are related as follows.
[0052] (1)
[0053] Edge computing sinks the capabilities of cloud computing to the edge and device sides, enabling hierarchical computing in a system where the cloud and the edge coexist. To meet the actual needs of distributed power grid connection for regulation, this patent establishes an information-energy coupling architecture and designs a collaborative "cloud-edge" optimization strategy based on the information layer. Each edge agent corresponds to each unit, and its role is to output the participation factor of each unit. The product of the total regulation command and the participation factor is the regulation command for each adjustable resource, and each adjustable resource adjusts the output of the adjustable resource according to the regulation command. Within the LFC framework of edge-cloud collaboration, secondary and tertiary regulations can be coordinated simultaneously to achieve the optimal combination of frequency deviation and generation cost.
[0054] The following are described separately:
[0055] (1)MDP (Markov Decision Process) model based on a learner
[0056] The objective function of the cloud agent is as follows. For each unit, to maximize its output power, it is necessary to ensure the minimum frequency deviation and the maximum output power of itself. Therefore, the objective function of the i-th edge agent is as shown. The specific situation is as follows:
[0057] (2)
[0058] In the formula, is the objective function of the cloud agent, is the objective function of the i-th edge agent, is the frequency deviation, is the unit output, T is the total operation time, is the total power generation cost, 、 、 are the conversion factors of the power generation cost.
[0059] The constraint conditions of the cloud agent are as follows:
[0060] (3)
[0061] In the formula is the total power command, and are the upper and lower limits of the generator set respectively, is the unit ramp rate, is the power command input to the i-th unit, where is the additional output of the i-th unit.
[0062] (2)Meta-learner MDP Modeling
[0063] In this patent, meta-reinforcement learning is introduced, and its goal is to learn a policy that can quickly adapt to new tasks. This patent also introduces a meta-learner and a base-learner. The meta-learner is responsible for the initialization of the algorithm to explore the noise variance, while the base-learner is responsible for the training of specific tasks. Due to the structure of the base-learner and the meta-learner, it is necessary to model the base-learner and the meta-learner separately in MDP modeling. MDP modeling includes three typical tuples: the control space, the action space, and the reward function. The MDP modeling is as Figure 3As shown in the figure. In order to enable the cloud agent to obtain a more accurate detection range during detection, the control space is set according to formula (4). At this time, the state space of the cloud agent is shown in formula (5). For the edge agent of any unit, the control space is set as shown in formula (9). According to the frequency deviation and the total adjustment instruction sent by the cloud agent, the state space of this unit is generated as shown in formula (10). The cloud agent collects the frequency deviation of the distribution network in the substation area, then outputs the adjustment command in the distribution network in the substation area to each edge agent, and finally each edge agent outputs the state space of each unit. In order to maximize the objective function (10) of the edge agent and balance the global benefit and the unit benefit, a penalty term is introduced into the reward function, as shown in formula (11).
[0064] The cloud agent outputs the total adjustment command of each unit in the distribution network in the substation area by collecting the frequency of the distribution network in the substation area and the output state of the unit, and the control interval is 4s. The cloud agent is responsible for outputting the total adjustment command. In order to enable the cloud agent to obtain a smaller detection range during detection to help the agent obtain better detection results, the action space is set to 1% of the total power generation instruction. Its action space is as follows:
[0065] (4)
[0066] The state space includes the frequency deviation of the microgrid and its integral, as well as the total output power of the unit:
[0067] (5)
[0068] where is the frequency deviation.
[0069] is the reward value of the meta-learner, is the average reward of the i-th edge agent when the episode is equal to 90, is the average reward of the cloud agent when the situation is equal to 90.
[0070] The reward function of the basic learner is as follows:
[0071] (6)
[0072] (7)
[0073] where is the reward function of the cloud-based agent, is the penalty function.
[0074] The i-th edge agent is responsible for outputting the participation factor of the i-th unit. Therefore, the adjustment instruction of the i-th unit is as follows
[0075] (8)
[0076] In the formula, is the participation factor of the i-th unit, is the adjustment instruction of the i-th agent.
[0077] Set the action space of the i-th agent in the first n - 1 units as the participation factor of the i-th unit. The action space is as follows:
[0078] (9)
[0079] where is the action of the i-th agent. Only outputting the allocation factors of n - 1 adjustment units can satisfy the constraints of formula (1).
[0080] The edge agent of the i-th unit generates the participation factor of this unit according to the frequency deviation and the total adjustment instruction sent by the cloud agent.
[0081] (10)
[0082] In the formula, is the state of the i-th edge agent, is the adjustment output power of the i-th unit.
[0083] The reward function of the i-th agent in the edge agent is as follows:
[0084] (11)
[0085] where and are weight coefficients, is the penalty term.
[0086] Among them
[0087] (12)
[0088] where is the unit overshoot penalty term of the i-th edge agent.
[0089] (3) Reinforcement learning
[0090] To solve the problems existing in current deep reinforcement learning, a meta-reinforcement learning algorithm is designed to learn useful meta-knowledge from a set of related tasks, enabling the agent to acquire the ability to improve the learning speed of new tasks and reduce the sample complexity, so as to achieve highly robust LFC. The training of this patent is divided into out-loop training and in-loop training. Out-loop training means that the meta-learner adopts the traditional DQN (Deep Q Network) learning algorithm. Through reasonable reward function settings, it uses the greedy strategy to explore the most suitable noise variance in different random environments, and finally obtains a strategy that can output the corresponding strategy of the most suitable OU noise variance in different environments. In-loop training refers to that while performing out-loop training, the centralized training strategy in the base learner is used in the MADDPG algorithm to train the cloud agent with n - 1 edge agents. This enables each unit in the offline training to obtain a policy function that is beneficial to global frequency difference optimization and its own interest maximization. Each unit no longer needs to communicate in the online application.
[0091] Reinforcement learning includes both a meta-learner and a base learner, specifically as follows:
[0092] 1) Meta-learner training
[0093] Meta-learning can achieve various hyperparameter adjustments: training of hyperparameters, structure of neural networks, initialization of neural networks, selection of optimizers, parameters of neural networks, definition of loss functions, backpropagation, etc. This patent designs a meta-reinforcement learning algorithm to solve the disadvantages of poor convergence of neural network algorithms and that the training results are only applicable to current tasks and environments. Its basic idea is to design an internal and external meta-parameter update rule to obtain a set of meta-exploration noise variances, improving the convergence speed of the model for different tasks and environmental adaptability.
[0094] A distributed meta-deep reinforcement learning training framework is introduced in the training of noise agents. In the training of deep reinforcement learning, the initialization parameter that has the greatest impact on algorithm convergence and trainability is the exploration noise variance. According to the idea of deep meta-reinforcement learning, in order to enable the algorithm to obtain a more universal strategy, the meta-learner adjusts the initialization exploration noise variance of the algorithm and names it the noise agent. The structure of the complete DDQN algorithm agent (selecting exploration actions through the greedy strategy) is included in the noise agent, as follows.
[0095] (13)
[0096] Among them, is the role of the noise agent in the (k + 1)-th step in the i-th parallel system, is the role of the noise agent in the i-th parallel system, is the Q function of the noise agent in the (k + 1)-th step, is the random action, It is an action set.
[0097] According to the idea of deep meta-reinforcement learning, in order to enable the algorithm to obtain a more general strategy, the meta-learner adjusts the initial exploration noise variance of the algorithm and names it the noise proxy. The training process of the noise proxy is as Figure 4 shown. Initialize the ECCDD-LFC proxy. The noise proxy is executed according to formula (13) to obtain random interference, calculate the loss function according to formula (16) to update the parameters, collect qualified samples, and finally output the parameters.
[0098] In the main network of DDQN, select the action with the largest Q value, that is
[0099] (14)
[0100] Then calculate the target Q value in the target network according to the selected agent:
[0101] (15)
[0102] where γ is the discount factor, used to balance the current and future rewards. ω is the weight parameter in the main network, and ω′ is the weight parameter in the target network.
[0103] Calculate the loss function according to Bellman's idea:
[0104] (16)
[0105] 2) Basic learner training
[0106] The training of the ECLSMA-DMDPG basic learner is as Figure 5 shown. Obtain the trained noise proxy through the meta-learner, generate noise variance after being disturbed, and input it into 16 parallel systems to obtain noise samples. This parallel system includes various explorers as well as the MDP model and the reward function. At the same time, the obtained samples continuously update the network parameters in the Leader.
[0107] The structure of the ECLSMA-DMDPG algorithm includes a parallel system, an explorer, an emergency agent, a leader, and an experience pool. The overview of this algorithm is as follows:
[0108] Parallel system: This method includes 16 parallel systems. Each parallel system includes a distributed LFC environment for the distribution network of a substation area. Parallel systems 1-8 include 1 cloud agent and n - 1 edge agents. Parallel systems 9-16 include 1 conventional LFC controller and 1 power distributor. In each parallel system, the environment corresponds to different random interferences to enrich the diversity of samples.
[0109] Emergency Agents: The main responsibility of these emergency agents is to learn by effectively interacting with the environment, obtain emergency samples and provide them to the leader. These samples constitute a high-value sample that can effectively guide the training of the leader. In each parallel system, the emergency agents are divided into two types: control emergency agents containing different controllers; distributed emergency agents containing different main power distributors, and these power distributors give reasonable actions according to the interaction with the environment, generate emergent samples and put them into the experience pool. Each emergency control uses controllers with various different principles, including fuzzy PID and fuzzy fractional-order PID. The controller coefficients are set manually and by coefficient optimization; the control emergency agents based on the fuzzy PID control algorithm are used in parallel systems 11 - 13, and the control emergency agents based on the fuzzy FOPID control algorithm are used in parallel systems 14 - 16.
[0110] Explorers: The main responsibility of the explorers is to explore the parallel system environment they are in. Each explorer only contains a participant network with a different network model, and different explorers adopt different exploration principles. The explorers in parallel systems 1 to 3 use Gaussian noise for environment exploration, and these explorers are named Gaussian explorers, and their actions are shown as follows.
[0111] (17)
[0112] where is the policy function of the i-th Gaussian explorer.
[0113] The explorers in parallel systems 4 to 6 adopt the greedy strategy, and this type of explorer is named ε-explorer. Its actions are discretized, and the greedy strategy is used to sample the discrete random actions, where the interval between different actions is 0.1, the action space range is [-60, 60], and the total power generation command range corresponding to the output of this type of explorer through the following exploration actions is [-6000, 6000] (MW).
[0114] (18)
[0115] where is the total number of actions that can be executed in the st state, ε is the greedy coefficient, is the policy function of the l-th ε-explorer.
[0116] OU noise is used to explore the environment in parallel systems 7 to 9, and it is named OU explorer. All explorers deposit the explored samples into the common experience pool, so as to provide more diverse samples for the leader and improve the training efficiency. The explorers regularly obtain the latest network parameters from the leader and update their own actor network parameters.
[0117] (19)
[0118] Leader: The leader updates its network parameters at each step and then periodically sends these network parameters to the resource manager for parameter update.
[0119] Emergency agent:
[0120] 1) Control the emergency agent
[0121] Control the emergency agent to execute this scenario. Each emergency control uses controllers with various different principles, including fuzzy PID and fuzzy fractional-order PID. The controller coefficients are set manually and by coefficient optimization; the latter is mainly considered within the international environmental assessment index area:
[0122] (20)
[0123] Where is the fitness function of the control emergency agent at time t, and t is the discrete time. The control emergency agent based on the fuzzy PID control algorithm is used for the parallel systems 11 - 13, and the control emergency agent based on the fuzzy FOPID control algorithm is used for the parallel systems 14 - 16.
[0124] 2) Dispatch the emergency agent
[0125] Emergency events containing different heuristic algorithms are called scheduling emergency events. Multiple heuristic algorithms are used as the scheduling algorithms for the AGC power distribution system. For optimization, the fitness functions of the selected algorithms are as follows:
[0126] (21)
[0127] Where, is the fitness function of the assigned emergency event at time t, and T is the total optimization step size.
[0128] The overall training process is as Figure 6 shown. Initialize the explorer parameters, obtain the initial state through the noise agent policy. The explorer and the emergency agent act according to formulas (17) - (19) and formulas (20) - (21) respectively. After obtaining the samples, calculate the target value according to formula (15) and update the parameters according to its gradient to obtain qualified samples.
[0129] To verify the effectiveness of the proposed algorithm, the present invention presents an LFC model based on the Zhuzhou microgrid system.
[0130] First, ECCDD-LFC was tested in four algorithms (ECLSMA-DMDPG, Ape-x-MADDPG, MATD3, MADDPG). Additionally, 16 LFC algorithms were given. The controller includes the particle swarm optimization fuzzy PI control algorithm (PSO-fuzzy-PI), the particle swarm optimization fuzzy FOPI control algorithm (PSO-fuzzy-FOPI), the genetic algorithm fuzzy PI control algorithm (GA-fuzzy-PI), and the particle swarm optimization algorithm (PSO-PI). The distributor includes the particle swarm optimization algorithm (PSO), the genetic algorithm (GA), the grey wolf optimization algorithm (GWO), and the proportional algorithm (PROP).
[0131]
[0132] Table 1
[0133] As shown in Table 1, the frequency deviation of the ECLSMA-DMDPG algorithm is 0.00864 Hz, and the total cost is $426.1. Compared with the other 19 algorithms, the ECLSMA-DMDPG algorithm reduces the frequency deviation by 37.04% - 66.55% and reduces the power generation cost by 0.75% - 6.80%.
[0134] This is because ECLSMA-DMDPG adopts multiple strategies, including a training strategy of multi-agent mutual cooperation. This embodiment is an emergency operation strategy of a meta-learning technology that improves training efficiency and enhances robustness; the diversified technology enables multi-agent contracts to cooperate and obtains an LFC cooperation strategy with high robustness performance.
[0135] From the above technical solutions, it can be seen that the embodiment of the present invention adopts the edge cloud collaborative data-driven load frequency control method (ECCDD-LFC), which fully mobilizes the node computing resources and the global optimization ability of the cloud. Each adjustment unit is set as an independent decision-making edge agent, and the edge cloud collaborative framework is combined with deep meta-reinforcement learning to enable interaction between units and achieve cloud-edge collaborative control. At the same time, the embodiment of the present invention proposes an emergent computing large-scale multi-agent deep meta-deterministic policy gradient (ECLSMA-DMDPG). This strategy adopts a centralized training and decentralized execution strategy to achieve multi-agent cooperation, adopts meta-learning technology to improve the robustness of the algorithm, and adopts emergent computing to generate high-value samples to improve performance.
[0136] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the above method.
[0137] In another aspect, the present invention also discloses a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the above method.
[0138] In yet another embodiment provided by the present application, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute any of the load frequency control methods for the low-voltage distribution network in the above embodiments.
[0139] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples and beneficial effects of related content, reference can be made to the corresponding parts in the above method.
[0140] The embodiments of the present application also provide an electronic device, including a processor, a communication interface, a memory and a communication bus. Among them, the processor, the communication interface and the memory complete communication with each other through the communication bus.
[0141] The memory is used to store a computer program.
[0142] The processor is used to implement the above load frequency control method for the low-voltage distribution network when executing the program stored in the memory.
[0143] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0144] The communication interface is used for communication between the above electronic device and other devices.
[0145] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0146] The above-mentioned processor may be a general-purpose processor, including a central processing unit (English: Central Processing Unit, abbreviated: CPU), a network processor (English: Network Processor, abbreviated: NP), etc.; it may also be a digital signal processor (English: Digital Signal Processing, abbreviated: DSP), an application-specific integrated circuit (English: Application Specific Integrated Circuit, abbreviated: ASIC), a field-programmable gate array (English: Field-Programmable Gate Array, abbreviated: FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0147] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media may be magnetic media (for example, floppy disks, hard disks, magnetic tapes), optical media (for example, DVDs), or semiconductor media (for example, solid state disks (SSDs)).
[0148] It should be noted that in the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0149] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.
[0150] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A load frequency control method for a low-voltage distribution network in a low-voltage substation area, characterized in that, Implement cloud-edge collaborative control by using a pre-constructed load frequency control model. The construction steps of the load frequency control model are as follows: S1. First, propose the LFC modeling of the distribution network in the substation area, and model the microcomputer, distributed power supply, wind power generation, and photovoltaic power generation in the system unit of the distribution network in the substation area; S2. Introduce edge cloud collaborative data-driven on the basis of the LFC built in S1, build the ECCDD-LFC framework, use the cloud agent to collect the frequency deviation data of the distribution network in the substation area and output the total adjustment command, and use the edge agents of each unit to output the participation variables of each unit; S3. Considering the interaction between each agent and its environment, introduce meta-reinforcement learning in the ECCDD-LFC framework, and model the basic learner and the meta-learner in the MDP modeling. The meta-learner is responsible for the initialization of the algorithm to explore the noise variance, and the basic learner is responsible for the training of specific tasks; S4. In the training of the basic learner, through the MADDPG algorithm, change the gradients of the local policies of all agents to optimize the global objective, and through the ECLSMA-DMDPG algorithm, use a large number of samples to update its own network parameters; The specific content of S2 includes: Establish an ECCDD-LFC transfer function model. Edge computing sinks the capabilities of cloud computing to the edge side and the device side to achieve hierarchical computing in a system where the cloud and the edge coexist; Establish an information-energy coupling architecture and design a collaborative "cloud-edge" optimization strategy based on the information level; the cloud-edge collaborative LFC framework includes a cloud agent and n edge agents; the role of the cloud agent is to collect the frequency deviation data of the distribution network in the island substation area and output the total adjustment command; Each edge agent corresponds to each unit, and its role is to output the participation factor of each unit; the product of the total adjustment command and the participation factor is the adjustment command of each adjustable resource, and each adjustable resource adjusts the output of the adjustable resource according to the adjustment command; within the LFC framework of cloud-edge collaboration, secondary and tertiary adjustments can be coordinated simultaneously to achieve the optimal combination of frequency deviation and generation cost.
2. The load frequency control method for the low-voltage substation distribution network according to claim 1, wherein: The objective function of the cloud proxy is as follows. For each unit, in order to maximize its output power, it is necessary to ensure the minimum frequency deviation and the maximum output power of itself. Therefore, the objective function of the i-th edge proxy is as shown, and the specific situation is as follows: (2) In the formula, is the objective function of the cloud proxy, is the objective function of the i-th edge proxy, is the frequency deviation, is the unit output, T is the total operation time, is the total power generation cost, , , are the power generation cost conversion factors; The constraint conditions of the cloud agent are as follows: (3) where is the total power command, and are the upper and lower limits of the generating unit respectively, is the unit climbing rate, is the power command input to the i-th unit, where is the additional output of the i-th unit.
3. The load frequency control method for the low-voltage substation distribution network according to claim 2, wherein: The MDP modeling steps in step S3 include three typical tuples: the control space, the action space, and the reward function: Its action space is as follows: (4) The state space includes the frequency deviation of the microgrid and its integral, as well as the total output power per unit: (5) wherein is the frequency deviation; is the reward value for the meta-learner, is the average reward of the i-th edge agent when the episode is equal to 90, is the average reward of the cloud agent when the situation is equal to 90; The reward function of the basic learner is as follows: (6) (7) Among them is the reward function for the cloud-based agent, is the penalty function; The i-th edge agent is responsible for outputting the participation factor of the i-th unit. Therefore, the adjustment instruction of the i-th unit is as follows: (8) wherein, is the participation factor of the i-th unit, is the adjustment instruction of the i-th agent; Set the action space of the i-th agent in the first n - 1 units as the participation factor of the i-th unit, and the action space is as follows: (9) wherein is the action of the i-th agent; only the allocation factors of n - 1 adjustment units need to be output to satisfy the constraints of formula (1); The edge agent of the i-th unit generates the participation factor of its own unit according to the frequency deviation and the total adjustment instruction sent by the cloud agent; (10) In the formula, is the state of the i-th edge agent, is the regulated output power of the i-th unit; The reward function of the i-th agent in the edge agent is as follows: (11); wherein and are weight coefficients, is a penalty term; Among them (12) Among them is the unit overshoot penalty term for the i-th edge agent.
4. The load frequency control method for low-voltage distribution network in low-voltage substation area according to claim 3, wherein: The reinforcement learning in step S3 includes the training of the meta-learner and the training of the basic learner, which are specifically as follows: 1) Meta-learner training A distributed meta-depth reinforcement learning training framework is introduced in the training of the noise proxy. In the training of deep reinforcement learning, the initialization parameter that has the greatest impact on the algorithm's convergence and trainability is the exploration noise variance. According to the idea of deep meta-reinforcement learning, in order to enable the algorithm to obtain a more general strategy, the meta-learner adjusts the initialization exploration noise variance of the algorithm and names it the noise proxy; the complete DDQN algorithm proxy, that is, the structure that selects exploration actions through the greedy strategy, is included in the noise proxy, as follows: (13) wherein, is the effect of the noise agent in the (k + 1)-th step in the i-th parallel system, is the effect of the noise agent in the i-th parallel system, is the Q function of the noise agent in the (k + 1)-th step, is the random effect, is the effect set; In the main network of DDQN, select the action with the largest Q value, that is (14) Then calculate the target Q value in the target network according to the selected agent: (15) where γ is the discount factor used to balance the current and future rewards, ω is the weight parameter in the main network, and ω′ is the weight parameter in the target network; Calculate the loss function according to Bellman's idea: (16) 2) Basic learner training The algorithm structure of the ECLSMA-DMDPG basic learner training includes a parallel system, an explorer, an emergency agent, a leader, and an experience pool; Parallel system: This method contains 16 parallel systems; each parallel system includes a distributed LFC environment for a distribution network in a substation area. Parallel systems 1-8 include 1 cloud agent and n-1 edge agents, and parallel systems 9-16 include 1 conventional LFC controller and 1 power distributor. In each parallel system, the environment corresponds to different random perturbations to enrich the diversity of samples; Emergency agent: In each parallel system, the emergency agent is divided into two types: a control emergency agent containing different controllers; a distributed emergency intelligent agent containing different main power distributors, and these power distributors give reasonable actions according to the interaction with the environment, generate emergent samples and put them into the experience pool; each emergency control uses various controllers with different principles, including fuzzy PID and fuzzy fractional-order PID; the controller coefficients are set manually and by coefficient optimization; the control emergency agent based on the fuzzy PID control algorithm is used for parallel systems 11-13, and the control emergency agent based on the fuzzy FOPID control algorithm is used for parallel systems 14-16; Explorer: Each explorer only contains a participant network with a different network model, and different explorers adopt different exploration principles. The explorers in parallel systems 1 to 3 use Gaussian noise to explore the environment, and these explorers are named Gaussian explorers. Their actions are as follows: (17) wherein is the policy function of the i-th Gaussian explorer; The explorers in parallel systems 4 to 6 adopt the greedy strategy and name this type of explorer ε-explorer; their actions are discretized, and the greedy strategy is used to sample the discrete random actions, where the interval between different actions is 0.1, and the action space range is [-60, 60]. The total power generation command range corresponding to this type of explorer through the following exploration action output is [-6000, 6000]; (18) where is the total number of actions that can be executed in the st state, ε is the greedy coefficient, is the policy function of the l-th ε-explorer; OU noise is used to explore the environment in parallel systems 7 to 9 and is named OU explorer; all explorers deposit the explored samples into the common experience pool. The explorers regularly obtain the latest network parameters from the leader and update their actor network parameters; (19) Leader: The leader updates its network parameters at each step and then periodically sends these network parameters to the resource manager for parameter update.
5. The load frequency control method for low-voltage distribution network in a low-voltage substation area according to claim 4, characterized in that: The emergency agent includes: 1) Controlling the emergency agent The control emergency agent executes this scenario; each emergency control uses controllers with various different principles, including fuzzy PID and fuzzy fractional-order PID; the controller coefficients are set manually and by coefficient optimization; (20) where is the fitness function for controlling the emergency agent at time t, where t is discrete time; the control emergency agent based on the fuzzy PID control algorithm is used for the parallel systems 11-13, and the control emergency agent based on the fuzzy FOPID control algorithm is used for the parallel systems 14-16; 2) Dispatching the emergency agent Emergency events containing different heuristic algorithms are called scheduling emergency events; multiple heuristic algorithms are used as the scheduling algorithms for the AGC power distribution system; the fitness functions of the selected algorithms are as follows: (21) In the formula, is the fitness function of the allocated emergency event at time t, and T is the total optimization step length.
6. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the method according to any one of claims 1 to 5.
7. A computer device comprising a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to execute the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Full-distributed load frequency control method for island microgrid
CN116565952A
Intelligent power grid system based on cloud edge fusion architecture and scheduling method
CN116739236A