Interaction control method and device for distributed power system with industrial load and power grid
Patent Information
- Application Number
- CN202211625787.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-12-16
AI Technical Summary
虽然经典优化方法、基于规划的方法和启发式算法是实现工业生产经济运行与优化调度的常用方法,但是工业生产运行过程中的环境条件变化复杂,算法需要处理的数据量大,经典优化方法、基于规划的方法和启发式算法都难以实时高效的实现工业生产过程优化控制
(1)在满足工业生产过程运行约束的前提下,建立工业生产过程的运行与成本模型;
Smart Images

Figure CN115912486B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of magnesium oxide production and energy optimization technology, and in particular to a method and apparatus for interactive control of a distributed power supply system with industrial load and the power grid. Background Technology
[0002] Modern industrial production processes involve various industrial distributed power sources, energy storage, loads, and related monitoring and protection devices, enabling self-control and management as well as interaction with the power grid. Existing research and practice demonstrate that interactive control between industrial distributed power sources and the power grid is an effective way to maximize the efficiency of industrial distributed power systems. In recent years, countries worldwide have increased their attention and development efforts regarding the interactive control technology between industrial distributed power systems and the power grid.
[0003] Currently, the technology for interaction between industrial distributed power generation systems and the power grid is developing rapidly, and its achievements are being vigorously promoted and applied. The interaction control of industrial distributed power generation systems and the power grid integrates advanced information technology, control technology, and power technology. While improving the utilization efficiency of distributed renewable energy and providing diversified energy supply forms, it optimizes energy efficiency, economic efficiency, and environmental benefits compared to traditional centralized power grids. As an important component of the smart grid, the main goal of the smart distribution system is to solve the operational problems of a large number of dispersed industrial distributed power sources in the power system, in order to adapt to the needs of power grid construction in the new era, which is characterized by greenness, efficiency, and harmony. The interaction control technology between industrial distributed power generation systems and the power grid acts as a link between industrial distributed power sources and the distribution network. It enables the large-scale integration of industrial distributed power sources, maximizing the utilization of renewable energy, while avoiding the impact of intermittent power sources on the safe operation of the distribution network and ensuring power quality for users within the network. Furthermore, the interaction control technology between industrial distributed power generation systems and the power grid is also an important pathway for future smart distribution networks to achieve self-healing, user-side interaction, and demand response.
[0004] However, traditional centralized control methods for industrial distributed power sources suffer from low flexibility and poor robustness. Furthermore, industrial production processes involve various types of energy inputs and outputs, encompassing multiple energy conversion units. Both energy input and consumption exhibit volatility and randomness. Coupled with differences in response time across various stages, imbalances exist at different spatial and temporal scales in energy input, conversion, distribution, and consumption within industrial production processes. This often results in a mixed-integer nonlinear programming (MINLP) problem, posing significant challenges to energy management and optimal scheduling in industrial production. While classical optimization methods, planning-based methods, and heuristic algorithms are commonly used for achieving economical operation and optimal scheduling in industrial production, the complex environmental conditions and large data volumes required for processing make it difficult for these methods to achieve real-time and efficient optimal control. Therefore, seeking effective distributed control methods and optimal control strategies to achieve high-quality and efficient energy management is crucial for improving system performance. Summary of the Invention
[0005] To address the technical problems raised in the background, this invention provides a method and apparatus for interactive control of a distributed power supply system with industrial loads and the power grid. A two-layer control architecture for interaction between the industrial distributed power supply system and the power grid is designed. In the lower-layer control, a distributed control method using a multi-agent consensus algorithm is established to optimize the industrial production process, maximizing the utilization of renewable energy in the industrial production process and enabling "plug-and-play" operation of distributed gas turbines and distributed energy storage devices. Simultaneously, the upper-layer control uses a distributed deep reinforcement learning algorithm to minimize the operating costs of industrial production, achieving economical operation of industrial production while ensuring safe operation.
[0006] To achieve the above objectives, the present invention employs the following technical solution: An interactive control method for a distributed power supply system with industrial loads and the power grid, the control method comprising the following steps: Step 1: Establish the operation model and cost model of the distributed power system in the industrial production process. The distributed power system in the industrial production process is the industrial distributed power system, which includes photovoltaic, wind power, distributed gas turbines and distributed energy storage. Step 2: Establish communication between adjacent distributed gas turbines and adjacent distributed energy storage, so that adjacent gas turbines can exchange power information and adjacent energy storage can exchange power and state of charge information. Step 3: Use the multi-agent consensus algorithm to make the output power of the distributed gas turbine and the distributed energy storage consistent with the state of charge of the distributed energy storage. Step 4: Design the state space, action space, and reward function for the interaction between the industrial distributed power supply system and the power grid, and establish a Markov model for the control of the interaction between the industrial distributed power supply system and the power grid. Step 5: Use the TD3 algorithm to establish a deep reinforcement learning agent for distributed gas turbines and distributed energy storage; Step 6: Train the deep reinforcement learning agent offline using a centralized training method; Step 7: Implement interactive control between the industrial distributed power system and the power grid using distributed decision-making.
[0007] Furthermore, the control method includes lower-level control and upper-level control; The lower-level control targets industrial distributed power sources of the same type, taking into account the weak communication connections between them. It uses a multi-agent consensus distributed control method to automatically allocate the output power of industrial distributed power sources in the industrial production process and realize the "plug-and-play" characteristic of industrial distributed power sources, thereby improving the flexibility and robustness of the industrial production process. For different types of industrial distributed power sources, the upper-level control adopts a distributed deep reinforcement learning method to optimize and schedule the output of controllable industrial distributed power sources with multiple energy forms in industrial production, while ensuring the safe operation of industrial production and minimizing the operating cost of industrial production.
[0008] Furthermore, the specific steps of the multi-agent consensus algorithm are as follows: Step 1: Adjacent distributed gas turbines and adjacent distributed energy storage systems communicate with each other to exchange power information and state of charge information; Step 2: The control method includes upper-level control, which sends control signals to a distributed gas turbine and a distributed energy storage device in the industrial production process. The gas turbine and the energy storage device update their output power according to the control signals. Step 3: Other distributed gas turbines update their output power based on the power information of adjacent gas turbines and themselves; other distributed energy storage updates their output power based on the power and state of charge information of adjacent energy storage devices and themselves.
[0009] Furthermore, based on the operating costs of industrial distributed power sources and distributed energy storage, industrial load electricity prices, external grid time-of-use electricity prices, and the operating modes of various equipment in the industrial production process, the state space, action space, and reward function of the deep reinforcement learning agent for the interaction between the industrial distributed power source system and the power grid are designed, and a Markov model for the interactive control of the industrial distributed power source system and the power grid is established accordingly.
[0010] Furthermore, the distributed deep reinforcement learning method employs the TD3 deep reinforcement learning algorithm, which improves the DDPG training process using the following scheme: First, a truncated double-Q learning method is used to set up two value networks and two target value networks to alleviate the overestimation problem that occurs during the training of value networks; Second, noise is added to the target policy network to smooth it out. Third, reduce the update frequency of the policy network and the three target networks. That is, update the value network once per round, but update the policy network and the three target networks once every k rounds.
[0011] Furthermore, two different deep reinforcement learning agents are used to control the distributed gas turbine and distributed energy storage respectively; a distributed deep reinforcement learning architecture with centralized training and distributed decision-making is adopted, in which each deep reinforcement learning agent only needs to observe the operating information of the equipment near the agent to make control decisions, thereby reducing communication costs in the industrial production process and improving the computation speed of deep reinforcement learning algorithms.
[0012] Furthermore, in the centralized training and distributed execution distributed deep reinforcement learning architecture: during the training process, the observations and actions of all deep reinforcement learning agents are centrally collected into the experience pool, and the policy network of each deep reinforcement learning agent only relies on the local observations obtained from the environment. i Make an action a i The value network updates its parameters based on the observations *s* and actions *a* of all agents obtained from the experience pool, and generates actions *a* to be issued by the policy network. i Evaluation q (s, a, w) i Finally, the policy network evaluates q(s, a, w). i The deep reinforcement learning agent updates its own neural network parameters θ. After the deep reinforcement learning agent has been trained, each deep reinforcement learning agent can make distributed decisions based solely on its local observations.
[0013] The present invention also provides an apparatus for implementing the interactive control method between the distributed power supply system containing industrial loads and the power grid, comprising a processor and a memory; The processor is configured to execute the interactive control method between the distributed power supply system containing industrial loads and the power grid. The memory is used to store the executable instructions of the processor.
[0014] The present invention also provides a computer storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the aforementioned interactive control method between a distributed power supply system with industrial load and the power grid.
[0015] Compared with the prior art, the beneficial effects of the present invention are: (1) Under the premise of satisfying the operational constraints of the industrial production process, establish an operation and cost model for the industrial production process; (2) Compared with other control methods, this control method does not rely on a central controller. Each industrial distributed power source can control all industrial distributed power sources by communicating with adjacent industrial distributed power sources. (3) An optimization scheduling method based on multi-agent deep reinforcement learning was adopted. Each deep reinforcement learning agent controls different types of industrial distributed power sources and can only observe some information about the industrial production environment. Therefore, the optimal strategy for the operation of industrial distributed power sources at different time scales can be calculated with the goal of minimizing operating costs. Attached Figure Description
[0016] Figure 1 is a basic control framework diagram of the industrial distributed power supply system and grid interaction control method in the embodiment of the present invention; Figure 2 is a diagram of the multi-agent consensus control structure of industrial distributed power supply based on graph theory in an embodiment of the present invention. Figure 3 is a flowchart of the neural network update process of the TD3 algorithm in an embodiment of the present invention; Figure 4 is a diagram of the distributed deep reinforcement learning structure of "centralized training and distributed decision-making" in the embodiment of the present invention. Detailed Implementation
[0017] The specific embodiments provided by the present invention will be described in detail below with reference to the accompanying drawings.
[0018] Example 1 A method for interactive control between a distributed power supply system with industrial loads and the power grid, the control method comprising the following steps: Step 1: Establish the operation model and cost model of the distributed power system in the industrial production process. The distributed power system in the industrial production process is the industrial distributed power system, which includes photovoltaic, wind power, distributed gas turbines and distributed energy storage. Step 2: Establish communication between adjacent distributed gas turbines and adjacent distributed energy storage, so that adjacent gas turbines can exchange power information and adjacent energy storage can exchange power and state of charge information. Step 3: Use the multi-agent consensus algorithm to make the output power of the distributed gas turbine and the distributed energy storage consistent with the state of charge of the distributed energy storage. Step 4: Design the state space, action space, and reward function for the interaction between the industrial distributed power supply system and the power grid, and establish a Markov model for the control of the interaction between the industrial distributed power supply system and the power grid. Step 5: Use the TD3 algorithm to establish a deep reinforcement learning agent for distributed gas turbines and distributed energy storage; Step 6: Train the deep reinforcement learning agent offline using a centralized training method; Step 7: Implement interactive control between the industrial distributed power system and the power grid using distributed decision-making.
[0019] In this invention, both the distributed gas turbines and distributed energy storage utilize PQ control. The power reference signal is provided externally, with the external power grid providing voltage and frequency support for industrial production. Different gas turbines and energy storage devices are connected via a consensus algorithm. Two deep reinforcement learning agents control one gas turbine and one energy storage device, respectively referred to as the gas turbine deep reinforcement learning agent and the energy storage device deep reinforcement learning agent. Other gas turbines and energy storage devices will follow the actions of the gas turbines and energy storage devices controlled by the deep reinforcement learning agents through the consensus algorithm. The basic control architecture is as follows: Figure 1 As shown. Specific embodiments of the present invention are as follows: Step 1: Establish an operational and cost model for the industrial production process. The power output of energy sources such as photovoltaics and wind power is highly volatile and uncertain. To ensure the safe operation of industrial production, all industrial distributed power supply equipment and buses in the industrial production process must meet the operational constraints of industrial production. Establishing an operation and cost model for the industrial production process is the foundation for studying the interaction and control between industrial distributed power supply systems and the power grid.
[0020] Step 1.1: Establish an operational model for industrial production, taking into account the upper and lower limits of the output power of distributed gas turbines, the upper and lower limits of the charging and discharging power of distributed energy storage devices, the state of charge of energy storage devices, and the power balance of the power grid bus.
[0021] Step 1.2: Establish an industrial production process cost model considering fuel costs, gas turbine operating costs, energy storage equipment operating costs, grid electricity purchase and sale costs, and revenue from selling electricity to industrial loads during the industrial production process.
[0022] This article covers industrial production processes including photovoltaic power generation, wind power generation, distributed gas turbines, distributed energy storage, and industrial applications. Composition of industrial load.
[0023] First, we model the distributed gas turbines, distributed energy storage, and grid bus in the industrial production process: 1) Gas turbine: Gas turbines provide adjustable power supply for industrial production by burning natural gas, effectively reducing the dependence of industrial production processes on external power grids. Its fuel cost is expressed as a quadratic function, as shown in the following formula.
[0024]
[0025] In the formula: C MT (t) represents the total fuel cost of the distributed gas turbine during time period t; P MT, i (t) represents the electrical power output (in kW) of the i-th gas turbine during time period t; a MT, i b MT, i c MT, i Let be the fuel cost coefficient for the i-th gas turbine.
[0026] The output power of the gas turbine meets the constraints:
[0027] In the formula: , These represent the maximum and minimum output power of the gas turbine, respectively.
[0028] 2) Energy Storage: Energy storage devices consist of batteries, which can operate in coordination with renewable energy sources that exhibit randomness and volatility, playing a role in peak shaving and valley filling to ensure the reliability and economy of industrial production processes. Considering the charging and discharging power of the batteries and the state of charge (SOC) of the energy storage, the charging and discharging expression for the energy storage device is:
[0029] Where: SOC i (t) represents the state of charge of the i-th energy storage device at time t; P ES,i (t) represents the electrical power output or absorbed by the i-th energy storage device during time period t (unit: kW); η and ζ are the charging and discharging efficiencies of the energy storage device, respectively; S es This refers to the rated capacity of the energy storage device.
[0030] The cost of energy storage consists of capacity cost and power cost, as shown below:
[0031] In the formula: C ES (t) represents the total cost of distributed energy storage during time period t; E i g represents the rated capacity of the i-th energy storage device; EThe capacity factor is the energy storage capacity cost; g P The power factor is the cost of energy storage power.
[0032] The energy storage device's charging and discharging meet the constraints:
[0033] SOCmin < SOCi (t) < SOCmax In the formula: P ch. max P dis . max The maximum power for charging and discharging energy storage devices, respectively; SOC max SOC min Storage respectively The maximum and minimum states of charge of the energy device.
[0034] Grid bus: The power grid contains a large number of renewable energy power sources, such as wind turbines and solar power. To achieve full utilization of renewable energy, it is assumed that all wind turbine and solar power is connected to the grid. The grid bus must maintain power balance, therefore the grid bus power is expressed as:
[0035] In the formula: P grid (t) Power exchanged between industrial production processes and the external power grid; P PV (t) 、P WT (t) represents the output power of photovoltaic power generation and wind power generation, respectively; P L (t) represents the power consumed by the load during industrial production.
[0036] The cost of purchasing electricity from an external power grid or the revenue from selling electricity during industrial production can be expressed as:
[0037] In the formula: C grid (t) represents the cost of energy exchange with the external power grid during industrial production; σ b (t) 、σ s (t) represents the electricity price for industrial production processes to buy and sell electricity to the external power grid.
[0038] Therefore, the total cost of the industrial production process proposed in this paper can be expressed as:
[0039] In the formula: F is the total operating cost of the industrial production process, σ L The price of electricity sold by an industrial production process to its internal industrial loads.
[0040] Photovoltaic power generation and wind power generation, at their maximum output, along with distributed gas turbines and distributed energy storage, execute the control steps 2 and 3.
[0041] II. Lower-level control of interaction between industrial distributed power systems and the power grid: Step 2: Establish communication between adjacent distributed gas turbines and adjacent distributed energy storage, so that adjacent gas turbines can exchange power information and adjacent energy storage can exchange power and state of charge information.
[0042] For the same type of industrial distributed power source, their communication relationship can be represented by Figure G. μ (V μ , ψ μ ,K μ B μ ) to represent, where V represents a set of nodes, where each node can represent an industrial distributed power source. μ , l Indicates the leader node. The set of edges represents the communication lines between industrial distributed power sources.
[0043] Represents the weight of the edge, if (V μ , i V μ , j ) ∈ V μ If there is a communication connection between them ,on the contrary .
[0044] Represents the leading adjacency matrix, if V μ , i ∈ V μ Capable of receiving top-level deep reinforcement learning agents Control signals ,on the contrary .
[0045] Step 3: Utilize a multi-agent consensus algorithm to optimize the output power of the distributed gas turbine and distributed energy storage. The state of charge of the energy reaches a consistency.
[0046] The control objective of distributed gas turbines is to ensure that the output power of all gas turbines is equal to the power specified by the control signal, that is:
[0047] In the formula: P MT,DRL This represents a reference signal for the gas turbine power provided by a deep reinforcement learning agent, distributed gas turbine... The control expression for a gas turbine can be represented in the following form:
[0048] Where: n MT This represents the number of distributed gas turbines in the industrial production process.
[0049] The control objective of distributed energy storage is to ensure that all energy storage units have equal output power while simultaneously matching the control signal with their output or input power.
[0050] In the formula: P MT,DRL This represents a reference signal for the gas turbine power provided by a deep reinforcement learning agent, distributed storage. The control expression for energy can be represented in the following form:
[0051] Where: n ES This refers to the number of distributed energy storage devices used in industrial production processes.
[0052] Multi-agent consensus control of industrial distributed power sources can be represented using graph theory, and its structure diagram is as follows: Figure 2 As shown, each industrial distributed power source acts as an agent, communicating with its neighboring industrial distributed power sources. A multi-agent consensus algorithm is used to ensure consistency between the output power of each industrial distributed power source and energy storage device and the state of charge of the energy storage device, minimizing power mismatch between industrial distributed power sources and energy storage, and facilitating better interaction and control between the industrial distributed power source system and the power grid.
[0053] Step 3.1: The distributed deep reinforcement learning agent sends control signals to a distributed gas turbine and a distributed energy storage device, respectively; Step 3.2: Upon receiving the control signal, the distributed gas turbine and distributed energy storage device follow the control signal and operate accordingly; Step 3.3: Each distributed gas turbine and distributed energy storage device is regarded as a node in graph theory, and power information and energy storage state of charge information are exchanged between each adjacent node; Step 3.4: Each distributed gas turbine and distributed energy storage device updates itself using a multi-agent consensus algorithm. . output power.
[0054] III. Using distributed deep reinforcement learning algorithms for upper-level optimization control of industrial distributed power systems and grid interaction: Modern industrial production processes involve various controllable forms of distributed power sources, such as gas turbines and energy storage, which increases the difficulty of controlling the interaction between industrial distributed power systems and the power grid. The control conditions and methods differ for different types of industrial distributed power sources. Therefore, it is necessary to establish an industrial production process optimization control model based on distributed deep reinforcement learning, which can improve computational efficiency and reduce communication and operating costs in industrial production processes.
[0055] Step 4: Establish a Markov model of the interaction between the industrial distributed power system and the power grid.
[0056] Step 4.1: The system running time, grid output, output of the l-th gas turbine, and time-of-use electricity price in the industrial production process constitute the state space of the gas turbine deep reinforcement learning agent; the photovoltaic output, wind turbine output, user load, output of the l-th energy storage device, and energy storage charge state in the industrial production process constitute the state space of the energy storage device deep reinforcement learning agent. Step 4.2: In the industrial production process, the output power of the distributed gas turbine and the output power of the distributed energy storage device constitute the action space of the deep reinforcement learning agent of the gas turbine and the deep reinforcement learning agent of the energy storage device, respectively. Step 4.3: Determine the reward function for the distributed deep reinforcement learning agent based on the operating costs of the gas turbine, the operating costs of the energy storage equipment, the time-of-use electricity price for purchasing and selling electricity from the external power grid, and the electricity price for supplying electricity to the internal industrial load.
[0057] Step 4.4: Based on the state space, action space, and reward function, establish a Markov model of the interaction between the industrial distributed power system and the power grid.
[0058] Deep reinforcement learning can often be described as a Markov decision process (MDP). An MDP typically consists of five elements: {S, A, P}. S ,S ' , r, g}, where: S is the state space, representing the set of environmental state information that the agent can observe; A is the action space, representing the set of actions taken by the agent; P s ,s ' Let be the state transition probability, representing the probability that the environment transitions from state S to state S' after the agent takes action 'a'; r is the immediate reward, representing the immediate reward given to the agent by the environment after the agent takes action 'a' in state S; g is the discount factor, representing the magnitude of the influence of the action taken at the current moment on the agent's reward at a future moment. For the interaction process between the industrial distributed power system and the power grid discussed in this paper, its state space, action space, and reward function can be designed as follows: (1) State space The state space refers to the set of environmental information that a deep reinforcement learning agent can observe, including runtime, user load, wind turbine output, photovoltaic output, grid output, gas turbine output, energy storage output, energy storage state of charge, and time-of-use pricing. The gas turbine and energy storage are controlled by two separate agents, and the environmental variables observed by these two agents are... They are not the same. Assuming the gas turbine is located near the grid interface, the gas turbine agent can only observe the operating time, grid output, the output of the l-th gas turbine, and the time-of-use electricity price. Similarly, if the energy storage device is located near the renewable energy generation equipment, the energy storage agent can only observe the photovoltaic output, wind turbine output, user load, the output of the l-th energy storage device, and the energy storage state of charge. Therefore, the state space can be described as:
[0059] In the formula: S MT S represents the state space of the gas turbine intelligent agent; ES This represents the state space of the energy storage intelligent agent.
[0060] (2) Action space During the interaction between industrial distributed power systems and the power grid, control actions include the output power of the gas turbine and the charging and discharging power of the energy storage. The action space of the intelligent agent can be defined as follows: A MT = P MT,DRL (t) A ES = P ES, DRL (t) In the formula: A MT A represents the action space of a gas turbine intelligent agent; ES This represents the action space of the energy storage intelligent agent.
[0061] (3) Reward function After an agent selects any action, the environment rewards it. However, if the agent's chosen action causes the industrial production operation to exceed environmental constraints, the environment will impose a penalty. In this paper, the environmental constraint penalty is derived from the state of charge constraint of energy storage, and the penalty expression is as follows:
[0062] In the formula: λ represents the penalty for exceeding the constraint; λ represents the penalty coefficient for environmental constraint penalties.
[0063] For energy storage, we want the state of charge (SOC) of the stored energy to be as equal as possible to its initial state at the last moment of dispatch. Therefore, we set an exponential reset penalty for energy storage: ; In the formula: The reset penalty for energy storage is represented by ε; the reset penalty coefficient is represented by t; and the reset penalty exponential coefficient is represented by t. In a scheduling cycle, the reset penalty for energy storage is small at the beginning, but it increases over time, reaching its maximum at the end of the cycle.
[0064] In summary, this paper primarily focuses on economic efficiency, considering how to minimize the operating cost of industrial production within the scheduling period T by rationally controlling the output of controllable industrial distributed power sources and energy storage devices. Therefore, the total reward function of the deep reinforcement learning agent for the interaction control of the industrial distributed power source system and the power grid described in this paper can be expressed in the following form:
[0065] In the formula: R represents the total reward of the deep reinforcement learning agent for the interaction control of industrial distributed power systems and power grids.
[0066] Step 5: Calculate the optimal solution for the interaction between the industrial distributed power system and the power grid using the TD3 deep reinforcement learning algorithm.
[0067] The TD3 algorithm is an improvement on the DDPG algorithm, primarily addressing the overestimation problem inherent in the DDPG algorithm. The convergence speed of the DDPG algorithm is improved. The TD3 algorithm contains 6 neural networks. First, a multivariate array (S, a, r, S') is obtained by random sampling from the experience replay pool. The actor network generates the action signal a using the state S, and the critic network generates the evaluation Q(S, a) of action a using the state S. The actor network updates the neural network parameters θ using gradient ascent with the goal of maximizing Q(S, a). To solve the overestimation problem that is prone to occur during the iterative update process, TD3 adopts a target network and double Q learning method to improve the above process. The target actor network generates action a' using the state S', and the target critic1 network and target critic2 network generate the evaluations Q1(S', a') and Q2(S', a') of action a' using the state S', respectively. Loss functions y1 and y2 can be calculated using Q(S, a), Q1(S', a'), and Q2(S', a'), and the expression for the loss function yi is shown in the following equation. The minimum value of y1 and y2 is taken, and the parameters w of the critic neural network are updated using gradient descent. Finally, the network parameters of actor, critic1, and critic2 are used to perform a soft update on the corresponding target network parameters, as shown in the following equation.
[0068]
[0069] θ T (t + 1) = θ(t) + (1-t)θ T (t) w T (t + 1) = w(t) + (1-t)w T (t) In the formula: θ T The neural network parameters representing the target actor network; w T represents the neural network parameters of the target critic network; t is the target smoothness factor, representing the update speed of the target neural network parameters.
[0070] In addition to the improvements mentioned above, the TD3 algorithm also adds noise signals following a truncated normal distribution to the output action signals of the target actor network. The truncated normal distribution can be denoted as CN(0, σ). 2 , - z, z ) indicates that the variable follows a mean of 0 and a variance of σ. 2 The value network follows a normal distribution, but the probability of the variable falling outside [-z, z] is 0. Using this probability distribution can prevent excessive noise. The TD3 algorithm also adopts a strategy of delayed updates to the policy network and target networks, that is, the value network is updated once in each round, but the policy network and the three target networks are updated once every β rounds.
[0071] The TD3 deep reinforcement learning algorithm consists of 6 neural networks, and the relationships between these 6 neural networks are shown in Figure 3. The training process of the TD3 deep reinforcement learning algorithm is as follows: Step 5.1: The TD3 deep reinforcement learning agent interacts with the environment and stores the observed information in the experience pool in the form of quadruplets (S, a, r, S'); Step 5.2: Have the target network make predictions: Each element in the vector ξ is independently derived from the truncated normal distribution CN(0, σ). 2 Extract from , -c, c).
[0072] Step 5.3: Have the two target networks make predictions:
[0073] Step 5.4: Calculate the TD target:
[0074] Step 5.5: Have the two value networks make predictions:
[0075] Step 5.6: Calculate the TD error:
[0076] Step 5.7: Update the value network:
[0077] Step 5.8: Update the policy network and the three target networks every k rounds:
[0078] Step 6: Train the distributed deep reinforcement learning agent using a distributed deep reinforcement learning architecture that employs "centralized training and distributed execution".
[0079] Because the control rules for gas turbines and energy storage differ, the power output of a gas turbine at time t0 has little correlation with the power output at the next time t1. However, due to the limitations of State of Charge (SOC), as shown in the equation, the power absorbed or emitted by energy storage at time t0 directly affects the power absorbed or emitted by energy storage thereafter. Therefore, when using deep reinforcement learning algorithms, gas turbines and energy storage should not use the same neural networks and algorithm parameters. Using two separate agents to control the distributed gas turbine and distributed energy storage will significantly improve the computational speed of deep reinforcement learning algorithms.
[0080] Centralized training and distributed execution in multi-agent reinforcement learning do not require communication during decision-making; only the agent needs to make decisions. The advantage of enabling real-time decision-making based on ground observation information has been well-received. This paper extends the TD3 algorithm to the MATD3 algorithm, which is characterized by the centralized collection of all agent observations and actions into an experience pool s = (o1, o2, ...o2) during the training process. N ), a = (a1, a2, ...a N Each agent's policy network obtains local observations from the environment. i Make an action a i The value network updates its parameters based on the observations *s* and actions *a* of all agents obtained from the experience pool, and generates actions *a* to issue to the policy network. i Evaluation q (s,a, w) i Finally, the policy network evaluates q(s, a, w). iThe agent updates its own neural network parameters θ. After the deep reinforcement learning agent has been trained, it can make distributed decisions based solely on local observations.
[0081] The distributed deep reinforcement learning structure of "centralized training and distributed decision-making" is shown in Figure 4. The specific steps are as follows: Step 6.1: During training, collect all observations and actions of the deep reinforcement learning agents into the experience pool, as shown in the figure: s = (o1, o2, ..., o N ), a = (a1, a2, ..., a N ) ; Step 6.2: The policy network of each deep reinforcement learning agent is based on local observations obtained from the environment. i Make an action a i ; Step 6.3: The value network adjusts the network parameters w based on the observations s and actions a of all deep reinforcement learning agents obtained from the experience pool. i Update; Step 6.4: The value network generates an action a to the policy network. i Evaluation q(s, a, w) i ) ; Step 6.5: The policy network evaluates q(s, a, w) i Update its own neural network parameters θ.
[0082] Step 7: Implement interactive control between the industrial distributed power system and the power grid using distributed decision-making.
[0083] This invention designs a control method for the interaction between a distributed power supply system (DPS) with industrial loads and the power grid. The method employs a two-layer control architecture for the interaction between the DPS and the power grid. The lower-layer control utilizes a distributed control approach based on multi-agent consensus. Compared to other control methods, this method does not rely on a central controller; each DPS communicates with its neighboring DPS to control all of them. The upper-layer control employs an optimization scheduling method based on multi-agent deep reinforcement learning. Each deep reinforcement learning agent controls different types of DPS and can only observe a portion of the environmental information during the production process. It calculates the optimal strategy for DPS operation at different time scales with the goal of minimizing operating costs and provides reference signals for the lower-layer control. Compared to other control methods, this method has greater applicability and superiority in handling the interaction control problem between DPS and the power grid, which involves high uncertainty, high real-time requirements, and a large parameter space. This project solves the problem of economic operation and optimal scheduling in industrial production processes and is of great significance for power grid planning and safe and stable operation.
[0084] Example 2 This invention also provides an apparatus for implementing the interactive control method between a distributed power supply system with industrial load and the power grid, comprising a processor and a memory; The processor is configured to execute the interactive control method between the distributed power supply system containing industrial loads and the power grid. The memory is used to store the executable instructions of the processor.
[0085] Example 3 As a third embodiment of the present invention, a computer storage medium is also provided, on which a computer program is stored, the computer program being executed by a processor to implement the aforementioned interactive control method between a distributed power supply system containing industrial loads and the power grid.
[0086] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0087] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0088] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0089] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 indivual The steps of the function specified in one or more boxes.
[0090] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0091] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for interactive control between a distributed power supply system containing industrial loads and the power grid, characterized in that, The control method includes lower-level control and upper-level control; The lower-level control targets industrial distributed power sources of the same type, taking into account the weak communication connections between industrial distributed power sources. It realizes the automatic allocation of the output power of industrial distributed power sources in the industrial production process through a multi-agent consensus distributed control method, and realizes the plug-and-play functionality of industrial distributed power sources. For different types of industrial distributed power sources, the upper-level control adopts a distributed deep reinforcement learning method to optimize and schedule the output of controllable industrial distributed power source equipment of various energy forms in industrial production, with the goal of minimizing the operating cost of industrial production while ensuring the safe operation of industrial production. The control method includes the following steps: Step 1: Establish the operation model and cost model of the distributed power system in the industrial production process. The distributed power system in the industrial production process is the industrial distributed power system, which includes photovoltaic, wind power, distributed gas turbines and distributed energy storage. Step 2: Establish communication between adjacent distributed gas turbines and adjacent distributed energy storage, so that adjacent gas turbines can exchange power information and adjacent energy storage can exchange power and state of charge information. Step 3: Use a multi-agent consensus algorithm to make the output power of the distributed gas turbine and the distributed energy storage consistent with the state of charge of the distributed energy storage. Step 4: Design the state space, action space, and reward function for the interaction between the industrial distributed power supply system and the power grid, and establish a Markov model for the interaction control between the industrial distributed power supply system and the power grid. Step 5: Use the TD3 algorithm to establish a deep reinforcement learning agent for distributed gas turbines and distributed energy storage; Step 6: Train the deep reinforcement learning agent offline using a centralized training method; After the deep reinforcement learning agent has been trained, each deep reinforcement learning agent can make distributed decisions based solely on local observations. Step 7: Implement interactive control between the industrial distributed power system and the power grid using distributed decision-making.
2. The interactive control method between a distributed power supply system containing industrial loads and the power grid according to claim 1, characterized in that, The specific steps of the multi-agent consensus algorithm are as follows: Step 1: Adjacent distributed gas turbines and adjacent distributed energy storage systems communicate with each other to exchange power information and state of charge information; Step 2: The upper-level control sends control signals to a distributed gas turbine and a distributed energy storage device in the industrial production process. The gas turbine and energy storage device update their output power according to the control signals. Step 3: Other distributed gas turbines update their output power based on the power information of adjacent gas turbines and themselves; other distributed energy storage updates their output power based on the power and state of charge information of adjacent energy storage devices and themselves.
3. The interactive control method between a distributed power supply system containing industrial loads and the power grid according to claim 1, characterized in that, Based on the operating costs of industrial distributed power sources and distributed energy storage, industrial load electricity prices, external grid time-of-use electricity prices, and the operating modes of various equipment in the industrial production process, the state space, action space, and reward function of the deep reinforcement learning agent for the interaction between the industrial distributed power source system and the power grid are designed, and a Markov model for the interactive control of the industrial distributed power source system and the power grid is established accordingly.
4. The interactive control method between a distributed power supply system containing industrial loads and the power grid according to claim 1, characterized in that, The distributed deep reinforcement learning method described above employs the TD3 deep reinforcement learning algorithm, which improves the DDPG training process using the following scheme: First, a truncated double-Q learning method is used to set up two value networks and two target value networks to alleviate the overestimation problem that occurs during the training of value networks; Second, noise is added to the target policy network to smooth it out. Third, reduce the update frequency of the policy network and the three target networks; that is, update the value network only once per round, but every... The policy network and three target networks are updated once per round.
5. The interactive control method between a distributed power supply system containing industrial loads and the power grid according to claim 1, characterized in that, Two different deep reinforcement learning agents are used to control the distributed gas turbine and distributed energy storage respectively. A distributed deep reinforcement learning architecture with centralized training and distributed decision-making is adopted. Each deep reinforcement learning agent only needs to observe the operating information of the equipment in its vicinity to make control decisions, which reduces the communication cost in the industrial production process and improves the computation speed of the deep reinforcement learning algorithm.
6. The interactive control method between a distributed power supply system containing industrial loads and the power grid according to claim 5, characterized in that, In the centralized training and distributed execution distributed deep reinforcement learning architecture: during training, the observations and actions of all deep reinforcement learning agents are centrally collected into an experience pool, and the policy network of each deep reinforcement learning agent only relies on local observations obtained from the environment. Make an action The value network, on the other hand, is based on observations of all agents obtained from the experience pool. and amount of movement Update the parameters of the value neural network and generate actions for the policy network. Evaluation Finally, the strategy network is evaluated. Update its own neural network parameters .
7. An apparatus for implementing the interactive control method between a distributed power supply system containing industrial loads and the power grid as described in any one of claims 1 to 6, characterized in that, Including the processor and memory; The processor is configured to execute the grid interaction control method for a distributed power supply system with industrial load as described in any one of claims 1 to 6. The memory is used to store the executable instructions of the processor.
8. A computer storage medium, characterized in that, It stores a computer program, which is executed by a processor to implement the method for interactive control between a distributed power supply system with industrial load as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-agent-based micro power supply decentralized coordination control method
CN106877398A
Power distribution network optimization method based on multi-agent deep reinforcement learning
CN114725936A