Cooperative control method and device among multiple energy systems
By building a physical simulation model of multi-energy systems and building security constraints, and using reinforcement learning agents for training, the problem of reasonable control between different energy systems is solved, and the safe operation and stability of multi-energy systems are achieved.
Patent Information
- Application Number
- CN202510088813.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-17
AI Technical Summary
The existing technology has not yet proposed an effective solution to how to achieve reasonable control between different energy systems and ensure the safe operation and stability of multi-energy systems.
By building a physical simulation model of multi-energy systems, security constraints and prior constraints are built, and training is performed using reinforcement learning agents to ensure that collaborative control between multi-energy systems is performed on the basis of meeting security constraints.
The hard constraint construction of safety constraints and prior constraints is realized, the flexibility and initial effectiveness of multi-energy system control is improved, and the problem of reasonable control between different energy systems is solved.
Smart Images

Figure CN120163554A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of energy system management, and particularly to a collaborative control method and device between multiple energy systems. Background Art
[0002] Currently, the interconnections between different energy systems are continuously strengthening, and an integrated control strategy can further improve the overall efficiency and performance of these multi-energy and multi-carrier systems. Such an integrated control strategy usually has multi-objective functions oriented towards economy, environment, etc. To ensure that the multi-energy system reaches the optimal operating level, it is necessary to establish specific setpoints to maintain the expected goals on the basis of satisfying all system constraints. However, since the flexibility between systems is reflected in dynamic behavior and multiple systems need to be linked between consecutive time steps, it is necessary to optimize the strategy between multiple time steps, and it is also necessary to manage various uncertain factors, such as changes in demand, price, and weather, to maintain the stability of the multi-energy system.
[0003] Currently, for the problem of how to achieve reasonable control between different energy systems in the related art, no effective solution has been proposed. Summary of the Invention
[0004] Embodiments of the present application provide a collaborative control method and device between multiple energy systems to at least solve the problem of how to achieve reasonable control between different energy systems in the related art.
[0005] In a first aspect, embodiments of the present application provide a collaborative control method between multiple energy systems, the method comprising:
[0006] Build a physical simulation model of the multi-energy system;
[0007] Based on the physical simulation model, construct safety constraints for ensuring the safe operation of the system;
[0008] Construct prior constraints, where the prior constraints are used to constrain the derivation of the policy of the reinforcement learning agent of the multi-energy system;
[0009] Based on the prior constraints, train the reinforcement learning agent of the built multi-energy system;
[0010] On the basis of satisfying the safety constraints, perform collaborative control between the multi-energy systems through the trained reinforcement learning agent.
[0011] In some of these embodiments, training the reinforcement learning agent of the built multi-energy system based on the prior constraints includes:
[0012] Train the reinforcement learning agent of the multi - energy system through the dual - delay deep deterministic policy gradient algorithm;
[0013] During the training process, constrain the policy of the reinforcement learning agent based on the prior constraints.
[0014] In some embodiments, during the training process, constraining the policy of the reinforcement learning agent based on the prior constraints includes:
[0015] Detect whether the action a selected by the reinforcement learning agent in state s satisfies the prior constraints;
[0016] If it is satisfied, then the action a is a safe action a safe And execute it to obtain the next state s′, reward r, and completion signal d, forming an experience tuple (s, a, r, s′, d) of the policy of the reinforcement learning agent;
[0017] If it is not satisfied, then use the safe action a safe To replace the action a, generating a feasible experience tuple (s, a safe , r, s′, d) and an infeasible experience tuple (s, a, r - c, s′, d), where c is an additional negative reward.
[0018] In some embodiments, based on the physical simulation model, constructing safety constraints for ensuring the safe operation of the system includes:
[0019] Based on the physical simulation model, construct safety constraints for ensuring the safe operation of the system, where the safety constraints include safety constraints of natural gas boilers, safety constraints of heat pump devices, safety constraints of combined heat and power units, safety constraints of thermal energy storage systems, safety constraints of battery energy storage systems, and total heat balance constraints of the system.
[0020] In some embodiments, based on the physical simulation model, constructing the total heat balance constraint of the system for ensuring the safe operation of the system includes:
[0021] Construct the total heat balance constraint of the system for ensuring the safe operation of the system:
[0022]
[0023] Wherein, is the total output heat of the multi - energy system, is the total required heat of the multi - energy system.
[0024] In some embodiments, after constructing the total heat balance constraint of the system for ensuring the safe operation of the system, the method includes:
[0025] Convert the overall heat balance constraint of the system into an inequality representation:
[0026]
[0027] Among them, is the heat of the natural gas boiler, is the heat of the heat pump device, is the heat of the combined heat and power unit, is the heat of the thermal energy storage system, is the total required heat of the multi - energy system, Q tol is the preset heat threshold.
[0028] In some embodiments, based on the physical simulation model, the safety constraints for ensuring the safe operation of the system, including the safety constraints of the natural gas boiler, the heat pump device, the combined heat and power unit, the thermal energy storage system, and the battery energy storage system, are constructed as follows:
[0029] Construct the safety constraints of the natural gas boiler for ensuring the safe operation of the system:
[0030]
[0031] Among them, is the heat of the natural gas boiler, is the preset maximum heat of the natural gas boiler, is the preset minimum heat of the natural gas boiler, is the binary control variable;
[0032] Construct the safety constraints of the heat pump device for ensuring the safe operation of the system:
[0033]
[0034] Among them, is the heat of the heat pump device, is the electric energy of the heat pump device, is the preset maximum heat of the heat pump device, is the preset minimum heat of the heat pump device, is the preset maximum electric energy of the heat pump device, is the preset minimum electric energy of the heat pump device, is the binary control variable;
[0035] Construct the safety constraints of the combined heat and power unit for ensuring the safe operation of the system:
[0036]
[0037] Among them, is the heat of the combined heat and power unit, is the electric energy of the combined heat and power unit, is the preset maximum heat of the combined heat and power unit, is the preset minimum heat of the combined heat and power unit, is the preset maximum electric energy of the combined heat and power unit, is the preset minimum electric energy of the combined heat and power unit, is a binary control variable;
[0038] Construct the safety constraints of the thermal energy storage system for ensuring the safe operation of the system:
[0039]
[0040] Among them, is the heat of the thermal energy storage system, is the preset maximum heat of the thermal energy storage system, is the preset minimum heat of the thermal energy storage system;
[0041] Construct the safety constraints of the battery energy storage system for ensuring the safe operation of the system:
[0042]
[0043] Among them, is the electric energy of the battery energy storage system, is the preset maximum electric energy of the battery energy storage system, is the preset minimum electric energy of the battery energy storage system.
[0044] In some embodiments, before training the reinforcement learning agent of the constructed multi - energy system based on the prior constraints, the method includes:
[0045] Construct the reinforcement learning agent of the multi - energy system, where the reinforcement learning agent includes a state space S, an action space A, and a reward function R a .
[0046] In some embodiments, constructing the physical simulation model of the multi - energy system includes:
[0047] Construct the physical simulation model of the multi - energy system, where the multi - energy system includes a transformer, a wind turbine generator set, a photovoltaic device, a natural gas boiler, a heat pump device, a combined heat and power unit, a thermal energy storage system, and a battery energy storage system.
[0048] In a second aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect above is implemented.
[0049] Compared with the related art, an embodiment of the present application provides a collaborative control method and device between multiple energy systems. The method includes building a physical simulation model of the multiple energy systems; based on the physical simulation model, constructing safety constraints for ensuring the safe operation of the system; constructing prior constraints, where the prior constraints are used to constrain the derivation of the policy of the reinforcement learning agent of the multiple energy systems; based on the prior constraints, training the reinforcement learning agent of the built multiple energy systems; on the basis of satisfying the safety constraints, performing collaborative control between the multiple energy systems through the trained reinforcement learning agent, realizing the construction of hard constraints of safety constraints and prior constraints, rather than the guarantee of soft constraints and accidental constraints, having a higher initial utility, decoupling the safety constraints from the training process of the reinforcement learning agent, improving the flexibility of the control of the multiple energy systems, and solving the problem of how to achieve reasonable control between different energy systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0051] Figure 1 is a flowchart of the steps of the collaborative control method between multiple energy systems according to an embodiment of the present application;
[0052] Figure 2 is an internal structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts fall within the scope of protection of the present application.
[0054] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0055] In the present application, the mention of "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0056] Unless otherwise defined, the technical terms or scientific terms involved in the present application should have the ordinary meaning understood by those of ordinary skill in the technical field to which the present application belongs. The words "a", "an", "one kind", "the" and other similar words involved in the present application do not represent a limitation in quantity and can represent singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and other similar words involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0057] The embodiment of the present application provides a collaborative control method between multiple energy systems. Figure 1 It is a flowchart of the steps of the collaborative control method between multiple energy systems according to the embodiment of the present application, as Figure 1As shown, the method includes the following steps:
[0058] Step S102, build a physical simulation model of the multi - energy system;
[0059] Specifically, in step S102, build a physical simulation model of the multi - energy system, where the multi - energy system includes a transformer, a wind turbine generator set, a photovoltaic device, a natural gas boiler, a heat pump device, a combined heat and power unit, a heat energy storage system, and a battery energy storage system.
[0060] It should be noted that the multi - energy system does not include a gas turbine, that is, a standby generator set. Although the physical simulation model is a detailed differential - algebraic equation system (with an equal number of variables), it does not include any virtual control modules (such as a PID controller). In other words, in this embodiment, the control system is abstracted, and the actions executed in the simulation environment will remain unchanged within the considered control range (15 minutes), and there is no continuous unconstrained error processing, that is, the minimization of the difference between the desired set point and the measured process variable.
[0061] Step S104, based on the physical simulation model, construct safety constraints for ensuring the safe operation of the system;
[0062] Specifically, in step S104, based on the physical simulation model, construct safety constraints for ensuring the safe operation of the system, where the safety constraints include the safety constraints of the natural gas boiler, the safety constraints of the heat pump device, the safety constraints of the combined heat and power unit, the safety constraints of the heat energy storage system, the safety constraints of the battery energy storage system, and the overall heat balance constraint of the system.
[0063] Preferably in step S104:
[0064] ①Construct the overall heat balance constraint of the system for ensuring the safe operation of the system:
[0065]
[0066] Wherein, is the total output heat of the multi - energy system, is the total required heat of the multi - energy system.
[0067] Furthermore, convert the overall heat balance constraint of the system into an inequality representation:
[0068]
[0069] Wherein, is the heat of the natural gas boiler, is the heat of the heat pump device, is the heat of the combined heat and power unit, is the heat of the heat energy storage system, The total heat demand for the multi - energy system, Q tol is the preset heat threshold.
[0070] Furthermore, assuming that the grid connection is large enough, although the power balance constraint is always satisfied, special attention needs to be paid to the heat balance constraint. In this embodiment, additional constraints such as ramp rate, minimum operation, and downtime are not considered because the energy balance equation is considered the most restrictive constraint in the energy management problem. To write the heat balance in more detail, the equality constraint formula is relaxed to an inequality constraint, that is, this inequality can also be converted to:
[0071]
[0072] where A t is the action, η is the energy efficiency, COP is the coefficient of performance, SOC is the state of charge, and T is various specific temperatures, that is is the return water temperature of the boiler, is the evaporation equipment temperature of the heat pump, is the condenser temperature of the heat pump, is the ambient air temperature, is the average temperature of the stratified thermal storage water tank.
[0073] It should be noted that to maintain a low computational complexity, Q tol can be set to 15.0% of
[0074] ②Construct the safety constraint of the natural gas boiler for ensuring the safe operation of the system:
[0075]
[0076] where is the heat of the natural gas boiler, is the preset maximum heat of the natural gas boiler, is the preset minimum heat of the natural gas boiler, is the binary control variable, that is its value determines the opening or closing of this safety constraint;
[0077] ③Construct the safety constraint of the heat pump device for ensuring the safe operation of the system:
[0078]
[0079] where is the heat of the heat pump device, is the electric energy of the heat pump device, is the preset maximum heat of the heat pump device, is the preset minimum heat of the heat pump device, is the preset maximum electric energy of the heat pump device, is the preset minimum electric energy of the heat pump device, is a binary control variable, that is, its value determines the opening or closing of this safety constraint;
[0080] ④Construct the safety constraints of the cogeneration unit for ensuring the safe operation of the system:
[0081]
[0082] Among them, is the heat of the cogeneration unit, is the electric energy of the cogeneration unit, is the preset maximum heat of the cogeneration unit, is the preset minimum heat of the cogeneration unit, is the preset maximum electric energy of the cogeneration unit, is the preset minimum electric energy of the cogeneration unit, γ c t hp is a binary control variable, that is, its value determines the opening or closing of this safety constraint;
[0083] ⑤Construct the safety constraints of the thermal energy storage system for ensuring the safe operation of the system:
[0084]
[0085] Among them, is the heat of the thermal energy storage system, is the preset maximum heat of the thermal energy storage system, is the preset minimum heat of the thermal energy storage system;
[0086] ⑥Construct the safety constraints of the battery energy storage system for ensuring the safe operation of the system:
[0087]
[0088] Among them, is the electric energy of the battery energy storage system, is the preset maximum electric energy of the battery energy storage system, is the preset minimum electric energy of the battery energy storage system.
[0089] Step S106, construct prior constraints, where the prior constraints are used to constrain the policy of the reinforcement learning agent of the multi - energy system to be obtained;
[0090] Specifically, in step S106, the reinforcement learning agent of the multi - energy system follows the standard reinforcement learning formula of the state - value function and extends it to the constraint conditions. The goal is to find a policy π, which is a mapping from the state (s ∈ S) to the action (a ∈ A(s)), that maximizes the expected sum of discounted rewards and is subject to the constraint sets X and U (i.e., prior constraints):
[0091]
[0092] where, E π is the expected value of the infinite sum of the discounted factor γ times the reward R minus the cost c at each time step after taking the policy π. The first equation above is a discrete - time - varying infinite - horizon stochastic optimal control problem, but due to the constraint sets X and U (i.e., prior constraints) being different from the standard formulation of reinforcement learning. In this embodiment, the prior constraints can be decoupled from the reinforcement learning agent so that any new reinforcement learning algorithm can be used while always ensuring that the hard constraints are satisfied.
[0093] After step S106, the method further includes step S107 of building a reinforcement learning agent for the multi - energy system, where the reinforcement learning agent includes a state space S, an action space A, and a reward function R a .
[0094] Specifically, in step S107, the state space, action space, and reward function in the safe energy management decision - making process of the multi - energy system are defined to build a Markov decision process; the Markov decision process under fully observable discrete time is represented as <S, A, P a , R a >, specifically:
[0095]
[0096] where, is the heat demand, is the electricity demand, is the wind power input, is the solar power input, is the electricity price signal (i.e., the day - ahead spot price), is the state - of - charge (SOC) of the thermal energy storage system, is the state - of - charge of the battery energy storage system. h t is the hour of the day, d t is the day of the week. These observable variables constitute the state space S. The action space A includes the action from the natural gas boiler the action from the heat pump the action from the combined heat and power unit Actions from the thermal energy storage system Actions from the battery energy storage system
[0097] The goal of the reinforcement learning agent is given by the reward function R a which is composed of the energy cost loss at the t-th time step Comfort loss scalarized weights a and b, and an additional cost c after violating the expected constraints. The goal maximizes the negative function, thus minimizing the positive function. The comfort loss is defined as where Q t is the thermal energy output. In the Modelica model, assuming that the grid connection is large enough, the electricity demand and natural gas consumption can always be met by purchasing from the respective infinitely large main grids. The comfort loss is limited by a tolerance. Therefore, this term in the reward function can serve as a fine-tuning mechanism to further reduce the comfort loss within this bound, guiding the reinforcement learning agent to take safer actions and reducing the modeling error of the constraints themselves.
[0098] Step S108, training the reinforcement learning agent of the constructed multi-energy system based on prior constraints;
[0099] Step S108 specifically includes the following steps:
[0100] Step S1081, training the reinforcement learning agent of the multi-energy system through the twin-delayed deep deterministic policy gradient algorithm;
[0101] It should be noted that in this embodiment, the twin-delayed deep deterministic policy gradient algorithm (TD3) is used, which is one of the most advanced reinforcement learning algorithms currently. The reinforcement learning agents themselves do not need to use a prior dataset for "offline" pre-training, and they can be safely trained "online". Therefore, the reinforcement learning agents do not need any prior knowledge about the environment at the beginning because their dataset (experience tuples) is generated when interacting with the environment (multi-energy management system). The pseudo-code implementation of the twin-delayed deep deterministic policy gradient algorithm is as follows:
[0102]
[0103]
[0104] Step S1082, during the training process, constraining the policy of the reinforcement learning agent based on prior constraints.
[0105] Specifically, in step S1082, it is detected whether the action a selected by the reinforcement learning agent in state s satisfies the prior constraint;
[0106] If it is satisfied, then the action a is a safe action a safe And execute it to obtain the next state s′, reward r, and completion signal d, forming an experience tuple (s, a, r, s′, d) of the policy of the reinforcement learning agent;
[0107] If it is not satisfied, then use the safe action a safe To replace the action a, generating a feasible experience tuple (s, a safe , r, s′, d) and an infeasible experience tuple (s, a, r - c, s′, d), where c is an additional negative reward.
[0108] It should be noted that step S1082 is a safety fallback policy π proposed in this embodiment that depends on prior constraints safe , which can usually be derived in the form of a set of hard - coded rules through classical control theory, such as a priority - based energy management policy. Such simple rule - based policies are usually easy to construct. When the constraint function (prior constraint) is given, it can simply be checked whether the action a selected at state s satisfies the constraint condition.
[0109] When the constraint condition is satisfied, the selected action a is considered a safe action a safe And execute it in the current environment s, then obtain the next state s′, reward r, and completion signal d, forming a regular experience tuple (s, a, r, s′, d) of the reinforcement learning agent; however, if the constraint condition is violated, then use the safe action a of the prior safety fallback policy safe To replace the selected action a, not only forming a feasible experience tuple (s, a safe , r, s′, d), but also forming an infeasible experience tuple (s, a, r - c, s′, d) with c being an additional negative reward (i.e., cost).
[0110] The pseudo - code implementation of the safety fallback policy π safe is as follows:
[0111]
[0112]
[0113] In addition to the safety fallback policy π that depends on prior constraints proposed in step S1082 safe , in some embodiments, a safety - given policy that does not depend on prior constraints is also proposed, which depends on the reinforcement learning agent itself to deliver the safe action a safe, which can be checked again through the given constraints. If the action a selected in state s passes the constraint check, then the safe action a is executed in the environment safe , and receive the regular experience tuple (s, a safe , r, s′, d). However, if the constraint is violated, since the infeasible action is not executed, the transition to the next state cannot be observed, and an additional negative reward c is given. The reinforcement learning agent will receive the experience tuple (s, a, c, s, d). The reinforcement learning agent will then select a new action a and check if the constraint is satisfied, and this process will repeat until the constraint check passes and the selected action is considered the safe action a safe . Then this safe action is executed in the environment, and the regular experience tuple is received.
[0114] The pseudo-code implementation of the safe given policy is as follows:
[0115]
[0116]
[0117] Step S110, based on satisfying the safety constraints, perform the coordinated control between multiple energy systems through the trained reinforcement learning agent.
[0118] It should be noted that before using the reinforcement learning agent to perform coordinated control, it also needs to be evaluated. Specifically, on the premise of satisfying the thermal comfort constraints, with the goal of minimizing the energy cost, compare the performance of the safety fallback algorithm and the safety given algorithm in the proposed method with the unconstrained reinforcement learning agent, the safe and unsafe random agents. The random agents are the lowest learning benchmarks for the relevant algorithms, that is, how much the agents learn when interacting safely with their environment, and also the performance of the algorithms before any training. Therefore, these random agents are not intended to be the best feasible controllers. This study uses a one-year training environment, a 15-minute control horizon (i.e., the energy management reinforcement learning agent can select a new action every 15 minutes), and a one-week evaluation environment, and participates in the day-ahead electricity market at the same time. Any uncertainties, such as those from demand, price, or wind and solar power generation, can be handled by the reinforcement learning agent because it is formulated as a discrete time-varying infinite horizon stochastic optimal control problem.
[0119] Through the above steps in the embodiments of the present application, the construction of hard constraints for safety constraints and prior constraints is realized, rather than the guarantee of soft constraints and accidental constraints, with higher initial utility. Decouple the safety constraints from the training process of the reinforcement learning agent, improve the flexibility of multi-energy system control, and solve the problem of how to achieve reasonable control between different energy systems.
[0120] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0121] This embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.
[0122] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0123] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.
[0124] In addition, in combination with the collaborative control method between multi-energy systems in the above embodiments, an embodiment of the present application can provide a storage medium to implement. A computer program is stored on the storage medium; when the computer program is executed by a processor, it implements any one of the collaborative control methods between multi-energy systems in the above embodiments.
[0125] In one embodiment, a computer device is provided. The computer device can be a terminal. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a collaborative control method between multi-energy systems. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0126] In one embodiment, Figure 2 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. As Figure 2 shown, an electronic device is provided. The electronic device can be a server, and its internal structure diagram can be as Figure 2As shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected by an internal bus. Among them, the non-volatile memory stores an operating system, computer programs, and a database. The processor is used to provide computing and control capabilities. The network interface is used to communicate with external terminals through a network connection. The internal memory is used to provide an environment for the operation of the operating system and computer programs. The computer program, when executed by the processor, implements a collaborative control method between multiple energy systems. The database is used to store data.
[0127] Those skilled in the art can understand that Figure 2 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0128] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0129] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0130] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A collaborative control method between multiple energy systems, characterized in that: The method comprises: Build a physical simulation model of a multi-energy system; Based on the physical simulation model, construct safety constraints for ensuring safe operation of the system; Constructing a priori constraints, wherein the a priori constraints are used to constrain the strategy derivation of the reinforcement learning agent of the multi-energy system; Based on the prior constraints, training the reinforcement learning agent of the constructed multi-energy system; On the basis of satisfying the safety constraints, the coordinated control among the multi-energy systems is performed through the trained reinforcement learning agent.
2. The method according to claim 1, characterized in that Based on the prior constraints, training the reinforcement learning agent of the constructed multi-energy system includes: Training a reinforcement learning agent for the multi-energy system using a double-delayed deep deterministic policy gradient algorithm; During the training process, the strategy of the reinforcement learning agent is constrained based on the prior constraints.
3. The method according to claim 2, characterized in that During the training process, constraining the strategy of the reinforcement learning agent based on the prior constraints includes: Detecting whether the action a selected by the reinforcement learning agent in state s satisfies the prior constraint; If satisfied, then the action a is a safe action a safe And execute, get the next state s′, reward r and completion signal d, and form the experience tuple (s, a, r, s′, d) of the reinforcement learning agent’s strategy; If not satisfied, use safety action a safe Substituting the action a, a feasible experience tuple (s, a) of the strategy of the reinforcement learning agent is generated safe ,r,s′,d) and the infeasible experience tuple (s,a,rc,s′,d), where c is the additional negative reward.
4. The method according to claim 1, characterized in that: Based on the physical simulation model, safety constraints for ensuring safe operation of the system are constructed, including: Based on the physical simulation model, safety constraints are constructed to ensure the safe operation of the system, wherein the safety constraints include safety constraints of natural gas boilers, safety constraints of heat pump devices, safety constraints of cogeneration units, safety constraints of thermal energy storage systems, safety constraints of battery energy storage systems, and overall heat balance constraints of the system.
5. The method according to claim 4, characterized in that Based on the physical simulation model, the overall heat balance constraints of the system for ensuring safe operation of the system include: Construct the overall thermal balance constraints of the system to ensure safe operation of the system: in, is the total heat output of the multi-energy system, is the total heat demand of the multi-energy system.
6. The method according to claim 5, characterized in that After constructing the overall heat balance constraint of the system for ensuring safe operation of the system, the method includes: The total heat balance constraint of the system is converted into an inequality expression: in, is the heat of the natural gas boiler, is the heat of the heat pump device, is the heat of the combined heat and power unit, is the heat of the thermal energy storage system, is the total heat demand of the multi-energy system, Q tol is the preset heat threshold.
7. The method according to claim 4, characterized in that Based on the physical simulation model, the safety constraints of the natural gas boiler, the safety constraints of the heat pump device, the safety constraints of the cogeneration unit, the safety constraints of the thermal energy storage system, and the safety constraints of the battery energy storage system are constructed to ensure the safe operation of the system, including: Construct safety constraints for natural gas boilers to ensure safe system operation: in, is the heat of the natural gas boiler, The preset maximum heat of the natural gas boiler, The preset minimum heat for the natural gas boiler. is a binary control variable; Construct safety constraints for the heat pump plant to ensure safe operation of the system: in, is the heat of the heat pump device, is the electrical energy of the heat pump device, The preset maximum heat of the heat pump device. The preset minimum heat of the heat pump device, is the preset maximum electrical energy of the heat pump device, is the preset minimum electrical energy for the heat pump unit, is a binary control variable; Construct the safety constraints of the CHP unit to ensure safe operation of the system: in, is the heat of the combined heat and power unit, is the electricity of the combined heat and power unit, The preset maximum heat of the combined heat and power unit, The preset minimum heat for the CHP unit, is the preset maximum power of the cogeneration unit, The preset minimum power for the CHP unit, is a binary control variable; Construct safety constraints for thermal energy storage systems to ensure safe operation of the system: in, is the heat of the thermal energy storage system, The preset maximum heat of the thermal energy storage system, A preset minimum heat for the thermal energy storage system; Construct safety constraints for the battery energy storage system to ensure safe operation of the system: in, is the electric energy of the battery energy storage system, is the preset maximum power of the battery energy storage system. The preset minimum power of the battery energy storage system.
8. The method according to claim 1, characterized in that Before training the reinforcement learning agent of the constructed multi-energy system based on the prior constraints, the method includes: Build a reinforcement learning agent for the multi-energy system, wherein the reinforcement learning agent includes a state space S, an action space A and a reward function R a .
9. The method according to claim 1, characterized in that: Building a physical simulation model for a multi-energy system includes: Build a physical simulation model of a multi-energy system, wherein the multi-energy system includes a transformer, a wind turbine, a photovoltaic device, a natural gas boiler, a heat pump device, a cogeneration unit, a thermal energy storage system and a battery energy storage system.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 9.
Citation Information
Cited By
Heating furnace air-fuel ratio self-optimization control device and method
CN121430348A