Real-time scheduling method of high-proportion new energy power system and related equipment
By building a virtual simulation environment and a safety reinforcement learning algorithm in a high proportion of new energy power systems, combined with the safety constraints of the power system, the problem of lack of safety constraints in the traditional reinforcement learning framework is solved, and the security training and efficient scheduling of the agent are realized.
Patent Information
- Application Number
- CN202510091686.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The traditional reinforcement learning framework lacks clear security constraints in real-time scheduling of high proportion new energy power systems, resulting in the agents that may take behaviors that produce unsafe results.
By constructing a virtual simulation environment and a security reinforcement learning algorithm, combined with the safe operation constraints of the power system, it is transformed into a Markov decision-making process, and the state space, action space, reward function and punishment function of the agent are constructed, and iterative training is carried out to obtain the trained agent.
It effectively avoids the insecure behavior of the agent during the exploration process, improves the security, stability and robustness of the scheduling strategy, and achieves a good real-time scheduling strategy generation capability of a high proportion of new energy power system.
Smart Images

Figure CN120185084A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of scheduling, and particularly to a real-time scheduling method and related devices for a high-proportion new energy power system. Background Art
[0002] With the increasing large-scale grid connection and operation of clean energy clusters such as wind power and photovoltaic power, traditional reinforcement learning, as a data-driven learning method, has been widely used in the real-time scheduling of high-proportion new energy power systems. Through the interaction between the agent and the environment, it can autonomously learn from historical experience, search for the optimal control strategy, and maximize the cumulative reward. However, there are certain deficiencies in the traditional reinforcement learning framework in terms of security. The actions of the agent are not clearly safety-constrained, resulting in the agent possibly taking actions that produce unsafe results. Summary of the Invention
[0003] This application provides a real-time scheduling method and related devices for a high-proportion new energy power system to solve the problem that in the prior art, the actions of the agent are not clearly safety-constrained, resulting in the agent possibly taking actions that produce unsafe results.
[0004] To achieve the above object, an embodiment of this application provides a real-time scheduling method for a high-proportion new energy power system, including:
[0005] Construct a virtual simulation environment of the power system according to the component information and historical operation information of the power system;
[0006] Taking thermal power units, energy storage units, wind turbines, and photovoltaic units as regulation objects, and combining the safe operation constraint conditions of the power system, construct a power system scheduling model with the goal of minimizing the sum of the regulation cost and the curtailment penalty of wind and light of the power system; wherein, the safe operation constraint conditions include: power balance constraint, line transmission capacity constraint, thermal power unit operation constraint, energy storage unit operation constraint, and maximum curtailment of wind and light constraint;
[0007] Convert the power system scheduling model into a Markov decision process to construct the state space, action space, reward function, and penalty function of the agent;
[0008] Interact the agent with the virtual simulation environment, and based on the set constrained optimization goal of the agent, use a safe reinforcement learning algorithm to iteratively train the agent to obtain a trained agent;
[0009] Use the trained agent to perform real-time scheduling of the power system.
[0010] As an improvement of the above solution, the objective function of the power system scheduling model is expressed as:
[0011]
[0012] In the formula, and are decision variables, representing the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the curtailed power of the i-th wind turbine at time t, and the curtailed power of the i-th photovoltaic unit at time t, respectively; is the regulation cost of the power system; is the penalty for wind and photovoltaic curtailment of the power system;
[0013]
[0014] In the formula, N G is the number of thermal power units in the power system, is the unit regulation cost of the i-th thermal power unit, is the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, is the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t; N S is the number of energy storage units in the power system, is the unit regulation cost of the i-th energy storage unit, is the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, is the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t;
[0015]
[0016] In the formula, N Wind is the number of wind turbines in the power system, is the unit curtailment penalty of the i-th wind turbine, is the curtailed power of the i-th wind turbine at time t; N Solar is the number of photovoltaic units in the power system, is the unit curtailment penalty of the i-th photovoltaic unit, is the curtailed power of the i-th photovoltaic unit at time t.
[0017] As an improvement of the above solution, the power balance constraint is expressed as:
[0018]
[0019] Among them, N G is the number of thermal power units in the power system, is the real-time plan of the i-th thermal power unit at time t; N Wind is the number of wind turbines in the power system, is the real-time plan of the i-th wind turbine at time t; N Solar is the number of photovoltaic units in the power system, is the real-time plan of the i-th photovoltaic unit at time t; N S is the number of energy storage units in the power system, is the real-time plan of the i-th energy storage unit at time t; N Load is the number of loads in the power system, is the real-time demand of the i-th load at time t;
[0020] The line transmission capacity constraint is expressed as:
[0021]
[0022] where, N Bus is the number of buses in the power system; G i-j is the influence of the injection power per unit of the j-th bus on the transmission power of the i-th line; is the lower limit of the transmission capacity of the i-th line, is the upper limit of the transmission capacity of the i-th line; is the real-time plan of the j-th thermal power unit at time t, is the real-time plan of the j-th wind turbine at time t, is the real-time plan of the j-th photovoltaic unit at time t, is the real-time plan of the j-th energy storage unit at time t, is the real-time demand of the j-th load at time t;
[0023] The operation constraint of the thermal power unit is expressed as:
[0024]
[0025] where, is the minimum allowable output power of the i-th thermal power unit, is the maximum allowable output power of the i-th thermal power unit, is the real-time plan of the i-th thermal power unit at time t, is the maximum allowable ramp rate of the i-th thermal power unit;
[0026] The operation constraint of the energy storage unit is expressed as:
[0027]
[0028] where, is the minimum allowable output power of the i-th energy storage unit, is the real-time plan of the i-th energy storage unit, is the maximum allowable output power of the i-th energy storage unit; is the minimum allowable state of charge of the i-th energy storage unit, is the real-time state of charge of the i-th energy storage unit, is the maximum allowable state of charge of the i-th energy storage unit; is a binary variable for the charging state, is a binary variable for the discharging state, is the real-time plan of the i-th energy storage unit at time t, is the charge-discharge efficiency; ΔT is the scheduling time interval, is the capacity of the i-th energy storage unit;
[0029] The maximum wind and PV curtailment constraint is expressed as:
[0030]
[0031] where, is the curtailment of the i-th wind turbine, is the intra-day predicted power generation of the i-th wind turbine at time t; is the curtailment of the i-th PV unit at time t, is the intra-day predicted power generation of the i-th PV unit at time t.
[0032] As an improvement to the above solution, the state space includes: the current time, the active power of the thermal power unit at time t, the active power of the energy storage unit at time t, the active power of the wind turbine at time t, the active power of the PV unit at time t, the active power of the load at time t, the upper adjustable space of the thermal power unit at time t, the lower adjustable space of the thermal power unit at time t, the upper adjustable space of the energy storage unit at time t, the lower adjustable space of the energy storage unit at time t, the line load rate at time t, the day-ahead plan of the thermal power unit at time t+1, the day-ahead plan of the energy storage unit at time t+1, the intra-day predicted demand of the load at time t+1, the intra-day predicted power generation of the wind turbine at time t+1, the intra-day predicted power generation of the PV unit at time t+1;
[0033] The action space includes: the adjustment amount of the output of the thermal power unit in two adjacent time periods and the ratio of the maximum allowable output of the energy storage unit;
[0034] The reward function is expressed as:
[0035]
[0036] where, N Gis the number of thermal power units in the power system, is the unit regulation cost of the i-th thermal power unit, is the real-time plan of the i-th thermal power unit at time t, is the day-ahead plan of the i-th thermal power unit at time t; N S is the number of energy storage units in the power system, is the unit regulation cost of the i-th energy storage unit, is the real-time plan of the i-th energy storage unit at time t, is the day-ahead plan of the i-th energy storage unit at time t;
[0037] The penalty function is expressed as:
[0038]
[0039] where α rho is the weight factor for line overload penalty, α slack is the weight factor for balancing machine overload penalty; N Line is the number of lines in the power system; ρ i,t is the load factor of the i-th line at time t; is the output of the balancing machine at time t, is the upper limit of the output allowed for the balancing machine, is the lower limit of the output allowed for the balancing machine.
[0040] As an improvement to the above solution, the interaction between the agent and the virtual simulation environment, based on the constrained optimization objective of the set agent, and the use of the safe reinforcement learning algorithm to iteratively train the agent to obtain a trained agent includes:
[0041] Set the constrained optimization objective of the agent:
[0042]
[0043] In the formula, a represents the agent's action, π represents the agent's policy, s represents the environmental state observed by the agent, E a : π(s) represents taking the expectation of the agent's policy π, γ represents the discount factor, t represents the time, r t represents the immediate reward received by the agent, s0 represents the initial environmental state, c t represents the immediate penalty received by the agent, c represents the penalty threshold, represents the target entropy;
[0044] Convert the optimization objective into an unconstrained optimization problem, that is:
[0045]
[0046] In the formula, is the augmented objective function, f(π) = E a : π(s) [∑ t γ t r t |s0 = s], λ is the adaptive weight coefficient of the agent action safety term; β is the adaptive weight coefficient of the agent action random term; π represents the agent policy;
[0047] Construct a reward evaluation network and a punishment evaluation network; among them, the loss function of the reward evaluation network is expressed as: The loss function of the punishment evaluation network is expressed as:
[0048] In the formula, φ r is the parameter of the reward evaluation network, φ c is the parameter of the punishment evaluation network, τ is the interaction trajectory sampled from the experience pool, is the experience pool, represents taking the expectation on the sampled trajectory, Q r (s t ,a t ) is the reward value function evaluated by the reward evaluation network when the agent applies action a t under the current state s t ; Q c (s t ,a t ) is the punishment value function evaluated by the punishment evaluation network when the agent applies action a t under the current state s t ; r t represents the immediate reward received by the agent, c t represents the immediate punishment received by the agent, γ is the discount factor, is the output value of the target reward evaluation network, is the output value of the target punishment reward evaluation network;
[0049] Construct an action network; among them, the loss function of the action network is expressed as:
[0050] In the formula, φ π is the parameter of the action network, τ is the interaction trajectory sampled from the experience pool, is the experience pool, represents taking the expectation on the sampled trajectory, Q r (s t ,a t ) is the reward evaluation network evaluating the current state st The lower agent applies action a t The subsequent reward value function, where λ is the adaptive weight coefficient of the agent action safety term, β is the adaptive weight coefficient of the agent action random term, and π is the agent policy;
[0051] Interact the agent with the virtual simulation environment. During the interaction process, use the unconstrained optimization problem, the reward evaluation network, the penalty evaluation network, and the action network to iteratively train the agent to obtain a trained agent.
[0052] As an improvement to the above solution, constructing the virtual simulation environment of the power system according to the component information and historical operation information of the power system includes:
[0053] Obtain the original power grid data of the power system;
[0054] Parse and preprocess the original power grid data to obtain component information and historical operation information;
[0055] Construct a power flow calculation simulation model of the power system according to the component information;
[0056] Construct an operation scenario set of the power system according to the historical operation information;
[0057] Construct the virtual simulation environment of the power system according to the power flow calculation simulation model and the operation scenario set.
[0058] To achieve the above object, an embodiment of the present application also provides a real-time scheduling device for a high-proportion new energy power system, including:
[0059] A simulation environment construction module for constructing the virtual simulation environment of the power system according to the component information and historical operation information of the power system;
[0060] A model construction module for taking thermal power units, energy storage units, wind turbines, and photovoltaic units as regulation objects, and constructing a power system scheduling model with the goal of minimizing the sum of the regulation cost and the curtailment penalty of the power system in combination with the safe operation constraint conditions of the power system; where the safe operation constraint conditions include: power balance constraint, line transmission capacity constraint, thermal power unit operation constraint, energy storage unit operation constraint, and maximum curtailment constraint;
[0061] An agent construction module for converting the power system scheduling model into a Markov decision process to construct the state space, action space, reward function, and penalty function of the agent;
[0062] An agent training module for interacting the agent with the virtual simulation environment, and iteratively training the agent using a safe reinforcement learning algorithm based on a set constrained optimization objective of the agent to obtain a trained agent;
[0063] A real-time scheduling module for performing real-time scheduling of the power system using the trained agent.
[0064] To achieve the above object, an embodiment of the present application further provides a real-time scheduling device for a high-proportion new energy power system, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the real-time scheduling method for a high-proportion new energy power system as described above is implemented.
[0065] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the real-time scheduling method for a high-proportion new energy power system as described above.
[0066] To achieve the above object, an embodiment of the present application further provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the real-time scheduling method for a high-proportion new energy power system as described above.
[0067] Compared with the prior art, a real-time scheduling method, device, equipment, storage medium, and product for a high-proportion new energy power system provided by an embodiment of the present application construct a virtual simulation environment of the power system according to the component information and historical operation information of the power system; take thermal power units, energy storage units, wind turbines, and photovoltaic units as control objects, and combine the safe operation constraint conditions of the power system to construct a power system scheduling model with the objective of minimizing the sum of the regulation cost and curtailment penalty of the power system; convert the power system scheduling model into a Markov decision process to construct the state space, action space, reward function, and penalty function of the agent; interact the agent with the virtual simulation environment, and iteratively train the agent using a safe reinforcement learning algorithm based on a set constrained optimization objective of the agent to obtain a trained agent; use the trained agent to perform real-time scheduling of the power system, which can effectively avoid the agent from taking actions that produce unsafe results during the exploration process, thereby improving safety, stability, and robustness while ensuring the performance of the policy, and finally achieving a good ability to generate real-time scheduling strategies for high-proportion new energy power systems. Description of the Drawings
[0068] Figure 1It is a flowchart of a real-time scheduling method for a high-proportion new energy power system provided by an embodiment of the present application;
[0069] Figure 2 It is a structural block diagram of a real-time scheduling device for a high-proportion new energy power system provided by an embodiment of the present application;
[0070] Figure 3 It is a structural block diagram of a real-time scheduling device for a high-proportion new energy power system provided by an embodiment of the present application. Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0072] See Figure 1 , Figure 1 It is a flowchart of a real-time scheduling method for a high-proportion new energy power system provided by an embodiment of the present application. The real-time scheduling method for the high-proportion new energy power system includes:
[0073] S1. Construct a virtual simulation environment of the power system according to the component information and historical operation information of the power system;
[0074] S2. Take thermal power units, energy storage units, wind turbines and photovoltaic units as regulation objects, and combine the safe operation constraint conditions of the power system to construct a power system scheduling model with the minimum sum of the regulation cost and curtailment penalty of the power system as the goal; wherein, the safe operation constraint conditions include: power balance constraint, line transmission capacity constraint, thermal power unit operation constraint, energy storage unit operation constraint and maximum curtailment constraint;
[0075] S3. Convert the power system scheduling model into a Markov decision process to construct the state space, action space, reward function and penalty function of the intelligent agent;
[0076] S4. Interact the intelligent agent with the virtual simulation environment, and based on the set constrained optimization goal of the intelligent agent, use a safe reinforcement learning algorithm to iteratively train the intelligent agent to obtain a trained intelligent agent;
[0077] S5. Use the trained intelligent agent to perform real-time scheduling on the power system.
[0078] It should be noted that the power system is a high - proportion new - energy power system. First, the present application constructs a virtual simulation environment for real - time scheduling of a high - proportion new - energy power system, simulates and executes the scheduling instructions of the intelligent agent, conducts simulation and deduction of the operating state of the power system, and returns new state information and adjustment costs. Second, based on the data - driven safety reinforcement learning algorithm, it effectively absorbs the advantages of the neural network's self - learning, high computational efficiency, and independence from accurate mathematical models, and overcomes the problems of a large number of controllable objects and difficult - to - accurately - construct mechanism models faced in the real - time scheduling of a high - proportion new - energy power system. Finally, by constructing a constrained Markov decision process, a penalty evaluation network is established to evaluate the safety of the intelligent agent's actions, guiding the intelligent agent to gradually meet the constraints during training and improving the usability of intelligent decision - making. Ultimately, it realizes a good ability to generate real - time scheduling strategies for a high - proportion new - energy power system, with high computational efficiency, safety, and stability.
[0079] In an alternative embodiment, constructing the virtual simulation environment of the power system according to the component information and historical operation information of the power system includes:
[0080] S11. Obtain the original grid data of the power system;
[0081] It should be noted that the original grid data is saved in formats such as CIM / XML, CIM / E, QS, etc. Taking the QS file as an example, this data is the system operation state measurement information, specifically including system information, substation information, bus information, generator information, load information, capacitance - reactance information, AC line information, transformer information, switch information, disconnecting - switch information, earthing - disconnecting - switch information, DC switch information, converter information, DC line information, etc. The original grid data has diverse types and complex content, involving a large amount of information in different dimensions, usually including real - time monitoring information from multiple sources and devices, and has the characteristics of high frequency, timeliness, and multi - dimension. The QS format is an efficient and compressed storage form. Although it helps to reduce storage space and speed up data transmission, it is not suitable for direct data analysis. Therefore, before analyzing and applying the original grid data, it is necessary to perform pre - parsing and pre - processing operations on this type of QS - formatted data packet based on Python.
[0082] S12. Parse and pre - process the original grid data to obtain component information and historical operation information;
[0083] Specifically, the original grid data includes original component information and original historical operation information. Parse and pre - process the original grid data to obtain component information and historical operation information.
[0084] Exemplarily, parsing includes decompression and decoding. Since the data packets in QS format are stored in a compressed manner, they need to be decompressed first. After decompression, decoding is also required to convert them into a readable structured data form.
[0085] Exemplarily, preprocessing includes: column splitting and cleaning, format conversion and standardization, time series alignment, and integration. Specifically, perform column splitting and cleaning, format conversion and standardization, and integration operations on the parsed original component information; perform column splitting and cleaning, format conversion and standardization, time series alignment, and integration operations on the parsed original historical operation information.
[0086] Column splitting and cleaning: The parsed data is often a complex multi-column data set that may contain redundant information, missing values, or outliers. Therefore, through column splitting, the data needs to be disassembled according to different dimensions and fields, and data cleaning is performed to remove noise and incorrect data.
[0087] Format conversion and standardization: Different data sources may use different data formats and units. Therefore, the data needs to be standardized to ensure that all data conforms to a unified analysis standard. This can improve data compatibility and facilitate subsequent modeling and analysis work.
[0088] Time series alignment: Since power grid data has strong time series characteristics, the data collected at different time points needs to be aligned to ensure consistency in the time dimension.
[0089] Integration: Integrate the data from different systems or devices to comprehensively reflect the operating status of the high-proportion new energy system.
[0090] Through parsing and preprocessing operations, the original power grid data packets can be converted into structured data suitable for algorithmic program processing, laying a data foundation for the construction of subsequent power flow calculation simulation models and operation scenario sets.
[0091] S13. Construct a power flow calculation simulation model of the power system according to the component information;
[0092] Exemplarily, use the Pandapower simulation software to construct the power flow calculation simulation model of the power system. Pandapower is a power system analysis tool based on the Python language, supporting multiple functions such as power flow calculation, short-circuit analysis, and optimal power flow.
[0093] Bus: Accurately model the bus to reflect the voltage level, geographical distribution, and connected equipment of each bus.
[0094] Line: Based on the parsed QS data, define the parameters of each line, such as resistance, reactance, admittance, and the start and end nodes of the line.
[0095] Load: For load nodes below 110 kV, through the method of equivalent aggregation, multiple loads at low voltage levels are integrated into a unified load, which is then connected to the corresponding 110 kV node. This processing method not only preserves the accuracy of the overall load characteristics of the power grid but also simplifies the complexity of the model, ensuring the efficiency and accuracy of power grid simulation calculations.
[0096] Generator: It details information on various types of generator sets, such as hydropower, thermal power, wind power, and photovoltaic power, fully reflecting the diversified energy structure and actual operation conditions of the power grid.
[0097] Static Generator: For small power sources below 110 kV, through the method of equivalent aggregation, multiple power sources at low voltage levels are integrated into a unified generator, which is then connected to the corresponding 110 kV node and regarded as an uncontrollable power source.
[0098] DC Line: The DC lines in the power grid are modeled, and information such as the voltage level, power transmission capacity, and start and end nodes of the DC lines is set.
[0099] Through the comprehensive modeling of these six types of key equipment, the power flow calculation simulation model can accurately simulate the actual operation of the power system and support multiple functions such as power flow calculation and optimal power flow. Power flow calculation can help analyze the power distribution and voltage conditions in the power system, laying a foundation for the subsequent interaction simulation of the reinforcement learning agent in the environment.
[0100] S14. Construct an operation scenario set for the power system according to the historical operation information;
[0101] Based on historical operation information, such as historical time series data of wind power, photovoltaic power output, and load power consumption, this embodiment of the present application constructs an operation scenario set for the power system. This operation scenario set comprehensively reflects the operation status and power supply and demand characteristics under different seasons and different time conditions. The main process is as follows:
[0102] Construct data for the day-ahead stage: Integrate the maintenance conditions of various equipment, the day-ahead plans of generator sets, and the day-ahead prediction information of new energy and loads;
[0103] Construct data for the intra-day stage: Integrate the intra-day prediction information of new energy and loads;
[0104] Construct data for the real-time stage: Integrate the real-time data of new energy generation and load power consumption;
[0105] The above data forms the original operation scenario set of the power system. The operation scenarios in the original operation scenario set are verified and corrected to obtain the operation scenario set of the power system:
[0106] Verification and correction: On the three time scales of day-ahead, intra-day, and real-time, check whether the unit output in each original operation scenario set is within the upper and lower limits of the unit's allowable output, and correct the part exceeding the upper and lower limits back to the upper and lower limit boundaries; check whether the unit output between adjacent time sections meets the ramping constraint, and correct the part not meeting the ramping constraint back to the upper and lower bounds of the ramping constraint; for operation scenarios with minor imbalance amounts, correct them in the form of unit proportional sharing, and at the same time delete operation scenarios with large imbalance amounts, and finally form the operation scenario set of the power system.
[0107] S15. Based on the power flow calculation simulation model and the operation scenario set, construct the virtual simulation environment of the power system based on the Python language.
[0108] The embodiment of the present application constructs the virtual simulation environment of the power system, which supports the functions of regulation strategy execution, system state update, observation information return, and reward function calculation:
[0109] Regulation strategy execution: After the intelligent agent outputs an intelligent regulation action, the virtual simulation environment will execute the corresponding system dispatching instructions, such as adjusting the output of the generating unit, the charge and discharge state and power of the energy storage, etc., and modify the power flow calculation simulation model.
[0110] System state update: Conduct power flow simulation calculation for the modified power flow calculation simulation model, the system state will be updated according to the modified power flow calculation simulation model, and the virtual simulation environment will refresh the new state information of the system in real time, such as voltage change, power flow distribution, load power consumption, etc., for the intelligent agent to make the next round of decisions.
[0111] Observation information return: The virtual simulation environment returns the new state information and real-time observation data of the system. The intelligent agent uses it as input and feeds it back into the action network, and can continuously adjust its strategy to adapt to the dynamic changes of the system.
[0112] Reward function calculation and feedback: The virtual simulation environment calculates the reward value and penalty value according to the system operation state and feeds them back to the intelligent agent. The reward function is mainly used to evaluate the economy of the regulation decision, and the penalty function is mainly used to evaluate the safety of the regulation decision.
[0113] In an optional embodiment, the objective function of the power system scheduling model is expressed as:
[0114]
[0115] In the formula, and are decision variables, representing respectively the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the curtailed power of the i-th wind power unit at time t, and the curtailed power of the i-th photovoltaic unit at time t; is the regulation cost of the power system; is the penalty for wind and photovoltaic curtailment in the power system;
[0116]
[0117] In the formula, N G is the number of thermal power units in the power system, is the unit regulation cost of the i-th thermal power unit, is the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, is the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t; N S is the number of energy storage units in the power system, is the unit regulation cost of the i-th energy storage unit, is the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, is the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t; among them, these adjustment amounts are output by the agent.
[0118]
[0119] In the formula, N Wind is the number of wind power units in the power system, is the unit curtailment penalty of the i-th wind power unit, is the curtailed power of the i-th wind power unit at time t; N Solar is the number of photovoltaic units in the power system, is the unit curtailment penalty of the i-th photovoltaic unit, is the curtailed power of the i-th photovoltaic unit at time t.
[0120] In an optional embodiment, the power balance constraint is expressed as:
[0121]
[0122] Among them, N G is the number of thermal power units in the power system, is the real-time plan of the i-th thermal power unit at time t; N Wind is the number of wind power units in the power system, is the real-time plan of the i-th wind turbine at time t; N Solar is the number of photovoltaic units in the power system, is the real-time plan of the i-th photovoltaic unit at time t; N S is the number of energy storage units in the power system, is the real-time plan of the i-th energy storage unit at time t; N Load is the number of loads in the power system, is the real-time demand of the i-th load at time t;
[0123] Specifically,
[0124] In the formula, is the day-ahead plan of the i-th thermal power unit at time t, is the positive adjustment amount of the day-ahead plan of the i-th thermal power unit at time t, is the negative adjustment amount of the day-ahead plan of the i-th thermal power unit at time t;
[0125] Specifically,
[0126] In the formula, is the intra-day predicted power generation of the i-th wind turbine at time t, is the curtailment power of the i-th wind turbine at time t;
[0127] Specifically,
[0128] In the formula, is the intra-day predicted power generation of the i-th photovoltaic unit at time t, is the curtailment power of the i-th photovoltaic unit at time t;
[0129] Specifically,
[0130] In the formula, is the day-ahead plan of the i-th energy storage unit at time t, is the positive adjustment amount of the day-ahead plan of the i-th energy storage unit at time t, is the negative adjustment amount of the day-ahead plan of the i-th energy storage unit at time t;
[0131] The line transmission capacity constraint calculates the line transmission power using the power transfer distribution factor, and the line transmission capacity constraint is expressed as:
[0132]
[0133] Among them, N Bus is the number of buses in the power system; G i-jThe influence of the injected power of the j-th bus unit on the transmission power of the i-th line, which belongs to the physical parameters of the power system; The lower limit of the transmission capacity of the i-th line, The upper limit of the transmission capacity of the i-th line; The real-time plan of the j-th thermal power unit at time t, The real-time plan of the j-th wind power unit at time t, The real-time plan of the j-th photovoltaic unit at time t, The real-time plan of the j-th energy storage unit at time t, The real-time demand of the j-th load at time t;
[0134] The operating constraints of the thermal power unit include: thermal power unit capacity constraint, ramp rate constraint; the operating constraints of the thermal power unit are expressed as:
[0135]
[0136] Among them, The minimum allowable output power of the i-th thermal power unit, The maximum allowable output power of the i-th thermal power unit, The real-time plan of the i-th thermal power unit at time t, The maximum allowable ramp rate of the i-th thermal power unit;
[0137] The operating constraints of the energy storage unit include: charge and discharge power constraint, upper and lower limits of the state of charge of the energy storage, state of charge constraint of the energy storage, charge and discharge state constraint; the operating constraints of the energy storage unit are expressed as:
[0138]
[0139] Among them, The minimum allowable output power of the i-th energy storage unit, The real-time plan of the i-th energy storage unit, The maximum allowable output power of the i-th energy storage unit; The minimum allowable state of charge of the i-th energy storage unit, The real-time state of charge of the i-th energy storage unit, The maximum allowable state of charge of the i-th energy storage unit; The binary variable of the charging state, The binary variable of the discharging state, The real-time plan of the i-th energy storage unit at time t, The charge and discharge efficiency; ΔT is the scheduling time interval, The capacity of the i-th energy storage unit;
[0140] The maximum wind and solar curtailment constraint: Considering the continuous deviation of new energy prediction in the intraday stage, when the dispatchable resources are exhausted and the power balance still cannot be met, the operation of curtailing new energy is allowed; the maximum wind and solar curtailment constraint is expressed as:
[0141]
[0142] where is the curtailed power of the i-th wind turbine; is the intraday predicted power generation of the i-th wind turbine at time t; is the curtailed power of the i-th photovoltaic unit at time t; is the intraday predicted power generation of the i-th photovoltaic unit at time t.
[0143] In an optional embodiment, the state space includes: the current time, the active power of the thermal power unit at time t, the active power of the energy storage unit at time t, the active power of the wind turbine at time t, the active power of the photovoltaic unit at time t, the active power of the load at time t, the upper adjustable space of the thermal power unit at time t, the lower adjustable space of the thermal power unit at time t, the upper adjustable space of the energy storage unit at time t, the lower adjustable space of the energy storage unit at time t, the line load rate at time t, the day-ahead plan of the thermal power unit at time t + 1, the day-ahead plan of the energy storage unit at time t + 1, the intraday predicted demand of the load at time t + 1, the intraday predicted power generation of the wind turbine at time t + 1, and the intraday predicted power generation of the photovoltaic unit at time t + 1;
[0144] The action space includes: the adjustment amount of the output of the thermal power unit in two adjacent time periods and the ratio of the maximum allowable output of the energy storage unit;
[0145] It should be noted that this application adopts a continuous action space with an action range of [-1, 1] to continuously adjust the thermal power output and the energy storage charge-discharge power: where where is the scheduling action of the thermal power agent; is the scheduling action of the energy storage agent; is the scheduling action of the first thermal power agent; is the scheduling action of the N G -th thermal power agent; is the scheduling action of the first energy storage agent; is the scheduling action of the N G -th energy storage agent;
[0146] For the thermal power unit, the action output by the agent is defined as the adjustment amount of the output of the thermal power unit in two adjacent time periods, that is wherein is the real-time plan of the i-th thermal power unit at time t, is the action instruction of the i-th thermal power agent at time t, is the maximum allowable ramp rate of the i-th thermal power unit. For the energy storage unit, the agent outputs the action is defined as the ratio of the maximum allowable output of the energy storage unit, that is wherein is the real-time plan of the i-th energy storage unit at time t, is the action instruction of the i-th energy storage agent at time t, is the maximum allowable output power of the i-th energy storage unit.
[0147] Since real-time scheduling is an optimization problem with constraints, the present application decouples the reward and penalty signals based on the constrained Markov decision process, more precisely describes the objective function and the constraint conditions, so as to improve the learning performance and convergence of the agent.
[0148] The reward function is negatively correlated with the regulation cost, and the reward function is expressed as:
[0149]
[0150] where N G is the number of thermal power units in the power system, is the unit regulation cost of the i-th thermal power unit, is the real-time plan of the i-th thermal power unit at time t, is the day-ahead plan of the i-th thermal power unit at time t; N S is the number of energy storage units in the power system, is the unit regulation cost of the i-th energy storage unit, is the real-time plan of the i-th energy storage unit at time t, is the day-ahead plan of the i-th energy storage unit at time t;
[0151] The penalty function mainly evaluates the satisfaction of the line transmission capacity constraint and the upper and lower limits of the balancing machine output constraint, and the penalty function is expressed as:
[0152]
[0153] where α rho is the weight factor for line overlimit penalty, α slack is the weight factor for balancing machine overlimit penalty; N Line is; ρ i,t is the load rate of the i-th line at time t; the balancing machine is a certain thermal power unit, is the output of the balancing machine at time t, is the upper limit of the allowable output of the balancing machine, is the lower limit of the allowable output of the balancing machine.
[0154] In the safety-enhanced learning algorithm (SAC-Lagrangian) of the embodiment of the present application, a maximum entropy framework is adopted. In addition to maximizing the cumulative reward, the entropy of the policy at each moment is also maximized to increase the randomness of the policy and prevent the policy from converging to the local optimum prematurely.
[0155] Considering that the importance of entropy varies in different states and training stages. For states where the optimal action has not been determined, the entropy should be relatively large, and for states where the optimal action can be clearly defined, the entropy should take a relatively small value. Therefore, the constrained optimization objective of the agent is set as:
[0156]
[0157] In the formula, a represents the agent's action, π represents the agent's policy, s represents the environmental state observed by the agent, E a : π(s) represents taking the expectation with respect to the agent's policy π, γ represents the discount factor, t represents the time, r t represents the immediate reward received by the agent, s0 represents the initial state of the environment, c t represents the immediate penalty received by the agent, c represents the penalty threshold, represents the target entropy, which can be preset;
[0158] Based on the primal-dual safety-enhanced learning method, the optimization objective is transformed into an unconstrained optimization problem by introducing dual variables, that is:
[0159]
[0160] In the formula, is the augmented objective function, f(π) = E a : π(s) [∑ t γ t r t |s0 = s], λ is the adaptive weight coefficient of the agent's action safety term; β is the adaptive weight coefficient of the agent's action random term, which is used to control the importance between entropy and reward in different states; π represents the agent's policy;
[0161] Construct a reward evaluation network and a penalty evaluation network; among them, the loss function of the reward evaluation network is expressed as: The loss function of the penalty evaluation network is expressed as:
[0162] In the formula, φ ris the reward evaluation network parameter, φ c is the penalty evaluation network parameter, τ is the interaction trajectory sampled from the experience pool, For the experience pool, Indicates the expectation on the sampling trajectory, Q r (s t ,a t ) is the reward evaluation network to evaluate the current state s t The next agent applies action a t The reward value function after c (s t ,a t ) is the penalty evaluation network evaluation of the current state s t The next agent applies action a t The penalty value function after t represents the immediate reward received by the agent, c t represents the immediate penalty received by the agent, γ is the discount factor, is the target reward evaluation network output value, The output value of the target penalty reward evaluation network;
[0163] It is worth noting that the reward evaluation network is used to evaluate the economic efficiency of the strategy, and the penalty evaluation network is used to evaluate the security of the strategy. In addition, by constructing an additional target network The target Q value is estimated to stabilize the training and improve the convergence of the algorithm.
[0164] Construct an action network; wherein the loss function of the action network is expressed as:
[0165] In the formula, φ π is the action network parameter, τ is the interaction trajectory sampled from the experience pool, For the experience pool, Indicates the expectation on the sampling trajectory, Q r (s t ,a t ) is the reward evaluation network to evaluate the current state s t The next agent applies action a t The reward value function after the action is: λ is the adaptive weight coefficient of the agent's action safety term, β is the adaptive weight coefficient of the agent's action random term, and π is the agent's strategy;
[0166] The intelligent agent interacts with the virtual simulation environment. During the interaction, the unconstrained optimization problem, the reward evaluation network, the penalty evaluation network and the action network are used to iteratively train the intelligent agent to obtain a trained intelligent agent.
[0167] A real-time scheduling method for a high-proportion new energy power system provided by an embodiment of the present application constructs a virtual simulation environment of the power system according to the component information and historical operation information of the power system; takes thermal power units, energy storage units, wind turbines, and photovoltaic units as control objects, and constructs a power system scheduling model with the goal of minimizing the sum of the adjustment cost and the curtailment penalty of wind and light of the power system in combination with the safe operation constraint conditions of the power system; converts the power system scheduling model into a Markov decision process to construct the state space, action space, reward function, and penalty function of the intelligent agent; interacts the intelligent agent with the virtual simulation environment, and based on the set constrained optimization goal of the intelligent agent, uses a safe reinforcement learning algorithm to iteratively train the intelligent agent to obtain a trained intelligent agent; uses the trained intelligent agent to perform real-time scheduling of the power system, which can effectively avoid the intelligent agent from taking actions that produce unsafe results during the exploration process, thereby improving safety, stability, and robustness while ensuring the performance of the strategy, and finally realizing a good ability to generate real-time scheduling strategies for high-proportion new energy power systems.
[0168] See Figure 2 , Figure 2 FIG. is a structural block diagram of a real-time scheduling device 10 for a high-proportion new energy power system provided by an embodiment of the present application. The real-time scheduling device for a high-proportion new energy power system includes:
[0169] A simulation environment construction module 11, configured to construct a virtual simulation environment of the power system according to the component information and historical operation information of the power system;
[0170] A model construction module 12, configured to take thermal power units, energy storage units, wind turbines, and photovoltaic units as control objects, and construct a power system scheduling model with the goal of minimizing the sum of the adjustment cost and the curtailment penalty of wind and light of the power system in combination with the safe operation constraint conditions of the power system; wherein, the safe operation constraint conditions include: power balance constraint, line transmission capacity constraint, thermal power unit operation constraint, energy storage unit operation constraint, and maximum curtailment of wind and light constraint;
[0171] An intelligent agent construction module 13, configured to convert the power system scheduling model into a Markov decision process to construct the state space, action space, reward function, and penalty function of the intelligent agent;
[0172] An intelligent agent training module 14, configured to interact the intelligent agent with the virtual simulation environment, and based on the set constrained optimization goal of the intelligent agent, use a safe reinforcement learning algorithm to iteratively train the intelligent agent to obtain a trained intelligent agent;
[0173] A real-time scheduling module 15, configured to perform real-time scheduling of the power system by using the trained intelligent agent.
[0174] It should be noted that the working processes of the various modules in the real-time scheduling device 10 of the high-proportion new energy power system described in the embodiments of the present application can refer to the working process of the real-time scheduling method of the high-proportion new energy power system described in the above embodiments, and will not be elaborated here.
[0175] The real-time scheduling device 10 of a high-proportion new energy power system provided by the embodiments of the present application constructs a virtual simulation environment of the power system according to the component information and historical operation information of the power system; takes thermal power units, energy storage units, wind turbine units, and photovoltaic units as control objects, and constructs a power system scheduling model with the goal of minimizing the sum of the regulation cost and curtailment penalty of the power system in combination with the safe operation constraint conditions of the power system; converts the power system scheduling model into a Markov decision process to construct the state space, action space, reward function, and penalty function of the intelligent agent; interacts the intelligent agent with the virtual simulation environment, and based on the set constrained optimization goal of the intelligent agent, uses a safe reinforcement learning algorithm to iteratively train the intelligent agent to obtain a trained intelligent agent; uses the trained intelligent agent to perform real-time scheduling of the power system, which can effectively avoid the intelligent agent from taking actions that produce unsafe results during the exploration process, thereby improving safety, stability, and robustness while ensuring the performance of the strategy, and finally realizing a good ability to generate real-time scheduling strategies for high-proportion new energy power systems.
[0176] In addition, the embodiments of the present application also provide a computer-readable storage medium, which includes a stored computer program; wherein, the computer program controls the device where the computer-readable storage medium is located to execute the real-time scheduling method of the high-proportion new energy power system described in any of the above embodiments when running.
[0177] In addition, the embodiments of the present application also provide a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, they implement the real-time scheduling method of the high-proportion new energy power system described in any of the above embodiments.
[0178] See Figure 3 , Figure 3 is a structural block diagram of a real-time scheduling device 20 of a high-proportion new energy power system provided by the embodiments of the present application. The real-time scheduling device 20 of the high-proportion new energy power system includes: a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, it implements the steps in the embodiments of the above real-time scheduling method of the high-proportion new energy power system. Or, when the processor 21 executes the computer program, it implements the functions of the various modules / units in the above device embodiments.
[0179] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 22 and executed by the processor 21 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the real-time scheduling device 20 of the high-proportion new energy power system.
[0180] The real-time scheduling device 20 of the high-proportion new energy power system may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the real-time scheduling device 20 of the high-proportion new energy power system, and does not constitute a limitation on the real-time scheduling device 20 of the high-proportion new energy power system. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the real-time scheduling device 20 of the high-proportion new energy power system may also include input / output devices, network access devices, buses, etc.
[0181] The processor 21 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 21 is the control center of the real-time scheduling device 20 of the high-proportion new energy power system, and connects various parts of the entire real-time scheduling device 20 of the high-proportion new energy power system through various interfaces and lines.
[0182] The memory 22 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 22, and invoking the data stored in the memory 22, the processor 21 realizes various functions of the real-time scheduling device 20 of the high-proportion new energy power system. The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0183] Among them, if the modules / units integrated in the real-time scheduling device 20 of the high-proportion new energy power system are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0184] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0185] The above is the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of this application.
Claims
1. A real-time dispatching method for a high-proportion renewable energy power system, characterized in that: include: Constructing a virtual simulation environment of the power system based on component information and historical operation information of the power system; Taking thermal power units, energy storage units, wind power units and photovoltaic units as the control objects, and combining the safe operation constraints of the power system, a power system dispatching model is constructed with the goal of minimizing the sum of the regulation cost of the power system and the penalty for wind and solar power abandonment; wherein the safe operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints and maximum wind and solar power abandonment constraints; Convert the power system dispatch model into a Markov decision process to construct the state space, action space, reward function and penalty function of the intelligent agent; The intelligent agent interacts with the virtual simulation environment, and based on the set constrained optimization goal of the intelligent agent, the intelligent agent is iteratively trained using a safe reinforcement learning algorithm to obtain a trained intelligent agent; The trained intelligent agent is used to perform real-time dispatch of the power system.
2. The real-time dispatching method for a high-proportion new energy power system according to claim 1, characterized in that: The objective function of the power system dispatch model is expressed as: In the formula, and are decision variables, which are respectively represented by the positive adjustment of the i-th thermal power unit to the day-ahead plan at time t, the negative adjustment of the i-th thermal power unit to the day-ahead plan at time t, the positive adjustment of the i-th energy storage unit to the day-ahead plan at time t, the negative adjustment of the i-th energy storage unit to the day-ahead plan at time t, the power abandonment of the i-th wind turbine unit at time t, and the power abandonment of the i-th photovoltaic unit at time t; The regulation cost of the power system; Penalize wind and solar power curtailment in power systems; Where N G is the number of thermal power units in the power system, is the unit regulation cost of the i-th thermal power unit, is the positive adjustment of the i-th thermal power unit to the day-ahead plan at time t, is the negative adjustment of the i-th thermal power unit to the day-ahead plan at time t; N S is the number of energy storage units in the power system, is the unit regulation cost of the i-th energy storage unit, is the positive adjustment of the i-th energy storage unit to the day-ahead plan at time t, is the negative adjustment of the i-th energy storage unit to the day-ahead plan at time t; Where N Wind is the number of wind turbines in the power system, is the unit power abandonment penalty of the i-th wind turbine, is the amount of power abandoned by the i-th wind turbine at time t; N Solar is the number of photovoltaic units in the power system, is the unit power abandonment penalty of the i-th PV unit, is the amount of power abandoned by the ith PV unit at time t.
3. The real-time dispatching method for a high-proportion new energy power system according to claim 1, characterized in that: The power balance constraint is expressed as: Among them, N G is the number of thermal power units in the power system, is the real-time plan of the i-th thermal power unit at time t; N Wind is the number of wind turbines in the power system, is the real-time plan of the i-th wind turbine at time t; N Solar is the number of photovoltaic units in the power system, is the real-time plan of the ith PV unit at time t; N S is the number of energy storage units in the power system, is the real-time plan of the i-th energy storage unit at time t; N Load is the number of loads in the power system, is the real-time demand of the i-th load at time t; The line transmission capacity constraint is expressed as: Among them, N Bus is the number of buses in the power system; G i-j The impact of the j-th bus unit injection power on the i-th line transmission power; is the lower limit of the transmission capacity of the ith line, is the upper limit of the transmission capacity of the ith line; is the real-time plan of the jth thermal power unit at time t, is the real-time plan of the j-th wind turbine at time t, is the real-time plan of the j-th PV unit at time t, is the real-time plan of the j-th energy storage unit at time t, is the real-time demand of the jth load at time t; The operation constraints of the thermal power unit are expressed as: in, is the minimum output power allowed for the i-th thermal power unit, is the maximum output power allowed for the i-th thermal power unit, is the real-time plan of the i-th thermal power unit at time t, is the maximum ramp rate allowed for the i-th thermal power unit; The energy storage unit operation constraints are expressed as: in, is the minimum output power allowed for the i-th energy storage unit, is the real-time plan of the i-th energy storage unit, is the maximum output power allowed by the i-th energy storage unit; is the minimum state of charge allowed for the i-th energy storage unit, is the real-time charge state of the i-th energy storage unit, is the maximum state of charge allowed for the i-th energy storage unit; is a binary variable of charging status, is a binary variable of the discharge state, is the real-time plan of the i-th energy storage unit at time t, is the charging and discharging efficiency; ΔT is the scheduling time interval, is the capacity of the i-th energy storage unit; The maximum wind and solar curtailment constraint is expressed as: in, is the power abandonment of the i-th wind turbine, The daily predicted power generation of the i-th wind turbine at time t; is the amount of power abandoned by the ith PV unit at time t, The daily power generation forecast for the i-th photovoltaic unit at time t.
4. The real-time dispatching method for a high-proportion new energy power system according to claim 1, characterized in that: The state space includes: the current time, the active power of the thermal power unit at time t, the active power of the energy storage unit at time t, the active power of the wind power unit at time t, the active power of the photovoltaic unit at time t, the active power of the load at time t, the upper adjustable space of the thermal power unit at time t, the lower adjustable space of the thermal power unit at time t, the upper adjustable space of the energy storage unit at time t, the lower adjustable space of the energy storage unit at time t, the line load rate at time t, the day-ahead plan of the thermal power unit at time t+1, the day-ahead plan of the energy storage unit at time t+1, the intraday forecast demand of the load at time t+1, the intraday forecast power generation of the wind power unit at time t+1, and the intraday forecast power generation of the photovoltaic unit at time t+1; The action space includes: the ratio of the output adjustment amount of the thermal power unit in two adjacent time periods to the maximum output allowed by the energy storage unit; The reward function is expressed as: Among them, N G is the number of thermal power units in the power system, is the unit regulation cost of the i-th thermal power unit, is the real-time plan of the i-th thermal power unit at time t, is the day-ahead plan of the i-th thermal power unit at time t; N s is the number of energy storage units in the power system, is the unit regulation cost of the i-th energy storage unit, is the real-time plan of the i-th energy storage unit at time t, is the day-ahead plan of the i-th energy storage unit at time t; The penalty function is expressed as: Among them, α rho is the weight factor of line limit penalty, α slack N is the weight factor of the balance machine's over-limit penalty; Line is the number of lines in the power system; ρ i,t is the load rate of the ith line at time t; is the output of the balancing machine at time t, The upper limit of the output allowed by the balancing machine. It is the lower limit of the output allowed by the balancing machine.
5. The real-time dispatching method for a high-proportion new energy power system according to claim 1, characterized in that: The intelligent agent is interacted with the virtual simulation environment, and based on the set constraint optimization goal of the intelligent agent, a safe reinforcement learning algorithm is used to iteratively train the intelligent agent to obtain a trained intelligent agent, including: Set the agent's constrained optimization goal: In the formula, a represents the agent action, π represents the agent strategy, s represents the environment state observed by the agent, and E a : π(s) represents the expectation of the agent strategy π, γ represents the discount factor, t represents the time, r t represents the immediate reward received by the agent, s0 represents the initial state of the environment, c t represents the immediate penalty received by the agent, c represents the penalty threshold, represents the target entropy; The optimization objective is transformed into an unconstrained optimization problem, namely: In the formula, is the augmented objective function, f(π)=E a : π(s) [∑ t γ t r t |s0=s], λ is the adaptive weight coefficient of the agent action safety term; β is the adaptive weight coefficient of the agent action random term; π represents the agent strategy; Construct a reward evaluation network and a penalty evaluation network; wherein the loss function of the reward evaluation network is expressed as: The loss function of the penalty evaluation network is expressed as: In the formula, φ r is the reward evaluation network parameter, φ c is the penalty evaluation network parameter, τ is the interaction trajectory sampled from the experience pool, For the experience pool, Indicates the expectation on the sampling trajectory, Q r (s t ,a t ) is the reward evaluation network to evaluate the current state s t The next agent applies action a t The reward value function after c (s t ,a t ) is the penalty evaluation network evaluation of the current state s t The next agent applies action a t The penalty value function after t represents the immediate reward received by the agent, c t represents the immediate penalty received by the agent, γ is the discount factor, The output value of the target reward evaluation network, The output value of the target penalty reward evaluation network; Construct an action network; wherein the loss function of the action network is expressed as: In the formula, φ π is the action network parameter, τ is the interaction trajectory sampled from the experience pool, For the experience pool, Indicates the expectation on the sampling trajectory, Q r (s t ,a t ) is the reward evaluation network to evaluate the current state s t The next agent applies action a t The reward value function after the action is: λ is the adaptive weight coefficient of the agent's action safety term, β is the adaptive weight coefficient of the agent's action random term, and π is the agent's strategy; The intelligent agent interacts with the virtual simulation environment. During the interaction, the unconstrained optimization problem, the reward evaluation network, the penalty evaluation network and the action network are used to iteratively train the intelligent agent to obtain a trained intelligent agent.
6. The real-time dispatching method for a high-proportion new energy power system according to claim 1, characterized in that: The step of constructing a virtual simulation environment of the power system according to the component information and historical operation information of the power system includes: Acquiring original grid data of the power system; Parsing and preprocessing the original power grid data to obtain component information and historical operation information; Constructing a power flow calculation simulation model of the power system according to the component information; constructing an operation scenario set of the power system according to the historical operation information; A virtual simulation environment of the power system is constructed according to the power flow calculation simulation model and the operation scenario set.
7. A real-time dispatching device for a high-proportion renewable energy power system, characterized in that: include: A simulation environment construction module, used to construct a virtual simulation environment of the power system according to component information and historical operation information of the power system; A model building module is used to take thermal power units, energy storage units, wind power units and photovoltaic units as control objects, and in combination with the safe operation constraints of the power system, to build a power system dispatching model with the goal of minimizing the sum of the regulation cost of the power system and the penalty for wind and solar power abandonment; wherein the safe operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints and maximum wind and solar power abandonment constraints; An agent building module, used to convert the power system dispatch model into a Markov decision process to construct the state space, action space, reward function and penalty function of the agent; An agent training module is used to interact the agent with the virtual simulation environment, and based on the set constraint optimization goal of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent; The real-time dispatching module is used to utilize the trained intelligent agent to perform real-time dispatching of the power system.
8. A real-time dispatching device for a high-proportion renewable energy power system, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the real-time scheduling method of the high-proportion new energy power system as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the real-time scheduling method for a high-proportion new energy power system as described in any one of claims 1 to 6.
10. A computer program product, characterized in that It comprises a computer program / instruction, which, when executed by a processor, implements the real-time dispatching method for a high-proportion new energy power system as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Power distribution network scheduling method and device based on deep reinforcement learning, and medium
CN115169957A
Active power distribution network real-time scheduling method and device based on safety reinforcement learning
CN115714382A
Economic dispatching method of household energy management system and related device
CN118611081A
Multi-agent reinforcement learning scheduling method based on optimality guarantee
CN119150497A
Comprehensive energy system economic dispatching model method based on deep reinforcement learning
CN119273066A