A real-time scheduling method for a high-proportion new energy power system and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]本申请提供一种高比例新能源电力系统的实时调度方法及相关设备,以解决现有技术中智能体的动作没有被明确安全约束,导致智能体可能会采取产生不安全结果的行为
[0067] Compared with existing technologies, the real-time scheduling method, apparatus, equipment, storage medium, and product for a high-proportion renewable energy power system provided in this application embodiment constructs a virtual simulation environment for the power system based on the component information and historical operation information of the power system; taking thermal power units, energy storage units, wind power units, and photovoltaic units as the control objects, and combining the safe operation constraints of the power system, a power system scheduling model is constructed with the objective of minimizing the sum of the power system's regulation cost and the penalty for wind and solar curtailment; the power system scheduling model is converted into a Markov decision process to construct the state space, action space, reward function, and penalty function of the agent; the agent interacts with the virtual simulation environment, and based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent; using the trained agent to perform real-time scheduling of the power system can effectively avoid the agent from taking actions that produce unsafe results during the exploration process, thereby improving safety, stability, and robustness while ensuring strategy performance, and ultimately achieving a good real-time scheduling strategy generation capability for a high-proportion renewable energy power system.
Smart Images

Figure CN120185084B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of dispatching technology, and in particular to a real-time dispatching method and related equipment for a high-proportion renewable energy power system. Background Technology
[0002] With the increasing number of large-scale grid-connected clean energy clusters such as wind power and photovoltaics, traditional reinforcement learning, as a data-driven learning method, is widely used in the real-time scheduling of high-proportion renewable energy power systems. Through the interaction between the agent and the environment, it can autonomously learn from historical experience, search for optimal control strategies, and maximize cumulative rewards. However, traditional reinforcement learning frameworks have certain shortcomings in terms of security. The agent's actions are not subject to explicit safety constraints, which may lead to actions that produce unsafe outcomes. Summary of the Invention
[0003] This application provides a real-time scheduling method and related equipment for a high-proportion renewable energy power system to address the problem that in the prior art, the actions of intelligent agents are not subject to explicit safety constraints, which may lead to actions by intelligent agents that produce unsafe results.
[0004] To achieve the above objectives, embodiments of this application provide a real-time dispatching method for a high-proportion renewable energy power system, comprising:
[0005] Based on the component information and historical operation information of the power system, a virtual simulation environment for the power system is constructed.
[0006] Using thermal power units, energy storage units, wind power units, and photovoltaic units as the control targets, and combining the safety operation constraints of the power system, a power system dispatch model is constructed with the objective of minimizing the sum of the power system's regulation costs and wind and solar curtailment penalties. The safety operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints, and maximum wind and solar curtailment constraints.
[0007] The power system scheduling model is transformed into a Markov decision process to construct the agent's state space, action space, reward function, and penalty function.
[0008] The agent interacts with the virtual simulation environment. Based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent.
[0009] The trained intelligent agent is used to perform real-time scheduling of the power system.
[0010] As an improvement to the above scheme, the objective function of the power system dispatching model is expressed as:
[0011]
[0012] In the formula, and Let be the decision variables, representing the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the curtailment amount of the i-th wind power unit at time t, and the curtailment amount of the i-th photovoltaic unit at time t. For the regulation costs of the power system; Penalties for curtailment of wind and solar power in the power system;
[0013]
[0014] In the formula, N G This refers to the number of thermal power units in the power system. Let be the unit regulation cost of the i-th thermal power unit. Let be the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t. N represents the negative adjustment of the i-th thermal power unit to the day-ahead plan at time t; S The number of energy storage units in the power system. Let i be the unit regulation cost of the i-th energy storage unit. Let be the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t. Let be the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t;
[0015]
[0016] In the formula, N Wind The number of wind turbines in the power system. The unit curtailment penalty for the i-th wind turbine is... N represents the amount of power wasted by the i-th wind turbine at time t; Solar The number of photovoltaic units in the power system. The unit curtailment penalty for the i-th photovoltaic unit is... Let represent the amount of electricity wasted by the i-th photovoltaic unit at time t.
[0017] As an improvement to the above scheme, the power balance constraint is expressed as:
[0018]
[0019] Where, N G This refers to the number of thermal power units in the power system. N represents the real-time plan for the i-th thermal power unit at time t; Wind The number of wind turbines in the power system. N represents the real-time plan for the i-th wind turbine at time t; Solar The number of photovoltaic units in the power system. N represents the real-time plan for the i-th photovoltaic unit at time t; S The number of energy storage units in the power system. N represents the real-time plan for the i-th energy storage unit at time t; Load The number of loads in the power system. Let be the real-time demand of the i-th load at time t;
[0020] The line transmission capacity constraint is expressed as follows:
[0021]
[0022] Where, N Bus G represents the number of busbars in the power system. i-j The effect of injecting unit power into the j-th bus on the transmission power of the i-th line; This is the lower limit of the transmission capacity of the i-th line. This represents the upper limit of the transmission capacity of the i-th line; For the real-time plan of the j-th thermal power unit at time t, For the j-th wind turbine, the real-time plan at time t is... For the real-time plan of the j-th photovoltaic unit at time t, For the j-th energy storage unit at time t, Let be the real-time demand of the j-th load at time t;
[0023] The operating constraints of the thermal power unit are expressed as follows:
[0024]
[0025] in, Let be the minimum allowable output power of the i-th thermal power unit. Let i be the maximum allowable output power of the i-th thermal power unit. Let be the real-time plan for the i-th thermal power unit at time t. Let be the maximum allowed climbing rate for the i-th thermal power unit;
[0026] The operating constraints of the energy storage unit are expressed as follows:
[0027]
[0028] in, Let i be the minimum allowable output power of the i-th energy storage unit. For the real-time plan of the i-th energy storage unit, This represents the maximum allowable output power of the i-th energy storage unit; This represents the minimum permissible state of charge for the i-th energy storage unit. This represents the real-time state of charge of the i-th energy storage unit. This represents the maximum allowable state of charge for the i-th energy storage unit. A binary variable representing the charging state. This is a binary variable representing the discharge state. For the i-th energy storage unit at time t, ΔT represents the charge / discharge efficiency; ΔT represents the scheduling time interval. Let i be the capacity of the i-th energy storage unit;
[0029] The maximum wind and solar curtailment constraint is expressed as:
[0030]
[0031] in, Let i be the amount of abandoned power of the i-th wind turbine. For the i-th wind turbine, predict the daily power generation at time t; Let be the amount of power curtailed by the i-th photovoltaic unit at time t. The predicted daily power generation of the i-th photovoltaic unit at time t.
[0032] As an improvement to the above scheme, the state space includes: the active power of the thermal power unit at the current time, the active power of the energy storage unit at the time of t, the active power of the wind power unit at the time of t, the active power of the photovoltaic unit at the time of t, the active power of the load at the time of t, the upper adjustable space of the thermal power unit at the time of t, the lower adjustable space of the thermal power unit at the time of t, the upper adjustable space of the energy storage unit at the time of t, the lower adjustable space of the energy storage unit at the time of t, the line load rate at the time of t, the daily plan of the thermal power unit at the time of t+1, the daily plan of the energy storage unit at the time of t+1, the intraday predicted demand of the load at the time of t+1, the intraday predicted power generation of the wind power unit at the time of t+1, and the intraday predicted power generation of the photovoltaic unit at the time of t+1.
[0033] The operational space includes: the output adjustment amount of the thermal power unit in two adjacent time periods and the ratio of the maximum allowable output of the energy storage unit;
[0034] The reward function is expressed as follows:
[0035]
[0036] Where, N GThis refers to the number of thermal power units in the power system. Let be the unit regulation cost of the i-th thermal power unit. Let be the real-time plan for the i-th thermal power unit at time t. Let N be the day-ahead schedule for the i-th thermal power unit at time t; S The number of energy storage units in the power system. Let i be the unit regulation cost of the i-th energy storage unit. For the i-th energy storage unit at time t, Let i be the day-ahead plan for the i-th energy storage unit at time t;
[0037] The penalty function is expressed as follows:
[0038]
[0039] Where, α rho α is the weighting factor for the line over-limit penalty. slack N is the weighting factor for the balancing machine's over-limit penalty; Line ρ represents the number of lines in the power system. i,t Let be the load rate of the i-th line at time t; Let be the output force of the balancing machine at time t. This is the upper limit of the output allowed by the balancing machine. This is the lower limit of the allowable output of the balancing machine.
[0040] As an improvement to the above scheme, the interaction between the intelligent agent and the virtual simulation environment, and the iterative training of the intelligent agent using a safe reinforcement learning algorithm based on a set constrained optimization objective, to obtain a trained intelligent agent, includes:
[0041] Define a constrained optimization objective for the agent:
[0042]
[0043] In the formula, a represents the agent's action, π represents the agent's policy, s represents the environmental state observed by the agent, and E a : π(s) Let γ represent the expectation of the agent's policy π, γ represent the discount factor, t represent time, and r represent the time interval. t s0 represents the immediate reward received by the agent, c represents the initial state of the environment, and s0 represents the initial state of the environment. t This represents the immediate penalty received by the agent, where c represents the penalty threshold. Represents the target entropy;
[0044] The optimization objective is transformed into an unconstrained optimization problem, namely:
[0045]
[0046] In the formula, To augment the objective function, f(π) = E a : π(s) [∑ t γ t r t |s0=s], λ is the adaptive weighting coefficient of the agent's action safety term; β is the adaptive weighting coefficient of the agent's action random term; π represents the agent's policy;
[0047] Construct a reward evaluation network and a penalty evaluation network; wherein the loss function of the reward evaluation network is expressed as: The loss function of the penalty evaluation network is expressed as:
[0048] In the formula, φ r To reward the evaluation of network parameters, φ c To penalize the evaluation network parameters, τ represents the interaction trajectory sampled from the experience pool. For experience pool, This represents calculating the expectation along the sampling trajectory, Q. r (s t ,a t The reward evaluation network assesses the current state s. t The lower agent applies action a t The reward value function after that, Q c (s t ,a t The network evaluates the current state s as a penalty evaluation network. t The lower agent applies action a t The subsequent penalty value function, r t c represents the immediate reward received by the agent. t This represents the immediate penalty received by the agent, where γ is the discount factor. The network output value is evaluated based on the target reward. The network output value is used to evaluate the target penalty and reward.
[0049] Construct an action network; wherein the loss function of the action network is expressed as:
[0050] In the formula, φ π Here are the action network parameters, and τ is the interaction trajectory sampled from the experience pool. For experience pool, This represents calculating the expectation along the sampling trajectory, Q. r (s t ,a t The reward evaluation network assesses the current state s.t The lower agent applies action a t The reward value function is denoted by λ, which is the adaptive weight coefficient of the agent's action safety term, β, which is the adaptive weight coefficient of the agent's action random term, and π, which is the agent's policy.
[0051] The agent interacts with the virtual simulation environment. During the interaction, the agent is iteratively trained using the unconstrained optimization problem, the reward evaluation network, the penalty evaluation network, and the action network to obtain a trained agent.
[0052] As an improvement to the above solution, the step of constructing a virtual simulation environment for the power system based on the component information and historical operating information of the power system includes:
[0053] Obtain the raw power grid data of the power system;
[0054] The raw power grid data is parsed and preprocessed to obtain component information and historical operation information;
[0055] Based on the component information, a power flow calculation simulation model of the power system is constructed;
[0056] Based on the historical operating information, a set of operating scenarios for the power system is constructed;
[0057] Based on the power flow calculation simulation model and the set of operating scenarios, a virtual simulation environment for the power system is constructed.
[0058] To achieve the above objectives, embodiments of this application also provide a real-time dispatching device for a high-proportion renewable energy power system, comprising:
[0059] The simulation environment construction module is used to construct a virtual simulation environment for the power system based on the component information and historical operating information of the power system.
[0060] The model building module is used to construct a power system dispatch model with the objectives of minimizing the sum of the power system's regulation costs and wind / solar curtailment penalties, taking thermal power units, energy storage units, wind power units, and photovoltaic units as the control objects and combining them with the power system's safe operation constraints. The safe operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints, and maximum wind / solar curtailment constraints.
[0061] The agent construction module is used to convert the power system scheduling model into a Markov decision process to construct the agent's state space, action space, reward function, and penalty function.
[0062] The agent training module is used to enable the agent to interact with the virtual simulation environment. Based on the set constrained optimization objectives of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent.
[0063] The real-time scheduling module is used to perform real-time scheduling of the power system using a trained intelligent agent.
[0064] To achieve the above objectives, this application also provides a real-time scheduling device for a high-proportion renewable energy power system, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the real-time scheduling method for the high-proportion renewable energy power system as described above.
[0065] To achieve the above objectives, embodiments of this application also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the real-time scheduling method for a high-proportion renewable energy power system as described above.
[0066] To achieve the above objectives, this application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the real-time scheduling method for a high-proportion renewable energy power system as described above.
[0067] Compared with existing technologies, the real-time scheduling method, apparatus, equipment, storage medium, and product for a high-proportion renewable energy power system provided in this application embodiment constructs a virtual simulation environment for the power system based on the component information and historical operation information of the power system; taking thermal power units, energy storage units, wind power units, and photovoltaic units as the control objects, and combining the safe operation constraints of the power system, a power system scheduling model is constructed with the objective of minimizing the sum of the power system's regulation cost and the penalty for wind and solar curtailment; the power system scheduling model is converted into a Markov decision process to construct the state space, action space, reward function, and penalty function of the agent; the agent interacts with the virtual simulation environment, and based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent; using the trained agent to perform real-time scheduling of the power system can effectively avoid the agent from taking actions that produce unsafe results during the exploration process, thereby improving safety, stability, and robustness while ensuring strategy performance, and ultimately achieving a good real-time scheduling strategy generation capability for a high-proportion renewable energy power system. Attached Figure Description
[0068] Figure 1This is a flowchart of a real-time dispatching method for a high-proportion renewable energy power system provided in an embodiment of this application;
[0069] Figure 2 This is a structural block diagram of a real-time dispatching device for a high-proportion renewable energy power system provided in an embodiment of this application;
[0070] Figure 3 This is a structural block diagram of a real-time dispatching device for a high-proportion new energy power system provided in an embodiment of this application. Detailed Implementation
[0071] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0072] See Figure 1 , Figure 1 This is a flowchart illustrating a real-time dispatching method for a high-proportion renewable energy power system provided in an embodiment of this application. The real-time dispatching method for a high-proportion renewable energy power system includes:
[0073] S1. Construct a virtual simulation environment for the power system based on the component information and historical operating information of the power system;
[0074] S2. Taking thermal power units, energy storage units, wind power units, and photovoltaic units as the control objects, and combining the safety operation constraints of the power system, construct a power system dispatch model with the objective of minimizing the sum of the power system's regulation cost and wind / solar curtailment penalties; wherein, the safety operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints, and maximum wind / solar curtailment constraints;
[0075] S3. The power system scheduling model is converted into a Markov decision process to construct the state space, action space, reward function and penalty function of the agent;
[0076] S4. The agent interacts with the virtual simulation environment. Based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent.
[0077] S5. The trained intelligent agent is used to perform real-time scheduling of the power system.
[0078] It is worth noting that the power system in question is a high-proportion renewable energy power system. This application first constructs a virtual simulation environment for real-time dispatching of such a system, simulating the execution of dispatching instructions by an intelligent agent, simulating and extrapolating the power system's operating state, and returning new state information and adjustment costs. Second, based on a data-driven secure reinforcement learning algorithm, it effectively leverages the advantages of neural networks—self-learning, high computational efficiency, and independence from precise mathematical models—overcoming the challenge of accurately constructing massive controllable objects and mechanistic models for real-time dispatching of high-proportion renewable energy power systems. Finally, by constructing a constrained Markov decision process and establishing a penalty evaluation network to assess the safety of the agent's actions, the application guides the agent to gradually satisfy constraints during training, improving the usability of intelligent decision-making. Ultimately, it achieves a good real-time dispatching strategy generation capability for high-proportion renewable energy power systems, exhibiting high computational efficiency, security, and stability.
[0079] In one optional embodiment, constructing the virtual simulation environment of the power system based on the component information and historical operating information of the power system includes:
[0080] S11. Obtain the raw power grid data of the power system;
[0081] It is worth noting that raw power grid data is saved in formats such as CIM / XML, CIM / E, and QS. Taking QS files as an example, this data contains system operating status measurement information, specifically including system information, substation information, bus information, generator information, load information, capacitor and reactance information, AC line information, transformer information, switch information, disconnector information, grounding disconnector information, DC switch information, converter information, and DC line information. The raw power grid data is diverse and complex, involving a large amount of information from different dimensions. It typically includes real-time monitoring information from multiple sources and devices, characterized by high frequency, timeliness, and multi-dimensionality. While the QS format is an efficient and compressed storage format that helps reduce storage space and speed up data transmission, it is not suitable for direct data analysis. Therefore, before analyzing and applying raw power grid data, it is necessary to pre-parse and preprocess these QS format data packets using Python.
[0082] S12. The original power grid data is parsed and preprocessed to obtain component information and historical operation information;
[0083] Specifically, the raw power grid data includes raw component information and raw historical operating information. The raw power grid data is parsed and preprocessed to obtain component information and historical operating information.
[0084] For example, parsing includes decompression and decoding. Since QS format data packets are stored in a compressed manner, they first need to be decompressed. After decompression, decoding is also required to convert them into a readable structured data format.
[0085] For example, preprocessing includes: splitting and cleaning, format conversion and standardization, timing alignment, and integration. Specifically, the parsed raw component information is split and cleaned, format converted and standardized, and integrated; the parsed raw historical running information is also split and cleaned, format converted and standardized, and timing aligned and integrated.
[0086] Splitting and Cleaning: The parsed data is often a complex multi-column dataset that may contain redundant information, missing values, or outliers. Therefore, it is necessary to split the data according to different dimensions and fields through splitting, and then perform data cleaning to remove noise and erroneous data.
[0087] Format conversion and standardization: Different data sources may use different data formats and units, so data standardization is necessary to ensure that all data conforms to a unified analysis standard. This improves data compatibility and facilitates subsequent modeling and analysis.
[0088] Time series alignment: Since power grid data has strong time series characteristics, it is necessary to align data collected at different time points to ensure consistency in the time dimension.
[0089] Integration: Integrating data from different systems or devices to comprehensively reflect the operating status of high-proportion renewable energy systems.
[0090] Through parsing and preprocessing, the raw power grid data packets can be transformed into structured data suitable for algorithm processing, laying the data foundation for the subsequent construction of power flow calculation simulation models and operating scenario sets.
[0091] S13. Based on the component information, construct a power flow calculation simulation model for the power system;
[0092] For example, a power flow calculation simulation model of the power system is constructed using Pandapower simulation software. Pandapower is a power system analysis tool based on the Python language, supporting multiple functions such as power flow calculation, short-circuit analysis, and optimal power flow.
[0093] Busbars: Accurate modeling of busbars reflects the voltage level, geographical distribution, and connected equipment of each busbar.
[0094] Lines: Based on the parsed QS data, the parameters of each line are defined, such as resistance, reactance, admittance, and the start and end nodes of the line.
[0095] Load: For load nodes below 110kV, multiple low-voltage loads are integrated into a single unified load using an equivalent aggregation method and then connected to the corresponding 110kV node. This approach preserves the accuracy of the overall load characteristics of the power grid while simplifying the model's complexity, ensuring the efficiency and accuracy of power grid simulation calculations.
[0096] Generators: This section includes detailed information on various types of generator sets, such as hydropower, thermal power, wind power, and photovoltaic power, fully reflecting the diversified energy structure and actual operation of the power grid.
[0097] Static generator: For small power sources below 110kV, multiple low-voltage power sources are integrated into a unified generator through the equivalent aggregation method, and then connected to the corresponding 110kV node, which is regarded as an uncontrollable power source.
[0098] DC lines: A model was created for DC lines in the power grid, and information such as voltage level, power transmission capacity and start and end nodes of the DC lines was set.
[0099] Through comprehensive modeling of these six key equipment categories, the power flow simulation model can accurately simulate the actual operation of the power system, supporting multiple functions such as power flow calculation and optimal power flow. Power flow calculation helps analyze the power distribution and voltage status in the power system, laying the foundation for subsequent simulation of reinforcement learning agents interacting in the environment.
[0100] S14. Based on the historical operation information, construct the operation scenario set of the power system;
[0101] This application embodiment constructs an operational scenario set for the power system based on historical operational information, such as historical time-series data of wind power, photovoltaic power output, and load electricity consumption. This operational scenario set comprehensively reflects the operational status and power supply and demand characteristics under different seasons and time conditions. The main process is as follows:
[0102] Construct day-ahead data: integrate maintenance status of various equipment, day-ahead plans of generator sets, and day-ahead forecast information of new energy sources and loads;
[0103] Constructing intraday phase data: Integrating intraday forecast information for new energy sources and load;
[0104] Constructing real-time phased data: Integrating real-time data on new energy power generation and load power consumption;
[0105] The above data forms the original operating scenario set of the power system. The operating scenarios in the original operating scenario set are then verified and corrected to obtain the final operating scenario set of the power system.
[0106] Verification and Correction: At the three time scales of day-ahead, intraday, and real-time, check whether the unit output under each original operating scenario set is within the upper and lower limits of the unit's allowed output, and correct the part exceeding the upper and lower limits back to the upper and lower limits boundary; check whether the unit output between adjacent time segments meets the ramping constraint, and correct the part that does not meet the ramping constraint back to the upper and lower limits of the ramping constraint; for operating scenarios with minor imbalance, correction is carried out by unit proportional allocation, and operating scenarios with large imbalance are deleted, finally forming the operating scenario set of the power system.
[0107] S15. Based on the power flow calculation simulation model and the set of operating scenarios, construct a virtual simulation environment for the power system using the Python language.
[0108] This application embodiment constructs a virtual simulation environment for a power system, supporting functions such as control strategy execution, system state updating, observation information return, and reward function calculation.
[0109] Control strategy execution: When the intelligent agent outputs intelligent control actions, the virtual simulation environment will execute the corresponding system scheduling instructions, such as adjusting the output of generator sets, the charging and discharging state and power of energy storage, etc., and modifying the power flow calculation simulation model.
[0110] System status update: Power flow simulation calculations are performed on the modified power flow calculation simulation model. The system status will be updated according to the modified power flow calculation simulation model. The virtual simulation environment will refresh the new system status information in real time, such as voltage changes, power flow distribution, load power consumption, etc., to provide the agent with the decision-making information for the next round.
[0111] Observational Information Return: The virtual simulation environment returns new state information and real-time observation data of the system. The agent uses this as input to feed back into the action network, enabling it to continuously adjust its strategy to adapt to the dynamic changes in the system.
[0112] Reward function calculation and feedback: The virtual simulation environment calculates reward and penalty values based on the system's operating state and feeds them back to the agent. The reward function is mainly used to evaluate the economic efficiency of control decisions, while the penalty function is mainly used to evaluate the safety of control decisions.
[0113] In an optional embodiment, the objective function of the power system dispatch model is expressed as:
[0114]
[0115] In the formula, and Let be the decision variables, representing the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t, the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t, the curtailment amount of the i-th wind power unit at time t, and the curtailment amount of the i-th photovoltaic unit at time t. For the regulation costs of the power system; Penalties for curtailment of wind and solar power in the power system;
[0116]
[0117] In the formula, N G This refers to the number of thermal power units in the power system. Let be the unit regulation cost of the i-th thermal power unit. Let be the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t. N represents the negative adjustment of the i-th thermal power unit to the day-ahead plan at time t; S The number of energy storage units in the power system. Let i be the unit regulation cost of the i-th energy storage unit. Let be the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t. Let be the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t; where these adjustments are output by the agent.
[0118]
[0119] In the formula, N Wind The number of wind turbines in the power system. The unit curtailment penalty for the i-th wind turbine is... N represents the amount of power wasted by the i-th wind turbine at time t; Solar The number of photovoltaic units in the power system. The unit curtailment penalty for the i-th photovoltaic unit is... Let represent the amount of electricity wasted by the i-th photovoltaic unit at time t.
[0120] In an optional embodiment, the power balance constraint is expressed as:
[0121]
[0122] Where, N G This refers to the number of thermal power units in the power system. N represents the real-time plan for the i-th thermal power unit at time t; Wind The number of wind turbines in the power system. N represents the real-time plan for the i-th wind turbine at time t; Solar The number of photovoltaic units in the power system. N represents the real-time plan for the i-th photovoltaic unit at time t; S The number of energy storage units in the power system. N represents the real-time plan for the i-th energy storage unit at time t; Load The number of loads in the power system. Let be the real-time demand of the i-th load at time t;
[0123] Specifically,
[0124] In the formula, Let i be the day-ahead schedule for the i-th thermal power unit at time t. Let be the positive adjustment amount of the i-th thermal power unit to the day-ahead plan at time t. Let be the negative adjustment amount of the i-th thermal power unit to the day-ahead plan at time t;
[0125] Specifically,
[0126] In the formula, For the predicted daily power generation of the i-th wind turbine at time t, Let represent the amount of power wasted by the i-th wind turbine at time t.
[0127] Specifically,
[0128] In the formula, Let i be the predicted daily power generation of the i-th photovoltaic unit at time t. Let be the amount of electricity wasted by the i-th photovoltaic unit at time t;
[0129] Specifically,
[0130] In the formula, Let i be the day-ahead plan for the i-th energy storage unit at time t. Let be the positive adjustment amount of the i-th energy storage unit to the day-ahead plan at time t. Let be the negative adjustment amount of the i-th energy storage unit to the day-ahead plan at time t;
[0131] The line transmission capacity constraint is calculated using the power transmission distribution factor, and the line transmission capacity constraint is expressed as follows:
[0132]
[0133] Where, N Bus G represents the number of busbars in the power system. i-jThe impact of injecting unit power into the j-th bus on the transmission power of the i-th line is a physical parameter of the power system. This is the lower limit of the transmission capacity of the i-th line. This represents the upper limit of the transmission capacity of the i-th line; For the real-time plan of the j-th thermal power unit at time t, For the j-th wind turbine, the real-time plan at time t is... For the real-time plan of the j-th photovoltaic unit at time t, For the j-th energy storage unit at time t, Let be the real-time demand of the j-th load at time t;
[0134] The operating constraints of the thermal power units include: thermal power unit capacity constraints and ramping constraints; the operating constraints of the thermal power units are expressed as follows:
[0135]
[0136] in, Let be the minimum allowable output power of the i-th thermal power unit. Let i be the maximum allowable output power of the i-th thermal power unit. Let be the real-time plan for the i-th thermal power unit at time t. Let be the maximum allowed climbing rate for the i-th thermal power unit;
[0137] The operating constraints of the energy storage unit include: charge / discharge power constraints, upper and lower limits of energy storage state of charge constraints, energy storage state of charge constraints, and charge / discharge state constraints; the operating constraints of the energy storage unit are expressed as follows:
[0138]
[0139] in, Let i be the minimum allowable output power of the i-th energy storage unit. For the real-time plan of the i-th energy storage unit, This represents the maximum allowable output power of the i-th energy storage unit; This represents the minimum permissible state of charge for the i-th energy storage unit. This represents the real-time state of charge of the i-th energy storage unit. This represents the maximum allowable state of charge for the i-th energy storage unit. A binary variable representing the charging state. This is a binary variable representing the discharge state. For the i-th energy storage unit at time t, ΔT represents the charge / discharge efficiency; ΔT represents the scheduling time interval. Let i be the capacity of the i-th energy storage unit;
[0140] The maximum wind and solar curtailment constraint: Considering the continuous deviation in renewable energy forecasts during the day, if dispatchable resources are exhausted and power balance cannot be met, renewable energy curtailment is permitted; the maximum wind and solar curtailment constraint is expressed as follows:
[0141]
[0142] in, Let i be the amount of abandoned power of the i-th wind turbine. For the i-th wind turbine, predict the daily power generation at time t; Let be the amount of power curtailed by the i-th photovoltaic unit at time t. The predicted daily power generation of the i-th photovoltaic unit at time t.
[0143] In an optional embodiment, the state space includes: the current time, the active power of the thermal power unit at time t, the active power of the energy storage unit at time t, the active power of the wind power unit at time t, the active power of the photovoltaic unit at time t, the active power of the load at time t, the upper adjustable space of the thermal power unit at time t, the lower adjustable space of the thermal power unit at time t, the upper adjustable space of the energy storage unit at time t, the lower adjustable space of the energy storage unit at time t, the line load rate at time t, the daily plan of the thermal power unit at time t+1, the daily plan of the energy storage unit at time t+1, the intraday predicted demand of the load at time t+1, the intraday predicted power generation of the wind power unit at time t+1, and the intraday predicted power generation of the photovoltaic unit at time t+1.
[0144] The operational space includes: the output adjustment amount of the thermal power unit in two adjacent time periods and the ratio of the maximum allowable output of the energy storage unit;
[0145] It is worth noting that this application employs a continuous operating space with an operating range of [-1,1] to continuously adjust the output of thermal power and the charging and discharging power of energy storage: in in, For the scheduling actions of the thermal power intelligent agent, For the scheduling actions of energy storage intelligent agents, This is the scheduling action of the first thermal power intelligent agent. For the Nth G The scheduling actions of a thermal power intelligent agent This is the scheduling action of the first energy storage agent. For the Nth G The scheduling actions of an energy storage intelligent agent;
[0146] For thermal power units, the intelligent agent outputs actions. Defined as the output adjustment of a thermal power unit between two adjacent time periods, i.e. In the formula, Let be the real-time plan for the i-th thermal power unit at time t. This is the action instruction for the i-th thermal power intelligent agent at time t. Let be the maximum allowed ramp rate for the i-th thermal power unit. For energy storage units, the agent outputs the action. Defined as the proportion of the maximum allowable output of the energy storage unit, i.e. In the formula, For the i-th energy storage unit at time t, This is the action command for the i-th energy storage agent at time t. Let be the maximum allowable output power of the i-th energy storage unit.
[0147] Since real-time scheduling is a constrained optimization problem, this application decouples reward and punishment signals based on constrained Markov decision processes to more accurately describe the objective function and constraints, thereby improving the learning performance and convergence of the agent.
[0148] The reward function is negatively correlated with the adjustment cost, and the reward function is expressed as:
[0149]
[0150] Where, N G This refers to the number of thermal power units in the power system. Let be the unit regulation cost of the i-th thermal power unit. Let be the real-time plan for the i-th thermal power unit at time t. Let N be the day-ahead schedule for the i-th thermal power unit at time t; S The number of energy storage units in the power system. Let i be the unit regulation cost of the i-th energy storage unit. For the i-th energy storage unit at time t, Let i be the day-ahead plan for the i-th energy storage unit at time t;
[0151] The penalty function primarily evaluates the satisfaction of the line transmission capacity constraint and the upper and lower limits of the balancing machine output constraint. The penalty function is expressed as follows:
[0152]
[0153] Where, α rho α is the weighting factor for the line over-limit penalty. slack N is the weighting factor for the balancing machine's over-limit penalty; Line For; ρ i,t Let be the load rate of the i-th line at time t; the balancing machine is a certain thermal power unit. Let be the output force of the balancing machine at time t. This is the upper limit of the output allowed by the balancing machine. This is the lower limit of the allowable output of the balancing machine.
[0154] The Secure Reinforcement Learning Algorithm (SAC-Lagrangian) in this application adopts a maximum entropy framework. In addition to maximizing the cumulative reward, it also maximizes the entropy of the policy at each time step to increase the randomness of the policy and prevent the policy from converging to a local optimum too early.
[0155] Considering the varying degrees of importance placed on entropy in different states and training phases, entropy should be relatively high for states where the optimal action is not yet determined, and relatively low for states where the optimal action is clear. Therefore, the agent's constrained optimization objective is set as follows:
[0156]
[0157] In the formula, a represents the agent's action, π represents the agent's policy, s represents the environmental state observed by the agent, and E a : π(s) Let γ represent the expectation of the agent's policy π, γ represent the discount factor, t represent time, and r represent the time interval. t s0 represents the immediate reward received by the agent, c represents the initial state of the environment, and s0 represents the initial state of the environment. t This represents the immediate penalty received by the agent, where c represents the penalty threshold. The target entropy can be preset;
[0158] Based on the primal-dual secure reinforcement learning method, the optimization objective is transformed into an unconstrained optimization problem by introducing dual variables, i.e.:
[0159]
[0160] In the formula, To augment the objective function, f(π) = E a : π(s) [∑ t γ t r t |s0=s], λ is the adaptive weighting coefficient of the agent's action safety term; β is the adaptive weighting coefficient of the agent's action random term, used to control the importance of entropy and reward in different states; π represents the agent's policy.
[0161] Construct a reward evaluation network and a penalty evaluation network; wherein the loss function of the reward evaluation network is expressed as: The loss function of the penalty evaluation network is expressed as:
[0162] In the formula, φ rTo reward the evaluation of network parameters, φ c To penalize the evaluation network parameters, τ represents the interaction trajectory sampled from the experience pool. For experience pool, This represents calculating the expectation along the sampling trajectory, Q. r (s t ,a t The reward evaluation network assesses the current state s. t The lower agent applies action a t The reward value function after that, Q c (s t ,a t The network evaluates the current state s as a penalty evaluation network. t The lower agent applies action a t The subsequent penalty value function, r t c represents the immediate reward received by the agent. t This represents the immediate penalty received by the agent, where γ is the discount factor. The network output value is evaluated based on the target reward. The network output value is used to evaluate the target penalty and reward.
[0163] It's worth noting that the reward evaluation network is used to assess the economics of the policy, while the penalty evaluation network is used to assess the safety of the policy. Furthermore, by constructing an additional target network... Estimate the target Q-value to stabilize training and improve the convergence of the algorithm.
[0164] Construct an action network; wherein the loss function of the action network is expressed as:
[0165] In the formula, φ π Here are the action network parameters, and τ is the interaction trajectory sampled from the experience pool. For experience pool, This represents calculating the expectation along the sampling trajectory, Q. r (s t ,a t The reward evaluation network assesses the current state s. t The lower agent applies action a t The reward value function is denoted by λ, which is the adaptive weight coefficient of the agent's action safety term, β, which is the adaptive weight coefficient of the agent's action random term, and π, which is the agent's policy.
[0166] The agent interacts with the virtual simulation environment. During the interaction, the agent is iteratively trained using the unconstrained optimization problem, the reward evaluation network, the penalty evaluation network, and the action network to obtain a trained agent.
[0167] This application provides a real-time scheduling method for a high-proportion renewable energy power system. It constructs a virtual simulation environment for the power system based on component information and historical operating information. Using thermal power units, energy storage units, wind power units, and photovoltaic units as control objects, and combining these with the power system's safety operation constraints, it constructs a power system scheduling model with the objective of minimizing the sum of the power system's regulation costs and wind / solar curtailment penalties. The power system scheduling model is then transformed into a Markov decision process to construct the agent's state space, action space, reward function, and penalty function. The agent interacts with the virtual simulation environment, and based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent, resulting in a well-trained agent. Using this trained agent to perform real-time scheduling of the power system effectively avoids actions that lead to unsafe outcomes during the exploration process. This ensures strategy performance while improving safety, stability, and robustness, ultimately achieving a good real-time scheduling strategy generation capability for a high-proportion renewable energy power system.
[0168] See Figure 2 , Figure 2 This is a structural block diagram of a real-time dispatching device 10 for a high-proportion renewable energy power system provided in an embodiment of this application. The real-time dispatching device for the high-proportion renewable energy power system includes:
[0169] The simulation environment construction module 11 is used to construct a virtual simulation environment for the power system based on the component information and historical operation information of the power system.
[0170] The model building module 12 is used to construct a power system dispatch model with the objectives of minimizing the sum of the power system's regulation cost and wind / solar curtailment penalties, taking thermal power units, energy storage units, wind power units, and photovoltaic units as the control objects and combining them with the power system's safe operation constraints. The safe operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints, and maximum wind / solar curtailment constraints.
[0171] The agent construction module 13 is used to convert the power system scheduling model into a Markov decision process to construct the agent's state space, action space, reward function and penalty function;
[0172] The agent training module 14 is used to enable the agent to interact with the virtual simulation environment, and to iteratively train the agent using a safe reinforcement learning algorithm based on the set constrained optimization objective of the agent to obtain a trained agent.
[0173] The real-time scheduling module 15 is used to perform real-time scheduling of the power system using a trained intelligent agent.
[0174] It is worth noting that the working process of each module in the real-time dispatching device 10 of the high-proportion renewable energy power system described in this application embodiment can refer to the working process of the real-time dispatching method of the high-proportion renewable energy power system described in the above embodiment, and will not be repeated here.
[0175] This application provides a real-time dispatching device 10 for a high-proportion renewable energy power system. It constructs a virtual simulation environment for the power system based on component information and historical operating information. Using thermal power units, energy storage units, wind power units, and photovoltaic units as control objects, and combining the power system's safety operation constraints, it constructs a power system dispatching model with the objective of minimizing the sum of the power system's regulation cost and the penalty for wind and solar curtailment. The power system dispatching model is then converted into a Markov decision process to construct the agent's state space, action space, reward function, and penalty function. The agent interacts with the virtual simulation environment, and based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent, resulting in a well-trained agent. Using the trained agent to perform real-time dispatching of the power system effectively avoids actions that lead to unsafe outcomes during the exploration process. This ensures strategy performance while improving safety, stability, and robustness, ultimately achieving a good real-time dispatching strategy generation capability for a high-proportion renewable energy power system.
[0176] Furthermore, this application also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the real-time scheduling method for a high-proportion renewable energy power system as described in any of the above embodiments.
[0177] Furthermore, this application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the real-time scheduling method for a high-proportion renewable energy power system as described in any of the above embodiments.
[0178] See Figure 3 , Figure 3 This is a structural block diagram of a real-time dispatching device 20 for a high-proportion renewable energy power system provided in this application embodiment. The real-time dispatching device 20 for a high-proportion renewable energy power system includes: a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, it implements the steps in the above-described real-time dispatching method embodiment for a high-proportion renewable energy power system. Alternatively, when the processor 21 executes the computer program, it implements the functions of each module / unit in the above-described device embodiments.
[0179] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 22 and executed by the processor 21 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the real-time dispatching device 20 of the high-proportion renewable energy power system.
[0180] The real-time dispatching device 20 for the high-proportion renewable energy power system may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will understand that the schematic diagram is merely an example of the real-time dispatching device 20 for the high-proportion renewable energy power system and does not constitute a limitation on the device. It may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the real-time dispatching device 20 for the high-proportion renewable energy power system may also include input / output devices, network access devices, buses, etc.
[0181] The processor 21 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 21 is the control center of the real-time dispatching equipment 20 of the high-proportion renewable energy power system, connecting various parts of the real-time dispatching equipment 20 of the entire high-proportion renewable energy power system through various interfaces and lines.
[0182] The memory 22 can be used to store the computer programs and / or modules. The processor 21 implements various functions of the real-time dispatching equipment 20 of the high-proportion new energy power system by running or executing the computer programs and / or modules stored in the memory 22 and calling the data stored in the memory 22. The memory 22 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0183] The modules / units integrated into the real-time dispatching equipment 20 of the high-proportion new energy power system, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by the processor 21, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0184] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0185] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A real-time dispatching method for a high-proportion renewable energy power system, characterized in that, include: Based on the component information and historical operation information of the power system, a virtual simulation environment for the power system is constructed. Using thermal power units, energy storage units, wind power units, and photovoltaic units as the control targets, and combining the safety operation constraints of the power system, a power system dispatch model is constructed with the objective of minimizing the sum of the power system's regulation costs and wind and solar curtailment penalties. The safety operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints, and maximum wind and solar curtailment constraints. The power system scheduling model is transformed into a Markov decision process to construct the agent's state space, action space, reward function, and penalty function. The agent interacts with the virtual simulation environment. Based on the set constrained optimization objective of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent. The trained intelligent agent is used to schedule the power system in real time. The interaction between the agent and the virtual simulation environment, and the iterative training of the agent using a safe reinforcement learning algorithm based on a set constrained optimization objective, to obtain a trained agent, includes: Define a constrained optimization objective for the agent: ; ; In the formula, Indicates the action of the intelligent agent. Represents the agent's policy. This represents the environmental state observed by the agent. Indicates the policy of the agent Seeking expectations, Indicates the discount factor. Indicates time, This represents the immediate reward received by the intelligent agent. Indicates the initial state of the environment. This represents the immediate punishment received by the intelligent agent. Indicates the penalty threshold. Represents the target entropy; The optimization objective is transformed into an unconstrained optimization problem, namely: In the formula, To broaden the objective function, , , ; For the adaptive weighting coefficients of the agent's action safety term; For the adaptive weighting coefficients of the random terms of the agent's actions; This represents the agent's policy.
2. The real-time dispatching method for a high-proportion renewable energy power system as described in claim 1, characterized in that, The objective function of the power system dispatch model is expressed as: ; In the formula, and Let be the decision variables, representing the positive adjustment of the i-th thermal power unit to the day-ahead plan at time t, and the th... i Each thermal power unit t The negative adjustment amount of the current day's plan at all times, the first i One energy storage unit in t The amount of positive adjustment to the current day's plan at all times, the first i One energy storage unit in t The negative adjustment amount of the current day's plan at all times, the first i Each wind turbine unit t The amount of power wasted at each moment and the first i A photovoltaic unit in t The amount of electricity wasted at any given moment; For the regulation costs of the power system; Penalties for curtailment of wind and solar power in the power system; ; In the formula, This refers to the number of thermal power units in the power system. For the first i The unit regulation cost of a thermal power unit For the first i Each thermal power unit t The amount of positive adjustment to the current plan at all times. For the first i Each thermal power unit t The amount of negative adjustment to the current day's plan at any time; The number of energy storage units in the power system. For the first i The unit regulation cost of an energy storage unit For the first i One energy storage unit in t The amount of positive adjustment to the current plan at all times. For the first i One energy storage unit in t The amount of negative adjustment to the current day's plan at any time; ; In the formula, The number of wind turbines in the power system. For the first i Unit curtailment penalty for wind turbines For the first i Each wind turbine unit t The amount of electricity wasted at any given moment; The number of photovoltaic units in the power system. For the first i Unit curtailment penalty for photovoltaic power generation For the first i A photovoltaic unit in t The amount of electricity wasted at any given moment.
3. The real-time dispatching method for a high-proportion renewable energy power system as described in claim 1, characterized in that, The power balance constraint is expressed as: ; in, This refers to the number of thermal power units in the power system. For the first i Each thermal power unit t Real-time planning at any given moment; The number of wind turbines in the power system. For the first i Each wind turbine unit t Real-time planning at any given moment; The number of photovoltaic units in the power system. For the first i A photovoltaic unit in t Real-time planning at any given moment; The number of energy storage units in the power system. For the first i One energy storage unit in t Real-time planning at any given moment; The number of loads in the power system. For the first i A load in t Real-time requirements at any given moment; The line transmission capacity constraint is expressed as follows: ; in, This refers to the number of busbars in the power system. For the first j The unit injected power of the busbar is related to the first i The impact of line transmission power; For the first i The lower limit of the transmission capacity of a single line. For the first i The upper limit of the transmission capacity of each line; For the first j Each thermal power unit t Real-time planning at any moment For the first j Each wind turbine unit t Real-time planning at any moment For the first j A photovoltaic unit in t Real-time planning at any moment For the first j One energy storage unit in t Real-time planning at any moment For the first j A load in t Real-time requirements at any given moment; The operating constraints of the thermal power unit are expressed as follows: ; ; in, For the first i The minimum allowable output power of a thermal power unit. For the first i The maximum allowable output power of each thermal power unit For the first i Each thermal power unit t Real-time planning at any moment For the first i The maximum permissible ramp rate for each thermal power unit; The operating constraints of the energy storage unit are expressed as follows: ; ; ; ; in, For the first i The minimum allowable output power of each energy storage unit. For the first i Real-time planning for each energy storage unit For the first i The maximum allowable output power of each energy storage unit; For the first i The minimum allowable state of charge for each energy storage unit For the first i Real-time status of charge of each energy storage unit For the first i The maximum allowable state of charge for each energy storage unit; A binary variable representing the charging state. This is a binary variable representing the discharge state. For the first i One energy storage unit in t Real-time planning at any moment For charge and discharge efficiency; For scheduling time intervals, For the first i The capacity of each energy storage unit; The maximum wind and solar curtailment constraint is expressed as: ; ; in, For the first i The amount of electricity wasted by each wind turbine unit For the first i Each wind turbine unit t Forecasted daily power generation at any given time; For the first i A photovoltaic unit in t The amount of electricity wasted at any given moment For the first i A photovoltaic unit in t Forecasted daily power generation at any given time.
4. The real-time dispatching method for a high-proportion renewable energy power system as described in claim 1, characterized in that, The state space includes: the current time, t Active power of thermal power units at any given time t The active power of the energy storage unit at any given time t The active power of the wind turbine at any given time t Active power of photovoltaic units at any given time t Active power of load at any given time t Adjustable space of thermal power units at all times t The adjustable space of thermal power units at any time t The adjustable space of the energy storage unit at any time t The adjustable range of the energy storage unit at any time t Line load rate at any time t The daily schedule of thermal power units at +1 hour. t +1 hour energy storage unit day-ahead schedule, t The intraday forecast demand for load at +1 time. t The predicted daily power generation of the wind turbine at time +1 t The daily forecast power generation of the photovoltaic unit at time +1; The operational space includes: the output adjustment amount of the thermal power unit in two adjacent time periods and the ratio of the maximum allowable output of the energy storage unit; The reward function is expressed as follows: ; in, This refers to the number of thermal power units in the power system. For the first i The unit regulation cost of a thermal power unit For the first i Each thermal power unit t Real-time planning at any moment For the first i Each thermal power unit t The current day's schedule; The number of energy storage units in the power system. For the first i The unit regulation cost of an energy storage unit For the first i One energy storage unit in t Real-time planning at any moment For the first i One energy storage unit in t The current day's schedule; The penalty function is expressed as follows: ; in, The weighting factor for the line exceeding the limit penalty. The weighting factor for the penalty of exceeding the limit by the balancing machine; The number of lines in the power system; For the first i The lines are in t Load rate at any given time; For balancing machine t Constant effort This is the upper limit of the output allowed by the balancing machine. This is the lower limit of the allowable output of the balancing machine.
5. The real-time dispatching method for a high-proportion renewable energy power system as described in claim 1, characterized in that, The step of interacting the intelligent agent with the virtual simulation environment, and iteratively training the intelligent agent using a safe reinforcement learning algorithm based on a set constrained optimization objective to obtain a trained intelligent agent, further includes: Construct a reward evaluation network and a penalty evaluation network; wherein the loss function of the reward evaluation network is expressed as: The loss function of the penalty evaluation network is expressed as: ; In the formula, To reward the evaluation of network parameters, To penalize the evaluation network parameters, Interaction trajectories sampled from the experience pool, For experience pool, This means calculating the expectation along the sampling trajectory. To reward the network's assessment of the current state The lower agent applies actions The subsequent reward value function, To penalize the evaluation network for assessing the current state The lower agent applies actions The subsequent penalty value function, This represents the immediate reward received by the intelligent agent. This represents the immediate punishment received by the intelligent agent. As a discount factor, The network output value is evaluated based on the target reward. The network output value is used to evaluate the target penalty and reward. Construct an action network; wherein the loss function of the action network is expressed as: ; In the formula, For action network parameters, Interaction trajectories sampled from the experience pool, For experience pool, This means calculating the expectation along the sampling trajectory. To reward the network's assessment of the current state The lower agent applies actions The subsequent reward value function, For the adaptive weighting coefficients of the agent's action safety term, For the adaptive weighting coefficients of the random term of the agent's actions, For agent policies; The agent interacts with the virtual simulation environment. During the interaction, the agent is iteratively trained using the unconstrained optimization problem, the reward evaluation network, the penalty evaluation network, and the action network to obtain a trained agent.
6. The real-time dispatching method for a high-proportion renewable energy power system as described in claim 1, characterized in that, The construction of a virtual simulation environment for the power system based on its component information and historical operating information includes: Obtain the raw power grid data of the power system; The raw power grid data is parsed and preprocessed to obtain component information and historical operation information; Based on the component information, a power flow calculation simulation model of the power system is constructed; Based on the historical operating information, a set of operating scenarios for the power system is constructed; Based on the power flow calculation simulation model and the set of operating scenarios, a virtual simulation environment for the power system is constructed.
7. A real-time dispatching device for a high-proportion renewable energy power system, characterized in that, include: The simulation environment construction module is used to construct a virtual simulation environment for the power system based on the component information and historical operating information of the power system. The model building module is used to construct a power system dispatch model with the objectives of minimizing the sum of the power system's regulation costs and wind / solar curtailment penalties, taking thermal power units, energy storage units, wind power units, and photovoltaic units as the control objects and combining them with the power system's safe operation constraints. The safe operation constraints include: power balance constraints, line transmission capacity constraints, thermal power unit operation constraints, energy storage unit operation constraints, and maximum wind / solar curtailment constraints. The agent construction module is used to convert the power system scheduling model into a Markov decision process to construct the agent's state space, action space, reward function, and penalty function. The agent training module is used to enable the agent to interact with the virtual simulation environment. Based on the set constrained optimization objectives of the agent, a safe reinforcement learning algorithm is used to iteratively train the agent to obtain a trained agent. A real-time scheduling module is used to schedule the power system in real time using a trained intelligent agent; The interaction between the agent and the virtual simulation environment, and the iterative training of the agent using a safe reinforcement learning algorithm based on a set constrained optimization objective, to obtain a trained agent, includes: Define a constrained optimization objective for the agent: ; ; In the formula, Indicates the action of the intelligent agent. Represents the agent's policy. This represents the environmental state observed by the agent. Indicates the policy of the agent Seeking expectations, Indicates the discount factor. Indicates time, This represents the immediate reward received by the intelligent agent. Indicates the initial state of the environment. This represents the immediate punishment received by the intelligent agent. Indicates the penalty threshold. Represents the target entropy; The optimization objective is transformed into an unconstrained optimization problem, namely: In the formula, To broaden the objective function, , , ; For the adaptive weighting coefficients of the agent's action safety term; For the adaptive weighting coefficients of the random terms of the agent's actions; This represents the agent's policy.
8. A real-time dispatching device for a high-proportion renewable energy power system, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the real-time scheduling method for a high-proportion renewable energy power system as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the real-time scheduling method for a high-proportion new energy power system as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes a computer program / instruction that, when executed by a processor, implements the real-time scheduling method for a high-proportion renewable energy power system as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Active power distribution network real-time scheduling method and device based on safety reinforcement learning
CN115714382A