Multi-energy power system scheduling method, system, and device based on reinforcement learning

By building a multi-energy power system scheduling model and using reinforcement learning and opposing learning optimization algorithms to generate optimal scheduling solutions, the power system scheduling challenges brought about by new energy grid connection are solved, and the reduction of power grid net load fluctuations and stable absorption of new energy is achieved.

CN120013198BActive Publication Date: 2025-07-22YANTAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510457557.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-22
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The uncertainty and volatility of output caused by the high proportion of new energy grid connections has brought challenges to power system scheduling, especially the difficulty of flexibly adjusting thermal power units, resulting in difficulty in power balance, widening of the net load peak-to-valley difference, and increasing the cost of deep peak-shaving and auxiliary service of thermal power units.

Method used

Build a multi-energy power system scheduling model, adopt reinforcement learning and opposing learning to initialize populations, learn optimized behavior strategies through reinforcement learning optimization algorithms, generate optimal scheduling solutions, use energy storage systems to balance new energy output, and reduce power grid net load fluctuations.

Benefits of technology

It effectively reduces the net load fluctuations of the power grid, improves the capacity for new energy consumption, stabilizes the operation of the power system, and reduces the peak shaving pressure and auxiliary service costs of thermal power units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_10
    Figure QLYQS_10
  • Figure QLYQS_23
    Figure QLYQS_23
  • Figure QLYQS_30
    Figure QLYQS_30
Patent Text Reader

Abstract

The present invention belongs to the technical field of power system dispatching, and specifically relates to a multi-energy power system dispatching method, system, and device based on reinforcement learning. Aiming at the problem of multi-energy power system dispatching, the present invention first constructs a multi-energy power system dispatching model, with the goal of minimizing the variance of the remaining load fluctuation, and uses energy storage to absorb more wind and solar power output; then designs an optimization algorithm to obtain a dispatching plan. Among them, opposition-based learning is used to initialize the population, improve the diversity of the initial population, which is beneficial to avoiding the premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution; the Q value of reinforcement learning is introduced to learn the optimal behavior strategy of the optimization algorithm and improve the effectiveness of the optimization algorithm to obtain a better dispatching plan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system dispatching, and particularly relates to a multi-energy power system dispatching method, system, and device based on reinforcement learning. Background Art

[0002] At present, a new power system with new energy as the main body is accelerating its formation. Among them, the installed capacity ratio and power generation ratio of wind power and solar power generation will gradually increase. However, although the access of a high proportion of new energy has alleviated the environmental pressure on the power system to a certain extent, the randomness and volatility of its output have brought unprecedented challenges to the dispatching operation of the power system. The uncertainty and unpredictability of new energy output make the fluctuation of the grid net load intensify. Especially for a power supply structure mainly based on thermal power, lacking a power source that can flexibly adjust with the fluctuation of new energy, it will make it extremely difficult to achieve power balance when a high proportion of new energy is connected to the grid. If full consumption of new energy is to be achieved, the peak-valley difference of the grid net load will expand, and thermal power units will face the heavy burden of deep peak shaving. During the low-load period, the auxiliary service costs borne by new energy units will also increase sharply.

[0003] Therefore, the output uncertainty brought by the grid connection of a high proportion of wind and light seriously affects the consumption of new energy and the stable operation of the power system. Summary of the Invention

[0004] The present invention provides a multi-energy power system dispatching method, system, and device based on reinforcement learning.

[0005] The technical solution of the present invention is as follows:

[0006] The present invention provides a multi-energy power system dispatching method based on reinforcement learning, including the following steps:

[0007] S1: Construct a thermal power model according to the output energy and ramp rate limit of thermal power units during dispatching;

[0008] Construct a wind and light model according to the output limit of wind turbines and the output power limit of photovoltaic power generation systems;

[0009] Construct a energy storage model according to the thermal power output power, actual wind power output value, actual photovoltaic output value, charge and discharge power of energy storage batteries, initial grid load value, and rated capacity of energy storage power stations;

[0010] Determine an objective function according to the thermal power model, wind and light model, and energy storage model, and construct a multi-energy power system dispatching model;

[0011] S2: Based on the multi-energy power system dispatching model, initialize individuals in a random and opposition-based learning manner to obtain a population;

[0012] Initialize the parameters of reinforcement learning, including the Q-table and reward values;

[0013] S3: Use the reinforcement learning method to learn the individual behaviors and roles of the population. Calculate the Q-value of each behavior according to the current role and optional behaviors of the individual, select and execute the behavior with the maximum Q-value from the Q-table, and generate a new individual position according to the selected role and behavior;

[0014] S4: Calculate the fitness value of the new individual position. If the fitness value of the new position is greater than that of the original position, the individual is updated from the original position to the new position;

[0015] S5: After updating the parameters according to the fitness value of the new individual position, execute S3 until the preset condition is reached to obtain the optimal individual position as the scheduling scheme.

[0016] In the said S1, according to the output energy and ramp rate limit of the thermal power unit during scheduling, a thermal power model is constructed as follows according to the formula: , realized as;

[0017] In the formula, , respectively represent the output energy of the thermal power unit at moment and moment, is the ramp rate limit of the thermal power unit j during scheduling.

[0018] In the said S1, according to the output limit of the wind power unit and the output power limit of the photovoltaic power generation system, a wind-solar model is constructed as follows according to the formula: , realized as;

[0019] In the formula, is the actual wind power output value at moment, , respectively are the minimum and maximum output values of the wind power unit at moment, is the actual photovoltaic output value at moment, , respectively are the minimum and maximum output power values of the photovoltaic power generation system at moment.

[0020] In the said S1, according to the thermal power output, actual wind power output value, actual photovoltaic output value, charge and discharge power of the energy storage battery, and the initial grid load value, and the rated capacity of the energy storage power station, an energy storage model is constructed as follows according to the formula: , and , realized as;

[0021] In the formula, The thermal power output power at time is the actual wind power output value at time is the actual photovoltaic output value at time is the discharge power of the energy storage battery at time is the charging power of the energy storage battery at time is the initial grid load value at time is the rated capacity of the energy storage power station is the scheduling period

[0022] In S1, the objective function is the minimum variance of the grid net load before peak shaving trading; the net load is the generated electricity borne only by thermal power after reducing the combined output of wind, light, and storage in the multi - energy power system

[0023] The expression of the objective function is ;

[0024] Among them, ; ;

[0025] In the formula, represents the objective function, is the scheduling period, is the grid net load at time before peak shaving trading, is the average value of the net load before peak shaving trading, , , , , are respectively the initial grid load value, the actual wind power output value, the actual photovoltaic output value, the charging power of the energy storage battery, the discharge power of the energy storage battery at time , are respectively the charging efficiency and discharge efficiency of the energy storage battery

[0026] In S2, the individuals are initialized by using the random and opposition - based learning method, that is, the first preset number of individuals are generated by the random method, and the second preset number of individuals are generated by the opposition - based learning method

[0027] In S3, the individual behaviors and roles of the population are learned by using the reinforcement learning method. The individual behaviors are local search and global search, and the roles are the leader and the follower

[0028] The present invention also provides a multi - energy power system scheduling system based on reinforcement learning, including:

[0029] Multi - energy power system scheduling model construction module: used to construct a thermal power model according to the output energy and ramp rate limit of thermal power units during scheduling;

[0030] Construct a wind - solar model according to the output limit of wind turbines and the output power limit of photovoltaic power generation systems;

[0031] Construct a energy storage model according to the thermal power output power, actual wind power output value, actual photovoltaic power output value, charge - discharge power of energy storage batteries, initial grid load value, and rated capacity of the energy storage power station;

[0032] Determine the objective function according to the thermal power model, wind - solar model and energy storage model, and construct a multi - energy power system scheduling model;

[0033] Initialization module: Based on the multi - energy power system scheduling model, initialize individuals in a random and opposition - based learning manner to obtain a population;

[0034] Initialize the parameters of reinforcement learning, including the Q - table and reward value;

[0035] Individual position generation module: used to learn the individual behaviors and roles of the population using reinforcement learning methods, calculate the Q - value of each behavior according to the current role and optional behaviors of the individual, select and execute the behavior with the maximum Q - value from the Q - table, and generate a new individual position according to the selected role and behavior;

[0036] Individual position update module: used to calculate the fitness value of the new individual position. If the fitness value of the new position is greater than that of the original position, the individual is updated from the original position to the new position;

[0037] Scheduling plan generation module: used to update the parameters according to the fitness value of the new individual position, then enter the individual position generation module until a preset condition is reached, obtain the optimal individual position as the scheduling plan.

[0038] The present invention also provides a multi - energy power system scheduling device based on reinforcement learning, including a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the multi - energy power system scheduling method based on reinforcement learning.

[0039] Beneficial effects

[0040] For the problem of multi - energy power system scheduling, the present invention first constructs a multi - energy power system scheduling model, aiming at minimizing the variance of the remaining load fluctuation, and utilizes energy storage to absorb more wind and photovoltaic power output; then designs an optimization algorithm to obtain a scheduling scheme. Among them, opposition - based learning is used to initialize the population, improving the diversity of the initial population, which is beneficial to avoiding premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution; the Q - value of reinforcement learning is introduced to learn the optimal behavior strategy of the optimization algorithm, improving the effectiveness of the optimization algorithm to obtain a better scheduling scheme. Detailed implementation mode

[0041] The following embodiments are intended to illustrate the present invention rather than further limit the present invention.

[0042] The present invention provides a multi - energy power system scheduling method based on reinforcement learning, including the following steps:

[0043] S1: Construct a thermal power model according to the output energy and ramp rate limit of the thermal power unit during scheduling;

[0044] Construct a wind - solar model according to the output power limit of the wind turbine and the output power limit of the photovoltaic power generation system;

[0045] Construct an energy storage model according to the thermal power output power, the actual wind power output value, the actual photovoltaic power output value, the charge - discharge power of the energy storage battery, the initial grid load value, and the rated capacity of the energy storage power station;

[0046] Determine the objective function according to the thermal power model, the wind - solar model and the energy storage model, and construct a multi - energy power system scheduling model.

[0047] Among them, constructing a thermal power model according to the output energy and ramp rate limit of the thermal power unit during scheduling is realized according to the formula: , as the ramp constraint condition;

[0048] In the formula, , respectively represent the output energy of the thermal power unit at time and time , is the ramp rate limit of the thermal power unit j during scheduling.

[0049] Preferably, constructing a wind - solar model according to the output power limit of the wind turbine and the output power limit of the photovoltaic power generation system is realized according to the formula: , as the new - energy output constraint condition;

[0050] In the formula, is the actual wind power output value at time , , are respectively the minimum and maximum output powers of the wind turbine at a certain moment, is the actual output value of the photovoltaic power at a certain moment, , are respectively the minimum and maximum output powers of the photovoltaic power generation system at a certain moment.

[0051] Preferably, according to the thermal power output power, the actual wind power output value, the actual photovoltaic power output value, the charge and discharge power of the energy storage battery, the initial grid load value, and the rated capacity of the energy storage power station, an energy storage model is constructed, which is based on the formulas: , and , to be realized, and are respectively used as the system supply-demand balance constraint condition and the energy storage operation constraint condition;

[0052] In the formula, is the thermal power output power at a certain moment, is the actual wind power output value at a certain moment, is the actual photovoltaic power output value at a certain moment, is the discharge power of the energy storage battery at a certain moment, is the charge power of the energy storage battery at a certain moment, is the initial grid load value at a certain moment, is the rated capacity of the energy storage power station, is the scheduling period.

[0053] In addition, in addition to the above 4 constraint conditions, the objective function is the minimum variance of the grid net load before the peak shaving transaction; the net load is the power generation borne only by the thermal power after the combined output of the wind, light, storage, and savings in the multi-energy power system is reduced. The variance of the net load reflects the fluctuation of the load borne by the thermal power unit. The smaller the variance, the smaller the fluctuation range of the load, and the more stable the operation mode of the corresponding thermal power unit.

[0054] Furthermore, the expression of the objective function is: ;

[0055] Among them, ; ;

[0056] In the formula, represents the objective function, is the scheduling period, is before the peak shaving transaction the grid net load at a certain moment, is the average value of the net load before peak shaving trading, , , , , are respectively the initial grid load value, the actual wind power output value, the actual photovoltaic power output value, the charging power of the energy storage battery, and the discharging power of the energy storage battery at the moment, , are respectively the charging efficiency and discharging efficiency of the energy storage battery.

[0057] After constructing the multi - energy power system scheduling model, reinforcement learning is used to obtain the optimal scheduling plan, and the operation is as follows:

[0058] S2: Based on the multi - energy power system scheduling model, initialize individuals in a random and opposition - based learning manner to obtain a population;

[0059] Initialize the parameters of reinforcement learning, including the Q - table and the reward value.

[0060] Preferably, initialize individuals in a random and opposition - based learning manner, that is, generate the first preset number of individuals in a random manner and generate the second preset number of individuals in an opposition - based learning manner.

[0061] For example, for the problem of multi - energy power system scheduling, initialize a population composed of N individuals, and each individual represents a scheduling operation strategy for the scheduling problem. N / 2 individuals can be generated in a random manner, and the other N / 2 individuals are generated in an opposition - based learning manner.

[0062] The formula for opposition - based learning is as follows:

[0063] ;

[0064] In the formula, is the j - th dimension of the position of the i - th individual generated in an opposition - based learning manner, is the j - th dimension of the position of the i - th individual generated randomly, . , are respectively the lower bound and upper bound of the search space.

[0065] The present invention uses opposition - based learning to initialize the population, improves the diversity of the initial population, is beneficial to avoiding premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution.

[0066] S3: Use the reinforcement learning method to learn the individual behaviors and roles of the population, calculate the Q - value of each behavior according to the current role and optional behaviors of the individual, select and execute the behavior with the maximum Q - value from the Q - table, and generate a new individual position according to the selected role and behavior.

[0067] Preferably, the individual behaviors and roles of the population are learned using a reinforcement learning method. The individual behaviors are local search and global search, and the roles are leader and follower.

[0068] The position of the leader individual is expressed as follows:

[0069] ;

[0070] In the formula, is the position of the leader individual, is the position of an individual in the population generated by using random and opposition-based learning, represents the j-th dimensional value of the optimal individual in the current iteration, is a random number between 0 and 1.

[0071] The position of the follower individual is expressed as follows:

[0072] ;

[0073] ;

[0074] In the formula, is the position of the follower individual, refers to the j-th dimensional value of the average position information of a randomly selected individual, are random numbers between 0 and 1 respectively, is the current iteration number, is the maximum iteration number.

[0075] The position of an individual for global search is expressed as follows:

[0076] ;

[0077] ;

[0078] In the formula, is the position of the individual using global search, is the position information of a randomly generated predator, is a random number generated by the levy distribution, is a random number between 2 and 4, is a random number between 1 and 1.5, is a random number between 2 and 3, is a random number between -1 and 1, is a random vector, is the distance between the individual and the randomly generated predator, 、 They are the fitness values of the individual and the predator respectively.

[0079] The position for the individual to perform local search is expressed as follows:

[0080] ;

[0081] ;

[0082] In the formula, is the position of the individual adopting local search, , are respectively the lower bound and the upper bound of the local search range, is a random number between 0 and 1.

[0083] Regarding learning the individual behavior of the population using the reinforcement learning method, specifically:

[0084] According to the formula: , to proceed;

[0085] In the formula, , respectively represent the current state and behavior; represents the next state; represents the discount factor, and its value range is (0, 1); represents the learning rate, and its value range is (0, 1); is the total cumulative reward; is the total cumulative reward of the next generation; is to select the optimal action in the next state of the Q value; is the immediate reward obtained by executing the behavior , and is defined as follows:

[0086] ;

[0087] Among them, represents the th individual; represents the current iteration number; is the fitness value of the th individual in the population; is the reward decay factor to prevent the cumulative reward from increasing infinitely, resulting in overly greedy behavior selection.

[0088] The present invention selects the optimal individual update method of the optimization algorithm according to the state and reward of the reinforcement learning method, and updates the state and reward according to the action result selected by the individual, so that the action that can obtain the maximum reward is taken in any state.

[0089] S4: Calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than that of the original position, the individual is updated from the original position to the new position; otherwise, the individual remains at the original position unchanged.

[0090] S5: After updating the parameters according to the fitness value of the individual's new position (including updating the Q value of the current state in the Q table according to the fitness value to improve subsequent decisions and obtain rewards), execute S3 until a preset condition (such as reaching the maximum number of iterations) is met, obtain the optimal individual position as the scheduling scheme. And the obtained optimal fitness value represents the expected value of the variance of the net load of the power grid before peak shaving trading.

[0091] For the problem of multi - energy power system scheduling, the present invention first constructs a multi - energy power system scheduling model with the goal of minimizing the variance of the remaining load fluctuation, and uses energy storage to absorb more wind and photovoltaic power output; then designs an optimization algorithm to obtain the scheduling scheme. Among them, opposition - based learning is used to initialize the population to improve the diversity of the initial population, which is beneficial to avoiding premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution; the Q value of reinforcement learning is introduced to learn the optimal behavior strategy of the optimization algorithm and improve the effectiveness of the optimization algorithm to obtain a better scheduling scheme.

[0092] The present invention also provides a multi - energy power system scheduling system based on reinforcement learning, including:

[0093] Multi - energy power system scheduling model construction module: used to construct a thermal power model according to the output energy and ramp rate limit of thermal power units during scheduling;

[0094] Construct a wind - solar model according to the output limit of wind turbines and the output power limit of photovoltaic power generation systems;

[0095] Construct an energy storage model according to the thermal power output, actual wind power output value, actual photovoltaic power output value, charge - discharge power of energy storage batteries, initial load value of the power grid, and rated capacity of the energy storage power station;

[0096] Determine the objective function according to the thermal power model, wind - solar model and energy storage model, and construct a multi - energy power system scheduling model;

[0097] Initialization module: Based on the multi - energy power system scheduling model, initialize individuals in a random and opposition - based learning manner to obtain a population;

[0098] Initialize the parameters of reinforcement learning, including the Q table and reward value;

[0099] Individual position generation module: used to learn the individual behaviors and roles of the population using reinforcement learning methods, calculate the Q value of each behavior according to the current role and optional behaviors of the individual, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position according to the selected role and behavior.

[0100] Individual position update module: It is used to calculate the fitness value of the new position of an individual. If the fitness value of the new position is greater than that of the original position, the individual is updated from the original position to the new position.

[0101] Scheduling scheme generation module: After updating the parameters according to the fitness value of the new position of the individual, it enters the individual position generation module until the preset conditions are met, and the optimal individual position is obtained as the scheduling scheme.

[0102] The present invention also provides a multi-energy power system scheduling device based on reinforcement learning, including a processor and a memory. Wherein, when the processor executes the computer program stored in the memory, the multi-energy power system scheduling method based on reinforcement learning as described above is implemented.

Claims

1. A multi-energy power system scheduling method based on reinforcement learning, characterized in that, Including: S1: According to the output energy and ramp rate limit of a thermal power unit during scheduling, build a thermal power model using the formula: , where , represent the output energy of the thermal power unit at time and time respectively, and is the ramp rate limit of the thermal power unit during scheduling; According to the output limit of the wind turbine and the output power limit of the photovoltaic power generation system, the formula: is used to construct the wind-solar model; in the formula, is the actual output value of the wind power at time , are respectively the minimum and maximum output values of the wind turbine at time is the actual output value of the photovoltaic power at time , are respectively the minimum and maximum output power values of the photovoltaic power generation system at time According to the thermal power output , the actual output value of wind power , the actual output value of photovoltaic , the discharge power of the energy storage battery , the charging power of the energy storage battery and the initial load value of the power grid , and the rated capacity of the energy storage power station , from the formula: , and , construct an energy storage model; in the formula, is the scheduling period; Determine the objective function according to the thermal power model, wind-solar model and energy storage model, and construct a multi-energy power system scheduling model; The expression of the objective function is as follows: ; Among them, ; ; In the formula, represents the objective function, is the net grid load before the peaking transaction at the moment, is the average value of the net load before the peaking transaction, and are the charging efficiency and discharging efficiency of the energy storage battery respectively; S2: Based on the multi-energy power system scheduling model, initialize individuals in a random and opposition-based learning manner to obtain a population; Initialize the parameters of reinforcement learning, including the Q-table and the reward value; S3: Use the reinforcement learning method to learn the individual behaviors and roles of the population. The individual behaviors are local search and global search, and the roles are the leader and the follower; calculate the Q-value of each behavior according to the current role and optional behaviors of the individual, select and execute the behavior with the maximum Q-value from the Q-table, and generate a new individual position according to the selected role and behavior; S4: Calculate the fitness value of the new individual position. If the fitness value of the new position is greater than that of the original position, the individual is updated from the original position to the new position; S5: After updating the parameters according to the fitness value of the new individual position, execute S3 until a preset condition is reached, obtain the optimal individual position, and use it as the scheduling plan.

2. The multi-energy power system scheduling method based on reinforcement learning according to claim 1, wherein, In S2, initializing individuals in a random and opposition-based learning manner means generating a first preset number of individuals in a random manner and generating a second preset number of individuals in an opposition-based learning manner.

3. A multi-energy power system scheduling system based on reinforcement learning, characterized in that, Including: Multi - energy power system scheduling model construction module: used to construct a thermal power model according to the output energy and ramp rate limit of thermal power units during scheduling, by the formula: , where , represent the output energy of the thermal power unit at time and time respectively, and is the ramp rate limit of the thermal power unit during scheduling; According to the output limit of the wind turbine and the output power limit of the photovoltaic power generation system, the formula: is used to construct the wind-solar model; in the formula, is the actual wind power output value at time , , are respectively the minimum and maximum output powers of the wind turbine at time , is the actual photovoltaic output value at time , , are respectively the minimum and maximum output powers of the photovoltaic power generation system at time ; According to the thermal power output , the actual output value of wind power , the actual output value of photovoltaic power , the discharge power of the energy storage battery , the charging power of the energy storage battery , and the initial load value of the power grid , and the rated capacity of the energy storage power station , from the formula: , and , construct an energy storage model; in the formula, is the scheduling period; Determine the objective function according to the thermal power model, wind-solar model and energy storage model, and construct a multi-energy power system scheduling model; The expression of the objective function is as follows: ; Among them, ; ; In the formula, represents the objective function, is the net grid load before the peak shaving transaction at the moment, is the average value of the net load before the peak shaving transaction, , are the charging efficiency and discharging efficiency of the energy storage battery respectively; Initialization module: Based on the multi-energy power system scheduling model, initialize individuals in a random and opposition-based learning manner to obtain a population; Initialize the parameters of reinforcement learning, including the Q-table and the reward value; Individual position generation module: Used to learn the individual behaviors and roles of the population by using the reinforcement learning method. The individual behaviors are local search and global search, and the roles are the leader and the follower; calculate the Q-value of each behavior according to the current role and optional behaviors of the individual, select and execute the behavior with the maximum Q-value from the Q-table, and generate a new individual position according to the selected role and behavior; Individual position update module: Used to calculate the fitness value of the new individual position. If the fitness value of the new position is greater than that of the original position, the individual is updated from the original position to the new position; Scheduling plan generation module: Used to update the parameters according to the fitness value of the new individual position, then enter the individual position generation module until a preset condition is reached, obtain the optimal individual position, and use it as the scheduling plan.

4. A multi-energy power system scheduling device based on reinforcement learning, characterized in that, Including a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the multi-energy power system scheduling method based on reinforcement learning as described in claim 1 or 2.

Citation Information

Patent Citations

  • Multi-agent power generation optimal scheduling method based on reinforcement learning

    CN110728406A