Multi-energy power system scheduling method, system and device based on reinforcement learning

Through the multi-energy power system scheduling method based on reinforcement learning, thermal power, wind and energy storage models are constructed, and the net load fluctuations of the power grid are optimized, which solves the power system scheduling difficulties caused by high proportion of new energy access, and has achieved improvement in the efficiency of new energy consumption and reduction of peak shaving burden of thermal power units.

CN120013198AActive Publication Date: 2025-05-16YANTAI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510457557.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

High proportion of new energy access leads to difficulties in scheduling the power system. The randomness and volatility of new energy output have intensified the fluctuation of net load in the power grid. Thermal power units need to be deeply peak-shaved, and the new energy units have increased the cost of auxiliary services.

Method used

Using a multi-energy power system scheduling method based on reinforcement learning, the objective function is determined as the minimum variance of the net load of the power grid before peak-shaving transactions by constructing thermal power, wind and light, and energy storage models. The reinforcement learning optimization algorithm is used to initialize the population and learn individual behaviors and roles to generate the optimal scheduling scheme.

Benefits of technology

It effectively reduces the fluctuations in the net load of the power grid, improves the efficiency of new energy consumption, reduces the peak shaving burden of thermal power units, and reduces the auxiliary service costs of new energy units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_8
    Figure QLYQS_8
  • Figure QLYQS_20
    Figure QLYQS_20
  • Figure QLYQS_35
    Figure QLYQS_35
Patent Text Reader

Abstract

The invention belongs to the technical field of electric power system dispatching, and particularly relates to a multi-energy electric power system dispatching method, system and device based on reinforcement learning, aiming at the problem of multi-energy electric power system dispatching, the multi-energy electric power system dispatching method comprises the following steps: firstly, constructing a multi-energy electric power system dispatching model, taking the minimum residual load fluctuation variance as a target, and utilizing stored energy to absorb more wind and light output; secondly, an optimization algorithm is designed to obtain a scheduling scheme, opposite learning is adopted to initialize a population, the diversity of the initial population is improved, premature convergence of a local optimal solution is avoided, and recognition of a global optimal solution by the algorithm is accelerated; the Q value of reinforcement learning is introduced, the optimal behavior strategy of the optimization algorithm is learned, the effectiveness of the optimization algorithm is improved, and a better scheduling scheme is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system dispatching, and specifically relates to a multi-energy power system dispatching method, system, and device based on reinforcement learning. Background Art

[0002] At present, a new power system with new energy as the main body is accelerating its formation. Among them, the proportion of wind power and solar power generation installed capacity and power generation will gradually increase. However, although the access of a high proportion of new energy has alleviated the environmental pressure of the power system to a certain extent, the randomness and volatility of its output have brought unprecedented challenges to the dispatching and operation of the power system. The uncertainty and unpredictability of new energy output have intensified the fluctuation of the net load of the power grid. In particular, the power structure dominated by thermal power lacks power sources that can be flexibly adjusted with the fluctuation of new energy, which will make it extremely difficult to balance the power when a high proportion of new energy is connected to the grid. If the full absorption of new energy is to be achieved, the peak-to-valley difference of the net load of the power grid will be enlarged, and the thermal power units will face the heavy burden of deep peak regulation. In the low load period, the auxiliary service costs that the new energy units need to bear will also increase sharply.

[0003] Therefore, the output uncertainty brought about by the high proportion of wind and solar power connected to the grid has seriously affected the absorption of new energy and the stable operation of the power system. Summary of the invention

[0004] The present invention provides a multi-energy power system scheduling method, system and device based on reinforcement learning.

[0005] The technical solution of the present invention is as follows: The present invention provides a multi-energy power system scheduling method based on reinforcement learning, comprising the following steps: S1: Construct a thermal power model based on the output energy and ramp rate limit of the thermal power units during the dispatch period; Construct a wind-solar model based on the output limit of wind turbines and the output power limit of photovoltaic power generation systems; Construct an energy storage model based on the output power of thermal power, actual output value of wind power, actual output value of photovoltaic power, charging and discharging power of energy storage batteries, initial load value of power grid, and rated capacity of energy storage power station; According to the thermal power model, wind and solar power model and energy storage model, the objective function is determined and the multi-energy power system dispatch model is constructed; S2: Based on the multi-energy power system dispatch model, individuals are initialized using a random, adversarial learning method to obtain a population; Initialize the parameters of reinforcement learning, including Q table and reward value; S3: Use reinforcement learning methods to learn the individual behaviors and roles of the population, calculate the Q value of each behavior based on the individual's current role and optional behaviors, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position based on the selected role and behavior; S4: Calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than the fitness value of the original position, the individual is updated from the original position to the new position. S5: After updating the parameters according to the fitness value of the individual's new position, execute S3 until the preset conditions are met and the optimal individual position is obtained as the scheduling solution.

[0006] S1, according to the output energy and ramp rate limit of the thermal power unit during the scheduling period, constructs a thermal power model according to the formula: ,accomplish; In the formula, , Respectively represent thermal power units exist time, Output energy at any time, is the ramp rate limit of thermal power unit j during the scheduling period.

[0007] S1, according to the wind turbine output limit and the photovoltaic power generation system output power limit, constructs a wind-solar model according to the formula: ,accomplish; In the formula, for The actual wind power output value at the moment, , They are The wind turbine output is minimum and maximum at the moment. for The actual photovoltaic output value at the moment, , They are The minimum and maximum output power of the photovoltaic power generation system at each moment.

[0008] S1 constructs an energy storage model according to the thermal power output power, the actual wind power output value, the actual photovoltaic output value, the charging and discharging power of the energy storage battery and the initial load value of the power grid, and the rated capacity of the energy storage power station, which is based on the formula: ,and ,accomplish; In the formula, for The thermal power output power at the moment, for The actual wind power output value at the moment, for The actual photovoltaic output value at the moment, for The discharge power of the energy storage battery at all times, for The charging power of the energy storage battery at all times, for The initial load value of the power grid at time is the rated capacity of the energy storage power station, The scheduling cycle.

[0009] The S1, the objective function is the minimum variance of the net load of the power grid before the peak-shaving transaction; the net load is the power generation borne only by thermal power after the multi-energy power system reduces the combined output of wind, solar and energy storage.

[0010] The expression of the objective function is: ; in, ; ; In the formula, represents the objective function, is the scheduling period, Before peak-shaving transactions Net load of the grid at any moment, is the average net load before peak load regulation transaction, , , , , They are The initial load value of the power grid at the moment, the actual output value of wind power, the actual output value of photovoltaic power, the charging power of the energy storage battery, and the discharging power of the energy storage battery, , The charging efficiency and discharging efficiency of the energy storage battery respectively.

[0011] The S2, initializing individuals by random and adversarial learning methods, is to generate a first preset number of individuals by random method and to generate a second preset number of individuals by adversarial learning method.

[0012] The S3 uses a reinforcement learning method to learn the individual behaviors and roles of the population, where the individual behaviors are local search and global search, and the roles are leader and follower.

[0013] The present invention also provides a multi-energy power system dispatching system based on reinforcement learning, comprising: Multi-energy power system dispatch model construction module: used to construct a thermal power model according to the output energy and ramp rate limit of the thermal power units during the dispatch period; Construct a wind-solar model based on the output limit of wind turbines and the output power limit of photovoltaic power generation systems; Construct an energy storage model based on the output power of thermal power, actual output value of wind power, actual output value of photovoltaic power, charging and discharging power of energy storage batteries, initial load value of power grid, and rated capacity of energy storage power station; According to the thermal power model, wind and solar power model and energy storage model, the objective function is determined and the multi-energy power system dispatch model is constructed; Initialization module: Based on the multi-energy power system dispatch model, individuals are initialized using a random, adversarial learning method to obtain a population; Initialize the parameters of reinforcement learning, including Q table and reward value; Individual position generation module: used to learn the individual behaviors and roles of the population using reinforcement learning methods, calculate the Q value of each behavior based on the individual's current role and optional behaviors, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position based on the selected role and behavior; Individual position update module: used to calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than the fitness value of the original position, the individual is updated from the original position to the new position. Scheduling scheme generation module: It is used to update the parameters according to the fitness value of the individual new position, and then enter the individual position generation module until the preset conditions are met to obtain the optimal individual position as the scheduling scheme.

[0014] The present invention also provides a multi-energy power system scheduling device based on reinforcement learning, comprising a processor and a memory, wherein the processor implements the multi-energy power system scheduling method based on reinforcement learning when executing a computer program stored in the memory.

[0015] Beneficial Effects Aiming at the problem of multi-energy power system scheduling, the present invention first constructs a multi-energy power system scheduling model, takes the minimum variance of residual load fluctuation as the goal, and uses energy storage to absorb more wind and solar power output; then designs an optimization algorithm to obtain a scheduling plan, wherein adversarial learning is used to initialize the population to improve the diversity of the initial population, which is conducive to avoiding the premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution; the Q value of reinforcement learning is introduced to learn the optimal behavior strategy of the optimization algorithm, improve the effectiveness of the optimization algorithm, and obtain a better scheduling plan. DETAILED DESCRIPTION

[0016] The following examples are intended to illustrate the present invention rather than to further limit the present invention.

[0017] The present invention provides a multi-energy power system scheduling method based on reinforcement learning, comprising the following steps: S1: Construct a thermal power model based on the output energy and ramp rate limit of the thermal power units during the dispatch period; Construct a wind-solar model based on the output limit of wind turbines and the output power limit of photovoltaic power generation systems; Construct an energy storage model based on the output power of thermal power, actual output value of wind power, actual output value of photovoltaic power, charging and discharging power of energy storage batteries, initial load value of power grid, and rated capacity of energy storage power station; According to the thermal power model, wind and solar power model and energy storage model, the objective function is determined and a multi-energy power system dispatching model is constructed.

[0018] Among them, according to the output energy and ramp rate limit of the thermal power unit during the dispatch period, the thermal power model is constructed according to the formula: , to achieve, as a climbing constraint; In the formula, , Respectively represent thermal power units exist time, Output energy at any time, For thermal power units j Ramp rate limit during dispatch.

[0019] Preferably, a wind-solar model is constructed according to the output limit of the wind turbine group and the output power limit of the photovoltaic power generation system, which is based on the formula: , to achieve, as the constraint condition of new energy output; In the formula, for The actual wind power output value at the moment, , They are The wind turbine output is minimum and maximum at the moment. for The actual photovoltaic output value at the moment, , They are The minimum and maximum output power of the photovoltaic power generation system at each moment.

[0020] Preferably, the energy storage model is constructed according to the thermal power output power, the actual wind power output value, the actual photovoltaic output value, the charging and discharging power of the energy storage battery and the initial load value of the power grid, and the rated capacity of the energy storage power station, which is based on the formula: ,and , to be implemented as the system supply and demand balance constraint and energy storage operation constraint respectively; In the formula, for The thermal power output power at the moment, for The actual wind power output value at the moment, for The actual photovoltaic output value at the moment, for The discharge power of the energy storage battery at all times, for The charging power of the energy storage battery at all times, for The initial load value of the power grid at time is the rated capacity of the energy storage power station, The scheduling cycle.

[0021] In addition to the above four constraints, the objective function is the minimum variance of the net load of the power grid before the peak-shaving transaction; the net load is the power generation borne only by thermal power after the multi-energy power system reduces the combined output of wind, solar and energy storage. The variance of the net load reflects the fluctuation of the load borne by the thermal power unit. The smaller the variance, the smaller the fluctuation of the load, and the more stable the operation mode of the corresponding thermal power unit.

[0022] Furthermore, the objective function is expressed as: ; in, ; ; In the formula, represents the objective function, is the scheduling period, Before peak-shaving transactions Net grid load at any moment, is the average net load before peak load regulation transaction, , , , , They are The initial load value of the power grid at the moment, the actual output value of wind power, the actual output value of photovoltaic power, the charging power of the energy storage battery, and the discharging power of the energy storage battery, , The charging efficiency and discharging efficiency of the energy storage battery respectively.

[0023] After constructing the multi-energy power system dispatch model, reinforcement learning is used to obtain the optimal dispatch plan. The operation is as follows: S2: Based on the multi-energy power system dispatch model, individuals are initialized using a random, adversarial learning method to obtain a population; Initialize the parameters of reinforcement learning, including Q table and reward value.

[0024] Preferably, individuals are initialized in a random and adversarial learning manner, that is, a first preset number of individuals are generated in a random manner, and a second preset number of individuals are generated in an adversarial learning manner.

[0025] For example, for the problem of multi-energy power system scheduling, a population of N individuals is initialized, each of which represents a scheduling operation strategy for the scheduling problem. N / 2 individuals can be generated randomly, and the other N / 2 individuals can be generated using adversarial learning.

[0026] The formula for adversarial learning is as follows: ; In the formula, is the jth dimension of the i-th individual position generated by adversarial learning. is the jth dimension of the randomly generated i-th individual position, . , are the lower and upper bounds of the search space respectively.

[0027] The present invention uses adversarial learning to initialize the population, thereby improving the diversity of the initial population, which is beneficial to avoiding premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution.

[0028] S3: Use reinforcement learning methods to learn the individual behaviors and roles of the population, calculate the Q value of each behavior based on the individual's current role and optional behaviors, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position based on the selected role and behavior.

[0029] Preferably, the individual behaviors and roles of the population are learned using a reinforcement learning method, where the individual behaviors are local search and global search, and the roles are leader and follower.

[0030] The position of the dominant individual is expressed as follows: ; In the formula, The position of the dominant individual, is the individual position in the population generated using random and adversarial learning, represents the j-th dimension value of the best individual in the current iteration, is a random number between 0 and 1.

[0031] The position of the follower individual is expressed as follows: ; ; In the formula, is the position of the follower individual, refers to the j-th dimension value of the average position information of a randomly selected individual, are random numbers between 0 and 1, is the current iteration number, is the maximum number of iterations.

[0032] The position of the global search for an individual is expressed as follows: ; ; In the formula, is the individual position using global search, is the randomly generated predator's location information, is a random number generated by Levy distribution, is a random number between 2 and 4, is a random number between 1 and 1.5, is a random number between 2 and 3, is a random number between -1 and 1. is a random vector, is the distance between the individual and the randomly generated predator, , are the fitness values ​​of the individual and the predator, respectively.

[0033] The position of the local search for an individual is expressed as follows: ; ; In the formula, is the individual position using local search, , are the lower and upper bounds of the local search range, respectively. is a random number between 0 and 1.

[0034] Regarding the use of reinforcement learning methods to learn the individual behaviors of the population, specifically: According to the formula: , to carry out; In the formula, , Represents the current state and behavior respectively; Indicates the next state; Represents the discount coefficient, the value range is (0,1); Represents the learning rate, the value range is (0,1); is the total amount of accumulated rewards; is the total amount of accumulated rewards for the next generation; is in the next state Choose the best action Q value; It is an execution behavior The instant reward obtained is defined as follows: ; in, Indicates individual; Indicates the current iteration number; Is the first in the group The fitness value of each individual; It is a reward decay factor to prevent the cumulative reward from increasing infinitely, leading to overly greedy behavior choices.

[0035] The present invention selects the best individual update mode of the optimization algorithm according to the state and reward of the reinforcement learning method, updates the state and reward according to the action results selected by the individual, so that the action that can obtain the maximum reward is taken in any state.

[0036] S4: Calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than the fitness value of the original position, the individual is updated from the original position to the new position; otherwise, the individual remains in the original position unchanged.

[0037] S5: After updating the parameters according to the fitness value of the new individual position (including updating the Q value of the current state in the Q table according to the fitness value to improve subsequent decisions and obtain rewards), execute S3 until the preset conditions are met (such as reaching the maximum number of iterations) and the optimal individual position is obtained as the scheduling solution. The optimal fitness value obtained represents the expected value of the net load variance of the power grid before the peak-shaving transaction.

[0038] Aiming at the problem of multi-energy power system scheduling, the present invention first constructs a multi-energy power system scheduling model, takes the minimum variance of residual load fluctuation as the goal, and uses energy storage to absorb more wind and solar power output; then designs an optimization algorithm to obtain a scheduling plan, wherein adversarial learning is used to initialize the population to improve the diversity of the initial population, which is conducive to avoiding the premature convergence of local optimal solutions and accelerating the algorithm to identify the global optimal solution; the Q value of reinforcement learning is introduced to learn the optimal behavior strategy of the optimization algorithm, improve the effectiveness of the optimization algorithm, and obtain a better scheduling plan.

[0039] The present invention also provides a multi-energy power system dispatching system based on reinforcement learning, comprising: Multi-energy power system dispatch model construction module: used to construct a thermal power model according to the output energy and ramp rate limit of the thermal power units during the dispatch period; Construct a wind-solar model based on the output limit of wind turbines and the output power limit of photovoltaic power generation systems; Construct an energy storage model based on the output power of thermal power, actual output value of wind power, actual output value of photovoltaic power, charging and discharging power of energy storage batteries, initial load value of power grid, and rated capacity of energy storage power station; According to the thermal power model, wind and solar power model and energy storage model, the objective function is determined and the multi-energy power system dispatch model is constructed; Initialization module: Based on the multi-energy power system dispatch model, individuals are initialized using a random, adversarial learning method to obtain a population; Initialize the parameters of reinforcement learning, including Q table and reward value; Individual position generation module: used to learn the individual behaviors and roles of the population using reinforcement learning methods, calculate the Q value of each behavior based on the individual's current role and optional behaviors, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position based on the selected role and behavior; Individual position update module: used to calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than the fitness value of the original position, the individual is updated from the original position to the new position. Scheduling scheme generation module: It is used to update the parameters according to the fitness value of the individual new position, and then enter the individual position generation module until the preset conditions are met to obtain the optimal individual position as the scheduling scheme.

[0040] The present invention also provides a multi-energy power system scheduling device based on reinforcement learning, comprising a processor and a memory, wherein the processor implements the multi-energy power system scheduling method based on reinforcement learning when executing a computer program stored in the memory.

Claims

1. A multi-energy power system dispatching method based on reinforcement learning, characterized in that: The following steps are involved: S1: Construct a thermal power model based on the output energy and ramp rate limit of the thermal power units during the dispatch period; Construct a wind-solar model based on the output limit of wind turbines and the output power limit of photovoltaic power generation systems; Construct an energy storage model based on the output power of thermal power, actual output value of wind power, actual output value of photovoltaic power, charging and discharging power of energy storage batteries, initial load value of power grid, and rated capacity of energy storage power station; According to the thermal power model, wind and solar power model and energy storage model, the objective function is determined and the multi-energy power system dispatch model is constructed; S2: Based on the multi-energy power system dispatch model, individuals are initialized using a random, adversarial learning method to obtain a population; Initialize the parameters of reinforcement learning, including Q table and reward value; S3: Use reinforcement learning methods to learn the individual behaviors and roles of the population, calculate the Q value of each behavior based on the individual's current role and optional behaviors, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position based on the selected role and behavior; S4: Calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than the fitness value of the original position, the individual is updated from the original position to the new position. S5: After updating the parameters according to the fitness value of the individual's new position, execute S3 until the preset conditions are met and the optimal individual position is obtained as the scheduling solution.

2. The multi-energy power system dispatching method based on reinforcement learning according to claim 1 is characterized in that: The S1 constructs a thermal power model according to the output energy and ramp rate limit of the thermal power unit during the scheduling period, which is based on the formula: ,accomplish; In the formula, , Represents thermal power units exist time, Output energy at any moment, is the ramp rate limit of thermal power unit j during the scheduling period.

3. The multi-energy power system dispatching method based on reinforcement learning according to claim 1 is characterized in that: S1 constructs a wind-solar model according to the output limit of the wind turbine and the output power limit of the photovoltaic power generation system, which is based on the formula: ,accomplish; In the formula, for The actual wind power output value at the moment, , They are The wind turbine output is minimum and maximum at the moment. for The actual photovoltaic output value at the moment, , They are The minimum and maximum output power of the photovoltaic power generation system at each moment.

4. The multi-energy power system dispatching method based on reinforcement learning according to claim 1 is characterized in that: S1 constructs an energy storage model according to the thermal power output power, the actual wind power output value, the actual photovoltaic output value, the charging and discharging power of the energy storage battery and the initial load value of the power grid, and the rated capacity of the energy storage power station, which is based on the formula: ,and ,accomplish; In the formula, for The thermal power output power at the moment, for The actual wind power output value at the moment, for The actual photovoltaic output value at the moment, for The discharge power of the energy storage battery at all times, for The charging power of the energy storage battery at all times, for The initial load value of the power grid at time is the rated capacity of the energy storage power station, The scheduling cycle.

5. The multi-energy power system dispatching method based on reinforcement learning according to claim 1 is characterized in that: The S1 objective function is the minimum variance of the net load of the power grid before peak-shaving transactions; the net load is the power generation borne only by thermal power after the multi-energy power system reduces the combined output of wind, solar and energy storage.

6. The multi-energy power system dispatching method based on reinforcement learning according to claim 5 is characterized in that: The expression of the objective function is: ; in, ; ; In the formula, represents the objective function, is the scheduling period, Before peak-shaving transactions Net grid load at any moment, is the average net load before peak load regulation transaction, , , , , They are The initial load value of the power grid at the moment, the actual output value of wind power, the actual output value of photovoltaic power, the charging power of the energy storage battery, and the discharging power of the energy storage battery, , The charging efficiency and discharging efficiency of the energy storage battery respectively.

7. The multi-energy power system dispatching method based on reinforcement learning according to claim 1 is characterized in that: The S2, initializing individuals by random and adversarial learning methods, is to generate a first preset number of individuals by random method and to generate a second preset number of individuals by adversarial learning method.

8. The multi-energy power system dispatching method based on reinforcement learning according to claim 1 is characterized in that: The S3 uses a reinforcement learning method to learn the individual behaviors and roles of the population, where the individual behaviors are local search and global search, and the roles are leader and follower.

9. A multi-energy power system dispatching system based on reinforcement learning, characterized in that: include: Multi-energy power system dispatch model construction module: used to construct a thermal power model according to the output energy and ramp rate limit of the thermal power units during the dispatch period; Construct a wind-solar model based on the output limit of wind turbines and the output power limit of photovoltaic power generation systems; Construct an energy storage model based on the output power of thermal power, actual output value of wind power, actual output value of photovoltaic power, charging and discharging power of energy storage batteries, initial load value of power grid, and rated capacity of energy storage power station; According to the thermal power model, wind and solar power model and energy storage model, the objective function is determined and the multi-energy power system dispatch model is constructed; Initialization module: Based on the multi-energy power system dispatch model, individuals are initialized using a random, adversarial learning method to obtain a population; Initialize the parameters of reinforcement learning, including Q table and reward value; Individual position generation module: used to learn the individual behaviors and roles of the population using reinforcement learning methods, calculate the Q value of each behavior based on the individual's current role and optional behaviors, select and execute the behavior with the maximum Q value from the Q table, and generate a new individual position based on the selected role and behavior; Individual position update module: used to calculate the fitness value of the individual's new position. If the fitness value of the new position is greater than the fitness value of the original position, the individual is updated from the original position to the new position. Scheduling scheme generation module: It is used to update the parameters according to the fitness value of the individual new position, and then enter the individual position generation module until the preset conditions are met to obtain the optimal individual position as the scheduling scheme.

10. A multi-energy power system dispatching device based on reinforcement learning, characterized in that: It comprises a processor and a memory, wherein when the processor executes the computer program stored in the memory, it implements the multi-energy power system scheduling method based on reinforcement learning as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-agent power generation optimal scheduling method based on reinforcement learning

    CN110728406A

  • Wind-photovoltaic-gas-storage combined dynamic economic dispatching optimization method based on Q learning

    CN111064229A

  • Unmanned aerial vehicle task allocation method considering operation environment and performance in collaborative search and rescue

    CN112230675A

  • Wind and light storage-considered power plant system optimization scheduling method, system and equipment and storage medium

    CN118300092A

  • Wind, light, water and fire storage day-ahead optimization scheduling method and system

    CN119726834A