Power optimization scheduling method and system based on reinforcement learning
By adopting deep reinforcement learning and multi-objective optimization technologies in power system scheduling, the problems of single-objective optimization and low computing efficiency in the existing methods are solved, and the efficient, clean and reliable operation of the power system is achieved.
Patent Information
- Application Number
- CN202510247347.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The existing power system scheduling methods based on reinforcement learning have problems such as single-objective optimization, simplified model, low computing efficiency and lack of dynamic characteristics of the power market, which is difficult to meet the needs of power grid complexity and real-time scheduling.
The power system scheduling optimization method based on deep reinforcement learning is adopted to build a multi-objective optimization model, and optimized scheduling strategies are generated using deep Q network and Actor-Critic structure, and the computing efficiency and adaptability of strategies are improved through distributed computing and power system environment simulators.
It has achieved the best balance between power generation costs, environmental impacts and system stability, significantly improved the accuracy and adaptability of scheduling strategies, and improved the optimization efficiency and real-time scheduling capabilities of large-scale power systems.
Smart Images

Figure CN120181464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power dispatching, and more specifically, to a power optimization dispatching method and system based on reinforcement learning. Background Art
[0002] With the continuous expansion of the scale and the increase in complexity of the power system, traditional power system dispatching methods have been difficult to meet the current operation requirements of the power grid. Especially in the context of large-scale grid connection of renewable energy, the uncertainty and volatility of the power system have increased significantly, posing great challenges to dispatching optimization.
[0003] In recent years, with the rapid development of artificial intelligence technology, power system dispatching methods based on machine learning have gradually attracted the attention of the academic and industrial communities. Among them, reinforcement learning, as a method that can autonomously learn and adapt to complex environments, has shown great potential in the field of power system dispatching optimization. However, there are still some problems in existing power system dispatching methods based on reinforcement learning.
[0004] First of all, most existing methods only consider a single objective, such as minimizing the generation cost or maximizing the utilization rate of renewable energy, while ignoring the multi-objective nature of power system dispatching. This single-objective optimization often makes it difficult to achieve a good balance among economy, environmental protection, and reliability, and cannot meet the requirements of actual power system operation.
[0005] Secondly, existing methods usually adopt a simplified power system model, which cannot fully reflect the complexity and dynamic characteristics of the actual power grid. Although this simplified model can reduce the computational complexity, it will also lead to a large deviation between the optimization result and the actual situation, affecting the practicality and reliability of the dispatching strategy.
[0006] In addition, existing methods usually face the problem of low computational efficiency when dealing with large-scale power systems. As the scale of the power grid expands, the state space and action space increase sharply, resulting in a slow convergence rate of the reinforcement learning algorithm and making it difficult to meet the requirements of real-time dispatching.
[0007] Finally, existing methods generally lack consideration of the dynamic characteristics of the power market. The electricity price, as an important economic signal, has an important impact on dispatching decisions. However, most methods adopt static or simplified electricity price models, which cannot accurately reflect the volatility and uncertainty of the power market. Summary of the Invention
[0008] In view of the above problems, the present invention proposes a power system scheduling optimization method and system based on deep reinforcement learning. The method aims to establish a comprehensive, efficient and adaptable power system scheduling optimization framework, which can simultaneously consider multiple objectives such as economy, environmental protection and reliability, make full use of deep learning technology to process complex power grid models, and improve the optimization efficiency of large-scale systems through distributed computing.
[0009] The present invention provides a power system scheduling optimization method based on deep reinforcement learning, including:
[0010] An acquisition step, including:
[0011] Acquire the operation data, load prediction data and renewable energy output data of the power system;
[0012] A processing step, including:
[0013] Based on the operation data, load prediction data and renewable energy output data, construct a power system scheduling optimization model;
[0014] Using a deep reinforcement learning algorithm, based on the power system scheduling optimization model, generate an optimized scheduling strategy;
[0015] According to the optimized scheduling strategy, calculate the scheduling cost, environmental impact and system reliability index of the power system;
[0016] An output step, including:
[0017] Output the optimized scheduling strategy and the corresponding scheduling cost, environmental impact and system reliability index.
[0018] Preferably, the acquisition step specifically includes:
[0019] Acquire power grid topology structure data, generator set parameters, load demand data and renewable energy power generation prediction data;
[0020] Based on historical data, generate short-term, medium-term and long-term load prediction data;
[0021] Acquire real-time meteorological data and environmental monitoring data.
[0022] Preferably, the construction of the power system scheduling optimization model specifically includes:
[0023] Establish a multi-objective optimization model including economic objectives, environmental protection objectives and reliability objectives;
[0024] Among them, the economic objectives include minimizing the total generation cost, pollutant emission cost and network loss cost;
[0025] The environmental protection objectives include minimizing greenhouse gas emissions and resource consumption;
[0026] The reliability objectives include maximizing system stability and power supply reliability.
[0027] Preferably, the specific steps of generating an optimized scheduling strategy using the deep reinforcement learning algorithm include:
[0028] Construct a deep Q-network (DQN) as a value function approximator;
[0029] Design a policy gradient algorithm based on the Actor-Critic structure;
[0030] Adopt an experience replay technique to store and randomly sample historical state-action-reward data;
[0031] Continuously optimize the scheduling strategy through interactive learning.
[0032] Preferably, it also includes:
[0033] Design a reward function considering the characteristics of the power system, including factors such as economic benefits, environmental impacts, and system reliability;
[0034] Adopt an ε-greedy strategy to balance exploration and exploitation;
[0035] Use the target network technique to improve learning stability.
[0036] Preferably, it also includes:
[0037] Construct a distributed parallel computing architecture, and divide the power system into multiple subsystems;
[0038] Allocate an independent deep reinforcement learning agent to each subsystem;
[0039] Design a cooperation mechanism among agents to achieve global optimal scheduling.
[0040] Preferably, it also includes:
[0041] Design an adaptive parameter adjustment strategy to dynamically adjust the neural network structure, learning rate, and weight initialization method;
[0042] Based on the model performance evaluation results, automatically adjust the computing resource allocation and agent cooperation strategy.
[0043] Preferably, it also includes:
[0044] Construct a power system environment simulator, and regard load forecasting and the operating status of generating units as time series;
[0045] Design a electricity price information estimation module to dynamically adjust electricity price forecasting according to the environmental simulation results;
[0046] Evaluate and optimize the scheduling strategy based on the simulation environment and the estimated electricity price.
[0047] Preferably, it further includes:
[0048] Design a comprehensive evaluation index system, including economic, environmental protection and reliability indexes;
[0049] Adopt a multi-objective optimization algorithm, such as NSGA-II, to solve the Pareto optimal solution set;
[0050] Based on the decision maker's preference, select the final scheduling scheme from the Pareto optimal solution set.
[0051] A power system scheduling optimization system based on deep reinforcement learning for executing the above method, including:
[0052] A data acquisition module, used to obtain the operation data, load prediction data and renewable energy output data of the power system;
[0053] A model construction module, used to construct a power system scheduling optimization model based on the operation data, load prediction data and renewable energy output data;
[0054] A deep reinforcement learning module, used to generate an optimized scheduling strategy based on the power system scheduling optimization model by using the deep reinforcement learning algorithm;
[0055] An evaluation calculation module, used to calculate the scheduling cost, environmental impact and system reliability indexes of the power system according to the optimized scheduling strategy;
[0056] A result output module, used to output the optimized scheduling strategy and the corresponding scheduling cost, environmental impact and system reliability indexes;
[0057] Among them, the deep reinforcement learning module includes:
[0058] A deep Q-network sub-module, used to construct a deep Q-network as a value function approximator;
[0059] An Actor-Critic policy gradient sub-module, used to realize policy optimization based on the Actor-Critic structure;
[0060] An experience replay sub-module, used to store and randomly extract historical state-action-reward data;
[0061] An environment interaction sub-module, used to interact with the power system environment simulator to continuously optimize the scheduling strategy.
[0062] The method of the present invention has the following remarkable beneficial effects:
[0063] First, by constructing a multi-objective optimization model including economy, environmental protection, and reliability, the present invention can find the best balance point among power generation cost, environmental impact, and system stability. This comprehensive optimization method can not only reduce the operating cost of the power system, but also improve the utilization rate of renewable energy while ensuring the safe and stable operation of the power grid.
[0064] Secondly, the present invention adopts deep reinforcement learning algorithms, especially deep Q-network and Actor-Critic structures, which can effectively handle high-dimensional state spaces and complex power grid dynamic characteristics. This method significantly improves the accuracy and adaptability of the scheduling strategy and can better cope with various uncertainties and fluctuations in power grid operation.
[0065] Furthermore, the distributed parallel computing architecture designed by the present invention greatly improves the optimization efficiency, enabling this method to be applied to the real-time scheduling of large-scale power systems. By dividing the power grid into multiple subsystems and equipping each subsystem with an independent deep reinforcement learning agent, this method realizes the efficient utilization of computing resources and the parallel processing of optimization tasks.
[0066] In addition, the present invention introduces a power system environment simulator and a electricity price information estimation module, which can more accurately simulate the power grid operation environment and predict electricity price changes. This design significantly improves the practicality and economic benefits of the scheduling strategy, enabling the system to better adapt to the dynamic changes in the power market.
[0067] Finally, the comprehensive evaluation index system and the multi-objective optimization method based on the NSGA-II algorithm proposed by the present invention provide a comprehensive and objective scheduling scheme evaluation tool for decision-makers. This not only helps to select the optimal scheduling strategy, but also can flexibly adjust the weights of various indicators according to actual needs to meet the scheduling requirements in different scenarios.
[0068] In summary, the power system scheduling optimization method and system based on deep reinforcement learning proposed by the present invention realize the intelligentization and high efficiency of power system scheduling through the organic combination of innovative technologies such as multi-objective optimization, deep learning, and distributed computing. This method can not only significantly improve the economic and environmental benefits of the power grid, but also enhance the reliability and flexibility of the power system, providing strong support for building a clean, efficient, and intelligent modern power system. Brief Description of the Drawings
[0069] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiments
[0070] Please refer to Figure 1, the present invention provides a power system scheduling optimization method and system based on deep reinforcement learning. The method obtains the operation data, load prediction data, and renewable energy output data of the power system, constructs a power system scheduling optimization model, and uses a deep reinforcement learning algorithm to generate an optimized scheduling strategy, thereby achieving efficient scheduling of the power system.
[0071] Specifically, the method of the present invention includes the following steps: First, in the acquisition step, the method obtains the operation data, load prediction data, and renewable energy output data of the power system. These data are the basis for power system scheduling optimization. Preferably, in an embodiment of the present invention, the operation data includes the power grid topology structure, generator set parameters, etc.; the load prediction data includes short-term, medium-term, and long-term load predictions; and the renewable energy output data includes the predicted outputs of wind power generation and solar power generation.
[0072] Next, in the processing step, the method constructs a power system scheduling optimization model based on the acquired data. This model is the basis for subsequent deep reinforcement learning. Preferably, in another embodiment of the present invention, the model considers multiple objectives such as economy, environmental protection, and reliability, so that the final scheduling strategy can achieve a balance in multiple aspects.
[0073] Then, the method uses a deep reinforcement learning algorithm to generate an optimized scheduling strategy based on the constructed power system scheduling optimization model. The use of the deep reinforcement learning algorithm enables the method to continuously learn and adapt to the complex power system environment, generating a more intelligent and efficient scheduling strategy.
[0074] After generating the scheduling strategy, the method calculates the scheduling cost, environmental impact, and system reliability index of the power system according to the strategy. The calculation of these indexes provides a basis for evaluating the effect of the scheduling strategy.
[0075] Finally, in the output step, the method outputs the optimized scheduling strategy and the corresponding scheduling cost, environmental impact, and system reliability index. These output results provide an important reference for power system scheduling decisions.
[0076] Furthermore, in the acquisition step of the method of the present invention, it also includes obtaining power grid topology structure data, generator set parameters, load demand data, and renewable energy power generation prediction data. These detailed data enable the method to more comprehensively understand the operating conditions of the power system. Preferably, in an embodiment of the present invention, the power grid topology structure data includes the number of nodes, line connection conditions, etc.; the generator set parameters include the unit type, capacity, efficiency curve, etc.; the load demand data includes the predicted electricity demand of each node; and the renewable energy power generation prediction data includes the output predictions of wind farms and photovoltaic power stations.
[0077] In addition, this method also generates short-term, medium-term, and long-term load forecasting data based on historical data. Preferably, in another embodiment of the present invention, the short-term load forecasting targets the electricity demand for the next 24 hours to one week; the medium-term load forecasting targets the electricity demand for the next month to one quarter; and the long-term load forecasting targets the electricity demand for the next year or longer. This multi-scale load forecasting enables this method to better meet the scheduling requirements of different time scales.
[0078] Meanwhile, this method also acquires real-time meteorological data and environmental monitoring data. These data play an important role in accurately predicting the output of renewable energy and evaluating environmental impacts. Preferably, in one embodiment of the present invention, the meteorological data includes wind speed, sunshine intensity, temperature, etc.; the environmental monitoring data includes atmospheric pollutant concentration, greenhouse gas emissions, etc.
[0079] When constructing the power system scheduling optimization model, this method establishes a multi-objective optimization model including economic objectives, environmental protection objectives, and reliability objectives. This multi-objective optimization model can comprehensively consider all aspects of power system scheduling and achieve a more balanced and sustainable scheduling strategy.
[0080] Specifically, the economic objectives include minimizing the total generation cost, pollutant emission cost, and network loss cost. Preferably, in one embodiment of the present invention, the total generation cost can be expressed as:
[0081]
[0082] where C total is the total generation cost, N is the number of generating units, T is the scheduling period, P i,t is the output of the i-th unit at time t, a i is the quadratic cost coefficient of the i-th unit; b i is the linear cost coefficient of the i-th unit; c i is the fixed cost coefficient of the i-th unit;
[0083] The pollutant emission cost can be expressed as:
[0084]
[0085] where C emission is the pollutant emission cost, M is the number of pollutant types, α j is the unit emission cost of the j-th pollutant, E i,j (P i,t ) is the emission of the j-th pollutant from the i-th unit when the output is P i,t .
[0086] The network loss cost can be expressed as:
[0087]
[0088] Among them, C loss is the network loss cost, β is the electricity price, and P loss,t is the network loss power at time t.
[0089] The environmental protection objectives include minimizing greenhouse gas emissions and resource consumption. Preferably, in another embodiment of the present invention, the greenhouse gas emissions can be expressed as:
[0090]
[0091] Among them, E GHG is the greenhouse gas emissions, and γ i is the greenhouse gas emission coefficient per unit of power generation of the i-th unit.
[0092] The resource consumption can be expressed as:
[0093]
[0094] Among them, R consumption is the resource consumption, and δ i is the resource consumption coefficient per unit of power generation of the i-th unit.
[0095] The reliability objectives include maximizing system stability and power supply reliability. Preferably, in an embodiment of the present invention, the system stability can be measured by frequency deviation and voltage deviation:
[0096]
[0097] Among them, S stability is the system stability index, w f and w v are the weight coefficients of frequency deviation and voltage deviation respectively, and Δf t and ΔV t are the frequency deviation and voltage deviation at time t respectively.
[0098] The value range of the frequency deviation Δf t is:
[0099] The frequency deviation Δf t = f t - f nominal where f t is the actual frequency, and f nominal is the nominal frequency (usually 50Hz or 60Hz). According to the operation standards of the power system, the allowable range of frequency deviation is usually:
[0100] -Δf max ≤Δft ≤Δf max
[0101] where Δf max is the maximum allowable frequency deviation (e.g., ±0.2 Hz or ±0.5 Hz).
[0102] The voltage deviation ΔV t has a value range of:
[0103] The voltage deviation ΔV t = V t - V nominal , where V t is the actual voltage and V nominal is the nominal voltage; according to the operation standards of the power system, the allowable range of voltage deviation is usually:
[0104] -ΔV max ≤ΔV t ≤ΔV max ,
[0105] where ΔV max is the maximum allowable voltage deviation (e.g., ±5% or ±10% of the nominal voltage).
[0106] Power supply reliability can be measured by the expected load loss:
[0107]
[0108] where, R reliability is the power supply reliability index and EENS t is the expected unsupplied electricity at time t.
[0109] By establishing such a multi-objective optimization model, the method of the present invention can find the best balance among economy, environmental protection and reliability, and achieve the efficient, clean and reliable operation of the power system. After constructing the power system dispatching optimization model, the method of the present invention uses the deep reinforcement learning algorithm to generate the optimized dispatching strategy. Specifically, this method constructs a deep Q-network (DQN) as the value function approximator, designs a policy gradient algorithm based on the Actor-Critic structure, and adopts the experience replay technique to continuously optimize the dispatching strategy through interactive learning. DQN is a method based on value function approximation, which selects the optimal action by learning the action value function (Q function).
[0110] In power system dispatch optimization, DQN can be used to quickly generate preliminary dispatch strategies, especially performing well in discrete action spaces. Its advantages are high computational efficiency and suitability for solving simple or moderately complex dispatch problems. Its disadvantage is limited processing ability for high-dimensional continuous action spaces. Actor-Critic is a reinforcement learning algorithm that combines policy gradient methods and value function estimation. Among them, the "Actor" is responsible for generating policies (i.e., selecting actions), while the "Critic" is responsible for evaluating the quality of the current policy (i.e., estimating the value function). In power system dispatch optimization, Actor-Critic is more suitable for handling high-dimensional continuous action spaces and can generate more refined and flexible dispatch strategies. Its advantages are applicability to complex dynamic environments and the ability to achieve better long-term performance. Its disadvantage is that the training process may be slower and require more computational resources.
[0111] In practical applications, DQN and Actor-Critic can be interactively collaborated or switched in the following ways:
[0112] (1) Use in stages
[0113] Initial stage: Use DQN to quickly generate a preliminary dispatch strategy. Due to the fast learning speed of DQN, a feasible solution can be provided in the early stage.
[0114] Optimization stage: Switch to the Actor-Critic algorithm to further optimize the dispatch strategy. Actor-Critic can better explore the action space and thus find the global optimal solution.
[0115] (2) Dynamic switching mechanism
[0116] Dynamically select the appropriate algorithm according to the system state:
[0117] When the system state changes little, use DQN for quick adjustment;
[0118] When the system state is complex or the uncertainty is high, switch to Actor-Critic to obtain a more accurate strategy.
[0119] (3) Use in combination
[0120] Use DQN as the Critic part in Actor-Critic, responsible for estimating the action value function; at the same time, the Actor part still uses the policy gradient method to generate actions. This combined method combines the advantages of both: DQN provides efficient value function estimation, while Actor-Critic realizes flexible action selection.
[0121] If the system uses both DQN and Actor-Critic algorithms, it is necessary to clarify their division of labor and cooperation methods according to the specific scenario:
[0122] (1) Simple scheduling scenario
[0123] If the scale of the power system is small and the action space is simple, DQN can be used alone to complete the scheduling task. For example, small distributed generation systems and local power grid scheduling.
[0124] (2) Complex scheduling scenario
[0125] For large-scale power systems or multi-objective optimization problems, it is recommended to use the Actor-Critic algorithm, or combine DQN with Actor-Critic, such as regional power grid scheduling and multi-energy collaborative optimization.
[0126] (3) Real-time scheduling and long-term optimization
[0127] Real-time scheduling: Use DQN to quickly respond to changes in the system state and generate short-term scheduling strategies.
[0128] Long-term optimization: Use the Actor-Critic algorithm to optimize long-term scheduling strategies to ensure the economy and reliability of the system.
[0129] 4. Functional redundancy analysis of parallel solutions
[0130] If DQN and Actor-Critic are used as parallel solutions, the following points need to be noted to avoid functional redundancy: DQN focuses on quickly generating preliminary scheduling strategies and is suitable for short-term optimization; Actor-Critic focuses on fine-tuning and long-term optimization and is suitable for complex scenarios. In parallel solutions, DQN and Actor-Critic can share the Experience Replay Buffer to improve data utilization. For example, the experience data generated by DQN can be used as training samples for Actor-Critic, and vice versa. Regularly compare the performance metrics of the two algorithms (such as total cost, pollutant emissions, system stability, etc.) and dynamically adjust the weights or switch algorithms according to the results.
[0131] Through reasonable design of the interaction mechanism or clear division of labor, DQN and Actor-Critic can complement and cooperate in the scheduling optimization of the power system and avoid functional redundancy. The following are the recommended cooperation methods:
[0132] 1. Use in stages: Use DQN to quickly generate preliminary strategies in the initial stage and switch to Actor-Critic for refined optimization in the later stage.
[0133] 2. Dynamic switching: Dynamically select the appropriate algorithm according to the system state.
[0134] 3. Integrated use: Use DQN as the Critic part and Actor-Critic as the overall framework to achieve efficient value function estimation and flexible action selection.
[0135] This collaborative mechanism can not only improve the quality of the scheduling strategy, but also significantly reduce the computational cost, providing strong support for the efficient, clean and reliable operation of the power system.
[0136] Preferably, in an embodiment of the present invention, the structure of the deep Q-network is as follows: The input layer receives the state of the current power system, including the output of the generator set, the load demand, the output of renewable energy, etc.; the hidden layer uses a multi-layer fully connected neural network, and the activation function is selected as ReLU; the output layer corresponds to the Q values of different scheduling actions. The loss function of DQN is defined as:
[0137]
[0138] where θ is the parameter of DQN, D is the experience replay pool, r is the reward, γ is the discount factor, and θ - is the parameter of the target network. The method of the present invention also designs a policy gradient algorithm based on the Actor-Critic structure. The Actor network is responsible for generating the scheduling strategy, and the Critic network is responsible for evaluating the value of the strategy. Preferably, in another embodiment of the present invention, the loss function of the Actor network is defined as:
[0139]
[0140] where φ is the parameter of the Actor network, ρ π is the state distribution under the policy π, and Q w is the Critic network. The loss function of the Critic network is defined as:
[0141] L Critic (w) = E (s,a,r,s′)~D (r + γQ w (s', π φ (s′)) - Q w (s, a)) 2
[0142] where w is the parameter of the Critic network.
[0143] To improve the learning efficiency and stability, the method of the present invention adopts the experience replay technique. Preferably, in an embodiment of the present invention, the capacity of the experience replay pool is set to 500000 or 1000000 or higher, and 1024 samples are randomly sampled from it for learning each time. This method can break the correlation between samples and improve the learning effect.
[0144] In addition, the method of the present invention also designs a reward function considering the characteristics of the power system, including factors such as economic benefits, environmental impact, and system reliability. Preferably, in another embodiment of the present invention, the reward function is defined as:
[0145] R = w1R econ + w2R env + w3R rel ,
[0146] wherein, R econ , R env and R rel are the rewards for economic benefits, environmental impact, and system reliability respectively, and w1, w2, and w3 are the corresponding weight coefficients.
[0147] To balance exploration and exploitation, this method adopts the ε-greedy strategy. Preferably, in one embodiment of the present invention, the initial value of ε is set to 0.9 and gradually decays to 0.1 as the training progresses, with a decay rate of 0.995. This allows for sufficient exploration in the initial stage of training and more utilization of the learned knowledge in the later stage. If the decay rate is too small (e.g., 0.9), the value of ε will rapidly decrease in a short time, potentially resulting in insufficient exploration and the model being unable to fully understand the environment.
[0148] If the decay rate is too large (e.g., close to 1), the value of ε hardly changes, causing the model to maintain a high exploration ratio for a long time and affecting the convergence speed.
[0149] The decay rate of 0.995 provides a relatively smooth transition method, which can gradually reduce the exploration ratio during training while retaining a certain degree of flexibility.
[0150] In power system dispatch optimization, the state space and action space can be very complex, and sufficient time for exploration is required to cover diverse operating modes.
[0151] The combination of the initial value ε = 0.9 and the decay rate of 0.995 can ensure that the model has a high exploration ratio in the initial stage of training, thus better adapting to the complex environment.
[0152] As the training progresses, the model gradually accumulates knowledge, the importance of exploration decreases, and the importance of exploitation increases.
[0153] The decay rate of 0.995 enables the value of ε to decrease slowly, relying more on the learned knowledge in the later stage, thereby accelerating convergence.
[0154] The selection of the decay rate is usually based on empirical tuning and experimental verification. In practical applications, researchers will test different decay rates according to the complexity of the task and the data scale, and finally select a value that can ensure both exploration effectiveness and accelerated convergence.
[0155] For complex tasks such as power system dispatch optimization, 0.995 is a proven reasonable choice.
[0156] To improve the stability of learning, the method of the present invention uses the target network technology. Preferably, in another embodiment of the present invention, the parameters of the target network are updated every 100 training steps, and the update method adopts soft update:
[0157] θ - ← τθ+(1 - τ)θ - ,
[0158] where τ is the soft update coefficient, and its value is 0.01. τ = 0.01 does not cause the target network parameters to be updated too quickly, but is a proven reasonable value. In power system dispatch optimization, this setting can reflect the changes of the online network at an appropriate speed while ensuring the stability of the target network. If there is concern about the update speed being too fast, τ can be further reduced according to the specific task requirements (such as 0.005 or 0.001), but it should be noted that this may increase the training time. The final choice should be based on experimental results and the specific requirements of the task. In practical applications, researchers usually test different τ values through methods such as grid search or random search to find the optimal configuration.
[0159] Through the above design, the method of the present invention can effectively learn complex power system dispatch strategies, adapt to different operating environments, and achieve efficient, clean, and reliable power system dispatch.
[0160] To further improve the calculation efficiency and optimization effect, the method of the present invention constructs a distributed parallel computing architecture. Specifically, this method divides the power system into multiple subsystems, assigns independent deep reinforcement learning agents to each subsystem, and designs a cooperation mechanism between the agents to achieve global optimal dispatch.
[0161] Preferably, in one embodiment of the present invention, the division of subsystems is based on the network topology structure and load distribution characteristics of the power system. For example, the subsystems can be divided according to voltage levels, geographical locations, or load centers. Each subsystem is equipped with a deep reinforcement learning agent responsible for the dispatch optimization within the subsystem.
[0162] To achieve global optimal dispatch, the method of the present invention designs a cooperation mechanism between the agents. Preferably, in another embodiment of the present invention, a multi-agent reinforcement learning algorithm based on a shared experience pool is adopted. Specifically, all agents share a global experience replay pool, and each agent not only uses its own experience when learning, but also can learn the experience of other agents. This method can accelerate the learning process and improve the global dispatch effect.
[0163] In addition, the method of the present invention also designs an adaptive parameter adjustment strategy to dynamically adjust the neural network structure, learning rate, and weight initialization method. Preferably, in one embodiment of the present invention, a hyperparameter adjustment method based on Bayesian optimization is adopted. Specifically, a hyperparameter search space is defined, including parameters such as the number of neural network layers, the number of neurons in each layer, and the learning rate, and then the Bayesian optimization algorithm is used to find the optimal hyperparameter combination in this search space.
[0164] Preferably, in another embodiment of the present invention, the learning rate is adjusted using the adaptive learning rate algorithm Adam. The update rule of the Adam algorithm is as follows:
[0165] m t = β1m t-1 + (1 - β1)g t ,
[0166]
[0167]
[0168] where m t and v t are the first-order moment estimate and second-order moment estimate of the gradient respectively, β1 and β2 are momentum parameters, α is the learning rate, and ∈ is a small constant for numerical stability. The method of the present invention also automatically adjusts the computing resource allocation and agent cooperation strategy based on the model performance evaluation results.
[0169] Preferably, in one embodiment of the present invention, a resource allocation method based on the performance contribution degree is adopted. Specifically, the contribution degree of each agent to the global performance is evaluated regularly, and then the computing resources allocated to each agent are dynamically adjusted according to the contribution degree. The calculation formula of the contribution degree is as follows:
[0170]
[0171] where C i is the contribution degree of the i-th agent, ΔP global is the improvement of the global performance, and ΔP i is the improvement of the performance of the i-th agent.
[0172] Through the above design, the method of the present invention can make full use of distributed computing resources to achieve the efficient scheduling optimization of large-scale power systems. At the same time, through adaptive parameter adjustment and agent cooperation, the optimization effect is continuously improved. To further improve the effect of power system scheduling optimization, the method of the present invention constructs a power system environment simulator, regarding load forecasting and the operating state of power generation units as time series. Such a simulator can provide a realistic training environment for deep reinforcement learning algorithms, which helps to improve the learning effect and the practicality of the strategy.
[0173] Preferably, in one embodiment of the present invention, the environmental simulator employs a time series prediction model based on Long Short-Term Memory (LSTM) network. The structure of this model is as follows:
[0174] f t = σ(W f · [h t-1 , x t + b f ),
[0175] i t = σ(W i · [h t-1 , x t + b i ),
[0176]
[0177]
[0178] o t = σ(W o · [h t-1 , x t + b o ),
[0179] h t = o t * tanh(C t ),
[0180] where f t , i t and o t are the forget gate, input gate, and output gate respectively, C t is the cell state, h t is the hidden state, W and b are model parameters, and σ is the sigmoid function.
[0181] LSTM is a neural network structure specifically designed for processing time series data, capable of effectively capturing long-term dependencies. In power system scheduling optimization, LSTM is typically used to predict key variables such as load demand, power generation output, or new energy output in the future for a period of time.
[0182] To ensure the effectiveness of the model, the time series features of the input data need to be properly preprocessed and embedded so that LSTM can fully learn the patterns in these features.
[0183] In the power system scheduling scenario, common input data and their time series features include:
[0184] Load demand: Reflects the variation of users' electricity demand over time.
[0185] Power generation output: Reflects the actual output power of various types of generating units.
[0186] New energy output: Such as wind power, photovoltaic power, etc., has strong randomness and volatility.
[0187] Weather data: Such as temperature, wind speed, light intensity, etc., has a significant impact on load and new energy output.
[0188] Historical data: Such as the operation records over a past period of time, used to capture periodic patterns.
[0189] Multi-dimensionality: The input data may contain multiple dimensions (such as load in different regions, different types of generating units).
[0190] Non-linearity: There may be complex non-linear relationships in time series data.
[0191] Long-term dependence: Some patterns may span a relatively long time period (such as seasonal variations).
[0192] In order to clearly embed the above time series features into the LSTM model, the following methods can be adopted:
[0193] (1) Data standardization:
[0194] Before inputting into the LSTM, perform standardization or normalization on the time series data to eliminate the dimension difference and improve the model convergence speed.
[0195] Common methods include Min-Max standardization and Z-Score standardization.
[0196] (2) Feature engineering:
[0197] Extract features related to the task, such as:
[0198] Timestamp features (hour, date, day of the week, etc.) to capture periodic patterns;
[0199] Sliding window features (such as data of the past N time steps) to enhance the model's learning ability for short-term trends;
[0200] Difference features (such as the difference between the current value and the previous moment value) to reduce the influence of trends and seasonality.
[0201] (3) Multi-dimensional input processing:
[0202] If the input data contains multiple dimensions (such as load, power generation output, weather data, etc.), it can be concatenated into a multi-dimensional time series and used as the input of the LSTM.
[0203] For example, assume that the input at each time step contains the following features:
[0204] Xt = [loadt, wind power outputt, photovoltaic power outputt, temperaturet, wind speedt] Xt = [loadt, wind power outputt, photovoltaic power outputt, temperaturet, wind speedt]
[0205] (4) Embedding Layer
[0206] For discrete or categorical features (such as weather conditions, holiday flags, etc.), an embedding layer can be used to transform them into low-dimensional dense vectors for easier learning by the LSTM.
[0207] (5) Time Step Selection
[0208] Determine the appropriate input time step (sequence length), that is, the length of the time series received by the LSTM each time.
[0209] A larger time step helps capture long-term dependencies but increases computational complexity; a smaller time step is more suitable for short-term prediction.
[0210] Through these methods, time series features can be clearly embedded into the LSTM model, thereby improving prediction accuracy and model stability.
[0211] The method of the present invention also designs an electricity price information estimation module to dynamically adjust electricity price prediction according to the environmental simulation results. This design can better reflect the dynamic characteristics of the electricity market and provide more accurate price signals for scheduling optimization.
[0212] Preferably, in another embodiment of the present invention, the electricity price information estimation module adopts a prediction method based on support vector regression (SVR). The objective function of SVR is defined as:
[0213]
[0214] where W and b are model parameters, ξ i and are slack variables, C is the penalty coefficient, ∈ is the parameter of the insensitive loss function, and φ(x) is the kernel function.
[0215] Based on the simulated environment and the estimated electricity price, the method of the present invention can more accurately evaluate and optimize the scheduling strategy. This method can not only improve the accuracy and reliability of the scheduling strategy but also help the power system better cope with market changes and uncertainties.
[0216] To comprehensively evaluate the effectiveness of the scheduling strategy, the method of the present invention designs a comprehensive evaluation index system, including economic, environmental protection, and reliability indicators. Such a multi-dimensional evaluation system can comprehensively reflect the performance of the scheduling strategy and provide a more comprehensive basis for decision-making.
[0217] Preferably, in one embodiment of the present invention, the economic indicators include the total scheduling cost, power generation cost, and transmission cost; the environmental protection indicators include carbon emissions and renewable energy utilization rate; the reliability indicators include the load shedding probability and system frequency deviation.
[0218] To find the best balance among multiple objectives, the method of the present invention uses the multi-objective optimization algorithm NSGA-II (Non-dominated Sorting Genetic Algorithm II) to solve the Pareto optimal solution set. The main steps of the NSGA-II algorithm include:
[0219] 1. Initialize the population;
[0220] 2. Fast non-dominated sorting;
[0221] 3. Calculate the crowding degree;
[0222] 4. Selection operation;
[0223] 5. Crossover and mutation operations;
[0224] 6. Elite retention strategy;
[0225] Preferably, in another embodiment of the present invention, the parameter settings of the NSGA-II algorithm are as follows: the population size is 100, the crossover probability is 0.9, the mutation probability is 0.1, and the maximum number of iterations is 1000.
[0226] After obtaining the Pareto optimal solution set, the method of the present invention selects the final scheduling plan from the Pareto optimal solution set based on the decision maker's preference. Preferably, in one embodiment of the present invention, the fuzzy comprehensive evaluation method is used to select the final plan. The specific steps are as follows:
[0227] 1. Establish an evaluation index system;
[0228] 2. Determine the weights of the evaluation indicators;
[0229] 3. Construct a fuzzy relation matrix;
[0230] 4. Calculate the fuzzy comprehensive evaluation result;
[0231] 5. Select the plan with the highest comprehensive evaluation value as the final scheduling plan.
[0232] Through the above design, the method of the present invention can select the most practical scheduling plan considering multiple objectives, and achieve the efficient, clean and reliable operation of the power system.
[0233] Deep reinforcement learning generates scheduling strategies through interactive learning, which can adapt to complex dynamic environments and adjust decisions in real time. In power system scheduling, DRL can quickly respond to dynamic changes such as load fluctuations and uncertainties in new energy output. Multi-objective optimization algorithms such as NSGA-II are used to find the Pareto optimal solution set among multiple objectives, ensuring that the scheduling plan reaches a balance in multiple dimensions such as economy, environmental protection and reliability. This method is more suitable for dealing with static or quasi-static problems and can generate a set of candidate scheduling plans in the offline stage. Deep reinforcement learning is good at dealing with dynamic problems, but it may be difficult to directly balance the trade-offs between multiple objectives; multi-objective optimization can provide the global optimal solution set, but it may lack the ability to adapt to dynamic environments. Therefore, combining the two can achieve complementary advantages: DRL provides dynamic adaptation ability, and NSGA-II provides global optimization ability.
[0234] To bridge the logical gap between deep reinforcement learning and multi-objective optimization, the following collaborative verification mechanism can be designed:
[0235] (1) Policy initialization based on the Pareto optimal solution set
[0236] Select several representative scheduling plans from the Pareto optimal solution set of NSGA-II as the initial policies of deep reinforcement learning.
[0237] These initial policies provide a good starting point for DRL, avoiding training from random policies, thus accelerating the convergence speed.
[0238] (2) Dynamically adjust the objective weights
[0239] During the deep reinforcement learning process, dynamically adjust the objective weights according to the current system state (such as the trade-off between economy, environmental protection and reliability).
[0240] The objective weights can refer to the evaluation index system and fuzzy comprehensive evaluation results in NSGA-II to ensure that the policy generation process of DRL is consistent with the objectives of multi-objective optimization.
[0241] (3) Regular verification and feedback
[0242] During the training process of deep reinforcement learning, regularly use NSGA-II to verify the generated policies:
[0243] Map the policies generated by DRL into the space of multi-objective optimization and calculate their positions in the Pareto optimal solution set;
[0244] If the policy deviates from the Pareto optimal solution set, the reward function or parameter settings of DRL are adjusted through the feedback mechanism.
[0245] (4) Incorporate fuzzy comprehensive evaluation
[0246] In the final solution selection stage, the results of DRL and NSGA-II are comprehensively evaluated by combining the fuzzy comprehensive evaluation method: the fuzzy comprehensive evaluation method is used to score each candidate solution; the scheduling solution that best meets the actual requirements is selected according to the comprehensive evaluation value.
[0247] The following are the specific implementation steps for combining deep reinforcement learning with multi-objective optimization:
[0248] Step1: Initialization
[0249] Use NSGA-II to generate the Pareto optimal solution set and select several representative scheduling solutions from it as the initial policy of DRL.
[0250] Step2: Deep reinforcement learning training
[0251] During the DRL training process, dynamically adjust the objective weights to ensure that the policy generation process is consistent with the objectives of multi-objective optimization.
[0252] Regularly submit the generated policy to NSGA-II for verification to ensure that the policy does not deviate from the Pareto optimal solution set.
[0253] Step3: Comprehensive evaluation
[0254] After the training is completed, use the fuzzy comprehensive evaluation method to comprehensively evaluate the results of DRL and NSGA-II:
[0255] 1. Establish an evaluation index system (such as economy, environmental protection, reliability);
[0256] 2. Determine the weights of the evaluation indexes;
[0257] 3. Construct a fuzzy relation matrix;
[0258] 4. Calculate the fuzzy comprehensive evaluation result;
[0259] 5. Select the solution with the highest comprehensive evaluation value as the final scheduling solution.
[0260] Step4: Verification and iteration
[0261] Simulate and verify the final scheduling solution to evaluate its performance in actual operation.
[0262] If the result is not satisfactory, adjust the parameter settings of DRL or NSGA-II and re-execute the above steps.
[0263] Suppose we need to design a dispatching strategy for a regional power grid. The specific steps are as follows:
[0264] 1. NSGA-II initialization:
[0265] Set the population size to 100, the crossover probability to 0.9, the mutation probability to 0.1, and the maximum number of iterations to 1000;
[0266] Generate a Pareto optimal solution set containing 50 candidate dispatching schemes.
[0267] 2. DRL training:
[0268] Select 5 representative schemes from the Pareto optimal solution set as the initial strategy of DRL;
[0269] During the training process, dynamically adjust the weights of economy, environmental protection, and reliability;
[0270] Submit the generated strategy to NSGA-II for verification every 100 steps to ensure that the strategy does not deviate from the Pareto optimal solution set.
[0271] 3. Fuzzy comprehensive evaluation:
[0272] Establish an evaluation index system, including the total cost, pollutant emissions, and system stability;
[0273] Use the expert scoring method to determine the index weights;
[0274] Construct a fuzzy relation matrix and calculate the comprehensive evaluation value;
[0275] Select the scheme with the highest comprehensive evaluation value as the final dispatching scheme.
[0276] By introducing a collaborative verification mechanism, the logical gap between deep reinforcement learning and multi-objective optimization can be effectively bridged. Specifically:
[0277] NSGA-II provides a global optimal solution set, providing the initial strategy and verification criteria for DRL;
[0278] DRL provides dynamic adaptability to ensure that the dispatching strategy can cope with complex dynamic environments;
[0279] The fuzzy comprehensive evaluation method is used to comprehensively evaluate the results of the two methods and select the dispatching scheme that best meets the actual needs.
[0280] This combination method can not only improve the quality of the dispatching strategy, but also significantly improve the stability and applicability of the model, providing strong support for the efficient, clean, and reliable operation of the power system.
[0281] Finally, the present invention also provides a power system scheduling optimization system based on deep reinforcement learning, which system includes a data acquisition module 1, a model construction module 2, a deep reinforcement learning module 3, an evaluation calculation module 4, and a result output module 5.
[0282] The data acquisition module 1 is used to obtain the operation data, load prediction data, and renewable energy output data of the power system. Preferably, in an embodiment of the present invention, the data acquisition module 1 includes a real-time data acquisition sub-module 11 and a historical data processing sub-module 12. The real-time data acquisition sub-module 11 is responsible for obtaining the real-time operation data from the power system SCADA system, while the historical data processing sub-module 12 is responsible for processing and analyzing the historical operation data to provide support for prediction and optimization.
[0283] The model construction module 2 is used to construct a power system scheduling optimization model based on the operation data, load prediction data, and renewable energy output data. Preferably, in another embodiment of the present invention, the model construction module 2 includes an economic model sub-module 21, an environmental protection model sub-module 22, and a reliability model sub-module 23. These three sub-modules are respectively responsible for constructing the mathematical models of the economic objective, environmental protection objective, and reliability objective.
[0284] The deep reinforcement learning module 3 is used to generate an optimized scheduling strategy based on the power system scheduling optimization model by using the deep reinforcement learning algorithm. This module is the core of the system and includes a deep Q-network sub-module 31, an Actor-Critic policy gradient sub-module 32, an experience replay sub-module 33, and an environment interaction sub-module 34. The deep Q-network sub-module 31 is responsible for constructing and training the deep Q-network; the Actor-Critic policy gradient sub-module 32 realizes the policy optimization based on the Actor-Critic structure; the experience replay sub-module 33 is used to store and randomly extract the historical state-action-reward data; and the environment interaction sub-module 34 is responsible for interacting with the power system environment simulator to continuously optimize the scheduling strategy.
[0285] The evaluation calculation module 4 is used to calculate the scheduling cost, environmental impact, and system reliability index of the power system according to the optimized scheduling strategy. Preferably, in an embodiment of the present invention, the evaluation calculation module 4 includes a cost calculation sub-module 41, an environmental impact evaluation sub-module 42, and a reliability analysis sub-module 43. These three sub-modules are respectively responsible for calculating and evaluating the performance of the scheduling strategy in terms of economy, environmental protection, and reliability.
[0286] The result output module 5 is used to output the optimized scheduling strategy and the corresponding scheduling costs, environmental impacts, and system reliability indicators. Preferably, in another embodiment of the present invention, the result output module 5 includes a visualization sub-module 51 and a report generation sub-module 52. The visualization sub-module 51 is responsible for intuitively displaying the optimization results in the form of charts, while the report generation sub-module 52 is responsible for generating a detailed optimization report to provide comprehensive information support for decision-makers.
[0287] Through the collaborative work of the above modules, the system of the present invention can achieve the intelligentization and automation of power system scheduling, greatly improving the scheduling efficiency and optimization effect, and providing strong support for the economic, clean, and reliable operation of the power system.
[0288] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A power system dispatch optimization method based on deep reinforcement learning, characterized in that: include: The acquisition steps include: Obtain power system operation data, load forecast data and renewable energy output data; Processing steps include: Based on the operation data, load forecast data and renewable energy output data, a power system dispatch optimization model is constructed; Using a deep reinforcement learning algorithm, based on the power system dispatch optimization model, an optimized dispatch strategy is generated; Calculating the dispatch cost, environmental impact and system reliability index of the power system according to the optimized dispatch strategy; Output steps include: Output the optimized scheduling strategy and corresponding scheduling cost, environmental impact and system reliability indicators.
2. The method according to claim 1, characterized in that The acquisition step specifically includes: Obtain grid topology data, generator set parameters, load demand data, and renewable energy generation forecast data; Generate short-term, medium-term and long-term load forecast data based on historical data; Get real-time weather data and environmental monitoring data.
3. The method according to claim 1, characterized in that The construction of the power system dispatch optimization model specifically includes: Establish a multi-objective optimization model including economic objectives, environmental objectives and reliability objectives; The economic objectives include minimizing the total power generation cost, pollutant emission cost and network loss cost; The environmental goals include minimizing greenhouse gas emissions and resource consumption; The reliability objectives include maximizing system stability and power supply reliability.
4. The method according to claim 1, characterized in that: The use of a deep reinforcement learning algorithm to generate an optimized scheduling strategy specifically includes: Construct a deep Q network DQN as a value function approximator; Design a policy gradient algorithm based on the Actor-Critic structure; Use experience replay technology to store and randomly extract historical state-action-reward data; Continuously optimize scheduling strategies through interactive learning.
5. The method according to claim 4, characterized in that Also includes: Design reward functions that take into account the characteristics of the power system, including economic benefits, environmental impacts, and system reliability; Use the ε-greedy strategy to balance exploration and exploitation; Improving learning stability using target network techniques.
6. The method according to claim 1, characterized in that Also includes: Build a distributed parallel computing architecture to divide the power system into multiple subsystems; Assign an independent deep reinforcement learning agent to each subsystem; Design a collaboration mechanism between agents to achieve global optimal scheduling.
7. The method according to claim 6, characterized in that Also includes: Design adaptive parameter adjustment strategies to dynamically adjust the neural network structure, learning rate, and weight initialization method; Based on the model performance evaluation results, computing resource allocation and agent collaboration strategies are automatically adjusted.
8. The method according to claim 1, characterized in that Also includes: Build a power system environment simulator that considers load forecasts and power generation unit operating states as time series; Design an electricity price information estimation module to dynamically adjust electricity price forecasts based on environmental simulation results; Evaluate and optimize dispatch strategies based on simulated environments and estimated electricity prices.
9. The method according to claim 1, characterized in that: Also includes: Design a comprehensive evaluation index system, including economic, environmental and reliability indicators; Use multi-objective optimization algorithms, such as NSGA-II, to solve the Pareto optimal solution set; Based on the decision maker's preference, the final scheduling solution is selected from the Pareto optimal solution set.
10. A power system dispatch optimization system based on deep reinforcement learning that executes the method according to any one of claims 1 to 9, characterized in that: include: Data acquisition module, used to obtain power system operation data, load forecast data and renewable energy output data; A model building module, used to build a power system dispatch optimization model based on the operation data, load forecast data and renewable energy output data; A deep reinforcement learning module, used to generate an optimized dispatching strategy based on the power system dispatching optimization model using a deep reinforcement learning algorithm; An evaluation and calculation module, used to calculate the dispatch cost, environmental impact and system reliability index of the power system according to the optimized dispatch strategy; A result output module, used to output the optimized scheduling strategy and corresponding scheduling cost, environmental impact and system reliability indicators; Wherein, the deep reinforcement learning module includes: The Deep Q-Network submodule is used to build a deep Q-network as a value function approximator; Actor-Critic policy gradient submodule, used to implement policy optimization based on the Actor-Critic structure; The experience replay submodule is used to store and randomly extract historical state-action-reward data; The environmental interaction submodule is used to interact with the power system environmental simulator and continuously optimize the dispatching strategy.
Citation Information
Cited By
Community power dispatching optimization method, system and device and medium
CN120914788A
Power grid load side response intelligent decision-making system and method adopting deep reinforcement learning
CN120914792A
Water affair system operation data model optimization scheduling method based on reinforcement learning
CN121212713A
Water system operation data model optimization scheduling method based on reinforcement learning
CN121212713B