Power grid load optimization regulation and control method and system

By adopting adaptive prediction module, intelligent regulation system and power outage risk assessment module in the power grid, combined with deep reinforcement learning and Monte Carlo simulation method, real-time prediction and dynamic regulation of grid load are achieved, solving the problems of inflexible regulation and lack of feedback in traditional technologies, and improving the safety and stability of the power grid.

CN120049422APending Publication Date: 2025-05-27POWER DISPATCHING CONTROL CENT OF GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510158927.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The traditional grid load regulation method cannot adapt to changes in the grid operation state, resulting in inflexible regulation strategies and inability to achieve optimal regulation, and lack of feedback mechanisms, resulting in poor regulation results.

Method used

Adaptive prediction module, intelligent regulation system and power outage risk assessment module are adopted to dynamically predict the power grid load changes through deep reinforcement learning (DQN) and Monte Carlo simulation methods, target load regulation strategies are generated, and the regulation strategies are continuously optimized based on real-time data.

Benefits of technology

Real-time prediction and dynamic regulation of grid loads are realized, the problems of lagging response and inflexible regulation in traditional technologies are overcome, and the safety, stability and regulation efficiency of the power grid are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120049422A_ABST
    Figure CN120049422A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid load optimization regulation and control method and system, and the method comprises the steps: constructing a self-adaptive prediction module, an intelligent regulation and control system and a power failure risk assessment module, and enabling the self-adaptive prediction module to carry out the dynamic prediction of the operation state of a power grid based on the current state data of a power grid system, and generating a target load regulation and control strategy of the power grid; the intelligent regulation and control system generates an updated state reward value according to the target load regulation and control strategy and a state reward function used for evaluating load balance of the power grid, and regulates and controls the power grid according to the target load regulation and control strategy when it is judged that the updated state reward value is smaller than a reward value threshold value; and the power failure risk assessment module is used for identifying different operation modes of the power grid, assessing power failure risks in the modes and feeding back the power failure risks to the adaptive prediction model so as to continuously optimize the regulation and control strategy. According to the invention, the problems of untimely monitoring, slow regulation and control response, rigid scheduling strategy and the like of stable section management during the power failure of the power grid in the prior art are solved, the reliable operation of the stable section during the power failure of the power grid is ensured, and the efficiency and flexibility of power grid scheduling are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power grid load regulation and control, and in particular to a power grid load optimization regulation and control method and system. Background Art

[0002] When there is a partial power outage in the power grid, load regulation can be used to restore the normal operation of the power grid or reduce the impact of the power outage on the power grid, that is, to adjust the load distribution of each node, enable backup power supply and other control measures. For example, the load in the power outage area can be transferred to other normal power supply lines or equipment to reduce the scope of the power outage; in the peak load area adjacent to the power outage area, the load pressure of the power grid can be reduced by adjusting the load distribution and limiting the power consumption of some users; or backup power supplies, such as diesel generators and UPS power supplies, can be enabled in the power outage area to provide temporary power supply.

[0003] In traditional control methods, preset fixed strategies are often used to deal with different power outage events. For example, a series of fixed control strategies for power outage events are pre-set. When a power outage occurs, load transfer and backup power activation are performed according to the preset strategies and the data of the current section. However, the structure and operating conditions of the power grid often change over time, and when a local power outage occurs, it may involve changes in meteorological conditions and uneven load distribution. Since traditional control methods cannot adapt to the actual operating status and changes of the power grid, it is impossible to flexibly derive the optimal control strategy that can adapt to different situations. In addition, traditional control methods cannot provide feedback and continuous optimization based on the control effect, and lack a certain feedback mechanism, resulting in poor control effect, and cannot guarantee the stable operation of the power grid after a power outage or failure. Summary of the invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a method and system for optimizing and controlling power grid load, so as to ensure the reliable operation of stable sections during power grid outages and improve the efficiency and flexibility of power grid dispatching.

[0005] Technical solution: To achieve the above-mentioned purpose, a power grid load optimization and control method described in the present invention includes constructing an adaptive prediction module, an intelligent control system, and a power outage risk assessment module, wherein the adaptive prediction module dynamically predicts the power grid operation status based on the current state data of the power grid system, and generates a target load control strategy for the power grid; the intelligent control system generates an updated state reward value according to the target load control strategy and a state reward function for evaluating the load balancing of the power grid, and when it is determined that the updated state reward value is less than the reward value threshold, the power grid is controlled according to the target load control strategy; the power outage risk assessment module is used to identify different operating modes of the power grid, evaluate the power outage risk under the mode and feed it back to the adaptive prediction model to continuously optimize the control strategy.

[0006] The status data includes power grid operation data and meteorological data, wherein the power grid operation data includes voltage data, current data and load data; the meteorological data includes temperature and wind speed.

[0007] The adaptive prediction module generates a target load control strategy for the power grid, including using a deep Q network DQN to train an agent, so that the agent can predict future load demand by learning power grid operation data, and select the optimal control strategy as the target load control strategy based on the predicted load demand and the current power grid state, specifically:

[0008] The current state data is used as the state space, and an action space is generated according to different preset load control strategies; data features in the state space are extracted, and a strategy execution reward value corresponding to each preset load control strategy in the action space is generated according to an action reward function and various data features; wherein the action reward function is used to evaluate the frequency volatility of the power grid; and the preset load control strategy corresponding to the maximum strategy execution reward value is used as the target load control strategy of the power grid.

[0009] The step of generating an updated state reward value according to the target load regulation strategy and the state reward function for evaluating the load balancing of the power grid includes:

[0010] Input the target load regulation strategy and the current state data into a preset probability prediction model to generate future state data and state transition probability; wherein the future state data is used to characterize the future meteorological data of the power grid and the future operation data of each section after the load regulation strategy is executed; the state transition probability is used to characterize the probability of the state data of the power grid being converted from the current state data to the future state data;

[0011] Based on the Monte Carlo simulation method, the future state data is simulated to generate several expected reward values ​​corresponding to the future state data; according to the state reward function, the expected reward values, the strategy execution reward value corresponding to the target load control strategy, and the state transition probability corresponding to the target load control strategy, an updated state reward value is generated.

[0012] Wherein, the probability prediction model includes a state prediction layer and an output layer;

[0013] The state prediction layer is used to capture the complex dependencies of the power grid system equipment in the long-term operation, and generate future state data corresponding to the current state data according to the dependencies;

[0014] The output layer is used to construct a state transfer matrix according to the future state data and the current state data, and generate a state transfer probability according to each probability value in the state transfer matrix; wherein each element in the state transfer matrix represents a probability value of transferring from the current state to the future state.

[0015] Among them, the state prediction layer generates future state data corresponding to the current state data according to the dependency relationship, including generating the degree of forgetting corresponding to the current state data at the next moment, the degree of updating corresponding to the next moment, and the degree of contribution corresponding to the next moment based on various data features; based on the degree of forgetting, the degree of updating, and the degree of contribution, the future state data corresponding to the current state data is generated.

[0016] Among them, the degree of forgetting corresponding to the next moment is obtained according to the following formula:

[0017] f t =σ(W f ·[h t-1 ,x t-1 ]+b f );

[0018] Among them, f t is the degree of forgetting corresponding to the next moment, σ is the activation function, W f is the weight matrix corresponding to the forget gate, h t-1 is the meteorological characteristics corresponding to the current state data, x t-1 is the cross-sectional feature corresponding to the current state data, b f is the bias vector corresponding to the forget gate;

[0019] The update degree corresponding to the next moment is obtained according to the following formula:

[0020] i t =σ(W i ·[h t-1 ,x t ]+b i );

[0021] Among them, i t is the update degree corresponding to the next moment, σ is the activation function, W i is the weight matrix corresponding to the input gate, b i is the bias vector corresponding to the input gate;

[0022] The contribution degree corresponding to the next moment is obtained according to the following formula:

[0023] o t =σ(W o ·[h t-1 ,x t ]+bo );

[0024] Among them, t is the contribution degree corresponding to the next moment, W o is the weight matrix corresponding to the output gate, b o is the bias vector corresponding to the output gate.

[0025] The step of generating an updated state reward value according to the state reward function, each expected reward value, the strategy execution reward value corresponding to the target load regulation strategy, and the state transition probability corresponding to the target load regulation strategy includes:

[0026] According to the following state reward function, the updated state reward value is calculated:

[0027]

[0028] Among them, V(S t ) represents the updated state reward value, R(S t ,A t ) is the strategy execution reward value corresponding to the target load control strategy, γ is the discount factor, P(S t+1 |S t ,A t ) is the state transition probability, V(S t+1 ) is the expected reward value; It represents the sum of the products of different expected reward values ​​and state transition probabilities, and represents the weighted average of the expected reward values ​​transferred from the current state data to the future state data according to the target load control strategy.

[0029] The power outage risk assessment module uses the K-means clustering algorithm to divide the operation data of the power grid into different clusters, each cluster represents a specific operation mode or load state, and the formula is as follows:

[0030]

[0031] Where J represents the clustering objective function, which is used to measure the power grid operation data point x i To the center μ of its cluster j The sum of the squared distances of data point x i Represents the multi-dimensional characteristic data of power grid operation, such as load, voltage, current, frequency, etc. j Represents the typical operating state of the power grid; r ij Used to determine the data point x i To which cluster it belongs, that is, the operating mode; N represents the number of power grid operation data points; k represents the number of clusters.

[0032] A power grid load optimization control system according to the present invention comprises:

[0033] Adaptive prediction module: dynamically predicts the operation status of the power grid based on the current status data of the power grid system and generates the target load control strategy of the power grid;

[0034] Intelligent control system: generating an updated state reward value according to the target load control strategy and the state reward function for evaluating the load balancing of the power grid, and when it is determined that the updated state reward value is less than the reward value threshold, controlling the power grid according to the target load control strategy;

[0035] Power outage risk assessment module: used to identify different operating modes of the power grid, assess the power outage risk under this mode and feed it back to the adaptive prediction model to continuously optimize the control strategy.

[0036] Beneficial effects: The present invention has the following advantages: 1. The present invention can predict the load change trend during power outage in real time and dynamically adjust the load distribution of the section, thus overcoming the problems of delayed response and inflexible regulation in traditional technologies;

[0037] 2. The present invention can continuously optimize the control strategy based on real-time data, identify and prevent power grid instability in advance, and avoid the risk of large-scale power outages;

[0038] 3. Compared with the existing technology, the present invention has achieved significant improvements in power grid security, stability and regulation efficiency, and is particularly suitable for stable section management and power outage scheduling in complex power grid environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of the load optimization and control process of the present invention. DETAILED DESCRIPTION

[0040] The technical solution of the present invention is described in detail below in conjunction with the embodiments and drawings.

[0041] like Figure 1As shown, the method described in the present invention includes constructing an adaptive prediction module, an intelligent control system, and a power outage risk assessment module, wherein the adaptive prediction module can dynamically predict the operation state of the power grid based on the current state data of the power grid system (current meteorological data and operation data), identify potential stability section problems, and generate a target load control strategy for the power grid; the intelligent control system generates and executes the optimal control strategy based on the target load control strategy. In order to evaluate the effectiveness of the target load control strategy of the current iteration, the intelligent control system also introduces a state reward function, through which the advantages and disadvantages of the control strategy can be evaluated based on the load balance of the power grid, and a corresponding reward value is given. Then, through continuous iteration and feedback, the optimal control strategy can be gradually found; and the power outage risk assessment module is used to analyze real-time data and feed it back to the adaptive prediction model to continuously optimize the control strategy and improve the accuracy of the power grid operation state prediction and the control effect.

[0042] 1. Adaptive Prediction Module

[0043] In the complex environment of power grid operation, power outages are random and uncertain, and traditional methods often have problems such as delayed response and inflexible regulation. Deep reinforcement learning (DRL) can handle the complexity, randomness and uncertainty of power grid operation through continuous learning and adjustment.

[0044] The present invention constructs an adaptive prediction module based on a deep Q network (DQN), and trains agents to enable them to perform load forecasting (predicting future load demand by learning data such as power grid operation) and section control (selecting the optimal control strategy as the target load control strategy based on the predicted load demand and current power grid status). As the decision-making subject in DRL, the agent learns the optimal decision-making strategy through interaction with the power grid environment, thereby dynamically adapting to the complexity and uncertainty in power grid operation and optimizing the load distribution and stability of the power grid.

[0045] Specifically, DQN optimizes the load control strategy by constructing the state space S, action space A and reward function R. The state space S is used to reflect the operation status of the power grid in real time, including key parameters such as voltage, current, frequency and load (combined with external data such as weather and load fluctuations), to help the agent understand the current state of the power grid. The action space A defines the load control strategies that the power grid dispatching system can adopt, covering load adjustment, backup power activation and other control operations. The reward function R defines the reward based on the stability of the power grid operation and the degree of load balance, with the aim of guiding the system to optimize load distribution and reduce overload or section imbalance.

[0046] The learning rate α and the discount factor γ are key parameters for controlling the learning process. The learning rate α is used to control the learning speed of the model, ensuring that the model can balance the short-term and long-term regulation effects. The discount factor γ determines the impact of future returns on current decisions, ensuring that the model focuses on long-term grid stability during the regulation process. By setting these parameters reasonably, DQN can learn and optimize load regulation strategies to adapt to different grid operating states and external environmental changes, thereby effectively overcoming the limitations of traditional methods.

[0047] The formula for using a deep Q network (DQN) to train agents for load forecasting and section control is as follows:

[0048]

[0049] Among them, the Q value function represents the current power grid state s t Next, take action a t The expected return after t+1 is the immediate reward, which is used to measure the effect of the current regulation; α is used to control the learning speed of the model; γ is the discount factor, which determines the impact of future returns on the current decision Represents the future state s t+1 The Q value with the maximum expected reward among all possible actions.

[0050] Using the adaptive prediction module, the process of obtaining the target load control strategy is as follows:

[0051] The voltage data, current data, load data, temperature and wind speed in the current state data are used as the state space, and the action space is generated according to different preset load control strategies;

[0052] Extracting meteorological features and cross-sectional features in the state space; wherein the meteorological features include: temperature features and wind speed features; the cross-sectional features include: voltage features, current features and load features;

[0053] Generate a strategy execution reward value corresponding to each preset load regulation strategy in the action space according to the action reward function, each meteorological feature and each section feature; wherein the action reward function is used to evaluate the frequency volatility of the power grid;

[0054] The preset load control strategy corresponding to the maximum strategy execution reward value is used as the target load control strategy of the power grid.

[0055] 2. Intelligent Control System

[0056] The intelligent control system obtains the state reward value after each execution of the target load control strategy through the Markov decision process (MDP) to determine whether the load control requirements are met (that is, whether the updated state reward value is less than the reward value threshold). In the entire Markov decision process, each execution action, that is, the determination of the target load control strategy, is predicted by the adaptive prediction model. The adaptive prediction model can adapt to the actual operating status and changes of the power grid, avoiding the limitations of the traditional method that relies on preset fixed strategies and cannot analyze and optimize the current status.

[0057] After obtaining the target load regulation strategy, the preset state reward function is used to evaluate the balance benefit of executing the target load regulation strategy on the power grid, so as to find an optimal strategy that maximizes the cumulative reward obtained after taking actions according to each strategy starting from the initial state.

[0058] Then, according to the target load regulation strategy and the state reward function for evaluating the load balancing of the power grid, an updated state reward value is generated, specifically including:

[0059] Input the target load regulation strategy and the current state data into a preset probability prediction model, so that the probability prediction model generates future state data and state transition probability according to the meteorological characteristics of the current meteorological data and the section characteristics corresponding to each current operation data; wherein the future state data is used to characterize the future meteorological data of the power grid and the future operation data of each section after the load regulation strategy is executed; and the state transition probability is used to characterize the probability of the state data of the power grid being converted from the current state data to the future state data;

[0060] Simulate the future state data based on the Monte Carlo simulation method to generate several expected reward values ​​corresponding to the future state data;

[0061] Generate an updated state reward value according to the state reward function, each expected reward value, the strategy execution reward value corresponding to the target load regulation strategy, and the state transition probability corresponding to the target load regulation strategy;

[0062] The training of the probability prediction model includes:

[0063] Using the meteorological data samples and the operation data samples of the sections corresponding to the meteorological data samples as the second training samples;

[0064] Taking several second training samples and the actual future state data and the actual state transition probability of each second training sample as input, and taking the predicted future state data and the predicted state transition probability of each second training sample as output, the probability prediction model to be trained is iteratively trained until the model converges to generate a preset probability prediction model.

[0065] In the calculation process of maximizing the cumulative reward, the basic components of the Markov decision process include:

[0066] State set S: represents the current operating state of the power grid equipment, such as voltage data, current data, load data, temperature and wind speed in the current state data;

[0067] Action set A: represents the actions that the system can take in a certain state, such as the target load regulation strategy, the corresponding line load distribution value, and the activation position of the backup power supply. For each state s∈S, there is a corresponding action set A(s).

[0068] State transition probability p: represents the probability that the system will transition from the current state s to the next state s' after taking an action a. Schematically, the embodiment of the present invention can predict the state transition probability of the system transitioning from the current state s to the next state s' through a preset probability prediction model.

[0069] State reward function R(s,a): represents the benefit or cost of taking an action a in the current state s. The reward function defines the goal of the system, such as keeping the load of the power grid as balanced as possible.

[0070] Strategy π: represents the mapping from state to action, that is, for each state s, strategy π will choose an action a=π(s).

[0071] The present invention can generate an updated state reward value according to the state reward function, each expected reward value, the strategy execution reward value corresponding to the target load regulation strategy, and the state transition probability corresponding to the target load regulation strategy, and the state reward function is:

[0072]

[0073] Among them, V(S t ) represents the updated state reward value, R(S t ,A t ) is the strategy execution reward value corresponding to the target load control strategy, γ is the discount factor, P(S t+1 |S t ,A t ) is the state transition probability, V(S t+1 ) is the expected reward value; It represents the sum of the products of different expected reward values ​​and state transition probabilities, and represents the weighted average of the expected reward values ​​transferred from the current state data to the future state data according to the target load control strategy.

[0074] In the process of generating the updated state reward value, the target load regulation strategy is used to transfer to each expected reward value corresponding to the future state data. In order to better obtain each expected reward value, the present invention can simulate the future state data based on the Monte Carlo simulation method to generate several expected reward values ​​corresponding to the future state data. In MDP, the expected reward value refers to the average reward that the system may obtain in the future after adopting a certain strategy starting from the current state. By comparing the expected reward values ​​obtained by different strategies in the future state, the optimal strategy can be gradually selected.

[0075] Based on accurate future state data, an accurate expected reward value can be simulated. In order to obtain a more accurate state reward value based on the expected reward value, the probability prediction model of the present invention can accurately predict the future state data. The probability prediction model of the present invention includes: a state prediction layer and an output layer;

[0076] The probability prediction model generates future state data and state transition probability according to the meteorological characteristics of the current meteorological data and the cross-sectional characteristics corresponding to each current operation data, including:

[0077] The state prediction layer of the probability prediction model captures the meteorological dependency between the current meteorological data and the current operation data, and generates meteorological features and cross-sectional features according to the dependency; then, based on each meteorological feature and each cross-sectional feature, generates future state data corresponding to the current state data;

[0078] Through the output layer of the probability prediction model, a state transfer matrix is ​​constructed according to the future state data and the current state data, and the state transfer probability is generated according to each probability value in the state transfer matrix; wherein each element in the state transfer matrix represents the probability value of transferring from the current state to the future state.

[0079] When generating future state data corresponding to the current state data according to various meteorological characteristics and various cross-sectional characteristics, it can be specifically as follows:

[0080] Through the state prediction layer of the probability prediction model, based on various meteorological characteristics and various cross-sectional characteristics, the forgetting degree corresponding to the current state data at the next moment, the updating degree corresponding to the next moment, and the contribution degree corresponding to the next moment are generated;

[0081] According to the forgetting degree, updating degree and contribution degree, future state data corresponding to the current state data is generated.

[0082] In a preferred embodiment, the forgetting degree corresponding to the next moment can be obtained according to the following formula:

[0083] f t =σ(W f ·[h t-1 ,x t-1 ]+b f );

[0084] Among them, f t is the degree of forgetting corresponding to the next moment, σ is the activation function, W f is the weight matrix corresponding to the forget gate, h t-1 is the meteorological characteristics corresponding to the current state data, x t-1 is the cross-sectional feature corresponding to the current state data, b f is the bias vector corresponding to the forget gate;

[0085] The update degree corresponding to the next moment is obtained according to the following formula:

[0086] i t =σ(W i ·[h t-1 ,x t ]+b i );

[0087] Among them, i t is the update degree corresponding to the next moment, σ is the activation function, W i is the weight matrix corresponding to the input gate, b i is the bias vector corresponding to the input gate;

[0088] The contribution degree corresponding to the next moment is obtained according to the following formula:

[0089] o t =σ(W o ·[h t-1 ,x t ]+b o );

[0090] Among them, t is the contribution degree corresponding to the next moment, W o is the weight matrix corresponding to the output gate, b o is the bias vector corresponding to the output gate.

[0091] Therefore, by accurately calculating the degree of forgetting, updating and contribution, the model can further improve the accuracy and reliability of the prediction. The present invention can capture the complex dependencies of the equipment during long-term operation through the state prediction layer and obtain accurate future state data corresponding to the current state data.

[0092] Through the collaborative work of the state prediction layer and the output layer, the probabilistic prediction model can capture the complex dependencies of equipment during long-term operation and generate accurate future state data and state transition probabilities, thereby further improving the calculation accuracy of the state reward value, thereby quickly and accurately finding a more effective load control strategy to achieve load optimization control of the power grid.

[0093] Therefore, the present invention is a continuous iterative optimization process, and the input of the adaptive prediction model is the real-time status information of the power grid equipment (such as voltage, current, meteorological data, etc.). By analyzing these states, the Q values ​​corresponding to different actions (i.e., the strategy execution reward values) are output, and the optimal control action (i.e., the target load control strategy) is selected according to the Q value output by DQN. By predicting the future state of the control action and the probability prediction, the benefit value of transferring to the next state (i.e., the state reward value) can be obtained, and the pros and cons of the current control action can be evaluated based on the benefit value.

[0094] The present invention can stop the load regulation operation on the power grid when it is determined that the updated state reward value is not less than the reward value threshold by continuously executing the regulation action and evaluating the state reward value, thereby realizing dynamic optimization regulation of the load.

[0095] 3. Power outage risk assessment module

[0096] The present invention combines a clustering-based risk assessment algorithm to provide an accurate assessment and dynamic adjustment scheme for the power outage risk of the power grid. The clustering algorithm is used to analyze the different operating states of the power grid, and combined with real-time data to identify risk patterns, the K-means clustering algorithm can be used to perform unsupervised learning analysis on the different operating states of the power grid. The clustering algorithm can automatically divide the operating data of the power grid into different clusters, each cluster representing a specific operating mode or load state. By identifying these patterns, the risk of power outage can be assessed in advance and provide a basis for optimization and regulation. The formula is as follows:

[0097]

[0098] Among them, the data point x i Represents the multi-dimensional characteristic data of power grid operation, such as load, voltage, current, frequency, etc. j Represents the typical operating state of the power grid; r ij Used to determine the data point x ibelongs to which cluster; J represents the clustering objective function (loss function), which is used to measure the power grid operation data point x i To the center μ of its cluster (operating mode) j The goal is to minimize this value so that the grid operation status of the same category is as close as possible, so as to more accurately identify the power outage risk pattern. N represents the number of grid operation data points, each of which represents the operation status of the grid at a certain moment (such as voltage, current, load, etc.); k represents the number of clusters, that is, the different possible operation modes of the grid.

[0099] The system divides the operation status of the power grid into k modes based on historical data and real-time data, such as normal operation mode, light overload mode, high-risk mode (may cause power outage), section imbalance mode, etc.

[0100] Through K-means clustering, the system can identify different operating modes of the power grid and evaluate the risk of power outages under these modes. When a certain operating mode frequently experiences load overload or voltage fluctuations, the mode will be marked as high-risk, allowing the system to make regulatory adjustments in advance.

[0101] An embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements a load optimization and control method for a power grid as described in any method embodiment of the present invention.

[0102] An embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a load optimization and control method for a power grid as described in any method embodiment of the present invention.

Claims

1. A method for optimizing and controlling power grid load, characterized in that: include: Construct an adaptive prediction module, an intelligent control system, and a power outage risk assessment module. The adaptive prediction module dynamically predicts the grid operation status based on the current status data of the grid system and generates a target load control strategy for the grid. The intelligent control system generates an updated state reward value according to the target load control strategy and the state reward function used to evaluate the load balancing of the power grid. When it is determined that the updated state reward value is less than the reward value threshold, the power grid is controlled according to the target load control strategy; the power outage risk assessment module is used to identify different operating modes of the power grid, evaluate the power outage risk under the mode and feed it back to the adaptive prediction model to continuously optimize the control strategy.

2. The power grid load optimization control method according to claim 1 is characterized in that: The state data includes power grid operation data and meteorological data, wherein the power grid operation data includes voltage data, current data and load data; the meteorological data includes temperature and wind speed.

3. The power grid load optimization control method according to claim 1, characterized in that: The adaptive prediction module generates a target load control strategy for the power grid, including using a deep Q network DQN training agent to enable the agent to predict future load demand by learning power grid operation data, and select the optimal control strategy as the target load control strategy based on the predicted load demand and the current power grid state, specifically: The current state data is used as the state space, and an action space is generated according to different preset load control strategies; data features in the state space are extracted, and a strategy execution reward value corresponding to each preset load control strategy in the action space is generated according to an action reward function and various data features; wherein the action reward function is used to evaluate the frequency volatility of the power grid; and the preset load control strategy corresponding to the maximum strategy execution reward value is used as the target load control strategy of the power grid.

4. The power grid load optimization control method according to claim 1, characterized in that: The step of generating an updated state reward value according to the target load regulation strategy and the state reward function for evaluating the load balancing of the power grid includes: Input the target load regulation strategy and the current state data into a preset probability prediction model to generate future state data and state transition probability; wherein the future state data is used to characterize the future meteorological data of the power grid and the future operation data of each section after the load regulation strategy is executed; the state transition probability is used to characterize the probability of the state data of the power grid being converted from the current state data to the future state data; Based on the Monte Carlo simulation method, the future state data is simulated to generate several expected reward values ​​corresponding to the future state data; according to the state reward function, the expected reward values, the strategy execution reward value corresponding to the target load control strategy, and the state transition probability corresponding to the target load control strategy, an updated state reward value is generated.

5. The method for optimizing and controlling power grid load according to claim 4, characterized in that: The probability prediction model includes a state prediction layer and an output layer; The state prediction layer is used to capture the complex dependencies of the power grid system equipment in the long-term operation, and generate future state data corresponding to the current state data according to the dependencies; The output layer is used to construct a state transfer matrix according to the future state data and the current state data, and generate a state transfer probability according to each probability value in the state transfer matrix; wherein each element in the state transfer matrix represents a probability value of transferring from the current state to the future state.

6. The method for optimizing and controlling power grid load according to claim 5, characterized in that: The state prediction layer generates future state data corresponding to the current state data according to the dependency relationship, including generating the degree of forgetting corresponding to the current state data at the next moment, the degree of updating corresponding to the next moment, and the degree of contribution corresponding to the next moment based on various data features; and generating future state data corresponding to the current state data according to the degree of forgetting, the degree of updating, and the degree of contribution.

7. The method for optimizing and controlling power grid load according to claim 6, characterized in that: The degree of forgetting corresponding to the next moment is obtained according to the following formula: f t =σ(W f ·[h t-1 ,x t-1 ]+b f ); Among them, f t is the degree of forgetting corresponding to the next moment, σ is the activation function, W f is the weight matrix corresponding to the forget gate, h t-1 is the meteorological characteristics corresponding to the current state data, x t-1 is the cross-sectional feature corresponding to the current state data, b f is the bias vector corresponding to the forget gate; The update degree corresponding to the next moment is obtained according to the following formula: i t =σ(W i ·[h t-1 ,x t ]+b i ); Among them, i t is the update degree corresponding to the next moment, σ is the activation function, W i is the weight matrix corresponding to the input gate, b i is the bias vector corresponding to the input gate; The contribution degree corresponding to the next moment is obtained according to the following formula: the t =σ(W o ·[h t-1 ,x t ]+b o ); Among them, t is the contribution degree corresponding to the next moment, W o is the weight matrix corresponding to the output gate, b o is the bias vector corresponding to the output gate.

8. The method for optimizing and controlling power grid load according to claim 7, characterized in that: The generating an updated state reward value according to the state reward function, each expected reward value, the strategy execution reward value corresponding to the target load regulation strategy, and the state transition probability corresponding to the target load regulation strategy includes: According to the following state reward function, the updated state reward value is calculated: Among them, V(S t ) represents the updated state reward value, R(S t ,A t ) is the strategy execution reward value corresponding to the target load control strategy, γ is the discount factor, P(S t+1 |S t ,A t ) is the state transition probability, V(S t+1 ) is the expected reward value; It represents the sum of the products of different expected reward values ​​and state transition probabilities, and represents the weighted average of the expected reward values ​​transferred from the current state data to the future state data according to the target load control strategy.

9. The method for optimizing and controlling power grid load according to claim 1, characterized in that: The power outage risk assessment module uses the K-means clustering algorithm to divide the operation data of the power grid into different clusters, each cluster represents a specific operation mode or load state, and the formula is as follows: Where J represents the clustering objective function, which is used to measure the power grid operation data point x i To the center μ of its cluster j The sum of the squared distances of data point x i Represents the multi-dimensional characteristic data of power grid operation, such as load, voltage, current, frequency, etc. j Represents the typical operating state of the power grid; r ij Used to determine the data point x i To which cluster it belongs, that is, the operating mode; N represents the number of power grid operation data points; k represents the number of clusters.

10. A power grid load optimization control system, characterized in that: include: Adaptive prediction module: dynamically predicts the grid operation status based on the current status data of the grid system and generates the target load control strategy of the grid; Intelligent control system: generating an updated state reward value according to the target load control strategy and the state reward function for evaluating the load balancing of the power grid, and when it is determined that the updated state reward value is less than the reward value threshold, controlling the power grid according to the target load control strategy; Power outage risk assessment module: used to identify different operating modes of the power grid, assess the power outage risk under this mode and feed it back to the adaptive prediction model to continuously optimize the control strategy.

Citation Information

Cited By

  • Electromechanical equipment adaptive control system based on depth deterministic strategy gradient algorithm

    CN120335318A

  • Adaptive control system for electromechanical equipment based on deep deterministic policy gradient algorithm

    CN120335318B