An intelligent substation with an automated control system
Through the Markov decision-making process and the automated control system of deep neural network algorithm, the reactive power balance and voltage stability problems of traditional power scheduling systems in complex power grid environments are solved, and efficient fault detection and optimization control of intelligent substations are realized.
Patent Information
- Application Number
- CN202510142791.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Traditional power scheduling systems are difficult to cope with complex and changeable power grid environments, resulting in uneven scheduling of power and power resources, unable to achieve reactive power balance and voltage stability control, and low fault detection efficiency.
An automated control system using Markov decision-making process (MDP) and deep neural network (DQN) algorithms is used to realize reactive power balance and voltage stability of the substation through data acquisition, communication and equipment execution layers, combining fault identification and positioning.
It improves the adaptability of the power grid and the accuracy of fault detection, reduces the false alarm rate, and realizes optimized control of reactive power balance and voltage stability.
Smart Images

Figure CN119602492B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of substation power dispatching, and specifically, relates to an intelligent substation with an automatic control system. Background Art
[0002] In order to cope with the increasingly complex power grid environment and the growing power demand, especially the large-scale access of renewable energy and the wide distribution of distributed generation resources, it has had a significant impact on the reactive power and voltage deviation characteristics of the power system. Reactive power and voltage deviation are not only related to the stability of the power system and the power quality, but also directly affect the transmission efficiency and economy of the power grid. Existing power grid dispatching strategies are unable to cope with these new characteristics, and it is urgent to explore new dispatching strategies to adapt to this change; the automatic control system of intelligent substations has emerged as the times require. However, the traditional power dispatching scheme has the following deficiencies, which limit its application in intelligent substations:
[0003] First: The traditional system usually conducts dispatching based on fixed models and historical data, and it is difficult to accurately respond to the complex and changeable power grid environment, which may lead to uneven scheduling and distribution of electric power and energy resources, and it is impossible to achieve optimal reactive power balance and voltage stability control.
[0004] Second: Different fault events often occur during the power dispatching process of traditional substations. The fault detection methods often rely on manual experience and complex data analysis, and it is difficult to accurately locate the fault point, which affects the fault handling efficiency. Summary of the Invention
[0005] To solve the technical problem that traditional substations are difficult to accurately respond to the complex and changeable power grid environment, resulting in uneven scheduling and distribution of electric power and energy resources, the present invention provides an intelligent substation with an automatic control system; an automatic optimization control strategy based on the Markov decision process (MDP) and the deep neural network (DQN) algorithm.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] An intelligent substation with an automatic control system, comprising a data acquisition layer, a communication layer, a decision-making control layer, and a device execution layer;
[0008] The data acquisition layer is used to collect line state parameters, load equipment operation actions, and electrical characteristics of the substation operation status;
[0009] Preferably, the line state parameters include current, voltage, active power, and reactive power of each node of the transmission line;
[0010] The operating actions of the load equipment include the breaker switch state of the load equipment, the tap position of the transformer, and the operating state of the reactive power compensation equipment;
[0011] The electrical characteristics include voltage deviation and reactive power deviation.
[0012] The communication layer is used for data transmission, uploading the collected operating state data to the decision control layer, and sending control instructions to the equipment execution layer; and sending control instructions to the equipment execution layer.
[0013] The decision control layer is used to analyze the collected operating state data, model the control strategy of the substation operating state using the Markov decision process, and at the same time learn the optimal control strategy by training the deep neural network algorithm, and generate control instructions for the corresponding load equipment to achieve the optimization of reactive power balance and voltage stability of the substation;
[0014] The equipment execution layer includes breakers, transformer tap switches, and reactive power compensation equipment; it is used to perform state operations on the corresponding load equipment according to the control instructions;
[0015] Preferably, the specific implementation process of the decision control layer includes:
[0016] S1) MDP modeling: Model the power grid dispatching problem of the substation as a Markov decision process (MDP), and define the state space, action space, and reward function;
[0017] Preferably, the state space includes the transmission line voltage, transmission line reactive power, transmission line active power, and voltage deviation of the substation;
[0018] The action space includes the operation control of the compensation capacitor bank, the adjustment of the transformer tap, the adjustment of the transformer voltage ratio, and the power output control of the reactive power compensation equipment (SVC);
[0019] The reward function includes a positive reward for reactive power balance, a positive reward for voltage stability, and a negative reward for the operating state where the reactive power deviation or voltage deviation exceeds the threshold.
[0020] Preferably, the reward function provides a positive reward for the target state through the balance of reactive power, encouraging the system to reach or approach the state of reactive power balance; the specific calculation method includes:
[0021] Define the reactive power deviation: Calculate the reactive power deviation ΔQ of the substation, that is, the difference between the power demand of the load line of the substation and the current transmission line power;
[0022] Reward design: Set the reward value according to the reactive power deviation value ΔQ; the smaller the deviation, the greater the reward; the reward function of reactive power is expressed as:
[0023] ;
[0024] Wherein, is the maximum positive reward; is the adjustment coefficient; is the reactive power deviation value.
[0025] Preferably, the reward function provides a positive reward for the target state through voltage stability, encouraging the system to maintain the voltage stability of each node; the specific calculation method includes:
[0026] Define the voltage deviation: Calculate the voltage deviation ΔV of each line node in the system, that is, the difference between the load line voltage of the substation and the target voltage;
[0027] Reward design: Set the reward value according to the voltage deviation ΔV; when the node voltage is close to the target value, a higher reward is given; the smaller the deviation, the greater the reward; the reward function of the voltage deviation is expressed as:
[0028] ;
[0029] Wherein, is the maximum positive reward; is the adjustment coefficient; is the voltage deviation value.
[0030] Preferably, the reward function comprehensively calculates the overall reward through reactive power balance, voltage stability and negative reward conditions;
[0031] Among them, if the reactive power deviation value ΔQ or the voltage deviation ΔV exceeds its set threshold, a negative reward is given; the overall reward function is expressed as:
[0032] ;
[0033] Wherein, , are the weights of each reward, and the weight parameters are set according to the operation priority; is the positive reward value of the reactive power deviation; is the positive reward value of the voltage deviation; is the negative reward of the reactive power deviation value or the voltage deviation.
[0034] S2) DQN algorithm construction: Construct a fully connected neural network (DQN) algorithm, and bring the reward function after MDP modeling into the Q-value function in the DQN algorithm. The Q-value function represents the expected cumulative reward for taking a certain action in a given state;
[0035] Among them, the input layer receives the current state vector, and the number of state vectors is used as the number of neurons;
[0036] Multiple fully connected layers process the input state vector and extract vector features; the number of fully connected layers is the number of actions;
[0037] The output layer outputs the Q-values of each possible action;
[0038] S3) Training process:
[0039] Through MDP modeling, samples of states, actions, rewards, and next states are collected and stored in the experience replay pool;
[0040] Initialize the DQN algorithm parameters, regularly draw samples from the experience replay pool, and use these samples to train the DQN algorithm;
[0041] In the specific implementation process,
[0042] Obtain the current state s, execute the action a, and obtain the reward R;
[0043] Obtain the next state s', and select the action a' according to the ε-greedy strategy;
[0044] Update the Q-value using the Bellman equation to maximize the expected cumulative reward strategy; at the same time, update the DQN algorithm parameters.
[0045] Preferably, the calculation formula for updating the Q-value by the Bellman equation is expressed as:
[0046] ;
[0047] In the formula, is the expected cumulative reward for taking a certain action in a given state; is the overall reward of the current action, that is, the immediate reward of the current action a; is the discount factor, which controls the influence degree of future rewards; is the Q-value of the optimal action a' in the next state s′, representing the strategy of maximizing the expected reward.
[0048] S4) Application and deployment:
[0049] After training is completed, deploy the DQN algorithm to the actual power grid dispatching system, input a new state vector into the DQN algorithm, and obtain the Q-values of each action; select the action with the largest Q-value as the control instruction.
[0050] Preferably, the control instruction is at least one action in the action space of the MDP modeling.
[0051] Preferably, the step S3) further includes:
[0052] DQN algorithm parameter optimization: The mean squared error (MSE) is used as the loss function to measure the gap between the expected Q value and the target Q value; the parameters of the DQN algorithm are optimized through the backpropagation algorithm and gradient descent.
[0053] Preferably, the decision control layer is also used for fault identification and rapid location of power anomalies in the substation; the specific implementation process includes:
[0054] Collect historical data of the substation system in different operating states (different power loads or active powers), and statistically calculate the Q-value distribution of corresponding actions;
[0055] Using the weighted average method, analyze the Q-value distribution of each action and set a normal fluctuation threshold.
[0056] When new state data is input into the DQN algorithm to calculate the Q values of each action; compare with the normal fluctuation threshold, if the Q value of an action exceeds its normal fluctuation threshold, it is marked as abnormal;
[0057] Further analyze the associated actions and load equipment with abnormal Q values, and then accurately locate the power anomaly point.
[0058] Advantages of the present invention:
[0059] 1. Using the Markov decision process to model the control strategy of the substation operating state, the state-action mapping of the MDP can help the system associate different operating states with corresponding control strategies; at the same time, by training the deep neural network (DQN) algorithm to learn the optimal control strategy and generate control instructions for the corresponding load equipment, to achieve the optimization of the reactive power balance and voltage stability of the substation.
[0060] 2. The DQN algorithm introduces an experience replay mechanism to enhance the robustness of the reinforcement learning process in a random sampling manner and avoid overfitting; coupled with the ε-greedy strategy to achieve a balance between exploration and exploitation, which helps to improve the training efficiency and makes the system more adaptable in practical applications.
[0061] 3. This solution sets a reasonable threshold range by analyzing the Q-value distribution of different actions, and then, after the input of a new state, it can judge in real time whether the Q value of an action exceeds the normal range, so as to accurately locate the abnormal point. It not only improves the speed of fault detection but also reduces the false alarm rate, making the detection result more reliable. Description of the Drawings
[0062] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0063] Figure 1 Schematic diagram of the framework structure of an intelligent substation with an automated control system according to the present invention;
[0064] Figure 2 Flowchart of the algorithm of the decision control layer in an intelligent substation with an automated control system according to the present invention. Detailed implementation manners
[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention.
[0066] Please refer to Figure 1 - Figure 2 As shown, an intelligent substation with an automated control system includes a data acquisition layer, a communication layer, a decision control layer, and a device execution layer;
[0067] The data acquisition layer is used to collect line state parameters, load equipment operation actions, and electrical characteristics of the substation operation state;
[0068] Furthermore, the line state parameters include current, voltage, active power, and reactive power of each node of the transmission line;
[0069] The load equipment operation actions include the circuit breaker switch state of the load equipment, the tap position of the transformer, and the operation state of the reactive power compensation equipment;
[0070] The electrical characteristics include voltage deviation and reactive power deviation.
[0071] Specifically, the intelligent substation includes a transmission line from a power station, load equipment, and a load line (supplying power to users); among them, the load equipment is a control means for the power energy scheduling and distribution of the substation and is crucial for achieving reactive power balance and voltage stability.
[0072] The data acquisition layer of the present invention, as the "perception layer" of the entire system, is used to obtain various operation data of the substation in real time. By deploying voltage transformers (VT) at the nodes of the transmission line to measure the transmission voltage; current transformers (CT) to measure the transmission current. At the same time, voltage transformers (VT) are deployed at the nodes of the load line to measure the load voltage; current transformers (CT) to measure the load current.
[0073] Deploy status indicators on load equipment to collect the operating status of each load equipment. Among them, the main function of the circuit breaker is to control the current path of the transmission line in the power system, execute the opening and closing states of the circuit breaker, and ensure the reactive power balance of the system when the power demand changes.
[0074] The adjustment of the transformer tap position directly affects the voltage regulation ability of the transformer output, and then ensures the stability of the grid voltage.
[0075] Reactive power compensation equipment (such as static var compensator SVC, capacitors, instrument transformers, etc.) is used in the power system to balance reactive power and stabilize voltage.
[0076] The programmable logic controller (PLC) prepares data acquisition tasks, collects data from each measurement device in real time, and at the same time preliminarily processes the aggregated data to calculate the reactive power deviation and voltage deviation of each node.
[0077] Among them, the reactive power deviation is the difference between the transmission line power and the current load line power; it can be calculated by using the measured values of voltage and current at the same time. The voltage deviation is the difference between the current load line voltage and the target voltage after transformation.
[0078] Power characteristics include: Reactive power balance: Using the balance of reactive power differences is also for reasonable scheduling and distribution of actual electricity demand, and the effect is more intuitive.
[0079] Voltage deviation: Voltage deviation refers to the difference between the voltage of a certain node in the power system and the rated voltage. Excessive voltage deviation will affect the normal operation of power equipment and even cause equipment damage.
[0080] As a more preferred implementation, the power characteristics can also include line overload: Line overload means that the current flowing through the power line exceeds its rated current. Line overload will cause the line to heat up and even burn out the line, resulting in a power outage accident. Therefore, the power characteristics also need to collect temperature data of relevant nodes (load circuits, transformers and other load equipment) as a basis.
[0081] The communication layer is used for data transmission, uploading the collected operating status data to the decision-making control layer, and sending control instructions to the device execution layer; and sending control instructions to the device execution layer.
[0082] Specifically, the IEC 61850 or Modbus communication protocol is adopted. The EC 61850 protocol is particularly suitable for the communication requirements of power systems, supports data modeling and device management, and helps to improve the reliability and efficiency of transmission. In terms of network transmission, it can be transmitted through industrial Ethernet or wireless sensor network (WSN) technology; ensuring the reliability and efficiency of data transmission. Industrial Ethernet is suitable for high-bandwidth and low-latency requirements, while WSN is more flexible in remote or difficult-to-wire areas. By combining the two, seamless connection of the substation is achieved.
[0083] The decision control layer is used to analyze the collected operation status data, model the control strategy of the substation operation status using the Markov decision process, and at the same time learn the optimal control strategy by training the deep neural network algorithm, and generate control instructions for the corresponding load equipment to achieve the optimization of the reactive power balance and voltage stability of the substation;
[0084] The device execution layer includes circuit breakers, transformer tap changers, and reactive power compensation equipment; it is used to perform state operations on the corresponding load equipment according to the control instructions;
[0085] Furthermore, the specific implementation process of the decision control layer includes:
[0086] S1) MDP modeling: Model the power grid scheduling problem of the substation as a Markov decision process (MDP), and define the state space, action space, and reward function;
[0087] Furthermore, the state space includes the transmission line voltage, transmission line reactive power, transmission line active power, and voltage deviation of the substation;
[0088] The action space includes the operation control of the compensation capacitor bank, the adjustment of the transformer tap, the adjustment of the voltage transformation ratio of the transformer, and the power output control of the reactive power compensation equipment (SVC);
[0089] The reward function includes a positive reward for reactive power balance, a positive reward for voltage stability, and a negative reward for the operation state where the reactive power deviation or voltage deviation exceeds the threshold.
[0090] Specifically, MDP modeling (Markov decision process modeling) is used to transform the power grid scheduling problem into a state-action-reward framework, so that the decision control layer can optimize the control strategy of the power system through the deep neural network algorithm.
[0091] Among them, first, according to the power dispatching requirements of the substation, define the state space of the system, mainly including the following parameters:
[0092] Voltage (V): The voltage value of the transmission line, used to judge voltage stability.
[0093] Reactive power (Q): The reactive power of the transmission line, which is used for reactive power balance monitoring.
[0094] Active power (P): The active power of the transmission line, which is used for the judgment of the load line.
[0095] Assume that the state space S includes the above several feature vectors, then the state space can be expressed as S = [V, Q, P], where t represents the current moment.
[0096] Next, in each state, the action space that the system can execute includes the following:
[0097] Adjust the reactive power compensation state (C): By accessing, increasing or decreasing the input power of reactive power compensation devices (such as SVC, capacitors, transformers, etc.), to meet the reactive power balance requirements of the power grid.
[0098] Adjust the tap position of the transformer (T): Change the tap position to adjust the voltage level.
[0099] Control the breaker state (K): Open or close the breaker to adjust the power distribution or isolate the faulty line.
[0100] Assume that the action space a contains these operation options, then the action space can be expressed as a = [C, T, K], representing reactive power compensation control, tap regulation, and breaker control respectively.
[0101] Finally, the reward function reflects the effect of the scheduling action. If the adjusted power system achieves reactive power balance or the voltage deviation is maintained within a reasonable range, the reward function gives a positive value, indicating that the operation is effective. If the parameters exceed the limits (such as the reactive power or voltage deviation exceeds the threshold), a negative reward is given.
[0102] Furthermore, the reward function provides a positive reward for the target state through the balance of reactive power, encouraging the system to reach or approach the state of reactive power balance; the specific calculation method includes:
[0103] Define the reactive power deviation: Calculate the reactive power deviation ΔQ of the substation, that is, the difference between the power demand of the load line of the substation and the current transmission line power;
[0104] Reward design: Set the reward value according to the reactive power deviation value ΔQ; the smaller the deviation, the greater the reward; the reward function of reactive power is expressed as:
[0105] ;
[0106] In the formula, is the maximum positive reward; is the adjustment coefficient; is the reactive power deviation value.
[0107] Furthermore, the reward function provides a positive reward for the target state through voltage stability, encouraging the system to maintain stable voltages at each node. The specific calculation method includes:
[0108] Define the voltage deviation: Calculate the voltage deviation ΔV of each line node in the system, which is the difference between the load line voltage of the substation and the target voltage.
[0109] Reward design: Set the reward value according to the voltage deviation ΔV. When the node voltage approaches the target value, a higher reward is given. The smaller the deviation, the greater the reward. The reward function for voltage deviation is expressed as:
[0110] ;
[0111] In the formula, is the maximum positive reward; is the adjustment coefficient; is the voltage deviation value.
[0112] Furthermore, the reward function comprehensively calculates the overall reward through reactive power balance, voltage stability, and negative reward conditions.
[0113] Among them, if the reactive power deviation value ΔQ or the voltage deviation ΔV exceeds its set threshold, a negative reward is given. The overall reward function is expressed as:
[0114] ;
[0115] In the formula, , are the weights of each reward, and the weight parameters are set according to the operation priority; is the positive reward value for reactive power deviation; is the positive reward value for voltage deviation; is the negative reward for reactive power deviation value or voltage deviation.
[0116] S2) DQN algorithm construction: Construct a fully connected neural network (DQN) algorithm, and substitute the reward function modeled by MDP into the Q-value function in the DQN algorithm. The Q-value function represents the expected cumulative reward for taking a certain action in a given state.
[0117] Among them, the input layer receives the current state vector, and the number of state vectors serves as the number of neurons;
[0118] Multiple fully connected layers process the input state vector and extract vector features. The number of fully connected layers is the number of actions;
[0119] The output layer outputs the Q-value of each possible action;
[0120] Specifically, DQN is a reinforcement learning algorithm that estimates the Q-value function through a fully connected neural network. It can predict the long-term cumulative reward for each possible action in a given state, and then determine the optimal action.
[0121] In the construction process of the DQN algorithm, the input layer is used to receive the state vector S of the current system. Among them, the state vector S contains 5 feature parameters, so the number of neurons in the input layer is 5.
[0122] Multiple fully connected layers (usually 2 - 3 layers) are used to extract features from the state vector and implement non-linear feature mapping to better learn the complex relationship between states and actions. An activation function (such as ReLU) can be applied between each fully connected layer to introduce non-linearity, enabling the network to express complex control strategies. The action space contains at least 3 scheduling operations:
[0123] 1. Reactive power compensation device adjustment: Increase the output quota increment of reactive power devices (e.g., from 20 Mvar to 30 Mvar).
[0124] 2. Transformer tap position adjustment: Raise the tap from position 2 to position 3.
[0125] 3. Circuit breaker state control: Disconnect a certain transmission line; to reduce the reactive power of the transmission line.
[0126] The output layer will have at least 3 nodes, and each node outputs a Q value. These Q values represent the expected cumulative reward for executing each action in the current state.
[0127] S3) Training process:
[0128] Through MDP modeling, collect samples of states, actions, rewards, and next states, and store them in the experience replay pool;
[0129] Initialize the DQN algorithm parameters, regularly extract samples from the experience replay pool, and use these samples to train the DQN algorithm;
[0130] In the specific implementation process,
[0131] Obtain the current state s, execute the action a, and obtain the reward R;
[0132] Obtain the next state s', and select the action a' according to the ε-greedy strategy;
[0133] Use the Bellman equation to update the Q value to maximize the strategy of the expected cumulative reward; at the same time, update the DQN algorithm parameters.
[0134] Furthermore, the calculation formula for updating the Q value using the Bellman equation is expressed as:
[0135] ;
[0136] Wherein, is the expected cumulative reward for taking a certain action in a given state; is the overall reward of the current action, that is, the immediate reward of the current action a; is the discount factor, which controls the influence degree of future rewards; is the Q-value of the optimal action a' in the next state s', representing the strategy that maximizes the expected reward.
[0137] S4) Application and Deployment:
[0138] After the training is completed, the DQN algorithm is deployed into the actual power grid dispatching system. A new state vector is input into the DQN algorithm to obtain the Q-values of each action; the action with the largest Q-value is selected as the control instruction.
[0139] Specifically, the decision-making and control layer is the "core" of the system. The key lies in using an intelligent dispatching algorithm to combine the Markov decision process (MDP) and the deep neural network (DQN) algorithm to maintain the reactive power balance and voltage stability of the system under changing load line power supply demands and environments. The specific implementation process is as follows:
[0140] MDP Modeling: Model the power dispatching problem of the substation as a Markov decision process, and define the state space, action space, and reward function. The state space includes the state information of transmission lines and compensation devices, the action space includes reactive power compensation devices and transformer adjustment operations, and the reward function reflects the reactive power balance and voltage stability requirements of the system.
[0141] Deep Learning Training: Use the DQN algorithm (specifically, a fully connected neural network) to approximate the Q-value function. The DQN algorithm gradually learns the optimal control strategy based on the action performance in different states. Here, the experience replay and Bellman equation are used to continuously optimize the Q-value function, so as to accurately evaluate and select the optimal control action under complex power grid states.
[0142] At the initial stage of training, the weight parameters of the DQN network are random values. The DQN starts to interact with the system environment through these initial parameters, continuously optimizes its parameters, and finally learns the optimal dispatching strategy.
[0143] In each training iteration, the ε-greedy strategy is also adopted to balance exploration and exploitation during the training process: randomly select an action with probability ϵ (exploration) to discover potential better strategies. Select the action with the largest Q-value in the current state with probability (1 - ϵ) (exploitation) to execute the known optimal strategy.
[0144] Control instruction generation and optimization: Evaluate the best action in the current state through the Q-value function, convert this action into a control instruction and send it to the device execution layer. Ensure that the control strategy has long-term stability and reliability.
[0145] Furthermore, the control instruction is at least one action in the action space in MDP modeling.
[0146] Furthermore, step S3) also includes:
[0147] DQN algorithm parameter optimization: Use the mean squared error (MSE) as the loss function to measure the gap between the expected Q-value and the target Q-value; Optimize the DQN algorithm parameters through the backpropagation algorithm and gradient descent.
[0148] Furthermore, the decision control layer is also used for fault identification and rapid location of power anomalies in the substation; The specific implementation process includes:
[0149] Collect historical data of the substation system in different operating states (different power loads or active powers), and statistically calculate the Q-value distribution of corresponding actions;
[0150] Adopt the weighted average method to analyze the Q-value distribution of each action and set a normal fluctuation threshold.
[0151] When new state data is input into the DQN algorithm to calculate the Q-values of each action; Compare with the normal fluctuation threshold. If the Q-value of a certain action exceeds its normal fluctuation threshold, it is marked as abnormal;
[0152] Further analyze the associated actions and load equipment with abnormal Q-values, and then accurately locate the power anomaly point.
[0153] Specifically, in different operating states of the substation (such as peak load, valley load or different active power levels), various historical data are collected. These data include the data of substation equipment in different states, and record the Q-values of each action in these states. Specifically, the Q-value distribution of each action (such as transformer tap position adjustment, reactive power compensation equipment start-stop, etc.) will obtain different distribution curves according to different historical loads and power conditions, and these curves can reflect the normal Q-value range of each action in different states.
[0154] Statistically analyze the Q-value distribution, and use the weighted average method to obtain the normal range of Q-values of each action, and then set a normal fluctuation threshold. Usually, if the Q-value of a certain action is distributed between [10, 15] in most states, then [10, 15] can be defined as the normal fluctuation range of this action, and the Q-value fluctuations within this range are considered normal.
[0155] Then, compare the Q-value of each action output by the DQN algorithm with the previously set normal fluctuation threshold. If the Q-value of a certain action exceeds its normal range, it is regarded that the device corresponding to this action may be abnormal.
[0156] When an abnormal Q-value is detected, the system will further analyze the specific affected objects of this action, such as associated load devices, transformers, etc. By tracing the action corresponding to the abnormal Q-value and its affected area, the system can accurately locate the specific power anomaly point. For example, if the output power of a reactive power compensation device changes and the Q-value is abnormal, it may indicate that the device has a fault or the current reactive power compensation is insufficient. By locating the anomaly point, the system can issue an alarm in time and perform corresponding maintenance and emergency operations.
[0157] In the description of the specification, if the method or process is implemented in the form of a software program, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. And the foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
[0158] In the description of the specification, the descriptions referring to terms such as "specifically", "further", etc. mean that the specific features, structures, materials, or characteristics described in connection with this embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0159] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all belong to the protection scope of the present invention.
Claims
1. An intelligent substation with an automatic control system, characterized in that: It includes a data acquisition layer, a communication layer, a decision-making and control layer, and a device execution layer; The data acquisition layer is used to collect line state parameters, load device operation actions, and power characteristics of the substation operation status; The communication layer is used for data transmission, uploading the collected operation status data to the decision-making and control layer, and sending control instructions to the device execution layer; The decision-making and control layer is used to analyze the collected operation status data, model the control strategy of the substation operation status using the Markov decision process, and at the same time learn the optimal control strategy by training the deep neural network algorithm, and generate control instructions for the corresponding load devices; The specific implementation process of the decision-making and control layer includes: S1) MDP modeling: Model the power grid dispatching problem of the substation as a Markov decision process, and define the state space, action space, and reward function; S2) DQN algorithm construction: Construct a fully connected neural network algorithm, and bring the reward function after MDP modeling into the Q-value function in the DQN algorithm. The Q-value function represents the expected cumulative reward for taking a certain action in a given state; Among them, the input layer receives the current state vector, and the number of state vectors is used as the number of neurons; Multiple fully connected layers process the input state vector and extract vector features; the number of fully connected layers is the number of actions; The output layer outputs the Q-values of each possible action; S3) Training process: Through MDP modeling, collect samples of states, actions, rewards, and next states, and store them in the experience replay pool; Initialize the DQN algorithm parameters, regularly extract samples from the experience replay pool, and use these samples to train the DQN algorithm; Obtain the current state s, execute the action a, and obtain the reward R; Obtain the next state s', and select the action a' according to the ε-greedy strategy; Use the Bellman equation to update the Q-value to maximize the expected cumulative reward strategy; at the same time, update the DQN algorithm parameters; S4) Application and deployment: After training, deploy the DQN algorithm to the actual power grid dispatching system, input a new state vector into the DQN algorithm, and obtain the Q-values of each action; select the action with the largest Q-value as the control instruction; The device execution layer includes circuit breakers, transformer tap changers, and reactive power compensation devices; it is used to execute the state operations of the corresponding load devices according to the control instructions; In step S1), the state space includes the transmission line voltage, transmission line reactive power, transmission line active power, and voltage deviation of the substation; The action space includes the operation control of the compensation capacitor bank, the adjustment of the transformer tap, the adjustment of the transformer voltage ratio, and the power output control of the reactive power compensation device; The reward function includes a positive reward for reactive power balance, a positive reward for voltage stability, and a negative reward for the operation state where the reactive power deviation or voltage deviation exceeds the threshold.
2. The intelligent substation with an automated control system according to claim 1, characterized in that: The line state parameters include the current, voltage, active power, and reactive power of each node of the transmission line; The device operation actions include the circuit breaker switch state, transformer tap position, and reactive power compensation device operation state of the load device; The power characteristics include voltage deviation and reactive power deviation.
3. The intelligent substation with an automatic control system according to claim 1, characterized in that: The reward function provides a positive reward for the target state based on the balance of reactive power, encouraging the system to reach or approach the state of reactive power balance; the specific calculation method includes: Define the reactive power deviation: Calculate the reactive power deviation ΔQ of the substation, which is the difference between the reactive power demand of the load lines of the substation and the current transmission line power; Reward design: Set the reward value according to the reactive power deviation value ΔQ; the smaller the deviation, the greater the reward; the reward function of reactive power is expressed as: ; wherein, is the maximum positive reward; is the adjustment coefficient; is the reactive power deviation value.
4. The intelligent substation with an automatic control system according to claim 1, characterized in that: The reward function provides a positive reward for the target state based on voltage stability, encouraging the system to maintain the voltage stability of each node; the specific calculation method includes: Define the voltage deviation: Calculate the voltage deviation ΔV of each line node in the system, which is the difference between the voltage of the load line of the substation and the target voltage; Reward design: Set the reward value according to the voltage deviation ΔV; when the node voltage is close to the target value, a higher reward is given; the smaller the deviation, the greater the reward; the reward function of voltage deviation is expressed as: ; In the formula, is the maximum positive reward; is the adjustment coefficient; is the voltage deviation value.
5. An intelligent substation with an automated control system according to claim 1, characterized in that: The reward function comprehensively calculates the overall reward through reactive power balance, voltage stability and negative reward conditions; Among them, if the reactive power deviation value ΔQ or the voltage deviation ΔV exceeds its set threshold, a negative reward is given; the overall reward function is expressed as: ; Wherein, , are the weights of each reward, and the weight parameters are set according to the operation priority; is the positive reward value of reactive power deviation; is the positive reward value of voltage deviation; is the negative reward of reactive power deviation value or voltage deviation.
6. An intelligent substation with an automated control system according to claim 1, characterized in that: The calculation formula for updating the Q value by the Bellman equation is expressed as: ; wherein, is the expected cumulative reward for taking a certain action in a given state; is the overall reward of the current action, that is, the immediate reward of the current action a; is the discount factor, which controls the influence degree of future rewards; is the Q value of the optimal action a' in the next state s', representing the strategy that maximizes the expected reward.
7. An intelligent substation with an automated control system according to claim 1, characterized in that: The step S3) also includes: DQN algorithm parameter optimization: Use the mean square error as the loss function to measure the gap between the expected Q value and the target Q value; optimize the DQN algorithm parameters through the backpropagation algorithm and gradient descent.
8. An intelligent substation with an automated control system according to any one of claims 1-7, characterized in that: The decision control layer is also used for fault identification and rapid location of power anomaly points in the substation; the specific implementation process includes: Collect the historical data of the substation system in different operating states and statistically calculate the Q value distribution of the corresponding actions; Adopt the weighted average method to analyze the Q value distribution of each action and set a normal fluctuation threshold; When the new state data is input into the DQN algorithm to calculate the Q values of each action; compare with the normal fluctuation threshold, if the Q value of an action exceeds its normal fluctuation threshold, it is marked as abnormal; Further analyze the associated actions and load equipment with abnormal Q values, and then accurately locate the power anomaly point.
Citation Information
Patent Citations
5G power distribution networking architecture system
CN117318306A
Active power correction control method and device for power system
CN118899917A