Reinforcement learning-based power system voltage control device and its operating method

The reinforcement learning-based system integrates heterogeneous control facilities in power systems, ensuring stability and efficiency by adhering to safety constraints, thus addressing control conflicts and adaptability issues.

KR102996760B1Active Publication Date: 2026-07-29CROCUS INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
CROCUS INC
Filing Date
2025-10-13
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing voltage control methods in power systems face challenges in integrating heterogeneous control facilities due to their different characteristics, leading to control conflicts and inefficiencies, and are limited by high computational complexity and lack of adaptability to dynamic changes.

Method used

A reinforcement learning-based system that includes a state observation unit, policy decision unit, and action execution unit, with a safety constraint unit to ensure compliance with physical, logical, and operational constraints, enabling integrated control of multiple heterogeneous control facilities.

Benefits of technology

The system prevents control conflicts, enhances system stability, reduces mechanical equipment wear, and optimizes voltage stability and power quality by adapting to dynamic changes in real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure R1020250146504_ABST
    Figure R1020250146504_ABST
Patent Text Reader

Abstract

The system voltage control device of the present invention, in a voltage control device of a power system system including a plurality of control facilities, may include a state observation unit that observes state information from the power system system, a policy decision unit that determines a proposed action for the plurality of control facilities based on the state information using a policy model learned based on reinforcement learning, and an action execution unit that outputs the action determined by the policy decision unit to the power system system.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a system voltage control device and a method for optimizing voltage stability and power quality of a power system by integrally controlling multiple heterogeneous control facilities within a power system using artificial intelligence-based reinforcement learning. Background Technology

[0002] Recently, as the complexity and uncertainty of power systems have increased due to the rise of renewable energy and electric vehicles, the importance of stable voltage control is growing.

[0003] Various equipment used for voltage control, such as OLTC, VSC, and ESS, have different characteristics, making integrated control difficult. Furthermore, existing rule-based or model-based control methods are difficult to apply in real time and may lack adaptability to changing environments.

[0004] In other words, existing voltage control methods, such as rule-based control or model-based optimization, require accurate grid models; however, in large-scale systems, the computational complexity required to derive optimal solutions is high, limiting real-time application and potentially lacking adaptability to the dynamic changes of the grid that occur moment by moment. Although there have been recent attempts to apply reinforcement learning to voltage control, its application to actual power grids remains limited due to the risk of violating safety regulations during the learning process or causing conflicting operations between equipment. The problem to be solved

[0005] The present invention aims to provide a reinforcement learning-based voltage control technology capable of integrated control of multiple heterogeneous control facilities without a model and autonomously adapting to dynamic changes in the system. means of solving the problem

[0006] The system voltage control device of the present invention, in a voltage control device of a power system system including a plurality of control facilities, may include a state observation unit that observes state information from the power system system, a policy decision unit that determines a proposed action for the plurality of control facilities based on the state information using a policy model learned based on reinforcement learning, and an action execution unit that outputs the action determined by the policy decision unit to the power system system. Effects of the invention

[0007] The present invention can prevent control conflicts and effect cancellations between multiple devices and significantly improve the stability of the system by applying constraints such as matching the direction of reactive power and preventing conflicting operations between heterogeneous equipment.

[0008] In addition, the present invention can protect the lifespan of mechanical equipment and reduce long-term operating and maintenance costs by applying constraints such as limiting the number of OLTC operations and guaranteeing a minimum operation interval.

[0009] In addition, the present invention can maximize grid operation efficiency by responding in real time to grid uncertainty, renewable energy output fluctuations, or dynamic load changes through reinforcement learning, and by integrally optimizing multiple objectives such as voltage stability, minimization of power loss, and power quality.

[0010] Furthermore, by introducing safety reinforcement learning and safety constraints, the present invention can resolve the issue of unguaranteed safety that may arise in reinforcement learning-based control. Through this, various constraints, such as physical, operational, and logical constraints, are guaranteed, enabling the reliable application of artificial intelligence technology to critical industrial infrastructure, such as power systems. Brief explanation of the drawing

[0011] FIG. 1 is an overall configuration diagram of a reinforcement learning-based power system voltage control system according to one embodiment of the present invention. FIG. 2 is a detailed structural block diagram of a control device including a reinforcement learning structure according to one embodiment of the present invention. FIG. 3 is an operation flowchart of a reinforcement learning-based power system voltage control method according to one embodiment of the present invention. FIG. 4 is a conceptual diagram of a specific processing example of logical constraints and operational constraints processed in the safety constraint section of the present invention. Specific details for implementing the invention

[0012] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings. However, the present invention is not limited to the embodiments disclosed below and can be implemented in various different forms.

[0013] The reinforcement learning (RL) of the present invention may refer to a machine learning methodology in which an agent learns a behavior policy in a direction that maximizes cumulative rewards through interaction with an environment (300), such as a power system. In particular, the present invention may include deep reinforcement learning (DRL) that utilizes an artificial neural network to handle the complex and continuous state space of a power system.

[0014] The Safe Reinforcement Learning (Safe RL) of the present invention ensures that the system does not violate predefined safety constraints during the learning and execution processes of reinforcement learning. The present invention goes beyond simply imposing a penalty on the reward when a constraint is violated and includes actively avoiding or correcting dangerous behavior itself, and may be required to apply reinforcement learning to safety-critical systems such as power grids.

[0015] Voltage-VAR Control (VVC) can refer to a technology that maintains power quality and system stability by controlling the voltage and reactive power of a power system.

[0016] Heterogeneous control equipment may refer to control devices of different characteristics mixed within a power system. For example, mechanically operated OLTC (310) and power electronic device-based VSC (320), ESS (330), STATCOM (340), etc., differ in response speed (slow / fast) and control method (discrete / continuous).

[0017] The present invention relates to power control technology for optimizing voltage stability and power quality of a power system. Recently, power systems have been difficult to control in an integrated manner due to the coexistence of various heterogeneous control facilities, and uncertainty in the system is increasing due to factors such as the increase in renewable energy sources.

[0018] Accordingly, the present invention aims to solve these problems by introducing reinforcement learning to adapt to dynamic changes in the system without a model and to integrally optimize multiple heterogeneous control facilities.

[0019] In one embodiment, the present invention can perform control based on Safe Reinforcement Learning (Safe RL) while satisfying not only physical limits of the power system but also logical constraints such as preventing conflicting operations between devices or operational constraints such as limiting the number of operations of specific equipment. To this end, the present invention may include a safety constraint unit (250) as a control structure.

[0020] FIG. 1 is a schematic diagram of the overall configuration of the reinforcement learning-based power grid voltage control system (100) of the present invention.

[0021] In FIG. 1, the grid voltage control device (100) may include a reinforcement learning-based control unit (200) or a power grid system (environment) (300). The present invention may adopt a centralized control structure for integrated optimization of the entire grid.

[0022] The power system (300) may correspond to an environment for reinforcement learning and may include a number of heterogeneous control facilities for controlling voltage or reactive power. In one embodiment, it may include at least one of an OLTC (310), VSC / SST (320), ESS (330), and STATCOM / SVC (340).

[0023] The OLTC (310) controls the voltage by mechanically adjusting the taps of the transformer and has slow response speed and discrete control characteristics.

[0024] VSC / SST (320), ESS (330), or STATCOM / SVC (340) is a power electronics-based device that has the ability to control active power and / or reactive power quickly and continuously.

[0025] In addition, the power system system (300) of the present invention may include a measuring instrument (350), such as a PMU and a SCADA data collection device, that measures the state of the system in real time.

[0026] The reinforcement learning-based control unit (200) is a Safe Reinforcement Learning (Safe RL)-based agent and can observe state information (S) of the entire power system (300) from a measuring instrument (350) or each control facility. The state information (S) may include the voltage, current, and total harmonics (THD) of the system, and the current state of each control facility, for example, the OLTC tap position, the SOC of the ESS, etc.

[0027] The control unit (200) can determine an optimal action command (A) based on observed state information (S) and transmit it to each control facility component (310 to 340) for execution. The action command (A) may refer to an integrated set of control commands including an OLTC tap adjustment command, a VSC / ESS / STATCOM setting value change command, etc.

[0028] FIG. 2 is a detailed configuration diagram of the reinforcement learning-based control unit (200) of the present invention, which may include a safety reinforcement learning structure.

[0029] Since power system control requires a high level of safety and reliability, the search process of the reinforcement learning agent must not cause instability in the system. To ensure this, the present invention may include a Safe RL structure that includes a safety constraint (250) as a component.

[0030] The reinforcement learning-based control unit (200) may include at least one of a state observation unit (210), a reward calculation unit (220), a policy decision unit (230), a learning update unit (240), a safety constraint unit (250), and an action execution unit (260).

[0031] The state observation unit (210) can observe or process raw data collected from the power system (300) into state information (S) that can be recognized by a reinforcement learning model.

[0032] The policy decision unit (230) receives real-time state information (S) and can determine a proposed action (A') expected to maximize the reward using a built-in policy model (e.g., a deep neural network). The proposed action (A') is the optimal action at the current time, but whether it satisfies safety constraints has not yet been verified.

[0033] The safety constraint unit (250) can verify whether the proposed action (A') output from the policy decision unit (230) satisfies predefined safety constraints. The safety constraints may include (1) logical constraints (e.g., prevention of conflicting actions between devices), (2) operational constraints (e.g., limiting the number of OLTC operations), (3) physical constraints (e.g., voltage / current range), etc.

[0034] If the proposed action (A') violates the constraint, the safety constraint unit (250) can correct or prohibit (Action Masking), that is, mask the action, to determine the final (verified) action (A) within a safe range. If the constraint is satisfied, the proposed action (A') can be confirmed as the final action (A).

[0035] Here, action masking can serve the role of concealing actions that violate safety constraints from the options altogether. For example, if a constraint such as the daily OLTC action limit is encountered, actions like tap up and tap down could be overlaid with very large negative values ​​or negative infinity. Thus, this can be interpreted as making the actions that violate the constraint unselectable among the calculated probability values ​​or Q values ​​(expected reward values) for the actions possible in the current state.

[0036] Additionally, behavior correction or projection may mean moving the proposed behavior (A') to the boundary of a safe area or modifying it according to predefined rules if the behavior violates safety constraints. For example, if a continuous behavior value exceeds the facility capacity limit (physical constraint), the value may be limited to the maximum capacity value, and if a reactive power direction mismatch (logical constraint) occurs, the behavior direction of a specific facility may be corrected to align with other facilities.

[0037] The action execution unit (260) can convert the final action (A) determined from the safety constraint unit (250) into specific control commands that each control component (310 to 340, etc.) can execute and apply them to the power system (300).

[0038] The compensation calculation unit (220) can calculate a compensation signal (R) based on multiple goals, such as voltage stability and minimization of power loss, based on the system state changed after performing an action.

[0039] The learning update unit (240) can store experience data (S, A, R, etc.) in the experience memory (241) and use it to continuously learn and update the policy model of the policy decision unit (230). The learning process can be performed through a DRL algorithm (e.g., PPO, SAC, etc.) and can be performed asynchronously with the real-time control process.

[0040] In order for the reinforcement learning-based control unit (200) to learn the complex dynamic characteristics of the power system and derive an optimal voltage control policy, it is necessary for the state space (S), the action space (A), or the compensation function (R) to be precisely designed to reflect the power characteristics.

[0041] The state space (S) needs to contain all the information necessary for the agent to accurately perceive the current situation of the power system (300) and make optimal decisions.

[0042] The state observation unit (210) can define a state (S) by including information such as system physical state, control equipment state, external environment and context information.

[0043] The physical state of the system may include information such as (1) voltage magnitude or phase angle of major buses, (2) current magnitude or active power (P) and reactive power (Q) flow through major lines, (3) total system load level and distribution state, and (4) total harmonic distortion (THD) and system frequency, which are power quality indicators.

[0044] Information such as the status of the control equipment, (1) the current tap position of the OLTC (310), (2) the number of daily accumulated tap operations of the OLTC (310) for determining operational constraints, (3) the amount of reactive power currently being output or the set value of the VSC / SST (320) and STATCOM / SVC (340), (4) the current charge state (SOC), charge / discharge power amount (P, Q), etc. of the ESS (330) may be included.

[0045] External environment and context information may include (1) time information (season, day of the week, time of day) for reflecting load patterns or rate plans, (2) short-term load forecast or renewable energy generation forecast information, etc.

[0046] The state space (S) may be a high-dimensional space containing multiple continuous variables and some discrete variables. The present invention may utilize Deep Reinforcement Learning (DRL) to process such a complex state space (S).

[0047] In processing state information (S), state variables may have different ranges and units (e.g., pu, MVar, number of times, etc.), and continuous and discrete variables may be mixed. The state observation unit (210) may perform a preprocessing process on the collected state information to increase the learning stability and efficiency of the deep neural network embedded in the policy decision unit (230). This may include a normalization or standardization process to adjust the scale of the data.

[0048] The action space (A) may be a set of possible control actions that an agent can take on the controlled equipment (310 to 340). The present invention may introduce a hybrid action space by reflecting the characteristics of heterogeneous control equipment.

[0049] Discrete actions may include tap adjustment commands of the OLTC (310) (e.g., -1 (tap down), 0 (hold), +1 (tap up), etc.).

[0050] The continuous action may include (1) a control setting value of VSC / SST (320) (e.g., amount of reactive power injection / absorption, reference value of the connection point voltage, etc.), (2) a setting value of active power (P) and reactive power (Q) of ESS (330), (3) a setting value of reactive power (Q) of STATCOM / SVC (340), etc.

[0051] The policy decision unit (230) can search for an optimal combination within this hybrid behavior space and output a proposed behavior (A'). To this end, an actor-critic-based advanced DRL algorithm (e.g., SAC, PPO, etc.) may be utilized.

[0052] In the implementation of the hybrid action space of the present invention, the policy decision unit (230) may use a specialized neural network structure to effectively process the hybrid action space in which discrete actions and continuous actions are mixed. In one embodiment, the output layer of the policy model (neural network) may be configured by separating it into a part that determines discrete actions (e.g., OLTC tap command) and a part that determines continuous actions (e.g., VSC / ESS / STATCOM setting value).

[0053] The compensation function (R) can quantify various goals of power system operation to guide the learning direction of the agent. The compensation calculation unit (220) uses a multi-goal compensation function that considers the following elements.

[0054] The total reward (R_total) can be given as follows in one example.

[0055] R_total = (w1 * R_volt) + (w2 * R_loss) + (w3 * R_quality)

[0056] Here, R_volt is a voltage stability compensation, and a higher compensation can be applied as the voltage of all buses in the system can be maintained within a specified range (e.g., 0.95 to 1.05 pu) and the deviation from the reference voltage (1.0 pu) is smaller. For example, it can be defined using the sum of squared deviations between the voltages of N buses in the system and the reference voltage.

[0057] R_loss may mean efficiency compensation, and a higher compensation may be granted as the overall power loss of the system of the present invention decreases.

[0058] R_quality can refer to power quality compensation, and higher compensation is granted as power quality indicators such as total harmonic distortion (THD) improve.

[0059] Additionally, an element " + w4 * R_cost " can be added to R_total, where R_cost represents compensation for operating costs. The operating cost of the control equipment can be the cost of ESS degradation or maintenance costs due to frequent control, and higher compensation can be granted as the operating cost of the control equipment is lower or the economic efficiency is higher.

[0060] The weights (at least one of w1, w2, w3, and w4) are set according to the priority of the operational goal and can be dynamically adjusted according to system conditions such as normal / emergency mode. The present invention may be characterized by ensuring hard constraints through a safety constraint unit (250), and the compensation function may focus mainly on performance optimization goals.

[0061] Additionally, the present invention may include specific safety constraints implemented in the safety constraint section (250). The safety constraints may include operating rules based on power system knowledge, going beyond the physical limitations primarily addressed by existing technologies.

[0062] The safety constraints of the present invention may include at least one of logical / coordination constraints, operational / lifetime constraints, and physical / regulatory constraints.

[0063] Logical / cooperative constraints may be constraints designed to prevent inefficient operation in which multiple heterogeneous control devices conflict with each other or cancel out control effects.

[0064] For example, logical / coordination constraints may include directional alignment between reactive power compensation devices, or prevention of conflicting operations between devices with different response speeds.

[0065] In terms of directional alignment between reactive power compensation devices, components capable of controlling reactive power (Q), such as VSC (320), ESS (330), and STATCOM (340), may be required to operate in the same direction at the same time. That is, all can inject reactive power or all can absorb it. This can prevent the cancellation of control effects between components or system vibration.

[0066] In preventing conflicting operations between devices with different response speeds, it is possible to prohibit an OLTC (310) having a slow response characteristic and a power electronic-based device (320 to 340, etc.) having a fast response characteristic from controlling voltage in opposite directions.

[0067] For example, it may be prohibited for the OLTC (310) to operate in a direction that increases the voltage (tap increase) while the VSC (320) operates in a direction that decreases the voltage (reactive power absorption increase). This prevents the control system from failing to stably reach the target value and from becoming unstable, constantly fluctuating up and down around the target value.

[0068] Operational / lifetime constraints may be time constraints that take into account the mechanical lifespan or operational limitations of the equipment. Operational / lifetime constraints may include limiting the number of daily operations of the OLTC (310) or guaranteeing a minimum operation interval.

[0069] In the OLTC daily operation limit, the OLTC (310) may be limited to a maximum number of daily operations due to mechanical wear. The safety constraint unit (250) tracks the cumulative number of operations and may prohibit the OLTC operation command when the limit is exceeded.

[0070] In guaranteeing a minimum operating interval, a minimum time interval between consecutive operations can be guaranteed to prevent frequent operation of specific equipment such as OLTCs and circuit breakers.

[0071] Physical / regulatory constraints are physical limits of the power system and facilities or essential constraints under operating regulations, and may include at least one of system voltage range limits, line current capacity limits, and facility capacity limits.

[0072] In the system voltage range limit, the voltage of all busbars cannot exceed the specified range (e.g., ±10% of the rated voltage or 0.9 to 0.1 pu).

[0073] In line current capacity limits, the current flowing through any line cannot exceed the thermal allowable capacity of that line.

[0074] In terms of facility capacity limitations, each control facility component must operate only within its rated capacity range (e.g., minimum / maximum SOC range of the ESS, maximum output range of the VSC / STATCOM).

[0075] The safety constraint unit (250) can convert the proposed action (A') of the policy decision unit (230) into a final action (A) by utilizing behavior masking, behavior correction / projection, etc., to ensure various types of constraints such as logical, operational, and physical conditions.

[0076] Figure 3 is an operation flowchart of the reinforcement learning-based power system voltage control method of the present invention.

[0077] First, the state observation unit (210) can observe current system state information (S) from the power system system (300) (S201).

[0078] The policy decision unit (230) can determine a proposed action (A') corresponding to the current state (S) based on the learned policy model (S202).

[0079] The safety constraint unit (250) can verify whether the proposed action (A') satisfies predefined safety constraints (logical, operational, physical constraints, etc.) (S203).

[0080] If, as a result of verification, a constraint is violated (No), the safety constraint unit (250) can determine a safe alternative behavior (A) by correcting or masking the corresponding behavior (S204). For example, the direction of the behavior can be modified when a logical constraint is violated, or a specific behavior can be prohibited (masked) when an operational constraint is violated.

[0081] If the proposed action (A') satisfies all constraints (Yes), the proposed action (A') can be confirmed as the final action (A) (A=A').

[0082] The confirmed final action (A) can be transmitted to each control facility of the power system (300) through the action execution unit (260) and executed (S205).

[0083] After the action is executed, the reward calculation unit (220) can calculate a reward (R) according to multiple goals based on the changed system state and can store the corresponding experience data (S, A, R, S', etc.) in the experience memory (241) (S206). The subsequent process returns to step S201 and is repeated.

[0084] Meanwhile, in one embodiment, the learning update unit (240) can continuously learn and update the policy model of the policy decision unit (230) using data accumulated in the experience memory (241) asynchronously with the real-time control flow.

[0085] The asynchronous learning process can be performed in parallel independently of the real-time control method (e.g., S201 to S206). This ensures that computational delays caused by the learning of complex DRL models do not affect real-time control responsiveness and can satisfy the rapid response speed essential for power system control.

[0086] FIG. 4 shows a specific operational example of the safety constraint unit (250), which is a core component of the present invention. How the present invention handles logical constraints or operational constraints is explained through specific scenarios.

[0087] Case 1 of FIG. 4 describes the matching of the reactive power compensation direction as an example of logical constraint processing.

[0088] The policy decision unit (230) can determine a proposed action (A') that commands reactive power injection (+) to VSC (320) or ESS (330) and reactive power absorption (-) to STATCOM (340) based on the current grid state, thereby assuming a collision occurs.

[0089] The safety constraint unit (250) can detect that this combination of actions violates the reactive power compensation directionality alignment.

[0090] The safety constraint unit (250) can correct the corresponding action. For example, the direction can be aligned by changing the command of STATCOM (340) from absorption (-) to injection (+) to match the direction (injection) of multiple devices and determining the final action (A). This prevents the cancellation of control effects between control components or system instability.

[0091] Case 2 of FIG. 4 is an example of an operational constraint processing example in which the number of OLTC operations can be limited. Assuming that the daily operation limit of the OLTC (310) is 100 times and the current cumulative number of operations has reached 100 times, it can be assumed that the policy decision unit (230) has decided on a proposed action (A') including raising the tap (+1) of the OLTC (310) for voltage adjustment, which is a state of exceeding the limit.

[0092] The safety constraint unit (250) can detect that this action violates the constraint of the daily operation limit of the OLTC, and the safety constraint unit (250) can mask the action. That is, the operation command of the OLTC (310) can be changed to maintain (0) to confirm the final action (A). That is, the operation can be prohibited. Through this, the lifespan of mechanical equipment such as the OLTC (310) can be protected and long-term reliability can be ensured.

[0093] The dynamic switching of the operating mode of the present invention is described.

[0094] The reinforcement learning-based control unit (200) can be implemented to dynamically switch the operating mode according to the current system state (S). The reinforcement learning-based control unit (200) can be implemented to dynamically switch the operating mode according to the current system state (S) or an external command. This can be implemented by the compensation calculation unit (220) dynamically adjusting the weights (w1, w2, etc.) of the multiple target compensation function according to the operating mode.

[0095] The operating mode may include at least one of normal mode, calibration mode, and prediction mode.

[0096] Normal mode can assign high weight to efficiency (R_loss) and economics.

[0097] The calibration mode can assign a high weight to voltage stability (R_volt) to rapidly restore stability in the event of external disturbances.

[0098] The prediction mode can perform preemptive control to prevent potential future problems in advance by utilizing short-term prediction information.

[0099] In addition, the present invention may introduce additional logical / operational constraints. Various constraints may be added to the safety constraint section (250) according to a specific grid environment or operation policy. For example, it may include a limit on the SOC operating interval considering the lifespan of the ESS (330), a constraint to prioritize the voltage stability of critical loads, etc.

[0100] In addition to a centralized structure, the present invention can be extended to a distributed or hierarchical structure based on multi-agent reinforcement learning. In this case as well, a configuration performing the same function as the safety constraint (250) is included within each agent or parent agent to ensure the safety of the entire system.

[0101] The reinforcement learning-based control unit (200) may be implemented as a computing device comprising at least one of one or more processors (e.g., CPU, GPU, NPU), memory, storage device, and communication interface.

[0102] To effectively handle the hybrid behavioral space and constraints of the present invention, an Actor-Critic structure-based DRL algorithm (e.g., PPO, SAC, or an extended version of the constraint optimization thereof) may be utilized.

[0103] Before actual application to the grid, initial training can be performed in an offline environment linked with power system simulators such as OpenDSS and PowerFactory.

[0104] The control unit (200) exchanges data in real time with a measuring instrument (350) and control components (310 to 340, etc.) using SCADA or EMS infrastructure, etc., and may include an interface that supports related communication protocols such as DNP3, IEC 61850, etc. Explanation of the symbols

[0105] 100... System voltage control unit 200... Reinforcement learning-based control unit 210... Status Observation Unit 220... Compensation Calculation Unit 230... Policy Decision Department 240... Learning Update Department 241... Experience Memory 250... Safety Constraint Department 260.. Action Execution Unit 300... Power System (Environment) 310... OLTC 320... VSC / SST 330... ESS 340... STATCOM / SVC 350... Instrument S... Status Information A'... Proposed action A... Final (verified) action R... reward signal S201... System status information observation stage S202... Proposed Action Decision Stage S203... Safety Constraint Verification Step S204... Behavior correction or masking step S205... Final action execution phase S206... Compensation calculation and experience data storage step

Claims

Claim 1 A voltage control device for a power system including heterogeneous control equipment, such as an OLTC and a power electronics-based control equipment, having different response speeds, comprising: a state observation unit that observes state information including voltage, current, and the state of each control equipment from the power system; a policy decision unit that determines a proposed action for the heterogeneous control equipment based on the state information using a policy model learned based on reinforcement learning; a safety constraint unit that verifies whether the proposed action determined by the policy decision unit satisfies a predefined safety constraint, and if the proposed action violates the safety constraint, performs action masking to make the action unselectable or corrects the proposed action to within a safe range to determine a final action; and an action execution unit that outputs the final action determined by the safety constraint unit to the power system; wherein the safety constraint includes a logical constraint that prohibits conflicting operations in which the OLTC with a slow response speed and the power electronics-based control equipment with a fast response speed control the voltage in opposite directions. Claim 2 In claim 1, the policy model of the policy decision unit is a system voltage control device that learns in a direction to maximize a compensation function calculated based on at least one of the voltage stability, power loss, and power quality including THD of the power system. Claim 3 A system voltage control device according to claim 1, comprising: a compensation calculation unit that calculates a compensation signal based on the state information and the action determined by the policy decision unit; and a learning update unit that stores experience data and learns and updates the policy model using the experience data. Claim 4 A system voltage control device according to claim 1, comprising: a learning update unit that stores experience data and learns and updates the policy model using the experience data; wherein the learning process of the policy model by the learning update unit is performed asynchronously with the real-time control operation by the policy decision unit and the action execution unit. Claim 5 A system voltage control device according to claim 2, comprising: a compensation calculation unit that calculates a compensation signal based on the state information and the action determined by the policy decision unit; wherein the compensation calculation unit calculates the compensation signal by adjusting the weights of the compensation function according to a predefined operating mode. Claim 6 A system voltage control device according to claim 1, wherein a plurality of control facilities include an OLTC, and the status information includes the current tap position of the OLTC and the cumulative number of tap operations per unit time. Claim 7 delete Claim 8 delete Claim 9 In claim 1, the logical constraint comprises a system voltage control device including a condition that stipulates that a control facility capable of reactive power control performs injection or absorption operations in the same direction at the same time. Claim 10 A grid voltage control device according to claim 1, wherein the safety constraint includes an operational constraint, and the operational constraint includes a condition that limits the maximum number of operations per unit time of the OLTC. Claim 11 A system voltage control device according to claim 1, wherein the safety constraint includes an operational constraint, and the operational constraint includes a condition that ensures a minimum time interval between consecutive operations to prevent frequent operation of a plurality of control facilities. Claim 12 delete Claim 13 delete Claim 14 In claim 1, the action determined by the policy decision unit is a grid voltage control device determined within a hybrid action space comprising a discrete control command including OLTC tap adjustment and a continuous control command including at least one set value among VSC, ESS, and STATCOM.